Background
During my formal Sanskrit training — from 2008 to 2022 at Cornell, Harvard, Vienna, and Leipzig — my experience with searchable Sanskrit e-texts centered on open collections created and guided by western philologists: mostly GRETIL, Muktabodha, and DCS, supplemented by additional, smaller open collections like SARIT. You can see that same body of material today still comprising the bulk of my own SETI inventory, or Sebastian's Dharmamitra database, or Harry Spier's SALSa aggregation platform.
Second, of course, everyone has always used whatever private Dropbox, Google Drive, and USB thumb-drive collections they happen to have been invited to share. And good ol' Google (feeling lucky?)
And third, there were other collections, too, but of decidedly lower priority for my own academic community of origin: things like Jain eLibrary and Sanskrit Documents. I would check these as well, but only after exhausting other possibilities, for a number of reasons: either their focus did not quite match my own, or texts needed to be accessed one at a time on a difficult-to-use website, or maybe I had doubts about curation or digitization quality — things like metadata (e.g., which edition used), formatting choices, proofreading status, etc.
Today's post argues that those "other collections" have now become the main game in town, unbeknownst to the scholars I talk with the most. In my previous post about Sanskrit Wikisource, I called that collection one of several "icebergs," whose visible extent belies a far greater latent magnitude, and which potential users are consequently underutilizing or disregarding altogether. In the weeks following, I worked hard to get my head around two more such collections, which now have their own Atlases:
I'll discuss these in a second. But the more exciting news is that I found a way to tie it all together in a mutually illuminating way.
Sāgarasaṅgama
A few years ago with a colleague, I was reading Samudrasaṅgama, the Sanskrit version of Dārā Shukoh's Majma-ul-Bahrain, which attempts a syncretism of Sufi and Vedāntic thought. I really liked the metaphor of intellectual oceans coming together. Similarly, Sanskrit grammar is sometimes described as a vast ocean —
अनन्तपारं किल शब्दशास्त्रं
स्वल्पं तथायुर्बहवश्च विघ्नाः ।
सारं ततो ग्राह्यमपास्य फल्गु
हंसैर्यथा क्षीरमिवाम्बुमध्यात् ॥
— so how much more the entirety of Sanskrit literature, which these e-text mega-collections attempt to represent?
If each mega-collection constitutes its own imposing "ocean" (samudra, sāgara) that requires an Atlas to navigate, then something that represents them together in one place would be a virtual "confluence" or "meeting" (saṅgama) of those oceans. And now that exists, too.
sagarasangama.info
On this new project's About page, I give a top-level comparison of the component collections, currently three in number:
As these figures show, here's the story no one has been telling, but which Sāgarasaṅgama is finally able to tell: In 2016, Sanskrit Wikisource outstripped GRETIL in size (see previous post for numbers). Then, in 2021, E-bharatisampat rocketed past both and has continued climbing at an astonishing rate, apparently with much greater use of OCR than any other related project. These materials often appear to be of good, usable quality, even if metadata is lacking in some cases. We all need to be checking these collections from now on.
Fortunately, doing so is now a lot easier. Using the Sāgarasaṅgama main page, you can search titles and authors across all available collections, nearly 19,000 items. Like in the Atlases, search results display relevant metadata like category, date, and size.
Wait, what?
Yeah! Let me back up a bit. In case you missed it — many have! — E-bharatisampat is a project by Samskrita Bharati with significant funding by the Indian government. It officially started in 2017 (for late-2019 About page vibes, see the Wayback Machine) and seems to have taken on its mature form by mid-2021. For its part, as I discussed earlier this summer, Sanskrit Wikisource is a crowdsourced library that has been growing steadily for over 20 years. (Sanskrit Documents, which needs no introduction here, is a smaller story. Suffice it to say that it, too, is larger than most realize, but its narrow genre coverage makes it less generally useful to scholars.) Because these websites are hard to use and don't provide for easy mass-download of their contents, they haven't penetrated into systematic scholarly use. But the more I use them — specifically with the help of the tools I've built to do so — the more I believe these collections need to be appreciated and celebrated. So, in my capacity as a librarian, I've essentially made a "finding aid" for making better use of this important body of material by quickly identifying where to look closer, sort of like WorldCat.
Instead of offering some sort of quantitative argument to support my claim that these materials are of sufficiently high quality — and let's remember that GRETIL, in its own right, was never perfect — I encourage students and scholars to just start engaging with them, using the tools I've built, which I hope make for a much smoother and less confusing navigation experience.
Atlas specifics
Having done this three times now, I now see the "Atlas" concept coming into sharper focus. It's still the case that an Atlas is not a mirror, in that it does not aim to re-host material in its entirety as a sort of backup. Instead, the Atlas focuses on reorganizing the material's metadata and representing it more legibly, through an alternative front-end overlay, so that users get a clearer sense of scope and can better navigate the relevant structure.
Some interesting patterns I've noticed:
- No mega-collection explicitly clarifies its extent in bytes. (Recall the traditional practice of measuring manuscripts in granthas or groups of 32 akṣaras.)
- No collection explicitly charts its own history.
- No collection explicitly advertises recent changes. (Sanskrit Wikisource's Recent Changes page is hardly an exception.)
- Each collection has its own unique approach to categorization.
- Each collection has unique structural faults (e.g., duplication, broken links, miscategorized material).
My goal is to achieve a practical, synoptic understanding of these collections that is not only synchronic (capturing internal complexity at this particular moment in time) but also diachronic (historical, over time). The method: zealous dissection with the help of AI and scraping software. Far from being flippant about the latter, I am very mindful of my activities' effect on the respective servers, and so I scrape carefully and at a polite pace, and I make sure that my scraping code does not get shared publicly. I also recognize that maintainers of these sites may not approve of my snooping. Ultimately, though, I believe that my alternative interfaces will help more people come and use these sites, which is a good thing for them. So after wrestling with some doubt, I went ahead.
My research findings into each collection are published in three main forms:
- The Atlas main page and its particular features (e.g., browsing options).
- The Atlas About page sections "Practical Intro to XYZ Structure" and "Obtaining and Representing XYZ Data".
- The Atlas About page sections "Data Quality" (synchronic audit of structural faults) and "Data Quantity" (diachronic insights).
Let's get into some specifics:
Sanskrit Wikisource Atlas
- Curation is decentralized, leading to a wide range of structural inconsistencies; the Atlas does its own abstraction to "text" (e.g., by mining breadcrumb links) and surfaces category orphans with a synthetic asaṃbaddhavargīkṛta ("miscategorized") bucket.
- Many but not all items were created using Wikisource's inbuilt OCR pipeline, which also retains source images; the Atlas distinguishes provenance type and provides links as possible.
- Inherent Wikimedia versioning structure makes downloading very easy and facilitates exceptionally fine-grained diachronic exploration.
E-bharatisampat Atlas
- Most items are public domain editions, with images uploaded in full; the e-texts thereof, whether by OCR, manual typing, or both, often appear to have been proofread fairly well.
- Categorization is two-level and consistent, and each item has a unique serial ID.
- Browsing on the source website is distinct for e-texts ("Unicode Book") vs. PDFs ("E-Book") and inadvertently conceals a full third of e-texts.
- A few dozen more items are either empty or improperly encoded, leaving them unusable.
- A minority of e-texts are structured and presented in a distinct "Read Chapters" interface.
- Date information must be inferred from a stray timestamp on thumbnail image files.
Sanskrit Documents Atlas
- The source site has three overlapping categorization layers (folders, topics, and category tags), plus a nav menu that is internally redundant, presenting the same topic in more than one place. The Atlas clearly exposes all three layers: folders and topics as alternative trees, category tags as an optional additional filter.
- Some PDFs are hosted by the source site, but most are provided in the form of external links, e.g., to Archive.org.
- Only a single last-updated date is provided for each item, which conceals the true age of many if not most.
More Atlases to come
An Atlas for GRETIL? Definitely. For Jain eLibrary? I'd love to, but we'll see. After that, I'm not currently aware of other mega-collections that would warrant the Atlas treatment — with one interesting exception. SETI inventorying is sufficient for all the small collections whose structure just isn't complicated enough to need illuminating. In order to integrate SETI into Sāgarasaṅgama, I have in mind to give SETI a pseudo-Atlas that simply aggregates the overall stats. (Any collection represented both by SETI and by its own Atlas would be deduplicated first.) In that way, small collections contribute metadata via SETI, and large collections contribute metadata via their Atlases, and Saṅgama can sum over them all.
Finally, in the long-term (3–5 years), I'll expand SETI to represent all of the Atlas content, update Pandit Project accordingly, and enable comprehensive visualization in Pāṇḍitya — my one true love.
❤️