A set of questions came up recently on the Indology mailing list about Sanskrit Wikisource — what is it, who made it, and how useful is it? As it happens, I spent time a few months ago investigating this very thing, and I came up with a prototype solution that addresses some of the problems I found. As a result of the Indology exchange, I dug in deeper and developed the prototype further, which I'm now ready to share more widely. First things first, though: let's try to answer those basic questions.

What is it?

Wikisource is part of the Wikimedia family — same foundation as Wikipedia — and like its siblings, it is a radically open, crowdsourced project with public revision history. Whereas Wikipedia aims to be "The Free Encyclopedia," Wikisource aspires to be "the free library that anyone can improve," and it has a few hundred subdomains for different languages. The Sanskrit subdomain, started in 2004, is a Sanskrit-only space containing a wide range of Sanskrit literature.

Who made it?

Lacking any single creator or curator, Sanskrit Wikisource has been used for over two decades by motivated individuals and groups to work on and share Sanskrit e-text material that might otherwise never have circulated without such an easy-to-edit platform. These contributors are noted in the item revision histories, but you have to go looking for them. To take one random example: on the page याज्ञवल्क्यस्मृतिः/आचाराध्यायः/उपोद्घातप्रकरणम्, clicking "इतिहासः दृश्यताम्" shows you the revision history, and then you can click on the specific username Sbblr0803 to learn that this is a person named Abhiram, see what else he has done, etc.

How useful is it?

On multiple occasions, I've found very interesting things there, like first-online full-text postings of Gaṅgeśa's Tattvacintāmaṇi and Śrīharṣa's Naiṣadhīyacarita. As far as I know, no one has undertaken a rigorous evaluation of its data quality relative to other well-known collections like GRETIL, but I think it is much more potentially useful than many people realize.

However.

The Wikisource platform has inherent pros and cons for hosting an evolving e-text collection like this, and on balance, I think the cons win out, with the result that the collection has not been living up to its potential. It's hard enough to keep a sprawling, decentralized project on track, and from what I've seen, the Sanskrit domain shows a worrying lack of curatorial vision — frankly just a lot of mess, which I'll discuss more below. As a result, it can be very hard to plumb its depths — a fact ironically reinforced by the Wikisource logo itself.

Wikisource's iceberg logo
Wikisource's iceberg logo

Did you know that Sanskrit Wikisource has hundreds of megabytes of text data? If you've ever tried to browse it, did you feel inspired to come back again the next time you need something? No? You're not alone.

Metadata Problems

The discerning user will not take long to notice that most Sanskrit Wikisource content lacks any usable metadata at all beyond revision history. Specifically, not knowing what edition or editions a given text is based on can make it extremely difficult to use responsibly. Users are left with the implicit claim that "this just is the Pañcadaśī by Vidyāraṇya." An enterprising researcher may well figure it out by comparing against available printed editions. But why the guessing game?

The exception is items produced through Wikisource's in-built Proofread OCR pipeline, whose Index pages hold the source PDF. On the Sanskrit domain, these Index pages are called "अनुक्रमणिका;" see for example this Nyāyamañjarī from the Kashi Sanskrit Series.

This widespread lack of metadata is a real problem. If it's going to be fixed anywhere, I think it should be done on-platform, in collaboration with the contributors who continue to add content.

Another set of problems, however, can be addressed relatively quickly with a light-touch technological solution.

Navigation Problems

The best way to navigate Sanskrit Wikisource isn't clear from its home page.

Wikisource's home page
Bird's-eye view of the Sanskrit Wikisource home page

First, not everyone is comfortable with an all-Sanskrit, all-Devanāgarī interface. Second, this page is a cluttered mess. If you look past the less useful stuff, the वर्गसर्वस्वम् section offers some links to category pages, but clicking them gives you only a narrow view of one category at a time.

Pro tip: There's a separate वर्गः:वर्गसर्वस्वम् page a few wiki-clicks away. There, you'll find a dynamic dropdown tree structure that is already much better for browsing.

Sanskrit Wikisource's hidden Vargasarvasvam page
The somewhat hidden Vargasarvasvam page, Sanskrit Wikisource's own best way to browse

As for the categorical breakdown itself, it's mostly unobjectionable, but as is always the case, there can be a learning curve if you come from a different taxonomic perspective. And there is undoubtedly mess here, too. By convention on Wikisource, categories are supposed to form a single, cohesive tree covering everything, but some categories are misplaced (e.g., धर्मशास्त्रम् as a sibling of ग्रन्थाः rather than a child); some are unhelpfully separated due to naming differences (ज्योतिष्यम् 9.9 MB and ज्यौतिषम् 1.6 MB); some sit outside of the central tree (ज्यौतिषम् is within वेदाङ्गानि, but ज्योतिष्यम् is free-floating); and many items lack categories altogether.

And what is an "item" deserving of categorization anyway? It's quite subtle, actually. Abstraction to a "text" level is purely manual on Wikisource and can be done in various ways. One Wikisource editor may choose to put an entire Sanskrit text on a single page. Another may choose to spread it across dozens or hundreds of pages connected by nested tables of contents. Sometimes, the top-level table-of-contents page is not labeled with a Category, but second-tier (e.g., chapter) table-of-contents pages are. Furthermore, an editor may sometimes wish to rename one or more pages, but because it's hard to chase down and fix all inbound connections, they can just set up redirect pages, adding further layers of navigational complexity. Users are at the mercy of all this curatorial diversity, which may or may not be implemented correctly.

Granted, there are other, non-categorical access options, as well: a search bar, and an alphabetical page index (e.g., ). But I find both of these to be myopically restrictive and disorienting in their own ways, too. My bet is that, even after trying all of the above, a typical user may still be shrugging their shoulders about whether this collection has material that might be useful to them — and might I add, probably feeling exhausted by the austere wiki aesthetics of it all. After all the work that went in, by so many people, that's a shame.

A Better Interface

Fortunately, it's possible to create another way in, even without modifying Wikisource. Based on my understanding of the collection's structure, I made an alternate interface that I personally find much more user-friendly and informative.

https://tylergneill.github.io/sanskrit-wikisource-atlas

My Sanskrit Wikisource Atlas

I call this new interface an "atlas" because it not only charts the navigable waters but also offers analytic insights not directly readable from the surface. It is purely a navigational layer, hosting no text of its own, so all of its links bring you to sa.wikisource.org. It understands and compensates for many of the issues listed above, giving you: 1) cleaner browsing and search in a single-page display with a sidebar; 2) no distracting Wikimedia artifacts (e.g., "निष्कासनाय", "to be deleted"); 3) transliteration options; 4) more intuitive abstraction to the "text" level; and 5) clearer stats, specifically page and text counts, plus size and last-updated info throughout.

The way I did it: 1) download the full data; 2) process it to calculate relationships and counts; 3) boil that info down to a single main data file; and then 4) build a simple website on top of that file, which GitHub Pages easily hosts for free.

The main reason I do not re-host the full text data myself, as a true "mirror" would do, is that the collection is demonstrably still living and growing. The Atlas's last-updated info for each item clearly shows this, as does Wikisource's own "सद्यपरिवर्तनानि," i.e., Recent Changes page. Like Pāṇḍitya's use of Pandit Project, this Atlas starts from a data snapshot, builds functionality on top, and points back to the source for further reference. That does mean that it needs to be periodically refreshed to stay current. Fortunately, an automated pipeline  makes that easy.

What the snapshot numbers from Summer 2026 reveal: Sanskrit Wikisource holds over 700 MB of plain-text content (measured in IAST) in something like 2,000–3,000 texts across over 26,000 individual web pages. For comparison, GRETIL is around 350 MB before controlling for duplication, roughly 175 MB and 700 unique works after (see my previous post). In short, Sanskrit Wikisource is definitely huge — roughly 2–4x the size of GRETIL — and it's very interesting.

It's Actually More Complicated...

This is the part where I could digress badly trying to communicate all the rough edges that the Atlas is smoothing out — it's almost too user-friendly in some respects. But I'm going to resist that urge. If you really want, you can go read the Atlas's comprehensive About page. Here's what I give there:

Here's a screenshot of those last two, which I honestly think are pretty cool:

Sanskrit Wikisource Atlas About page data reports
Atlas About page data reports, for understanding and improving the source collection

Charting a Course

Apart from creating an additional navigation layer like I've done here, I can think of two other ways excited newcomers to Sanskrit Wikisource may want to improve upon it.

Option 1: Duplicate the Data and Don't Look Back

Some people will see the value of Sanskrit Wikisource as content they can take and add to their own aggregation effort. To be sure, doing so opens up interesting possibilities for federated processing, like search. Examples are Vishvas Vasuki's raw_etexts GitHub repo and Harry Spier's Searchable Aggregate Library of Sanskrit. A separate question is whether one will also feed eventual improvements back upstream, to help the source project itself. Wikisource's default license is CC BY-SA 4.0, which means that you can do whatever you want with its material as long as you 1) make a properly detailed acknowledgment — the BY part — and 2) release your adaptation under either CC BY-SA 4.0 exactly or another license that Creative Commons has officially designated as one-way or two-way compatible with CC BY-SA 4.0 — the SA part.

That all said, I hope to have shown in this post why this kind of copy-and-go strategy is not a properly scholarly approach in this particular case, for three reasons. First, as a whole, Sanskrit Wikisource is riddled with organizational issues, with the result that an extraction effort will very likely make interpretive mistakes. Second, the individual items badly need more metadata for them to be of good scholarly use, and if such philological attention is going to be paid, it should be done at the source, rather than re-derived privately by each person who makes a copy. And third, the collection is still being edited all the time, so any wholesale snapshot will become misleadingly stale before long.

As an interesting aside, the Atlas pipeline system also surfaces provenance insights about borrowing in the other direction. For at least 244 pages (about 5% of the total collection by content size), Sanskrit Wikisource has itself directly borrowed from other collections, including Sanskrit Documents, Muktabodha, GRETIL, and others.

Option 2: Adopt the Best, Clean Up the Rest

Instead, my opinion is that we as a scholarly community should approach items one by one, study them carefully, and where appropriate, transplant them onto other platforms where they may receive more care and attention. That is what I intend to do in select cases with HANSEL.

AND I believe we should invest effort improving Sanskrit Wikisource in-place, as well, whether at the top, navigational level, or at lower levels like categories and metadata. For example, I voiced a lot of dissatisfaction with the home page, but I could also try to change it. Similarly, if I see something miscategorized, why not just submit an edit and see what happens? My honest answer: Before having this Atlas as a tool, I did not feel very motivated to do this on any sort of scale. Not yet being a part of the editing community myself, but having had certain experiences on Wikipedia, I know that it's possible that my changes may (and should!) be reverted if I don't approach things properly. More crucially, though, I didn't have enough perspective to know whether my effort would make much of a difference in the grand scheme of things.

Now though, with the Atlas as a tool to illuminate these otherwise forbidding depths, I have much more confidence that I understand the finite number of critical edits needed to fix the site's major problems. A concerted improvement campaign is now possible because I know the extent of the job: what each step consists in, and when I'll be able to say that I'm done. Also, since the Atlas is public and didactic in nature, anyone else who is curious can immediately get the full perspective and be motivated to pitch in, too.

Coming Back to the Surface

As intriguing as I find this Sanskrit Wikisource case, it's actually just the beginning. There are several other similarly underestimated Sanskrit icebergs out there. My plan is to extend the alternate-interface treatment to them, and to consolidate these interfaces into a single portal, which I will maintain as an additional Kalpataru Grove project named Sāgarasaṅgama, "Meeting of the Oceans." As these collections become better understood, their metadata can then be ingested into SETI, and their prosopographical aspects represented in Pandit Project. Finally, the items will become available for visualization in Pāṇḍitya. That's the grand vision, which I'll aim to realize in the next three to five years.

← It's Raining Links AgainJuly 24, 2026