A multilingual alignment index for preserving language and culture, and for connecting people through them.
LOL gathers translations of the same text across more than 1,500 languages and aligns them, so that a sentence in one language can be found beside its counterparts in all the others. It records where every piece of text came from and on what terms. Today it builds that skeleton from Bible translations; it is designed to grow into a shared index of meaning that serves research, localisation, accessibility and the preservation of languages and cultures.
Languages carry cultures. When a language loses speakers, what goes with it is not only vocabulary but a way of dividing up the world. Parallel text, the same content in many languages, is one of the few resources that lets people, tools and researchers move between languages at scale, and it is often the only substantial written material a small language has.
LOL starts from the Bible for a practical reason, not a religious one. It is the most widely translated text in existence, and its chapter-and-verse numbering is a ready-made alignment key: verse 3 of chapter 16 means the same place in every language. That makes it the fastest way to build a skeleton covering as many languages as possible. The skeleton is the starting point, not the destination.
The project began as a rewrite of Ehsaneddin Asgari’s 1000Langs super-parallel corpus crawler (LMU Munich), whose work showed how much can be built this way.
| Use | What LOL provides |
|---|---|
Language and culture research |
Aligned text across 1,500+ languages for comparative, typological, historical and anthropological work. |
Localisation and internationalisation |
A back end that tools such as pandoc and i18n pipelines can ask: here is a phrase, what are its aligned counterparts in language X, and on what terms may they be used? |
Accessibility and inclusion |
Reach into languages that mainstream tools ignore, so documents and interfaces can be offered in a reader’s own language. |
Archives and document processing |
Language and script identification keyed to stable language identities, for pipelines such as docudactyl; in return, digitised archives become a source of material for low-resource languages. |
-
An index, not a warehouse. LOL holds alignments, provenance and terms of use, and fetches or links to source text when it is needed. It does not try to copy every corpus into itself.
-
Provenance and consent travel with the data. Every item records its source, its licence or terms, and any community conditions on its use. LOL will not serve material for a use its terms exclude.
-
Disagreement is part of the record. When sources differ, as dialects, variant readings and rival translations do, LOL keeps every version with its source. It does not vote one winner and discard the rest.
-
Equivalence is stated, not assumed. "These two texts correspond" always says in what sense: same verse, same concept, same gloss. Correspondence is a claim with a basis, not a fact about strings.
-
Sources earn their place. A new source is added only if it brings a new alignment key, new languages or communities, or a new medium (such as speech), together with clear provenance and terms. More of the same text under the same key adds bulk, not reach.
LOL is a crawler and alignment toolkit. It currently draws on these sources (language counts as stated by each source):
| Source | Languages (approx.) | Access |
|---|---|---|
1,500+ |
API |
|
2,000+ |
Scraper |
|
800+ |
Download |
|
1,000+ |
Download |
|
1,200+ |
API |
What is built, tested and proved, and the evidence for each claim, is recorded in EXPLAINME.
LOL crawls translations; it does not redistribute them. Many translations, particularly on commercial Bible sites, are under copyright and covered by those sites' terms of service. Before publishing any corpus built with LOL, check the licence of each translation in it.
eBible.org hosts many public-domain and openly licensed translations, and the eBible corpus records the licence of each one. It is the safest default source.
Material from Indigenous and other communities is subject to their own authority over it. LOL is adopting the Local Contexts Traditional Knowledge and Biocultural labels and the CARE principles for Indigenous data governance, so that community terms are machine-readable and respected (see Roadmap).
Prerequisites: Node.js 20+, just, Podman; Nix optional.
git clone https://github.com/hyperpolymath/LOL.git
cd LOL
just build # compile
just test # run the test suites
just prove # run the proof checks (see EXPLAINME for what they cover)
just crawl-all # run every crawler| Path | Contents |
|---|---|
|
Crawlers (Bible Cloud, Bible.com, PNG Scriptures and others), alignment, language-code utilities |
|
Test suites |
|
Proof checks for alignment properties |
|
Source and crawl configuration |
|
Idris2 ABI package |
Items are marked with the estate’s statuses: DESIGNED means specified but not built; OPEN means not yet designed.
| Step | Status |
|---|---|
Key every item to a Glottolog language identity (and ISO 639-3 where one exists) |
DESIGNED |
Adopt CLDF (Cross-Linguistic Data Formats) as the exchange format, so LOL links to Lexibank, Grambank, WALS and Concepticon instead of copying them |
DESIGNED |
Add provenance, licence and Local Contexts label fields to the data model, before any new source is added |
DESIGNED |
First non-religious alignment key: the Universal Declaration of Human Rights in its many translations |
DESIGNED |
Concept-level alignment via Concepticon concept sets, alongside verse-level alignment |
OPEN |
Cultural and environmental context via D-PLACE, keyed to the same language identities |
OPEN |
Locale data from Unicode CLDR, so LOL can serve as a localisation back end |
OPEN |
Speech: links to Mozilla Common Voice and to endangered-language archives, respecting their access terms |
OPEN |
Align with Palimpsest and the Content Provenance Protocol, so provenance and consent are carried in a standard form |
OPEN |
-
Ehsaneddin Asgari, 1000Langs (Apache-2.0): the crawler LOL was rewritten from.
-
Mayer and Cysouw, Creating a massively parallel Bible corpus (LREC 2014).
-
Christodoulopoulos and Steedman, A massively parallel corpus: the Bible in 100 languages (Edinburgh).
-
McCarthy et al., The Johns Hopkins University Bible Corpus: 1600+ tongues for typological exploration (LREC 2020).
-
Akerman et al., The eBible Corpus (2023).
-
D-PLACE: cultural, linguistic and environmental diversity, keyed to Glottolog.
LOL is by Jonathan D.A. Jewell. It began as a rewrite of 1000Langs by Ehsaneddin Asgari (LMU Munich), licensed under Apache-2.0; how much of LOL derives from that code is being documented in EXPLAINME, and any derived parts keep their Apache-2.0 terms and notices.
LOL’s own work is dual-licensed under the Palimpsest-MPL-1.0 License and the Palimpsest License v0.8. Commercial use with attribution is permitted; proprietary AI training without attribution is prohibited. See LICENSE.
Source texts remain under their own licences and terms; see Data and rights.