Skip to content
hyperpolymathPublic

About

Parallel corpus crawler and alignment index across 1,500+ languages, for preserving languages and cultures and connecting people through them. Starts from Bible translations for their unmatched coverage; designed to carry provenance and terms of use with every text. For research, localisation and accessibility.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

90 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LOL (Loads of Languages)

A multilingual alignment index for preserving language and culture, and for connecting people through them.

LOL gathers translations of the same text across more than 1,500 languages and aligns them, so that a sentence in one language can be found beside its counterparts in all the others. It records where every piece of text came from and on what terms. Today it builds that skeleton from Bible translations; it is designed to grow into a shared index of meaning that serves research, localisation, accessibility and the preservation of languages and cultures.

Why LOL exists

Languages carry cultures. When a language loses speakers, what goes with it is not only vocabulary but a way of dividing up the world. Parallel text, the same content in many languages, is one of the few resources that lets people, tools and researchers move between languages at scale, and it is often the only substantial written material a small language has.

LOL starts from the Bible for a practical reason, not a religious one. It is the most widely translated text in existence, and its chapter-and-verse numbering is a ready-made alignment key: verse 3 of chapter 16 means the same place in every language. That makes it the fastest way to build a skeleton covering as many languages as possible. The skeleton is the starting point, not the destination.

The project began as a rewrite of Ehsaneddin Asgari’s 1000Langs super-parallel corpus crawler (LMU Munich), whose work showed how much can be built this way.

What LOL is for

Use What LOL provides

Language and culture research

Aligned text across 1,500+ languages for comparative, typological, historical and anthropological work.

Localisation and internationalisation

A back end that tools such as pandoc and i18n pipelines can ask: here is a phrase, what are its aligned counterparts in language X, and on what terms may they be used?

Accessibility and inclusion

Reach into languages that mainstream tools ignore, so documents and interfaces can be offered in a reader’s own language.

Archives and document processing

Language and script identification keyed to stable language identities, for pipelines such as docudactyl; in return, digitised archives become a source of material for low-resource languages.

Principles

  1. An index, not a warehouse. LOL holds alignments, provenance and terms of use, and fetches or links to source text when it is needed. It does not try to copy every corpus into itself.

  2. Provenance and consent travel with the data. Every item records its source, its licence or terms, and any community conditions on its use. LOL will not serve material for a use its terms exclude.

  3. Disagreement is part of the record. When sources differ, as dialects, variant readings and rival translations do, LOL keeps every version with its source. It does not vote one winner and discard the rest.

  4. Equivalence is stated, not assumed. "These two texts correspond" always says in what sense: same verse, same concept, same gloss. Correspondence is a claim with a basis, not a fact about strings.

  5. Sources earn their place. A new source is added only if it brings a new alignment key, new languages or communities, or a new medium (such as speech), together with clear provenance and terms. More of the same text under the same key adds bulk, not reach.

Current state

LOL is a crawler and alignment toolkit. It currently draws on these sources (language counts as stated by each source):

Source Languages (approx.) Access

Bible Cloud

1,500+

API

Bible.com

2,000+

Scraper

PNG Scriptures

800+

Download

eBible

1,000+

Download

Find.Bible

1,200+

API

What is built, tested and proved, and the evidence for each claim, is recorded in EXPLAINME.

Data and rights

LOL crawls translations; it does not redistribute them. Many translations, particularly on commercial Bible sites, are under copyright and covered by those sites' terms of service. Before publishing any corpus built with LOL, check the licence of each translation in it.

eBible.org hosts many public-domain and openly licensed translations, and the eBible corpus records the licence of each one. It is the safest default source.

Material from Indigenous and other communities is subject to their own authority over it. LOL is adopting the Local Contexts Traditional Knowledge and Biocultural labels and the CARE principles for Indigenous data governance, so that community terms are machine-readable and respected (see Roadmap).

Quick start

Prerequisites: Node.js 20+, just, Podman; Nix optional.

git clone https://github.com/hyperpolymath/LOL.git
cd LOL

just build        # compile
just test         # run the test suites
just prove        # run the proof checks (see EXPLAINME for what they cover)
just crawl-all    # run every crawler

Layout

Path Contents

src/

Crawlers (Bible Cloud, Bible.com, PNG Scriptures and others), alignment, language-code utilities

test/

Test suites

proofs/

Proof checks for alignment properties

config/

Source and crawl configuration

lol-abi.ipkg

Idris2 ABI package

Roadmap

Items are marked with the estate’s statuses: DESIGNED means specified but not built; OPEN means not yet designed.

Step Status

Key every item to a Glottolog language identity (and ISO 639-3 where one exists)

DESIGNED

Adopt CLDF (Cross-Linguistic Data Formats) as the exchange format, so LOL links to Lexibank, Grambank, WALS and Concepticon instead of copying them

DESIGNED

Add provenance, licence and Local Contexts label fields to the data model, before any new source is added

DESIGNED

First non-religious alignment key: the Universal Declaration of Human Rights in its many translations

DESIGNED

Concept-level alignment via Concepticon concept sets, alongside verse-level alignment

OPEN

Cultural and environmental context via D-PLACE, keyed to the same language identities

OPEN

Locale data from Unicode CLDR, so LOL can serve as a localisation back end

OPEN

Speech: links to Mozilla Common Voice and to endangered-language archives, respecting their access terms

OPEN

Align with Palimpsest and the Content Provenance Protocol, so provenance and consent are carried in a standard form

OPEN

Credits and licence

LOL is by Jonathan D.A. Jewell. It began as a rewrite of 1000Langs by Ehsaneddin Asgari (LMU Munich), licensed under Apache-2.0; how much of LOL derives from that code is being documented in EXPLAINME, and any derived parts keep their Apache-2.0 terms and notices.

LOL’s own work is dual-licensed under the Palimpsest-MPL-1.0 License and the Palimpsest License v0.8. Commercial use with attribution is permitted; proprietary AI training without attribution is prohibited. See LICENSE.

Source texts remain under their own licences and terms; see Data and rights.

About

Parallel corpus crawler and alignment index across 1,500+ languages, for preserving languages and cultures and connecting people through them. Starts from Bible translations for their unmatched coverage; designed to carry provenance and terms of use with every text. For research, localisation and accessibility.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages