Greenland: A Multilingual Linguistic Database and Toolchain

Hide hard words

The Problem

Building a language-learning°hard word app is a data problem disguised°hard word as a software problem.

Trakaido°hard word is a language-learning°hard word application°hard word. To teach someone Lithuanian, or Swahili°hard word, or Japanese°hard word, the app needs more than a translation°hard word dictionary°hard word. It needs to know that Lithuanian nouns°hard word decline°hard word into seven grammatical°hard word cases, and which declension°hard word class each noun°hard word belongs to. It needs to know that Chinese°hard word has no verb°hard word conjugation°hard word but every noun°hard word requires°hard word a specific°hard word measure word — you count flat things differently from round things differently from long things — and that the measure word depends°hard word on the individual°hard word noun°hard word, not just its shape. It needs to know that German nouns°hard word have grammatical°hard word gender°hard word that doesn't track biological°hard word sex or English°hard word intuition°hard word: die Sonne (the sun) is feminine°hard word, der Mond (the moon) is masculine°hard word, das Kind (the child) is neuter°hard word. It needs to know that Swahili°hard word verb°hard word infinitives°hard word take a ku- prefix°hard word. These aren't edge cases — they are the grammar°hard word, and the app needs them for every word.

It also needs example sentences, IPA°hard word pronunciations°hard word, synonyms°hard word, difficulty°hard word ratings°hard word, and translations°hard word — not into one other language, but into fourteen simultaneously°hard word, with more handled°hard word on the margins°hard word. And it needs all of this organized°hard word: not just a flat list of words, but words grouped by semantic°hard word category°hard word (food nouns°hard word, motion verbs°hard word, perception°hard word adjectives°hard word) and difficulty°hard word level, so the app can present vocabulary°hard word in a sensible°hard word pedagogical°hard word order.

That data does not exist in one place, in one consistent°hard word format°hard word, at the required°hard word level of detail, for all of these languages. So we built Greenland°hard word to create it.

What Greenland°hard word Is

Greenland°hard word is the backend°hard word data and tooling repository°hard word for Trakaido°hard word. It builds, maintains°hard word, and exports°hard word a multilingual°hard word linguistic°hard word database covering fourteen primary°hard word languages — English°hard word, Lithuanian, Chinese°hard word (Simplified°hard word), French°hard word, German, Spanish°hard word, Portuguese°hard word, Korean°hard word, Swahili°hard word, Vietnamese°hard word, Japanese°hard word, Italian, Dutch°hard word, and Swedish°hard word — plus°hard word dozens of additional°hard word languages handled°hard word for source data and supplementary°hard word content°hard word.

The database currently holds around 2,500 words across 90 semantic°hard word categories°hard word. For each word, it stores translations°hard word into all primary°hard word languages, inflected°hard word forms (conjugations°hard word, declensions°hard word, plurals°hard word), IPA°hard word pronunciations°hard word, example sentences with translations°hard word, grammar°hard word facts, synonyms°hard word, and learner difficulty°hard word levels.

Words are organized°hard word by part of speech and semantic°hard word category°hard word: food nouns°hard word, motion verbs°hard word, manner adverbs°hard word, and so on. This organization matters both pedagogically°hard word — it shapes how the app sequences°hard word material — and operationally°hard word, since the tooling processes batches°hard word of related°hard word words together.

How It Works

The Data Layer

The canonical°hard word source of truth is a set of JSONL°hard word files°hard word in the repository°hard word — one file°hard word per°hard word semantic°hard word category°hard word, one JSON°hard word record per°hard word word. Because this lives in version°hard word control, every word added, changed, or removed°hard word is visible°hard word as a code°hard word change and can be reviewed°hard word like any other commit°hard word.

The live working database is SQLite°hard word. It extends°hard word the JSONL°hard word baseline°hard word with everything generated by the automation°hard word pipeline°hard word: full morphological°hard word forms, sentences, pronunciations°hard word, grammar°hard word facts, operation logs°hard word. It can be rebuilt°hard word from the JSONL°hard word files°hard word at any time, and changes made through the tools can be exported°hard word back to JSONL°hard word to keep the repository°hard word current.

The Automation Agents

A fleet of autonomous°hard word scripts°hard word performs°hard word the bulk°hard word of data generation°hard word and quality assurance°hard word. Each agent°hard word is named after a Lithuanian animal and performs°hard word a specific°hard word task.

Voras°hard word (the Spider) generates and validates°hard word translations°hard word. Vilkas°hard word (the Wolf) generates word forms — conjugations°hard word, declensions°hard word, plurals°hard word — for six languages. Žvirblis°hard word (the Sparrow°hard word) generates example sentences. Lape°hard word (the Fox) generates language-specific°hard word grammar°hard word facts: which measure word a Chinese°hard word noun°hard word takes, the grammatical°hard word gender°hard word of French°hard word and German nouns°hard word, the declension°hard word class of Lithuanian nouns°hard word. Lokys°hard word (the Bear) validates°hard word English°hard word lemma°hard word forms and definitions°hard word. Papuga°hard word (the Parrot) generates IPA°hard word pronunciations°hard word. Šernas°hard word (the Boar°hard word) generates synonyms°hard word and alternative°hard word forms. Sarka°hard word (the Magpie°hard word) generates conversational°hard word example dialogues°hard word.

On the quality and export°hard word side: Bebras°hard word (the Beaver°hard word) checks database integrity°hard word. Povas°hard word (the Peacock) generates HTML°hard word reports. Ungurys°hard word (the Eel°hard word) and Elnias°hard word (the Deer) export°hard word the database in the formats°hard word the app consumes°hard word. Strazdas°hard word (the Thrush°hard word) and Vieversys°hard word (the Lark°hard word) generate audio°hard word files°hard word via°hard word eSpeak-NG°hard word and OpenAI°hard word TTS°hard word respectively°hard word.

The agents°hard word run in a pipeline°hard word: initialization°hard word, then°hard word translations°hard word, then°hard word word forms and grammar°hard word facts, then°hard word sentences, then°hard word audio°hard word, then°hard word export°hard word. Each supports a read-only°hard word check mode°hard word, a dry-run°hard word preview°hard word, and a fix mode°hard word that writes changes. They are idempotent°hard word and safe to run repeatedly°hard word as the database grows.

Generation°hard word is done°hard word by LLMs°hard word — OpenAI°hard word, Anthropic°hard word, and Google°hard word models, as well as local°hard word models via°hard word Ollama°hard word and LM°hard word Studio. Results are stored°hard word so each word is only processed once; humans review°hard word and correct output°hard word through the web°hard word editor°hard word.

The Benchmark Suite

Greenland°hard word includes°hard word a framework°hard word for evaluating°hard word the LLMs°hard word it relies°hard word on. Over 40 benchmarks°hard word test different skills°hard word:

- *Token and word processing*: syllable°hard word counting, spell checking, IPA°hard word transcription°hard word, plural°hard word generation°hard word

- *Lexical semantics*: antonyms°hard word, synonyms°hard word, definitions°hard word, part-of-speech°hard word tagging°hard word, lemma°hard word identification°hard word

- *Morphology*: verb°hard word form generation°hard word, lemma°hard word validation°hard word

- *Translation*: nine language pairs°hard word spanning°hard word European°hard word and Asian°hard word languages

- *Mathematics*: arithmetic°hard word, algebra°hard word, geometry°hard word, unit conversion°hard word, word problems, time arithmetic°hard word

- *General knowledge*: geography, historical°hard word dates, book-author°hard word matching, food classification°hard word, syllogism°hard word validity°hard word

Some benchmarks°hard word directly test agent°hard word decisions — for example, whether Lokys's°hard word lemma°hard word validation°hard word judgments°hard word match known-correct°hard word answers. This closes a quality loop°hard word: the models used to build the database are evaluated°hard word against ground truth, so systematic°hard word errors can be detected before they propagate°hard word through thousands°hard word of words.

Barsukas: The Web Editor

Barsukas°hard word is a local°hard word Flask°hard word application°hard word for human curators°hard word to interact°hard word with the database. It provides°hard word a full editing°hard word interface°hard word: browsing°hard word and searching°hard word the lemma°hard word inventory°hard word, editing°hard word translations°hard word and definitions°hard word, reviewing°hard word and correcting example sentences, managing°hard word pronunciations°hard word.

A sync°hard word interface°hard word compares the live database against the JSONL°hard word files°hard word — showing additions, removals°hard word, difficulty°hard word changes, and translation°hard word changes — and lets curators°hard word apply or reject°hard word each category°hard word. Selected°hard word agent°hard word workflows°hard word can be triggered°hard word from within°hard word Barsukas°hard word, so a curator°hard word can generate a translation°hard word or sentence for a specific°hard word word without going to the command line. Operation logs°hard word record what changed, when, and by which agent°hard word or user°hard word, making it possible to trace°hard word the provenance°hard word of any piece of data.

The Export Layer

Ungurys°hard word exports°hard word the database into WireWord°hard word format°hard word — the JSON°hard word structure the app consumes°hard word — organized°hard word by difficulty°hard word level and part of speech. Elnias°hard word produces a minimal°hard word bootstrap°hard word format°hard word. Audio°hard word files°hard word, generated by Strazdas°hard word and Vieversys°hard word, are uploaded°hard word to S3.

Why This Approach°hard word

The core°hard word design choice is to use LLMs°hard word for generation°hard word and humans for judgment°hard word. LLMs°hard word handle the volume: translating 2,500 words into 14 languages, generating every conjugation°hard word of every verb°hard word in six languages, identifying°hard word the measure word for every Chinese°hard word noun°hard word. Humans handle the exceptions°hard word, corrections°hard word, and quality bar°hard word — through the web°hard word editor°hard word, through reviewing°hard word commits°hard word, through the benchmark°hard word suite°hard word catching systematic°hard word failures°hard word.

The JSONL°hard word files°hard word in the repository°hard word make this sustainable°hard word. Every word has a traceable°hard word record. The database can be audited°hard word, reconstructed°hard word, and extended°hard word without losing history. New languages, new word categories°hard word, or new types of linguistic°hard word data can be added incrementally°hard word without disrupting°hard word what's already there.