УкраїнськоюEnglish
Ukrainian Frequency / Text profiler

How much of this text can your learners read?

Paste a Ukrainian text. The profiler lemmatizes it against the 12,000-lemma list and reports what share of the running words a reader knows, either at a validated CEFR level or against the exact words your class has covered. What is left comes back as a gloss list you can export.

Nothing you paste leaves your browser: the lexicon is downloaded once and the analysis runs on your device. No text is stored, logged, or sent anywhere.

Lexicon bundle (open data)Back to the frequency list
01

Measure against

Level

The validated PULS map places 6,284 of the 12,000 lemmas at a level. Words it has not reached count as unknown at every level, so this reading is a floor.

02

Text to profile

1
04

Is it right? The validation run

The profiler was run in level mode over 29 Ukrainian texts whose level is asserted by an independent source (A1: 9, A2: 10, B1: 10). For each text: strict coverage against every validated level, the first level that reaches the 95 percent target, and whether that matches the claim. The result is published as it came out.

What the texts are

placement in the course's core CEFR track (a1 / a2 / b1), aligned by the course to the Ukrainian State Standard 2024. The course modules are machine-drafted and audited by that project's own review pipeline, not graded readers vetted by a testing body. The level is the course's assertion, independent of this profiler, and weaker than a published graded reader's.

learn-ukrainian (Krisztian Koos and contributors), the open A1 to C2 Ukrainian course, CC-BY-SA-4.0 (curriculum content; LICENSE-CONTENT.md in the course repository).

For each thematic module page, keep the Ukrainian-language prose blocks (a block with more than 25 Cyrillic letters and more than four times as many Cyrillic as Latin letters) outside the learner-path, self-check, textbook-check, error-work and summary sections; drop fill-in blanks and cross-references; join in page order; em dashes normalized to a hyphen (punctuation is not profiled). Modules yielding fewer than 60 Ukrainian words are skipped, at most 10 per level.

What came out

Strict coverage never reaches the target: 0 of 29 texts match their asserted level and 28 reach 95 percent at no validated level, not even with every A1 to C1 word known. The ceiling is the map: on these texts 8 percent of running words are off the 12,000-lemma list and 6.4 percent are on it without a validated level.

The coarse ordering holds, and only the coarse one. At a fixed level (A1 to B1 known) the A1 texts read at 84.3 percent, the A2 texts at 84.8 and the B1 texts at 77.4. The B1 texts sit 6.9 points below the easier two, but A1 and A2 are 0.5 apart, which nine and ten texts cannot separate. This set tells B1 from the A levels; it does not tell A1 from A2.

On the conditional reading (only the words that carry a validated level in the denominator) 2 of 29 match exactly and 26 are within one level: the profile reads almost every text one level above the course's claim. Either the course's texts run a level ahead of the PULS onset levels, or 95 percent is too strict a bar for a lesson text. This set cannot tell which, and the page does not pretend it can.

Validation set 2026-09-13, run 2026-09-13, threshold 95 percent.validation.json
05

How it counts, and where it is not to be trusted

Coverage is running-text (token) coverage: known words over all Ukrainian words, proper names set aside. Latin words and numbers are not counted. A word is known when its lemma is in the known set; an inflected form is mapped to its lemma through the Wiktionary paradigm tables of thread 04.

That table covers 70.6 percent of the lemmas. A real form of a lemma without a table reads as off-list, so off-list counts are an upper bound and coverage a lower bound. The source list is news-heavy: everyday words like снідати or сметана are not on it at all and stay off-list at every level.

Levels are the validated PULS onset levels only (6,284 of 12,000 lemmas). Nothing is guessed. Family links come from the Wiktionary derivational graph (1,917 families touching the list) and aspect links from the B1 relation table, so a verb outside that table has no partner here. Straddle warnings use the validated levels, so an unlabeled relative cannot raise one.

Two published sets ship: the Borodin A1 and A2 lexical minimums (CC BY 4.0). Multiword expressions (connectors, the Borodin phrases) are not matched yet; a phrase is counted word by word.

Laufer and Nation 1995 (the lexical frequency profile); Laufer 1989 and Hu and Nation 2000 (95 percent); Nation 2006 (98 percent).

Need something like this built?

I design and ship AI tools, full-stack apps, and data pipelines — end to end, to production. Tell me the problem in a sentence; I'll give you an honest read on fit within a day.

Work with me →