A lemma-based, coverage-graded frequency list for modern Ukrainian, built as open data. It answers a practical question with a real number: how much of what you read does a given amount of vocabulary actually unlock?
The report that started this series was blunt about the gap: for Ukrainian, the frequency lists a serious learner needs are thin, often wordform-based, and sometimes quietly calqued from Russian. A wordform list scatters a single word's frequency across a dozen or more inflected forms, so it understates how much a given amount of vocabulary buys you.
This list counts lemmas instead, the dictionary form a learner actually studies, with its part of speech. It is derived from a large, modern, native-authored corpus, so it reflects the Ukrainian people actually write today, not a mid-century textbook.
For comparison, in English the ~2800 words of the General Service List average about 82% coverage. Ukrainian needs far more lemmas to reach the same place. That is the inflection tax, measured.
The most frequent lemmas, with part of speech and per-mille frequency. The full 12,000 are in the download.
| # | Lemma | POS | per mille |
|---|---|---|---|
| 1 | в | ADP | 20.7489 |
| 2 | на | ADP | 20.6735 |
| 3 | у | ADP | 18.6486 |
| 4 | і | CCONJ | 16.5815 |
| 5 | з | ADP | 15.259 |
| 6 | бути | VERB | 10.6332 |
| 7 | не | PART | 10.4292 |
| 8 | що | SCONJ | 10.3296 |
| 9 | до | ADP | 9.1939 |
| 10 | за | ADP | 8.5376 |
| 11 | та | CCONJ | 8.4347 |
| 12 | який | DET | 7.9025 |
| 13 | рік | NOUN | 6.6948 |
| 14 | про | ADP | 6.2134 |
| 15 | він | PRON | 6.1849 |
| 16 | Україна | PROPN | 5.8085 |
| 17 | це | PRON | 5.7827 |
| 18 | а | CCONJ | 4.994 |
| 19 | цей | DET | 4.3466 |
| 20 | як | SCONJ | 4.203 |
| 21 | для | ADP | 4.1004 |
| 22 | від | ADP | 3.9154 |
| 23 | я | PRON | 3.642 |
| 24 | вони | PRON | 3.403 |
Part-of-speech make-up of the released list. Proper nouns are included and labeled, so you can filter them out.
Thread two set out to label every high-frequency lemma with a CEFR level, and now it can. There is exactly one corpus-based, expert-validated CEFR vocabulary profile for Ukrainian (Synchak and colleagues, 2025, on the PULS platform), and a co-author granted permission to redistribute its per-lemma A1 to C1 levels. Joined onto this frequency list, POS-aware, that validated map now labels 6,207 of the top 12,000 lemmas. It is the headline of this thread, below.
Two honest companions stand beside it. The paper's published distribution is still cited (a frozen view the live map already exceeds), cross-referenced against this list's coverage curve to show why a frequency rank is not a proficiency level. And the open, unvalidated course overlay stays on as a direct cross-check: now that a validated map exists, the free signal can be measured against it. The validated map is a dated snapshot, not a final one, because the profile is still being extended.
Lemmas placed by the expert-validated process (Synchak et al. 2025) against the project's per-level target. Complete through B1, preliminary above it.
The validated A1 to A2 core (2,363 lemmas) versus a same-size core taken straight off this frequency list, by running-text coverage. The two are measured on different corpora, so read the spread as indicative: pedagogical selection deliberately trades raw coverage for teachability.
Why they diverge: the validated profile front-loads communicatively useful words that are not the most frequent (борщ borscht, зошит workbook) and pushes high-frequency abstract nouns (наука science, освіта education) up to B1 for conceptual load. A frequency list optimizes coverage; a CEFR level encodes teachability. That is why you cannot mechanically relabel one as the other, and why a real CEFR overlay here waits on open data above B1.
This is the headline of thread two: the one corpus-based, expert-validated CEFR profile for Ukrainian (Synchak et al. 2025, the PULS platform), joined onto this frequency list per lemma. A co-author granted permission to redistribute the per-lemma A1 to C1 levels; the join is POS-aware, so a homograph is dropped rather than mislabeled. Each lemma's level is its onset level, the earliest at which the profile introduces it.
It is a dated snapshot, not the final map. The profile is still being extended (the live list already carries more B2 and C1 lemmas than the published paper), so treat these levels as validated but provisional.
6,207 of the top 12,000 lemmas now carry a validated CEFR level, A1 to C1
Only weakly, which is the finding. The rank correlation between validated level and corpus frequency is 0.379: the median frequency rank rises cleanly from A1 to C1, so the ordering is sane, but the fit is loose. A CEFR level and a frequency rank are simply not the same axis, which is why you cannot relabel one as the other.
Measured on the 1,965 lemmas that both the validated map and the open course overlay cover, comparing the course's raw first-occurrence guess against the validated level. A far stronger test than the small hand-curated gold sample: it is the whole overlap.
| # | Lemma | CEFR |
|---|---|---|
| 1 | в | A1 |
| 2 | на | A1 |
| 3 | у | A1 |
| 4 | і | A1 |
| 5 | з | A1 |
| 6 | бути | A1 |
| 8 | що | A2 |
| 11 | та | A2 |
| 20 | як | A2 |
| 38 | під | A2 |
| 45 | по | A2 |
| 47 | після | A2 |
| 61 | заявити | C1 |
| 69 | повідомити | B1 |
| 72 | рада | B2 |
| 73 | і | B1 |
| 75 | ж | B1 |
| 78 | й | B1 |
| 80 | при | B1 |
| 83 | повідомляти | B1 |
| 121 | зазначити | B2 |
| 155 | представник | B2 |
| 213 | верховний | B2 |
| 225 | управління | B2 |
| 252 | примітка | B2 |
| 292 | глава | C1 |
| 411 | чин | C1 |
| 429 | розслідування | C1 |
| 622 | виборчий | C1 |
| 752 | законодавство | C1 |
The validated map above needs the authors' permission to redistribute. This overlay is the fully open alternative: it labels our lemmas from the learn-ukrainian open course (Krisztian Koos, CC-BY-SA), an expert selection based on Ukraine's 2024 State Standard, no permission needed. A lemma's level here is the first course lesson that introduces it, so treat it as a rough teaching-order signal, not a proficiency label.
Its role now is a cross-check, not a fallback. The validated overlay above measures the two against each other on the full overlap; this card keeps the open overlay's own calibration against the published gold points. Neither the open map alone nor the paper's frozen tables is the finished thing a learner deserves, and closing that gap is why this series exists.
2,326 of the top 12,000 lemmas carry an open CEFR label, A1 to B2
Synchak et al. publish concrete CEFR levels for specific lemmas. Where those overlap this overlay, the disagreement is measured directly, so the overlay is calibrated, not just asserted. The sample is small (a few dozen gold points), so read the interval, not the point estimate.
| Lemma | Validated | Overlay | |
|---|---|---|---|
| бути | A1 | A1 | ✓ |
| рік | A1 | A1 | ✓ |
| він | A1 | A1 | ✓ |
| для | A1 | B1 | ≠ |
| я | A1 | A1 | ✓ |
| вони | A1 | A1 | ✓ |
| також | A1 | A1 | ✓ |
| вона | A1 | A1 | ✓ |
| час | A1 | A1 | ✓ |
| мати | A1 | A1 | ✓ |
| могти | A1 | A1 | ✓ |
| людина | A1 | A1 | ✓ |
| після | A2 | A1 | ≠ |
| новий | A1 | A1 | ✓ |
| перший | A1 | A2 | ≠ |
| країна | A1 | A1 | ✓ |
| сказати | A1 | A1 | ✓ |
| стати | A1 | A2 | ≠ |
| особа | A2 | A1 | ≠ |
| через | A1 | B1 | ≠ |
| область | A1 | B1 | ≠ |
| заявити | C1 | B1 | ≠ |
| або | A1 | A1 | ✓ |
| щоб | A2 | A1 | ≠ |
| голова | A1 | A1 | ✓ |
| два | A1 | A1 | ✓ |
| місто | A1 | A1 | ✓ |
| повідомити | B1 | A2 | ≠ |
| слово | A1 | A1 | ✓ |
| день | A1 | A1 | ✓ |
| місце | A1 | A1 | ✓ |
| робота | A1 | A1 | ✓ |
| повідомляти | B1 | B1 | ✓ |
| себе | A2 | A2 | ✓ |
| якщо | A1 | A1 | ✓ |
| питання | A1 | A1 | ✓ |
| великий | A1 | A1 | ✓ |
| щодо | B1 | B1 | ✓ |
| рішення | B1 | B1 | ✓ |
| компанія | A1 | A2 | ≠ |
| район | A1 | A1 | ✓ |
| раніше | A2 | A1 | ≠ |
| отримати | A1 | A2 | ≠ |
| частина | A1 | A1 | ✓ |
| партія | B1 | B2 | ≠ |
| хто | A1 | A1 | ✓ |
| працювати | A1 | A1 | ✓ |
| результат | A1 | A2 | ≠ |
| можна | A1 | A1 | ✓ |
| дитина | A1 | A1 | ✓ |
| головний | A1 | A2 | ≠ |
| село | A1 | A1 | ✓ |
| склад | B1 | A1 | ≠ |
| держава | A2 | B2 | ≠ |
| чоловік | A1 | A1 | ✓ |
| хотіти | A1 | A1 | ✓ |
| життя | A1 | A2 | ≠ |
| право | B1 | B2 | ≠ |
| сьогодні | A1 | A1 | ✓ |
| уряд | B1 | B2 | ≠ |
| між | A2 | A2 | ✓ |
| вибори | B1 | B1 | ✓ |
| тисяча | A1 | A1 | ✓ |
| знати | A1 | A1 | ✓ |
| дія | B1 | A1 | ≠ |
| ситуація | A1 | B2 | ≠ |
| тут | A1 | A1 | ✓ |
| зараз | A1 | A1 | ✓ |
| центр | A1 | A1 | ✓ |
| становити | B1 | B2 | ≠ |
| допомога | A1 | A1 | ✓ |
| три | A1 | A1 | ✓ |
| березень | A1 | A1 | ✓ |
| зробити | A1 | B1 | ≠ |
| місцевий | A2 | A2 | ✓ |
| кількість | B1 | A2 | ≠ |
| початок | A1 | A1 | ✓ |
| другий | A1 | B1 | ≠ |
| почати | A1 | B1 | ≠ |
| гривня | A1 | A1 | ✓ |
| стан | B1 | B1 | ✓ |
| там | A1 | A1 | ✓ |
| писати | A1 | A1 | ✓ |
| мета | B1 | A2 | ≠ |
| освіта | B1 | A2 | ≠ |
| жити | A1 | A1 | ✓ |
| написати | A1 | A2 | ≠ |
| чорний | A1 | A1 | ✓ |
| наука | B1 | A2 | ≠ |
| білий | A1 | A1 | ✓ |
| підписати | A2 | B1 | ≠ |
| займатися | A2 | A2 | ✓ |
| червоний | A1 | A1 | ✓ |
| стіл | A1 | A1 | ✓ |
| зелений | A1 | A1 | ✓ |
| традиція | A1 | A2 | ≠ |
| підпис | A1 | B1 | ≠ |
| ліжко | A1 | A1 | ✓ |
| синій | A1 | A1 | ✓ |
| крісло | A1 | A1 | ✓ |
| годинник | A1 | B1 | ≠ |
| помаранчевий | A2 | A1 | ≠ |
| цікавитися | A2 | A2 | ✓ |
| типовий | A2 | A1 | ≠ |
| стілець | A1 | A1 | ✓ |
| меблі | A1 | A2 | ≠ |
| захоплюватися | A2 | A2 | ✓ |
| посуд | A1 | B1 | ≠ |
| допис | B1 | B1 | ✓ |
| лампа | A1 | A1 | ✓ |
| шафа | A1 | B1 | ≠ |
| холодильник | A1 | A2 | ≠ |
| рушник | A1 | A1 | ✓ |
| диван | A1 | A1 | ✓ |
Only 114 of 200 gold points fall inside the top 12,000 and share the overlay's part of speech; the rest are rarer words, or homographs of a different sense. The misses run both ways: more often the course places a concrete but corpus-rare noun (шафа wardrobe, посуд dishes) higher than the validated profile, which front-loads it to A1 for a beginner; less often it puts an abstract noun (наука science, освіта education) a level lower. Much of that gap is not error but two philosophies: the course follows where a word is taught, the profile makes deliberate beginner-priority overrides.
The join is POS-aware, so 75 homographs (a fish and a quantifier that share a spelling, for instance) were dropped rather than mislabeled. And where the validated profile publishes a level, that level is used and marked validated, so 114 labels here are Synchak's own, not the heuristic's. The agreement above is measured on the heuristic BEFORE those overrides, so the number is not flattered by them.
Weakly, which is the honest answer. The rank correlation between assigned level and corpus frequency is 0.283 (a CEFR level is not frequency, so a loose fit is expected), but the median frequency rank does rise from A1 to B2, so the ordering is at least directionally sane.
| # | Lemma | CEFR |
|---|---|---|
| 1 | в | A1 |
| 2 | на | A1 |
| 3 | у | A1 |
| 4 | і | A1 |
| 5 | з | A1 |
| 6 | бути | A1 |
| 9 | до | B1 |
| 10 | за | A2 |
| 22 | від | B1 |
| 38 | під | A2 |
| 45 | по | B1 |
| 47 | після | A2 |
| 57 | особа | A2 |
| 61 | заявити | C1 |
| 63 | щоб | A2 |
| 69 | повідомити | B1 |
| 83 | повідомляти | B1 |
| 85 | себе | A2 |
| 90 | щодо | B1 |
| 123 | дані | B2 |
| 284 | матч | B2 |
| 330 | парламент | B2 |
| 337 | акція | B2 |
| 401 | порушення | B2 |
| 406 | належати | B2 |
The field census below used to list the "1000 i 1 slovo" A1 lexical minimum (Kseniia Borodin and Oksana Turkevych, School of Ukrainian Language and Culture, Ukrainian Catholic University) as walled: licensed CC-BY-4.0 but published only as a prose PDF an extractor could not reach, so the one independent, empirically grounded A1 guideline for Ukrainian could not be joined and the overlay's rigor model called that missing consensus a seam. This page asked the authors for a machine-readable table, and co-author Kseniia Borodin sent the word list herself. This card is that list, joined onto the frequency list and republished as attributed open data under the guide's own license.
The join is honest about its grain. The authors' unit is the published entry (a single word, an aspect pair, a masculine and feminine pair, sometimes a whole phrase), the list carries no part of speech, and the join is by lemma string at the best rank, coarser than the POS-aware overlays above. Phrase entries are counted, never force-joined, and a beginner word the corpus does not rank stays honestly off-list.
838 of the 929 unique A1 entries join the top 12,000, reaching 833 distinct lemmas
Of the 768 joined lemmas the validated PULS map also covers, 75.4 percent are validated A1 and 96.5 percent sit at A1 or A2. Two independent efforts, a corpus-validated profile and an expert lexical minimum, converge on what beginner Ukrainian is. The handful above A2 is mostly the price of a coarser join: автомат the A1 ticket machine shares a spelling with the C1 rifle, and this list cannot tell them apart without a part of speech.
The 833 matched lemmas cover about 34.7 percent of running text; the 833 highest-frequency lemmas would cover about 57.5 percent. That is the same finding as the validated core above, now measured on a source this page did not build: a pedagogical A1 selection deliberately trades raw coverage for concrete, teachable words, and 72 of its word entries (апельсин orange, виделка fork, бутерброд sandwich) do not rank in this news-heavy top 12,000 at all.
| # | Published entry | Lemma | PULS |
|---|---|---|---|
| 1 | в, у | в | A1 |
| 2 | на | на | A1 |
| 4 | і, й, та | і | A1 |
| 5 | з, із, зі | з | A1 |
| 6 | бути | бути | A1 |
| 7 | не | не | A1 |
| 1536 | готель | готель | A1 |
| 1545 | стать | стать | A2 |
| 1546 | дочка | дочка | A1 |
| 1549 | готувати | готувати | A1 |
| 1566 | папір | папір | A1 |
| 1570 | купувати – купити | купити | A1 |
| 11650 | цукерка | цукерка | A1 |
| 11758 | курка | курка | A1 |
| 11839 | крем | крем | A2 |
| 11871 | фотографувати | фотографувати | A1 |
The A1 card above closed a seam. Its provenance note also recorded an open one, in plain words: the A2 volume of the same series existed, but no machine-readable A2 list had been shared. That volume has since been published open access under CC-BY-4.0 by Kseniia Borodin and Olesia Lazarenko at the Europa-Universitat Viadrina, so the seam closes by publication rather than by favor. This card is its 1,000-entry register, extracted from the publication's own PDF, joined onto the frequency list the same coarse way, and republished as attributed open data.
Two things become answerable that were not before. Because this is the same authors' next volume, the two lists together measure what one CEFR step actually costs, in frequency ranks rather than in word counts. And because the A2 book also prints its whole vocabulary regrouped under its own 51 topic headings, this is the first source on this page that can say WHICH semantic fields a news-heavy frequency list misses, not merely how many words it misses.
708 of the 1,000 published A2 entries join the top 12,000, reaching 707 distinct lemmas
Only 45 of the A2 entries are words the same authors' A1 minimum already claimed, and those are almost all verbs collecting their perfective partner. The other 955 (95.5 percent) are new at A2, and they sit measurably deeper in the corpus: median rank 2,902 against 1,643 for the words carried up from A1. A CEFR step is not a bigger pile of the same words. It is a move outward into rarer ones.
The book groups its 1,124 thematic units under 51 topic headings. Joined against the frequency list, 21.6 percent of them do not rank in the top 12,000 at all, and the misses are not evenly spread. The worst-covered fields are the concrete, domestic, physical ones: food, personal hygiene, household technology, and the numerals, where both the collective forms and the ordinals the authors deliberately extended to thirtieth (so a learner could say a date) fall outside the list. A corpus of published Ukrainian news is a poor description of a kitchen.
Of the 656 joined lemmas the validated PULS map also covers, 70.3 percent sit at A1 or A2 and 29.7 percent sit above A2, nearly all at B1. That is a looser agreement than the A1 card found, and it should be: A2 is where two expert efforts start to disagree about the boundary, and the coarse lemma-string join adds noise of its own. The page reports the disagreement rather than tuning it away.
The 707 matched lemmas cover about 11.9 percent of running text; the 707 highest-frequency lemmas would cover about 55.1 percent. The trade the A1 card measured widens sharply at A2, and 229 of the word entries do not rank in this news-heavy top 12,000 at all, against 72 at A1. This is the honest ceiling on frequency-first learning: past the beginner core, a frequency list and a syllabus stop describing the same language.
| # | Published entry | Lemma | PULS |
|---|---|---|---|
| 9 | до | до | A1 |
| 21 | для | для | A1 |
| 22 | від | від | A1 |
| 25 | свій | свій | · |
| 27 | Бува́й(те)! | те | A1 |
| 35 | той | той | A1 |
| 2771 | дах | дах | A2 |
| 2814 | Бе́льгія | Бельгія | · |
| 2823 | звук (Р.в. одн. зву́ка) | звук | A2 |
| 2837 | відві́дувач / відві́дувачка | відвідувач | B1 |
| 2841 | зустріча́ти(ся) – зустрі́ти(ся) | зустріти | A1 |
| 2844 | підзе́мний | підземний | B1 |
| 11797 | А́рктика | Арктика | · |
| 11811 | волейбо́л | волейбол | A2 |
| 11925 | симпати́чний | симпатичний | B1 |
| 11969 | па́ска | паска | A2 |
The card above says 229 word entries of the A2 minimum do not rank in this list's top 12,000, and reads that as the ceiling on learning by frequency. That reading rests on something untested. This list is built on UberText 2.0's COMBINED corpus, and the largest thing in that corpus by a wide margin is news. So an off-list word might be rare in Ukrainian, or it might be ordinary and simply absent from journalism, which does not often discuss forks, pumpkins or neckties. Those two readings give opposite advice to anyone building a word list, and nothing above can tell them apart.
So we built two frequency lists instead of one, from UberText's own fiction and news subcorpora: same collection, same cleaning, same lemmatizer, same sampling shape, both capped to the same token count. Genre is the only thing that differs. The ranks below are internally comparable and deliberately NOT comparable to the main list on this page, which a different lemmatizer produced.
News reaches 658 of the entries (70.3 percent). Fiction reaches 648 (69.2 percent). Almost the same amount, and not the same words: together they reach 744 (79.5 percent). A second genre of the same size buys 9.2 points that more of the first genre would not have bought. For a pedagogical target, corpus diversity is worth more than corpus size.
The exchange is close to symmetric, which is what makes it a real result rather than a bigger hammer: fiction reaches 30.9 percent of what news misses, and news reaches 33.3 percent of what fiction misses. Neither genre is better. They are looking at different parts of the language.
Read the two rows together and the mechanism is plain. Journalism supplies the institutional and the geographic; fiction supplies the domestic, the bodily and the sensory. Neither is a description of the language on its own.
These ranks come from our own lemmatizer pass over two subcorpora, not from the lemmatizer behind the main list on this page. They are comparable to each other and to nothing else here, and no figure above mixes them with the main list. The lemmatizer also mis-handles a few high-frequency pronouns, identically in both genres, so it cannot manufacture a difference between them.
Thread two now joins a validated profile (by the authors' permission), an author-shared A1 lexical minimum, and an open course overlay. That it took two direct asks to get here is the point: the open, per-lemma, CEFR-graded word list, a standard resource for learning a language, has been built for about a dozen languages and never published as open Ukrainian data. Here is the whole field.
| Source | Access | Levels |
|---|---|---|
| Ukrainian Vocabulary Profile (PULS)validated | by permission | A1-C1 (to B1) |
| State Standard 2024 (SLSUFL) | documents only | A1-C2 |
| CEFR & Ukrainian-English Language Portfolios (PCUH) | documents only | A1-B1 |
| learn-ukrainian course | open, CC-BY-SA | A1-B2 |
| 1000+1 Words (Borodin, Turkevych) | shared by the authors | A1 |
| 1000+1 Words A2 (Borodin, Lazarenko) | open, CC-BY | A2 |
This page folds in an open Ukrainian CEFR word list the moment one exists, and that is not hypothetical: the A1 minimum above arrived exactly this way, because its authors answered an email. The schema it needs is small: lemma, part of speech, and CEFR level, as CSV or JSON.
A frequency list tells you which words to learn. It does not tell you that some of the words you will meet in real Ukrainian text are ones a careful editor would replace: russianisms, surzhyk, and calques that older teaching materials quietly passed on. Thread three flags them on the exact list you study from.
There is no single settled russianism list to join against. What there is, open and maintained, is the Ukrainian idiomatic-style tooling: LanguageTool's Ukrainian replacement tables (the brown-uk project, the same one behind the VESUM morphology). This thread joins that tooling to the frequency list and, for every top-12,000 lemma it flags, shows the native alternative it recommends.
Read this the way you read thread two's open CEFR overlay: it is an OPEN, joinable signal, but an unvalidated, prescriptive one. It is what a style checker would replace, an editorial opinion, not a settled linguistic verdict. Its core is derussification, but the tooling's scope is broader than strictly-Russian loans, and a few entries (поліцейський, безкоштовний) are live usage debates. Proper-noun and spelling fixes are dropped. Every suggested form is the tooling's own, kept verbatim.
87 of the top 12,000 lemmas carry a flag from the style tooling
Of the 74 flagged forms whose native replacement is a single word, 48 (64.9%) have the flagged form at least as common in real Ukrainian text as its replacement: the idiomatic word is rarer, or not in the top 12,000 at all. That is the empirical face of older materials teaching the calque, and the corpus still carrying it.
One row per frequency band, lowest rank first. Each flagged form, the native alternative the tooling recommends, and where that alternative sits on the list. The full list is in the download.
| # | Flagged form | Recommended | Native rank |
|---|---|---|---|
| 1,255 | поліцейськийNOUN | →поліційний | not in top 12,000 |
| 2,204 | прийомNOUN | →приймання | not in top 12,000 |
| 3,086 | безкоштовнийADJ | →безплатний | not in top 12,000 |
| 4,188 | безкоштовноADV | →безплатно | not in top 12,000 |
| 5,251 | передвиборнийADJ | →передвиборчий | #4,420↑ |
| 6,676 | недостовірнийADJ | →невірогідний | not in top 12,000 |
| 7,121 | місцезнаходженняNOUN | →місце | #81↑ |
| 8,004 | кримчанинNOUN | →кримець | not in top 12,000 |
| 9,167 | достовірнийADJ | →вірогідний | not in top 12,000 |
| 10,308 | подаліADV | →трохи далі | a construction |
| 11,027 | благополуччяNOUN | →добра доля | a construction |
11 of the flagged adjectives are the active-participle calque (діючий, існуючий): the -ючий forms Russian has and standard Ukrainian rewrites as a relative clause (що діє). The tooling lists that rewrite first.
Every thread so far has worked on the bare lemma, the dictionary headword. But a learner of a heavily inflected language does not meet headwords in text; they meet the forms. рік (year) shows up as року, роки, років, рокам. A verb has dozens of shapes. Thread four attaches those forms, and the sound, across the whole top-12,000 list.
The source is open and machine-readable: the Wiktionary declension and conjugation tables (via the kaikki.org extraction), plus the IPA and the native-speaker pronunciation audio that live in the same entries on Wikimedia Commons. Forms are kept verbatim, stress marks and all; the audio is linked, never re-hosted; and coverage is graded honestly rather than assumed.
Wiktionary + Wikimedia, CC BY-SAThe 8,339 lemmas in the top 12,000 that inflect expand to 125,380 distinct wordforms, an average of 15 each. Thread one measured the inflection tax as slow reading coverage; this is the same tax counted in the surface forms you actually have to recognize. A noun runs about ten forms, a verb closer to thirty.
A few high-frequency words, declined or conjugated, with IPA and open native audio. Pick a word; the forms are Wiktionary's own.
This is what turns the thread-one Anki seed from a bare headword list into a deck that teaches each word with its forms and its sound.
All five threads are released. Each names the open source it is built on, and every figure on this page re-derives from a committed release. No thread claims a result it cannot show.
The top 12000 lemmas of modern Ukrainian, ranked by corpus frequency, each with its part of speech, per-mille frequency, cumulative running-text coverage, and a thousand-rank frequency band. Published as CSV and JSON, plus an Anki-importable deck seed.
Study the words that actually pay off first, in the right order, as lemmas rather than scattered wordforms. See exactly how much reading coverage a given amount of vocabulary buys.
A validated per-lemma CEFR overlay: the only corpus-validated CEFR profile for Ukrainian (Synchak et al. 2025 / PULS) joined onto the frequency list, A1 to C1, for 6207 of the top 12000 lemmas, redistributed with the authors' permission as a dated snapshot. The published paper's distribution stays cited beside it, the open course overlay is kept as a measured cross-check, and the finding holds: the validated level tracks frequency only weakly, so a frequency rank is not a proficiency level.
Study the high-frequency words in validated CEFR order, see which half of the list has a validated level and how far it reaches, and see how weakly proficiency level and raw frequency actually line up.
Flag non-idiomatic forms on the list and pair each with the native alternative the open Ukrainian style tooling recommends (LanguageTool's replacement tables, brown-uk, backed by the VESUM morphology). The core is derussification (russianisms, surzhyk, the active-participle calque); the honest scope is broader, and it is fenced as an unvalidated, prescriptive signal, not a settled verdict.
Learn the idiomatic word, not the calqued one that older materials quietly teach, and see how often the calque is actually the more common form in real text.
Attach a pronunciation and the inflection paradigm to the top 12000 lemmas, from the open Wiktionary declension and conjugation tables (kaikki.org) and open Wikimedia pronunciation audio, so a lemma is taught as its forms rather than a bare headword. Coverage is graded honestly (it thins with rank), and the inflection multiplier is measured.
Hear the word, see how it changes across the seven cases and the aspect pairs (the part inflection makes hard), and see how many wordforms the words you are learning actually expand to.
A self-scoring yes or no vocabulary-size test drawn from this list: 60 real lemmas and 60 constructed pseudowords, stratified across the 12 frequency bands, scored with the standard false-alarm correction and mapped onto the coverage curve. No Ukrainian equivalent exists off the shelf.
Measure your vocabulary against a real proficiency proxy instead of guessing, and see roughly how much running text that vocabulary actually unlocks.
Everything on this page re-derives from these files. They are versioned and hashed. The list is CC-BY: use it, fork it, build on it.
The lemmas.anki.tsv file imports straight into Anki as a frequency deck seed.
Lawrence, J. (2026). Ukrainian Frequency: a lemma-based, coverage-graded open frequency list (v0.1.0). jakelawrence.xyz/research/ukrainian-frequency. Derived from UberText 2.0 (Chaplynskyi, 2023).
These checks re-run in your browser against the shipped numbers. If any fails, the release is inconsistent and the page says so.
The frequency signal is the UberText 2.0 lemma frequency dictionary (lang-uk). This release re-publishes derived aggregate counts, which are facts, not the source corpus text. Rows are filtered to Ukrainian-Cyrillic-alphabetic lemmas carrying a linguistic Universal Dependencies part of speech; punctuation, symbols, digit tokens, foreign or unclassified tokens, and single-character proper nouns are dropped.
Coverage is recomputed here from raw occurrence counts, not read from the source file's own frequency columns, whose normalization is undocumented. Frequency bands are transparent thousand-rank bins and are NOT CEFR levels; validated CEFR alignment is thread two, which profiles the one validated resource rather than inventing levels. The corpus is news and reference heavy, so war-era and civic vocabulary rank prominently. That is disclosed, not smoothed away.
Part of a standing effort to close the digital-tooling gap for the Ukrainian language.
I design and ship AI tools, full-stack apps, and data pipelines — end to end, to production. Tell me the problem in a sentence; I'll give you an honest read on fit within a day.
Work with me →