EnglishУкраїнською
Ukrainian Frequency / Field Data

The words that pay off first

A lemma-based, coverage-graded frequency list for modern Ukrainian, built as open data. It answers a practical question with a real number: how much of what you read does a given amount of vocabulary actually unlock?

Open datasetRefreshed 2026-07-28 · v0.1.0 · CC-BY-4.0

Why a frequency list, and why lemmas

The report that started this series was blunt about the gap: for Ukrainian, the frequency lists a serious learner needs are thin, often wordform-based, and sometimes quietly calqued from Russian. A wordform list scatters a single word's frequency across a dozen or more inflected forms, so it understates how much a given amount of vocabulary buys you.

This list counts lemmas instead, the dictionary form a learner actually studies, with its part of speech. It is derived from a large, modern, native-authored corpus, so it reflects the Ukrainian people actually write today, not a mid-century textbook.

12,000
lemmas ranked
1.5B
word tokens in the source corpus
60.74%
covered by the top 1,000 lemmas
82%
covered by the top 5,000 lemmas
Coverage

The coverage curve

0%25%50%75%100%1005001,0005,00012,000Lemmas learned (most frequent first)
What share of running word tokens the top N lemmas account for. The curve is the honest answer to how far a vocabulary goes.

For comparison, in English the ~2800 words of the General Service List average about 82% coverage. Ukrainian needs far more lemmas to reach the same place. That is the inflection tax, measured.

The list

The list itself

The most frequent lemmas, with part of speech and per-mille frequency. The full 12,000 are in the download.

#LemmaPOSper mille
1вADP20.7489
2наADP20.6735
3уADP18.6486
4іCCONJ16.5815
5зADP15.259
6бутиVERB10.6332
7неPART10.4292
8щоSCONJ10.3296
9доADP9.1939
10заADP8.5376
11таCCONJ8.4347
12якийDET7.9025
13рікNOUN6.6948
14проADP6.2134
15вінPRON6.1849
16УкраїнаPROPN5.8085
17цеPRON5.7827
18аCCONJ4.994
19цейDET4.3466
20якSCONJ4.203
21дляADP4.1004
22відADP3.9154
23яPRON3.642
24вониPRON3.403
Where the milestones land
60.74%
Top 1,000 lemmas: 60.74% of text
#1,000: тощо PART
70.5%
Top 2,000 lemmas: 70.5% of text
#2,000: ліквідація NOUN
82%
Top 5,000 lemmas: 82% of text
#5,000: перевозити VERB
88.39%
Top 10,000 lemmas: 88.39% of text
#10,000: безсмертний ADJ

What kind of words they are

Part-of-speech make-up of the released list. Proper nouns are included and labeled, so you can filter them out.

NOUN5,154 (43%)
ADJ2,418 (20.2%)
VERB2,109 (17.6%)
PROPN1,335 (11.1%)
ADV641 (5.3%)
PART87 (0.7%)
ADP80 (0.7%)
DET48 (0.4%)
SCONJ32 (0.3%)
NUM31 (0.3%)
PRON30 (0.3%)
CCONJ25 (0.2%)
INTJ10 (0.1%)
Ukrainian Frequency / CEFR alignment

What a validated CEFR map can and cannot tell you

Thread two set out to label every high-frequency lemma with a CEFR level, and now it can. There is exactly one corpus-based, expert-validated CEFR vocabulary profile for Ukrainian (Synchak and colleagues, 2025, on the PULS platform), and a co-author granted permission to redistribute its per-lemma A1 to C1 levels. Joined onto this frequency list, POS-aware, that validated map now labels 6,207 of the top 12,000 lemmas. It is the headline of this thread, below.

Two honest companions stand beside it. The paper's published distribution is still cited (a frozen view the live map already exceeds), cross-referenced against this list's coverage curve to show why a frequency rank is not a proficiency level. And the open, unvalidated course overlay stays on as a direct cross-check: now that a validated map exists, the free signal can be measured against it. The validated map is a dated snapshot, not a final one, because the profile is still being extended.

The validated profile, by level
A1969 / 700
A21,394 / 1,000
B12,141 / 1,500
B21,140 / 1,900
C1220 / 2,300
C227 / 2,600
placed (A1–B1) placed (B2–C2) target

Lemmas placed by the expert-validated process (Synchak et al. 2025) against the project's per-level target. Complete through B1, preliminary above it.

A frequency rank is not a proficiency level

The validated A1 to A2 core (2,363 lemmas) versus a same-size core taken straight off this frequency list, by running-text coverage. The two are measured on different corpora, so read the spread as indicative: pedagogical selection deliberately trades raw coverage for teachability.

Why they diverge: the validated profile front-loads communicatively useful words that are not the most frequent (борщ borscht, зошит workbook) and pushes high-frequency abstract nouns (наука science, освіта education) up to B1 for conceptual load. A frequency list optimizes coverage; a CEFR level encodes teachability. That is why you cannot mechanically relabel one as the other, and why a real CEFR overlay here waits on open data above B1.

Validated / Snapshot

The validated CEFR map, joined onto the list

Snapshot 2026-07-23

This is the headline of thread two: the one corpus-based, expert-validated CEFR profile for Ukrainian (Synchak et al. 2025, the PULS platform), joined onto this frequency list per lemma. A co-author granted permission to redistribute the per-lemma A1 to C1 levels; the join is POS-aware, so a homograph is dropped rather than mislabeled. Each lemma's level is its onset level, the earliest at which the profile introduces it.

It is a dated snapshot, not the final map. The profile is still being extended (the live list already carries more B2 and C1 lemmas than the published paper), so treat these levels as validated but provisional.

6,207 of the top 12,000 lemmas now carry a validated CEFR level, A1 to C1

Validated labels by level
A1812
A2956
B11,564
B22,130
C1745
How much of the top N is validated
Top 1,000: 827 validated (82.7%)
Top 2,000: 1,612 validated (80.6%)
Top 5,000: 3,495 validated (69.9%)
Top 12,000: 6,207 validated (51.7%)
Does a validated level track frequency?

Only weakly, which is the finding. The rank correlation between validated level and corpus frequency is 0.379: the median frequency rank rises cleanly from A1 to C1, so the ordering is sane, but the fit is loose. A CEFR level and a frequency rank are simply not the same axis, which is why you cannot relabel one as the other.

A11,595
A23,185
B13,788
B25,093
C17,550
median rank

How well does the free overlay agree with it?

Measured on the 1,965 lemmas that both the validated map and the open course overlay cover, comparing the course's raw first-occurrence guess against the validated level. A far stronger test than the small hand-curated gold sample: it is the whole overlap.

45.6%
exact match (896/1,965)
86.6%
within one level (1,701/1,965)
A sample of the validated overlay
#LemmaCEFR
1вA1
2наA1
3уA1
4іA1
5зA1
6бутиA1
8щоA2
11таA2
20якA2
38підA2
45поA2
47післяA2
61заявитиC1
69повідомитиB1
72радаB2
73іB1
75жB1
78йB1
80приB1
83повідомлятиB1
121зазначитиB2
155представникB2
213верховнийB2
225управлінняB2
252приміткаB2
292главаC1
411чинC1
429розслідуванняC1
622виборчийC1
752законодавствоC1
The validated overlay reconciles
Open / Unvalidated cross-check

The free overlay, now a cross-check on the validated map

The validated map above needs the authors' permission to redistribute. This overlay is the fully open alternative: it labels our lemmas from the learn-ukrainian open course (Krisztian Koos, CC-BY-SA), an expert selection based on Ukraine's 2024 State Standard, no permission needed. A lemma's level here is the first course lesson that introduces it, so treat it as a rough teaching-order signal, not a proficiency label.

Its role now is a cross-check, not a fallback. The validated overlay above measures the two against each other on the full overlap; this card keeps the open overlay's own calibration against the published gold points. Neither the open map alone nor the paper's frozen tables is the finished thing a learner deserves, and closing that gap is why this series exists.

2,326 of the top 12,000 lemmas carry an open CEFR label, A1 to B2

How right is it? Checked against the validated gold standard

Synchak et al. publish concrete CEFR levels for specific lemmas. Where those overlap this overlay, the disagreement is measured directly, so the overlay is calibrated, not just asserted. The sample is small (a few dozen gold points), so read the interval, not the point estimate.

63.2% ± 8.9
exact match (72/114)
86.8%
within one level (99/114)
LemmaValidatedOverlay
бутиA1A1
рікA1A1
вінA1A1
дляA1B1
яA1A1
вониA1A1
такожA1A1
вонаA1A1
часA1A1
матиA1A1
могтиA1A1
людинаA1A1
післяA2A1
новийA1A1
першийA1A2
країнаA1A1
сказатиA1A1
статиA1A2
особаA2A1
черезA1B1
областьA1B1
заявитиC1B1
абоA1A1
щобA2A1
головаA1A1
дваA1A1
містоA1A1
повідомитиB1A2
словоA1A1
деньA1A1
місцеA1A1
роботаA1A1
повідомлятиB1B1
себеA2A2
якщоA1A1
питанняA1A1
великийA1A1
щодоB1B1
рішенняB1B1
компаніяA1A2
районA1A1
ранішеA2A1
отриматиA1A2
частинаA1A1
партіяB1B2
хтоA1A1
працюватиA1A1
результатA1A2
можнаA1A1
дитинаA1A1
головнийA1A2
селоA1A1
складB1A1
державаA2B2
чоловікA1A1
хотітиA1A1
життяA1A2
правоB1B2
сьогодніA1A1
урядB1B2
міжA2A2
вибориB1B1
тисячаA1A1
знатиA1A1
діяB1A1
ситуаціяA1B2
тутA1A1
заразA1A1
центрA1A1
становитиB1B2
допомогаA1A1
триA1A1
березеньA1A1
зробитиA1B1
місцевийA2A2
кількістьB1A2
початокA1A1
другийA1B1
початиA1B1
гривняA1A1
станB1B1
тамA1A1
писатиA1A1
метаB1A2
освітаB1A2
житиA1A1
написатиA1A2
чорнийA1A1
наукаB1A2
білийA1A1
підписатиA2B1
займатисяA2A2
червонийA1A1
стілA1A1
зеленийA1A1
традиціяA1A2
підписA1B1
ліжкоA1A1
синійA1A1
кріслоA1A1
годинникA1B1
помаранчевийA2A1
цікавитисяA2A2
типовийA2A1
стілецьA1A1
мебліA1A2
захоплюватисяA2A2
посудA1B1
дописB1B1
лампаA1A1
шафаA1B1
холодильникA1A2
рушникA1A1
диванA1A1

Only 114 of 200 gold points fall inside the top 12,000 and share the overlay's part of speech; the rest are rarer words, or homographs of a different sense. The misses run both ways: more often the course places a concrete but corpus-rare noun (шафа wardrobe, посуд dishes) higher than the validated profile, which front-loads it to A1 for a beginner; less often it puts an abstract noun (наука science, освіта education) a level lower. Much of that gap is not error but two philosophies: the course follows where a word is taught, the profile makes deliberate beginner-priority overrides.

The join is POS-aware, so 75 homographs (a fish and a quantifier that share a spelling, for instance) were dropped rather than mislabeled. And where the validated profile publishes a level, that level is used and marked validated, so 114 labels here are Synchak's own, not the heuristic's. The agreement above is measured on the heuristic BEFORE those overrides, so the number is not flattered by them.

Do the levels track frequency at all?

Weakly, which is the honest answer. The rank correlation between assigned level and corpus frequency is 0.283 (a CEFR level is not frequency, so a loose fit is expected), but the median frequency rank does rise from A1 to B2, so the ordering is at least directionally sane.

A11,612
A22,526
B13,570
B24,197
median rank
Labels by level
A1639
A2422
B1821
B2443
C11
How much of the top N is labeled
Top 1,000: 522 labeled (52.2%)
Top 2,000: 893 labeled (44.7%)
Top 5,000: 1,588 labeled (31.8%)
Top 12,000: 2,326 labeled (19.4%)
Every label is graded for confidence
validated (Synchak's level) 114frequency-typical 2,184frequency outlier 28
A sample of the overlay
#LemmaCEFR
1вA1
2наA1
3уA1
4іA1
5зA1
6бутиA1
9доB1
10заA2
22відB1
38підA2
45поB1
47післяA2
57особаA2
61заявитиC1
63щобA2
69повідомитиB1
83повідомлятиB1
85себеA2
90щодоB1
123даніB2
284матчB2
330парламентB2
337акціяB2
401порушенняB2
406належатиB2
The overlay reconciles
Author-shared / A1 consensus

The second A1 source: the authors sent the list

Snapshot 2026-07-27

The field census below used to list the "1000 i 1 slovo" A1 lexical minimum (Kseniia Borodin and Oksana Turkevych, School of Ukrainian Language and Culture, Ukrainian Catholic University) as walled: licensed CC-BY-4.0 but published only as a prose PDF an extractor could not reach, so the one independent, empirically grounded A1 guideline for Ukrainian could not be joined and the overlay's rigor model called that missing consensus a seam. This page asked the authors for a machine-readable table, and co-author Kseniia Borodin sent the word list herself. This card is that list, joined onto the frequency list and republished as attributed open data under the guide's own license.

The join is honest about its grain. The authors' unit is the published entry (a single word, an aspect pair, a masculine and feminine pair, sometimes a whole phrase), the list carries no part of speech, and the join is by lemma string at the best rank, coarser than the POS-aware overlays above. Phrase entries are counted, never force-joined, and a beginner word the corpus does not rank stays honestly off-list.

838 of the 929 unique A1 entries join the top 12,000, reaching 833 distinct lemmas

Where the validated map places these words
A1579
A2162
B125
B21
C11
Where the A1 minimum sits on the list
Top 1,000: 307 of its lemmas (30.7%)
Top 2,000: 480 of its lemmas (24%)
Top 5,000: 688 of its lemmas (13.8%)
Top 12,000: 833 of its lemmas (6.9%)

Of the 768 joined lemmas the validated PULS map also covers, 75.4 percent are validated A1 and 96.5 percent sit at A1 or A2. Two independent efforts, a corpus-validated profile and an expert lexical minimum, converge on what beginner Ukrainian is. The handful above A2 is mostly the price of a coarser join: автомат the A1 ticket machine shares a spelling with the C1 rifle, and this list cannot tell them apart without a part of speech.

The teachability trade, on an independent instrument

The 833 matched lemmas cover about 34.7 percent of running text; the 833 highest-frequency lemmas would cover about 57.5 percent. That is the same finding as the validated core above, now measured on a source this page did not build: a pedagogical A1 selection deliberately trades raw coverage for concrete, teachable words, and 72 of its word entries (апельсин orange, виделка fork, бутерброд sandwich) do not rank in this news-heavy top 12,000 at all.

Beginner words the corpus does not rank
автовокзалалергіяалфавітапельсинбананблокнотблузкабутербродвареникивегетаріанськийвечерятивиделка
A sample of the join
#Published entryLemmaPULS
1в, увA1
2нанаA1
4і, й, таіA1
5з, із, зізA1
6бутибутиA1
7ненеA1
1536готельготельA1
1545статьстатьA2
1546дочкадочкаA1
1549готуватиготуватиA1
1566папірпапірA1
1570купувати – купитикупитиA1
11650цукеркацукеркаA1
11758куркакуркаA1
11839кремкремA2
11871фотографуватифотографуватиA1
The join reconciles
Open access / A2 and the thematic axis

The next level up, and the first look at which fields the corpus misses

Snapshot 2026-07-28

The A1 card above closed a seam. Its provenance note also recorded an open one, in plain words: the A2 volume of the same series existed, but no machine-readable A2 list had been shared. That volume has since been published open access under CC-BY-4.0 by Kseniia Borodin and Olesia Lazarenko at the Europa-Universitat Viadrina, so the seam closes by publication rather than by favor. This card is its 1,000-entry register, extracted from the publication's own PDF, joined onto the frequency list the same coarse way, and republished as attributed open data.

Two things become answerable that were not before. Because this is the same authors' next volume, the two lists together measure what one CEFR step actually costs, in frequency ranks rather than in word counts. And because the A2 book also prints its whole vocabulary regrouped under its own 51 topic headings, this is the first source on this page that can say WHICH semantic fields a news-heavy frequency list misses, not merely how many words it misses.

708 of the 1,000 published A2 entries join the top 12,000, reaching 707 distinct lemmas

What one CEFR step costs

Only 45 of the A2 entries are words the same authors' A1 minimum already claimed, and those are almost all verbs collecting their perfective partner. The other 955 (95.5 percent) are new at A2, and they sit measurably deeper in the corpus: median rank 2,902 against 1,643 for the words carried up from A1. A CEFR step is not a bigger pile of the same words. It is a move outward into rarer ones.

Which fields the frequency list misses

The book groups its 1,124 thematic units under 51 topic headings. Joined against the frequency list, 21.6 percent of them do not rank in the top 12,000 at all, and the misses are not evenly spread. The worst-covered fields are the concrete, domestic, physical ones: food, personal hygiene, household technology, and the numerals, where both the collective forms and the ordinals the authors deliberately extended to thirtieth (so a learner could say a date) fall outside the list. A corpus of published Ukrainian news is a poor description of a kitchen.

Їжа69%
Числівники63.2%
Особиста гігієна60%
Технології59.1%
Українська культура58.3%
Способи приготування їжі57.1%
Одяг, взуття, аксесуари56.4%
Продукти, фрукти, овочі53.8%
Topic / Off-list
Where the validated map places these words
A1148
A2313
B1161
B232
C12
A2 words the corpus does not rank
абе́ткаабрико́саквапа́ркАлло́!анана́сапо́строфбанду́рабезалкого́льнийбезглюте́новийбезлакто́знийбіле́тблонди́н / блонди́нка

Of the 656 joined lemmas the validated PULS map also covers, 70.3 percent sit at A1 or A2 and 29.7 percent sit above A2, nearly all at B1. That is a looser agreement than the A1 card found, and it should be: A2 is where two expert efforts start to disagree about the boundary, and the coarse lemma-string join adds noise of its own. The page reports the disagreement rather than tuning it away.

The teachability trade, one level up

The 707 matched lemmas cover about 11.9 percent of running text; the 707 highest-frequency lemmas would cover about 55.1 percent. The trade the A1 card measured widens sharply at A2, and 229 of the word entries do not rank in this news-heavy top 12,000 at all, against 72 at A1. This is the honest ceiling on frequency-first learning: past the beginner core, a frequency list and a syllabus stop describing the same language.

A sample of the join
#Published entryLemmaPULS
9додоA1
21длядляA1
22відвідA1
25свійсвій·
27Бува́й(те)!теA1
35тойтойA1
2771дахдахA2
2814Бе́льгіяБельгія·
2823звук (Р.в. одн. зву́ка)звукA2
2837відві́дувач / відві́дувачкавідвідувачB1
2841зустріча́ти(ся) – зустрі́ти(ся)зустрітиA1
2844підзе́мнийпідземнийB1
11797А́рктикаАрктика·
11811волейбо́лволейболA2
11925симпати́чнийсимпатичнийB1
11969па́скапаскаA2
The join reconciles
Thread 06 / Genre

A frequency list inherits the genre of its corpus

Snapshot 2026-07-28

The card above says 229 word entries of the A2 minimum do not rank in this list's top 12,000, and reads that as the ceiling on learning by frequency. That reading rests on something untested. This list is built on UberText 2.0's COMBINED corpus, and the largest thing in that corpus by a wide margin is news. So an off-list word might be rare in Ukrainian, or it might be ordinary and simply absent from journalism, which does not often discuss forks, pumpkins or neckties. Those two readings give opposite advice to anyone building a word list, and nothing above can tell them apart.

So we built two frequency lists instead of one, from UberText's own fiction and news subcorpora: same collection, same cleaning, same lemmatizer, same sampling shape, both capped to the same token count. Genre is the only thing that differs. The ranks below are internally comparable and deliberately NOT comparable to the main list on this page, which a different lemmatizer produced.

Where the A2 minimum falls in each genre
in both genres 562news only 96fiction only 86in neither 192

News reaches 658 of the entries (70.3 percent). Fiction reaches 648 (69.2 percent). Almost the same amount, and not the same words: together they reach 744 (79.5 percent). A second genre of the same size buys 9.2 points that more of the first genre would not have bought. For a pedagogical target, corpus diversity is worth more than corpus size.

The exchange is close to symmetric, which is what makes it a real result rather than a bigger hammer: fiction reaches 30.9 percent of what news misses, and news reaches 33.3 percent of what fiction misses. Neither genre is better. They are looking at different parts of the language.

Where the two genres disagree most
Способи приготування їжі14.3%85.7%
Сімʼя та родина57.1%100%
Людина: зовнішність, характер, емоції та стани57.7%88.5%
Частини тіла75%100%
Посуд та інше40%60%
Природа: рослини і тварини81.8%100%
Людина: особиста інформація50%66.7%
Комунікативні фрази7.7%23.1%
News / Fiction
Only news ranks these
анке́таАнтаркти́даА́рктикабале́тбанкома́тбезкошто́внийБе́льгіябланкбрита́нець / брита́нкабуря́квиши́ва́нкавідмі́нно
Only fiction ranks these
банду́рабігборода́брова́бу́ряваре́нийвари́ти – звари́тивго́лос, уго́лосвесе́лкави́рібви́шнявідро́

Read the two rows together and the mechanism is plain. Journalism supplies the institutional and the geographic; fiction supplies the domestic, the bodily and the sensory. Neither is a description of the language on its own.

What this measurement is not

These ranks come from our own lemmatizer pass over two subcorpora, not from the lemmatizer behind the main list on this page. They are comparable to each other and to nothing else here, and no figure above mixes them with the main list. The lemmatizer also mis-handles a few high-frequency pronouns, identically in both genres, so it cannot manufacture a difference between them.

The split reconciles
The figures reconcile

Why this is so thin: there is still no open Ukrainian CEFRLex

Thread two now joins a validated profile (by the authors' permission), an author-shared A1 lexical minimum, and an open course overlay. That it took two direct asks to get here is the point: the open, per-lemma, CEFR-graded word list, a standard resource for learning a language, has been built for about a dozen languages and never published as open Ukrainian data. Here is the whole field.

What Ukrainian has
SourceAccessLevels
Ukrainian Vocabulary Profile (PULS)validatedby permissionA1-C1 (to B1)
State Standard 2024 (SLSUFL)documents onlyA1-C2
CEFR & Ukrainian-English Language Portfolios (PCUH)documents onlyA1-B1
learn-ukrainian courseopen, CC-BY-SAA1-B2
1000+1 Words (Borodin, Turkevych)shared by the authorsA1
1000+1 Words A2 (Borodin, Lazarenko)open, CC-BYA2
Built for other languages, not Ukrainian
KELLY · 9 languages, no Ukrainian edition
CEFRLex · several languages, no Ukrainian edition
UniversalCEFR · 13 languages, no Ukrainian edition
Have the missing piece?

This page folds in an open Ukrainian CEFR word list the moment one exists, and that is not hypothetical: the A1 minimum above arrived exactly this way, because its authors answered an email. The schema it needs is small: lemma, part of speech, and CEFR level, as CSV or JSON.

lemmapart of speechCEFR levelGet in touch
Ukrainian Frequency / Derussification

The word older materials quietly got wrong

A frequency list tells you which words to learn. It does not tell you that some of the words you will meet in real Ukrainian text are ones a careful editor would replace: russianisms, surzhyk, and calques that older teaching materials quietly passed on. Thread three flags them on the exact list you study from.

There is no single settled russianism list to join against. What there is, open and maintained, is the Ukrainian idiomatic-style tooling: LanguageTool's Ukrainian replacement tables (the brown-uk project, the same one behind the VESUM morphology). This thread joins that tooling to the frequency list and, for every top-12,000 lemma it flags, shows the native alternative it recommends.

Open / Unvalidated

Read this the way you read thread two's open CEFR overlay: it is an OPEN, joinable signal, but an unvalidated, prescriptive one. It is what a style checker would replace, an editorial opinion, not a settled linguistic verdict. Its core is derussification, but the tooling's scope is broader than strictly-Russian loans, and a few entries (поліцейський, безкоштовний) are live usage debates. Proper-noun and spelling fixes are dropped. Every suggested form is the tooling's own, kept verbatim.

87 of the top 12,000 lemmas carry a flag from the style tooling

The finding: the calque is often the common one

64.9%48 / 74

Of the 74 flagged forms whose native replacement is a single word, 48 (64.9%) have the flagged form at least as common in real Ukrainian text as its replacement: the idiomatic word is rarer, or not in the top 12,000 at all. That is the empirical face of older materials teaching the calque, and the corpus still carrying it.

Where the native form sits (single-word swaps)
The pairs

One row per frequency band, lowest rank first. Each flagged form, the native alternative the tooling recommends, and where that alternative sits on the list. The full list is in the download.

#Flagged formRecommendedNative rank
1,255поліцейськийNOUNполіційнийnot in top 12,000
2,204прийомNOUNприйманняnot in top 12,000
3,086безкоштовнийADJбезплатнийnot in top 12,000
4,188безкоштовноADVбезплатноnot in top 12,000
5,251передвиборнийADJпередвиборчий#4,420
6,676недостовірнийADJневірогіднийnot in top 12,000
7,121місцезнаходженняNOUNмісце#81
8,004кримчанинNOUNкримецьnot in top 12,000
9,167достовірнийADJвірогіднийnot in top 12,000
10,308подаліADVтрохи даліa construction
11,027благополуччяNOUNдобра доляa construction

11 of the flagged adjectives are the active-participle calque (діючий, існуючий): the -ючий forms Russian has and standard Ukrainian rewrites as a relative clause (що діє). The tooling lists that rewrite first.

How firmly the tooling flags it
standard replace 66soft suggestion 21
How much of the top N is flagged
Top 1,000: 0 flagged (0%)
Top 2,000: 2 flagged (0.1%)
Top 5,000: 24 flagged (0.5%)
Top 12,000: 87 flagged (0.7%)
The join reconciles
Ukrainian Frequency / Audio and paradigms

One lemma is many words, and now you can hear it

Every thread so far has worked on the bare lemma, the dictionary headword. But a learner of a heavily inflected language does not meet headwords in text; they meet the forms. рік (year) shows up as року, роки, років, рокам. A verb has dozens of shapes. Thread four attaches those forms, and the sound, across the whole top-12,000 list.

The source is open and machine-readable: the Wiktionary declension and conjugation tables (via the kaikki.org extraction), plus the IPA and the native-speaker pronunciation audio that live in the same entries on Wikimedia Commons. Forms are kept verbatim, stress marks and all; the audio is linked, never re-hosted; and coverage is graded honestly rather than assumed.

Wiktionary + Wikimedia, CC BY-SA

The inflection tax, counted in forms

15×125,380 forms / 8,339 lemmas

The 8,339 lemmas in the top 12,000 that inflect expand to 125,380 distinct wordforms, an average of 15 each. Thread one measured the inflection tax as slow reading coverage; this is the same tax counted in the surface forms you actually have to recognize. A noun runs about ten forms, a verb closer to thirty.

Forms per lemma, by kind of word
verbs30 / 1,910 lem
pronouns23.1 / 35 lem
adjectives14.4 / 1,727 lem
nouns9.2 / 4,601 lem
other6 / 1 lem
numerals5.8 / 21 lem
adverbs2.2 / 44 lem

See a paradigm

A few high-frequency words, declined or conjugated, with IPA and open native audio. Pick a word; the forms are Wiktionary's own.

бути[ˈbute]
infinitiveбу́ти
1sg presentє
2sg presentє
3sg presentє
3pl presentє
past mбув
past fбула́
imperative 2sgбудь
Audio: Wikimedia Commons contributors, CC BY-SA

This is what turns the thread-one Anki seed from a bare headword list into a deck that teaches each word with its forms and its sound.

How much we could attach, of the top 12,000
carry a full paradigm69.5% (8,339)
carry an IPA transcription75% (9,002)
carry open native-speaker audio61.7% (7,398)
The join reconciles
Roadmap

The series roadmap

All five threads are released. Each names the open source it is built on, and every figure on this page re-derives from a committed release. No thread claims a result it cannot show.

01

The lemma frequency list

Released

The top 12000 lemmas of modern Ukrainian, ranked by corpus frequency, each with its part of speech, per-mille frequency, cumulative running-text coverage, and a thousand-rank frequency band. Published as CSV and JSON, plus an Anki-importable deck seed.

Study the words that actually pay off first, in the right order, as lemmas rather than scattered wordforms. See exactly how much reading coverage a given amount of vocabulary buys.

For: learners, teachers, tool builders, researcherslang-ukDownload the release
02

CEFR alignment

Released

A validated per-lemma CEFR overlay: the only corpus-validated CEFR profile for Ukrainian (Synchak et al. 2025 / PULS) joined onto the frequency list, A1 to C1, for 6207 of the top 12000 lemmas, redistributed with the authors' permission as a dated snapshot. The published paper's distribution stays cited beside it, the open course overlay is kept as a measured cross-check, and the finding holds: the validated level tracks frequency only weakly, so a frequency rank is not a proficiency level.

Study the high-frequency words in validated CEFR order, see which half of the list has a validated level and how far it reaches, and see how weakly proficiency level and raw frequency actually line up.

03

The derussification pass

Released

Flag non-idiomatic forms on the list and pair each with the native alternative the open Ukrainian style tooling recommends (LanguageTool's replacement tables, brown-uk, backed by the VESUM morphology). The core is derussification (russianisms, surzhyk, the active-participle calque); the honest scope is broader, and it is fenced as an unvalidated, prescriptive signal, not a settled verdict.

Learn the idiomatic word, not the calqued one that older materials quietly teach, and see how often the calque is actually the more common form in real text.

04

Audio and paradigms

Released

Attach a pronunciation and the inflection paradigm to the top 12000 lemmas, from the open Wiktionary declension and conjugation tables (kaikki.org) and open Wikimedia pronunciation audio, so a lemma is taught as its forms rather than a bare headword. Coverage is graded honestly (it thins with rank), and the inflection multiplier is measured.

Hear the word, see how it changes across the seven cases and the aspect pairs (the part inflection makes hard), and see how many wordforms the words you are learning actually expand to.

05

The vocabulary-size test

Released

A self-scoring yes or no vocabulary-size test drawn from this list: 60 real lemmas and 60 constructed pseudowords, stratified across the 12 frequency bands, scored with the standard false-alarm correction and mapped onto the coverage curve. No Ukrainian equivalent exists off the shelf.

Measure your vocabulary against a real proficiency proxy instead of guessing, and see roughly how much running text that vocabulary actually unlocks.

Download

The open-data release

Everything on this page re-derives from these files. They are versioned and hashed. The list is CC-BY: use it, fork it, build on it.

The lemmas.anki.tsv file imports straight into Anki as a frequency deck seed.

Versionv0.1.0
LicenseCC-BY-4.0
Cite
Lawrence, J. (2026). Ukrainian Frequency: a lemma-based, coverage-graded open frequency list (v0.1.0). jakelawrence.xyz/research/ukrainian-frequency. Derived from UberText 2.0 (Chaplynskyi, 2023).
Verify

Verify it yourself

These checks re-run in your browser against the shipped numbers. If any fails, the release is inconsistent and the page says so.

Methods

Methods and provenance

The frequency signal is the UberText 2.0 lemma frequency dictionary (lang-uk). This release re-publishes derived aggregate counts, which are facts, not the source corpus text. Rows are filtered to Ukrainian-Cyrillic-alphabetic lemmas carrying a linguistic Universal Dependencies part of speech; punctuation, symbols, digit tokens, foreign or unclassified tokens, and single-character proper nouns are dropped.

Coverage is recomputed here from raw occurrence counts, not read from the source file's own frequency columns, whose normalization is undocumented. Frequency bands are transparent thousand-rank bins and are NOT CEFR levels; validated CEFR alignment is thread two, which profiles the one validated resource rather than inventing levels. The corpus is news and reference heavy, so war-era and civic vocabulary rank prominently. That is disclosed, not smoothed away.

Sources
Universal Dependencies (UD) part-of-speech scheme · Universal Dependencies · 2026-07-14
Corpus-Based Vocabulary Profiling for Ukrainian: From Lexical Analysis to CEFR · Synchak, Starko, Burak and Svystun, eLex 2025 · 2026-07-14
The Checklist Method for Testing Vocabulary Size (X_Lex / Y_Lex) · Meara, P. and Buffery, C. · 2026-07-14
dict_uk (Ukrainian hunspell dictionary, via LibreOffice dictionaries) · Andriy Rysin and contributors, brown-uk / dict_uk · 2026-07-14
LanguageTool Ukrainian replacement tables (replace.txt, replace_soft.txt) · brown-uk (Andriy Rysin and contributors), via LanguageTool · 2026-07-22
VESUM: a large morphological dictionary of Ukrainian · brown-uk (Andriy Rysin, Vasyl Starko and contributors) · 2026-07-22
Wiktionary Ukrainian, machine-readable extraction (wiktextract / kaikki.org) · Wiktionary contributors; extraction by Tatu Ylonen (kaikki.org) · 2026-07-22
Wikimedia Commons Ukrainian pronunciation audio · Wikimedia Commons contributors · 2026-07-22
1000 and 1 Words: Lexical Minimum in Ukrainian (A1) · Borodin, K. and Turkevych, O., Slavic Language Education 3 (LANGUAGE-LAB) · 2026-07-27
1000 and 1 Words: Lexical Minimum in Ukrainian (A2) · Borodin, K. and Lazarenko, O., Europa-Universitat Viadrina Frankfurt (Oder) · 2026-07-28
UberText 2.0 subcorpora: fiction and news, read separately · lang-uk (Chaplynskyi, D.) · 2026-07-28
Lexical Minimums for Levels A1 and A2 in Learning Ukrainian: Methodology and Practical Challenges · Borodin, K. and Lazarenko, O., in New Directions in Slavic Studies (T. Portnova, I. A. Gonzalez-Hidalgo, M. P. Pena-Molina, eds.), Editorial Universidad de Granada, ISBN 978-84-338-7814-4, pp. 41-52 · 2026-07-28
State Language Standard: Ukrainian as a Foreign Language, Levels A1 to C2 (SLSUFL) · Ministry of Education and Science of Ukraine · 2026-07-22
Teacher's Guide: Using the CEFR and Ukrainian-English Language Portfolios · Prokopchuk, N. and Huzar, O., Prairie Centre for the Study of Ukrainian Heritage (PCUH), University of Saskatchewan · 2026-07-27
KELLY project: keywords for language learning (CEFR-graded lists) · Sprakbanken, University of Gothenburg · 2026-07-22
CEFRLex: CEFR-graded receptive lexical resources · CENTAL, UCLouvain · 2026-07-22
Related

Part of a standing effort to close the digital-tooling gap for the Ukrainian language.

Need something like this built?

I design and ship AI tools, full-stack apps, and data pipelines — end to end, to production. Tell me the problem in a sentence; I'll give you an honest read on fit within a day.

Work with me →