Stress on 12,000 Ukrainian words was a lookup, except for 246 arguments
An open stress dictionary settled 9,674 of the 9,920 in-scope words on my Ukrainian frequency list in one pass. The other 246 needed arguments: names taken out of scope, doublets sent to a printed dictionary, homographs decided by counting 2.3 billion words of text, and a native-speaker audit written down before anyone has run it.
Ukrainian print never marks stress, and the stress moves: за́мок is a castle and замо́к is a lock. A learner reading a frequency list has to be told where it falls, so the list needed a stress column. A friend who works on Ukrainian language tooling pointed me at ukrainian-word-stress, lang-uk's open stress dictionary, and I ran it over the 12,000 lemmas on my frequency list. Of the 9,920 in-scope words with more than one syllable, 9,674 came out with one stress every source agreed on. This post is about the other 246, which is where all of the work was, and about the check I wrote down before anyone has run it.
The lookup was the easy part
The dictionary carried 9,487 rows on its own. en.wiktionary and uk.wiktionary filled 293 more, a handful of explicit rules covered compounds and abbreviations, and 732 one-syllable words need no mark at all.
One change mattered more than any source. Called the documented way, the library runs a part-of-speech tagger over its input, and a tagger handed a bare headword with no sentence around it is guessing. The list already knows each word's part of speech, so I read the dictionary's trie directly and let the list's own tag pick the reading. Only 16 rows still rest on the tagger, where the dictionary itself is ambiguous for that part of speech, and they are labeled that way.
What was left fell into kinds, and each kind needed a different argument. The bar below is the whole list at one scale, with the arguments zoomed in underneath, because at one scale they are a sliver you can barely press.
All 12,000 words on the list
The 246 that needed an argument, zoomed in
Two words share a spelling and differ by stress. Their inflected forms that are spelled differently were counted across 2.3 billion words of UberText, and the commoner reading became the primary.
пора → пора́
Names are not a pronunciation rule
1,348 rows are proper nouns: cities, rivers, surnames, first names. I took them out of scope. They stay on their rows, with whatever stress the sources give, but they are not counted in any rate. A surname ending in -енко is stressed the way its family says it, and the rule that predicts the common pattern is wrong often enough that a column claiming it would be claiming something it cannot check. Most learner lists skip names anyway. I am asking the friend who sent me the library what he would do with them, and his answer can still change this.
Two stresses, and who gets to break the tie
88 words have two stresses that are both accepted (та́кож and тако́ж). Those rows carry both, one listed first.
49 are different. The dictionary lists two stresses, no other open source covers the word, and a dictionary listing two is not saying which the norm prefers. The tie goes to a printed orthoepic dictionary, Погрібний's, cited row by row with a page number, and until I have looked each one up the row is labeled a doublet rather than quietly resolved. An online dictionary could have settled them faster. Its license is closed, and a column under CC-BY-SA should not lean on a source nobody can redistribute.
Counting the homographs
The hard rows are the 88 homographs: two words with one spelling and different stress, like прохо́дити (to pass) and проходи́ти (to walk a distance). A headword on a list has no sentence around it, so meaning cannot pick. Usage can.
Most of these pairs have inflected forms that are spelled differently between the readings, and a form spelled one way can only belong to one of them. I counted those forms across all five subcorpora of UberText 2.0, about 2.3 billion words of news, Wikipedia, fiction, court rulings and social media, and the commoner reading became the row's primary stress when it had at least 20 hits and led by at least half again.
The trap is that a form can belong to a third word. заходи is the imperative of заходи́ти and also the plural of за́хід, so counting it would credit the wrong verb. Every candidate form went through the site's own lemmatizer, and any form it read as another word was struck: 79 of them.
32 homographs could be decided by spelling and 31 were; відрізати came out 1,423 to 1,407, too close to call. Four came out against en.wiktionary's first-listed entry, which is what the column would otherwise have shown: сходи́ти over схо́дити, 653 to 225; виріза́ти over ви́різати, 1,307 to 787; зго́дитися over згоди́тися; тверди́ти over тве́рдити. The other 56 cannot be counted at all, because every form of о́рган is spelled exactly like every form of орга́н. Those rows keep both stresses, list en.wiktionary's first entry first, and the note on the row says that is all the choice means.
Rebuilding it from the raw dumps
The column started as a handover of already extracted data, so the extractor had never been run end to end in this repo. A separate session ran it from the raw dumps and found five bugs in the ported stages, including a uk.wiktionary reader that matched a page layout the dump does not use and so read nothing.
With those fixed, every in-scope primary stress re-derives except one compound, півде́нно-за́хідний, whose notation depends on which side the tagger lands. 201 rows keep their stress under a more accurate label for where it came from. The rebuild's check now fails on a changed stress and only reports a changed label, which is the line between an error and bookkeeping.
Agreement is not correctness
Every number so far says the sources agreed. None of them says a native speaker would. So before any answer exists, the check is written down and the sample is drawn: 348 rows, seeded so the draw can be repeated, 200 of the 9,674 confident rows at random, 60 from the 137 free variants and doublets, and every one of the 88 homographs. Each row goes to a paid native-speaker reviewer through the site's validation desk, one verdict per row.
The confident tier counts as validated when at least 150 rows have been judged and the lower bound of the 95 percent Wilson interval on its accuracy is at least 0.95. If it falls short, the post that reports it will say so. A wrong verdict flags the row for the ledger of manual decisions; it never edits the column by itself.
When I would not do this
If you are learning Ukrainian and want stress on the words in front of you, a stressed dictionary or a stressed reader does that today, and none of this is necessary. The work above only pays when the column is something other people reuse: then every guess becomes theirs, and a label that says how a stress was settled is worth more than one more filled cell. And if you cannot name the source that breaks a tie, the honest move is to leave both stresses on the row.
Where it lives
The column is on the frequency hub under CC-BY-SA, archived on Zenodo with its own DOI (10.5281/zenodo.23092388), and installable in Python as ukvocab 0.4.0. Two errors turned up in the dictionary itself, його stored as йо́го and підозрюватися as підозрюва́тися, and I have written them up to send upstream along with the compounds whose main and secondary stress it lists as alternatives. The last post in this series, about a number that had to come down, was the list refusing itself. This one is the list admitting what it does not know yet.
Get the next one
An occasional note when something genuinely new ships here — essays, free tools, projects. No schedule, no filler, easy out.
Need something like this built?
I build the AI tools and automations a team adopts, and the full-stack apps and data pipelines under them, and ship them to production. Tell me the problem in a sentence and I'll give you an honest read on fit within a day.
Work with me →