2026-09-01 · 6 min read
Translating subtitles into Bulgarian: what the alphabet does not tell you
Bulgarian was this pipeline's first language and the source of its first regression test. Five things that go wrong when English subtitles are translated into Bulgarian, each with the check that now catches it.
Bulgarian ships in Cyrillic, which makes one failure trivially visible: an English line left in a Bulgarian file stands out on sight. That is the only easy part. Everything else about translating subtitles into Bulgarian is a way for output to look right and be wrong, and each of these shipped at least once before it became a check.
1. Russian words spelled entirely in Bulgarian letters
медленно, глубоко, сейчас. Every letter in those words exists in the Bulgarian alphabet. A character-set test passes them cleanly, and a large model trained on far more Russian than Bulgarian reaches for them under pressure. The only defence is a wordlist of Russianisms, grown from real files. The Bulgarian profile carries 47 of them plus 63 Bulgarian function words the translation is required to contain, so a cue that is Cyrillic but not Bulgarian is flagged rather than shipped.
2. ти and вие
English has one you. Bulgarian has two, and the choice is a statement about the relationship between two characters that has to stay consistent for a whole season. The model is told to infer it from surrounding cues and hold it for a given pair of speakers, and cues travel in batches of 32 so it can see the conversation, not one line at a time. A planned dossier of who addresses whom was measured against real output and found unnecessary: across an episode, zero cues mixed forms of address.
3. Quotation marks
Bulgarian uses „…“. A translator that emits "…" is not wrong exactly, but it reads as machine output to a Bulgarian viewer in the first minute. The style instruction names the marks explicitly, and the dangling-sentence check counts the closing “ as terminal punctuation so a correctly quoted line is not flagged as unfinished.
4. Sound labels
[LAUGHS], [DOOR CLOSES], [PHONE RINGS]. Bulgarian subtitling renders these as short verb phrases, [СЕ СМЕЕ], [ВРАТАТА СЕ ЗАТВАРЯ], [ТЕЛЕФОНЪТ ЗВЪНИ], and the model does not get to write them. They resolve from a fixed table before the text is sent, so a whole class of confident nonsense never has a chance to appear. A label the table does not know is reported, never guessed at; that is a design rule across every language here.
5. The file you were given is probably cp1251
Most Bulgarian subtitle files in circulation are Windows-1251, not UTF-8, and many exist only as frame-based MicroDVD .sub files. Read as UTF-8 they do not fail, they turn into replacement characters. The reader here scores candidate encodings on how plausible the decoded text looks, and reads the frame rate a .sub file declares in its first line rather than guessing one, because a wrong frame rate does not shift the timing, it stretches it.
How good is it, measured
Against a human Bulgarian translation of four episodes of a US sitcom, the pipeline scores 48.94 chrF++, the mean of three identical runs whose spread was 0.18. That number is a baseline for Bulgarian and nothing else: it is what any change to the prompt has to beat, and it is not comparable to the figure for another language. Cost to run was about $0.44 per episode in API calls, which is why an episode sells for $1.25.
Bulgarian is the language this was built for, and the page for it lists exactly what the profile checks. The first episode is free.
ProvenSubs translates subtitles and checks every cue before you see it.
Translate an episode free →