Translate subtitles
into 36 languages
Every language runs the same pipeline and the same thirteen checks. What differs is the profile: the script, the reading speed, the quotation marks, the sound-label table and the guard against an untranslated line. Each page below says exactly what its profile does, and whether the language has a human-scored baseline yet.
Translate subtitles to Bulgarian
The language this pipeline was built for: a Russianism wordlist, ти/вие held consistent across a season, „…“ quotes, and sound labels from a table.
Translate subtitles to Spanish
Neutral Latin American Spanish with ¿ and ¡ understood, a 75-word stopword guard against English slipping through in a shared alphabet, and repairs for the length Spanish adds.
Translate subtitles to Portuguese
Brazilian Portuguese, with the profile told never to calque from Spanish, a 49-word stopword guard, and sound labels as short verb phrases.
Translate subtitles to French
Idiomatic French with « … » quotes, a 57-word stopword guard, and labels like [LA PORTE SE FERME] from a table rather than the model.
Translate subtitles to German
The densest of the measured languages: German says the same thing in more characters, so it gets the most reading-speed repairs and the ceiling is deliberately not relaxed.
Translate subtitles to Italian
Idiomatic Italian with « … » quotes, the largest human reference set of any language here, and labels like [SQUILLA IL TELEFONO] from a table.
Translate subtitles to Polish
Polish with „…” quotes and sound labels as verbal nouns, [ŚMIECH] rather than a conjugated clause, which is how Polish subtitling writes them and sidesteps aspect and agreement entirely.
Translate subtitles to Russian
Russian with «…» quotes and noun-form sound labels, [СМЕХ], [ВЫСТРЕЛ]. Beta: the metric punishes an inflected language for near misses, and the profile is smaller than the tuned six.
Translate subtitles to Japanese
Japanese measured in columns, not characters: 32 columns is the conventional 16 full-width glyphs, reading speed is 7 per second, and lines break between characters with kinsoku rules.
Translate subtitles to Chinese
Simplified Chinese, Mainland usage, full-width punctuation, 32 columns of double-width glyphs, and a length check that knows Chinese is shorter than English.
Translate subtitles to Korean
Korean is spaced but double-width, so it wraps at word boundaries measured in columns. Politeness level is inferred per speaker and held.
Translate subtitles to Arabic
Modern Standard Arabic in logical order, with ؟ and ، as sentence punctuation, Western digits so numbers can be checked, and sound labels as verbal nouns after a bug that settled the design.
Translate subtitles to Hebrew
Contemporary Israeli Hebrew in the spoken register, masculine and feminine agreement inferred from context, logical-order text the player renders right-to-left.
Translate subtitles to Turkish
Turkish shares the Latin alphabet with English, so a 65-word stopword list is the guard against an English line shipping unnoticed. sen and siz are held consistent per speaker.
Translate subtitles to Thai
Thai has no sentence punctuation and no spaces between words, so the dangling-sentence check is off and lines wrap between characters. Politeness particles are matched to each speaker.
Translate subtitles to Hindi
Spoken Hindustani rather than Sanskritised Hindi, sentences ended with a danda, and the virama counted as zero width so conjuncts measure correctly.
Translate subtitles to Dutch
Dutch as a Netherlands subtitler writes it, je/u held consistent per speaker, a 70-word stopword guard against English in a shared alphabet, and labels as bare verbs: [LACHT], [ZUCHT].
Translate subtitles to Swedish
Swedish for an audience that watches everything subtitled. Danish and Norwegian drift is blocked where the spelling differs, and labels are bare verbs: [SKRATTAR], [SUCKAR].
Translate subtitles to Norwegian
Norwegian Bokmål, the standard subtitles use. Swedish drift is blocked where the spelling differs; Danish is too close to Bokmål to block safely, so it is not. Labels as bare verbs: [LER], [SUKKER].
Translate subtitles to Danish
Danish with Swedish drift blocked where the spelling differs, »…« or ”…” quotes used consistently, and labels as bare verbs: [GRINER], [SUKKER].
Translate subtitles to Finnish
Finnish is agglutinative, so one word does the work of three or four English ones and the length checks are cut for it. Spoken register where the source is spoken; labels conjugated for singular and plural.
Translate subtitles to Greek
Modern Greek in its own script, with the Greek question mark (;) understood so a correct question is never flagged as a cut-off sentence, and «…» quotes.
Translate subtitles to Czech
Czech with obecná čeština where the source is colloquial, ty/vy held per speaker, Slovak blocked where a word never occurs in Czech, and labels as nouns: [SMÍCH], [POVZDECH].
Translate subtitles to Romanian
Romanian for a country that subtitles everything: comma-below diacritics (ș, ț) rather than the cedilla forms, tu/dumneavoastră held per speaker, and labels like [UȘA SE ÎNCHIDE] from a table.
Translate subtitles to Hungarian
Hungarian is agglutinative and related to nothing around it, so a model has nowhere to drift to; what goes wrong is length and register. te/ön/maga held per speaker, labels conjugated for singular and plural.
Translate subtitles to Ukrainian
Ukrainian with two guards against Russian: four letters Russian has and Ukrainian does not are flagged on sight, and a blocklist catches Russian words spelled in letters Ukrainian shares. Labels as nouns: [СМІХ], [ПОСТРІЛ].
Translate subtitles to Indonesian
Indonesian with aku/saya and kamu/Anda matched to the speakers and held, a stopword guard against English in a shared alphabet, and sound labels from a table: [TERTAWA], [PINTU TERTUTUP].
Translate subtitles to Vietnamese
Vietnamese, where the hard part is that there is no neutral you: the pronoun pair a translator picks states the age and relationship of the speakers. Labels from a table: [CƯỜI], [THỞ DÀI].
Translate subtitles to Slovak
Slovak with Czech drift blocked where the spelling differs (děkuji, ano, jsem, není), ty/vy held per speaker, and labels as nouns: [SMIECH], [POVZDYCH].
Translate subtitles to Croatian
Standard Croatian, ijekavian, with ti/Vi held per speaker and labels as nouns: [SMIJEH], [UZDAH], [ZATVARANJE VRATA]. No blocklist, because Serbian and Bosnian are too close for one to be safe.
Translate subtitles to Slovenian
Slovenian with the dual named in the instruction, because a model that writes plural for two people is wrong in a way no check can see. Labels as nouns: [SMEH], [VZDIH].
Translate subtitles to Catalan
Catalan with Spanish drift blocked where a word has no Catalan homograph, tu/vostè held per speaker, «…» quotes, and labels as verbs: [RIU], [SOSPIRA].
Translate subtitles to Persian
Persian (Farsi) in logical order with ؟ and ، understood and Western digits, plus a guard Persian needs and Arabic does not: the Arabic letters ك and ي, which a model writes for the Persian ک and ی, are flagged.
Translate subtitles to Estonian
Estonian with agglutinative length ratios as for Finnish, sina/teie held per speaker, and labels as verbs: [NAERAB], [OHKAB], [UKS SULGUB].
Translate subtitles to Latvian
Latvian for a country that subtitles everything, tu/jūs held per speaker, and labels as verbs: [SMEJAS], [NOPŪŠAS], [AIZVERAS DURVIS].
Translate subtitles to Lithuanian
Lithuanian, heavily inflected, with tu/jūs held per speaker and labels as verbs: [JUOKIASI], [ATSIDŪSTA], [UŽSIDARO DURYS].