2026-09-01 · 5 min read
Why sound labels come from a table, not from the model
[LAUGHS], [DOOR CLOSES], [PHONE RINGS]. Subtitles for the deaf and hard of hearing are full of them, and a language model gets them subtly wrong in ways no reader can catch. One table per language, and the Arabic bug that settled the design.
An SDH subtitle file describes sound as well as speech. In a 600-cue episode, twenty or thirty cues are labels rather than dialogue: who laughs, what closes, what rings. They are short, formulaic and completely conventional, and they are the cues a language model is most likely to render as confident nonsense.
What goes wrong
The model is asked to translate a sentence and a label is not a sentence. It has no subject, no tense and no context beyond its brackets. Asked to translate [LAUGHS] into Arabic on its own, a model produced هي يضحك, which is she followed by a masculine verb: she he-laughs. It was grammatical nowhere. It shipped, because nothing in a character-set check or a length check could see it.
Slavic languages have the opposite problem. A verb needs an aspect and a subject that agrees in gender and number, so [SHE LAUGHS] rendered as a full clause is a small grammar exam on every label, and a subtly wrong table emits confidently broken text with no model in the loop to catch it. That reasoning kept Polish and Russian without label tables for months.
Name the sound instead
Adding Arabic, Japanese, Korean and Chinese showed the way round it. Those languages bracket a noun, not a sentence: the action is named and the subject is dropped. A noun has no aspect to get wrong and no case to agree. And it turns out that is what Slavic subtitling does anyway: [ŚMIECH], not [ONA SIĘ ŚMIEJE]. Both Polish and Russian now have tables built on that convention.
The same label, from each language's table:
- Bulgarian [СЕ СМЕЕ] · Russian [СМЕХ] · Polish [ŚMIECH]
- Spanish [RÍE] · Portuguese [RI] · French [RIT] · Italian [RIDE] · German [LACHT]
- Japanese [笑う] · Korean [웃음] · Chinese [笑]
- Arabic [يضحك] · Hebrew [צחוק] · Hindi [हँसी] · Thai [เสียงหัวเราะ] · Turkish [GÜLER]
How the table is applied
Labels are resolved before the text ever reaches the model. A resolved label is wrapped so the model reproduces it byte for byte, and the surrounding dialogue is translated normally. A cue that is nothing but a label never costs an API call at all.
A label the table does not know is left alone and reported in the quality report, never approximated. That is a deliberate choice: the tables grow from real files, and an unresolved label is a request to add one. Guessing would be faster and would put the Arabic bug back.
Capitalisation is part of the label
Sources vary: [LAUGHS], (laughs), [Laughs]. The table keeps the bracket style and the capitalisation of the source, so a file that arrived in ALL CAPS leaves in ALL CAPS and a file that used parentheses keeps them. It is the sort of detail that a viewer never notices when it is right and notices immediately when it is not.
Every language on the languages page shows three of its own label renderings, straight from its table.
ProvenSubs translates subtitles and checks every cue before you see it.
Translate an episode free →