2026-09-01 · 8 min read
How we score a machine translation against a human subtitler
Every prompt change here is measured against human-translated subtitles before it ships. How the reference is aligned, what chrF++ does and does not tell you, how big the noise floor is, and why the numbers cannot be compared across languages.
Without a score there is no way to tell an improvement from a regression except reading a language you may not speak. So every change to this pipeline runs against a gold set: an English episode paired with a professional human translation of the same episode, scored cue by cue. The scores are published on the languages page. This is what is behind them.
Aligning the reference
Human subtitlers merge and split lines freely, so the reference never has the same cues as the source. The aligner pairs them by timecode overlap, and scores units rather than cues: where the human merged two source cues into one, the pipeline's two cues are concatenated and scored as one. Scoring per cue would penalise the translation for the reference's line breaking.
Timecode overlap alone is not enough. Two episodes of the same show still align 47% of their cues against each other, because dialogue is dense everywhere. The discriminator is digits: numbers survive translation, so the right episode agrees on 12 of 12 and the wrong one on 0 of 3. A pairing with low digit agreement is refused and nothing is written.
The reference that drifts
A reference timed for 25 fps against a 23.976 fps master runs 4.3% long. The opening scene aligns and the end of a two-hour film is five minutes out; no single offset can fix it. The aligner searches the ratios between real frame rates and stretches before it shifts, falling back to no stretch whenever a plain offset explains things equally well. On the day it landed it rescued 13 references that had been quarantined as unusable and took Italian from four scorable episodes to ten.
What chrF++ measures
chrF++ counts overlapping character n-grams (1 to 6) and word n-grams (1 to 2) between the hypothesis and the reference. It rewards a translation that shares wording with the human one and is tolerant of small morphological differences, which is why it is the standard for morphologically rich languages. It does not read for meaning: two valid translations that say the same thing in different words score poorly against each other.
The noise floor
The model used here runs at temperature 1, so two identical runs do not give the same score. Three identical Bulgarian runs on four episodes scored 48.92, 48.86 and 49.04: a standard deviation of 0.09. Doubling the gold set from two episodes to four did not reduce that, and the reasoning that said it would was wrong. chrF++ is a corpus metric, so a noisy episode is added in, not averaged away, and the fourth episode was both the hardest and the least stable.
The practical rule is that a change is real when the mean of two runs moves by more than about 0.13. A single run that moves 0.2 is not a finding. That discipline retracted one earlier conclusion: a mined glossary that appeared to gain 0.19 on one run was withdrawn when a second run came in at 0.09, and it does not ship.
Why the numbers cannot be compared across languages
- Morphology. Word n-grams punish a heavily inflected language like Russian for near misses that an analytic language never pays for. Russian's 39.72 and Italian's 50.70 say nothing about which translation is better.
- Writing system. Character overlap between two valid Spanish sentences is high because Spanish has 26 letters. Between two valid Japanese sentences it is low because Japanese draws on thousands. Japanese and Chinese are also scored with chrF rather than chrF++, because they have no spaces and the word-level component only adds a penalty: 4.7 points of it for Japanese, purely as a tokenisation artifact.
- Reference style. One human translator spells numbers out and converts units; Spanish subtitlers write setenta y cinco dólares for 75 bucks and 5 kilos for 11 pounds. That translator's choices are the reference, and the digit check that works for Bulgarian rejects correct Spanish pairs.
- Alignment coverage. German kept 798 units where Italian kept 1220. Fewer units is a different sample, not a lower score.
The reference that was thrown out
An Arabic reference scored 26.99, and the number is published nowhere as an Arabic baseline because it measures the file, not the translation. That file stores right-to-left text in visual order, so 55% of its units put the sentence mark at the start of the string against 0% of ours, and every such unit mismatches at both ends. A guard now refuses any reference whose units are more than a quarter start-punctuated. It would have caught this before the run cost a dollar.
What a score is for
A regression baseline for its own language. Each of the eight measured languages has one number that any change to the prompt, the checks or the repair loop has to beat or explain. That is the whole use. The languages page shows them unranked, with the amount of human reference behind each, because ranking sixteen languages on a metric that cannot rank them would be a claim the measurements themselves disown.
ProvenSubs translates subtitles and checks every cue before you see it.
Translate an episode free →