The last 10%: what the machine can't time for you
AI transcription gets subtitles 90% right. The final 10% — timing, line breaks, taste — is where captions get good. A look at the human part of subtitling for video editors and content creators.
Automatic transcription is genuinely a marvel now. Feed a video to a modern speech model and it comes back with the words, roughly timed, in seconds. Ten years ago that was a research paper with a demo that mostly worked; today it’s a checkbox in a dozen apps. So it’s fair to ask the uncomfortable question: if the machine does the captions, what’s actually left for a human to do?
About 10%. And it turns out that 10% is the entire difference between “captioned” and “good.”
What the machine genuinely nails
Let’s give the robots their due, because SubSlap uses the same 90% as a starting point — we’re not pretending to have reinvented transcription.
- The words. On clean audio, modern models transcribe at near-human accuracy, punctuation and all.
- Rough timing. Start and end times, close enough to follow along.
- A first-pass split. Lines chopped into readable-ish lengths.
For a lot of casual content, that’s the whole job. Publish and move on. The trouble starts when “readable-ish” isn’t good enough.
What it structurally can’t do
Here’s the part no bigger model fixes, because it isn’t a knowledge problem — it’s a taste-and-context problem:
- Break on the breath, not the count. Machines split by character length; people split by thought. “the best part about this / whole thing” reads wrong even though every single word is correct, because it severs one idea across two lines.
- Hold for effect. A speaker pauses for a beat before the punchline. The model doesn’t know the punchline is coming — you do, because you watched the video and felt it land.
- Keep a voice. Should “gonna” stay “gonna,” or become “going to”? Is this a breezy caption or a formal subtitle? That’s an editorial call about tone, and tone isn’t in the audio.
- Respect the cut. A caption that spills one word past a hard cut looks like a mistake. The model can’t see your edit; it only hears the words.
Bong Joon-ho, accepting his Oscar for Parasite, called subtitles “the one-inch-tall barrier.” Cross that barrier well and a viewer forgets they’re reading at all; cross it badly and they’re pulled out of the film every few seconds. That crossing is a frame-by-frame craft decision — exactly the part a model can’t feel, because it has never sat in a dark room watching.
Why this matters more than it used to
Two things changed. First, most social video is now watched on mute, so the captions aren’t an accessibility nicety — they’re the performance itself. Second, the bar for “looks professional” quietly rose, because audiences see clean captions constantly. A line that breaks mid-thought doesn’t crash anything. It just makes the video feel slightly amateur, and the viewer usually can’t tell you why. That “can’t tell you why” is the dangerous part — it costs you trust without ever surfacing as a complaint.
How to actually own the 10%
The move is to treat that last 10% as its own step, not something you eyeball. In practice that means:
- Read the captions like a viewer, not the writer. Where does your eye stall?
- Fix the breaks first — most “off” feelings are line breaks, not timing.
- Then tune the timing — nudge cues onto beats and off of cuts.
- Then decide voice — contractions, punctuation, tone — consistently across the whole video.
That’s the entire reason Control Freak Mode exists in SubSlap: frame-level timing and split control, so you fix the model’s guesses instead of shipping them. The machine gets it 90% right. The last 10% is where captions get good — and it’s the part that’s still, gloriously, yours.