← All posts

The last 10%: what the machine can't time for you

AI transcription gets subtitles 90% right. The final 10% — timing, line breaks, taste — is where captions get good. A look at the human part of subtitling for video editors and content creators.

Automatic transcription is genuinely a marvel now. Feed a video to a modern speech model and it comes back with the words, roughly timed, in seconds. Ten years ago that was a research paper with a demo that mostly worked; today it’s a checkbox in a dozen apps. So it’s fair to ask the uncomfortable question: if the machine does the captions, what’s actually left for a human to do?

About 10%. And it turns out that 10% is the entire difference between “captioned” and “good.”

What the machine genuinely nails

Let’s give the robots their due, because SubSlap uses the same 90% as a starting point — we’re not pretending to have reinvented transcription.

  • The words. On clean audio, modern models transcribe at near-human accuracy, punctuation and all.
  • Rough timing. Start and end times, close enough to follow along.
  • A first-pass split. Lines chopped into readable-ish lengths.

For a lot of casual content, that’s the whole job. Publish and move on. The trouble starts when “readable-ish” isn’t good enough.

What it structurally can’t do

Here’s the part no bigger model fixes, because it isn’t a knowledge problem — it’s a taste-and-context problem:

  • Break on the breath, not the count. Machines split by character length; people split by thought. “the best part about this / whole thing” reads wrong even though every single word is correct, because it severs one idea across two lines.
  • Hold for effect. A speaker pauses for a beat before the punchline. The model doesn’t know the punchline is coming — you do, because you watched the video and felt it land.
  • Keep a voice. Should “gonna” stay “gonna,” or become “going to”? Is this a breezy caption or a formal subtitle? That’s an editorial call about tone, and tone isn’t in the audio.
  • Respect the cut. A caption that spills one word past a hard cut looks like a mistake. The model can’t see your edit; it only hears the words.

Bong Joon-ho, accepting his Oscar for Parasite, called subtitles “the one-inch-tall barrier.” Cross that barrier well and a viewer forgets they’re reading at all; cross it badly and they’re pulled out of the film every few seconds. That crossing is a frame-by-frame craft decision — exactly the part a model can’t feel, because it has never sat in a dark room watching.

Why this matters more than it used to

Two things changed. First, most social video is now watched on mute, so the captions aren’t an accessibility nicety — they’re the performance itself. Second, the bar for “looks professional” quietly rose, because audiences see clean captions constantly. A line that breaks mid-thought doesn’t crash anything. It just makes the video feel slightly amateur, and the viewer usually can’t tell you why. That “can’t tell you why” is the dangerous part — it costs you trust without ever surfacing as a complaint.

How to actually own the 10%

The move is to treat that last 10% as its own step, not something you eyeball. In practice that means:

  1. Read the captions like a viewer, not the writer. Where does your eye stall?
  2. Fix the breaks first — most “off” feelings are line breaks, not timing.
  3. Then tune the timing — nudge cues onto beats and off of cuts.
  4. Then decide voice — contractions, punctuation, tone — consistently across the whole video.

That’s the entire reason Control Freak Mode exists in SubSlap: frame-level timing and split control, so you fix the model’s guesses instead of shipping them. The machine gets it 90% right. The last 10% is where captions get good — and it’s the part that’s still, gloriously, yours.