The Caption Generation

Captions were won by deaf audiences after decades of pressure. Now the hearing majority has discovered them — and a feature built for one population is quietly being rebuilt for another.

A friend of mine watches everything with the subtitles on. Prestige dramas, sitcoms she has seen four times, cooking videos where the only dialogue is “and now we wait.” She is not hard of hearing. She just finds it easier — the dialogue in modern shows is mixed low, the actors mumble artfully, and her apartment shares a wall with someone who owns a subwoofer. She is, by now, entirely typical. In an AP-NORC poll from August 2025, a third of American adults said they always or often watch TV and movies with subtitles, and the habit skews sharply young. A YouGov survey two years earlier put the number at 63 percent among adults under 30.

This is usually told as a happy story, and it is half of one. Captions exist at scale because deaf and hard-of-hearing audiences spent decades forcing them into existence — lobbying broadcasters, suing streaming services, pushing laws through Congress while television spent its first thirty years treating them as invisible. What those activists won was then handed to everyone, and everyone turned out to want it. The curb-cut effect, in the classic formulation: build the ramp for wheelchair users and watch the strollers, suitcases and delivery carts roll through. It is one of the few genuinely comforting stories in accessibility. But the caption version has a complication the curb-cut version does not, and the complication is the interesting part.

A curb cut serves everyone identically. Concrete does not care who rolls over it. Captions are not concrete. They are text produced by someone, or increasingly something, under time and cost pressure, and the quality of that text is a dial that can be turned. When captions were a legal obligation serving a small, politically organized audience, the dial was set by that audience’s needs — because they were the only ones reading. Now that the majority of the people reading captions can also hear the audio, the dial is being reset by a different question: what is good enough for a hearing viewer who is half-watching in a loud room?

The answer, it turns out, is: quite bad is good enough. If you can hear the show, captions are a supplement. A dropped word, a mangled name, a missing speaker label, the phrase “[inaudible]” where a plot point used to be — these are mild annoyances, quickly patched over by the ear. The hearing reader is proofreading against the audio without noticing. For a deaf viewer there is nothing to proofread against. The captions are not a supplement to the program; they are the program. The same failure that costs one audience a shrug costs the other the scene.

You can watch this optimization happen in the product itself. Auto-generated captions, the kind speech recognition produces for free at upload time, are built and tuned for the volume use case: billions of videos, watched mostly by people who can hear, in conditions where “close enough” closes the ticket. They are genuinely useful — I use them constantly — and they are also measurably worse for some people than for others. A 2020 study in PNAS by Koenecke and colleagues tested five major commercial speech recognition systems and found an average word error rate of 0.35 for Black speakers against 0.19 for white speakers — a gap traced to the acoustic models, which is to say to whose voices the machines were taught on. That study measured conversational speech rather than broadcast captioning, but the mechanism is the point: systems trained on the majority fail the minority in proportion to their distance from the training data. Scale does not average this out. Scale encodes it.

The regulatory framework, meanwhile, has not caught up with the audience shift. The FCC’s caption quality rules, in force since 2014, require that television captions be accurate, synchronous, complete and properly placed — but there is no numeric accuracy threshold, despite the “99 percent accurate” figure that circulates as if it were law. The rules were written for a world where captions were a compliance deliverable: a box to check, audited rarely, enforced mostly by complaint. And large parts of the caption universe — social video, creator content, much of what young people actually watch — sit outside that regime entirely, governed by whatever the platform’s free auto-captioner produces and whether the uploader bothered to review it. Most do not bother, because their audience, statistically, can hear.

This is the general fate of accessibility features that go mainstream, and it cuts in both directions at once, which is what makes it hard to complain about cleanly. Mass adoption is a genuine victory. It normalizes the feature, kills the stigma of asking for it, funds its spread into videoconferencing and lecture halls and cinema apps. The World Health Organization estimates that more than 5 percent of the world’s population — some 430 million people — have disabling hearing loss, and only about one in five people who could benefit from a hearing aid uses one; captions flowing into every corner of digital life is real progress for a population that spent decades begging for basic access. Availability has never been better. It is specifically quality that is being re-optimized, because quality is where the two audiences diverge, and a market serving two populations will tune itself to the larger one every time.

The uncomfortable conclusion is that the curb-cut analogy flatters us. It suggests that inclusion scales frictionlessly — that once the ramp is poured, everyone benefits equally and forever. Captions suggest something more conditional: mainstream adoption gets you ubiquity, and then immediately starts eroding the standard the original users fought for, because the new majority cannot perceive the erosion. The hearing viewer with subtitles on is not freeloading on a deaf accommodation; she is actively repricing it. Her presence is why captions are everywhere, and also why “everywhere” increasingly means auto-generated, unreviewed, roughly-right.

The fix is not to resent the caption generation — the genie is out, and honestly the genie is great. It is to notice that “good enough for most viewers” has become the de facto industry standard precisely because most viewers can now decide what good enough means. The people for whom captions are the soundtrack never got that vote. They had to win quality through law, and the law has not been updated for a world where the feature’s biggest fans cannot hear its failures.