accessibility
Captions, transcripts and audio descriptions
What's the difference between captions, transcripts and audio descriptions?
These are three ways to make audio-visual content reach people who can't fully see or hear it. Captions put the audio (speech and meaningful sounds) on screen, in sync, for people who can't hear. A transcript is the whole thing as readable text. Audio description narrates the important visuals, in the gaps, for people who can't see. Each serves a different barrier - and most media needs more than one.
Also known as: captions, transcripts, audio description, subtitles, media accessibility
Prefer to watch?
Watch the recap 0:49
The demo
A short clip - someone can't hear it, someone else can't see it. Turn each track on and off, and notice which barrier each one removes, and who's still left out without it.
Audio description: “The chef sprinkles salt from a height, then winks.”
[upbeat music] Narrator: “…and a pinch of salt to finish.” [The chef sprinkles salt from a height, then winks at the camera.]
What this demo shows (text version)
A mock cooking clip has three independent tracks you can switch on. Captions put the spoken words and meaningful sounds ("[upbeat music]") on screen, in sync, for someone who can't hear. Audio description adds narration of the visual action - the chef sprinkling salt and winking - for someone who can't see. The transcript gives the entire thing as readable text, serving both groups plus anyone who'd rather read or search.
The point is that each removes a different barrier: captions and transcripts carry the audio to people who can't hear; audio description and transcripts carry the visuals to people who can't see. No single track does it all, so most media needs captions plus a transcript, with audio description added where the visuals carry information the soundtrack doesn't. And auto-captions need correcting - they tend to mangle the very words that matter.
Match the alternative to the barrier: captions and transcripts carry the audio to people who can't hear it; audio description and transcripts carry the visuals to people who can't see them. Captions must be synced and include meaningful non-speech sound; a transcript is the cheap, search-friendly catch-all; audio description is needed only when visuals carry information the soundtrack doesn't. Auto-captions are a starting point, not a finish line - correct them.
Toggle each track and you can feel who it's for: captions hand the dialogue to someone who can't hear it, audio description hands the on-screen action to someone who can't see it, and the transcript gives everyone the whole thing as text. One video, three different barriers - and no single track removes them all.
Captions are for people who can't hear the audio: synchronised on-screen text of the speech plus meaningful sounds ("[door slams]", "[ominous music]"). They differ from subtitles, which translate dialogue for hearing viewers and skip non-speech sound. Captions also help in noisy or sound-off settings - which is most of social media.
Audio description is for people who can't see the visuals: a narrator describes the important on-screen action in the natural pauses ("she slips the note into his pocket"). It's only needed when meaning is carried visually and isn't already clear from the soundtrack - so a talking-head clip may need none, while a silent visual gag needs a lot.
A transcript is the workhorse: the full content as text, serving both groups at once, plus anyone who'd rather read, skim or search - and it's the cheapest to produce and great for SEO. The practical rule: caption everything with speech, transcribe everything, and add audio description where visuals carry information. And fix auto-captions - "good enough" machine captions routinely mangle exactly the words that matter.