Skip to content

Accessibility and Inclusive Content

Caption and Transcript Review

Verify speech, speaker identity, important sounds, timing, synchronization, non-speech information, formatting, links, and media-text consistency.

Free editable Markdown · Media teams, accessibility reviewers, and editors ·

Download Markdown

Accessible HTML preview

Blank template

The downloaded file contains the same fields in editable Markdown.

Media context

Media title and ID
[Stable identifier]
Final version
[Export/build/date]
Duration and language
[Time; source language]
Caption type
[Closed, open, live, translated]
Transcript location
[URL or file]
Approved names and terms
[Glossary/source]
Style guide
[Caption conventions]
Reviewer and environment
[Name; player/device]

Review issue

Timestamp
[Start–end]
Current text
[Caption/transcript]
Audio or visual source
[What occurs]
Issue
[Accuracy, speaker, sound, timing, line break, visual information]
Consequence
[Meaning, identity, action, accessibility]
Revision
[Exact text or timing]
Verification source
[Speaker, script, document, visual]
Locale implication
[Translation or timing]

Completion checks

  • Review uses the locked final media.
  • Names, numbers, quotations, and technical terms are verified.
  • Speakers are identifiable when visuals are insufficient.
  • Meaningful sounds and on-screen information have equivalents.
  • Segmentation and timing follow natural phrasing.
  • Captions avoid obscuring important visuals.
  • Transcript headings and links support navigation.
  • Localized tracks and live published assets are verified.

How to use this template

  1. Lock the media version and collect approved names, terminology, script, sources, and style requirements.
  2. Compare every caption or transcript passage with speech and visual information.
  3. Add speaker identification and meaningful non-speech sound without unnecessary noise.
  4. Test synchronization, segmentation, display, player controls, transcript structure, and mobile use.
  5. Resolve uncertainty, complete language review, and link release ownership to the final media version.

Match the final media exactly

Work from the final export or a version-locked cut. Compare spoken words, names, quotations, numbers, and terminology against approved sources. Remove artifacts from deleted scenes and add material introduced late. Preserve meaningful false starts or emphasis when they affect interpretation, while following the chosen caption style for ordinary fillers. Do not “correct” a speaker into saying something substantively different. Record uncertain audio and obtain confirmation rather than guessing.

Identify speakers and meaningful sound

Use clear speaker labels when the visual context does not make identity obvious, and keep labels consistent. Include non-speech audio that carries information, such as a warning tone, audience reaction, change in music, or off-screen action. Avoid describing every background noise. For a transcript, add headings, speaker turns, links, and descriptions that make the media structure navigable. Ensure information shown only on screen—titles, charts, steps, URLs, or demonstrations—has an accessible textual equivalent.

Check timing and reading conditions

Captions should synchronize closely enough that readers can connect text, speaker, and action. Segment at natural phrase boundaries and give sufficient display time without covering essential visuals. Test small screens, high playback speed, pause, full-screen mode, contrast, and player controls. Translated captions need linguistic and timing review because sentence length changes. Publish a transcript in an accessible format near the media and give it the same update ownership as the video.

See the fields in context

Fictional example: museum map demonstration

Harbor Lens Museum and the video are invented for this caption review.

  • Timestamp: 02:14–02:21 in the fictional final cut.
  • Issue: Automated text writes “east wing,” while the speaker says “west wing,” changing directions.
  • Visual information: A cursor highlights the invented west-wing route; the transcript adds that action.
  • Sound: A meaningful confirmation tone is captioned once as “[route saved].”
  • Resolution: Correct direction, verify the fictional map label, and retime the caption before live player testing.

Frequently asked questions

Should captions be verbatim?

They should preserve the speaker's meaning and relevant speech according to the applicable caption style. Do not edit substance for polish.

Which sounds need captions?

Include sounds necessary to understand action, mood, identity, warning, or meaning. Routine background noise may not need description.

Is a transcript enough without captions?

Not for synchronized access to video in contexts where captions are needed. Provide the formats required by the media and applicable standards.

Can automatically generated captions be published directly?

They require human review. Automated output commonly misidentifies speakers, names, numbers, accents, and specialized terms.

File details

File name
caption-and-transcript-review.md
Format
Markdown (.md)
Size
2 KB
Designed for
Media teams, accessibility reviewers, and editors

Usage note: Review captions while watching and listening to the final edited media, and review the transcript as a standalone resource. Follow applicable accessibility requirements and involve qualified captioners, translators, or audio-description specialists when needed. Automated speech recognition can create a starting draft, but names, numbers, accents, technical terms, speakers, and consequential wording require human verification.