Accessibility and Inclusive Content
Caption and Transcript Review
Verify speech, speaker identity, important sounds, timing, synchronization, non-speech information, formatting, links, and media-text consistency.
Free editable Markdown · Media teams, accessibility reviewers, and editors ·
Accessible HTML preview
Blank template
The downloaded file contains the same fields in editable Markdown.
Media context
- Media title and ID
- [Stable identifier]
- Final version
- [Export/build/date]
- Duration and language
- [Time; source language]
- Caption type
- [Closed, open, live, translated]
- Transcript location
- [URL or file]
- Approved names and terms
- [Glossary/source]
- Style guide
- [Caption conventions]
- Reviewer and environment
- [Name; player/device]
Review issue
- Timestamp
- [Start–end]
- Current text
- [Caption/transcript]
- Audio or visual source
- [What occurs]
- Issue
- [Accuracy, speaker, sound, timing, line break, visual information]
- Consequence
- [Meaning, identity, action, accessibility]
- Revision
- [Exact text or timing]
- Verification source
- [Speaker, script, document, visual]
- Locale implication
- [Translation or timing]
Completion checks
- Review uses the locked final media.
- Names, numbers, quotations, and technical terms are verified.
- Speakers are identifiable when visuals are insufficient.
- Meaningful sounds and on-screen information have equivalents.
- Segmentation and timing follow natural phrasing.
- Captions avoid obscuring important visuals.
- Transcript headings and links support navigation.
- Localized tracks and live published assets are verified.
How to use this template
- Lock the media version and collect approved names, terminology, script, sources, and style requirements.
- Compare every caption or transcript passage with speech and visual information.
- Add speaker identification and meaningful non-speech sound without unnecessary noise.
- Test synchronization, segmentation, display, player controls, transcript structure, and mobile use.
- Resolve uncertainty, complete language review, and link release ownership to the final media version.
Match the final media exactly
Work from the final export or a version-locked cut. Compare spoken words, names, quotations, numbers, and terminology against approved sources. Remove artifacts from deleted scenes and add material introduced late. Preserve meaningful false starts or emphasis when they affect interpretation, while following the chosen caption style for ordinary fillers. Do not “correct” a speaker into saying something substantively different. Record uncertain audio and obtain confirmation rather than guessing.
Identify speakers and meaningful sound
Use clear speaker labels when the visual context does not make identity obvious, and keep labels consistent. Include non-speech audio that carries information, such as a warning tone, audience reaction, change in music, or off-screen action. Avoid describing every background noise. For a transcript, add headings, speaker turns, links, and descriptions that make the media structure navigable. Ensure information shown only on screen—titles, charts, steps, URLs, or demonstrations—has an accessible textual equivalent.
Check timing and reading conditions
Captions should synchronize closely enough that readers can connect text, speaker, and action. Segment at natural phrase boundaries and give sufficient display time without covering essential visuals. Test small screens, high playback speed, pause, full-screen mode, contrast, and player controls. Translated captions need linguistic and timing review because sentence length changes. Publish a transcript in an accessible format near the media and give it the same update ownership as the video.
See the fields in context
Fictional example: museum map demonstration
Harbor Lens Museum and the video are invented for this caption review.
- Timestamp: 02:14–02:21 in the fictional final cut.
- Issue: Automated text writes “east wing,” while the speaker says “west wing,” changing directions.
- Visual information: A cursor highlights the invented west-wing route; the transcript adds that action.
- Sound: A meaningful confirmation tone is captioned once as “[route saved].”
- Resolution: Correct direction, verify the fictional map label, and retime the caption before live player testing.
Frequently asked questions
Should captions be verbatim?
They should preserve the speaker's meaning and relevant speech according to the applicable caption style. Do not edit substance for polish.
Which sounds need captions?
Include sounds necessary to understand action, mood, identity, warning, or meaning. Routine background noise may not need description.
Is a transcript enough without captions?
Not for synchronized access to video in contexts where captions are needed. Provide the formats required by the media and applicable standards.
Can automatically generated captions be published directly?
They require human review. Automated output commonly misidentifies speakers, names, numbers, accents, and specialized terms.