Certified Realtime Captioner (CRC) — Questions and Answers
Question 1: What format does the iTunes Store (Apple TV+) require for closed caption delivery?
- SRT files only
- iTunes Timed Text (iTT) format (Correct answer)
- WebVTT only
- CEA-608 raw binary data
Correct answer: iTunes Timed Text (iTT) format
Apple's iTunes Store and Apple TV+ require the iTunes Timed Text (iTT) format, which is an XML-based format developed by Apple for their content delivery pipeline.
Question 2: Which caption file format stores timing information using hours, minutes, seconds, and frames (HH:MM:SS;FF)?
- VTT
- SRT
- TTML
- SCC (Correct answer)
Correct answer: SCC
SCC (Scenarist Closed Captions) uses SMPTE timecode in HH:MM:SS;FF format to sync captions to video frames.
Question 3: What technique is used to keep captions readable when speech is faster than the maximum recommended caption rate?
- Increasing the caption font size
- Using abbreviations exclusively
- Splitting the caption across multiple screens
- Condensing or paraphrasing the spoken words (Correct answer)
Correct answer: Condensing or paraphrasing the spoken words
When speech exceeds the maximum readable rate, captioners condense dialogue while preserving meaning to keep captions readable.
Question 4: When a speaker's words are inaudible or unintelligible on the audio track, how should the captioner handle it?
- Use an ellipsis (…) to indicate missing content
- Leave a blank gap with no caption
- Write [INAUDIBLE] or [UNINTELLIGIBLE] in brackets (Correct answer)
- Insert a best guess in regular text format
Correct answer: Write [INAUDIBLE] or [UNINTELLIGIBLE] in brackets
Using [INAUDIBLE] or [UNINTELLIGIBLE] in brackets alerts viewers that audio was present but could not be understood, maintaining transparency.
Question 5: What does 'timecode' refer to in subtitle synchronization?
- The frame rate of the video file
- The export time for the subtitle file
- The runtime of the entire video
- The in and out points that determine when a subtitle appears and disappears (Correct answer)
Correct answer: The in and out points that determine when a subtitle appears and disappears
Timecode in subtitle files specifies the exact in-point (when a subtitle appears) and out-point (when it disappears) for each caption entry.
Question 6: What happens to subtitle timing if a video is edited (trimmed or rearranged) after subtitles have been created?
- Subtitles automatically adjust
- Subtitles will be out of sync and must be re-timed (Correct answer)
- Only the first subtitle will be affected
- Subtitles are permanently embedded so editing is impossible
Correct answer: Subtitles will be out of sync and must be re-timed
Any editing of the video timeline after subtitles are created will desynchronize the subtitles, requiring re-timing or offset adjustment.
Question 7: In live captioning, what technology do CART providers use to produce real-time captions?
- Voice recognition software only
- AI translation engines
- Stenography machines with specialized software (Correct answer)
- Keyboard transcription at high typing speeds
Correct answer: Stenography machines with specialized software
CART (Communication Access Realtime Translation) providers use stenography machines that allow trained stenographers to write phonetic shorthand translated into text in real time.
Question 8: What can be used to show the audience unfamiliar words in a documentary or film?
- Captions (Correct answer)
- Glossary
- Voice-over
- Subtitles
Correct answer: Captions
While both subtitles and captions display text, captions are often used more broadly to provide additional contextual information, including explanations of unfamiliar words, terms, or on-screen text that might appear in a documentary or film. They can clarify specific details beyond just dialogue translation, enriching the viewer's understanding. This makes captions particularly useful for educational or informative content.
Question 9: What is the key advantage of the TTML (Timed Text Markup Language) format over SRT?
- TTML is an XML-based standard that supports complex styling, regions, and metadata natively (Correct answer)
- TTML supports audio tracks in addition to text
- TTML has faster rendering speed in browsers
- TTML files are smaller than SRT files
Correct answer: TTML is an XML-based standard that supports complex styling, regions, and metadata natively
TTML uses XML markup allowing detailed control over styling, positioning, and metadata, making it suitable for professional broadcast and streaming workflows.
Question 10: What is 'pre-recorded' or 'offline' captioning compared to live captioning?
- Captions provided by volunteers watching the broadcast
- Captions generated automatically by AI in real time
- Captions created before broadcast from a script or transcript, allowing precise timing (Correct answer)
- Captions added by viewers using a remote control
Correct answer: Captions created before broadcast from a script or transcript, allowing precise timing
Offline captioning is produced in advance using the recorded video and transcript, enabling precise timing and accuracy.
Question 11: What does 'subtitle condensation' mean in translation?
- Compressing subtitle files to reduce file size
- Lowering the caption display speed
- Reducing the number of subtitle lines per scene
- Shortening translated text to fit timing and line length constraints without losing essential meaning (Correct answer)
Correct answer: Shortening translated text to fit timing and line length constraints without losing essential meaning
Subtitle condensation involves reducing the word count of translated text to fit within timing and character constraints while preserving the core meaning.
Question 12: What is the term for captions that are embedded within the video container file as a separate data track rather than burned into the picture?
- External captions
- Open captions
- Hard-coded captions
- Soft-coded captions (Correct answer)
Correct answer: Soft-coded captions
Soft-coded captions (also called soft subs) are embedded in the video container as a separate selectable track, while hard-coded (open) captions are burned into the picture permanently.
Question 13: What is the standard caption text color used against a black background box in US closed captions?
- White (Correct answer)
- Green
- Cyan
- Yellow
Correct answer: White
White text on a black background box is the standard used in US closed captions for maximum contrast and readability.
Question 14: Which subtitle file format stores timing in milliseconds using an arrow notation (-->) to separate in and out times?
- TTML
- WebVTT
- SRT (SubRip Text) (Correct answer)
- STL
Correct answer: SRT (SubRip Text)
SRT files use a simple format with a sequence number, timecode line using --> notation in HH:MM:SS,ms format, and text lines.
Question 15: What is the standard maximum caption reading speed recommended for adult audiences in the U.S.?
- 120 words per minute
- 250 words per minute
- 300 words per minute
- 180 words per minute (Correct answer)
Correct answer: 180 words per minute
Industry guidelines typically recommend a maximum of 180 words per minute for adult caption reading speed.
Question 16: How are sound effects typically indicated in US caption formatting?
- Using quotation marks
- Using bold text
- Using underlined text
- Using ALL CAPS in brackets or parentheses (Correct answer)
Correct answer: Using ALL CAPS in brackets or parentheses
Sound effects are written in brackets or parentheses in ALL CAPS (e.g., [THUNDER CLAP]) to distinguish them from dialogue.
Question 17: What does 'caption extraction' mean in post-production workflows?
- Automatically generating captions from a video file
- Pulling embedded caption data out of a video file into a separate text or subtitle file (Correct answer)
- Removing captions from a broadcast signal
- Converting captions from one language to another
Correct answer: Pulling embedded caption data out of a video file into a separate text or subtitle file
Caption extraction strips embedded caption data (like CEA-708 from MXF files) into standalone subtitle files for editing, archiving, or repurposing.
Question 18: Which law extended closed captioning requirements to online video distributed by TV broadcasters and cable operators?
- Telecommunications Act of 1996
- Twenty-First Century Communications and Video Accessibility Act (CVAA) (Correct answer)
- Rehabilitation Act Section 508
- Americans with Disabilities Act
Correct answer: Twenty-First Century Communications and Video Accessibility Act (CVAA)
The CVAA of 2010 extended FCC captioning requirements to IP-delivered video that previously aired on TV with captions.
Question 19: What is the recommended caption safe area to avoid text being cut off on older TV screens?
- 50% of the screen center only
- 10% inset from all screen edges (Correct answer)
- The top quarter of the screen
- The full screen edge to edge
Correct answer: 10% inset from all screen edges
A 10% safe area inset from all screen edges ensures caption text is not cut off by TV overscan on older display technology.
Question 20: What subtitle format is commonly required by Netflix for professional content delivery?
- Plain text (.txt)
- SRT only
- PDF captions
- TTML/IMSC (Timed Text Markup Language) (Correct answer)
Correct answer: TTML/IMSC (Timed Text Markup Language)
Netflix requires TTML/IMSC (also called DFXP) for professional content delivery because it supports precise formatting, positioning, and styling metadata.
Question 21: What is the typical accuracy rate goal for professional human captioning according to broadcast industry standards?
- 75% accuracy
- 99% accuracy (Correct answer)
- 85% accuracy
- 90% accuracy
Correct answer: 99% accuracy
Professional human captioning is expected to achieve 99% accuracy, which is the standard set by organizations such as the Described and Captioned Media Program (DCMP).
Question 22: What percentage of caption accuracy does the FCC generally expect for broadcast content?
- 85% accuracy
- 70% accuracy
- 100% accuracy is always required
- 99% accuracy (Correct answer)
Correct answer: 99% accuracy
The FCC expects closed captions to be accurate and virtually verbatim, with general industry expectations of 99% or higher word accuracy, though live content has more flexibility.
Question 23: On platforms like YouTube, what is the default auto-captioning technology called?
- FCC-compliant ASR engine
- YouTube Closed Captioning System (YCCS)
- Automatic captions powered by Google's speech recognition (Correct answer)
- Community Caption Generator
Correct answer: Automatic captions powered by Google's speech recognition
YouTube uses Google's speech recognition AI to automatically generate captions for uploaded videos, labeled as 'automatic captions' in the interface.
Question 24: What is a 'roll-up' caption style?
- Captions that spin onto the screen
- Captions that appear in a bubble
- Captions that continuously scroll upward one line at a time (Correct answer)
- Captions that roll in from the side
Correct answer: Captions that continuously scroll upward one line at a time
Roll-up captions scroll upward continuously, typically showing 1-3 lines at a time, and are commonly used for live television.
Question 25: How many characters should a caption ideally be?
- 100-110
- 200-250
- under 80
- 138 to 150 (Correct answer)
Correct answer: 138 to 150
While Instagram allows for longer captions, research and best practices suggest that an ideal length for maximum engagement and readability is typically between 138 and 150 characters. This concise length ensures that the core message is conveyed quickly and effectively, as only the first few lines are visible before users need to click 'more'.
Question 26: What is the primary limitation of the SRT format compared to WebVTT?
- SRT cannot store timecodes in milliseconds
- SRT does not natively support styling, positioning, or speaker identification metadata (Correct answer)
- SRT files are larger than WebVTT files
- SRT only works on Windows systems
Correct answer: SRT does not natively support styling, positioning, or speaker identification metadata
SRT is a plain text format that stores only timecodes and text without native support for styling (fonts, colors) or precise positioning, unlike WebVTT which supports CSS styling and cue positioning.
Question 27: What is 'subtitle latency' in live captioning?
- The delay between speech being spoken and the caption appearing on screen (Correct answer)
- The time it takes to export a subtitle file
- The gap between two subtitle entries
- The file size of subtitle data
Correct answer: The delay between speech being spoken and the caption appearing on screen
Subtitle latency refers to the unavoidable delay between a live speaker's words and when the corresponding captions appear on screen.
Question 28: In subtitle localization, what does 'lip sync' or 'dubbing sync' refer to?
- Ensuring audio and video are synchronized before subtitling
- Synchronizing two subtitle tracks in different languages
- Matching subtitle timing and word choice to the mouth movements in dubbed audio replacements (Correct answer)
- Matching the subtitle reading speed to the narrator's pace
Correct answer: Matching subtitle timing and word choice to the mouth movements in dubbed audio replacements
In dubbing sync subtitling, text timing and wording is adjusted to align with the lip movements of dubbed audio, common in dubbed foreign films.
Question 29: Under FCC rules, captions must be 'synchronous,' which means they must:
- Coincide with the corresponding spoken words and sounds (Correct answer)
- Match the video frame rate precisely
- Appear before the speaker begins talking
- Be delivered at exactly 130 words per minute
Correct answer: Coincide with the corresponding spoken words and sounds
FCC synchrony requirements mandate that captions appear at the same time as the corresponding audio they represent.
Question 30: What does 'Line 21' refer to in the context of US closed captioning?
- A paragraph formatting rule for captions
- A specific scan line in the NTSC analog video signal used to carry caption data (Correct answer)
- The 21st subtitle in a caption file
- The 21-character line limit for captions
Correct answer: A specific scan line in the NTSC analog video signal used to carry caption data
Line 21 is a specific horizontal scan line in the NTSC analog video signal's vertical blanking interval that carries CEA-608 closed caption data.
Question 31: What does Section 508 of the Rehabilitation Act require regarding captions?
- Federal agencies and federally funded programs must make electronic information, including video with captions, accessible to people with disabilities (Correct answer)
- All US websites must provide captions for all audio content
- Private companies with 50+ employees must caption all training videos
- All US manufacturers must include caption chips in TVs
Correct answer: Federal agencies and federally funded programs must make electronic information, including video with captions, accessible to people with disabilities
Section 508 requires federal agencies and entities receiving federal funds to ensure electronic and information technology, including video content, is accessible to people with disabilities, including through captions.
Question 32: What is the main accuracy limitation of Automatic Speech Recognition (ASR) captioning?
- It cannot process audio faster than 100 words per minute
- It struggles with accents, technical jargon, background noise, and multiple overlapping speakers (Correct answer)
- It requires an internet connection to function
- It only works with English-language content
Correct answer: It struggles with accents, technical jargon, background noise, and multiple overlapping speakers
ASR accuracy decreases significantly with accents, specialized terminology, background noise, and simultaneous speakers.
Question 33: In captioning technology, what does an encoder do in a broadcast workflow?
- Monitors caption quality and flags errors
- Transcribes audio to text automatically
- Inserts caption data into the video signal for transmission to viewers (Correct answer)
- Converts video from analog to digital format
Correct answer: Inserts caption data into the video signal for transmission to viewers
A caption encoder inserts the encoded caption data stream into the broadcast video signal so it can be decoded and displayed by viewers' televisions.
Question 34: In subtitle files, what does 'timecode' represent?
- The frame rate of the video file
- The language identifier of the subtitle track
- The in and out points that define when a subtitle appears and disappears (Correct answer)
- The font size metadata
Correct answer: The in and out points that define when a subtitle appears and disappears
Timecodes in subtitle files mark the start (in-point) and end (out-point) times for each subtitle line to appear on screen.
Question 35: What does 'burned-in subtitles' mean when exporting video?
- Subtitles that overheat the display device
- Subtitle text that has been rendered permanently into the video frames as open captions (Correct answer)
- Subtitles encoded with high color saturation
- Subtitles that have been translated and re-encoded
Correct answer: Subtitle text that has been rendered permanently into the video frames as open captions
Burned-in subtitles (also called hard-coded or open captions) are rendered as part of the video image itself, making them visible on all playback devices without needing a decoder.
Question 36: What technology is used to encode Line 21 closed captions in analog US television broadcasts?
- Audio description track encoding
- HDMI metadata packets
- Vertical Blanking Interval (VBI) data embedded in the video signal (Correct answer)
- Digital watermarking
Correct answer: Vertical Blanking Interval (VBI) data embedded in the video signal
Analog US closed captions are encoded in Line 21 of the Vertical Blanking Interval (VBI), which carries caption data invisible within the video signal.
Question 37: What is the recommended maximum gap between the end of one caption and the start of the next?
- No more than 2 seconds (Correct answer)
- No more than 5 seconds
- No limit as long as no speech occurs
- No more than 10 seconds
Correct answer: No more than 2 seconds
Best practices recommend gaps between captions not exceed 2 seconds to avoid confusing viewers about when speech resumes.
Question 38: Which of the following best describes the ADA (Americans with Disabilities Act) relationship to captioning for public accommodations?
- ADA requires effective communication accessibility, which may include captions for video content in public settings (Correct answer)
- ADA sets the same caption quality standards as the FCC
- ADA directly mandates specific caption formats for all video content
- ADA only applies to physical building accessibility, not video
Correct answer: ADA requires effective communication accessibility, which may include captions for video content in public settings
While the FCC regulates broadcast captioning, the ADA's effective communication requirements can extend captioning obligations to video in public accommodation contexts like theaters and public venues.
Question 39: What is the difference between a 'transcript' and 'synchronized captions' from an accessibility standpoint?
- There is no meaningful accessibility difference between the two
- Transcripts are more accessible than captions for all users
- Synchronized captions appear time-aligned with the video; transcripts are text that may not be synchronized (Correct answer)
- Transcripts are required by law; captions are optional
Correct answer: Synchronized captions appear time-aligned with the video; transcripts are text that may not be synchronized
Synchronized captions appear in real-time alignment with audio, allowing deaf viewers to follow the video; transcripts are separate text documents that may not indicate when in the video each word occurs.
Question 40: What is the purpose of the 'paint-on' caption style?
- To show captions in italics
- To reveal captions word by word from left to right (Correct answer)
- To display captions in a box
- To display captions at the top of the screen
Correct answer: To reveal captions word by word from left to right
Paint-on captions appear word by word from left to right, often used for live captioning to simulate real-time speech.
Question 41: What is the main drawback of open captions for online video platforms?
- They reduce video quality significantly
- They are harder to read than closed captions
- They are not supported by modern browsers
- They cannot be turned off and may be unwanted by some viewers (Correct answer)
Correct answer: They cannot be turned off and may be unwanted by some viewers
Open captions are permanently embedded in the video, so viewers who do not want captions cannot turn them off, which may reduce viewer satisfaction.
Question 42: What does 'speaker identification' in caption formatting refer to?
- Labeling the name of the show
- Identifying the captioning company
- Marking timestamps for speakers
- Indicating which character or person is speaking (Correct answer)
Correct answer: Indicating which character or person is speaking
Speaker identification labels (e.g., [JOHN:] or >> JOHN:) indicate which character or person is producing the dialogue in a caption.
Question 43: What is the difference between captions and subtitles in US terminology?
- Captions are only for online video; subtitles are only for broadcast TV
- Captions include non-speech audio cues; subtitles typically provide only dialogue (Correct answer)
- They are identical terms with no distinction
- Subtitles are for domestic audiences; captions are for foreign language viewers
Correct answer: Captions include non-speech audio cues; subtitles typically provide only dialogue
In US terminology, captions include non-speech audio descriptions (sound effects, speaker IDs) for deaf viewers, while subtitles typically only transcribe spoken dialogue for hearing viewers.
Question 44: What is 'described video' or 'audio description' and how does it differ from captioning?
- Audio description replaces captions for viewers with both hearing and vision loss
- Audio description is captions read aloud by a text-to-speech system
- Audio description is an older term for closed captioning
- Audio description narrates visual content for blind viewers; captioning converts audio to text for deaf viewers (Correct answer)
Correct answer: Audio description narrates visual content for blind viewers; captioning converts audio to text for deaf viewers
Audio description adds narration of visual elements (actions, expressions, scene changes) for visually impaired viewers, while captioning serves deaf or hard-of-hearing viewers by converting audio to text.
Question 45: A caption editor notices that two lines of text appear on screen simultaneously. What is the maximum recommended number of lines per caption frame according to best practices?
- 4 lines
- 3 lines
- 1 line
- 2 lines (Correct answer)
Correct answer: 2 lines
Best practices recommend no more than two lines of text per caption frame to avoid obstructing too much of the video.
Question 46: How should music or song lyrics be indicated in US caption formatting?
- Using ALL CAPS
- Using quotation marks
- Using parentheses
- Using music notes (♪) or angle brackets (Correct answer)
Correct answer: Using music notes (♪) or angle brackets
Music notes (♪ or ♫) are placed at the beginning and/or end of captioned lyrics to indicate singing or music content.
Question 47: What does the WebVTT format add over SRT that makes it better suited for web browsers?
- Built-in audio description tracks
- Support for styling, positioning, and speaker identification metadata within the file (Correct answer)
- Smaller file size through compression
- Faster caption loading speed
Correct answer: Support for styling, positioning, and speaker identification metadata within the file
WebVTT extends SRT's basic format with support for CSS-based styling, caption positioning, and rich metadata natively within the file.
Question 48: Which underlying technology powers YouTube's automatic captioning feature?
- CART transcription services provided by third parties
- Automatic Speech Recognition (ASR) technology (Correct answer)
- Human stenographers working remotely in real time
- Pre-loaded caption libraries matched to audio fingerprints
Correct answer: Automatic Speech Recognition (ASR) technology
YouTube's automatic captions are generated using Google's Automatic Speech Recognition (ASR) technology, which converts spoken audio to text.
Question 49: What does 'caption accuracy' specifically measure in quality assessment?
- The percentage of words correctly transcribed compared to actual spoken words (Correct answer)
- The number of caption frames per second
- Whether captions are positioned in the correct screen area
- How quickly captions appear relative to speech
Correct answer: The percentage of words correctly transcribed compared to actual spoken words
Caption accuracy is typically expressed as a percentage reflecting how closely the captioned text matches every spoken word in the audio.
Question 50: Which subtitle positioning is generally preferred to avoid obscuring important on-screen action?
- Bottom right corner
- Top left corner
- Top center of the screen
- Bottom center of the screen (Correct answer)
Correct answer: Bottom center of the screen
Subtitles are conventionally placed at the bottom center of the screen, where they are least likely to cover important visual content.
Certified Realtime Captioner (CRC)
The CRC certification exam tests professional knowledge of broadcast and CART captioning standards, caption formatting, accessibility and legal requirements, and the technical captioning environment. It is administered by the National Court Reporters Association (NCRA).
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds