On-screen text legibility for video is the difference between a caption a viewer absorbs in half a second and one they scroll past without a thought. Most clips fail at this not because the copy is wrong, but because the text is too small, too low in contrast, or pushed against an edge where a phone interface crops it. This guide breaks down the three levers that govern readability, size, contrast, and safe margins, and shows how an automated quality gate can verify every shot before it ships.
Why on-screen text fails so often
Video gets watched in hostile conditions. A viewer scrolls a muted feed on a bright sidewalk, glances at a six-inch screen, and decides in under a second whether your frame is worth stopping for. Text that looked crisp on your 27-inch editing monitor can dissolve into noise on that phone.
Three things conspire against you. Compression smears thin strokes and tight letterforms into mush. Busy or moving backgrounds rob characters of the contrast they need to separate from the image. And platform interfaces, caption toggles, profile avatars, progress bars, follow buttons, stamp themselves over the bottom and side thirds of the frame, exactly where designers love to put a lower-third title. Legibility is a brand problem as much as a design one, which is why it belongs in your broader brand foundations rather than being treated as a one-off styling choice.
How big should text be on a video?
There is no single pixel value that works everywhere, because big enough depends on the canvas and the viewing distance. The reliable approach is to size text relative to the frame height rather than in fixed pixels.
A practical rule of thumb: a primary headline or key caption should occupy roughly one-tenth to one-eighth of the frame height in cap height, and body or supporting text should rarely drop below about one-twentieth of frame height. On a 1080-pixel-tall vertical video, that puts your main line in the neighborhood of 90 to 110 pixels tall and your smallest supporting text above roughly 50 pixels. Translate the same proportions to 4K and the pixel numbers grow, but the on-screen size a viewer perceives stays the same.
Two more sizing habits pay off. Keep your line count low, one or two lines per beat, never a paragraph, so each line can be large. And design for the smallest screen in your audience, not the largest, because text that reads on a phone always reads on a laptop, but the reverse is not true.
Contrast: the make-or-break factor
Size buys you nothing if the text blends into the background. Contrast is the single biggest predictor of whether on-screen text is readable, and it is the one most creators underestimate.
Aim for the accessibility standard used across the web: a contrast ratio of at least 4.5:1 between text and its immediate background for normal text, and 3:1 for large, bold display text. The trouble with video is that the background changes every frame, so a ratio that holds during one moment collapses during the next. Defend against this with a consistent treatment:
- Add a scrim. A semi-transparent dark or light panel behind the text guarantees a stable background regardless of the footage underneath.
- Use a subtle stroke or drop shadow. A thin outline or soft shadow separates letterforms from whatever sits behind them without screaming for attention.
- Keep text off the busiest part of the frame. Faces, fine textures, and high-motion regions are the worst beds for type.
White text with a soft shadow over a darkened lower third is a cliché precisely because it works in nearly every lighting condition.
Safe margins and placement
Even perfectly sized, high-contrast text fails if it lands where the device or the platform covers it. This is what broadcast engineers solved decades ago with safe areas, and the concept transfers directly to social video.
Title-safe and action-safe zones
Keep all readable text inside a title-safe zone, roughly the central 90% of the frame, leaving a 5% margin on every edge. Keep important graphics and faces inside an action-safe zone of about 93%. These margins absorb edge cropping and give your composition breathing room so nothing feels jammed against the bezel.
On social platforms you need a second, tighter exclusion zone for interface chrome. On vertical video, the bottom 15 to 20% of the frame frequently hides behind captions, usernames, and the description; the right edge hides behind the action rail of icons. Treat those regions as off-limits for anything a viewer must read.
Typography choices that help or hurt legibility
Font choice is more than taste: it changes how fast text resolves at speed and at size. For motion, favor typefaces with a generous x-height, open counters, and a medium-to-bold weight; hairline weights vanish under compression. The serif-versus-sans question matters here too, and our breakdown of serif vs sans-serif fonts covers when each reads better on screen.
Spacing is the quiet half of legibility. Letters set too tightly clot together at small sizes, while lines set too close are hard to track. If those terms are unfamiliar, our guide to kerning, tracking, and leading explains how to tune them for the screen.
Timing: give viewers time to read
A line can be large, high-contrast, and perfectly placed and still fail if it flashes by. A common editing guideline is to hold a line of text long enough to be read at a comfortable pace, plan for a few words per second and add a beat of padding. Short, punchy captions that match the rhythm of the narration read best; if you write those captions from reusable on-brand scripts, the on-screen text and the spoken line stay in lockstep.
How CoreReflex checks caption legibility on every shot
This is exactly the kind of thing that should be verified, not eyeballed. In CoreReflex, the agentic Director runs a quality gate on every shot it produces, and on-screen text legibility is one of the concrete checks it scores, alongside prompt match, sharpness, motion coherence, on-brand consistency, and claims risk.
When a shot's text is too small, too low in contrast, or sitting inside an unsafe margin, the gate flags it and triggers selective regeneration, only the failing shot is redone, not the whole film. Because every generation carries a portable provenance trace (the model, prompt, parameters, and score), you can replay exactly why a caption passed or failed and reproduce the fix. The result is a finished, graded cut where the text was held to a standard on the way out the door rather than caught by a viewer who couldn't read it. The check criteria are spelled out in our documentation if you want the specifics.
Frequently asked questions
How big should text be on a video?
Size text relative to the frame, not in fixed pixels. A primary caption should sit around one-tenth to one-eighth of the frame height, and supporting text should stay above roughly one-twentieth. Design for the smallest screen your audience uses, usually a phone, and keep lines to one or two so each can be large.
Why is my video text hard to read?
Almost always it is low contrast, small size, a busy background, or placement under platform interface elements. Add a scrim or shadow to reach a 4.5:1 contrast ratio, increase the size relative to the frame, keep text inside the title-safe zone, and hold each line on screen long enough to read.
Does CoreReflex check if captions are legible?
Yes. The quality gate scores on-screen text legibility on every shot as one of its concrete checks. Shots that fail are automatically and selectively regenerated, and the pass-or-fail decision is recorded in a replayable provenance trace.
Ship text people can actually read
Legible on-screen text is not a finishing touch. It is the part of the frame that does the talking when the sound is off. Get size, contrast, and margins right, then let an automated gate confirm it on every shot so nothing readable-in-theory ships as unreadable-in-practice. Start free with no credit card and watch the Director hold your captions to the standard.