Back to Research
Research

Video vs. text: what actually gets remembered inside a company.

Not all video beats all text. The research is more specific than that — and the specifics are exactly what determine whether a leadership message survives contact with a busy employee's week.

Two channels are better than one

In 1971, psychologist Allan Paivio proposed dual coding theory: the mind processes verbal information (words) and nonverbal information (images) through two separate channels, and information encoded through both channels at once is easier to retrieve later than information encoded through only one. A memo activates the verbal channel alone. A person speaking on camera — voice, face, tone, gesture — activates both at the same time.

This is the mechanistic reason a documentary-style video tends to stick with people in a way a written announcement doesn't: it isn't a stylistic preference, it's a difference in how many retrieval paths the memory has.

Where the effect breaks down

The honest version of this research has a limit worth naming. Richard Mayer's multimedia learning research — collected in The Cambridge Handbook of Multimedia Learning — found that adding on-screen text that repeats what a narrator is already saying does not help retention, and in most tested cases hurts it. Forcing people to read and listen to the same words at once competes for the same limited working memory instead of using two separate channels.

The advantage isn't "video" in the abstract. It's specifically pairing images with narration — not stacking narration on top of text, and not replacing a real person with a slide deck someone reads aloud. A video of a CEO talking to camera activates dual coding. A video of a CEO reading bullet points off a slide mostly doesn't.

What this adds beyond memory

Dual coding explains why a message is retained. It doesn't fully explain why it's trusted — that comes from what text structurally can't carry: tone, facial expression, pacing, the visible difference between someone who believes what they're saying and someone reciting talking points. Those nonverbal signals are part of how people decide whether to believe a message, not just whether they remember it.

Video also compounds differently than a live-only format. A message that only existed in the room it was said in can't be revisited — an archived video can be. That rewatchability is what turns a single exposure into message compounding: the same piece keeps generating retrieval opportunities long after the day it was published.

What this means for your company

Real people on camera, not slides read aloud. The retention advantage comes from pairing a real voice with a real face, not from the video format by itself.

Don't caption over narration for emphasis. Duplicating the message in on-screen text tends to make it harder to retain, not easier.

Archive it so it can be rewatched. A single exposure is a single data point. A retrievable one keeps paying off.

Alejandro Butula García

About the author

Alejandro Butula García is co-founder of Hermann® and directs the audiovisual identity for every client. Twenty-eight years in advertising, documentary and Olympic-level production, including campaigns for Procter & Gamble and the Olympic Channel series Heroes of the Future.

Sources

  • Paivio, A. (1971). Imagery and Verbal Processes. New York: Holt, Rinehart & Winston.
  • Mayer, R.E. (Ed.). The Cambridge Handbook of Multimedia Learning. Cambridge University Press.

If you are not the decision-maker,
do not schedule.