Brain Dump
Research••

Voice vs. writing: what research says about spoken journaling

Speech runs at 153 words per minute against 52 for typing. But one study since 1994 has compared the two directly, and it found no winner.

Voice vs. writing: what research says about spoken journaling

Summary: The honest state of the evidence: one study, from 1994, has ever compared speaking to writing directly, and it found similar emotional processing in both. There is no outcome result that favours voice. What voice has is a measured speed advantage and a plausible answer to the paradigm's real failure mode, which is that people do not finish.

What the evidence actually is

The expressive-writing literature is a writing literature. Smyth's 1998 meta-analysis pooled 13 studies of written emotional expression. Frattaroli's 2006 meta-analysis pooled 146 randomised studies and reported r = .075, roughly d = 0.15. Every one of those participants used a pen or a keyboard.

Exactly one published study put speaking and writing head to head. Murray and Segal (1994) ran 20-minute vocal and written sessions about traumatic events over 4 days. The finding: "similar emotional processing was produced by vocal and written expression." Painfulness of the topic fell steadily across the 4 days in both conditions. Both groups ended up feeling better about the topic and about themselves. Content analysis suggested more overt expression of emotion in the vocal condition. Negative emotion rose after every session, in both.

That is the entire direct evidence base. One study, 1994, measuring emotional processing rather than health at follow-up. Pennebaker's own 1999 review treats writing and talking as the same paradigm, but a narrative review is not a trial.

So this page does not claim that voice works better. Nobody has shown that. What follows is what can be said.

The speed difference is measured

Ruan and colleagues (2018) tested speech recognition against the built-in touchscreen keyboard on an iPhone 6 Plus, under laboratory conditions favourable to both:

Language Speech Keyboard Ratio
English 153 wpm 52 wpm 2.93x
Mandarin 123 wpm 43 wpm 2.87x

Speech corrected fewer errors during entry (5.30% vs 11.22%) but left slightly more in the final text (1.30% vs 0.79%).

This is a text-entry result, not a journaling result. It says nothing about whether the entries help you. It says that in a fixed five minutes you can produce roughly three times as much material by speaking.

The argument that survives: adherence, not outcomes

The expressive-writing paradigm has a failure mode nobody advertises. Baikie and colleagues (2006) ran the standard four-session protocol in a primary care clinic. They recruited 53 people. Fourteen completed all measures. That is a 26% completion rate, and the trial found no significant health benefits, which is exactly what a study with 14 completers should be expected to find.

This is the honest case for voice, and it is a case about behaviour rather than biology: a protocol that produces a small effect when completed produces nothing at all when abandoned. Fifteen minutes of continuous writing is a real ask. Fifteen minutes of talking, hands free, no blank page, is a smaller one.

Nobody has tested whether voice actually improves completion rates. If you see that claim stated as fact, including on this site, it is an inference and not a finding.

What voice plausibly changes

Emotional prosody

When you speak, you carry emotion through channels text does not have: pitch, rhythm, pace, volume, pauses, voice quality. Scherer, Banse and Wallbott (2001) found that emotion inferences drawn from vocal expression correlate across languages and cultures, which is evidence that the signal is in the voice itself rather than in the words.

Whether that signal does anything for a journaler reviewing their own recordings has not been studied. Murray and Segal's content analysis is the closest thing: speakers expressed more overt emotion than writers did.

Less filtering

Speaking is faster than your internal editor. When you write, there is a pause between thought and word where you edit, rephrase and second-guess. Speaking narrows that gap.

The Pennebaker protocol instructs people to keep going without worrying about grammar or structure. Voice makes that instruction easier to obey. This is a mechanism story, not a result.

Accessibility

Voice journaling works when writing doesn't:

What writing captures that voice doesn't

Deliberate processing

Writing is slower. That's usually framed as a disadvantage, but for certain kinds of journaling it's the point.

When you write, the speed constraint forces you to choose words. That choosing is cognitive work: you evaluate, categorise and structure as you go. Whether the slowness itself produces any of the benefit is an open question. No study has isolated it.

One caution against reading too much into linguistic markers. In the Baikie pilot, more positive-emotion words were associated with improvement in depression and stress, but increasing use of cognitive-process words was associated with worsening depressive mood. The markers are not a scoreboard.

Visual structure

Writing creates a spatial artifact. You can see your thoughts arranged on a page, draw connections, create lists and tables, mark important passages. Methods like fear-setting or bullet journaling rely on visual layout that pure voice can't replicate.

Privacy

You can write silently. You can't speak silently. In shared living spaces, offices, or public transit, writing offers privacy that voice doesn't.

Permanence of format

A written journal is immediately readable. A voice recording requires playback (or transcription) to review. The friction of reviewing voice entries is higher than scanning written pages.

What Pennebaker says

James Pennebaker, whose expressive writing protocol is the most-studied journaling intervention in psychology, has written about talking as well as writing. His 1999 review is titled "the values of writing and talking about upsetting events," which tells you he considered both in scope.

That is a framing, not a result. He did not run the head-to-head. Murray and Segal did, five years earlier, and found the two comparable on emotional processing.

His text analysis tool, LIWC, runs on any text, transcripts included. That makes voice entries measurable by the same instrument. It does not make them equally effective, and no study has checked.

The honest comparison

Speed and error rates below are from the Ruan text-entry study. The rest is a description of the two media, not a summary of findings, because the findings do not exist.

Dimension Voice Writing
Entry rate 153 wpm (English, measured) 52 wpm on a phone keyboard (measured)
Uncorrected errors 1.30% (measured) 0.79% (measured)
Evidence base One 1994 comparison study 146 randomised studies pooled by Frattaroli
Emotional signal Prosody, tone, pacing Deliberate word choice
Friction Near-zero setup Requires pen or device
Review Needs transcription Immediately scannable
Privacy Audible to others Silent
Accessibility Hands-free, eyes-free Requires motor control
Visual structure None, audio is linear Spatial layouts, lists, tables

Neither modality has been shown better. The right choice depends on:

The hybrid approach

Most experienced journalers end up using both:

  1. Voice for capture: raw brain dump, emotional processing, on-the-go thoughts
  2. Transcript for review: read the transcription, highlight key insights, add structure
  3. Writing for synthesis: use the transcript as raw material for structured entries, lists, or action items

This gives you voice's speed and emotional fidelity at the capture stage, plus writing's structure and deliberateness at the review stage. The transcript is the bridge.

What's missing from the research

Searches of PubMed, Europe PMC and Crossref for this page returned nothing on the following. These are gaps, not oversights:

What is left is this. Writing produces a small pooled effect, r = .075 across 146 randomised studies. One 1994 study found speaking and writing produced similar emotional processing. A quarter of the people enrolled in one clinical pilot of the written protocol finished it.

Use the modality you will actually complete. That is not a hedge, it is the only claim the evidence supports.


Related:

References

  1. Emotional processing in vocal and written expression of feelings about traumatic experiences • https://pubmed.ncbi.nlm.nih.gov/8087401/ • Murray & Segal (1994), Journal of Traumatic Stress 7(3):391-405. The only direct head-to-head: 20-minute vocal vs written sessions over 4 days. Similar emotional processing in both; content analysis showed greater overt emotion in the vocal condition. Negative emotion rose after every session in both.
  2. The effects of traumatic disclosure on physical and mental health: the values of writing and talking about upsetting events • https://pubmed.ncbi.nlm.nih.gov/11227757/ • Pennebaker (1999), International Journal of Emergency Mental Health. A narrative review by the originator, framing writing and talking as one paradigm. Not a trial.
  3. Written emotional expression: effect sizes, outcome types, and moderating variables • https://pubmed.ncbi.nlm.nih.gov/9489272/ • Smyth (1998), J Consult Clin Psychol 66(1):174-84. The d=0.47 figure everyone quotes. 13 studies, fixed effects, and Smyth names publication status as a moderator. Written expression only.
  4. Experimental disclosure and its moderators: a meta-analysis • https://pubmed.ncbi.nlm.nih.gov/17073523/ • Frattaroli (2006), Psychological Bulletin 132(6):823-65. 146 randomised studies, random effects, r = .075 (roughly d = 0.15). The most defensible estimate, and it is small.
  5. Expressive writing for high-risk drug dependent patients in a primary care clinic: a pilot study • https://pubmed.ncbi.nlm.nih.gov/17112389/ • Baikie et al. (2006), Harm Reduction Journal. 53 recruited, 14 (26%) completed all measures. No significant health benefits. The clearest published record of the protocol's adherence problem.
  6. Comparing speech and keyboard text entry for short messages in two languages on touchscreen phones • https://doi.org/10.1145/3161187 • Ruan, Wobbrock, Liou, Ng & Landay (2018), Proc. ACM IMWUT. iPhone 6 Plus lab study: speech 153 wpm vs keyboard 52 wpm in English (2.93x), 123 vs 43 wpm in Mandarin. Speech left more uncorrected errors, 1.30% vs 0.79%.
  7. Emotion inferences from vocal expression correlate across languages and cultures • https://doi.org/10.1177/0022022101032001009 • Scherer, Banse & Wallbott (2001), Journal of Cross-Cultural Psychology 32(1). Emotion inferences drawn from vocal expression hold up across language and culture.
  8. Opening Up by Writing It Down (James Pennebaker & Joshua Smyth) • https://www.guilford.com/books/Opening-Up-by-Writing-It-Down/Pennebaker-Smyth/9781462524921 • Third edition. The general-audience account of the expressive-writing paradigm and the 4-day protocol.
  9. LIWC: Linguistic Inquiry and Word Count • https://www.liwc.app/ • Pennebaker's text analysis tool. Runs on any text, including transcripts. Listed as the measurement instrument used across this literature, not as evidence for any modality.