Back to teardowns Second Chaptr Creative Lab.

Wispr Flow

The obvious promise is speed. Customer evidence points to a sharper territory: long-form work without the rewrite.

The exact pipeline we'd run on this account: raw customer signal to a batch of testable creative, and the loop that keeps it going.

325Comments coded
6Tensions
1Hypothesis
6Hooks
4Assets tested
01Evidence · Diverge — go wide

We start with everything they're already telling you

We coded 325 comments and reviews line by line, each tagged to the tension underneath it. A slice of the raw pile:

RedditGoogle PlayApp StoreProduct HuntCompetitor reviews

325 comments, coded across Reddit, Google Play, the App Store and Product Hunt. The brass lines mark one tension: post-dictation cleanup and accuracy.

02Tensions · Synthesize — narrow

The noise collapses into a handful of tensions

A feature isn't a tension. Coded down, six knots kept surfacing.

01Reliability regression & trust erosion“Transcription quality has gone down so far in the last 30 days it's ridiculous.”Reddit
02Meaning distortion during AI cleanup“Instead of 'I want' it becomes 'I don't want', and vice versa.”Reddit
03Language & accent recognition gaps“I select Urdu but it cannot type Urdu. It is typing Hindi.”Google Play
04Desktop–mobile experience gap“I'm constantly switching between keyboards, which kills any efficiency gain on iPhone.”Reddit
05App reliability & dictation failure“Freezes on 'Rolling' then shows 'Fail to fetch'.”Google Play
06Cloud privacy & offline-control concerns“Having my voice recordings sitting on someone else's servers felt weird.”Reddit
03Personas · Converge — who it's for

First we converge on who

Before choosing a message, we choose a person. The same comments map to several personas. Four matter for a first campaign; we'd build for the two that lead.

Long-form knowledge worker AI power user / engineer Mobile / on-the-go Accessibility user
Primary target

Long-form knowledge worker

Professionals who write detailed emails, updates, notes and first drafts all day.

Typical use cases Long emails, detailed updates, work messages, planning, note-taking.

Primary desired outcome

Minimize the time it takes to turn a detailed thought into a message that’s ready to use as-is.

Main barrier Accuracy, and the editing burden after speaking.

“I can write out detailed responses to my team without having to call them.” Google Play

Secondary target

AI power user / engineer

People who use ChatGPT, Cursor or other LLMs heavily and regularly create long, context-rich prompts.

Typical use cases Long prompts, coding sessions, brainstorming, explaining logic to AI.

Primary desired outcome

Minimize the time it takes to get a full, context-rich prompt into an LLM without losing your meaning.

Main barrier Meaning-changing rewrites, and accuracy on long dictations.

“Where it's most useful is structuring a very long prompt to any LLM.” Reddit

04Territory + bet · Converge — narrow to one

We don't pick a favourite. We score, then bet.

We scored six messaging territories for the primary persona, the long-form knowledge worker, and built the reel on the one that came out on top.

How we scored. Frequency = how often the idea appears in persona-aligned comments. Severity = inferred from the actual wording ("game-changing", "dangerous", "repair every few sentences", abandonment). Priority = frequency 60% / severity 40%.

1
Long-form is where voice wins Our pick · the reel Freq5 Sev4 Priority4.6
Problem

Typing becomes the bottleneck when the message, update, note or draft gets long.

Promise

Speak the full thought at natural speed and get the long-form communication down faster than typing.

Pain / intensity evidence“responding to emails and Teams messages while walking or on the go”Reddit

2
No cleanup tax after dictation Our pick · the reel Freq4 Sev5 Priority4.4
Problem

If users must repair punctuation, wording or recognition errors, the time advantage disappears.

Promise

Turn natural speech into text close enough to use immediately, with minimal cleanup.

Pain / intensity evidence“it keeps up with normal speaking speed without garbling things”Product Hunt

3
Write when the keyboard is not available Freq4 Sev4 Priority4.0
Problem

Long replies and ideas get postponed until the user is back at a proper keyboard.

Promise

Respond, capture and draft while walking, driving or away from the desk.

Pain / intensity evidence“I can even use it while listening to music, i.e. driving”Google Play

4
Protect the thought flow Freq3 Sev4 Priority3.4
Problem

Turning an idea into typed language can interrupt the thinking itself.

Promise

Speak first drafts, explanations and ideas without breaking momentum to type.

Pain / intensity evidence“Instead of stopping my flow to type everything manually, I simply speak my thoughts”Product Hunt

5
Reliability becomes table stakes for long-form work Freq2 Sev5 Priority3.2
Problem

Once dictation enters real work, failed or incomplete long transcriptions can break the workflow.

Promise

Make long-form voice input dependable enough to become part of daily work.

Pain / intensity evidence“Very good when it works, but every single day I waste time troubleshooting”Google Play

6
One voice layer across the workday Freq3 Sev3 Priority3.0
Problem

Work spans email, chat, notes, browsers and editors; a single-purpose tool adds friction.

Promise

Use the same voice workflow across the apps where work already happens.

Pain / intensity evidence“It formats differently whether you're in Slack, an email, or a code editor”Reddit

The bet

The reel combines the top two territories. Voice wins the moment typing becomes the bottleneck on long emails, updates and drafts, but only if the output is clean enough to send without a cleanup tax. So we position Wispr Flow for long-form work, without the rewrite.

05Hooks · Diverge — six ways in

Six doors into one message

The first three seconds. The line that earns the rest of the ad. Same message, six different doors in.

The message they all frame: when a message gets long, speaking beats typing, but only if it comes out clean enough to send as-is.

  • 01
    Identity

    “I run a research agency, and a surprising amount of my day is just getting long thoughts out of my head.”

  • 02
    Threshold

    “Short replies I'll happily type. It's the long ones that used to sit in my drafts for hours.”

  • 03
    Confession

    “I've tried voice-to-text apps so many times, and I always ended up going back to my keyboard.”

  • 04
    Insight

    “Voice-to-text isn't really saving you time if speaking faster just means more editing afterward.”

  • 05
    Experiment

    “I wrote the same long message two ways, just to see which one I'd actually have to fix.”

  • 06
    Curiosity

    “The first time an app gave me a long message I could send without editing, I didn't trust it.”

The test set. The bank runs deeper; these are the first round.

06Visual · the argument before the words

What the visual proves before a word is spoken

Most of the scroll happens with the sound off. The first frame has to make the argument on its own.

The visual thesis

Ours puts the cleanup on screen: a messy dictated message, then the cursor deleting filler, fixing punctuation, rewriting a line.

The cleanup nightmare Our pick · the reel opens here

First three seconds. Open on a messy dictated email. The cursor rapidly deletes filler, fixes punctuation, and rewrites a sentence.

Why it works The pain is legible instantly, and it pairs with the cleanup-tax hooks.

Also considered

Split-screen before / after
Messy voice-to-text on the left, clean output from the same spoken thought on the right.
Speak → send
Dictate one sentence, clean text appears, hit Send immediately.
Timer contradiction
A timer runs as you dictate, then keeps running while you edit the output.
Hands off the keyboard
Both hands visibly away while a long email appears as you speak.
Result-first: “I didn’t type this”
Finished email on screen, hit Send, then the text overlay lands.

A coherence note. The cleanup frame is a Territory-2 visual, so it pairs with the cleanup-tax hooks. If a Territory-1 hook wins, we'd open on a long-form frame like "hands off the keyboard" instead, so the first frame and the first line pull together.

The reel

The finished asset, 46 seconds. The footnote to the thinking above, not a replacement for it.

07Experiment · one variable, real signal

One narrative. Hook as the only variable.

A hook test only teaches you something if one thing changes. We hold the narrative, visual and CTA constant and swap only the opening, so any difference is the door, not the edit.

Hook rate Did the hook earn the stop in the first three seconds?
CTR Did the story make them want to look closer?
Trial conversion rate Did it pull the right curiosity, not just any curiosity?
CPA Can it acquire a customer economically?
Illustrative
Hook (door)Hook rateCTRTrial conversion rateCPA
Identity32%1.7%9%$37
Threshold37%2.1%13%$26
Confession40%2.5%13%$25
Insight — winner43%2.8%15%$21
Experiment35%1.9%11%$30
Curiosity44%2.3%8%$38

How we read it. Getting people to stop is easy. Getting them to sign up is what matters. Curiosity had the most stops but the fewest sign-ups. Insight got the cheapest sign-ups, which proves the cleanup angle was right. Threshold came a close second, so we test more long-form hooks next.

08Learn · the loop

The winner writes the next batch

A win isn't the finish line. It's the brief for the next test.

Evidence Tensions Personas Territory Hooks Visual Experiment Learn
What the test taught us
The cleanup angle converts. Insight won on cost per sign-up. People will pay to stop editing what they dictate.
A high hook rate can lie. Curiosity stopped the most thumbs but converted the fewest. Stops are cheap; sign-ups are the signal.
Long-form is a second door. Threshold placed second on cost per sign-up without being led on. Territory 1 may be a market of its own.
The next batch
Scale the winner. More hooks and angles inside the cleanup-tax door, before it fatigues.
Give long-form its own test. A dedicated batch for the long-form threshold, to see if the runner-up becomes a winner.
Cut, then reload. Retire the low-converting framings. The winner's numbers and the comments it draws become the next round's evidence, and the loop runs again, sharper.

This is what working together looks like.

Start a project

← Back to all teardowns