A perfect transcript is a bad transcript

While testing our speech model, we caught it doing something almost funny. It transcribed every um and uh we fed it, in order, missing none. When an um happened to start a sentence, it typed “Um.” Capital U. Full stop. Perfect punctuation for a sound that isn’t a word.

Technically, that’s a good transcription. Every sound that left the speaker’s mouth landed on the page. It was faithful to a fault.

For most dictation, it’s also useless. You don’t want a message full of ums just because you said them. That leads somewhere uncomfortable. “Perfect transcription” usually means every word, exactly as spoken. But that’s often not what we want at all.

You talk in drafts

Listen closely to a recording of anyone speaking, yourself included. It’s not neat sentences. It’s “so I think we — no, wait — okay, the thing is…” Real speech is full of restarts, repairs, and placeholders. Talking and thinking happen at the same time. You speak in drafts, and the drafts overlap.

And that isn’t sloppiness. You’ve been listening to speech like that your whole life and never noticed. You don’t hear a friend’s ums until someone points them out — then you can’t stop. Your brain has been quietly cleaning the signal the entire time, discarding the scaffolding and keeping the meaning.

What you want from a transcript isn’t what you said. It’s what you meant to say. And those are different documents.

Something has to decide

Now the awkward question. Our machine hears your um perfectly; we tested it. If that um doesn’t appear on the page, then somewhere between your mouth and your cursor, the system decided to remove it.

Every dictation tool that produces clean text is making editorial decisions. Which fillers vanish. Which restarts get collapsed. Whether “no wait, Tuesday” becomes Tuesday. Most tools present the output as simply “what you said,” as if no judgment were involved. The edit is invisible, so you can’t inspect it, adjust it, or turn it off.

We’d rather show you the editor. In Clabrate, the choice is visible. Raw gives you the speech model’s output, apart from substitutions you defined. Pick an enhancement style and a model cleans up fillers, restarts, and grammar. That model can run on your Mac; cloud enhancement is optional and only used if you configure it.

And the honest option sits right at the top: Raw. Choose it at onboarding or any time after, and no cleanup model rewrites the transcript. Your ums and restarts remain. Sometimes the faithful document is the point — for a quote, a research note, or a journal entry — and a tool shouldn’t quietly overrule you.

The listener you already trust

The reframe I keep coming back to is this: a good dictation tool isn’t a microphone with a keyboard attached. It’s a listener. You already know what good listening looks like because your friends do it every day. They hear your restarts and placeholders, then respond to what you meant.

Nobody calls that lying. It’s comprehension.

The failure mode isn’t editing; it’s hidden editing. An edit you chose is a service. An edit you can’t see or refuse is the tool deciding it knows better. That’s true whether it’s deleting your ums or “improving” a word you meant.

Try the court-reporter test

Record thirty seconds of yourself explaining anything — how you make coffee, what your day was. Then play it back and write down every sound, verbatim, ums and restarts included.

Read the result. That’s a perfect transcript. Now imagine pasting it into an email.

Then notice what your brain did on playback. It heard the meaning straight through the mess, the way it always has. That’s the standard: not a stenographer, but a listener.

The machine finally hears every word you say. The craft is in knowing which ones you meant.

Try the tool this blog is about
Clabrate turns your voice into clean, finished text in any Mac app — private, on-device, no cloud required.
Get Clabrate