How to

How To Proofread Dictated Text

A typo looks wrong. A dictation error looks completely fine, which is exactly why it survives all the way to somebody else's inbox.

Guides

Why dictation errors are harder to see

Typing produces misspellings. Those are visually obvious, they get underlined, and your eye catches them without effort.

Dictation produces the wrong word, spelled perfectly. Nothing is underlined, nothing looks unusual, and the sentence often still parses. Your brain, helpfully, reads what you meant and skips over what is there, because you were the one who said it thirty seconds ago and the intended version is still in your head.

That combination is why people who are careful proofreaders of typed text miss dictation errors constantly. The usual technique, scanning for something that looks wrong, does not work when nothing looks wrong.

The message that already went out

This is the one that actually costs you something.

The failure mode particular to dictation is a message that reads perfectly and says the opposite of what you intended. A missing "not" in a sentence about a deadline. A price with the wrong digit in it. A colleague's name turned into an ordinary word. Nothing about any of those prompts a question from the recipient, because nothing looked broken on their screen. They simply get acted on.

When it does happen, and it will, the fix is the same as for a typo: a short follow-up saying what you meant. People send those all day and nobody thinks twice about it. The reaction matters more than the error did, and treating a misheard word as a catastrophe is how people talk themselves out of dictating at all.

But a misheard name gets queried by the person reading it, and a missing "not" does not. That asymmetry is the whole argument for spending twenty seconds before you send.

Where the errors actually cluster

You do not need to reread everything. Errors are not evenly distributed, and four categories account for most of them.

Proper nouns. Names of people, places, companies, products. The most likely to be wrong and the most embarrassing when they are, because somebody is reading their own name spelled as a different word.

Numbers and codes. Digits have no context to disambiguate them, so they are guessed from sound alone. A wrong number looks exactly as plausible as a right one.

Short words that change meaning. Negations especially. "Can" and "cannot", "is" and "is not", "now" and "not". A single missed word inverts a sentence and nothing about the result looks damaged.

Homophones. Their and there, to and too, and the equivalents in every other language. Grammatically fine, semantically wrong.

Two of those four are visible without reading a word of the message. Capitals and digits stand out on a screen, which is what makes the routine below quick.

A twenty-second routine

For an ordinary message, this is enough.

Scan for names and numbers first. Not reading, scanning. Your eye can find capitals and digits quickly, and those are the expensive errors.

Read the first and last sentences properly. The opening sets the tone and the closing usually contains the ask. Those two carry most of the meaning.

Look for negations. If the message contains a "not", read that sentence twice.

Then send.

For anything longer or more consequential, add one step: read it aloud. Reading aloud forces you to process the words that are on the screen instead of the words you intended, and it catches inverted meanings and missing words that silent reading slides past. It is slower than scanning, and it is the step to keep when you only have room for one.

One habit that makes all of this shorter: do the checking at the end, not during. Repairing words while you are still dictating switches you between producing and evaluating, and neither goes well interleaved.

Which messages deserve more than twenty seconds

Not every message earns the same care. Most messages do not need careful proofreading, and treating every text like a legal document is how people give up on dictation.

More care: anything with a number that matters, anything going to a client or a stranger, anything with a name in it, anything where a reversed meaning would cause a problem, and anything you cannot edit after sending.

Less care: notes to yourself, messages to people who know you, short confirmations, anything where the recipient will simply ask if something is unclear.

Notes to yourself are a different case. They genuinely do not need proofreading for spelling, but they do need enough detail to still make sense next week, and that is a different kind of checking, arguably a more important one.

Reading on a phone screen is part of the problem

This is a phone activity, and the screen makes it harder.

A message that fills three lines on a laptop fills half a screen on a phone, and you are usually reading it in a hurry, one-handed, in poor light, while walking. Those are bad conditions for catching a subtle error, and they are the exact conditions under which dictation is most useful.

Scroll back to the top. Checking only what is visible misses the beginning, and the beginning is where you were still warming up and the errors cluster. And if the message genuinely matters, wait until you are somewhere still before sending it, which costs a minute and removes most of the risk.

Fix the source, not the symptom

If you find yourself correcting the same word repeatedly, stop correcting it.

Proper nouns you use often can be added to the keyboard's dictionary, which puts them into the running when the recogniser is choosing between options. LocalType keeps such additions on the phone and gives them to the speech model as context, so the name comes out right from then on. An hour of proofreading spread over a month turns into one evening of typing in twenty names.

Other recurring errors point at other fixes. Consistently mangled endings usually mean you are trailing off. Errors clustered at the start of dictations mean you are beginning cold, without giving the system a run-up. Widespread errors in noisy places point at the room, not the app.

The tool does not do the checking. LocalType transcribes. It does not proofread, rewrite, tidy, summarise or flag anything as suspicious. What appears is what it heard, with punctuation and capitals inferred, and the checking is yours.

What it does offer is control over the error rate before you get there: three model sizes, at 60, 190 and 539 MB, where the larger ones handle ambiguity better, and a personal dictionary for the words that fail most. On the 4 GB test handset the 190 MB option wrote out 6.9 seconds of speech in 4.1 seconds.

Recognition happens on the handset from a model file no other app can open, the network is contacted once to collect that model, and no archive of your recordings is kept, so there is nothing to go back and listen to. Proofreading is against the text, not against the audio.

Know that before you build the habit. There is no recording to consult when a sentence reads oddly and you cannot remember what you said, so the decision has to be made from the words on screen. In practice that pushes you towards saying the doubtful sentence again rather than reconstructing it, which is faster anyway.

Get LocalType for Android

Private voice typing that runs on your phone, with speech recognition on the device itself.

Get it on Google Play