How to

How To Dictate A Long Message

Dictating one sentence is easy. Dictating six paragraphs is a different skill, and the people who find it frustrating are usually doing one specific thing wrong.

Guides

Why long dictations go wrong

Four failure modes, and they compound.

Cut-offs. Some apps limit how long a single recording can be, and you find out by losing the end of a paragraph. Whisper IME, for example, caps recordings at roughly half a minute, which is fine for messages and awkward for anything sustained.

Punctuation collapse. The longer you speak without pauses, the less evidence the system has for where sentences end, so a long delivery produces a wall of text with commas in surprising places.

Losing your place. You cannot see what has been written while you are speaking, so by minute two you have repeated a point, forgotten a point, and lost the thread of the sentence you were in the middle of.

Cursor accidents. On a small screen, a stray tap moves the insertion point, and the next paragraph lands in the middle of the previous one or inside a signature.

None of those is about recognition quality. They are about the shape of the activity.

Finding out whether your own app has a cap takes one attempt. Read a paragraph of something aloud into a notes field, keep going past a minute, and see where the text stops. Better to discover the limit against a page of a novel than against the only version of a message you have already composed in your head.

Work in blocks

The single technique that fixes most of it.

Do not attempt one continuous performance. Dictate a paragraph, stop, glance at the screen, then dictate the next one. Three or four sentences at a time is about right.

Recognition works with cleaner boundaries, punctuation lands better because each block ends with a real full stop, you see the text often enough to keep your bearings, and if something goes wrong you lose one paragraph, not the lot.

It is also the natural way to speak. Nobody delivers six paragraphs without pausing except when reading aloud, and reading aloud is not what you are doing.

Everything you do before you press the key

Three pieces of preparation, in the order they matter.

Decide the shape. List the three or four things the message has to contain, in order, and say the main point first. Spoken messages that build to a conclusion tend to lose the reader long before they arrive at it. If you cannot list the points before you start, the thinking is not done, and dictating will produce a long message that circles the point without landing on it.

Add the names. Long messages contain more proper nouns than short ones, and proper nouns are where dictation fails. Before a long dictation, put in the names it will contain: people, places, products, the acronyms specific to whatever you are writing about. LocalType keeps those additions on the phone and hands them to the speech model as context, so they start arriving spelled correctly. A name that appears in every second paragraph is otherwise a correction in every second paragraph.

Pick the model to match. With short replies the wait after you stop speaking is the thing you notice, and that argues for a smaller model. With long dictation you are looking at the room rather than the screen, so a second or two of processing costs you nothing and every avoided error is an edit saved. LocalType offers three sizes, at 60, 190 and 539 MB, and for sustained work the middle or largest is usually the better choice unless your phone is short of memory. On the 4 GB test handset the middle setting turned 6.9 seconds of speech into text in 4.1 seconds, while the largest takes roughly twice as long as you spent speaking.

Cursor discipline, and leaving errors alone

Two rules for the part where you are actually talking, and the second one is the harder of the two to keep.

Tap into the field and check the insertion point before you begin. In a reply, make sure you are above any quoted text and above your signature. Between blocks, glance; do not settle in and read. You are checking that the text landed where you expected, not proofreading. If you need to insert something in the middle later, place the cursor deliberately and dictate there, because words go where the insertion point is and that is a hazard as often as it is a feature.

The second rule is not to fix anything until you have finished. Stopping to repair a word breaks the rhythm the system was using to place punctuation, so everything after the interruption comes out slightly worse. It also switches you between producing and evaluating, which are different modes, and interleaving them does neither well. Say the whole thing, then read the whole thing, then fix everything in one pass.

The exception is a genuine derailment, where you have said something so wrong that continuing is pointless.

Recovering a paragraph that came out wrong

Two recoveries, since a long dictation has more opportunities to derail than a short one.

The block came out badly. Select that paragraph and delete it whole, then say it again from the beginning of the block. Picking at individual words is slower than re-saying three sentences, and patched paragraphs usually leave the seams visible: half of it in the register you were in at the time, half in the register you switched to while repairing it.

You lost your thread. Stop. Read what is there. Then continue from the last complete sentence on screen, not from where you thought you were. Reconstructing a half-finished sentence you spoke a minute and a half ago is harder than reading the one that actually got written down.

Both come back to the same habit: text on a screen is cheaper to work with than a sentence held in your head. When something goes wrong, look before you speak again.

Where a five-minute dictation actually happens

A long dictation is a long time to be talking out loud, which rules out most shared spaces. That is not a software limitation and no app solves it.

What it means in practice is that long dictations happen in the places where nobody is listening. Walking between appointments. In the car before you go in. On a platform at the wrong end of a station. At the kitchen table on a Sunday. Those are frequently places with a weak signal or none at all, which is a problem if the recognition is happening somewhere else.

LocalType does the work on the handset, using a model held in storage this app alone can reach, so bars on the screen make no difference to a five-minute dictation. Internet access is used for collecting the model you picked and for nothing else after that. No sign-up exists anywhere in it. There is no advertising in the product, no behavioural tracking inside it, and no archive of what you dictated is kept once the words are in the field. Android's cloud backup is held away from the app's data on purpose, which is why a new phone starts with an empty vocabulary. And in a password field the microphone key is not there at all.

So the practical version is a walk. Points listed before you set off, one block per point, a glance between blocks, no corrections until you sit down, and a single read-through at the end while the kettle boils.

Get LocalType for Android

Private voice typing that runs on your phone, with speech recognition on the device itself.

Get it on Google Play