How to

How To Improve Voice Typing Accuracy

Most accuracy problems are not the app. They are distance, noise, delivery and vocabulary, in roughly that order, and all four are things you can change today.

Guides

Start with distance, because it is free

The microphone picks up whatever reaches it, and how much of that is you comes down to physics. Software has little to do with it.

Held at arm's length in a room with any noise at all, your voice arrives as one signal among several. Held near the bottom of the phone at conversational distance, it dominates everything else. That difference alone fixes a surprising share of complaints.

Do not cover the microphone with a finger, which on most phones sits at the bottom edge and is exactly where people grip. And if you are outside, turn away from the wind, because wind noise is loud, broadband and almost impossible for any model to separate from speech.

Then delivery, which is a learnable skill

Dictation rewards a particular way of speaking, and it is not the way most people talk.

Complete sentences. Recognisers use the surrounding words to choose between options that sound alike. One word said alone gives them nothing to work with. So a single name comes out wrong on its own, while the same name inside a sentence comes out right.

Decide before you speak. Trailing off, restarting halfway and thinking aloud all remove the rhythm the system relies on, and punctuation degrades first.

Normal pace and volume. Slowing down dramatically or over-enunciating makes things worse, not better. Models are trained on ordinary speech, so unusually careful speech is unusual input.

Pause between sentences. That is the signal used to place a full stop.

None of this requires a different app, and it usually makes more difference than switching would.

Teach it your vocabulary

The words most likely to be misheard are the ones you use most: colleagues' surnames, client names, products, acronyms, the technical terms of whatever you do.

A model trained on general speech has never encountered any of them, so it substitutes something ordinary that sounds similar. No amount of speaking clearly fixes that, because the correct word is not in the running.

If your keyboard lets you add words, this is the single highest-value change of the lot. LocalType keeps additions on the handset and hands them to the speech model as context, so they start arriving spelled correctly, and they improve the typing suggestions at the same time.

Add them in a batch once rather than one at a time as they go wrong.

Check the language setting

An entire category of "it produces nonsense" is a mismatch between the language you are speaking and the language the app expects.

The failure is confusing because there is no error message. The system does its best to hear the language it was told to expect, and produces confident rubbish.

If you write in more than one language, either set it explicitly before a long session, or use an app that identifies the language from the recording itself. LocalType does the latter across eighteen languages and decides per recording, so switching between messages needs nothing from you. Setting a specific language instead moves the keyboard layout and the dictionary along with it, which is the better option if you will be in one language all day.

One caveat: accuracy is not identical across those eighteen. Training material is distributed unevenly, and widely spoken languages do better.

About accents

Accents are where dictation is least even, and the reason is mechanical.

A model performs best on speech resembling what it was trained on. If your accent was thinly represented in that material, results will be worse, and the cause sits in the training material, not in anything about how you speak.

Two things genuinely help. A larger model has more capacity for unusual cases, so try moving up a size before concluding the technology is not for you. And the vocabulary trick above matters more, not less, because a mispronounced-by-the-model proper noun is your most frequent error.

What does not help is exaggerating your pronunciation towards some imagined neutral accent. It produces speech the model has heard even less of.

Buy accuracy with patience, if the app lets you

Where an app offers a choice of model, that choice is an accuracy control and most people never touch it.

Capacity costs time. A model with more room to represent unusual pronunciations, unfamiliar names and speech competing with noise will get more of them right, and it will take longer doing it. LocalType offers 60, 190 and 539 MB for exactly that reason, with 4.1 seconds needed for 6.9 seconds of speech in the middle setting on the 4 GB test handset.

Match it to what you are dictating. Picking once and forgetting is the mistake. Firing off replies, the wait is the thing you notice. Talking through a long message, your eyes are on the room and not on the screen, and the errors you avoid are edits you never have to make.

The exception is a phone genuinely short of memory, where the largest model can end up slower and no more accurate than the smallest, because it never gets to run comfortably.

Give it a run-up

A small trick that costs nothing and helps more than it should.

Recognisers are weakest at the very start of a recording, before there is any context to lean on. If the first thing you say is the unusual word, an unfamiliar surname, a technical term, a place name, it has the least help available at precisely the moment it needs the most.

Starting with two or three ordinary words gives the system something to settle against before the difficult part arrives. "So the address is" before the street name. "Send this to" before the surname. It feels artificial for about a day and then becomes automatic.

The same logic explains why re-dictating a failed word on its own rarely works. Alone, it has even less to go on than the first time. Say the whole phrase again instead.

What no setting will fix

There is a ceiling here, and no setting raises it.

Several people talking at once. A phone on a table across the room. A café at full volume. Heavy music behind you. In all of those, speech and noise arrive mixed into the same measurements and cannot be separated afterwards, on any device, by any app.

Homophones are the other permanent limit. Where two phrasings sound identical and both make sense, the system is guessing, and it will sometimes guess wrong.

And no recogniser knows what you meant. It knows what reached the microphone, and delivery outranks everything technical for that reason.

Five steps, in order

Get closer to the microphone. Speak in complete sentences and pause between them. Add your own names and jargon. Check the language. Then, if it still is not good enough, move up a model size.

Most people stop being annoyed somewhere around step three.

Get LocalType for Android

Private voice typing that runs on your phone, with speech recognition on the device itself.

Get it on Google Play