Guide

Dictation For Non-Native Speakers

If dictation works less well for you than the reviews suggest, the reason is probably what the model was trained on rather than anything about how you speak.

Guides

The two situations, and they are different

Writing in your second language. You are dictating English, or German, or Dutch, with the accent of your first language. The model is working in a well-resourced language and struggling with the pronunciation.

Writing in your first language, which is thinly resourced. Here the model has less material for the language itself, and everybody using it has the same experience regardless of accent.

Both feel like "dictation is worse for me", and they have different fixes. The first is helped by capacity and vocabulary; the second by the passage of time and by choosing tools that cover your language properly at all.

Training material, not your pronunciation

A speech model recognises what it has heard a great deal of.

Training material for any language is dominated by speakers of that language who learned it first, recorded in broadcast conditions, using standard registers. Second-language speakers are present in that material but in far smaller quantities, and the particular way any given first language shapes a second one is a smaller quantity still.

So a model tends to do best on the accents it saw most, which is not a judgement about clarity. Plenty of second-language speakers are easier for a human to understand than many first-language speakers with strong regional accents, and the model still finds the regional accent easier, because it heard more of it.

Knowing that helps on its own. People conclude their English is not good enough, when what is actually happening is a statistical property of a data set.

Words your first language keeps borrowing

A specific failure that stays invisible until you see it.

Most languages have absorbed English technical vocabulary, and speakers switch into it mid-sentence without noticing: the name of a meeting tool, a file format, a piece of jargon from work. The sentence is Dutch or Spanish or Polish, and three words in it are English, pronounced the way your language pronounces them.

That is genuinely hard for a recogniser. The sentence is identified as one language, and inside it are words belonging to another, spoken with an accent belonging to neither.

The fix is the dictionary, and it is a job you do once. Add those borrowed words, spelled the way you want them written, and the worst category of error in bilingual working life stops being a category.

Five adjustments, and two to skip

Use a larger model. This matters more for you than for a first-language speaker. Extra capacity is precisely what handles pronunciation the model saw less of, so the difference between model sizes is larger in your case than in the reviews you have read. LocalType offers 60, 190 and 539 MB, and if your phone can hold the largest, it is worth the wait.

Add your vocabulary aggressively. Names from your first language, places, family names, the words you use daily that a general model has never encountered in any accent. Those additions stay on the handset and are given to the speech model as context, which counts for more when the sound itself is being read less confidently.

Speak in complete sentences. Neighbouring words do more work when the sound is harder to interpret, so a full sentence recovers far more than an isolated word. Of everything on this list it is the one that costs nothing and applies to every message you send.

Do not change how you speak. Exaggerating towards an imagined neutral accent produces something the model has heard even less of. Your ordinary voice is the one it has the best chance with.

Keep the phone in one position. Hold it where you would hold it for a phone call, and leave it there from the first word to the last, rather than moving it around while you think.

Two things people reach for that work against them.

Slowing down dramatically. Models are trained on speech at ordinary speed. Very slow, over-articulated speech is unusual input and accuracy drops.

Switching to a different language setting to "get closer". Dictating your second language while the system expects your first, or the reverse, produces confident nonsense and nothing like a better approximation.

Where automatic language detection helps you

If you move between languages during a day, understanding the mechanism changes how you work.

LocalType identifies which of eighteen languages you are speaking from the recording, leaving no menu to set, so a message in your first language followed by one in your second needs nothing from you in between.

Detection works on the recording as a whole, so a sentence that genuinely mixes two languages has to be assigned to one. And short utterances give it little to work with, so a two-word reply is where it is most likely to choose wrong.

Speaking full sentences fixes both, the same advice as everywhere else, for the same underlying reason.

None of it needs a connection. Recognition happens on the handset from a model kept in storage only the app can reach, and the app requests internet access for exactly one purpose, collecting the model you chose. That has a practical edge for anyone writing to family abroad on a foreign SIM: dictating consumes no data allowance, and no signal is required to do it.

There is no account, advertising is absent, no behavioural tracking runs inside the product, and no archive of your recordings is kept. Android's cloud backup is deliberately denied the app's data, which is also why the vocabulary you add stays on the phone you added it to. In a password field the microphone key is not offered.

Writing to people who will judge the writing

This changes how much checking is appropriate.

Anybody writing professionally in a second language is already conscious that errors get read as a verdict on competence, not as typos. A dictation mistake in that situation costs more than it does for a first-language speaker, who gets the benefit of the doubt.

The practical response is not to avoid dictation. It is to separate the passes: dictate freely to get the content down, then read it once deliberately before sending, paying attention to names and to any sentence containing a negation.

Negations are worth the extra second because a dropped "not" reverses a sentence without leaving any trace that something went wrong, and it is the one error a reader will never assume was the software's fault.

What to expect

Results will probably be somewhat below what a first-language speaker gets, on the same phone, with the same settings. That gap is real.

It is also usually smaller than people expect once the adjustments above are in place, and much smaller than the gap between dictating and thumb-typing on a phone, which is the comparison that actually matters.

Accuracy is also not identical across the eighteen languages LocalType covers, because training material is not evenly distributed between them either. Widely spoken languages come out best, and that is the same effect one level up.

So: choose the largest model your phone will hold comfortably, and spend an evening adding the names you write most often, in both languages. Then judge it after a week of ordinary messages rather than an afternoon of testing it deliberately, because the two are not the same material.

Keep a note of the words you have to correct during that week. If the same three names keep turning up on the list, they belong in the dictionary, and that is a fix you make rather than one you wait for.

Get LocalType for Android

Private voice typing that runs on your phone, with speech recognition on the device itself.

Get it on Google Play