Training material is the whole first answer
A speech model learns by being shown recorded speech alongside the correct text. The more it has seen of a language, the better it handles it.
That material is wildly unevenly distributed. English has a colossal amount of transcribed audio available: broadcasts, lectures, subtitles, audiobooks, public archives. A language with ten million speakers and less digitised media has a small fraction of that, and a regional language may have almost none.
The model is not making a judgement. It is reflecting what it was given, and what it was given reflects which languages produce the most transcribed audio, which is largely a matter of population and how much of the internet happens in that language.
This is the reason for the pattern people notice, and why the gap narrows slowly. No new release is going to close it in one step.
Not all speakers of a language are represented equally
Within a single language, the same problem repeats at a smaller scale.
Training material skews towards broadcast speech, which skews towards standard accents, prepared delivery and educated registers. Regional accents, dialects, younger speech, older speech and second-language speakers are all present in smaller quantities.
So two people speaking the same language can get noticeably different results, and the difference has nothing to do with how clearly either of them speaks. It is about how much speech resembling theirs was in the pile.
Knowing that saves you from blaming yourself or the microphone.
Some languages are structurally harder
Beyond data volume, a few properties genuinely make recognition more difficult.
Compound words. Languages that glue words together produce very long forms that appear rarely in any training set. Dutch and German do this constantly, and a compound the model has never seen has to be assembled from pieces, since there is nothing to recognise.
Rich inflection. If a noun has fifteen forms and a verb forty, each individual form appears less often than an English equivalent would. The same amount of material spreads more thinly.
Tone. In tonal languages, a change in pitch changes the meaning of a word, not just the emphasis, and that adds a dimension the acoustic model has to get right.
Writing systems. Languages written without spaces between words require the system to decide where words begin, which is a separate problem from hearing them. Languages with multiple scripts add another.
Loanwords and code-switching. In many languages, everyday speech is full of English technical terms. That mixture is common in life and rare in clean training data.
None of these are insurmountable, and all of them mean a given amount of material buys less accuracy than it would in a simpler language.
Why a multilingual model is a compromise
Most on-device dictation now uses a single model that covers many languages. Separate models for separate languages are the other way to do it. That choice has consequences.
The advantage is real: one download, no packs to manage, and the ability to work out which language you are speaking from the recording itself, with no menu involved. For anyone who lives in two languages, that is the difference between using dictation and not.
The cost is that capacity is shared. A model of a given size covering eighteen languages devotes less of itself to each than a model of the same size covering one. That is the trade behind the observation that a dedicated single-language tool sometimes does better on its one language.
There is also a benefit people do not expect. Languages that are related can help each other, because patterns learned from one transfer partly to another. A smaller language sitting near a well-resourced relative often does better in a multilingual model than its own data volume would suggest.
What actually improves your results
Three things, in order of effect, and all three are in your hands, not ours.
Speak in complete sentences. A full sentence does more work in a poorly-resourced language, not less, because the acoustic evidence is weaker and the surrounding words have to carry more.
Add your own vocabulary. Proper nouns are the worst case in every language and the best case for a personal dictionary. LocalType keeps such additions on the handset and hands them to the speech model as context.
Use a larger model if your phone can hold it. Capacity helps most exactly where the language is thinly represented. The difference between the model sizes is therefore far more noticeable in some languages than it is in English.
Reducing background noise helps too, and it helps more than it would in a language the model knows well, since there is less redundancy to fall back on.
What will and will not change
Will change: the amount of training material available, slowly, as more transcribed audio in more languages accumulates. Model architectures also continue to get more efficient, so a phone-sized model now represents more than it once did.
Will not change: the structural properties. A language with heavy inflection will always spread its material more thinly than one without. Tonal distinctions will always be an extra dimension.
Will not change quickly: the ordering. English is not about to stop being the best-served language, and any product claiming uniform accuracy across a hundred languages is describing an intention.
What this means when you compare products
Comparisons written in English tell you very little about your language. A review praising accuracy is describing the best-resourced case, and the ranking between products can be entirely different in a language with less material behind it.
So test in the language you will actually use, on the kind of sentences you actually write. Do not trust somebody else's verdict. Ten minutes with your own speech is worth more than any published comparison, and it is the only way to find out whether a given product covers your case.
The same applies to the number of supported languages on a product page. Supporting a language and being good at it are different claims, and only the first one appears in marketing. A tool listing a hundred languages and one listing eighteen may be equally good at yours, or not, and the list will not tell you.
How LocalType handles this
Eighteen languages from one multilingual model, identified from the recording as you speak, without you picking one, and three model sizes so you can spend capacity where you need it.
It says the same on the languages page: accuracy is not identical across those eighteen, because training material is not. Widely spoken languages come out best, and you would notice it by your third message anyway.
Everything else about the app is unaffected by which language you speak. Recognition happens on the handset from a model file no other app can open, the network is contacted once to collect the model, there is no account, no advertising, nothing recording your usage, and no stored audio.
If your language is one of the eighteen and the results disappoint you, the order to try things is the one above: full sentences, your own vocabulary, then a larger model. In a thinly-resourced language those three matter considerably more than they do in English.