Three products, one name
The phrase covers tools that barely resemble each other, which is why the search results feel scattered.
Transcription services take a recording, a meeting, an interview, a voice memo, and give you a document afterwards. The work happens after the fact, usually on a server, and the output is a file you can search and share. Good for documenting; useless for writing a message.
Dictation apps give you a window with a microphone. You speak, text appears in the app, and you copy it out to wherever it was meant to go. Good for long passages into a document; irritating for a two-line reply.
Keyboards with voice input put a microphone key on the keyboard itself, so the words land directly in whatever field you were already typing in. Good for everything short and frequent, which is most phone writing.
Three questions settle which one you are after, and answering them honestly takes less time than reading one review.
Do you want a document at the end? Then you want transcription, not dictation. If the value is being able to search what somebody said in a meeting last spring, no keyboard replaces that and none of them will try.
Are you producing long text into one place? A dictation app is fine, and the copy step only happens once, at the end, when the writing is already done.
Are you writing lots of short things in different apps? Messages, replies, notes, search boxes, form fields. Then you want a keyboard, because opening a separate app to say one sentence is more effort than typing the sentence.
Most people asking this question turn out to want the third kind and did not know it was a category with its own name.
Turn the network off and watch what happens
Once you know which type you need, one technical choice remains, and it has more practical consequences than any feature list on any product page.
Speech recognition runs either on a server or on your phone.
Server-side means your audio is uploaded, processed elsewhere and returned as text. The models can be very large, so accuracy on hard material is strong, and a great many languages are supported. It also means dictation stops when your connection does, and your voice is transmitted every time you use it.
On-device means a model file sits on your handset and does the work there. Nothing is uploaded, nothing waits for a network, and a plane or a tunnel is irrelevant. The limits are your hardware: a smaller model than a server would run, some storage occupied, and fewer languages covered.
Marketing will not tell you which one you have, and in fairness it is not required to. The behaviour will. Switch off mobile data and Wi-Fi, open a message and try to dictate a sentence. Either the words appear or they do not, and there is no third answer and nothing to interpret.
Why phone results feel worse than the demos
This disappointment comes up often, and the explanation is straightforward.
The accuracy figures people quote online usually come from large models running on desktop hardware, transcribing clean recordings that were made carefully. Your phone is doing a harder job: a smaller model, live speech with no chance of a second take, and whatever room you happen to be standing in at the time.
Expect a gap. It narrows with the practical things, holding the phone closer, speaking in whole sentences instead of fragments, adding your own names and jargon where the app allows it, and it does not close entirely on a handset.
The useful comparison is not the demo. It is the message you would otherwise have thumbed out at a bus stop, half of it autocorrected into something you did not write.
LocalType is the third kind
It is an Android keyboard with a microphone key, running the recognition on the phone.
Practically, that means the text arrives in the message, note, browser or form you were already in, with nothing copied and the clipboard untouched. It works anywhere a keyboard opens, on phones and tablets running Android 8.0 and newer, and it is a normal keyboard the rest of the time: autocorrect and next-word suggestions from a dictionary matched to your language, the standard layouts including QWERTZ and AZERTY, a full emoji set, and your own names and jargon added by hand, which the speech model is then given as context.
On the where-does-it-happen question there is only one answer in the app. The speech model is a file in storage only LocalType can read, and no cloud path exists to fall back on. Internet access is requested for exactly one purpose, fetching the model you choose, and never after that. No account exists, no advertising, no behavioural tracking inside the product, no archive of your recordings, and the app's data is deliberately excluded from Android's cloud backup. In password fields no microphone key is offered.
Three model sizes, at 60, 190 and 539 MB, let you set the speed and accuracy trade yourself. The middle one needed 4.1 seconds to write out 6.9 seconds of speech on the 4 GB test handset, and the largest takes roughly twice as long as you spent speaking. Eighteen languages are covered by a single multilingual model, and it works out which one you are speaking from the recording itself, with no menu involved anywhere.
Three things it does not do, so nobody installs the wrong thing. It does not transcribe recordings: point it at an hour of audio and nothing happens, because it works on speech as you produce it. It does not translate, so speak German and you get German text. And it is not an assistant, since it writes down what you said and there is no second half where it acts on it.
Four questions before you install anything
Whichever category you settled on, these are the ones that decide whether you keep the app past the first week.
Does it need an account? Server-side products usually do, because somebody has to be billed for the processing. Local ones often do not.
Does it work in the apps you actually use? A keyboard does by definition. A dictation app means copying and pasting every single time, which is fine once a day and wearing twenty times a day.
Which languages, and how does switching work? If you write in more than one, find out whether that means downloading a pack per language and choosing before you speak, or whether the app works it out from what it hears.
What happens to the audio? Not what the marketing says. What the app does when the network is off.
"Free" covers three different arrangements
Check this before comparing anything on price, because the same word describes three unrelated business models.
Free with advertising means the product is funded by showing you things, which usually implies a profile of some description sitting behind it.
Free as in open source means somebody built it and published the code, often without any funding at all. Excellent value, and frequently narrower in scope or in languages than a commercial product, for reasons that have nothing to do with skill.
Free as a front end to a paid service means the app itself costs nothing and the transcription is billed per minute, or capped per month, or subsidised while the service is still growing its numbers.
The word on the listing tells you which of the three only rarely, so look at what funds the thing before deciding it is the cheap option. That, and the network test above, will tell you more about an app in two minutes than the whole of its store page.