"It is always listening"
This is the most common reason people avoid dictation, and it belongs to a different product. An assistant listens for a wake word, and the whole of the concern comes from there.
Dictation opens the microphone when you press a key and closes it when you stop. Between those moments nothing is running, and Android shows an indicator whenever any app has the microphone open, so you can see the state for yourself without taking it on trust.
Turning off an assistant and turning off dictation are separate switches, and people frequently disable the wrong one.
"It is not accurate enough to be useful"
The belief usually dates from an experience five or more years old, when it was true. Modern recognition on clear speech, close to the microphone, in a quiet room, is good. Not perfect, and the errors it does make are specific and predictable: proper nouns, numbers, and homophones.
A realistic comparison is also not against flawless typing. It is against thumb-typing on glass, which produces its own steady stream of errors that autocorrect quietly fixes.
Two things people think they have to do while speaking
Both are wrong, and believing either one makes the results worse.
The first is speaking slowly and over-enunciating. Models are trained on ordinary speech, so exaggerated delivery produces input unlike anything in the training material and accuracy drops. Speak normally, in complete sentences, at your usual pace. What genuinely helps is not trailing off, not restarting halfway through a sentence, and pausing between them.
The second is saying the punctuation aloud. Older systems expected "comma" and "full stop"; modern ones work punctuation out from your rhythm and the shape of the sentence. Learn dictation on the older generation and saying "comma" now gets you the word comma, which is a bewildering first five minutes and sends a fair number of people back to typing before they find out why.
"It works the same in every language"
This one comes from marketing rather than folklore.
Speech models are trained on wildly uneven amounts of material per language. English is the best served by a large margin, and any product listing a hundred languages is describing what it will attempt and not what it does well.
That applies to us too. LocalType covers eighteen languages and results are not identical across them.
"On-device means it is as good as the cloud"
This one is inconvenient for us and true anyway. A server can run a far larger model than a phone can hold. On difficult material, strong accents, background noise, several speakers, that difference is real and measurable.
What on-device gives you instead is that nothing is transmitted, it works without a connection, and there is no account. Those are the reasons to choose it. Raw capability on hard audio is not one of them, and anyone claiming otherwise is not being straight with you.
"Offline dictation needs a new flagship phone"
No. Memory is the constraint here, not age, and the fix is a smaller model.
LocalType runs on Android 8.0 and newer and offers three sizes, at 60, 190 and 539 MB, precisely so that older handsets can use the smallest. On a phone that is short of memory, the smallest model often outperforms the largest in practice.
"It will make you faster at everything"
Here the exaggeration runs the other way. Speaking is much faster than typing for anything longer than a sentence. For very short messages, addresses, numbers, codes and anything with unusual formatting, typing wins, because the fixed overhead of starting a dictation and checking the result is longer than the typing it replaced.
People who get the most out of dictation use it for some things and type the rest, rather than converting entirely.
"A keyboard cannot see what you type if it is private"
A claim some products imply and none can deliver.
Every Android keyboard receives your keystrokes and can read the text around your cursor. That is what an input method is, and it is how autocorrect functions at all. Ours does it too.
The real question is not whether a keyboard can see your typing. It is whether any of it leaves the device, which you can establish by checking whether the app requests internet access and by using it with the network off.
"Dictated text sounds unprofessional"
In our experience of reading both, the reverse is closer to the truth. Typed messages drift towards padding and stock phrases because typing is slow. Dictated ones come out closer to how you would explain something in person: shorter, more direct, more concrete.
What dictation does do is make messages longer than intended, because speaking is easy. Cutting a paragraph before sending is the fix, and not switching back to typing.
"Everybody will hear you"
True, and more limiting than any of the technical objections above.
Dictation is spoken aloud, so it does not work in an open-plan office, a quiet carriage, a library or a room containing the person you are writing about. That rules out a substantial share of the moments when people write on a phone, and no software fixes it.
The constraint is about length, not about dictation as such. A short sentence sounds like half a phone call and attracts no attention. A three-paragraph message announces itself.
People who use dictation most are not braver in public than anyone else. They dictate while walking, driving between places, at home and in the car park, and type the rest.
"You need to learn commands"
Left over from professional dictation software, which had extensive vocabularies of spoken instructions for formatting, navigation and correction.
Phone dictation has none of that. There is a key, you speak, words appear. There is nothing to memorise, no mode to enter, and no manual.
That simplicity is also a limitation: without commands you cannot say "select the last sentence" or "make that a bullet list", so formatting remains a job you do by hand. Simpler to start, less controllable at the end.
Checking our own claims the same way
Everything above is checkable, and so is this. LocalType runs the recognition on the handset from a model kept in storage only the app can reach, and requests internet access for exactly one purpose, collecting the model you chose. There is no account. It carries no advertising, runs no behavioural tracking inside the product, and keeps no archive of your recordings. Android's backup never receives the app's data, and that is deliberate. Move into a password field and no microphone key is offered there.
Turn the network off and dictate a sentence. Open the app's storage permissions and look at what it asks for. Neither takes a minute, and either one settles the question faster than reading another page about it.