Same principle, different shape
FUTO Voice Input turns speech into text on the Android device itself, with nothing going to a server. It is open source, paid for once instead of monthly, and needs no account. If you have got this far you almost certainly know that, and you almost certainly think the idea is correct.
So does LocalType. The difference is not the principle but the packaging: LocalType is the keyboard, with dictation as a key on it, rather than a recognition service you attach to a keyboard of your choosing.
For the record: LocalType is a separate product, not affiliated with FUTO or its developers.
Two reasons to stay with FUTO
For a decent share of the people reading, the comparison ends right here.
If being able to read the source matters to you, FUTO Voice Input is the better fit and LocalType cannot compete on that point. LocalType is not open source. Somebody who wants to check what a microphone-adjacent app does, rather than be told, has one honest option of the two.
If you have a keyboard you are not prepared to give up, the same applies. FUTO's design lets the voice half plug into a keyboard you already know, and years of muscle memory and an accumulated dictionary are a real thing to weigh against a tidier install. Their own site says you can pair their voice input with a different keyboard, and that flexibility is the point of building it that way.
Either of those and you already have your answer. The rest is for the reader who wants the keyboard and the dictation to be one product.
One app, or two that have to agree
FUTO splits the work. Voice Input is one app, FUTO Keyboard is another, and the voice app plugs into Android's speech recognition service so a compatible keyboard can call it. Each half can be swapped out without touching the other.
LocalType takes the other road. One install, one settings screen, one product that owns both halves of the job. Nothing to pair, no compatibility question, and no situation where dictation stops working because the two pieces disagree about who is handling the microphone.
That shows up in the small moments rather than the feature list. Because the microphone key belongs to the same app as the letters, dictating is not a mode you enter. You are writing a message, you reach a sentence you would rather say than thumb out, you tap the key, you say it, and you go back to typing the next word. The layout does not shift under you and no other app takes over the screen in between.
Set a language and it moves the keyboard layout, the dictionary and the dictation language at the same time, because one product is deciding all three. Not splitting the job in two is what makes that possible. It works wherever a keyboard opens, on phones and tablets running Android 8.0 or newer.
Neither arrangement is the clever answer. They are different bets about what people want, and you already know which one you are.
The part where we simply agree
On the question this whole category exists for, there is no daylight between the two products. Your voice is turned into text on the hardware in your hand. Neither of us wants an account. Neither of us needs a connection once the model is in place.
The housekeeping follows from that stance: the model lives in the app's own storage, app data is kept out of Android's cloud backup, the microphone switches itself off in password fields, and words you add to the dictionary stay on the phone. No archive of your recordings is kept anywhere. There is no advertising, nothing inside the product tracks what you do with it, and internet access is requested for exactly one purpose, fetching the model you chose.
None of that is a point scored against FUTO. It is the same conclusion reached separately, not a difference between us.
Choosing how hard your phone works
Where the products do diverge again is what you are allowed to tune.
Three speech models ship with LocalType, and you pick which trade you want. Compact is 60 MB and returns text fastest, which matters most on an older phone. Standard is 190 MB and is what the app suggests unless you say otherwise. Large is 539 MB and reads accents and background noise best, at the cost of making you wait: on the test phone, one with 4 GB of memory, Standard converted 6.9 seconds of speech in 4.1 seconds, while Large takes roughly twice as long as you spent talking.
Only one model is kept at a time, so moving up or down a size costs you nothing in storage beyond the download itself. The sensible way to use that is to start on Standard for a week of ordinary writing, then move in whichever direction annoyed you: up if you kept correcting words, down if you kept waiting.
Eighteen languages from one download
Recognition covers eighteen languages, all from a single multilingual model. There is no pack to fetch per language and no menu to visit before switching, because the model works out which of the supported languages you are speaking from the recording itself. Start a message in Dutch, answer the next one in English, and it follows.
Results are not equally strong in all eighteen. Training material is unevenly distributed across languages, and the widely spoken ones fare better. Every speech model has this problem, ours included, and knowing about it helps before you judge any of them on a language they barely saw.
A keyboard first, with your own words in it
This is where owning both halves earns its keep. LocalType has to be a keyboard you would tolerate on the days you never dictate at all: autocorrect and next-word suggestions from a dictionary matched to the language you set, QWERTY and QWERTZ and AZERTY where each is the national standard, a complete emoji set sorted with recent picks first. All of it runs with the connection switched off, exactly like the dictation.
Every dictation tool falls over on the same words: surnames, product names, the jargon of whatever you happen to do for a living. Nothing trained on general speech has met your colleague Sjoerd or the internal system your team refers to by an acronym.
You can add those words yourself, and because one app owns both halves of the job they do double duty. The list feeds the keyboard's suggestions while you type, and it is handed to the speech model as context while you dictate, so a name you have added starts coming out spelled correctly instead of approximated. Those words stay on the phone; they are not uploaded to a LocalType service, because there is no such service to upload them to.
One thing to know before you invest an evening in it: the list does not travel. App data is deliberately excluded from Android's cloud backup, so a new phone starts with an empty vocabulary and you build it up again from the twenty or thirty names you use most. Refusing to copy your data off the device carries a cost, and this is where you pay it.