Handy is a desktop app, this is the phone answer
Handy (handy.computer) is a desktop app. It runs on macOS, Windows and Linux.
If you want the same thing on Android, LocalType is it. It is a keyboard with a microphone key, and the speech recognition runs on the phone itself. You speak, the phone turns it into text, and your audio does not need to be sent to a transcription service to come back as words. No account, no cloud transcription, no upload.
There is no affiliation between LocalType and Handy or its maintainers; they are separate products. Handy is mentioned here because it is the closest thing on the desktop to what LocalType does on a phone, and because people who use one usually want the other.
The two side by side
| Handy | LocalType | |
|---|---|---|
| Platform | macOS, Windows, Linux | Android 8.0 and newer |
| Shape | Desktop app with a shortcut | Keyboard with a microphone key |
| Where recognition runs | On your computer | On your phone |
| Account required | No | No |
| Works without a connection | Yes | Yes, once a model is downloaded |
| Price | Free, funded by donations | Free trial, then a yearly subscription |
| Source | Open source | Not open source |
Two rows in that table are the reason to compare them at all, and they are the two that match. The rest is what changes when the machine changes.
A shortcut on a laptop, a key on a keyboard
Handy is free and open source. You hold a keyboard shortcut, you speak, and the text goes into whatever field you were already in. It runs Whisper models locally with GPU acceleration where the hardware allows, alongside a CPU-friendly alternative that detects the language by itself. Nothing is uploaded, and it is funded by donations rather than a subscription.
Once you have dictated for a while with no account, no upload and no waiting on a server, cloud voice typing on a phone feels like a step backwards. That is the part people get attached to, and it has nothing to do with the interface.
LocalType is the same principle on a different machine, and the shape is different because Android is different. On a desktop a global shortcut is the natural hook, since any window can be interrupted. On Android the keyboard is replaceable, and every app that accepts text opens it. So the microphone belongs on the keyboard, right next to the keys you were already using.
That removes the part that makes phone dictation annoying. No separate app to open, no transcription screen, nothing to copy out and paste back. You type, you tap the microphone, you speak, and you carry on typing. It works in messaging, email, notes, your browser and forms, because it works wherever a keyboard opens.
A file that used to live on somebody's server
A speech model is a file. A fairly large one, trained to recognise speech, that the app loads and runs on your device. Cloud voice typing keeps that file on a server you never see. Out of sight, it never crosses your mind. Running it locally means it lives on your phone instead, taking up storage and memory while it works.
LocalType offers three, and you choose:
- Compact, 60 MB. The fastest, for older phones or phones with little memory to spare.
- Standard, 190 MB. What LocalType suggests by default, and what suits most phones.
- Large, 539 MB. The most accurate, for phones with room to spare, when you would rather wait a moment than correct a word.
On the test device, a phone with 4 GB of memory, Standard turned 6.9 seconds of speech into text in 4.1 seconds. Large takes roughly twice as long as you spent speaking. Your own phone will differ.
A phone has less memory and less cooling than a laptop, so the models that fit on it are smaller than the largest ones a desktop can run. Expect good transcription, not desktop-identical transcription. LocalType keeps one model on the device at a time, and downloading another replaces it, so swapping does not cost you extra storage.
Local, item by item
- Speech recognition runs locally on your Android phone.
- The app asks for internet access for one thing: downloading the speech model you choose. After that, transcription no longer needs a connection.
- There is no sign-up, no login and no profile.
- There is no advertising in LocalType.
- There is no behavioural tracking inside the product.
- LocalType app data is deliberately excluded from Android's cloud backup.
- The speech model sits in the app's own private storage.
- Voice input is automatically disabled in password fields, so a password cannot be dictated by accident.
- Words you add yourself, names and jargon, stay on the device and are given to the on-device model as context.
- No archive of your recordings is kept. The audio becomes text in the field you were typing in, and that is the end of its life.
There is no private mode to switch on. Local speech recognition is not a setting, it is how the app works, and there is no cloud mode for it to fall back to.
The visible half of that list is what happens with the connection switched off. Voice typing keeps working on a plane, on the underground, abroad without roaming data, in a basement, or anywhere the signal drops in and out. Cloud voice typing tends to fail quietly in those situations: it waits, times out, or loses what you said. Local transcription does not have that failure mode. The one step that needs a connection is the first download of a speech model, and doing it on Wi-Fi before you need it is the whole of the setup advice.
Eighteen languages, without switching first
Eighteen languages are covered, all by one multilingual model, so there is no separate download to fetch for each one. LocalType also works out which of the supported languages you are speaking, per recording, so it follows you when you switch language between messages.
Accuracy is not identical in every language. Speech models are trained on uneven amounts of material, and widely spoken languages generally come out best.
And it is a keyboard the rest of the time
You can use LocalType all day without ever touching the microphone. Autocorrect and next-word suggestions, a dictionary per language, QWERTY, QWERTZ and AZERTY where those are the national standard, your own words added to the dictionary, and the full emoji set. All of it works with the connection switched off, the same as the voice typing.
The words you add are worth ten minutes on the first evening. Surnames, street names, the products and acronyms of whatever you do for a living: they go into the typing suggestions and to the speech model at the same time, which is the one lever that reliably moves recognition quality in your favour.
Desk or pocket, free or paid
If you are at a desk, Handy is free and open source, and it is a good answer. Use it.
LocalType is a paid product after a free trial, and it is not open source. What you get for that is dictation on the device you actually carry, built and supported as a product, with the same principle underneath: the recognition happens where you are, not on someone else's server.