Explained

Dictation vs Voice Assistant

Both involve talking to a phone, and there the similarity stops. One is waiting for its name. The other only listens when you tell it to.

Guides

Why people confuse them

The confusion is understandable.

Both are triggered by speaking, so they feel like the same gesture. Some phones put an assistant button next to or inside the keyboard, which blurs the boundary further. And the same word, "voice", is used for both in every settings menu ever written.

The confusion has a real cost. People who would happily dictate avoid it because they believe it implies an always-on microphone, and people who worry about assistants sometimes disable dictation instead, which changes nothing about the thing that concerned them.

The two designs part company in exactly one place. A voice assistant is always available, so something has to be listening for the moment you want it. Dictation is not available until you ask for it: you press a key, the microphone opens, you speak, it closes.

Almost every worry people attach to talking to phones belongs to the first design and not to the second.

Inside a wake word, and inside a key press

Behind the phrase "always listening" is a mechanism less sinister than it suggests, and also less trivial than manufacturers imply.

A small detector runs continuously, comparing incoming sound against one short pattern: the wake phrase. It is deliberately tiny, it recognises effectively nothing else, and it exists to answer one question, and the question is whether those particular syllables just occurred. Everything else it hears is discarded as it arrives.

When it thinks it heard the phrase, the real system starts, and that is the point at which a proper recogniser gets involved.

The detector genuinely does have to process sound all the time, so nobody should call that concern imaginary. And it is also not transcribing your conversations, because it cannot: it is a pattern matcher for one phrase, not a speech recogniser.

False triggers are the awkward part. Something that sounds close enough starts the assistant, and then a fragment of whatever was being said is handled as though you had asked for it. This is also the failure people notice, because it happens in front of them, when a screen lights up in the middle of a conversation that had nothing to do with a phone.

Dictation has no detector, because there is nothing to detect.

Android shows an indicator whenever any app has the microphone open, so you can see the state rather than infer it. When you stop, it closes again.

None of that is a privacy feature bolted onto the design. It is simply what a keyboard is: you were already touching the screen to write, so there is no reason to listen for a summons.

The practical result is that dictation has no ambient state at all. Between the moment you stop speaking and the next time you press the key, nothing about the feature is running, and there is no equivalent of a false trigger, because nothing exists to be falsely triggered.

What each one is for

Separating them helps, because plenty of people use both and get more out of each by knowing which is which.

An assistant answers and acts. What is the weather, set a timer, call someone, play something, turn the lights off. It needs knowledge of the world and access to your services, so running entirely on the handset is mostly out of reach.

Dictation writes down what you said. That is the whole feature. It does not answer, look anything up, summarise or decide. Sound in, words out, into the field where your cursor was.

If you use both, the split falls where the output goes. Anything that has to happen in the world is the assistant's job. Anything that has to end up as text somebody will read is the keyboard's. Asking an assistant to send a message is the one place they overlap, and it is also the one place where the text leaves before you have seen it.

The narrowness is the reason dictation can run locally at all. Writing down a sentence you deliberately spoke is a bounded task with everything it needs already present. Answering an arbitrary question is not.

Why the distinction matters for where processing happens

The two categories have gone in different directions for a technical reason.

An assistant has to reach a server for almost any question you would actually ask. The knowledge is not on your phone, your calendar and music and messages live in services, and the answer depends on the state of the world right now. Local processing cannot solve that no matter how capable the hardware becomes.

Dictation has the opposite shape. Everything it needs is already present: your voice, and a model that turns sound into words. Nothing about the task requires knowing anything beyond the sentence you just said.

So on-device dictation exists as a serious product category, and on-device assistants largely do not. The narrowness people sometimes see as a limitation is exactly what makes it possible to keep the whole thing on the handset.

Dictation, with no assistant in it

LocalType sits on the second side of every one of those distinctions.

The microphone opens when you press the key on the keyboard and closes when you stop. There is no wake word, nothing running in the background, and nothing listening between dictations. It does not answer questions, write text on your behalf or act on what you said. It writes down your words and stops.

The recognition happens on the phone, from a speech model kept in storage only the app can read, so the audio is not transmitted anywhere while the microphone is open either. Internet access serves exactly one purpose, fetching the model you choose, and nothing calls home after that. There is no account, no advertising, no behavioural tracking inside the product, and no archive of your recordings. The app's data is deliberately excluded from Android's cloud backup, and in password fields the microphone does not appear at all.

Three model sizes, at 60, 190 and 539 MB, let you set speed against accuracy, with 4.1 seconds needed for 6.9 seconds of speech at the middle setting on the 4 GB test handset. Eighteen languages are covered by one multilingual model, identified from the recording each time, and no menu is involved. Words you add yourself stay on the handset and go to the model as context, so your own names arrive spelled correctly. The rest of the time it is an ordinary keyboard with autocorrect, layouts and emoji.

What to check on your own phone

Two minutes and you will know exactly what is running.

Look at which apps have microphone permission, and separately at whether a wake-word feature is enabled in your phone or assistant settings. Those are different switches and turning one off does not affect the other.

Watch the microphone indicator during ordinary use. If it appears while you are not deliberately speaking to something, that tells you more than any explanation.

And check what happens when you use a keyboard's dictation: the indicator should appear when you press the key and disappear when you finish. If it stays on, look harder.

Then one last check, because it settles the question most people arrive with. Turn the wake word off and dictate anyway. If dictation carries on working, the two were never wired together, and you have separated the setting that bothered you from the feature you wanted to keep.

Get LocalType for Android

Private voice typing that runs on your phone, with speech recognition on the device itself.

Get it on Google Play