The asymmetry nobody mentions
A voice note is cheap to make and expensive to receive.
Making one takes as long as saying it. Receiving one takes as long as listening to it, at your pace rather than theirs, in a place where audio is socially acceptable, with headphones if anyone else is around, in full, with no way to skim to the part that matters.
A forty-second note containing one useful sentence still costs forty seconds. Multiply that across a group chat and the arithmetic gets uncomfortable.
Dictation removes the asymmetry. Your effort is the same as it was, and what arrives is text: skimmable, searchable, readable in a meeting, quotable in a reply.
Where voice notes genuinely win
Voice notes are not a lesser form of communication. Tone that text destroys. An apology, bad news, something affectionate, anything that would read as curt in writing. Voices carry information that punctuation cannot reproduce.
Complicated explanations. Things you would be gesturing about if you were in the room. Directions, a piece of family history, why something went wrong.
When you genuinely cannot look at the screen. Walking somewhere in the dark, carrying something in both hands.
Between people who prefer them. Plenty of relationships run on voice notes happily, and the receiving cost is not a cost if the other person enjoys it.
The rule of thumb most people converge on: information as text, feeling as audio.
What text does that audio cannot
Four things, and they compound over time.
It can be skimmed. A reader gets the point in two seconds and reads the rest if they need it.
It can be searched. Six months later you can find the address, the date, the name. A voice note is invisible to search forever.
It can be quoted. Replying to a specific sentence is trivial in text and impossible in audio.
It can be read anywhere. In a meeting, on a train without headphones, next to a sleeping child.
For anything anybody might need to refer back to, text is not merely more convenient. It is the difference between information existing and not.
The middle option people forget
You can dictate a voice note.
That sounds absurd until you notice what people actually do: they record a rambling ninety-second note because speaking is easy, when the same content dictated would have been four sentences.
Dictation imposes a small useful discipline. Because you can see the text appearing, you notice when you are repeating yourself, and because it arrives as sentences you tend to speak in them.
What you send is usually shorter and clearer than the voice note you would have recorded, at the same cost to you.
Where each one fails
Voice notes fail in a group where half the people cannot listen right now, when the information will be needed again later, when the recipient is at work, and when your message contains an address or a time that has to be accurate.
Dictation fails in a room where speaking is awkward, when the message is emotionally delicate enough that tone matters more than precision, and when the content is dense with names and numbers that will come out wrong.
Neither list is a criticism. They are different tools and the failure modes are the reason to know both.
The practical compromise
Dictate anything with content: times, places, decisions, instructions, anything somebody might need to find again.
Record a voice note when the delivery is the point, and the information is secondary.
And when a message contains both, dictate the facts and record the rest, or simply say the affectionate part in text and accept that it reads slightly flatter than you meant.
What makes dictation practical for this
People default to voice notes because of friction, not because of preference. Recording is one button; dictating used to mean an app, a wait, a copy and a paste.
A keyboard removes that. The microphone is a key on the keyboard that already opened when you tapped the message box, so dictating is exactly as many taps as recording, and the words go straight into the conversation.
LocalType also runs the recognition on the handset from a model kept in storage no other app can open, which matters here for a specific reason: messages get written in tunnels, lifts, aeroplanes and foreign countries with the data switched off. A voice note records anywhere; dictation that depends on a connection does not, and that inconsistency is what pushes people back to recording.
Internet access is wanted for one thing only, collecting the model you chose. Nothing needs registering and advertising is absent. No behavioural tracking runs inside the product, and nothing keeps an archive of what you dictated. A voice note, by contrast, leaves an actual recording sitting in somebody else's chat history for as long as they keep it.
Three model sizes, at 60, 190 and 539 MB, let you trade speed against accuracy; for messaging the smaller two usually feel best. Over a password box no microphone key is offered.
What happens to the recording afterwards
A voice note is a file, and it outlives the conversation. It sits in the recipient's chat history, gets backed up with their phone, and is still there years later unless somebody deliberately deletes it. You have no control over any of that once it is sent, and neither does the recipient really, since most messaging apps keep media by default.
Dictated text is also permanent in the conversation, but it is text: smaller, searchable, and containing no recording of your voice at all. Whatever you said aloud stopped existing the moment it became words.
That is not a reason to panic about voice notes, and it is a reason to think before recording anything you would not put in writing. People are noticeably more candid into a microphone than into a keyboard, and the recording is the part that lasts.
Group chats change the maths again
The cost multiplies with the number of people in the room. A one-minute voice note in a group of eight costs eight minutes of other people's attention, and at least half of them will not be somewhere they can listen. What usually follows is that four people never hear it, two ask what it said, and the information ends up being retyped by somebody else.
Text in a group is read by everyone, at a glance, in whatever situation they are in. For anything logistical, a date, a place, who is bringing what, that is not a preference but a functional requirement.
Text if they need to act on it
If the other person will need to read it, refer back to it, or act on it, send text. If the point is how you sound, send audio.
Most messages are the first kind, and most people send the second because it was easier. It no longer is.