How to

Voice Typing In Noisy Places

Noise is the one dictation problem no setting solves, because by the time the sound reaches the microphone your voice and the room have already been added together.

Guides

Why noise is different from other problems

Most dictation errors have a fix. Names can be added to a dictionary, punctuation improves with pauses, wrong languages can be set correctly.

Noise is not like that. The microphone measures air pressure, and it cannot distinguish between pressure caused by you and pressure caused by the espresso machine. Both arrive as one stream of numbers, mixed together, and the model has to work out which parts were words.

Once they are mixed, they cannot be separated. That is a matter of physics and not a software limitation, and it is why the advice here is about the microphone and not about the app.

Not all noise is equally bad

People assume all noise is the same, and it is not.

Other people talking is the worst. Speech competes directly with speech, and a model trained to find words will find theirs as readily as yours. A café full of conversation is harder than a building site.

Music with vocals is nearly as bad, for the same reason. Instrumental music is considerably easier.

Steady mechanical noise is surprisingly manageable. A fan, an engine, a train, air conditioning. It is loud but it is constant and unlike speech, so a model can often see past it.

Sudden noises are worse than loud ones. A door slamming mid-sentence damages that sentence far more than a continuous hum at the same volume.

Wind is uniquely destructive. It is broadband, it strikes the microphone directly and does not travel through the air as sound, and no amount of speaking louder helps.

What actually works, in order

Get closer. By far the biggest lever, and it costs nothing. Doubling the distance between your mouth and the microphone dramatically reduces your voice relative to everything else in the room. Held near your face, you dominate. Held at arm's length, you are one sound among several.

Find the microphone. On most phones it sits on the bottom edge, and that is exactly where fingers go. Covering it while gripping the phone is a common and invisible cause of bad results.

Turn away from the source. Not from the microphone, from the noise. Your body blocks a surprising amount, and phones are directionally sensitive.

Shield it with your hand, cupped loosely, especially outdoors. Crude and effective against wind.

Speak normally, not louder. Raising your voice changes its character, and models are trained on ordinary speech. Getting closer achieves what shouting does not.

Shorten what you say. In a difficult environment, one sentence at a time limits the damage when something goes wrong and gives you a natural moment to check.

Environments, specifically

A café. The hard case, because of the conversation. Sit away from the counter if you can, hold the phone close, and keep it short. Expect to fix more than usual.

A car. Steady noise, so better than it sounds, but only if the microphone is near you. A phone in a cradle across the dashboard, or a car system microphone mounted in the headliner, is far-field and considerably worse than the handset held close.

A train. Mechanical noise plus intermittent announcements plus other people. Between stations is much better than in a station.

Outdoors. Wind is the whole problem. Turn your back to it, cup your hand, and accept that a gusty day is not a dictation day.

A busy street. Traffic is steady and manageable; a passing motorbike is not. Pause for the motorbike rather than talking over it.

At home with a television on. Speech from the television competes directly. Muting it for fifteen seconds does more than any technique can.

Where the model size matters

This is the one situation where paying for capacity is clearly worth it.

Larger models have more capacity to represent unusual and degraded audio, which is exactly what a noisy recording is. In a quiet room the difference between model sizes is modest; in a café it is the difference between usable and not.

Of the three LocalType ships, 60, 190 and 539 MB, the heaviest is the one built for this. If you regularly dictate in difficult places, it earns its storage even though it makes you wait: the measured middle case was 6.9 seconds of speech converted in 4.1 seconds on a 4 GB test phone, and the largest takes roughly double the time you spent speaking.

On a phone that cannot comfortably hold the largest model, the choice is different: a smaller model running properly beats a larger one being evicted from memory.

Your own vocabulary matters more here too

The effect is larger than people expect.

In clean audio, a recogniser has good acoustic evidence and leans on it. In noisy audio the evidence is weak, so the surrounding words and known vocabulary carry much more of the decision.

So the words you have added to your own dictionary help disproportionately in bad conditions. LocalType keeps those additions on the handset and gives them to the speech model as context. A degraded recording needs precisely that sort of extra help.

Headsets and external microphones

People spend money on the assumption that a headset is an upgrade.

A wired headset with the microphone near your mouth is genuinely better than a phone held at arm's length, for the simple reason that it is closer. Proximity is the whole benefit, and it is a real one.

A wireless headset is less predictable. Which microphone Android hands to an app depends on the device and the profile in use, and it does not always choose the one you expect. Some headsets also compress audio for calls in ways that suit speech to a human ear and not a recogniser.

One comparison settles the question quickly: dictate the same sentence three times, once with the headset, once with the phone held close, once with the phone at arm's length. On many combinations the middle option wins, which surprises people who assumed the accessory was an upgrade.

Built-in car microphones deserve particular scepticism here. They are mounted for calls, positioned across the cabin, and tuned for a human listener on a call, never for transcription.

What to accept

Some situations are not solvable, and knowing which saves you from concluding the software is broken.

Several people talking at once, a microphone across a room, live music, a bar at full volume, or wind on an exposed coast. In all of those, speech and noise arrive inseparably mixed and no app on any hardware recovers it.

In those places, type. It is slower and it works.

What is unchanged regardless of the noise

Everything else about the app behaves the same in a café as at a kitchen table. Recognition happens on the handset from a model kept in storage only the app can reach, and internet access is requested for exactly one purpose, collecting the model you chose. There is no account and no advertising. No behavioural tracking runs inside the product and no archive of your recordings is kept. The app's data is held back from Android's cloud backup by design. Over a password box the microphone key does not appear.

All of which leaves a noisy place with no signal behaving exactly like a noisy place with five bars, and the only variable left is the room.

Get LocalType for Android

Private voice typing that runs on your phone, with speech recognition on the device itself.

Get it on Google Play