Runs on your device

Choosing a speech model

A speech model is the part that turns your voice into text. Because it runs on your phone rather than on a server, its size actually matters — and you get to choose.

What a speech model is

It is a file. A fairly large one, trained to recognise speech, that LocalType loads and runs on your device. Cloud voice typing keeps this file on a server you never see, which is why you never think about it. Running it locally means it lives on your phone instead, taking up storage and memory while it works.

That is the whole trade. You give up some storage; you get transcription that does not depend on a connection or on sending your audio anywhere.

The three things that trade off

You cannot maximise all three at once. Which is why there is more than one model.

Speed

How quickly the text appears after you stop speaking. Smaller models return results faster, particularly on older hardware.

Accuracy

How well it copes with accents, names, background noise and people who talk quickly. Larger models are generally better at this.

Storage and memory

How much space the file takes up, and how much your phone has to hold in memory while transcribing.

How to pick one

You do not have to understand any of the above to get a sensible result.

  • LocalType can suggest a model based on what your device can comfortably run.
  • If dictation feels slow, move down a size before assuming the app is at fault.
  • If accuracy is the problem and your phone has room to spare, move up a size.
  • You can change your mind at any time: downloading another model replaces the one you have.

Switching models means downloading the new one, which needs a connection at that moment.

The models side by side

Compact 60 MB Fastest Fair Older phones, or phones with little memory to spare
Standard 190 MB Faster than you speak Good Most phones. This is the one LocalType suggests by default
Large 539 MB Slower than you speak Best Phones with memory to spare, when accuracy matters more than waiting

Measured on the test device, a phone with 4 GB of memory: Standard turned 6.9 seconds of speech into text in 4.1 seconds. Large takes roughly twice as long as you spent speaking. Your own phone will differ, which is one of the things the beta is for.

What about older phones?

This is the honest limit of on-device processing: a phone with little memory will struggle with a large model, and no amount of software gets round that. A smaller model is usually the answer, at some cost in accuracy. Working out where that line falls across real devices is one of the main reasons for the beta.

Help us test on your phone

Questions about models

Which model should I start with?

The one LocalType suggests for your device. Adjust from there if it feels slow or misses too much.

How much storage does a model need?

Between 60 MB for Compact and 539 MB for Large, with Standard at 190 MB in between. During installation you briefly need room for two copies.

Can I have more than one model installed?

No, one at a time. Downloading another replaces it, so you never have two taking up space.

Does a bigger model drain my battery?

Transcription only runs while you are dictating, so the effect depends on how much you use it. A larger model does more work per sentence.

Is the model updated automatically?

Downloading anything is something you trigger yourself, and it needs a connection at that moment.

LocalType is in testing

Join the beta and help work out which phones, languages and situations still need attention.

Join the beta