"Offline" is doing two jobs in that sentence
The phrase gets used for two entirely different things, and only one of them is what you are looking for.
Some apps mean your notes are saved on the device. Useful, but it says nothing about your voice: the audio can still be uploaded, transcribed on a server and sent back as text, with the result then stored locally. The privacy question is untouched and the app stops working the moment you lose signal.
What you want is the recognition itself happening on the handset. A model file sits on your phone, your speech is processed by it, and the network is not involved at any point after the initial download.
There is a two-minute test that settles it. Install the app, let it finish any first-run download, then switch on airplane mode with Wi-Fi off and dictate a sentence. Words appear, or they do not. Nothing else needs to be taken on trust.
Then there is a third use of the word, and it is the one that reads most reassuringly on a listing page: offline as a mode.
An app that offers offline as a mode is telling you something important: there is another mode, and something decides which one you are in. That decision can depend on your handset, the language you picked, whether a pack finished downloading, or a change in a future update that nobody announces.
Modes are fine if the reason you want offline is patchy signal. They are much weaker if the reason is that you would rather your voice did not travel, because then you are relying on a setting staying where you left it.
An app with no cloud path at all cannot have that problem. If you want to know which of the two you are holding, run the airplane mode test twice, a fortnight and an update apart, and see whether the answer is still the same.
Four questions that shorten the list
Answer these before you read another comparison table anywhere, and you will not need the table.
Do you write in more than one language? This narrows things faster than anything else. Several of the on-device options are English-only, because multilingual models are larger and a small project reasonably picks one language and does it well. If you switch between languages, that rules most of them out.
Do you need to read the source? If yes, the open-source options are the answer and nothing else competes. For some people this is a real requirement, and it should not be argued with.
Do you want a keyboard, or a component? A voice input service plus your existing keyboard keeps your typing exactly as it is. A combined keyboard means one install and one settings screen, and the dictation and the dictionary know about each other. Neither is cleverer than the other.
How old is your phone? On-device recognition is bounded by memory. A smaller model on a tired handset beats a large model that swaps and stalls. Any app that does not let you choose the size is choosing for you.
The apps that really do run on the handset
Specifications are easy to look up, and they tell you almost nothing about who an app suits.
FUTO Voice Input is for the person who has a keyboard they will not part with. It is a recognition service, not a keyboard, so it slots in behind whatever you already use. Open source, one payment instead of a subscription, no account, and a project with a real record on this principle.
Whisper IME is for the tinkerer who wants free and inspectable, and writes in short bursts. There is a limit on how long a single recording can be, which makes it comfortable for replies and awkward for paragraphs.
Transcribro is for the English writer who wants the code readable and would rather not install from Google's store. It provides a keyboard and also offers itself to other apps as a recognition service, so it can be wired in behind things.
Gboard with a downloaded language pack is for the person who wants none of this to be a project. It is already installed, and on many handsets it is quick. There is a caveat. Whether it stays local depends on your phone, your language and what is installed, so it is a mode rather than a guarantee.
LocalType is for the multilingual writer who wants one finished thing: a keyboard with dictation in it, no cloud path anywhere in the app, and nothing to assemble. Paid after a free trial, and closed source.
That list mixes two shapes of software, and the difference decides how you install. A recognition service registers itself with Android and is then offered to other apps, so your existing keyboard keeps its layout, its dictionary and its habits and only the listening changes hands. A keyboard replaces the input method outright, emoji and autocorrect and all. Transcribro does both, which is why it can be read either way.
Scoring LocalType on the same four
It answers the first question and the third, and loses on the second.
LocalType handles eighteen languages with a single multilingual model, and works out which of them you are speaking from the recording itself, with no menu to touch, so switching language between two messages needs nothing from you. Pin it to one language and the layout, the dictionary and the dictation language move together.
It is a keyboard first. Autocorrect and next-word suggestions from a dictionary matched to your language, QWERTY, QWERTZ and AZERTY where those are the national standard, a full emoji set, and your own surnames and jargon added by hand. Those added words stay on the handset and are handed to the speech model as context, so the names you actually use start coming out spelled correctly.
On the fourth question it refuses to decide for you. Three model sizes ship with the app, from a lightweight one built for handsets that are short of memory, to a middle option that is the default recommendation, to a large one for people who would rather wait a beat than fix a word afterwards. The sizes are 60, 190 and 539 MB, and the figures behind the middle one, on a test handset with 4 GB of memory, were 6.9 seconds of speech converted in 4.1 seconds. Only one model is kept at a time, so changing your mind swaps the file rather than adding to it.
On the second question it simply loses. The source is not published, and if that is your requirement you already have better answers above.
None of the four questions reaches what happens once the model is on the phone. Internet access is requested for exactly one purpose, and that purpose is the model download you already agreed to. Nothing registers you, so there is no account and no profile attached to a keyboard. Advertising is absent from the product and no behavioural tracking runs inside it. Audio becomes text and is then finished with, so no archive of your recordings exists to be asked for. The model file lives in storage no other app can open, and the app's data is kept out of Android's cloud backup by design, which is why a new phone starts you from scratch. Over a password field the microphone key is not offered at all.
Pick by the sentence that sounds like you
If you need the code readable and you write in English, take one of the open source options. If you are attached to your current keyboard, pair it with a voice input service. If your phone is a recent Pixel or Samsung and you have not thought about any of this, the built-in offline packs are already better than most people realise.
If you write in several languages and want dictation inside a keyboard you would use anyway, without assembling it from parts, that is the gap LocalType was built for.
And if two of these still look equally plausible, install both and run the airplane mode test on each in the same afternoon, using the sentences you actually write. Whichever one is still producing words with the radios off, in the language you were speaking, is the one to keep.