Two completely different approaches
If you have used dictation across a decade, you have met both, and knowing which one you are using explains most of your frustration.
Spoken commands. You say the punctuation aloud. "Full stop." "New paragraph." "Open quote." The system takes those words as instructions, never as text. Precise, controllable and tiring, because you are effectively reading out the mark-up as well as the words.
Inferred punctuation. You speak normally and the system decides where the punctuation belongs, from the rhythm of your delivery and the grammar of what you said. It produces the occasional comma in a place you would not have put one.
Most current systems, including LocalType, do the second. There is no list of spoken commands to memorise, because there are no spoken commands.
What the system is actually using
Two signals, working together.
Your delivery. A pause suggests a boundary. Falling intonation at the end of a phrase suggests a full stop rather than a comma. A rising one suggests a question. This is the acoustic half, and it is why speaking naturally matters more than speaking clearly.
The language itself. How a sentence is shaped implies where a comma belongs, regardless of how it was said. A subordinate clause, a list, a name being addressed: all of those have conventional punctuation that can be inferred from the words alone.
When the two agree, the result is usually right. When you deliver a sentence in an unusual rhythm, the acoustic signal contradicts the grammatical one and the system has to pick.
Why yours comes out unpunctuated
The commonest complaint, and it nearly always has the same cause.
Speaking in one continuous stream with no pauses removes the acoustic half of the evidence entirely. Grammar alone is all that is left to guess from, and against a long run-on it will often produce exactly that: a long run-on.
Speaking more slowly is not the fix here. Speaking in units is. Say a sentence, stop briefly, say the next one. That single pause is the strongest punctuation signal there is, and it costs nothing.
The habits that fix most of it
Pause at the end of sentences. Half a second is plenty. Decide the sentence before you start it. Trailing off and restarting gives contradictory signals, and punctuation is the first thing to degrade.
Do not say the punctuation. On an inferring system, "comma" produces the word comma. This catches out anyone who learned dictation on older software.
Keep sentences to a reasonable length. A sixty-word sentence is hard for a system to punctuate for the same reason it is hard for a reader.
Let capitals happen. Sentence capitals and most names are handled automatically. Trying to control them by emphasising a word does nothing useful.
What you will still have to fix by hand
There is a ceiling, and these are the things that sit above it.
Lists. Inferred punctuation produces sentences, not bullet points. Dictating a shopping list gives you a paragraph with commas, and turning it into a list is manual work.
Quotation marks. Quotes require knowing where speech starts and ends, which is not reliably audible. Expect to add these yourself.
Paragraph breaks. A long pause may or may not produce one. For anything structured, dictate in separate passes and break it up afterwards.
Anything technical. Brackets, colons, semicolons, dashes. These have weak acoustic signals and appear rarely in ordinary speech, so they are inferred badly.
Everything here points the same way as it does for writers generally: dictate the words, format afterwards. Fighting the system for a semicolon costs more than adding it later.
Different languages, different rules
This matters if you write in more than one.
Punctuation conventions differ by language, and so does how much of it a model has seen. Comma use in German is largely grammatical and heavily rule-bound; in English it is looser and more rhythmic. A model handles each according to what it learned. Punctuation quality is therefore not uniform across the languages, any more than word accuracy is.
LocalType covers eighteen languages with one multilingual model, and identifies which you are speaking from the recording, so the punctuation conventions follow the language you actually used, and not whatever the keyboard happens to be set to.
Editing after the fact is not a failure
This decides whether people stick with dictation.
Nobody types a perfect paragraph either. Typed text gets a comma moved, a sentence split, a word swapped. The difference is that typing errors are visible as you make them, so the correction happens continuously and you never notice it as a separate activity.
Dictation front-loads the words and defers the corrections, which makes the editing feel like a distinct chore even when the total time is lower. Knowing that in advance stops people concluding the tool is broken when it is simply distributing the work differently.
The practical version: dictate the whole thing before fixing anything. Stopping to correct a comma mid-flow destroys the rhythm that was producing decent punctuation in the first place, so you end up making the problem worse while trying to fix it.
Where LocalType stands
It infers punctuation and does not accept spoken commands. There is no list of phrases to learn and no setting that switches to command mode, because that mode does not exist in the app.
What it produces is a transcription of what you said, with punctuation and capitals worked out, dropped into the field you were typing in. It does not rewrite, tidy, shorten or remove filler words. If you said "um", it is doing its best with "um", and that is deliberate: the app writes down the words you actually said, not an improved version of them.
All three model sizes affect this as well as word accuracy, since a model with more capacity handles ambiguous rhythm better. On the 4 GB test handset the middle size needed 4.1 seconds for 6.9 seconds of speech, and the largest takes roughly twice the time you spent speaking.
All of it runs on the phone, so punctuation quality does not change with your signal. Internet access is used once, to fetch the model you chose. There is no account and no advertising. No behavioural tracking runs inside the product, and no archive of your recordings is kept. The model sits in storage only the app can read. The app's data is excluded from Android's cloud backup, and in password fields the microphone does not appear.
Pause, and stop saying comma
Pause between sentences and stop saying "comma". That is ninety per cent of it.
Ten per cent is left over, and it is accepting that lists, quotation marks and anything structural are yours to add afterwards. Dictate the words, format at the end, and the whole thing stops being a fight. The systems that made you say every mark aloud were more controllable, and almost nobody misses them.