Why it fits
Most social writing is a paragraph or less, written on a phone, often while doing something else. That is precisely the shape dictation handles best.
It also suits the register. Posts and comments are informal, and dictated text comes out closer to how you would say something than typed text does. Typing on a phone pushes people towards clipped, abbreviated sentences because each character costs effort; speaking removes that pressure and you end up with whole thoughts instead of fragments.
For anyone who writes captions regularly, the time saving compounds. A caption is thirty seconds of typing and eight seconds of speaking.
What dictation cannot do here
Three limits are specific to this kind of writing.
Hashtags. Saying "hashtag" produces the word. There is no reliable way for a system that infers punctuation to know you meant a symbol, so tags are a typing job. Dictate the caption, add the tags by hand.
Mentions and handles. Same problem, worse. An @ is a symbol, and the name after it is usually not a normal word, so both halves fail. Type these.
Emoji. Dictation writes words. Saying the name of an emoji gives you the name, not the picture. Use the keyboard's emoji picker afterwards.
The practical shape that emerges is the same one that works everywhere else: speak the sentences, type the symbols. It takes about a day to become automatic.
The part that makes this different from messaging
A wrong word in a message to a friend gets corrected in the next line. A wrong word in a public post stays there, gets seen by everybody, and quite often cannot be edited at all depending on where you posted it.
That raises the value of checking, without raising it to the level of a work email. A realistic routine takes fifteen seconds:
Read it once before posting, looking specifically at names and at any sentence containing a negation. Those are the two categories where a dictation error changes the meaning. Elsewhere it reads as a typo, and a reader forgives it or does not notice.
Names deserve particular attention if you are mentioning a person, a place or a business. Getting somebody's name wrong in public is a different order of mistake from a misplaced comma.
Where dictation quietly improves what you write
Two effects run in your favour here, and neither is obvious. Length becomes natural, where typing made it effortful. People write shorter posts than they mean to, because typing is slow and they run out of patience halfway. Spoken posts tend to contain the whole thought, including the part that would have been cut.
Tone gets warmer. Text typed with thumbs tends to read flatter than intended. So much short-form writing needs emoji for exactly that reason, to signal that nothing is wrong. Dictated sentences carry more of the rhythm of speech and read as friendlier without any decoration.
Both are more useful for captions and comments than for anything else, because that is where tone does most of the work.
Doing it somewhere you are not overheard
Social writing happens in public places: on transport, in queues, in company. All of those are exactly where speaking a post aloud is awkward, particularly when the post is about the situation you are in.
A workable pattern is to dictate the substance somewhere private and post later, or to dictate the harmless parts and type the rest. The microphone is a key on the keyboard, not a separate app, so moving between the two happens inside the same box without ever leaving the post.
Setting it up for this
Two adjustments pay off on the first post. Add the names you post about. People, places, businesses, products and the recurring in-jokes no general speech model has ever encountered. Anything you add is stored on the phone and passed to the recogniser as context, so it starts appearing spelled the way you meant it, and not approximated.
Pick a smaller model. Social writing is short and frequent, so the wait after each dictation is the thing you notice most. The 60 MB or 190 MB options usually feel better here than the largest, unless you are posting from noisy places. For scale, the middle option converted 6.9 seconds of speech in 4.1 seconds on a 4 GB test handset.
The offline part, which matters more than you would expect
Posting happens where the signal is worst: festivals, trains, stadiums, holidays abroad, anywhere with a lot of people competing for the same masts.
Cloud dictation stalls in exactly those places. LocalType runs the recognition on the handset from a model in storage no other app can open, so you can write the caption now and let the app upload it when the connection returns. There is exactly one thing internet access is used for, collecting the model you chose.
No account is involved. Advertising is absent from the product, no behavioural tracking runs inside it, and nothing keeps an archive of what you said. The app's data is left out of Android's cloud backup by design. In a password field the microphone key is withheld.
Comments are the underrated case
Captions get the attention, and comments are where dictation actually changes behaviour.
A reply to somebody else's post is the writing people abandon most often. It is low stakes, it takes longer than it feels worth, and half of them are never finished. Speaking one takes five seconds, which moves it below the threshold where you give up.
Replies to your own comments work the same way, and that is where conversations either continue or quietly stop. Anyone who has posted something and then not answered the twelve responses knows the pattern. Typing effort is the problem there, not a lack of interest.
Two cautions specific to comments. They are public in somebody else's space, so the checking above applies in full. And they are frequently written in situations where speaking aloud is conspicuous, which is the same constraint as everywhere else.
If you post for a business
Volume is the argument in favour. Anyone producing captions daily is doing the same short task dozens of times a week, and the saving compounds in a way it does not for personal posting.
Two things pull the other way. Business writing gets read more carefully by people looking for reasons to be unimpressed, so the checking step is not optional. And brand names, product names and campaign terms are exactly what a general speech model has never met. Adding the vocabulary is therefore the first job on the list, not an afterthought.
The workable arrangement is to dictate drafts and edit them at a desk, rather than dictating and posting in one motion. That keeps the speed and removes the risk of a public error nobody caught.
Speak the sentences, type the tags
Whichever you are writing, the division of labour is the same. Sentences get spoken, because that is what speech is good at. Tags, handles and anything with a symbol in it get typed, because saying "hashtag" out loud produces the word. Then read it once, checking names and negations, and post.
The reading step takes about five seconds and it is the one people skip. A missing "not" reads as the opposite of what you meant, and on a public post that is a correction in the replies rather than a private line to a friend.