The problem is switching, not translation
Switching between languages and translating between them get confused constantly, and they are not the same problem.
Translation turns what you said into a different language. It is a distinct product and a different problem.
Dictating bilingually means writing down whichever language you happened to speak, correctly, without telling the system first. You say a sentence in Dutch, it writes Dutch. Ten minutes later you say one in English, it writes English. Nothing is converted; the system simply has to recognise which language it is hearing.
Small as that sounds, it is the friction that decides whether you dictate at all on a day when both languages are in play.
Living in two languages is not the same as knowing two
The two lead to different expectations of a keyboard.
Someone who learned a second language at school and uses it occasionally is dictating in a language they speak carefully and slowly. That is easy material: deliberate pronunciation, standard vocabulary, complete sentences.
Someone who lives in two languages does something quite different. They speak both fluently, fast, with local phrasing and imported words, and they switch according to who they are talking to, and no plan comes into it.
All of that is harder for any recogniser, and it is also the situation where automatic detection pays for itself several times a day. The errors you notice first will be proper nouns crossing between the two languages: a Dutch surname inside an English sentence, an English product name inside a Dutch one. That is the category you can fix yourself, by hand and once.
Why most tools handle it badly
Traditional dictation assumes one language at a time, chosen in advance.
That assumption is reasonable for someone who only ever writes in one. For everyone else it produces a settings trip several times a day, and the settings trip is the reason people stop using dictation. Nobody changes a menu to answer a two-line message.
The failure when you forget is also unusually unhelpful. The system does not say "this is not the language I expected". It does its best to hear the language it was told to expect and produces confident nonsense, so you get a message that looks like it was written by somebody having a stroke. What you do not get is an error you can act on.
Systems that download a pack per language add a second problem: each language you add costs storage, and switching may mean waiting for something to load.
Detecting the language from the recording
The better approach is to work it out from the audio, per recording, and act accordingly.
LocalType does this across eighteen languages using a single multilingual model, so there is nothing to select before you speak and nothing to download per language. Start a message in one language and the next one in another, and it follows. The full list, and the details of how the setting behaves, live on the languages page.
That one download comes in three sizes, at 60, 190 and 539 MB, which is where you set speed against accuracy. On a 4 GB test handset the middle size turned 6.9 seconds of speech into text in 4.1 seconds. For bilingual use the middle or the large one is usually worth the storage, since working out the language and transcribing it draw on the same capacity.
On a bilingual day the practical effect is that dictation stops being something you configure and starts being something you use. Reply to your mother in one language, your colleague in another, a form in a third, without touching a setting.
The model runs on the phone, which matters more to you than to most readers. Living in two languages often means moving between two countries, and a foreign SIM with data switched off is exactly where a cloud dictation app goes quiet. Here the network is touched for exactly one purpose, pulling down the model you picked, and it is not asked for again.
The rest of the product does not vary with your language either. No account exists and nothing asks you to sign up, there is no advertising in the app and no behavioural tracking inside it, and your recordings are not archived, because the audio becomes text and is finished with. The model file is kept where only LocalType can read it, the app's data stays out of Android's cloud backup by design, and no microphone key is offered over a password field.
Where automatic detection is weakest
Detection is not magic, and it gives way in places you can predict.
Very short utterances. One or two words give the system almost nothing to identify a language from, so a single name or a "yes, tomorrow" is where it is most likely to guess wrong. Full sentences are far more reliable.
Languages that resemble each other. Closely related languages share sounds and structures, and short phrases in one can look plausible as the other.
Sentences containing both. Detection happens per recording rather than per word, so a sentence that genuinely mixes two languages has to be assigned to one of them. A borrowed English word inside a Dutch sentence is fine, since the sentence is still Dutch. Half and half is not.
Uneven accuracy between languages. Speech models are trained on very different quantities of material per language, so results are not identical across the eighteen. Widely spoken languages come out best.
The practical workaround for all four is the same: speak in complete sentences and avoid fragments, and keep each recording in one language.
The keyboard layout question
This is the part that catches people out, and you will do better understanding it than fighting it.
If you set a specific language, everything moves together: the keyboard layout, the dictionary used for autocorrect, and the dictation language. That setup is the right one if you will be writing in one language for a stretch, because typing and dictation agree with each other.
If you leave it on automatic, the dictation follows your voice but the layout stays where it was, and it has to, because the layout is on screen before you have said anything. There is no way for a keyboard to know which language you are about to speak.
The practical size of that is smaller than it sounds if your two languages share a layout, and larger if one of them is QWERTZ or AZERTY and the other is not, since then the accented characters and the punctuation move as well.
So the advice is: automatic for dictation-heavy days when you switch constantly, and a fixed language when you will be typing as much as speaking in one of them. Neither is a workaround; they are answers to different days.
A routine that works
Set a language when you sit down to work in it. If your whole afternoon is in one language, pin it. Autocorrect will stop fighting you.
Leave it automatic when you are moving between conversations. That is the setting to be on when the next message could be in either language.
Add names in both languages once. Bilingual life is full of proper nouns that belong to one language and get used in the other. Those additions stay on the handset and are handed to the speech model as context, so they start arriving correctly whichever language surrounds them.
Work through that last one properly before you judge anything else. Surnames of people you write to in both languages, the street you live on, your employer, the two or three places whose names your other language mangles, and any word you have already corrected twice by hand.
Say full sentences. It is the fix for most detection errors and it costs nothing.
Then give it a fortnight of ordinary messages before deciding. If you find yourself opening a settings menu during that fortnight, note down what you were about to write, because that is the case the routine above has not covered yet.