Languages
99 languages. 16 we can prove.
Every offline dictation tool built on an open speech model can list a hundred languages, because the model lists a hundred languages. Almost nobody tests them. We tested ours, and 16 passed on recordings of real native speakers; the error rate for every one of them is below. Nine more passed our synthetic tests and then missed on real speakers. They are named on this page instead of quietly dropped.
How this works in the app
- You choose the language. It stays chosen. Pick it once in Settings, or switch it in two clicks from the tray menu. The language you picked is named in the tray menu and on the tray tooltip, so you can never be dictating into a language you did not intend — and while automatic detection is still deciding, both read “Detecting language”.
- Nothing is mixed. One language is active at a time. InkBeepAI does not try to guess per sentence, because guessing needs several seconds of speech and a dictated phrase does not have them. We measured that: on dictations of 3 to 15 seconds, guessing cost nothing in 13 of 14 languages; on short phrases it made 10 of 14 worse, and Hindi was misread as Urdu often enough to ruin a session.
- Automatic detection is there if you want it, off unless you turn it on, and it works the safe way round: it keeps using the language you picked until one dictation has at least four seconds of actual speech, identifies the language from that one, switches to it, and then stays there. The tray menu shows what it picked.
- It never translates. You speak Spanish, you get Spanish text. Dictation, not interpretation.
- Recognition is still entirely on your PC. Both speech models ship inside the download. Adding 98 languages did not add a single network request — the licence check is the same one it always was.
The 16 we verified
The figure is the error rate: out of every hundred words spoken, how many came back wrong (for Japanese, written without spaces, out of every hundred characters). Lower is better and zero is perfect. It counts a missing word, an extra word and a wrong word alike, and it is measured against the reference transcript that came with the recordings — including their own conventions for numbers and names, which we did not tune for.
Measured on recordings of native speakers
| Language | Error rate | Tested against |
|---|---|---|
| Spanish español | 1.9% | Native speakers, Google FLEURS |
| Italian italiano | 3.7% | Native speakers, Google FLEURS |
| Portuguese português | 4.2% | Native speakers, Google FLEURS |
| Japanese 日本語 | 4.8% | Native speakers, Google FLEURS · per character, because Japanese is written without spaces |
| German Deutsch | 4.9% | Native speakers, Google FLEURS |
| Polish polski | 7.2% | Native speakers, Google FLEURS |
| English | 7.8% | Native speakers, Google FLEURS |
| Russian русский | 7.8% | Native speakers, Google FLEURS |
| Turkish Türkçe | 8.0% | Native speakers, Google FLEURS |
| Ukrainian українська | 8.1% | Native speakers, Google FLEURS |
| Catalan català | 8.6% | Native speakers, Google FLEURS |
| Vietnamese Tiếng Việt | 8.6% | Native speakers, Google FLEURS |
| Dutch Nederlands | 8.8% | Native speakers, Google FLEURS |
| French français | 8.9% | Native speakers, Google FLEURS |
| Indonesian Indonesia | 9.1% | Native speakers, Google FLEURS |
| Finnish suomi | 9.3% | Native speakers, Google FLEURS |
English’s 7.8% is for the small English-only model the app runs by default, because that is what an English user actually gets. The same recordings through the large multilingual model score 5.5%, and you can switch English to it in Settings. The fast one is the default deliberately: the gap in speed is far larger than the gap in accuracy, and dictation lives or dies on the wait.
Nine passed on synthetic speech and failed on real speakers
Every language below cleared our synthetic tests, several of them with almost no errors. On recordings of actual native speakers they missed the bar — some on the average, some because too many individual sentences came back wrong. They are still in the app, because they run and some people will want them; the app tells you they missed when you pick one.
| Language | Synthetic speech | Native speakers | Sentences under 15% |
|---|---|---|---|
| Romanian română | 2.2% | 11.4% | 65% |
| Swedish svenska | 9.1% | 11.6% | 60% |
| Slovak slovenčina | 0.0% | 13.1% | 65% |
| Hindi हिन्दी | 3.4% | 14.1% | 50% |
| Czech čeština | 4.6% | 15.8% | 50% |
| Danish dansk | 7.0% | 16.5% | 60% |
| Hungarian magyar | 1.7% | 16.5% | 50% |
| Estonian eesti | 4.1% | 17.2% | 45% |
| Latvian latviešu | 3.8% | 20.5% | 45% |
To be verified, a language needs an average under 15% on native speakers and at least 70% of its sentences individually under 15%.
That gap is the entire argument for the human test. Slovak scored 0.0% on synthetic speech and 13.1% on real speakers, with 7 of its 20 sentences outside the bar. It is why we do not simply repeat the model's own language list back to you as a feature.
How it was tested
Four passes, all on one machine, using the same speech engine and the same settings the shipped app uses.
| Pass | Audio | What it was for |
|---|---|---|
| Clean | 641 clips across 69 languages: up to ten native-written sentences each, read by synthetic voices | A wide first sweep, to find which languages were worth testing properly |
| Noisy | 358 of those clips, covering 37 languages, with microphone band-limiting and room noise mixed in at a measured level | Dictation happens in real rooms, not studios |
| Very noisy | The same 37 languages, with the noise raised until the speech was only just dominant | The point at which a language stops being usable rather than merely worse |
| Human | 500 clips of native speakers reading, 20 for each of 25 languages, from Google’s FLEURS set, with the transcripts that came with them | The pass that decides. Only a language that clears it is verified, and its result is the figure we publish |
A language has to be good on average and good sentence by sentence, so one lucky sentence cannot carry it. On native speakers that means an average under 15% and at least 70% of sentences under 15%. The synthetic sentences were chosen without digits, because “2026” against “twenty twenty-six” scores as an error without anybody having misheard anything. The native-speaker transcripts are used exactly as they came, so there a number the model writes as digits does count against it.
What this costs you
Said plainly, because you will find out anyway:
- The download is much bigger. Two speech models ship inside it — a small English one and a large multilingual one. It is a one-time download and it is what keeps everything on your own machine.
- Languages other than English want a graphics processor. The graphics built into the processor counts. We have measured speed on one laptop, with Intel Iris Xe graphics; older or weaker graphics will be slower. Without a usable one, the multilingual model runs on the processor alone and is far too slow to dictate with. InkBeepAI times the larger model on your own PC when it loads it, and if it is far too slow it says so once and offers the way out that applies — turning the graphics processor on, or switching back to English. It does not let you find out by waiting.
- English does not need a graphics processor. It keeps the same small English model it has always used. On a PC with a graphics processor it now runs on that; without one it runs on the processor, as it always did.
- 83 languages are not verified. They run, and the app tells you when you pick one. We would rather offer them with that caveat than pretend to a hundred verified languages.
The other 74
Like the nine above, none of these is verified. They run; the app labels them as not verified when you pick one.
Did not clear our synthetic screen (44)
Tested on synthetic speech: most came in outside the bar, some of them far outside, and two cleared the quiet pass and fell over once we added room noise. A few of these results are partly scoring artefacts — Chinese output mixes simplified and traditional characters, and Hebrew spelling varies — so read this as “not proven”, not “broken”.
Afrikaans, Albanian, Amharic, Arabic, Azerbaijani, Bengali, Bosnian, Bulgarian, Burmese, Chinese, Croatian, Galician, Georgian, Greek, Gujarati, Hebrew, Icelandic, Kannada, Kazakh, Khmer, Korean, Lao, Lithuanian, Macedonian, Malay, Malayalam, Maltese, Marathi, Mongolian, Nepali, Pashto, Persian, Serbian, Sinhala, Slovenian, Somali, Sundanese, Swahili, Tamil, Telugu, Thai, Urdu, Uzbek, Welsh.
Not in our tests (30)
The model accepts these, and we have no measurement to give you either way.
Armenian, Assamese, Bashkir, Basque, Belarusian, Breton, Faroese, Haitian Creole, Hausa, Hawaiian, Javanese, Latin, Lingala, Luxembourgish, Malagasy, Māori, Norwegian, Norwegian Nynorsk, Occitan, Punjabi, Sanskrit, Shona, Sindhi, Tagalog, Tajik, Tatar, Tibetan, Turkmen, Yiddish, Yoruba.
Common questions
Do I have to download anything extra for another language?
No. Both speech models are inside the installer. Choose a language in Settings or from the tray menu and it is ready — no account, no download, no network request.
Can it switch languages automatically?
It can, and it is off unless you turn it on. When it is on, InkBeepAI keeps dictating in the language you picked until one dictation has at least four seconds of actual speech, identifies the language from that one, switches to it, and stays on it. It deliberately does not re-decide on every sentence: a two-second phrase is not enough to identify a language reliably, and getting that wrong turns speech into confident nonsense with no error message.
Does it translate what I say?
No. It writes down what you said in the language you said it in.
My language is not verified. Should I buy?
Only after trying it, inside the refund window. It will run, and it may be good enough for your work — but we could not verify it, and this page says whether we measured it falling short or never measured it at all. We would rather say that than invent a number. There is a 30-day refund if it is not good enough.
Why do other tools claim more languages?
Because the open speech model underneath most offline dictation tools, including this one, advertises 99. Repeating that number is free. Our list is the same list — the difference is that 16 of ours went through four rounds of testing on this machine and passed the round that decides — recordings of native speakers — and the results, including the 9 that failed on real speakers, are on this page.
Is my audio still private?
Unchanged. Recognition runs on your own PC in every language. InkBeepAI talks to one server, its licence server: once when you activate your key, and then a licence check every two weeks, each carrying your licence key and up to three anonymous hardware IDs — never audio, never text. How that is built.
Dictate in your own language
One-time purchase, 30-day refund, no subscription, no account.