Languages, Translation & Voice

What Languages, Translation and Voice
Capabilities Run On Device?

Which languages the assistant reads and answers in, which speech capabilities execute on the machine, and what has actually been tested.

Built from real buyer questions in our sales meetings

Ask an executive team whether their assistant handles Spanish and everyone nods. Ask whether last Tuesday’s recorded meeting traveled to a speech service in another country and the room goes quiet. Two questions hide inside one, and only the second carries a security consequence.

Direct Answer

Text coverage belongs to the model; speech coverage belongs to the product. Which languages your assistant reads and answers in follows the model you load, not the application around it: Qwen suits Asia-Pacific, Mistral suits EMEA, Llama covers English. Speech is answered capability by capability — AirgapAI Transcribe records microphone and system audio and transcribes it 100% offline on the device, and AirgapAI Translator runs translation on the AI PC.

The limit: no single language number, and listed is not tested. Iternal publishes no headline figure for the translator, and there is no stable one to give — the working set moves as models are added and validated, so ask for the list by name for your build. On transcription the gap is explicit: the speech model behind AirgapAI Transcribe lists 99 languages, and about 25 of the most common have been tested. Rarer languages degrade because of that speech model itself. Speaker detection is not supported, so a recording returns as one undifferentiated transcript.

Coverage is not adequacy. A language can sit on the list and still fall short for the market you sell into; Iternal has held Spanish open pending confirmation that it is good enough for a region. Name your languages, ask which Iternal has tested, and run your own audio before you commit — the questions below get that in writing.

The model sets your text coverage; the recording sets your compliance exposure. For more information on the model choice, visit the choosing a local model page. For more information on recording a conversation, visit the records and privilege page.

Why the Answer Is a Named List, Not a Language Count

Language coverage is a narrow question that matters enormously to the few who raise it: multinational and public-sector teams carrying a named language and a deadline. For them a headline number answers a question they did not ask.

A count is the wrong unit. Iternal assembles the translator from a stack of models, with a separate model per language on the speech side, so the working set moves as models are added and validated. Figures spoken inside an engineering review scope what is enough to ship rather than claim capability: the framing for the first release was that the set in hand was sufficient and one missing language acceptable. That language was Pashto, requested by a customer, not working today, with substitutes under research at roughly 70% similar.

So the deliverable is a named list, per build, in writing. A number ages the moment a model changes; a named list does not.

Pin it down: questions for your evaluation
  • Send us the list of languages the translator handles today, by name, for our build.
    Replaces a figure that ages with a list you can check line by line.
  • Which of our languages has Iternal tested, and which are listed by the speech model but untested?
    Separates first-class support from best effort before a region depends on it.
  • During the pilot, can we run our own recordings and documents in our own languages?
    Adequacy on your own material, which outranks any list — including this one.

Transcription: What Is Listed Against What Is Tested

AirgapAI Transcribe is a different product from the chat assistant, and its language coverage counts separately. It records microphone and system audio, runs Whisper on the device NPU, then summarizes locally — nothing reaches a speech service, which is why European teams under recording and privacy obligations reach for it first. An NPU is a requirement, not a preference.

Whisper lists 99 languages. Iternal has tested about 25 of the most common. The names offered when buyers press — Hindi, Thai, Chinese, Japanese, German, French, English and Arabic — illustrate what works well; they are neither the tested set nor its boundary. Plan around the gap between listed and tested. Quality degrades for rarer languages because of the speech model itself, so no amount of local hardware rescues a language that model was thin on. One further limit rides along with every language: speaker detection is not supported, because the model that separates voices needs server-class hardware, so a multi-party recording returns as one continuous transcript.

Where Each Capability Executes, and What State It Is In

Two facts decide whether a capability survives a sealed environment: where it executes, and how far Iternal has taken it. Blanket answers hide both.

Capability Where it executes State today
Text
Reading and answering in a language On the device, in the assistant Follows the model you load; mainstream open-weight models cover most major languages
Document and text translation On the device, in AirgapAI Translator Shipping; the text translation model runs without a discrete GPU
Speech
Dictation into the assistant On the device Available — prompts can be spoken instead of typed
Meeting transcription and summary On the device, NPU required Available, fully offline; audio is transcribed after capture, not displayed live
Real-time translation On the AI PC, in AirgapAI Translator Demonstrated Arabic into Chinese, English and French
Live captioning inside the chat app Not built today
Speaker separation in a transcript Would need server-class hardware Not supported
Voiceover in a cloned voice Iternal content tooling, not the on-device assistant Not available today; Iternal has described cloning as coming
Voice agent answering inbound calls A hosted service, because a phone network is involved Deployable, and outside the on-device boundary

Read the middle column first. A capability that executes on the device inherits the security posture you bought the device for; one that reaches a service inherits somebody else’s.

For more information visit the secure and offline AI translation tools page.

Three Tiers, and the Route to a Missing Language

Ask for a tier, not a yes. Every language you care about sits in one of three states, carrying very different risk:

  • Tested. Iternal has run the language and validated the result. Plan on it, and confirm the tier in writing.
  • Listed but untested. The underlying model claims it; nobody has proven it on content like yours. Treat it as best effort and settle it in the pilot.
  • Not working today. Pashto is the worked example, with substitutes under research at roughly 70% similar. A near-neighbor language is a workaround, never a replacement.

The route to a missing language runs through the model, not a feature request. Covering a new language means running a model that already handles it, and the translator gains a language by validating one for it. Iternal maintains a curated library of open-weight models it tests against local hardware and refreshes as models change.

Answered elsewhere
FAQ

FAQ: Languages, Translation and Voice on Device

There is no single number, and Iternal publishes none for the translator — the working set moves as models are added and validated. Text coverage follows the model you load, and mainstream open-weight models cover most major languages. Ask for the current list by name, for your build.

On the device. AirgapAI Transcribe records microphone and system audio and transcribes it 100% offline, running Whisper on the NPU and summarizing locally, so the audio never reaches a speech service. A device with an NPU is required.

No. Speaker detection is not supported, so a recording returns as one continuous transcript of everything said. The model that separates speakers needs server-class hardware, and Iternal states that no laptop today can run it. Plan for a person to attribute the lines afterwards.

Not today. Generated voiceover lives in Iternal content tooling rather than in the on-device assistant, and narration in a named person’s cloned voice is not available there — Iternal has described cloning as coming rather than shipped.

Yes. AirgapAI accepts dictated input through the microphone, so prompts and answers can be spoken rather than typed. People who work that way describe voice as faster and better at surfacing ideas, then return to the keyboard where spelling must be exact.

AirgapAI Translator does, on the AI PC: Arabic into Chinese, English and French has been demonstrated in real time on device. Live captioning inside the chat application is a separate capability and is not built today.

Name the Languages, Then Test Them

A language list settles an argument; your own recordings settle the question. Load the model you would deploy, record a real meeting in the language your people work in, and read what comes back. Whatever survives that hour is what you have bought, and it stays on the machine either way.

John Byron Hanby IV
About the Author

John Byron Hanby IV

CEO & Founder, Iternal Technologies

John Byron Hanby IV is the founder and CEO of Iternal Technologies, a leading AI platform and consulting firm. He is the author of The AI Strategy Blueprint and The AI Partner Blueprint, the definitive playbooks for enterprise AI transformation and channel go-to-market. He advises Fortune 500 executives, federal agencies, and the world's largest systems integrators on AI strategy, governance, and deployment.