Solutions / PolyText
Understands the way people actually talk.
PolyText is the translation API with the world's widest coverage — 550+ languages of text, 1,600+ understood from speech, mixed languages and code-switching included — plus a realtime voice pipeline that turns ElevenLabs' voices into full conversations.
Text & voice · real-time, with one way to integrate.
People don't speak in neat, single languages. Most software assumes they do.
"Our customers mix two languages in one sentence and the system falls apart."
Real speech is full of code-switching, and most language tools simply break when it happens.
"The languages our users actually speak aren't even on the list."
Whole communities get left out when a platform only supports a handful of major languages.
"We wired up one vendor for text and another for voice, and a third per language."
Stitching together separate tools is slow, brittle and expensive to maintain.
One API: hear, translate, speak.
Send text or voice
Pass in whatever your users say — a message, a call, a recording — in one language or several mixed together. PolyText understands speech in 1,600+ languages.
It understands
PolyText detects the languages, follows the code-switching, and grasps what the person actually means — in real time.
You respond, anywhere
Reply in text in 550+ languages, or by voice in 70+ — one streaming pipeline from hearing to speaking, instead of stitching STT, translation and TTS vendors yourself.
Built around the outcomes you care about.
Real, mixed speech
Follows code-switching across languages in a single sentence, so people can speak how they naturally do instead of forcing themselves into one language.
550+ languages of text
Translation, chat and messaging in 550+ languages — more than any translation API (DeepL covers ~33, Google ~190, Azure ~135) — so whole communities are finally included, not just speakers of a few major languages.
1,600+ understood from speech
Voice notes, recordings and calls understood in 1,600+ languages — more than any voice platform — with replies in text in the speaker's language or by voice in 70+.
Realtime voice pipeline
Streaming speech-to-text, first-clause synthesis and sub-3-second cold starts — with always-warm tiers for integrators — so ElevenLabs' voices become full conversations, not clips. Language never becomes the thing people wait on.
One way to integrate
One API — hear, translate, speak, streaming end-to-end — replaces DIY stitching of separate STT, translation and TTS vendors. Less to wire up, less to maintain.
Built for trust
Handled with care and the guardrails serious teams need, so you can serve every language community with confidence.
Five numbers, five different things.
Understanding a recording, holding a written conversation and speaking on a live call are three different jobs. We count them separately — here is exactly what each number covers.
Understood
PolyText can turn speech into text in this language — voice notes, recordings, voicemail, message attachments. Not instant; it takes accuracy over speed.
Live understanding
Fast enough to understand the speaker during a phone call, in the moment — not after it.
Text conversation
Full written back-and-forth — chat, SMS, WhatsApp, email — in the customer's own language.
Voice
PolyText can speak the language aloud — voice notes, campaign audio, and recorded replies.
Live calls
A real phone conversation, spoken in realtime. Beyond these, synthesis is accurate but slower than the call — so the language is served by voice note or by bridge mode, not by speaking it live.
Bridge mode
When PolyText understands a language it cannot speak — Twi, Ga and Ewe today — the speaker still speaks naturally in their own language and is understood — every word — and the reply comes back aloud in a regional language they also speak, or in writing in their own. It is what turns 1,600 languages understood into 1,600 languages served.
Wherever people speak in more than one language.
"Our users stopped forcing themselves into one language. The product finally meets them where they are."
What product and support teams tell us when PolyText goes in: more users understood the first time, fewer dropped because the language wasn't there, and far less plumbing to maintain. (Early-adopter results; we'll publish named case studies as deployments mature — no invented numbers.)
Start free. Scale when it's working.
Try it at no cost, move to Pro as your volume grows, and talk to us when you need enterprise scale, controls and support.
- Text & voice understanding
- Core languages
- One simple integration
- Community support
- 550+ text languages · 1,600+ understood from speech
- Mixed-language & code-switching
- Real-time text & voice
- Higher volume & throughput
- High volume & SLAs
- Dedicated warm capacity — guaranteed-latency, always-warm voice pipeline for integrators
- Advanced controls & audit
- Dedicated support
- Private / sovereign deployment
The things teams ask us first.
What do you mean by "the way people actually talk"?
Real speech rarely stays in one neat language. People mix two or three in a sentence, switch mid-thought, and use local turns of phrase. PolyText follows that code-switching and understands what someone means — instead of breaking on it.
How many languages does it cover?
550+ languages of text for translation, chat and messaging — more than any translation API — and 1,600+ understood from speech. It speaks 70+ aloud and holds live calls in 30+.
Does it work for both text and voice?
Yes. The same understanding works across written messages and spoken conversation, in real time, wherever your users reach you.
How hard is it to integrate?
One simple integration covers text and voice in a single API — hear, translate, speak — replacing the usual patchwork of separate STT, translation and TTS vendors. Less to wire up, less to maintain.
Is it fast enough for live conversations?
Yes. The voice pipeline streams end-to-end — speech-to-text as the person talks, translation, then first-clause synthesis so the reply starts speaking before the full sentence is composed — with sub-3-second cold starts and always-warm tiers for integrators.
What does it cost?
Start free, move to Pro as your volume grows, and talk to us for an Enterprise plan with high volume, advanced controls and dedicated support.
Pairs naturally with
KasaFlow
Calls that understand mixed-language callers and switch language mid-conversation.
Learn moreNexaGrow
Build agents that understand and reply the way your users actually talk.
Learn moreSOWA
Teach and assess students in the language they think in, text or voice.
Learn moreReach your users in 550+ languages — and be understood in 1,600+.
Start free and add multilingual understanding to text and voice with one integration. Read the docs or book a walkthrough.