A Vietnamese voice that runs on your own device. Download it once and read any Vietnamese text aloud, with nothing sent to a server.
This page is the same text-to-speech studio, pre-set to the Vietnamese voice — useful if you're pasting Vietnamese text specifically and don't want to pick a language first. Everything below is about that one voice: what it sounds like, what it needs from your text, and where it comes from.
The Voice: Vais1000
Vais1000 is a single-speaker Piper voice — a female voice trained on the VAIS1000 Vietnamese speech dataset, tagged "medium" quality on Piper's own scale. That puts it a notch below the very largest commercial voices in smoothness, but clearly intelligible, with the fixed cost of a small download and no ongoing latency once it's loaded, since nothing round-trips to a server per sentence.
Getting Tone Marks Right
Vietnamese TTS lives or dies on diacritics. The model reads tone directly off the accented vowels in your text — huyền, sắc, hỏi, ngã, nặng each change how a word is pronounced and, in Vietnamese, often what it means too. Paste in text with full diacritics and the model has what it needs; paste in text typed without them and there is no tone information left to read, so results will sound off or simply wrong. If your source text lacks diacritics, restoring them before pasting in is worth the extra step.
Tip:If you're not sure your source has correct diacritics, generate a short sentence first — a mispronounced word is usually a missing or wrong accent mark in the source text, not a model limitation.
Where the Download Goes
Like the rest of this tool, the Vietnamese voice downloads once into your browser's own storage and is read from there on every later visit, never downloaded again — and nothing about your Vietnamese text is sent anywhere to be read aloud. See the main Text to Speech page for the full breakdown of what's downloaded, how big it is, and how to remove it again.
Frequently asked questions
Which Vietnamese voice does this use?
A single-speaker Piper voice trained on the VAIS1000 dataset — a female voice at what Piper calls 'medium' quality, roughly comparable to a clear radio announcer rather than a big commercial voice assistant. It's the same underlying voice model whether you reach it from here or from the main text-to-speech tool; this page just starts you off already set to it.
Is the accent Northern, Central, or Southern Vietnamese?
It's trained on standard pronunciation of the kind used in national broadcasting, which leans closer to a Northern (Hanoi-style) accent than a Central or Southern one. If your text is written in standard Vietnamese spelling, it will be read in that standard accent regardless of which region you're writing about.
Will it read tone marks (dấu) correctly?
Yes, for text typed with full diacritics — tone is not an add-on for this model, it's read directly from the accents on each vowel. What it can't do is guess: Vietnamese written without diacritics loses the information the tones depend on, so paste in fully accented text for an accurate result.
I don't have a Vietnamese keyboard — can I still type accented text?
Yes. Any Telex or VNI input method (built into Windows, macOS, and most phones) types full Vietnamese diacritics on an ordinary keyboard, and you can also just paste in text you already have from elsewhere — a document, a chat, a webpage. The tool only needs the text already accented; it doesn't add the accents for you.
What's this useful for?
Reading Vietnamese study material aloud while learning the language, checking how a piece of Vietnamese writing actually sounds before recording it yourself, generating narration for a video, or making Vietnamese text accessible to someone who reads better by listening. Because it runs locally, none of that text has to leave your device to be read aloud.