Text to speech (TTS)
Turn text into natural speech (TTS) in your browser — realistic on-device AI voices you can download as WAV or MP3, or your device’s built-in voices for instant multilingual playback.
Open Text to speech (TTS) →What is the text to speech tool?
A converter that turns any text into spoken audio, entirely in your browser. It has two engines: realistic neural AI voices (the Kokoro model, running on your device) that you can preview and download as WAV or MP3, and your device's own built-in voices for instant live playback with adjustable rate, pitch and volume. Nothing you type ever leaves your browser.
How to use Text to speech
- Pick an engine — In the sidebar, choose Neural AI voices for realistic, downloadable speech, or Device voices for instant playback in your device's installed languages.
- Enter your text — Type or paste the text you want spoken into the box.
- Choose a voice and delivery — In the sidebar, pick a neural voice (American or British English) and set the speed — or, with device voices, pick a voice and adjust rate, pitch and volume.
- Generate or play — Press Generate speech to create the audio on your device, or Play to hear it live through your speakers.
- Preview and download — With the neural engine, listen to the result in the built-in player, then save it as a WAV or MP3 file using the download button at the top of the page.
Frequently asked questions
Can I download the audio as a file?
Yes — with the Neural AI voices engine, every generation can be saved as a lossless WAV or a smaller MP3 from the download button at the top of the page. Device voices play live only, because the Web Speech API cannot capture or export its audio.
How realistic are the neural voices?
Very — they come from Kokoro, an open neural speech model with 28 named voices in American and British English, and sound far more natural than typical built-in system voices.
Does it need a download or an internet connection?
The neural voice model (about 92 MB) downloads once from the Hugging Face CDN the first time you press Generate, then it is cached in your browser and later generations work offline. Device voices need no download at all.
Is my text uploaded anywhere?
No. Both engines run entirely on your device — the neural model generates the audio locally in your browser, and nothing you type is ever sent to a server.
Can it speak languages other than English?
The neural voices speak English (US and UK accents) only. For other languages, switch to Device voices, which use whatever text-to-speech voices your operating system provides.
Why are no device voices showing?
If your device has no speech voices installed, the device-engine controls may stay disabled. Add at least one text-to-speech voice through your operating system's settings and reload.
Tips
- The voices marked best in the picker (Heart and Bella) have the highest quality grades — start with those.
- Download WAV for editing or archiving and MP3 for sharing — the MP3 is roughly a tenth of the size.
- Lower the speed slightly when proofreading — hearing text read slowly helps catch typos and clumsy phrasing.
- Long texts are generated sentence by sentence with a progress bar — you can keep working in another tab while it runs.