Text to Speech (free) — voice text in your browser, nothing uploaded

Bottom Line Up Front (BLUF): Text to speech turns typed text into voice right in your browser: enter the text, pick a language and a voice, and get an audio file. The text and the audio never leave your device, no tokens are charged and there is no paid API here. The model knows 600+ languages and Russian is fully supported, and a voice can be cloned from your own audio — an uploaded file or a microphone recording. The result plays in the player and downloads as a WAV file.

What text to speech does and when you need it

It is speech synthesis: text goes in, sound comes out. Use it to voice a video with your own or a chosen voice, build a draft voice-over track, check how a text sounds in another language, or simply listen to a long article instead of reading it.

Delivery is controlled with a text instruction: calm reading, a whisper, joy or sadness. The emotion is set together with the text rather than with sliders, so the result sounds more natural.

Ready-made voices and your own cloned voice

The tool ships with a library of ready voices — male and female, different timbres. Pick a language and a voice and start voicing.

If you need your own voice, clone it from a short sample: upload an audio file or record 15–30 seconds from the microphone. The tool analyses the sample and adds the voice to your list — after that it is chosen like any other. Created voices are kept in the browser and survive a page reload.

Everything runs on your device

The text and the audio never leave your computer: not a line goes to our server or to anyone else's. The model is loaded into the browser once (about 3 GB of weights) and is then taken from the cache, while the computation runs on your graphics card through WebGPU. If the graphics card is unavailable, the engine honestly falls back to the processor and tells you about it — that is noticeably slower.

No tokens are charged for speech: it is a free tool that runs on your hardware, not on ours.

Long text is voiced in parts

The input field holds up to 20,000 characters. Long text is split into parts at sentence boundaries and voiced one after another with short pauses — the pace stays even regardless of length and nothing is lost at the end. The progress bar shows which part is being read.

You can cancel at any step: after cancelling, a new run starts right away, with no page reload.

What you need

The tool works best in Chrome or Edge on a desktop, where WebGPU is available. In browsers without WebGPU the voice is generated on the processor (WASM): it still works, but noticeably slower, especially for long text.

The first run takes longer: the weights are downloaded and the voice is warmed up so that the very first result already sounds clean. Later runs are much faster.

Soalan lazim

Do the text and the audio really stay off the server?

Yes. The model and its weights are loaded into the browser and the computation runs on your device. The text, the microphone recording and the resulting audio are never uploaded — the tool has no paid API at all.

How much does the speech cost?

Nothing. No tokens are charged: it is a free tool that computes on your computer.

Can I voice a long text?

Yes. The input field holds up to 20,000 characters, and long text is voiced in consecutive parts so nothing is lost.

How do I create my own voice?

Upload an audio file with a voice sample or record 15–30 seconds from the microphone. The tool analyses the sample and the voice appears in your list — you can voice any text with it.

Why is the speech slow and why does it say the processor is working?

It means the browser could not use the graphics card (WebGPU) and the engine switched to the processor — slower, but the same result. You can reload the page and try again: the graphics card is often free by then.

Which languages are supported?

600+ languages, including Russian. You can pick the language manually or leave automatic detection on.

Model serupa