Type or paste your text - or upload a transcript file (.srt, .vtt, .sbv, .txt) - pick one of ten studio voices, and press Generate speech. The voice model runs inside this browser tab, so the same voice comes out on a phone, a laptop, or a work desktop, the audio is 44.1 kHz, and you can download it as MP3 or WAV. Nothing is uploaded.
The first run downloads the voice model, about 380 MB, and keeps it in this browser; after that the tool starts immediately and still works with the network off. If that download is too heavy right now, the quick device voice is one click away.
Prefer recording your own voice? Capture it with the voice recorder.
Text to Speech
Paste your text, choose one of ten voices, and press Generate speech: a neural voice model runs inside this tab and hands you a 44.1 kHz audio track you can play or download as MP3 or WAV.
Typical use: turning an article, a script, a lesson, or a subtitle file into narration you can listen to or drop into a video editor, without sending the words to anyone.
Ten voices that sound the same everywhere
The voices are M1 to M5 and F1 to F5, and they come from the downloaded model rather than from your operating system. That is the practical difference from the voices built into a browser: a narration you generate on Windows sounds identical when a colleague regenerates it on a Mac or an Android phone, so a series of clips stays consistent. Speed runs from 0.8x to 1.5x, and the quality control picks how many refinement steps the model takes - 4 for a quick draft, 8 for everyday use, 16 when the result is going into something published.
What the first load costs, and why it pays off
The model is about 380 MB across four files. The tool shows the real percentage as it downloads, then keeps the files in this browser's cache, so the second visit starts in seconds and synthesis continues to work with the network switched off. Generation runs on your GPU through WebGPU when the browser supports it and on the CPU through WebAssembly otherwise; the status line names which one is in use, because CPU mode is noticeably slower on long text. On a device that cannot hold the model - a low-memory phone, an older browser - the page says so and offers the quick device voice instead of pretending.
Read a transcript or subtitle file aloud
Upload a .srt, .vtt, .sbv, or plain .txt file and it is parsed here in the tab. Subtitle formats are split into timed cues, each listed with its time range, and the text lands in the box so you can edit it before speaking. Generate the whole transcript on its timeline goes one step further: every line is rendered and placed at its own timecode in a single audio file, so a two-line file whose second cue starts at five seconds produces audio where that line really begins at five seconds. That file drops straight onto a video timeline without re-syncing.
Languages and automatic detection
The model covers 31 languages, among them English, Spanish, Portuguese, French, German, Italian, Dutch, Polish, Czech, Romanian, Swedish, Finnish, Hungarian, Greek, Turkish, Ukrainian, Russian, Hindi, Arabic, Indonesian, Japanese, Korean, and Vietnamese. Leave the language on Detect automatically and the writing system plus common-word patterns pick one offline; set it by hand when the text is short or mixes languages, because a heuristic can get that wrong. Unlike the previous build, a downloadable file is available in every one of those languages, not only in the handful a compact engine shipped voices for.
Where the audio and the licence stand
Your text, your transcript, and the finished audio stay in the browser; there is no upload step and no account. The voice model is Supertonic 3 by Supertone, published under the OpenRAIL-M licence, which the panel links along with the credits file. That licence carries use restrictions worth respecting: do not use these voices to imitate a real person without their consent, and do not use them to mislead people about who is speaking.
Frequently Asked Questions
What does Text to Speech do?
Type or paste text - or upload a subtitle file - pick one of ten voices, and press Generate speech. A neural voice model runs inside this browser tab and produces a 44.1 kHz track you can play in the page or download as MP3 or WAV. Nothing is uploaded.
Why does the first run download 380 MB?
Because the voice itself is the download. The four model files (about 380 MB in total) are fetched once from our asset server, then kept in this browser's cache, so later visits start in seconds and synthesis keeps working with the network off. The tool shows the real percentage while it downloads, and the quick device voice is available if you would rather not fetch it now.
Do the voices sound the same on every device?
Yes. The ten styles - M1 to M5 and F1 to F5 - come from the downloaded model, not from your operating system, so a clip generated on Windows matches one regenerated on a Mac or an Android phone. That is the difference from the browser's built-in voices, which vary by platform.
Can I download the speech as an MP3 or WAV?
Yes, in every supported language. After generating, the download button produces a 128 kbps MP3 or a 44.1 kHz WAV, encoded here in the page. Nothing is sent anywhere to make the file.
Which languages are supported?
Thirty-one, including English, Spanish, Portuguese, French, German, Italian, Dutch, Polish, Czech, Romanian, Swedish, Finnish, Hungarian, Greek, Turkish, Ukrainian, Russian, Hindi, Arabic, Indonesian, Japanese, Korean, and Vietnamese, plus a neutral tag for anything else. Leave detection on automatic, or set the language by hand for short or mixed text.
Can I keep the original subtitle timing?
Yes. Upload a .srt, .vtt, or .sbv file and use Generate the whole transcript on its timeline: every line is rendered and placed at its own timecode in one audio file, so a cue that starts at five seconds really starts at five seconds. The file can go straight onto a video timeline.
What do the speed and quality controls change?
Speed runs from 0.8x to 1.5x and stretches or compresses the predicted timing. Quality picks how many refinement steps the model takes: 4 renders fastest for a draft, 8 is the everyday setting, and 16 gives the smoothest result but takes longest, especially in CPU mode.
What if my device cannot run the model?
The page tells you and shows the quick device voice instead: playback through the voices your browser already ships, plus a compact eSpeak file generator. It is the fallback, not the main path - the voice varies by device and sounds more synthetic.
Is there anything I should not do with these voices?
The model is Supertonic 3 by Supertone under the OpenRAIL-M licence, linked in the panel together with the credits. Its use restrictions apply: do not imitate a real person without their consent, and do not use generated speech to mislead people about who is speaking.
What complementary tools work well alongside text to speech?
If you need the reverse direction, look at the related-tools section for the SPEECH-to-TEXT converter.