Text to Speech
Convert text to speech.
How to Convert Text to Speech
Type or paste your text
Enter anything from a sentence to whole paragraphs.
Pick a voice and speed
Choose from the voices installed on your device and adjust rate and pitch.
Listen instantly
The browser's speech engine reads the text aloud โ no account, no quota.
The voices come from your device
The browser does not ship voices of its own. It exposes whatever the operating system provides, so the list you see depends on your platform, your language settings, and which voice packs are installed. The same page offers different voices on Windows, macOS, Android and iOS.
This is why a voice a colleague recommends may not appear for you, and why the quality varies so widely. Some system voices are modern neural models that sound close to human; others are decades-old formant synthesisers that sound exactly like what people picture when they hear 'text to speech'.
Matching the voice to the language
A voice is built for a specific language, and giving it text in another produces something between an accent and nonsense. Turkish read by an English voice mispronounces almost every word, because the letter-to-sound rules are entirely different.
Mixed-language text is the harder case with no good answer: a Turkish paragraph containing English product names will have one or the other pronounced wrongly, whichever voice you pick. Splitting the text and reading each part with a matching voice is the only reliable fix.
Rate, pitch and comprehension
The default speaking rate is deliberately slow for anyone familiar with synthetic speech. Regular users of screen readers often run at two or three times normal speed, because comprehension adapts quickly with practice.
Pitch adjustment is available but rarely improves anything โ pushed far in either direction it makes a voice harder to understand rather than more pleasant. Rate is the setting worth adjusting; pitch is mostly a novelty.
Where synthesis reliably stumbles
Abbreviations are ambiguous and always will be: whether 'Dr.' is doctor or drive, whether 'St.' is street or saint, cannot be determined without understanding the sentence. Numbers are similar โ a year, a quantity and a phone number are read differently and look identical.
Homographs are the deepest problem, since the correct pronunciation depends on meaning rather than spelling. Where a word must be said correctly, spelling it phonetically in the input text is cruder than it sounds and works better than anything else available.
Practical uses, and one limitation
Proofreading by ear is the use that surprises people most: hearing your own writing read aloud exposes clumsy sentences and missing words that reading silently glides over. It is also genuinely useful for language learning, for accessibility, and for consuming text while doing something else.
Synthesis happens on your device, so the text is not uploaded. The limitation is that browser speech synthesis plays audio rather than producing a file โ there is no download to keep. If you need an audio file, that requires software built for the purpose.