Every modern browser and phone can read text aloud in dozens of languages and turn your speech into text. These tools are free and always available, which makes them tempting pronunciation coaches. Used well, they are genuinely useful; used naively, they can give you false confidence or false alarms. This guide explains what they can tell you, how to get better voices, and five routines that make the most of them.
Text-to-speech: a model you can replay forever
Text-to-speech (TTS) gives you an instant model for any word or sentence. The best current voices, often labelled "natural", "neural" or "enhanced", are close to human quality, with realistic stress and intonation. Older voices are intelligible but flat.
What TTS is good for:
- hearing the pronunciation of a word you have only seen written;
- slowing down a sentence to hear every syllable;
- repetition without tiring a teacher;
- hearing the same text in different accents (for example Spain and Mexico, or Brazil and Portugal).
Its limits: TTS sometimes misreads homographs (the French est "east" vs "is", the Portuguese sede "thirst" vs "headquarters"), gives isolated words an unnatural list intonation, and occasionally stresses the wrong syllable in rare words. Always check against a dictionary for important words.
How to get better voices
| Device | Where to add voices |
|---|---|
| Windows 10/11 | Settings → Time & language → Speech (or Language & region) → add a language and its speech pack. Microsoft Edge also offers online "Natural" voices automatically. |
| macOS | System Settings → Accessibility → Spoken Content → System voice → Manage voices. Download "Enhanced" or "Premium" versions. |
| iPhone / iPad | Settings → Accessibility → Spoken Content → Voices. |
| Android | Settings → System → Languages → Text-to-speech output → Speech Services by Google → Install voice data. |
| Chrome (desktop) | Includes Google network voices for many languages; no installation needed, but an internet connection is required. |
After installing, reload the page. The practice drills on this site automatically pick the best available voice for each language and warn you if none is installed.
Speech recognition: an outside listener
Automatic speech recognition (ASR) converts your speech into text. For pronunciation practice, it acts as a rough stand-in for a listener who does not know what you meant to say. If it transcribes your sentence correctly, a native listener probably would too. If it consistently gets one word wrong, that word deserves attention.
But recognisers are not accent graders. Research has found that ASR systems are less accurate for non-native speech (Derwing, Munro and Carbonaro 2000) and for some groups of native speakers (Koenecke and colleagues 2020). They also use context: if you say perro with a weak trill in a sentence about dogs, the recogniser may still write perro. So:
- a correct transcription means "probably intelligible", not "perfect";
- a wrong transcription means "check this", not necessarily "you said it wrong";
- numbers, names and rare words are often transcribed unpredictably.
Studies of learners using ASR for self-practice (McCrocklin 2016) found that its main benefit is autonomy: it lets you practise and notice problems between lessons.
Five practice routines
1. Listen, then read aloud
Play a sentence twice, then read it yourself, trying to match rhythm and melody. Repeat with the slow speed for difficult sounds.
2. Minimal pair listening
Use the listening drills in each guide: hear a word, pick which one it was. Switching between available voices adds useful variety.
3. Dictation
Listen to a sentence and type what you hear. This trains the listening side of pronunciation, especially linking and reduced forms. The dictation trainer does this in six languages.
4. Say it and check
Read a sentence into a recogniser and compare its transcription with the original. Focus on words it consistently mishears. The speaking drills in the guides on this site work this way.
5. Record and compare
Record yourself with your phone after listening to the TTS model, then play both back to back. This is the most direct way to hear the difference. See how to record yourself.
Phrases in this drill:
- Buenos días, ¿cómo estás?
- Me gusta mucho aprender idiomas.
Phrases in this drill:
- Bonjour, comment ça va ?
- J'aime beaucoup apprendre les langues.
A note on privacy
Speech synthesis runs on your device or, for network voices, on the voice provider's servers. Speech recognition in Chrome sends audio to Google for processing; Safari and Edge use Apple's and Microsoft's services. This site does not record or store your voice. If privacy is a concern for a particular session, use TTS-only practice and a local voice recorder.
Frequently asked questions
Can a speech recogniser grade my accent?
Not reliably. A recogniser tries to guess what you said, using both the sound and a language model of likely sentences. It can confirm that a listener would probably understand you, but it cannot tell you precisely what is wrong, and it may "correct" errors using context.
Why does the audio on this site sound robotic in some languages?
The audio uses voices installed in your browser or operating system. Their quality varies: newer "natural" or "neural" voices sound very human, while older ones sound mechanical. If no voice is installed for a language, the browser may use a default voice with the wrong accent.
Is my voice uploaded when I use speech recognition?
It depends on the browser. Chrome's Web Speech API sends audio to Google's servers for recognition; some browsers and operating systems can recognise speech on the device. This site does not store your recordings, but the browser vendor's privacy policy applies to the recognition itself.
Sources and further reading
- World Wide Web Consortium Community Group (2023). Web Speech API (draft specification). W3C.
- Derwing, T. M., Munro, M. J. and Carbonaro, M. (2000). Does popular speech recognition software work with ESL speech? TESOL Quarterly, 34(3), 592–603.
- McCrocklin, S. M. (2016). Pronunciation learner autonomy: the potential of automatic speech recognition. System, 57, 25–42.
- Koenecke, A. et al. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences, 117(14), 7684–7689.
Disclosure: this page links to Preply, a tutoring marketplace. If you book a lesson through that link we may earn a commission, at no extra cost to you. It does not change what is recommended here. See our disclaimer and privacy policy.
Practise with dictation
The dictation trainer reads sentences in six languages with your browser's voice. Type what you hear and see a word-by-word comparison.
Open the dictation trainer