Most learners treat pronunciation as something that will sort itself out with time. For vocabulary and grammar that is partly true; for sounds it usually is not. Adults tend to hear a new language through the filter of their first one, and without deliberate practice many errors fossilise. The good news from decades of research is that a few targeted techniques work well at any age. This guide turns them into a method you can use for any language.
Why adults find new sounds hard
Your first language taught your brain which sound differences matter. English speakers learned that the difference between aspirated and unaspirated p does not change meaning, so they stopped paying attention to it. Hindi and Mandarin speakers learned the opposite. When you learn a new language, you initially perceive its sounds through your own categories, a process described in both Flege's Speech Learning Model and Best's Perceptual Assimilation Model.
The practical consequence is that sounds that are similar to a sound in your language, like French /u/ vs English /uː/, are often harder to master than sounds that are completely new, like the Xhosa clicks. Your brain files the similar sound in an existing category and stops listening closely. Flege's 1987 study of English learners of French showed this clearly: experienced learners produced the new vowel /y/ of tu fairly accurately, but kept an English-like quality in the similar vowel /u/ of tout.
Step 1: map the sound system
Before practising, find out which sounds of the target language differ from yours. A good language overview (like the guides on this site for Spanish, French, German, Portuguese, Mandarin and Hindi) should tell you:
- which vowels and consonants have no equivalent in your language;
- which contrasts your language ignores (length, aspiration, tone, nasality);
- how stress, rhythm and intonation work;
- which spelling conventions will mislead you.
Make a short list of your top five problems. Prioritise the ones that change meaning most often, a principle called functional load.
Step 2: train your ear first
The most robust finding in pronunciation research is that perception training works. In a classic series of studies, Japanese speakers who could barely tell English /r/ from /l/ improved substantially after a few weeks of listening to minimal pairs spoken by many different voices, and their pronunciation improved too, without any speaking practice (Bradlow and colleagues, 1997). This technique is called high variability phonetic training (HVPT). Its key ingredients are:
- Minimal pairs: words that differ only in the target contrast (tu/tout, pero/perro, mā/mà).
- Identification, not just listening: hear a word, choose which one it was, get immediate feedback.
- Many voices: different speakers, so you learn the category rather than one person's voice.
- Short, repeated sessions over several weeks.
The listening drills in the guides on this site follow this format. To add variability, switch between the different voices your browser or phone offers for the language.
Step 3: learn the articulation
Once you can hear a sound, learn how it is physically made: where the tongue goes, what the lips do, whether the vocal folds vibrate. Explicit articulatory instruction speeds up learning for many adults, especially for sounds such as retroflex consonants, front rounded vowels or the uvular R. Use a mirror, a finger on your throat to feel voicing, or a tissue in front of your mouth to check aspiration.
Step 4: produce in growing chunks
Practise a new sound in stages:
- on its own, if possible;
- in a single syllable (ru, ra);
- in short, common words;
- in phrases, where rhythm and linking affect it;
- in free speech, where you are thinking about meaning, not sounds.
Accuracy usually drops at each stage, and that is normal. A sound is only learned when it survives stage 5.
Step 5: record and compare
You cannot hear yourself accurately while speaking. Record yourself saying the same sentence as a model, then listen to both, one after the other. Focus on one feature at a time: vowel quality, then stress, then intonation. A speech recogniser can act as a rough "outside listener": if it consistently writes the wrong word, a native listener may struggle too. Treat it as a hint, not a verdict.
Step 6: shadow real speech
Shadowing means repeating a recording almost simultaneously, a fraction of a second behind the speaker. It trains rhythm, linking and intonation, which are hard to practise word by word. Start with short clips at slow speed, and use a transcript at first. See the shadowing guide.
What to prioritise
Research on intelligibility (Derwing and Munro 2015) suggests focusing on features that most affect whether you are understood: the sound contrasts with the highest functional load, word stress in stress languages, tones in tone languages, and overall rhythm. A slight accent on less important sounds rarely causes misunderstanding. See accent vs intelligibility.
A sample weekly routine
| Day | Activity (10–15 minutes) |
|---|---|
| Mon, Wed, Fri | Listening drill on one contrast (HVPT style), then 5 minutes of articulation practice |
| Tue, Thu | Record and compare 5 sentences containing the target sound |
| Sat | Shadow a 1–2 minute clip |
| Sun | Free conversation or speaking aloud; note sounds that broke down |
After two or three weeks, move on to the next item on your list, and return to earlier ones briefly each week.
Check your understanding
Choose an answer to see the explanation.
1. Which kind of new sound is often hardest for adult learners?
- Completely new sounds
- Sounds similar to one in their first language
- Sounds in unstressed syllables
Answer: Sounds similar to one in their first language. Flege's research suggests learners assimilate similar sounds to existing categories and struggle to form new ones.
2. What is a key feature of high variability phonetic training?
- Listening to one clear voice
- Identifying minimal pairs spoken by many voices
- Reading IPA transcriptions
Answer: Identifying minimal pairs spoken by many voices. Variability across speakers helps learners build a robust category.
3. When is a new sound truly learned?
- When you can say it alone
- When you can say it in words
- When it survives in spontaneous speech
Answer: When it survives in spontaneous speech. Accuracy usually drops as tasks become less controlled; the goal is automatic use in real speech.
Frequently asked questions
Is it too late to learn good pronunciation as an adult?
No. Adults rarely become indistinguishable from native speakers, but research consistently shows that adults can become highly intelligible and can improve specific sounds with targeted training at any age.
Should I learn to hear sounds before I try to say them?
Usually, yes. Studies of perception training show that learning to hear a contrast often improves production of it too, even without speaking practice. If you cannot hear the difference between two sounds, you will struggle to produce it consistently.
How long should I practise pronunciation each day?
Short, frequent sessions work better than long, rare ones. Ten to fifteen minutes of focused practice daily, combined with lots of listening, is a realistic and effective routine.
Sources and further reading
- Flege, J. E. (1995). Second language speech learning: theory, findings, and problems. In W. Strange (ed.), Speech Perception and Linguistic Experience, 233–277. York Press.
- Flege, J. E. (1987). The production of "new" and "similar" phones in a foreign language: evidence for the effect of equivalence classification. Journal of Phonetics, 15(1), 47–65.
- Best, C. T. and Tyler, M. D. (2007). Nonnative and second-language speech perception: commonalities and complementarities. In O.-S. Bohn and M. J. Munro (eds.), Language Experience in Second Language Speech Learning, 13–34. John Benjamins.
- Thomson, R. I. (2018). High variability [pronunciation] training (HVPT): a proven technique about which every language teacher and learner ought to know. Journal of Second Language Pronunciation, 4(2), 208–231.
- Derwing, T. M. and Munro, M. J. (2015). Pronunciation Fundamentals: Evidence-Based Perspectives for L2 Teaching and Research. John Benjamins.
- Bradlow, A. R., Pisoni, D. B., Akahane-Yamada, R. and Tohkura, Y. (1997). Training Japanese listeners to identify English /r/ and /l/: IV. Some effects of perceptual learning on speech production. Journal of the Acoustical Society of America, 101(4), 2299–2310.
Disclosure: this page links to Preply, a tutoring marketplace. If you book a lesson through that link we may earn a commission, at no extra cost to you. It does not change what is recommended here. See our disclaimer and privacy policy.
Start with a language you are learning
Each language page has a sound-system overview, audio, and a recorder that shows which words a speech recogniser understood.
Choose a language