Does listening to a language
help you learn it?
Yes for the sound system, far less for vocabulary, and the evidence separating the two is unusually clean. Nine-month-olds who heard Mandarin from a live person learned its phonetic contrasts in twelve sessions; infants given the identical material as audio-only recordings learned nothing measurable (Kuhl, Tsao & Liu, 2003). Following unscripted speech at 98% coverage takes roughly 6,000–7,000 word families (Nation, 2006), and a word needs five to twenty encounters before it stays (Nation & Wang, 1999) — which is the part an audio feed cannot schedule.
Every figure on this page is referenced at the bottom
What listening has actually been measured to do
“Just listen to the language” is advice given far more often than it is specified, which is a pity, because listening is one of the few learning behaviours that has been isolated in a laboratory and tested against itself. Researchers have compared a live speaker against a recording of the same speaker, counted the vocabulary that ordinary speech demands, measured what a short listening session leaves behind, and tested whether audio played during sleep does anything at all. The picture that comes out is sharper than either camp usually admits.
One caveat belongs at the top rather than in a footnote: the cleanest studies are about infants and phonology, not adults and vocabulary. A nine-month-old learning Mandarin tone contrasts is not a commuter learning Spanish nouns, and the incidental-listening studies that do use adults are short and mostly test recognition. What the numbers below are good for is the shape of the effect — which parts of a language arrive through the ear on their own and which never do — not a promise about your December.
| What was measured | What the study found | What it means for you |
|---|---|---|
| Whether audio alone teaches sound | Kuhl, Tsao & Liu (2003): nine-month-olds who heard Mandarin from a live person learned its phonetic contrasts; the groups who received the same material as audio-only or video recordings showed no measurable learning | The waveform was identical and only one condition worked. Audio playing in the room is the condition that taught nothing, which is worth knowing before you buy a year of it |
| What listening does teach, very fast | Saffran, Aslin & Newport (1996): eight-month-olds pulled word boundaries out of two minutes of continuous, meaningless speech, using nothing but the statistics of which syllable follows which | This is the genuine prize, and it is nearly free. Your ear starts extracting structure from a stream long before you understand a word of it |
| The vocabulary speech demands | Nation (2006): roughly 6,000–7,000 word families for 98% coverage of spoken text, and about 3,000 for 95% | There is a threshold and it dominates everything else. Below about 3,000 families, an episode is sound you cannot cut into words, however many hours you log |
| What an hour of listening leaves behind | van Zeeland & Schmitt (2013): incidental gains from listening to stories were real but modest, and a word's meaning was picked up far more readily than its form | You keep a hazy sense of words you already half-knew. The written form — the one you need to meet that word anywhere else — largely does not arrive |
| Whether the word can be recalled later | Karpicke & Roediger (2008): retrieval practice beat repeated study on later retention, by a margin that grew with the delay | Listening is re-exposure in bulk and recall never. That single gap explains the comfortable ear and the empty mouth better than any theory about talent |
Read the first two rows together. Listening is not weak — it is specific. It delivers the layer of a language that is almost impossible to study deliberately, and it is close to indifferent to the layer everyone actually measures progress by. The threshold in the third row is a number you can estimate for yourself: how many words do you need to speak a language?
Three claims travel under “just put on a podcast”
They rest on completely different evidence, which is why the advice manages to feel obviously true and quietly disappointing at the same time. Only the first two survive contact with the studies, and the second only above a threshold.
- The supported one: your ear rebuilds itself. This is the real payoff and it is not small. Extended listening trains you to cut a stream into words, to survive speed and accent, and to keep going when you miss something — and Saffran, Aslin and Newport (1996) showed the machinery running on two minutes of input in eight-month-olds. No flashcard has ever taught anyone this. If listening is your weak strand, audio is close to the ideal instrument.
- The conditional one: this is comprehensible input. True above the coverage threshold and false below it. Nation (2006) puts 95% coverage of speech at roughly 3,000 word families and 98% at 6,000–7,000; comprehension climbs steeply across that band. Below it the audio is not input in any useful sense, because you cannot recover a sentence when one word in ten is missing and you cannot hear where the others end.
- The false one: the words will arrive on their own. Pickup is frequency-gated. Peters and Webb (2018) found the number of times a word occurred predicted whether it was learned, and Nation and Wang (1999) put the encounters needed somewhere between five and twenty. A one-hour episode hands you most of its interesting vocabulary exactly once, which is precisely the case where the forgetting curve wins.
The general version of this question — what exposure of any kind delivers and where it stops — has its own page: does passive learning actually work?
Six things decide whether an hour of listening teaches you anything
Two people can listen to the same number of hours in a year and end up somewhere completely different. None of the six reasons is talent and only one of them is effort. The rest are properties of the material and the calendar — which is the good news, because those are the ones you can change.
| What decides it | What the evidence says | What to do |
|---|---|---|
| Whether anyone is actually there | Kuhl, Tsao & Liu (2003): live exposure produced phonetic learning; recorded exposure to the same material produced none | Treat recordings as practice, not instruction. Anything that makes you respond — a person, a pause, a question — is doing work the audio cannot |
| Your coverage of the material | Nation (2006): ~3,000 word families for 95% coverage of speech, ~6,000–7,000 for 98% | Under the threshold, spend part of the time building the frequent core. It is what makes every later hour of listening pay |
| How often a word comes round again | Peters & Webb (2018): frequency in the material predicted learning; Nation & Wang (1999): five to twenty encounters to acquire a word | Stay with one host, one series, one subject. Narrow listening turns single encounters into repeated ones at no extra cost |
| Whether you ever retrieve anything | Karpicke & Roediger (2008): retrieval practice beat repeated study on later retention | Audio is all confirmation and no recall. Something outside the feed has to ask you the question |
| The gap between encounters | Cepeda et al. (2006), 254 studies: around 47% recall spaced against 37% massed | What returns, and when, is decided by an editor's running order. No playlist has ever been sequenced by what you forgot |
| What happens between episodes | Murre & Dros (2015): the classic forgetting curve replicates; what is not retrieved decays fastest right after the encounter | The word you looked up on Tuesday has to come back before Friday, and next week's episode will not bring it |
Notice what is missing from that list: concentration, discipline and aptitude. None of this fails because you were not listening hard enough. It fails because a recording has no idea which word you missed — and that indifference is exactly what makes listening pleasant enough to keep doing for years.
The ear is real. The second meeting is what’s missing.
Listening gets right the thing almost every deliberate method gets wrong: it is sustainable. Nobody has to be talked into an episode on the commute, nobody breaks a streak, and the input arrives in volume, in a format the brain will accept for forty minutes at a stretch. Nation’s four strands (2007) put meaning-focused input at roughly a quarter of a balanced programme, and audio fills that quarter in the parts of a day — the walk, the dishes, the gym — where nothing else can reach. Keep it. The argument here is not that podcasts are a waste of time; it is that they are being asked to do a job they were never shaped for.
What they cannot do is come back. An episode decides what you hear next according to its own running order; it has no idea which word made you rewind, and no way to reintroduce that word three days later, which is when reintroducing it would be worth most. So the words the host happens to repeat are learned nearly for free, and the ones that appeared once — usually the interesting mid-frequency ones, the ones you actually needed — decay along the ordinary curve (Murre & Dros, 2015). A few hundred hours in, the profile is familiar to anyone who has done it: comfortable comprehension, an ear that copes with speed, and a vocabulary parked around whatever the show says every week.
The repair is small, and it is not more listening. It is a second meeting with the specific words that slipped, spaced out, and shaped as a question rather than a replay — the combination Cepeda and colleagues (2006) measured at around 47% recall against 37% for massed practice, and that Karpicke and Roediger (2008) found beat re-reading outright. The awkward part has always been where to put it, because a review session is one more thing to schedule and scheduling is the exact step people fail at. Which argues for putting it on a surface that already repeats itself without anyone’s cooperation: the phone gets picked up about 186 times a day (Reviews.org, 2026), and not one of those pickups has to be planned.
Podcasts, music, background radio, sleep audio
These get recommended in one breath as though they were the same intervention, and the research treats them very differently. Two of them are supported, one is supported for a reason people usually get backwards, and one works only in a sense narrow enough that it is worth stating precisely.
The pattern underneath is the same each time: the more the format asks something of you, the more of it you keep.
| Format | What the evidence supports | What it will not do |
|---|---|---|
| Podcasts and audiobooks at your level | van Zeeland & Schmitt (2013): real if modest incidental gains from listening to stories, above the coverage band Nation (2006) sets at ~3,000 word families for 95% | Hand you the written form, or bring back the word you missed. Meaning is picked up far more readily than form, and neither is scheduled |
| Music and song | Ludke, Ferreira & Schellenberg (2014): adults who sang Hungarian phrases recalled them better afterwards than adults who spoke the same phrases | Work with the music merely on. The active ingredient in that study was producing the language; lyrics in the background is the untested condition, not the tested one |
| Background radio while you work | Saffran et al. (1996): phonological structure is extracted from a stream remarkably cheaply, so familiarity with the sound of the language is a genuine gain | Teach vocabulary. This is closest to the recorded-audio condition in Kuhl et al. (2003), the one that produced no measurable learning |
| Audio played during sleep | Schreiner & Rasch (2015): replaying words participants had already studied before sleeping improved recall of exactly those words | Teach a word you have never met. That result is reactivation of an existing memory, so it can finish a job you started awake and cannot start one |
The sleep row is the one most often quoted without its condition, so it is worth repeating: the words worked because they had been studied first. If you want the timing question on its own, it has a page: should you study vocabulary before bed?
A place for the words the episode said once
That gap between hearing a word and being able to produce it is what LearnScreen was built for. Keep the podcast. The words you rewound for come back on a surface you were already going to look at.
The word you rewound for becomes a card
Add it in a few seconds from the Share sheet, or paste a whole list after an episode, each with a translation and an example sentence. It joins the same queue as the built-in lists — so the word the host said exactly once starts behaving like a word you meet often, and it finally arrives with the written form listening withheld.
It comes back as a question, not a replay
Using Apple’s Screen Time API, LearnScreen shields the apps you choose and puts the card where the feed would have been, answer hidden until you commit to an attempt. That turns each encounter into retrieval — the operation Karpicke and Roediger (2008) found beats re-exposure, and the one an audio feed structurally cannot perform.
A Leitner schedule decides what returns
Words you miss come back sooner; words you get right back off geometrically. Spacing follows how well you know each word rather than an editor’s running order — the part the distributed-practice research (Cepeda et al., 2006) says does the work, and the part no playlist has ever been able to supply.
- Nothing extra gets scheduled. The review lands on pickups you were already making, which is why it survives the weeks when an episode on the walk to work is the most you can manage. Defaults come to roughly 25 recall attempts a day — about two and a half minutes, spread across moments you were not using for anything.
- Start from the frequent core if you are under the threshold. 1,080+ curated words across 18 topics in 11 languages, ordered so the cheapest coverage comes first — the fastest route to the ~3,000 families that make ordinary speech legible (which words come first).
- Your own words, in bulk. Add one from the Share sheet mid-episode or paste the list you built on the commute; every entry takes a translation and an example sentence, so the card keeps the sense the episode gave it.
- Offline, no account. Cards and shields run without a network once installed, and iCloud backup writes to your own private database rather than our servers.
Related reading
- Does passive learning actually work? — what exposure of any kind delivers, measured, and where it stops.
- Does watching TV in another language help? — the same question when there is a picture and a subtitle track.
- Can you learn while doing something else? — which half of learning survives a divided attention.
- How many words do you need to speak a language? — the coverage thresholds that decide whether audio is input or noise.
- Active vs passive vocabulary — why the words you recognise outnumber the words you can say.
- Why do I forget vocabulary I’ve already learned? — what happens to a word between episodes.
- How often should you review vocabulary? — the intervals a running order will never supply.
Frequently asked questions
Sources
- Kuhl, P. K., Tsao, F.-M., & Liu, H.-M. (2003). Foreign-language experience in infancy: Effects of short-term exposure and social interaction on phonetic learning. PNAS, 100(15), 9096–9101.
- Saffran, J. R., Aslin, R. N., & Newport, E. L. (1996). Statistical learning by 8-month-old infants. Science, 274(5294), 1926–1928.
- van Zeeland, H., & Schmitt, N. (2013). Incidental vocabulary acquisition through L2 listening: A dimensions approach. System, 41(3), 609–624.
- Nation, I. S. P. (2006). How large a vocabulary is needed for reading and listening? Canadian Modern Language Review, 63(1), 59–82.
- Peters, E., & Webb, S. (2018). Incidental vocabulary acquisition through viewing L2 television and factors that affect learning. Studies in Second Language Acquisition, 40(3), 551–577.
- Nation, I. S. P., & Wang, K. (1999). Graded readers and vocabulary. Reading in a Foreign Language, 12(2), 355–380.
- Schreiner, T., & Rasch, B. (2015). Boosting vocabulary learning by verbal cueing during sleep. Cerebral Cortex, 25(11), 4169–4179.
- Ludke, K. M., Ferreira, F., & Schellenberg, E. G. (2014). Singing can facilitate foreign language learning. Memory & Cognition, 42(1), 41–52.
- Karpicke, J. D., & Roediger, H. L. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966–968.
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
- Murre, J. M. J., & Dros, J. (2015). Replication and analysis of Ebbinghaus’ forgetting curve. PLOS ONE, 10(7), e0120644.
- Nation, I. S. P. (2007). The four strands. Innovation in Language Learning and Teaching, 1(1), 2–13.
- Reviews.org (2026). Cell Phone Usage Stats. Survey of ~1,000 US adults, fielded Q4 2025. Report
Three limits are worth carrying away with the numbers. The phonetic-learning studies are about infants, and infants are not adults in a smaller format — what transfers is the comparison between conditions, not the age. The incidental-listening studies are short and test recognition rather than production, so they understate what years of listening does and overstate how measurable a single hour is. And the sleep finding applies only to words already studied before sleeping. Where this page states a number it comes from the study named beside it; where it generalises to your own listening, that is an inference, and it is marked as one.
Keep listening. Don’t let the words leave with the episode.
The word you rewound for, back as a question, on a screen you already unlock.
Download on the App Store