Skip to content
The listening research, applied to your word list

Does listening to a language
help you learn it?

Yes for the sound system, far less for vocabulary, and the evidence separating the two is unusually clean. Nine-month-olds who heard Mandarin from a live person learned its phonetic contrasts in twelve sessions; infants given the identical material as audio-only recordings learned nothing measurable (Kuhl, Tsao & Liu, 2003). Following unscripted speech at 98% coverage takes roughly 6,000–7,000 word families (Nation, 2006), and a word needs five to twenty encounters before it stays (Nation & Wang, 1999) — which is the part an audio feed cannot schedule.

Every figure on this page is referenced at the bottom

A vocabulary card shown on an iPhone lock screen at the moment a blocked app is opened

What listening has actually been measured to do

“Just listen to the language” is advice given far more often than it is specified, which is a pity, because listening is one of the few learning behaviours that has been isolated in a laboratory and tested against itself. Researchers have compared a live speaker against a recording of the same speaker, counted the vocabulary that ordinary speech demands, measured what a short listening session leaves behind, and tested whether audio played during sleep does anything at all. The picture that comes out is sharper than either camp usually admits.

One caveat belongs at the top rather than in a footnote: the cleanest studies are about infants and phonology, not adults and vocabulary. A nine-month-old learning Mandarin tone contrasts is not a commuter learning Spanish nouns, and the incidental-listening studies that do use adults are short and mostly test recognition. What the numbers below are good for is the shape of the effect — which parts of a language arrive through the ear on their own and which never do — not a promise about your December.

Five findings from research on learning a language by listening, and what each implies for a learner
What was measuredWhat the study foundWhat it means for you
Whether audio alone teaches soundKuhl, Tsao & Liu (2003): nine-month-olds who heard Mandarin from a live person learned its phonetic contrasts; the groups who received the same material as audio-only or video recordings showed no measurable learningThe waveform was identical and only one condition worked. Audio playing in the room is the condition that taught nothing, which is worth knowing before you buy a year of it
What listening does teach, very fastSaffran, Aslin & Newport (1996): eight-month-olds pulled word boundaries out of two minutes of continuous, meaningless speech, using nothing but the statistics of which syllable follows whichThis is the genuine prize, and it is nearly free. Your ear starts extracting structure from a stream long before you understand a word of it
The vocabulary speech demandsNation (2006): roughly 6,000–7,000 word families for 98% coverage of spoken text, and about 3,000 for 95%There is a threshold and it dominates everything else. Below about 3,000 families, an episode is sound you cannot cut into words, however many hours you log
What an hour of listening leaves behindvan Zeeland & Schmitt (2013): incidental gains from listening to stories were real but modest, and a word's meaning was picked up far more readily than its formYou keep a hazy sense of words you already half-knew. The written form — the one you need to meet that word anywhere else — largely does not arrive
Whether the word can be recalled laterKarpicke & Roediger (2008): retrieval practice beat repeated study on later retention, by a margin that grew with the delayListening is re-exposure in bulk and recall never. That single gap explains the comfortable ear and the empty mouth better than any theory about talent

Read the first two rows together. Listening is not weak — it is specific. It delivers the layer of a language that is almost impossible to study deliberately, and it is close to indifferent to the layer everyone actually measures progress by. The threshold in the third row is a number you can estimate for yourself: how many words do you need to speak a language?

Three claims travel under “just put on a podcast”

They rest on completely different evidence, which is why the advice manages to feel obviously true and quietly disappointing at the same time. Only the first two survive contact with the studies, and the second only above a threshold.

  • The supported one: your ear rebuilds itself. This is the real payoff and it is not small. Extended listening trains you to cut a stream into words, to survive speed and accent, and to keep going when you miss something — and Saffran, Aslin and Newport (1996) showed the machinery running on two minutes of input in eight-month-olds. No flashcard has ever taught anyone this. If listening is your weak strand, audio is close to the ideal instrument.
  • The conditional one: this is comprehensible input. True above the coverage threshold and false below it. Nation (2006) puts 95% coverage of speech at roughly 3,000 word families and 98% at 6,000–7,000; comprehension climbs steeply across that band. Below it the audio is not input in any useful sense, because you cannot recover a sentence when one word in ten is missing and you cannot hear where the others end.
  • The false one: the words will arrive on their own. Pickup is frequency-gated. Peters and Webb (2018) found the number of times a word occurred predicted whether it was learned, and Nation and Wang (1999) put the encounters needed somewhere between five and twenty. A one-hour episode hands you most of its interesting vocabulary exactly once, which is precisely the case where the forgetting curve wins.

The general version of this question — what exposure of any kind delivers and where it stops — has its own page: does passive learning actually work?

Six things decide whether an hour of listening teaches you anything

Two people can listen to the same number of hours in a year and end up somewhere completely different. None of the six reasons is talent and only one of them is effort. The rest are properties of the material and the calendar — which is the good news, because those are the ones you can change.

Six factors that decide how much you gain from listening, the evidence behind each, and what to do about it
What decides itWhat the evidence saysWhat to do
Whether anyone is actually thereKuhl, Tsao & Liu (2003): live exposure produced phonetic learning; recorded exposure to the same material produced noneTreat recordings as practice, not instruction. Anything that makes you respond — a person, a pause, a question — is doing work the audio cannot
Your coverage of the materialNation (2006): ~3,000 word families for 95% coverage of speech, ~6,000–7,000 for 98%Under the threshold, spend part of the time building the frequent core. It is what makes every later hour of listening pay
How often a word comes round againPeters & Webb (2018): frequency in the material predicted learning; Nation & Wang (1999): five to twenty encounters to acquire a wordStay with one host, one series, one subject. Narrow listening turns single encounters into repeated ones at no extra cost
Whether you ever retrieve anythingKarpicke & Roediger (2008): retrieval practice beat repeated study on later retentionAudio is all confirmation and no recall. Something outside the feed has to ask you the question
The gap between encountersCepeda et al. (2006), 254 studies: around 47% recall spaced against 37% massedWhat returns, and when, is decided by an editor's running order. No playlist has ever been sequenced by what you forgot
What happens between episodesMurre & Dros (2015): the classic forgetting curve replicates; what is not retrieved decays fastest right after the encounterThe word you looked up on Tuesday has to come back before Friday, and next week's episode will not bring it

Notice what is missing from that list: concentration, discipline and aptitude. None of this fails because you were not listening hard enough. It fails because a recording has no idea which word you missed — and that indifference is exactly what makes listening pleasant enough to keep doing for years.

The ear is real. The second meeting is what’s missing.

Listening gets right the thing almost every deliberate method gets wrong: it is sustainable. Nobody has to be talked into an episode on the commute, nobody breaks a streak, and the input arrives in volume, in a format the brain will accept for forty minutes at a stretch. Nation’s four strands (2007) put meaning-focused input at roughly a quarter of a balanced programme, and audio fills that quarter in the parts of a day — the walk, the dishes, the gym — where nothing else can reach. Keep it. The argument here is not that podcasts are a waste of time; it is that they are being asked to do a job they were never shaped for.

What they cannot do is come back. An episode decides what you hear next according to its own running order; it has no idea which word made you rewind, and no way to reintroduce that word three days later, which is when reintroducing it would be worth most. So the words the host happens to repeat are learned nearly for free, and the ones that appeared once — usually the interesting mid-frequency ones, the ones you actually needed — decay along the ordinary curve (Murre & Dros, 2015). A few hundred hours in, the profile is familiar to anyone who has done it: comfortable comprehension, an ear that copes with speed, and a vocabulary parked around whatever the show says every week.

The repair is small, and it is not more listening. It is a second meeting with the specific words that slipped, spaced out, and shaped as a question rather than a replay — the combination Cepeda and colleagues (2006) measured at around 47% recall against 37% for massed practice, and that Karpicke and Roediger (2008) found beat re-reading outright. The awkward part has always been where to put it, because a review session is one more thing to schedule and scheduling is the exact step people fail at. Which argues for putting it on a surface that already repeats itself without anyone’s cooperation: the phone gets picked up about 186 times a day (Reviews.org, 2026), and not one of those pickups has to be planned.

Podcasts, music, background radio, sleep audio

These get recommended in one breath as though they were the same intervention, and the research treats them very differently. Two of them are supported, one is supported for a reason people usually get backwards, and one works only in a sense narrow enough that it is worth stating precisely.

The pattern underneath is the same each time: the more the format asks something of you, the more of it you keep.

Four listening formats, the evidence supporting each, and what each one will not do
FormatWhat the evidence supportsWhat it will not do
Podcasts and audiobooks at your levelvan Zeeland & Schmitt (2013): real if modest incidental gains from listening to stories, above the coverage band Nation (2006) sets at ~3,000 word families for 95%Hand you the written form, or bring back the word you missed. Meaning is picked up far more readily than form, and neither is scheduled
Music and songLudke, Ferreira & Schellenberg (2014): adults who sang Hungarian phrases recalled them better afterwards than adults who spoke the same phrasesWork with the music merely on. The active ingredient in that study was producing the language; lyrics in the background is the untested condition, not the tested one
Background radio while you workSaffran et al. (1996): phonological structure is extracted from a stream remarkably cheaply, so familiarity with the sound of the language is a genuine gainTeach vocabulary. This is closest to the recorded-audio condition in Kuhl et al. (2003), the one that produced no measurable learning
Audio played during sleepSchreiner & Rasch (2015): replaying words participants had already studied before sleeping improved recall of exactly those wordsTeach a word you have never met. That result is reactivation of an existing memory, so it can finish a job you started awake and cannot start one

The sleep row is the one most often quoted without its condition, so it is worth repeating: the words worked because they had been studied first. If you want the timing question on its own, it has a page: should you study vocabulary before bed?

A place for the words the episode said once

That gap between hearing a word and being able to produce it is what LearnScreen was built for. Keep the podcast. The words you rewound for come back on a surface you were already going to look at.

1

The word you rewound for becomes a card

Add it in a few seconds from the Share sheet, or paste a whole list after an episode, each with a translation and an example sentence. It joins the same queue as the built-in lists — so the word the host said exactly once starts behaving like a word you meet often, and it finally arrives with the written form listening withheld.

2

It comes back as a question, not a replay

Using Apple’s Screen Time API, LearnScreen shields the apps you choose and puts the card where the feed would have been, answer hidden until you commit to an attempt. That turns each encounter into retrieval — the operation Karpicke and Roediger (2008) found beats re-exposure, and the one an audio feed structurally cannot perform.

3

A Leitner schedule decides what returns

Words you miss come back sooner; words you get right back off geometrically. Spacing follows how well you know each word rather than an editor’s running order — the part the distributed-practice research (Cepeda et al., 2006) says does the work, and the part no playlist has ever been able to supply.

  • Nothing extra gets scheduled. The review lands on pickups you were already making, which is why it survives the weeks when an episode on the walk to work is the most you can manage. Defaults come to roughly 25 recall attempts a day — about two and a half minutes, spread across moments you were not using for anything.
  • Start from the frequent core if you are under the threshold. 1,080+ curated words across 18 topics in 11 languages, ordered so the cheapest coverage comes first — the fastest route to the ~3,000 families that make ordinary speech legible (which words come first).
  • Your own words, in bulk. Add one from the Share sheet mid-episode or paste the list you built on the commute; every entry takes a translation and an example sentence, so the card keeps the sense the episode gave it.
  • Offline, no account. Cards and shields run without a network once installed, and iCloud backup writes to your own private database rather than our servers.

Where the words show up

Four screens from the app: the shield card, the answer with its example sentence, the word list and the progress view.

A blocked app showing a vocabulary card on the shield screen instead of the feed The answer revealed on the shield card, showing the translation and an example sentence using the word The vocabulary list showing curated topics alongside custom words added by the learner The progress screen showing how many words have moved up through the spaced-repetition queue

Related reading

Frequently asked questions

Can you learn a language just by listening to it?
You can learn a great deal of its sound system that way, and very little of its vocabulary. Listening trains segmentation, rhythm and phonetic contrasts, which is real and hard to get any other way. But vocabulary pickup is gated by coverage and repetition: below roughly 3,000 word families a recording is sound you cannot cut into words, and a word needs five to twenty encounters before it stays (Nation & Wang, 1999). Nothing in an audio feed is scheduled by what you forgot.
Do podcasts actually help you learn a language?
Yes, above a threshold, and mostly for listening rather than vocabulary. Nation (2006) puts 98% coverage of spoken text at roughly 6,000 to 7,000 word families and 95% at about 3,000. Above that band a podcast is comprehensible input and a good use of a commute. Below it, comprehension collapses and you are practising tolerating confusion. Either way the words the host said once will not come back on a schedule.
Does listening to music in another language help?
The evidence that exists is for singing, not for listening. Ludke, Ferreira and Schellenberg (2014) found that adults who sang Hungarian phrases recalled them better afterwards than adults who spoke the same phrases. The active ingredient was producing the language, not having music playing. Lyrics on in the background is closer to the condition that has repeatedly shown no measurable learning.
Does listening to a language while you sleep work?
Not for learning new words, and yes in one narrow sense. Schreiner and Rasch (2015) replayed foreign words that participants had already studied before sleeping and found improved recall for exactly those words. That is reactivation of a memory that already existed, not acquisition. You have to meet the word awake first, which makes sleep audio a possible finisher and never a substitute.
Why do I understand so much more than I can say?
Because listening supplies recognition and never asks for recall. Karpicke and Roediger (2008) found that retrieval practice beat repeated study on later retention, and hours of input are repeated study by definition. A comfortable ear and an empty mouth is the exact profile you would predict from years of high-quality exposure with no retrieval attached to it. The gap has its own page: active vs passive vocabulary.
How much listening a day is worth doing?
As much as you enjoy, since it is the strand that survives a bad week. Nation (2007) puts meaning-focused input at roughly a quarter of a balanced programme, and audio fills commutes and chores that nothing else can reach. The thing worth adding is small: a few short retrieval attempts for the specific words you missed, spaced out. Distributed practice recalled around 47% against 37% for massed practice (Cepeda et al., 2006).

Sources

  1. Kuhl, P. K., Tsao, F.-M., & Liu, H.-M. (2003). Foreign-language experience in infancy: Effects of short-term exposure and social interaction on phonetic learning. PNAS, 100(15), 9096–9101.
  2. Saffran, J. R., Aslin, R. N., & Newport, E. L. (1996). Statistical learning by 8-month-old infants. Science, 274(5294), 1926–1928.
  3. van Zeeland, H., & Schmitt, N. (2013). Incidental vocabulary acquisition through L2 listening: A dimensions approach. System, 41(3), 609–624.
  4. Nation, I. S. P. (2006). How large a vocabulary is needed for reading and listening? Canadian Modern Language Review, 63(1), 59–82.
  5. Peters, E., & Webb, S. (2018). Incidental vocabulary acquisition through viewing L2 television and factors that affect learning. Studies in Second Language Acquisition, 40(3), 551–577.
  6. Nation, I. S. P., & Wang, K. (1999). Graded readers and vocabulary. Reading in a Foreign Language, 12(2), 355–380.
  7. Schreiner, T., & Rasch, B. (2015). Boosting vocabulary learning by verbal cueing during sleep. Cerebral Cortex, 25(11), 4169–4179.
  8. Ludke, K. M., Ferreira, F., & Schellenberg, E. G. (2014). Singing can facilitate foreign language learning. Memory & Cognition, 42(1), 41–52.
  9. Karpicke, J. D., & Roediger, H. L. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966–968.
  10. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
  11. Murre, J. M. J., & Dros, J. (2015). Replication and analysis of Ebbinghaus’ forgetting curve. PLOS ONE, 10(7), e0120644.
  12. Nation, I. S. P. (2007). The four strands. Innovation in Language Learning and Teaching, 1(1), 2–13.
  13. Reviews.org (2026). Cell Phone Usage Stats. Survey of ~1,000 US adults, fielded Q4 2025. Report

Three limits are worth carrying away with the numbers. The phonetic-learning studies are about infants, and infants are not adults in a smaller format — what transfers is the comparison between conditions, not the age. The incidental-listening studies are short and test recognition rather than production, so they understate what years of listening does and overstate how measurable a single hour is. And the sleep finding applies only to words already studied before sleeping. Where this page states a number it comes from the study named beside it; where it generalises to your own listening, that is an inference, and it is marked as one.

Keep listening. Don’t let the words leave with the episode.

The word you rewound for, back as a question, on a screen you already unlock.

Download on the App Store