Skip to content
Efficacy research on language apps

Do language learning apps
actually work?

Yes, measurably, and mostly at the beginner end. About 34 hours of app study matched a first university semester of Spanish on a placement test — in a trial the app’s own maker commissioned (Vesselinov & Grego, 2012). An independent semester with the same app found real gains in reading and listening and almost none in speaking (Loewen et al., 2019). The variable that separates learners it works for from learners it doesn’t is not the app: it is how many sessions happen.

Every figure on this page is referenced at the end

A vocabulary card on an iPhone lock screen, shown at the moment a blocked app was opened

What the efficacy studies actually found

Language apps are one of the most downloaded categories on earth and one of the least rigorously evaluated. There are a handful of real studies, they are small, and their results are more interesting than either the advertising or the backlash. Read them together and a consistent shape appears: genuine gains, concentrated in recognition rather than production, with enormous variation between individuals.

The funding column matters as much as the findings column. The single most-quoted number in this space comes from a study the vendor paid for, which does not make it wrong but does make it a floor for scepticism rather than a ceiling for claims.

Five studies of language-app effectiveness: what each measured and what it found
StudyWhat it measuredWhat it found
Vesselinov & Grego (2012)Adult beginners in Spanish, tested on a university placement test before and after self-paced app study — commissioned by the app’s makerAn average of roughly 34 hours produced placement gains comparable to a first university semester, with very wide variation in how much each learner actually studied
Loewen et al. (2019)Independent semester-long study of university learners using the same app to study Turkish from scratchMeasurable gains in reading and listening, very little movement in speaking, and steep drop-off in use across the term
Rachels & Rockinson-Szapkiw (2018)A gamified Spanish app compared against face-to-face instruction in elementary classrooms, achievement and self-efficacy both measuredNo significant difference in achievement, and lower self-efficacy in the app group — equal results, less confidence
Golonka et al. (2014)Review of the whole technology landscape for language learning, sorted by technology type and strength of supporting evidencePlenty of enthusiasm, thin rigorous evidence that any given technology beats ordinary instruction; the best-supported uses were narrow and specific
Burston (2015)Two decades of mobile-assisted language learning implementations, analysed for what the reported learning outcomes actually rest onMost projects were very short and very small, so positive results describe weeks of novelty rather than terms of learning

None of these studies tested this app, and none of them measured fluency. They measure placement scores, receptive skills and classroom achievement over weeks or a term. Treat the category verdict as “does something real, at the beginner end, in recognition first” — and treat any number promising conversational ability from screen taps as unsupported.

Three different claims travel under “apps work”

They rest on completely different evidence, which is why the same person can reasonably believe apps are useful and that app marketing is nonsense. Only the first one is well supported.

  • The supported one: measurable beginner gains, mostly receptive. Every study above found something. Placement scores move, reading and listening move, and vocabulary is the part that moves most reliably — it is the component that responds best to spaced retrieval, which is the one thing software does better than a classroom.
  • The overstated one: an app replaces a course. Rachels and Rockinson-Szapkiw (2018) found parity with classroom teaching on achievement, not superiority, and Loewen and colleagues found the speaking gap the classroom exists to close. Nation’s (2007) four strands make the shape explicit: input, output, deliberate study and fluency development are separate jobs, and a tap-and-select app is strong at exactly one of them.
  • The false one: fluent in fifteen minutes a day. No study in this literature measured fluency at all, so no study supports a fluency claim. At ten minutes a day, the 34 hours in the best-known result takes about seven months — and that arithmetic assumes you never skip, which is precisely the assumption the next section takes apart.

Exposure without retrieval is a fourth claim again, with its own evidence base: does passive learning actually work?

What actually decides whether an app works for you

Six things, each taken from a study rather than from a feature list, and ordered by how much they move the result. The first one dwarfs the rest, which is inconvenient for everybody selling the other five.

Six factors that determine whether a language app produces results, the evidence behind each, and what each looks like in practice
What decides itWhat the evidence saysWhat that looks like
Whether the session happensNielson (2011): federal employees given licences, paid study time and tutor support still largely stopped early, and few reached meaningful usageThe failure mode is not a bad lesson, it is a lesson that never opened — if support and paid time were enough, that study would have ended differently
Whether you retrieve or recogniseRoediger & Karpicke (2006): being tested on material produced better long-term retention than restudying it for the same timePicking the right tile out of four is much easier than producing the word, and easier practice buys less; the exercise type matters more than the app
Whether practice is spread outCepeda et al. (2006), synthesising 254 studies: about 47% recall for spaced practice against 37% for the same practice massedTen minutes on six days beats an hour on Sunday, and the app that shows up daily wins on this axis before any content is compared
Whether each word gets enough meetingsNation & Wang (1999): a word generally needs somewhere around 8 to 12 spaced encounters before it staysA streak made of new material every day quietly never finishes anything; the words that stick are the ones scheduled to come back
Whether the words are worth the meetingsNation (2006): roughly the first 3,000 word families cover about 95% of everyday conversation, and coverage flattens sharply after thatFrequency order is most of the return in the first year, which is why a themed course of animal names feels productive and measures poorly
Where the minutes come fromReviews.org (2026): US adults report checking their phones about 186 times a day, roughly 11.6 times per waking hourA plan that needs a new slot in the day competes with everything else in it; one that rides an existing habit does not

Only rows two to five are about software quality, and they are the rows every app in the category has largely solved. Rows one and six are about delivery, and they are where the outcomes are actually decided — which is the finding that keeps making method comparisons come out flat.

The variable the reviews never measure

App comparisons are written about features: how good the speech recognition is, whether the grammar notes are any good, how the review queue is scheduled. Those are real differences and they are small ones. The difference that is not small is between a learner who did forty sessions and a learner who did four, and no review can tell you which of those you will be, because it is not a property of the app.

This is why the efficacy literature keeps producing flat results. Golonka and colleagues (2014) went looking for technologies with strong evidence of advantage and mostly found enthusiasm; Burston (2015) found that two decades of mobile-assisted projects were dominated by studies lasting weeks. When time-on-task varies by an order of magnitude between participants and the trial is short, a method difference has almost no room to show up. Meanwhile Nielson (2011) ran what should have been the best case for self-study software — volunteers, licences, protected time, tutor support — and watched most of them stop. Adherence is not a soft factor sitting alongside the hard ones. It is the hard one.

Which reframes the question people actually mean when they ask whether apps work. The honest version is: does this app produce enough sessions, on days when nothing is going well, for the spacing and retrieval effects to accumulate? Every technique on this page is worthless in an app you stopped opening in March, and even a mediocre technique compounds in one you never had to decide to open. The productive move is not to hunt for a better exercise. It is to attack the step where the losses actually are.

What quitting costs, in the units the research uses

“Be consistent” is advice nobody can act on. What follows is the same point with numbers attached: four specific things that stop happening when the sessions stop, each of them measured by somebody.

Notice that three of the four are not about learning less. They are about losing material you had already paid for.

What is lost when app sessions stop, with the measured size of each loss
What is lostEvidenceMeasured size
The spacing advantageCepeda et al. (2006), synthesising 254 studies of distributed practiceAbout 47% spaced against 37% massed — a gap you only collect by returning on a later day
The encounters a word still needsNation & Wang (1999), on how often a word must be met before it holdsRoughly 8 to 12 spaced meetings, so a word abandoned at four is not four-twelfths learned — it is gone
Everything already half-learnedMurre & Dros (2015), replicating Ebbinghaus’ forgetting curveSteep early loss without review — the queue you stop clearing decays fastest in the first days, before it slows
The finished courseNielson (2011), workplace self-study with licences, time and supportMost participants stopped early — the modal outcome of self-directed software is an unfinished course, not a bad one

This is not an argument for discipline. It is an argument for arranging things so that fewer decisions stand between you and a review — which is a design problem, and a solvable one. The habit side of it is on how to stop quitting a language.

Removing the step where the losses happen

If adherence is the variable that decides the outcome, the useful design question is not how to make a lesson better. It is how to make a review happen without anyone deciding to start one. LearnScreen is built around that single idea — here is the mechanism, plainly.

1

No session to start

Apple’s Screen Time API lets the app shield the apps you choose. When you open one, a word card appears where the feed would have been. Phones are checked around 186 times a day (Reviews.org, 2026), so the review rides a habit that already exists instead of asking for a new slot in the day.

2

Each card is a retrieval attempt

The answer stays hidden until you tap, so the card is a test rather than a reading — the direction Roediger and Karpicke (2006) found produced better long-term retention. Ambient delivery raises the number of attempts; it does not change what one attempt is worth.

3

A schedule decides what returns

A Leitner queue brings missed words back sooner and lets known words back off geometrically, so the 8 to 12 meetings a word needs (Nation & Wang, 1999) get spread across days rather than crammed into one sitting.

  • You set the dose. Words per session adjust from 3 to 20, as does how often the shield returns. The defaults come to roughly 25 recall attempts a day — about two and a half minutes spread across it, which is small enough that no other plan has to be cancelled to fit it.
  • It is not a course, and does not pretend to be one. This is deliberate vocabulary study, one of Nation’s (2007) four strands. Speaking practice, reading and listening are jobs it does not do, and the honest place for it is underneath them: active versus passive vocabulary.
  • Frequency-ordered lists, or your own words. Curated sets start where coverage is cheapest (Nation, 2006), and anything you add by hand or paste in bulk joins the same queue — which words to learn first.
  • It works offline, with no account. Cards and shields run without a network once installed; iCloud backup writes to your own private database rather than our servers, and there is nothing to sign up for.

Where the words show up

Four screens from the app: the shield card, the answer with its example sentence, the word list and the progress view.

A blocked app showing a vocabulary card on the shield screen instead of the feed The answer revealed on the shield card, showing the translation and an example sentence using the word The vocabulary list showing curated topics alongside custom words added by the learner The progress screen showing how many words have moved up through the spaced-repetition queue

Related reading

Frequently asked questions

Do language learning apps actually work?
Yes, for measurable beginner-level gains, mostly in reading and listening. The most-quoted figure is that about 34 hours of app study produced placement-test improvement comparable to a first university semester of Spanish, from a study commissioned by the app’s maker (Vesselinov and Grego, 2012). An independent semester-long study of learners studying Turkish with the same app found real receptive gains and very little movement in speaking (Loewen et al., 2019), and a controlled comparison with classroom instruction found no significant difference in achievement, alongside lower self-efficacy in the app group (Rachels and Rockinson-Szapkiw, 2018). Nothing in that literature supports fluency claims, and none of it applies to sessions that do not happen.
Is one app enough to learn a language on its own?
Not on its own, and the limitation is structural rather than a flaw in one product. Nation’s (2007) four strands framework holds that a course needs meaning-focused input, meaning-focused output, deliberate language study and fluency development in roughly equal measure. A tap-and-select app is very good at deliberate study and reasonably good at input; it is weak at output, which is why Loewen and colleagues (2019) measured gains in reading and listening and almost none in speaking. The realistic role is as the vocabulary and grammar engine underneath speaking practice you get somewhere else.
Why do studies comparing language apps keep coming out flat?
Because the differences between methods are small compared with the differences in how much people actually do. Reviews of language-learning technology have repeatedly found the evidence for any particular technology thin and the study designs short (Golonka et al., 2014), and Burston’s (2015) analysis of two decades of mobile-assisted language learning projects found most ran for weeks rather than terms. When study duration is short and time-on-task varies wildly between participants, method effects are swamped. The practical reading is that the app you will open is better than the app that tests marginally better.
How many hours on an app equal a semester of class?
About 34 hours, on one placement test, in one vendor-commissioned study of Spanish (Vesselinov and Grego, 2012) — and that number should be handled carefully. It measures placement-test gain rather than the ability to hold a conversation, it comes from learners who finished, and it was funded by the company whose product it evaluated. Treat it as evidence that the category does something rather than as a conversion rate, and note that at ten minutes a day, 34 hours takes about seven months.
What separates people an app works for from people it doesn’t?
Whether the sessions happen. Nielson (2011) gave federal employees language-learning software with paid time and tutor support, and even in that unusually favourable setting most participants stopped early and few reached meaningful usage. That result is about the delivery model, not the software: self-directed study asks you to decide to start, every day, and that decision is where the drop-off is. Everything else on this page is downstream of it.
Does passive or ambient review count as using an app?
It counts if a recall attempt happens. Exposure alone teaches slowly, and the testing effect is the part that pays: Roediger and Karpicke (2006) found being tested on material produced better long-term retention than restudying it, and Cepeda and colleagues (2006), across 254 studies, put spaced practice at roughly 47% recall against 37% for massed practice. What ambient delivery changes is the number of attempts that occur, not the value of each one. A card you glance past is worth much less than a card you tried to answer.

Sources

  1. Vesselinov, R., & Grego, J. (2012). Duolingo Effectiveness Study. City University of New York / University of South Carolina. Commissioned by Duolingo.
  2. Loewen, S., Crowther, D., Isbell, D. R., Kim, K. M., Maloney, J., Miller, Z. F., & Rawal, H. (2019). Mobile-assisted language learning: A Duolingo case study. ReCALL, 31(3), 293–311.
  3. Rachels, J. R., & Rockinson-Szapkiw, A. J. (2018). The effects of a mobile gamification app on elementary students’ Spanish achievement and self-efficacy. Computer Assisted Language Learning, 31(1–2), 72–89.
  4. Nielson, K. B. (2011). Self-study with language learning software in the workplace: What happens? Language Learning & Technology, 15(3), 110–129.
  5. Golonka, E. M., Bowles, A. R., Frank, V. M., Richardson, D. L., & Freynik, S. (2014). Technologies for foreign language learning: A review of technology types and their effectiveness. Computer Assisted Language Learning, 27(1), 70–105.
  6. Burston, J. (2015). Twenty years of MALL project implementation: A meta-analysis of learning outcomes. ReCALL, 27(1), 4–20.
  7. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
  8. Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
  9. Nation, I. S. P., & Wang, K. (1999). Graded readers and vocabulary. Reading in a Foreign Language, 12(2), 355–380.
  10. Nation, I. S. P. (2006). How large a vocabulary is needed for reading and listening? Canadian Modern Language Review, 63(1), 59–82.
  11. Nation, I. S. P. (2007). The four strands. Innovation in Language Learning and Teaching, 1(1), 2–13.
  12. Murre, J. M. J., & Dros, J. (2015). Replication and analysis of Ebbinghaus’ forgetting curve. PLOS ONE, 10(7), e0120644.
  13. Reviews.org (2026). Cell Phone Usage Stats. Survey of ~1,000 US adults, fielded Q4 2025. Report

These are small studies. Sample sizes run from single classrooms to low hundreds, durations from weeks to one term, and outcome measures are placement tests and skill batteries rather than real-world conversation. One of them was paid for by the company it evaluated, which is stated where it is quoted. LearnScreen has not been through an efficacy trial and this page does not claim otherwise; the app-specific figures here are arithmetic from its default settings and a roughly twenty-second recall attempt, not measured user data.

Fix the step that decides the outcome

If the research says adherence beats method, the fix is not a better lesson — it is a review that happens without a decision. LearnScreen puts one on the phone you were already unlocking, and asks for no new minutes at all.

Download on the App Store