Should you pay students to learn? The evidence on rewards and extrinsic motivation

October 4, 2026 · 12 min read

A student at GT School, one of Alpha School's sister schools, earns about 10 GT Bucks for finishing the day's required lessons, and a GT Buck is worth 10 cents. A parent who reviewed the school called this dollar a day its "secret sauce". The best-known American trials of paying students found less. Roland Fryer distributed $9.4 million in experiments covering about 27,000 students and estimated the effect on test scores as "statistically 0, in each city".

I build Understand, a tutor for adults, and I have no connection to Alpha School or to the researchers below. I have run no incentive experiment myself. This post rests on the papers, read in full where an open copy exists, and I say where I read only an abstract.

A dollar a day at Alpha School

Alpha's FAQ says students earn "Alphas" toward prizes "or larger rewards such as a laser tag party". The detail comes from the anonymous parent whose review Astral Codex Ten published in June 2025, and which I drew on in my Alpha School review. At GT School a lesson counted toward GT Bucks only if the child scored 80% or higher on its problem sets. One summer Alpha paid $1 per lesson in real dollars.

The reviewer thinks the rewards matter as much as the software: "If the 2-hour learning tool is the self-driving car, the incentives are the fuel". The same software in a home pilot, without guides or rewards, produced about 1x growth by the reviewer's account. The review cites Fryer as the research behind the design.

Roland Fryer's $9.4 million

Fryer's paper, Financial Incentives and Student Achievement, appeared in the Quarterly Journal of Economics in 2011. It reports randomized trials in 203 schools in Dallas, New York and Chicago during 2007-08 and 2008-09. Schools were randomly assigned to pay or not, and each city paid for something different.

Dallas paid second graders $2 for each book they read, up to 20 books a semester, once they passed a computer quiz on it. The average student earned $13.81 over the year.

New York paid fourth and seventh graders for their scores on ten interim tests, up to $250 a year in fourth grade and $500 in seventh.

Chicago is the trial of paying kids for grades: $50 for an A, $35 for a B and $20 for a C in five core courses, every five weeks. One of the five was gym, which 22% of ninth graders had failed the year before. The average student earned $695.61.

The effect of paying for books on reading scores was 0.012 standard deviations, with a standard error of 0.069. New York's estimates were as close to zero. In Chicago, grades rose by 0.093 standard deviations, which Fryer calls marginally significant, and the achievement test did not move.

The one significant result appears when Dallas is split by language. Students who took the English-language test gained 0.173 standard deviations in reading. Students in bilingual classes, who took a Spanish test, lost 0.118. Most of the books were in English, and Fryer suspects they crowded out the Spanish those children were being taught in.

The paper's cautions get quoted less often than its zero. The trials were built to detect effects of 0.15 standard deviations or more, so Fryer "cannot rule out" smaller gains that would still repay the cost. And on a survey measure of intrinsic motivation the treatment effect was -0.017, which he reads as "little evidence that incentives decrease intrinsic motivation".

Fryer also notes that Pizza Hut's Book It! program, which pays readers in personal pan pizzas, had been running for 25 years "but never credibly evaluated".

The abstract of Fryer's April 2010 working paper, from the NBER's public PDF, captured in October 2026.The abstract of Fryer's April 2010 working paper, from the NBER's public PDF, captured in October 2026.

Paying for inputs and paying for outputs

The usual summary of Fryer is that paying for inputs such as books read works and paying for outputs such as test scores fails. That sentence comes from the 2010 working paper, whose abstract says so outright.

The working paper had a fourth city. In Washington DC, middle schoolers earned $2 a point for attendance, behaviour, wearing a uniform and turning in homework, and averaged $532.85 for the year. The estimate for reading was -0.041 standard deviations without control variables and +0.152 with them, in 34 schools. Fryer wrote that the controls altered "the results considerably, which is troubling".

The published paper drops Washington and the claim in the abstract. What remains is offered as speculation: "incentives are not a panacea and more effective if tailored to appropriate inputs". A child offered money for a higher score has to know what produces one. In Dallas, he writes, students "only needed to know how to read books".

Outputs have one result of their own. Eric Bettinger studied Coshocton, Ohio, where a local manufacturer funded payments of $15 or $20 for each state test a child in grades three to six passed, up to $100 a year. Grades within schools were drawn by lottery, with a bingo cage. Math scores rose by about 0.15 standard deviations, and reading, science and social studies did not move.

Steven Levitt, John List, Susanne Neckermann and Sally Sadoff tested timing in "The Behavioralist Goes to School". They told more than 6,000 students near Chicago, minutes before a low-stakes test, that an improved score would earn $10, $20 or a trophy. Scores rose by about a tenth of a standard deviation when the reward was handed over at once. When it was promised for a month later, the effect vanished. That experiment measures effort on a test the students could already do. It says little about learning.

Uri Gneezy, Stephan Meier and Pedro Rey-Biel reviewed the field in 2011 and concluded that incentives in education "seem to have moderate success when the incentives are well-specified and well-targeted ('read these books' rather than 'read books')".

Later meta-analyses, of which I read only the abstracts, point the same way. Vi-Nhuan Le's of 21 studies in 2020 found a positive effect on math and none on reading. Tomáš Lintner's of 18 randomized trials with 20,286 university students in 2024 found more credits earned and grade averages "marginally improved".

StudySampleRewardedResultMain limit
Lepper, Greene and Nisbett 197351 preschoolers who already liked drawingA promised "Good Player Award" for drawingA week or two later they drew for 8.6% of free time, against 16.7% with no awardOne nursery school, groups of 15 to 18
Fryer 2011, Dallas3,718 second graders, 42 schools$2 per book readReading 0.012 SD overall; +0.173 on the English test, -0.118 on the SpanishThe gain is a subgroup; quiz software came bundled
Fryer 2011, New York15,883 fourth and seventh gradersScores on ten interim testsMath 0.004 SD, reading -0.031 in seventh gradeBuilt to detect only 0.15 SD or more
Fryer 2011, Chicago7,655 ninth gradersCourse grades, every five weeksGrades +0.093 SD; test scores flatTest taken four months after payments ended
Fryer 2010, Washington DC34 middle schoolsAttendance, behaviour, uniform, homeworkReading +0.152 SD with controls, negative withoutDropped from the published paper
Bettinger 2012, CoshoctonGrades 3 to 6 in four schools, three years$15 to $20 per state test passedMath +0.15 SD; other subjects flat16 school-grade units in each lottery
Levitt and others 2016Over 6,000 students near ChicagoA better score on a test announced minutes earlierAbout +0.1 SD if paid at once; nothing if paid a month laterMeasures effort on one test

Intrinsic vs extrinsic motivation and the Good Player Award

Extrinsic motivation means doing a thing for a reward outside it, and intrinsic motivation means doing it for itself. The fear about paying students is that the first destroys the second.

The founding experiment is by Mark Lepper, David Greene and Richard Nisbett, published in 1973. They picked children at the Bing Nursery School in Stanford, California, who chose to draw with felt-tip pens during free play. Some were promised a "Good Player Award" for drawing, some got it as a surprise, and some got nothing. One to two weeks later, the children who had been promised the award spent 8.59% of their free time drawing. The others spent 16.73% and 18.09%. The authors call this overjustification: the child concludes they drew for the prize, and stops when the prizes do. The final sample was 51 children, and the paper records that observation sessions were skipped for "reasons ranging from the unanticipated arrival of a goat in the classroom to equipment failure".

Edward Deci, Richard Koestner and Richard Ryan pooled 128 such experiments in 1999. Tangible rewards lowered later free-choice interest by 0.28 to 0.40 standard deviations depending on what the reward was contingent on. Positive feedback raised it by 0.33.

Judy Cameron and David Pierce had pooled 96 experiments in 1994 and concluded that "overall, reward does not decrease intrinsic motivation". Their 2001 reply, written with Katherine Banko, is subtitled "The myth continues". I read the abstracts of all three.

The two camps agree on more than the titles suggest. Both find harm when a tangible reward is promised in advance for a task the person already likes. That is Lepper's condition exactly. Cameron's group adds that rewards for low-interest tasks raise later free-choice motivation. Neither condition describes Fryer's second graders well, and his own survey found no change.

Adults, stakes and deadlines

An adult has no principal with a budget. The adult versions are a reward you grant yourself, a deadline you set yourself, and money you put at risk.

The best-known evidence for self-imposed deadlines was Dan Ariely and Klaus Wertenbroch's 2002 paper in Psychological Science. The journal retracted it on September 2, 2026. Kate Travis reported in Retraction Watch the next day that Wertenbroch had requested the retraction, and that the notice cites a replication which "raised questions about the underlying data in the original study" and analyses by the Data Colada blog. In that replication, Kyle Hyndman and Alberto Bisin found that changing the deadlines "had a negligible effect" on performance. Travis quotes Ariely's response: "whatever you believed in it, believe it less". The paper was on my list as the evidence for deadlines, so I am taking his advice.

Commitment contracts have a real trial behind them. Xavier Giné, Dean Karlan and Jonathan Zinman offered smokers a savings account that returned their deposit after six months only if they passed a urine test. Of those offered it, 11% signed up. Smokers offered the account were 3 percentage points more likely to pass, and the difference held in surprise tests at 12 months. Most people decline to bet against themselves, which keeps the overall effect small.

Anyone can buy the same arrangement. stickK calls itself "a free goal-setting platform created by behavioral economists at Yale University", and its home page counted 649,000 commitments in August 2026. Beeminder draws a line on a graph of your numbers: "If you cross the line, we charge your payment method!" I found no randomized trial of either service for learning.

Beeminder's home page in October 2026. Duolingo points are among its examples of things to put money on.Beeminder's home page in October 2026. Duolingo points are among its examples of things to put money on.

Streaks and XP

Michael Sailer and Lisa Homner's meta-analysis of gamification pooled experiments that added game elements to learning. It found small positive effects on what learners knew, on motivation and on behaviour. The authors add that behaviour was "almost exclusively measured during interventions", which means while the game was still on.

Figure 2a from Sailer and Homner, "The Gamification of Learning: a Meta-analysis", Educational Psychology Review, published online in 2019 and licensed CC BY 4.0: effect sizes on what learners knew, by study. The three largest were excluded from the final estimate of 0.49. Retrieved October 2026.Figure 2a from Sailer and Homner, "The Gamification of Learning: a Meta-analysis", Educational Psychology Review, published online in 2019 and licensed CC BY 4.0: effect sizes on what learners knew, by study. The three largest were excluded from the final estimate of 0.49. Retrieved October 2026.

Jackie Silverman and Alixandra Barasch ran seven studies on streaks, of which I read the abstract. People shown an intact streak in a log were more likely to repeat the behaviour than people shown a broken one, even when their past behaviour was the same. The authors' reason is that people "consider maintaining a logged streak to be a meaningful goal in and of itself".

A streak is an input incentive: it pays you, in a number, for showing up. It should get you to open the app. I found no trial showing that an adult with a longer streak understands more.

A question you want answered

My reading of the trials is that paying for inputs has more behind it than paying for scores, and neither has much. Inputs have Dallas's English readers and a Washington estimate its own author withdrew. Outputs have math in Coshocton.

A parent asking whether you should pay kids for good grades is in good company: 65% of the children Bettinger surveyed said their parents already did. Chicago suggests the grades will rise a little and the test scores will stay put.

Fryer's explanation matters more to me than his estimates. A reward helps only when the learner knows what to do next. An adult studying alone has a harder version of this. Nobody will pay you, and as I wrote in Alpha School for adults, a reward you can grant yourself at any time is a weak reward. Stakes help a little, for the minority willing to sign.

Karl Popper wanted a school "in which no unwanted answers to unasked questions would have to be listened to; in which one did not study for the sake of passing examinations". David Deutsch follows him in holding that knowledge grows by conjecture and criticism, and a conjecture is always an attempt on a problem somebody has. A school needs Alphas because the questions in a curriculum belong to someone else. An adult can skip that step and start from a problem of their own. Why interest rates move bond prices is a motive that survives a missed day, and it tells you what to do next: guess, and check the guess.

The counter-case is practice. Fluent calculation or vocabulary takes drills that few people enjoy, and curiosity runs out before they are done. Cameron's finding that rewards help on low-interest tasks fits here, and so does a stake.

Things to try

  • Put any reward or stake on something you can count tonight, such as 30 minutes or one chapter explained from memory. A goal like "understand statistics" gives you Fryer's New York problem.
  • If you use money, let someone else hold it. A friend who gets $20 for every week you skip costs nothing to set up.
  • Keep rewards for the dull parts. If you already like reading history, Lepper's result is a reason to leave it unpaid.
  • Make the reward immediate. A reward a month away did nothing for Levitt's students.
  • Treat a streak as an attendance record. Once a week, explain something from that week aloud without notes and see what is missing.
  • Write down the question you want answered before you start a course, and drop the units that do not bear on it.

Understand

Understand is a voice tutor for adults that you can interrupt. It reads a text aloud word for word and stops when you ask a question. It asks you to explain an idea back, and it keeps a record of what you explained that you can read and correct. It has no streaks, points or prizes, so if a counter is what gets you to the desk, Duolingo or Beeminder will serve you better.

It also has no courses, no problem sets, no spaced repetition and no Android app, and no trial has evaluated it. It is free for now.

Understand's web app, as shown on its landing page in October 2026: a learner's follow-up question about a diagram the tutor drew.Understand's web app, as shown on its landing page in October 2026: a learner's follow-up question about a diagram the tutor drew.