High-dosage tutoring: 0.37 in the trials, 0.06 at scale

October 4, 2026 · 13 min read

High-dosage tutoring is tutoring for at least 30 minutes a session, one on one or in a small group, three or more times a week, with a trained tutor. That is the definition in the National Center for Education Statistics' School Pulse survey, and in 96 randomized trials of tutoring, most of them at three or more days a week, test scores rose by 0.37 standard deviations.

When eight US school systems gave it to thousands of students in 2023-24, the gain was 0.06 to 0.09. Most of that gap is hours. The students who showed up got 13 to 17 hours of tutoring in the year, and the students in the Chicago trials that set the benchmark got 34 to 82.

I build Understand, a voice tutor for adults, and I have no connection to the researchers or programs below. I have never run or attended a high-dosage tutoring program. This post rests on the papers themselves, and I say where I read only an abstract.

The definition of high-dosage tutoring

I take the School Pulse wording from Kraft, Schueler and Falken, who quote it. It also asks for "educators or well-trained tutors" and "an evidence-based core curriculum or program".

The National Student Support Accelerator at Stanford (NSSA) calls the same thing high-impact tutoring. Its list: sessions three or more times a week for at least ten weeks, groups no larger than four with one-to-one as the optimum, "consistent, well-trained tutors", instruction driven by data, structured materials, and a slot inside the school day.

The National Student Support Accelerator's page defining high-impact tutoring, October 2026.The National Student Support Accelerator's page defining high-impact tutoring, October 2026.

The school-day clause matters because a session on the timetable happens, and a session after school depends on the student turning up.

Tutoring effect size in the best trials

Nickow, Oreopoulos and Quan (2020) pooled 96 randomized trials of tutoring from preschool to high school. The overall effect was 0.37 standard deviations. Their tables break it down:

  • Teachers as tutors 0.50, paraprofessionals 0.40, volunteers 0.21, parents 0.23.
  • Preschool and kindergarten 0.45, first grade 0.42, grades 2 to 5 0.29, grades 6 to 11 0.16.
  • Reading 0.35 and math 0.38. Reading tutoring works best in the earliest grades and math tutoring later.
  • During school 0.40, after school 0.21.
  • One or two days a week 0.24, three days 0.34, four or five days 0.41.

The authors add that there is "little evidence of once-weekly tutoring sessions generating large effect sizes".

The peer-reviewed version, in the American Educational Research Journal in 2024, reports a pooled effect of 0.288 in its abstract, which is all I read of it. Its title changed from "The Impressive Effects of Tutoring" to "The Promise of Tutoring", which is what peer review does to adjectives.

Bloom's 2 sigma problem

Bloom's number is larger. Benjamin Bloom reported in 1984 that tutored students scored two standard deviations above a conventional class. I went through the sources in my history of intelligent tutoring systems. The result comes from two dissertations in which students were taught for 11 periods over three weeks. Kurt VanLehn's 2011 review put human tutoring against no tutoring at 0.79. He attributes most of Bloom's number to a rule that tutored students score 90 percent before moving on, and concludes that "human tutoring is not usually 2 sigmas more effective than classroom instruction."

The estimate has shrunk with each larger sample: 2.0 in Bloom, 0.79 in VanLehn, 0.37 in Nickow's working paper, 0.29 in the published version.

Saga Education in Chicago

The best evidence on a single program is Guryan and colleagues, published in the American Economic Review in 2023. Saga Education put a tutor with two students for one class period every school day, on top of the regular math class. The tutors were mostly recent college graduates on a one-year stipend with about 100 hours of training.

Two randomized trials in Chicago public high schools, with 2,633 and 2,710 ninth and tenth graders, found math gains of 0.16 and 0.37 standard deviations for students who took part. The journal abstract gives 0.18 to 0.40. The authors say 0.16 "is about what the average high school student learns in a year". The cost was $3,500 to $4,300 per student per year.

Only 40.2 percent of the students assigned to tutoring in the first trial received any, and 36.9 percent in the second. Many had changed schools over the summer or had a timetable clash.

A follow-up trial with about 4,000 students cut the cost by roughly 30 percent. Four students shared a tutor, and each day two of them worked on software while the other two were tutored. The abstract, which is what I read, reports a gain of 0.23.

Reading for older students is harder. Fryer and Howard-Noveck (2020) randomized New York City middle schools to offer at least 130 hours of four-on-one reading tutoring. I read only the abstract. It reports a positive but statistically insignificant effect on reading scores, and a gain of 0.09 per year for Black students.

The studies side by side

StudyWhoDesignEffectCost or limit
Bloom 1984Grades 4, 5 and 8, two dissertationsRandom assignment, 11 periods2.0 SDTutored group had a higher mastery bar
Nickow, Oreopoulos, Quan 202096 trials, preschool to grade 11Meta-analysis0.37 SD (0.29 as published)Effects fall with grade level
Kraft, Schueler, Falken 2024265 trialsMeta-analysis by program size0.55 SD under 100 students, 0.14 at 1,000 or moreNine trials at 1,000 or more
Guryan et al. 2023 (Saga)2,633 and 2,710 Chicago ninth and tenth gradersTwo randomized trials0.16 and 0.37 SD in math$3,500 to $4,300 a year; about 40% took part
Bhatt et al. 2024 (Saga with software)About 4,000 students, two districtsRandomized trial0.23 SD in mathAbout 30% cheaper
Fryer, Howard-Noveck 2020New York City middle schoolsSchool-level randomized trial, readingNot significant overallI read the abstract only
Kraft, Edwards, Cannata 2024 (Nashville)More than 6,800 studentsExperimental and quasi-experimentalReading 0.04 to 0.09 SD, math noneAbout $1,500 a year
Personalized Learning Initiative 202517,330 students, eight sitesRandomized trial0.06 to 0.09 SD$1,200 to $2,000; preliminary
Wang et al. 2024 (Tutor CoPilot)900 tutors, 1,800 studentsRandomized trial, AI assisting tutors4 points more lesson masteryNo gain on year-end tests
Oreopoulos, Low 2026 (Khanmigo)18 Tennessee middle schoolsRandomized trial0.06 to 0.08 SD a yearSame as Khan Academy without AI
Liu et al. 20262,379 Maryland undergraduatesRandomized by instructorFinal grades down 0.27 to 0.37 SD15% used the tutor
Kestin et al. 2025194 Harvard physics studentsCrossover, AI tutor against active class0.73 to 1.3 SDTwo lessons, tested right away

Effects at scale

Matthew Kraft, Beth Schueler and Grace Falken collected 265 randomized trials (the cover page of the October 2024 version says 282) and sorted them by the number of students tutored.

Programs tutoring fewer than 100 students averaged 0.55. Those with 100 to 399 averaged 0.32, those with 400 to 999 averaged 0.25, and those with 1,000 or more averaged 0.14. Their preferred estimates for large US programs judged on standardized tests are 0.16 to 0.21, "a third to a half as large" as the average across their whole sample.

Programs with fewer than 100 students make up 59 percent of the trials, and nine trials cover programs of 1,000 or more.

The first page of Kraft, Schueler and Falken's working paper on tutoring at scale, Annenberg Institute at Brown University, October 2024.The first page of Kraft, Schueler and Falken's working paper on tutoring at scale, Annenberg Institute at Brown University, October 2024.

The post-pandemic programs came in below even that. The University of Chicago Education Lab's Personalized Learning Initiative randomized 17,330 students across eight school systems in 2023-24. Its June 2025 interim report puts the effect of taking part at 0.06 to 0.09 standard deviations, "approximately 1-2 months of additional learning".

In Nashville, Kraft, Edwards and Cannata followed a district program that delivered over 125,000 hours of tutoring to more than 6,800 students. Reading scores rose 0.04 to 0.09. Math scores did not move.

Dosage delivered

In the Education Lab's sites, 86 percent of students assigned to the standard model got at least one session, far above Saga's 40 percent. Those students averaged 29 sessions and just over 17 hours in the year. The Education Lab estimates that students in the original Saga trials received 34 to 82 hours. Schools set dosage targets below Saga's and then missed them, because they "felt they simply had too many competing demands on limited instructional time". Nashville designed for 90 minutes a week and reached 16 to 24 sessions a semester.

The Education Lab's main finding is that "the student learning per minute of tutoring is consistent across sites and studies". Its $1,200 models did as well as its $2,000 models, and virtual tutors as well as tutors in the room. The scaled programs had the same slope and fewer minutes.

Attendance and scheduling

Voluntary tutoring outside school hours mostly does not happen. New Mexico offered free evening and weekend tutoring to an estimated 34,262 eligible students. According to the Education Lab's March 2024 report, 527 signed up, 1.5 percent, and 326 of those attended. The state then moved tutoring into the school day.

Sessions that do happen lose minutes. Nashville tutors said technology problems affected about half of their virtual sessions, with students "often starting a 30-minute session 10 minutes late because they needed to find working headphones or a laptop charger".

Tutor quality and the comparison group

Tutor quality is the reason I expected, and the evidence for it is thin. Kraft and colleagues found that larger programs were slightly more likely to use teachers and paraprofessionals. They suspect that oversight gets worse with size, and note that most trials do not measure it.

The comparison group changed too. In Nashville, about 55 percent of tutored students were tutored in a class period that the others spent on adaptive software or in a small group with a teacher. A tutor shows a smaller effect against that than against a class of 30.

Some of the small-study effect may be publication bias. Trials in peer-reviewed journals average 0.45 in Kraft's sample and unpublished reports 0.22, a pattern you would expect if small null results stay in a drawer. The authors "interpret these results with caution".

AI tutoring research

The best result for AI in this literature has a human in the session. In the Tutor CoPilot trial, Rose Wang, Ana Ribeiro, Carly Robinson, Susanna Loeb and Dora Demszky randomized 900 online tutors, working with 1,800 students in grades 3 to 8, to get suggestions from a language model while they tutored. Students of tutors with the tool passed the end-of-lesson check 66 percent of the time against 62 percent. For the lowest-rated tutors the pass rate went from 56 to 65 percent. The authors put the cost at $20 per tutor per year.

Tutors opened the tool in 29 percent of sessions. The trial ran two months and found no significant gain on end-of-year math tests.

Figure 2 from Wang, Ribeiro, Robinson, Loeb and Demszky (2024), arXiv:2410.03017, showing exit-ticket pass rates by tutor rating and experience. Licence: CC BY 4.0.Figure 2 from Wang, Ribeiro, Robinson, Loeb and Demszky (2024), arXiv:2410.03017, showing exit-ticket pass rates by tutor rating and experience. Licence: CC BY 4.0.

The Google and Eedi trial of LearnLM, in my history post, also kept a human in charge: 165 students, with tutors approving every message the model drafted.

Where the AI tutor stands alone, the results look like after-school tutoring. In the two-year Khanmigo trial I covered in my Khanmigo review, scores rose 0.06 to 0.08 standard deviations a year, which the authors say resembles Khan Academy practice without AI. The median student messaged the tutor on a third of practice days.

This month Jing Liu and colleagues reported a trial on adults. Thirty instructors at the University of Maryland, teaching 2,379 undergraduates, were randomized to offer a course-integrated AI tutor or not in fall 2025. About 15 percent of the students with access used it even once. Final grades in those sections were 0.27 to 0.37 standard deviations lower, and participation in the course's online activities fell by as much as 0.90. The authors caution that most instructors left the tutor in its default mode, which gives answers directly.

Kestin's Harvard study, also in the history post, is the counterexample: an effect of 0.73 to 1.3 standard deviations from an AI tutor. There the AI tutor was the assigned lesson for the week.

NSSA's August 2026 brief states the mechanism. AI removes the staffing limit, "but it reintroduces an opt-in requirement that a live, school-day tutoring model eliminates". An AI tutor is available at 3 a.m., which is when nobody wants tutoring.

My reading of the evidence

The tutoring effect is real. It shows up across 265 randomized trials, and Kraft's team calls even 0.16 "very impressive for large-scale education interventions".

It is well under 2 sigma. For a program a school district can run, 0.1 is a good result.

I think most of it is time on task with someone who knows where you are. The Education Lab's constant gain per minute points that way, and so do the things that made no difference there, the price of the model and whether the tutor was on a screen.

Karl Popper's account of learning fits this. A learner guesses what an idea means and needs someone to criticise the guess. A tutor with two students hears every guess. A teacher with 30 students hears few of them, and a tutor nobody opens hears none.

An AI tutor fails the way human tutoring fails at scale, by not being used. In Tennessee, 14.5 percent of the messages students sent Khanmigo contained a mathematical question or a step of reasoning. In Maryland, 15 percent of students opened the tutor at all.

None of these trials studied adults being tutored by a person. Adults choose to be there, which removes the worst attendance problem and leaves the ordinary one of a busy week.

Tutoring an adult can buy

On Wyzant, tutors set their own rates and you pay after each lesson. The page of 25 Algebra 1 tutors I loaded on October 4 ran from $40 to $400 an hour, with $100 in the middle, so two lessons a week comes to about $800 a month. The first hour with a new tutor is free if you are not satisfied. Wyzant is a division of IXL Learning, whose main product I reviewed.

Languages are cheaper. Preply sells 50-minute lessons and says its Spanish lessons start at $3 with an average of $17. A table on the same page gives average rates of $25 to $27 an hour. You book a trial lesson first. Speaking is the skill that Duolingo's own studies barely test, as I found when I asked whether Duolingo works, and a human tutor supplies the correction an app cannot.

The AI options cost less and are covered in the best AI tutors in 2026: ChatGPT's Study Mode is free, Khanmigo is $4 a month for US adults working through Khan Academy, and Math Academy is $49 a month for a structured math course without a tutor.

My reading of the trials is that the schedule matters more than the choice. Book fixed slots with one tutor, three a week if you can and two if you cannot. A tutor you talk to twice a week will do more for you than a better one you open once, and the same goes for software.

Understand

Understand is a voice call with an AI tutor on any subject. You can interrupt it mid-sentence. It reads a saved article or PDF aloud word for word and stops when you ask a question. It asks you to explain an idea back, and it keeps a record of what you explained, which you can read and correct.

A call in Understand: the learner's spoken request, the tutor's diagram and answer, and the learner restating the idea to check it.A call in Understand: the learner's spoken request, the tutor's diagram and answer, and the learner restating the idea to check it.

By the definition at the top of this post it is not high-dosage tutoring. There is no human tutor and no timetable, and nobody notices if you skip a week. It has no courses, no problem sets, no spaced repetition and no Android app. It is for adults, and it is free for now.

No trial has evaluated it. On the evidence above, my prior for any opt-in AI tutor, mine included, is a small effect for the people who keep using it and none for the people who do not. If you can afford a human tutor three times a week, the trials on school students are the best evidence there is for that purchase, and there is none yet for mine.