Project Follow Through: Direct Instruction won the largest US teaching experiment, and little changed

October 4, 2026 · 12 min read

Project Follow Through was the largest education experiment the US government has run: in its main phase, from 1968 to 1977, it paid 22 sponsors to install their teaching models in kindergarten to third grade classrooms in more than 170 poor communities, and tested the children against comparison classes. In the official evaluation, which Abt Associates delivered in April 1977, the scripted Direct Instruction model from the University of Oregon came out ahead on basic skills, on the problem-solving measures and on self-esteem.

In 1978 the US Commissioner of Education, Ernest Boyer, confirmed this in a letter to a senator. He wrote that "only one of the 22 models which were assessed in the evaluation consistently produced positive outcomes", and then explained why his office would promote 21 individual school sites instead of that model.

I build Understand, a voice tutor for adults, and I have no connection to Direct Instruction, its publisher or any Follow Through sponsor. I could not open the Abt volumes: ERIC lists them without full text. This post rests on the excerpts and reanalyses I could open, which I name as I go.

The ERIC record for Volume IV-A of the Abt Associates evaluation, dated April 15, 1977, as shown in October 2026.The ERIC record for Volume IV-A of the Abt Associates evaluation, dated April 15, 1977, as shown in October 2026.

An experiment by budget cut

President Johnson asked Congress for $120 million in fiscal 1968 to serve up to 200,000 children leaving Head Start. Follow Through received $15 million. Both figures are in a 1977 paper by Eugene Tucker of the US Office of Education, who called the result "a program too small to be an ongoing service program and too large to be a workable research program."

The administrators turned it into research. John Evans, who had headed the Office of Education's planning and evaluation office, described that in 1981 as "a wise, indeed, brilliant decision", and spent the rest of his paper on what went wrong afterwards.

The design was called planned variation. Each sponsor, usually a university or a research laboratory, brought a model, and each community chose one. Stanford Research Institute collected the test data and, from 1972, Abt Associates analysed it.

The sources disagree on the size. Robert Egbert, the first director, wrote that at its peak the program had a budget of about $60 million and "approximately 60,000 children in 170 projects". Tucker gives 173 sites by 1970 and about 80,000 children a year. The National Institute for Direct Instruction says more than 200,000 students in 178 communities took part.

On cost, Evans puts program operations at "several hundred million dollars" and the evaluation at "in excess of $50 million". Tucker says "over 30 million". The copy of the report that ERIC catalogues was priced at $15.00, plus $1.50 if the order was not prepaid.

The models and the scores

Abt sorted the models by their stated aims. Basic skills models taught reading and arithmetic directly. Cognitive-conceptual models aimed at reasoning. Affective-cognitive models put self-concept first, on the theory that skills would follow.

The tests were sorted the same way. Most came from the Metropolitan Achievement Test. Raven's Coloured Progressive Matrices stood for problem solving, and the Coopersmith Self-Esteem Inventory and a scale of responsibility for one's own success stood for the affective side. A site scored a plus when its children beat the comparison group by a quarter of a standard deviation with statistical significance, and a minus for the reverse.

The table shows the nine major models. The descriptions come from Abt's model summaries, which each sponsor edited, as reprinted in a 1996 issue of Effective School Practices. The last three columns give the direction of Abt's index of pluses and minuses, read from Gary Adams's chart in that issue.

Model and sponsorAbt's typeIn the classroomBasic skillsCognitiveSelf-esteem
Direct Instruction, University of OregonBasic skillsScripted small-group lessons, "a fast moving series of programmed questions and answers"AboveAboveAbove
Parent Education, University of FloridaCognitiveParent educators assist in class and visit every homeSlightly aboveLevelAbove
Behavior Analysis, University of KansasBasic skillsProgrammed materials and a token economySlightly aboveBelowAbove
Language Development, Southwest LabBasic skillsBilingual teaching for Spanish-speaking childrenLevelBelowAbove
Bank Street CollegeAffectiveChildren plan their tasks with the teacherBelowBelowBelow
Responsive Education, Far West LabAffectiveSelf-paced exploration, a toy library for parentsBelowBelowBelow
Tucson Early Education Model, University of ArizonaCognitiveLanguage built from the children's own experiencesBelowBelowBelow
Cognitively Oriented Curriculum, High/ScopeCognitivePiaget-based; teachers act as "catalysts"BelowBelowBelow
Open Education, Education Development CenterAffectiveThe open classroom of the British infant schoolsBelowBelowBelow

"Above" and "below" are relative to comparison children in ordinary classrooms. Adams's chart and Cathy Watkins's 1997 monograph differ on details for the middle rows. They agree on the first row and the last five.

None of the affective models beat the comparison children on the self-esteem measures. Watkins reports that Direct Instruction and Behavior Analysis ranked first and second on them.

Tucker's own overview for the Office of Education counted 19 models and found two "generally effective", both in the structured category. He rated five "generally ineffective".

Table 1 of Eugene Tucker's April 1977 paper for the US Office of Education, from the public ERIC copy (document ED141449), retrieved October 2026.Table 1 of Eugene Tucker's April 1977 paper for the US Office of Education, from the public ERIC copy (document ED141449), retrieved October 2026.

The case against the evaluation

The Ford Foundation paid for a review before the Abt report was out. Ernest House, Gene Glass, Leslie McLean and Decker Walker published it in the Harvard Educational Review in 1978 as "No Simple Answer". Its abstract calls Follow Through "the largest and most expensive federal educational experiment in this country's history." I could open only that abstract, so I rely on a 1977 paper by House and Elizabeth Hutchins that makes the same case, and on passages quoted by the people who answered it.

Their facts are mostly right, and the Office of Education conceded the first one. Tucker wrote: "Random assignment of districts, schools, classes, or children was not attempted." Comparison classes were found afterwards and often differed from the Follow Through classes, so Abt adjusted for the differences statistically.

Sites varied more than models did. This was Abt's own first finding. House and Hutchins report that in "at least two or three" Direct Instruction sites the children did much worse than the comparison classes, and that even the best model accounted for less than ten percent of the variation in scores.

The tests suited some models better than others. After 1972 the evaluation was cut to 17 sponsors, 80 sites, about 20,000 children and four standardized instruments. House and Hutchins note that Direct Instruction's best result was on a language subtest about punctuation and sentence types, which looked a lot like its third-grade lessons. They also say the two affective instruments had "serious deficiencies". A model that promised creativity was graded on spelling.

When the House panel reran the comparison with sites as the unit, it got roughly the same ranking as Abt and no statistically significant differences.

The reply

The Abt authors answered in the same issue. Watkins quotes their best line: "Any program that wishes to rid itself forever of the discomforts of evaluation need only add to its list of objectives one metaphysical, obscure, or otherwise unmeasurable purpose".

The careful answer came in 1981 from Carl Bereiter and Midian Kurland in "A constructive look at Follow Through results", which I read in the 1996 reprint. Bereiter had led the Illinois preschool project where Siegfried Engelmann first taught, so he was not a neutral party.

The paper agrees with House that the site is the right unit. It then drops the comparison classes, which removes the worst of the mismatch problem, and compares the models with one another, using only models with six or more sites. Model differences were significant on every achievement subtest and explained "roughly between 17 and 55 per cent" of the variance left after adjusting for who the children were. Direct Instruction and Behavior Analysis were at or near the top on every subtest. Open Education and Responsive Education were at or near the bottom.

Bereiter and Kurland found no significant differences between models on Raven's, and none on the Coopersmith once reading ability was controlled. So the result that survives is narrower than the bar chart on NIFDI's page. Two structured models produced better reading, arithmetic, spelling and language scores by third grade than two child-centered ones, for poor children, on one test battery. The problem-solving and self-esteem advantages appear in Abt's counting and disappear in the stricter analysis.

Their paper also has the best sentence in the dispute: "Philosophies don't teach kids. Events teach kids".

The winner and the 21 sites

The first page of Commissioner Ernest Boyer's 1978 letter to Senator Bob Packwood, a US government document published on NIFDI's site, retrieved October 2026.The first page of Commissioner Ernest Boyer's 1978 letter to Senator Bob Packwood, a US government document published on NIFDI's site, retrieved October 2026.

Boyer's letter says the evaluation "forced us to shift attention more to successful individual projects." Because only one model had worked consistently, it would be "inappropriate and irresponsible to disseminate information on all the models". The office would fund 21 successful sites as demonstrations.

Schools applied one at a time to the Joint Dissemination Review Panel. Watkins reports that its criteria allowed gains in self-concept or in teacher behavior to count, and that 22 Follow Through projects were validated, some from models that had not raised achievement. Engelmann's own account says 21 Follow Through schools joined a network of 200 programs, and that three of the 21 used Direct Instruction. He remarks that the National Diffusion Network was well named.

The funding formula followed the same logic. Watkins, citing a 1983 account by Eugene Ramp, writes that from fiscal 1982 sponsors with validated projects received the lowest grants and sponsors with none received the highest, so that more projects could be validated. By her account none were the following year, and the policy was renewed.

Evans records that successive administrations tried to phase the program out and that Congress refused each time. Tucker saw why: Follow Through "continues on the strength of its political and grass roots support and not on the strength of its research findings." Bonnie Grossen, writing in the 1996 issue, says it ran until the summer of 1995 and cost about a billion dollars in all.

No conspiracy is needed to explain the outcome. The evaluation had real flaws, four methodologists said so in the Harvard Educational Review, and American school districts choose their own curricula. Education faculties had been cool to the model from the start. Grossen recounts that around 1970 Engelmann and Wesley Becker offered their grant of a million and a half dollars a year to 13 universities. Two replied, and the faculty of one of the two voted unanimously against.

NIFDI's public page on Project Follow Through, October 2026.NIFDI's public page on Project Follow Through, October 2026.

Direct instruction vs inquiry in later studies

Jean Stockard and colleagues pooled 328 studies of Engelmann's programs published from 1966 to 2016, with 413 designs and 3,999 effects. The average effect was 0.54 standard deviations, with a 95% interval from 0.49 to 0.59. Reading came in at 0.51, math at 0.55 and affective outcomes at 0.33. I read a preprint of the full paper. About a fifth of the effects came from designs with random selection or matched students. The paper's notes say some of the work was done while the authors were employed part-time by NIFDI.

John Hattie's database gives "direct instruction" a weighted mean of 0.56 across eight meta-analyses. His definition covers any structured, teacher-led sequence, and the eight range from 0.21 to 0.83, so it averages different things.

Capital-letter Direct Instruction means Engelmann's scripted programs. Lowercase direct instruction means explicit teaching in general, and Barak Rosenshine's "Principles of Instruction" is the usual summary: present new material in small steps, ask many questions, check every student's response, and aim for a success rate of about 80 percent during practice.

On the other side, Paul Kirschner, John Sweller and Richard Clark argued in 2006 that minimally guided instruction fails novices. Cindy Hmelo-Silver, Ravit Duncan and Clark Chinn replied that problem-based and inquiry learning are heavily scaffolded and had been lumped in with discovery learning unfairly. A 2011 meta-analysis by Louis Alfieri and colleagues of 164 studies supports both: explicit instruction beat unassisted discovery (d = -0.38 for discovery), and guided discovery with feedback and worked examples beat other instruction (d = 0.30). I read the abstracts of these three.

Manu Kapur's productive failure, which I covered in the Brilliant review, has students attempt a problem before being taught. Its meta-analysis found a moderate benefit overall and a reversed effect for second to fifth graders, which is close to the age Follow Through tested.

Scripts as conjecture and criticism

I think the evidence that explicit instruction with frequent checking works for novices is strong. Follow Through alone would not establish it. Together with the classroom studies Rosenshine draws on and Alfieri's 580 comparisons, it does.

My posts usually side with Karl Popper and David Deutsch, who hold that knowledge grows by conjecture and criticism and cannot be poured in. Read quickly, that sounds like an argument for the open classroom. I take it as an argument for the script. In a Direct Instruction lesson the teacher asks, the child answers aloud, and a wrong answer is corrected on the spot. The child has to produce the answer. Each response is a small conjecture and each correction is criticism, at a rate no lecture and no interest center matches. Abt's own description says the model aims at "intelligent behavior rather than specific pieces of information by rote memorization."

The open classroom gave six-year-olds problems and too little criticism. A child guessing at how letters map to sounds needs to be told quickly when the guess is wrong.

Popper also wanted a school with no "unwanted answers to unasked questions", as I noted in the Math Academy review. A script supplies the questions too. For a first grader who must learn to read, I think that is the right trade. An adult is in a different position. Curiosity decides what an adult wants to learn, and explicit instruction is then the efficient way through the prerequisites. Kirschner's abstract makes the same boundary from the other side: the advantage of guidance recedes once the learner knows enough to guide themselves.

Mastery learning is the other half of the method. Direct Instruction places children by test and holds them until they pass. Bloom's mastery classes, described in the history of tutoring systems, gained about a standard deviation from that rule alone, and Alpha School sets its threshold at 90 percent.

Follow Through says nothing about adults or about goals nobody measured.

Understand, and its limits

Understand is a voice tutor you can interrupt. It reads an article or PDF aloud word for word and stops when you ask a question. It asks you to explain an idea back, and it keeps a record of what you explained that you can read and correct.

Understand reading a text aloud, with the current word highlighted.Understand reading a text aloud, with the current word highlighted.

It is not Direct Instruction. It has no courses, no scripted sequence, no problem sets and no spaced repetition. For teaching a child to read, Engelmann's programs have fifty years of studies and Understand has none: no trial has evaluated it. There is no Android app. It is for adults and it is free for now.

Questions people ask

Did Direct Instruction win Project Follow Through?

On achievement tests, yes: in Abt's analysis, in the House panel's ranking and in Bereiter and Kurland's reanalysis. The self-esteem and problem-solving results depend on the analysis.

Is mastery learning the same as Direct Instruction?

No. Mastery learning, also sold as mastery-based learning, means passing each unit before moving on, with any teaching method. Direct Instruction uses mastery and adds scripted, field-tested lessons.