A history of intelligent tutoring systems, from PLATO to LLM tutors

October 3, 2026 · 14 min read

In December 2021 I wrote an essay on tools for thought with a short section on intelligent tutoring systems. It opened with Prometheus, the fictional AI in Max Tegmark's Life 3.0:

"Given any person's knowledge and abilities, Prometheus could determine the fastest way for them to learn any new subject in a manner that kept them highly engaged and motivated to continue and produce the corresponding optimized videos, reading materials, exercises, and other learning tools."

I noted that ALEKS, the most successful intelligent tutoring system of the time, is limited to predefined domains and content, and predicted "the first more generally intelligent, Prometheus-like tutoring systems in the coming decades", built on reinforcement learning.

I was wrong on the timing and the method. Khan Academy announced Khanmigo, a tutor built on GPT-4, on March 14, 2023, fifteen months after my essay. It came from a large language model, and reinforcement learning had little to do with it.

Tegmark's sentence has two halves. An LLM can produce an explanation or an exercise on any subject, which is the second half. The first half, "given any person's knowledge and abilities", is what fifty years of intelligent tutoring systems research was about, and the LLM tutors mostly dropped it.

Teaching machines before computers

Sidney Pressey, a psychologist at Ohio State, showed a working teaching machine at the 1924 meeting of the American Psychological Association. Ludy Benjamin's history of teaching machines describes it: a drum exposed a multiple-choice question in a small window and the student pressed one of four keys. In "teach" mode the drum would not advance until the student pressed the right key. Pressey published it in 1926 as "A simple apparatus which gives tests and scores - and teaches".

B. F. Skinner's machines of the 1950s made the student write the answer instead of picking it, then slide a panel to compare it with the correct one. His case for them is "Teaching Machines", in Science in 1958. In Benjamin's summary, Pressey's devices depended on trial and error and Skinner's gave the learner a series of small successes. Neither machine knew anything about the student beyond the last response.

PLATO

PLATO (Programmed Logic for Automatic Teaching Operations) started in 1960 at the University of Illinois at Urbana-Champaign, under Donald Bitzer. In 2021 I wrote that ARPA funded it. The sequence is longer: small grants from a combined Army, Navy and Air Force pool for the first three versions, steady National Science Foundation funding from 1967, and ARPA sponsorship of terminals at military sites in the 1970s. Brian Dear, who wrote The Friendly Orange Glow, calls ARPA and NSF "the major backers" in his talk at Google.

PLATO IV, in 1972, had Bitzer's orange plasma display with a 512 by 512 bitmap, an infrared touch panel, and a terminal price of about $12,000.

Its users built things nobody had funded. In 1973 David Woolley wrote PLATO Notes, a message board, and Doug Brown wrote Talkomatic, a live chat with up to five active participants.

The lessons themselves were computer-assisted instruction, a fixed program of questions with branches. Dear on what the Illinois group noticed about the students who used them:

"One thing that was discovered early on PLATO was that students who have those kind of difficulties do really well interacting with a computer because none of the social peer pressures of performing well in front of all of these other kids in their class, none of that existed."

A screen that says "no, try again" does not laugh at you. NovaNET, the commercial descendant of PLATO, found its market in remedial education and in prisons. I wrote in 2021 that it shut down in the early 2000s. It was acquired then, and Pearson ran it until August 31, 2015.

SCHOLAR and the first intelligent tutoring systems

In 1970 Jaime Carbonell published "AI in CAI: An Artificial-Intelligence Approach to Computer-Assisted Instruction". His program, SCHOLAR, taught the geography of South America. Its facts were stored in a semantic network, separate from the teaching strategy, so the program could answer questions nobody had scripted and could ask its own. SCHOLAR, built with Allan Collins, is usually counted as the first intelligent tutoring system, and it tagged the parts of the network a student had shown they knew.

John Seely Brown and Richard Burton pushed the student model further. SOPHIE (1975, with Alan Bell) let a student troubleshoot a simulated electronic circuit, with no real equipment to break. BUGGY (1978) treated a child's wrong answers in arithmetic as the output of a procedure with a specific bug in it. A child who answers 52 - 38 with 26 has subtracted the smaller digit from the larger in each column. Brown and Kurt VanLehn's later catalogue calls this bug Smaller-From-Larger. A tutor that knows the bug can say more than "wrong".

William Clancey's GUIDON (1979) took MYCIN, the Stanford expert system for infectious diseases, and tried to teach from its rules. It needed 200 additional tutorial rules for guiding the dialogue and modelling the student, kept separate from the medical knowledge.

Derek Sleeman and Brown collected this work in a 1982 book, Intelligent Tutoring Systems, and the title became the name of the field. An intelligent tutoring system from then on meant a program with a model of the domain, a model of the student and a model of how to teach.

Bloom's 2 sigma problem

Benjamin Bloom's 1984 paper "The 2 Sigma Problem" rests on the dissertations of two University of Chicago doctoral students, Anania and Burke. They randomly assigned students in grades four, five and eight to a conventional class of about 30, a mastery learning class of about 30 with corrective feedback after each test, or a tutor working with one to three students. The subjects were probability and cartography, taught for 11 periods over three weeks. Bloom reported:

"the average student under tutoring was about two standard deviations above the average of the control class (the average tutored student was above 98% of the students in the control class)."

Mastery learning alone gave about one standard deviation. The "problem" in the title is finding group instruction as effective as tutoring, which Bloom called "too costly for most societies to bear on a large scale".

Kurt VanLehn's 2011 review collected the experiments run since. Human tutoring against no tutoring averaged an effect size of 0.79. Step-based intelligent tutoring systems averaged 0.76, and answer-based systems of the PLATO kind 0.31. VanLehn attributes most of Bloom's number to the mastery criterion: tutored students had to score 90% before moving on, the mastery class 80%. In his words, "human tutoring is not usually 2 sigmas more effective than classroom instruction."

Paul von Hippel reread the dissertations for Education Next in 2024. The topics were chosen because the students had never seen them, and the tutored students got extra quizzes with feedback and, by Burke's account, about an hour more instruction per week. He cites meta-analyses that put the average effect of tutoring at 0.33 (Cohen, Kulik and Kulik, 1982) and 0.37 (Nickow, Oreopoulos and Quan, 2020) standard deviations.

Measured against the 0.79 that human tutors typically achieve, the step-based software of 2011 had already caught up.

Cognitive Tutors at Carnegie Mellon

John Anderson's group at Carnegie Mellon built tutors on ACT-R, his theory that a skill is a set of production rules, each an if-then step. Their ten-year review covers tutors for LISP, geometry and algebra. In model tracing, the tutor solves the problem alongside the student with the same rules and matches each step the student takes to a rule, correct or buggy. In the best cases students reached the proficiency of conventional instruction in a third of the time.

Albert Corbett and Anderson added knowledge tracing in 1995. For each of several hundred rules the tutor keeps a probability that the student has learned it, updated after every step, and it keeps assigning exercises until every rule counts as mastered. The model, now called Bayesian Knowledge Tracing, has four parameters per skill: prior knowledge, the chance of learning at each attempt, the chance of a lucky guess and the chance of a slip.

Carnegie Learning was founded in 1998 to sell the algebra tutor to schools, and RAND tested it at scale. Pane, Griffin, McCaffrey and Karam (2014) randomly assigned matched pairs of schools in seven states to adopt Cognitive Tutor Algebra I or keep their curriculum. There was no effect in the first year. In the second year high school students gained about 0.2 standard deviations on an algebra proficiency exam, roughly eight percentile points for the median student. The middle school estimate was similar in size and not statistically significant. VanLehn's average for step-based tutors was 0.76.

AutoTutor and tutoring by dialogue

Art Graesser's group at the University of Memphis built AutoTutor in the late 1990s to hold a conversation. It asks a question that needs a paragraph to answer, such as a problem in Newtonian physics. A curriculum script lists the "expectations" a good answer contains and the misconceptions students usually have. The tutor then pumps for more, hints, prompts for a missing word and corrects, until the student has stated each expectation in their own words. It spoke through an animated agent with synthesized speech, and each question needed a hand-written script.

Graesser's 2004 summary reports gains of about 0.7 standard deviations on deep comprehension questions against reading the textbook or reading nothing, two conditions that barely differed from each other.

ALEKS and knowledge space theory

Jean-Paul Doignon and Jean-Claude Falmagne introduced knowledge spaces in a 1985 paper, "Spaces for the assessment of knowledge". Falmagne started building software on it at UC Irvine in 1994 with a National Science Foundation grant (UCI News). That became ALEKS, Assessment and Learning in Knowledge Spaces, which McGraw-Hill Education agreed to acquire in June 2013.

Knowledge space theory describes a subject as a set of problem types. A knowledge state is the subset a particular student can solve. Most subsets cannot occur, because some items are prerequisites for others. Take fractions:

  • a: multiply single-digit numbers
  • b: add fractions with the same denominator
  • c: find a common denominator (needs a)
  • d: add fractions with different denominators (needs b and c)

Four items have 16 subsets. Only seven are feasible states: {}, {a}, {b}, {a, b}, {a, c}, {a, b, c} and {a, b, c, d}. Nobody is in state {d}. This structure is a learning space, meaning every state can be reached from the empty one by adding one item at a time. A student in state {a} is ready to learn b or c, which the theory calls the outer fringe of the state.

ALEKS describes Algebra 1 as about 350 concepts and millions of feasible states, and says its adaptive assessment locates a student among them in 25 to 30 questions. Then it offers the student only items from their outer fringe.

The learner model here is the most precise of any system in this history. It costs years of expert work per course, and it exists for school and college mathematics, chemistry and a few neighbours.

Khan Academy and knowledge tracing

Khan Academy, incorporated in 2008, is computer-assisted instruction with mastery learning: videos, then exercises, with each skill moving through levels from Familiar to Proficient to Mastered.

Deep Knowledge Tracing (Piech and colleagues, 2015) trained a recurrent neural network on 1.4 million exercise attempts by 47,495 of its students to predict whether the next answer would be correct. It reached an AUC of 0.85 against 0.68 for Bayesian Knowledge Tracing, without anyone encoding which exercise tests which skill. A year later Khajah, Lindsey and Mozer showed that Bayesian Knowledge Tracing with a few known extensions performs about as well. Both models predict whether the next answer will be right, which is a narrower thing than what the student understands.

LLM tutors and the evidence so far

After Khanmigo came Google's LearnLM in May 2024, a family of Gemini models fine-tuned for learning, and OpenAI's study mode for ChatGPT on July 29, 2025. OpenAI's announcement says study mode "is powered by custom system instructions" and personalises from "questions that assess skill level and memory from previous chats". The LearnLM paper treats pedagogy as instruction following, where a teacher or developer describes the behaviour they want in the system prompt.

The best result I know of is Kestin and colleagues (2025). In Harvard's introductory physics course for life science students, 194 students were taught two lessons in consecutive weeks, one by an AI tutor at home and one in an active-learning class, with the groups swapped in the second week. Learning gains with the AI tutor were over double, an effect of 0.73 to 1.3 standard deviations, in a median of 49 minutes against 60 in class. The limits are in the paper: two lessons, a test taken right after each, and "expert-crafted, question-specific prompts written by instructors" for every problem.

Bastani and colleagues (2025) gave nearly a thousand high school students GPT-4 during math practice. With an interface like standard ChatGPT their practice grades rose 48%. Once access was taken away they scored 17% below students who never had it. A version prompted to give teacher-designed hints instead of answers raised practice grades 127%, and the abstract says the harm was "largely mitigated". The authors' word for how students used the plain version is "crutch".

Google and Eedi ran an exploratory trial with 165 students in five UK secondary schools. Human tutors supervised every message LearnLM drafted and approved 76.4% with at most a character or two changed. Students tutored this way solved a novel problem on the next topic 66.2% of the time, against 60.7% with human tutors alone. That measures LearnLM as a drafting aid for a tutor.

Timeline of intelligent tutoring system examples

YearSystem or paperPeopleModel of the learner
1924Teaching machineSidney PresseyNone (last key pressed)
1960PLATODonald Bitzer, University of IllinoisBranching on answers
1970SCHOLARJaime Carbonell, Allan CollinsTags on a semantic network
1975SOPHIEBrown, Burton, BellA simulated circuit to troubleshoot
1978BUGGYBrown, BurtonThe buggy procedure behind wrong answers
1979GUIDONWilliam ClanceySubset of MYCIN's rules
1985Knowledge spacesDoignon, FalmagneFeasible knowledge states
1985LISP TutorAnderson, ReiserModel tracing over production rules
1994ALEKS (development begins)Falmagne, UC IrvineA state in a learning space
1995Knowledge tracingCorbett, AndersonProbability of mastery per rule
1999AutoTutorGraesserExpectations the student has stated
2015Deep Knowledge TracingPiech and colleaguesNeural prediction of the next answer
2023KhanmigoKhan Academy, on GPT-4Not described in the launch post
2024LearnLMGoogleThe conversation and a system prompt
2025ChatGPT study modeOpenAIThe conversation, plus chat memory

The learner model went missing

Every system before 2023 in the table above could say what a particular student knows, and could say it only inside a domain that experts had mapped by hand.

An LLM tutor has the opposite profile. It will teach mortgage amortisation or CRISPR on request. What it knows about the learner is the conversation so far and, in ChatGPT's case, free-text memories of earlier chats. I read OpenAI's study mode announcement and Google's LearnLM paper for a description of a per-concept learner model and found none.

The old learner models also measured something narrower than understanding. A knowledge state in ALEKS is a set of problem types answered correctly. David Deutsch, in the first chapter of The Fabric of Reality, writes that

"understanding does not depend on knowing a lot of facts as such, but on having the right concepts, explanations and theories."

Correct answers are evidence of an explanation in the learner's head, and weak evidence, since a memorised procedure produces them too. BUGGY was built on that observation.

Deutsch's The Beginning of Infinity follows Karl Popper in holding that knowledge is created by conjecture and criticism. In Objective Knowledge Popper called the alternative the bucket theory of the mind, in which knowledge pours in through the senses. A lecture assumes the bucket. So does an LLM that answers every question in four fluent paragraphs, and Bastani's exam results show what the learner keeps.

If Popper is right, a learner hearing an explanation has to guess what it means and test the guess. A tutor can make that step visible by asking the learner to explain the idea back. The learner's explanation is a conjecture and the tutor's job is to criticise it. AutoTutor did this inside its scripts. It is also the only kind of evidence about understanding that I trust: what the learner said, in their own words, when asked.

Understand

I am building Understand, a voice tutor, around that gap. ALEKS and its relatives serve students inside a fixed curriculum. An adult who reads about monetary policy this week and protein folding the next had no learner model at all.

Understand keeps a graph of concepts for each learner, built from conversation in any subject. Under each concept sit dated pieces of evidence, marked as a success or a struggle, with a note on the context. Evidence comes from what the learner said, and the pass that extracts it after a conversation discards any item whose quote does not appear word for word in the learner's own lines. Hearing an explanation counts for nothing. A spoken claim like "I build storage engines" is stored as a lead for the tutor to test.

The tutor treats a concept as something it can build on after two successes in two different contexts, and stops if the latest evidence is a struggle. The learner can open the graph, read every note, rename, merge, move or delete what is wrong, and mark a concept as known. I think a learner model its subject cannot read and correct is unfinished, since the tutor's model of you is a conjecture too.

It has no prerequisite structure like a learning space, and the two-successes rule is a heuristic I chose. I have no effect size to report. A randomised comparison against a plain LLM chat, with a test a week later and no tutor in the room, is the experiment that would tell me whether the graph earns its place.