I've touched on assessment a couple of times in recent posts: from broken grading metrics, to driving the wrong behaviors, to the case for personalized instruction. Time to dig into to this one on its own, because unlike almost everything else in the AI-in-education debate, this part shouldn't be controversial at all.

Forty years ago, Benjamin Bloom ran an experiment that education never quite got over. Teach half a group of students the normal way, one teacher, thirty kids. Give the other half one-on-one tutoring plus mastery learning, no moving on until you've actually mastered the material. Test both groups.

The tutored kids didn't do a little better. They did two standard deviations better. The average tutored student outscored 98% of the regular classroom.

Bloom called it the two-sigma problem, and "problem" is right, because it was a taunt, not a triumph. We'd found the single most effective intervention in education, and it was unaffordable at scale. Good tutoring runs $1,800 to $4,000 a kid per year. Scale that to every student who could use it and you're past $15 billion, every year, forever. No district has that. No country does.

So we settled for cheaper substitutes, including:

  • Smaller classes: 0.13 to 0.20 standard deviations.
  • Summer school: 0.08 to 0.09.
  • Extended days: 0.05.

Worth doing? Yes, but none of them come close to tutoring which clears 0.29 to 0.55, more than double the next best option, and we've known that since the 1980s.

And that's the challenge, the solution isn't a secret, every teacher already knows it in their bones: exactly which kid needs five more minutes on fractions, which one's ready to move on. It's just that they just have twenty-nine other kids and one of them. The national student-teacher ratio is 15:1, and even that's a school-wide average, not what one teacher faces when the door closes. One-on-one was never on the table. It was the perfect solution that nobody could afford to act on.

And that's the part that's changing, and it's not just about cost. It's about using the right tool for the job. Because knowing what tutoring a kid needs is an exercise in pattern matching, and pattern matching just happens to be what AI is really good at (actually, it's the thing it's best at, full stop). Take a stack of essays, a semester of quiz results, a portfolio of projects, and find the patterns a human would eventually find too, given hours nobody has. Which kid keeps missing the same category of word problem. Whose writing has strong ideas buried under shaky mechanics. Whose project shows someone ready to move faster than the syllabus assumes. That's not AI bolted onto education because it's trendy. It's using AI to do exactly what the technology was built for.

What could this look like in practice? A kid turns in a math problem set. Instead of a single grade landing three days later, the tool scores it and flags the pattern underneath the wrong answers: right every time except when a negative number shows up in front of a fraction, that's not one mistake, that's a rule that hasn't locked in. It hands back specific feedback, not a final grade, try these three, here's what to watch for, and the student has another go. Submit, feedback, revise, same afternoon instead of the following week when the teacher finally reaches problem set 24 of 30. The teacher isn't out of the loop, they're looking at the dashboard: six kids across two periods hitting that same wall, which tells them exactly what to teach Tuesday instead of moving on. The AI closed the small, mechanical gap, fast. The teacher made the call only they can make: what the whole class needs next. That's the same thing a great teaching assistant would do, if every classroom could afford one.

It's also dramatically cheaper. A recent wave of studies on AI-assisted tutoring puts numbers on it:

  • 0.22 SD in Italian math classrooms for about $55 a student.
  • 0.49 SD in wartime Ukraine for under $100.
  • Simple phone-based mentoring in Bangladesh hit 0.75 SD for $20 a kid.

None of those are two-sigma results, but they're a fraction of tutor cost, available to every kid, not just the ones whose parents can write a check. And the UK study I mentioned last time (AI-tutored students posting median gains more than double their classroom peers) is the closest thing yet to Bloom's original number, nowhere near $4,000 a year.

Forget "AI helps you write an essay faster." This is the actual prize! Every kid working at their own pace, caught the moment they misunderstand something instead of three weeks later on a test, with a teacher who has finally been gifted the bandwidth to notice.

Three in ten teachers already use AI weekly, saving close to six hours a week, six weeks a year, mostly on lesson prep today, because that's what the tools are best suited for right now. But the direction points straight at the problem-set example above: free up the diagnostic grunt work, and a teacher spends their finite attention on the calls only a human should make. That's the two-sigma outcome: time to do the parts of the job that made them become a teacher.

Before we go any further, it's worth being very explicit in separating two questions that keep getting mashed together: should a kid be handed an AI tool to use alone, and should a teacher use AI as part of how they teach. The first deserves real caution: a student working unsupervised needs guardrails they may not have yet. The second is a different animal. A teacher running quizzes through a diagnostic tool isn't handing a kid anything unsupervised, it's a professional using a tool to do their job better, same as a doctor with diagnostic software or an accountant with a spreadsheet. The risk profile isn't remotely the same, which cuts the opposite direction from how these policies treat it: the lower the risk, the stronger the case for moving fast, and teacher-side diagnostics are about as low-risk as AI in education gets. Lumping it in with "a kid had ChatGPT write the essay" throws out the safest, best-evidenced use because it shares a category with the riskiest one.

Which is what makes the bans I wrote about recently so utterly frustrating, and why NYC's is worse than LA's on exactly this point. NYC still lets teachers use AI for lesson planning, but walls it off from "any decision about a student," broad enough to swallow low-stakes diagnostic feedback along with high-stakes calls like promotion. That's precisely the use case this piece is about, banned by name, guilt by association. LAUSD's version is narrower and better aimed: it restricts what shows up on a kid's own device, but leaves teachers free to use AI exactly this way. Neither is a full embrace, but one does a better job at risk assessment.

Here's the bottom line. It is irresponsible to throw away, for a year at a minimum, the greatest opportunity with the best evidence behind it of anything we've tried: individualized diagnosis and feedback, at a cost that finally makes it possible for every kid, not just the ones whose families can afford a tutor. We spent forty years wishing we could afford Bloom's result. We're closer than we've ever been, and 600,000 NYC Pre-K through 8th graders, in one policy alone, won't get near it for at least a year because the pendulum swung too far. And with so many school districts under pressure to evaluate their use of AI, I fear that others may follow NYC's well-intended yet ill-conceived example. That's a distressing thought.

So yes, I'll say it, teachers need to use AI! Or rather, students need teachers who use AI!