It arrives at almost every leadership table eventually, and it usually comes from the person furthest from the day to day: a governor, a trustee, a board member who has just read something worrying. How do we compare with other schools on AI? Are we behind?
Here is the honest answer. You are probably behind where you would like to be. So is almost everyone you would measure yourself against, and that is precisely why the comparison will not help you decide anything.
A benchmark tells you how common your position is. It cannot tell you whether your position is good enough. In a field where hardly anyone is ready, those two things point in opposite directions, and a leadership team that confuses them will walk out of the meeting reassured and no better prepared.
What a benchmark actually measures
A benchmark measures the field, not the job. It answers one question well: how many others are where we are? It has nothing to say about the question underneath, which is whether where we are is a sensible place to be.
That distinction would be academic if the field were in decent shape. It is not. In IDC's 2026 AI MaturityScape Benchmark, published on August 11, 2026 by Xiao Liu and Andrea Siviero and covering 1,900 organizations across 20 markets, 3.1 percent of organizations had reached the highest maturity stage and more than 61 percent were still in the two least mature ones. That is a cross-industry picture rather than an education one, which makes it more useful here, not less: the shortfall is not a peculiarity of schools.
Education's own numbers say the same thing from a different angle. The IBM and Morning Consult study of AI readiness in U.S. schools, published on September 2, 2026 and based on a July 2026 survey of 1,019 K-12 education professionals and 1,029 parents, found 76 percent of middle school educators reporting AI use in their classroom at least weekly, while only 20 percent said they had received extensive training in it. Asked whether the U.S. education system was adapting "very well" to advances in AI, 24 percent of the educators said yes.
Put those together and the median is doing something at scale it has not been prepared for. Landing on that median is not a clean bill of health. It is the problem, shared.
The peer group decides the answer, and somebody chooses the peer group
Before you accept any comparison, look at who is in it. The composition of the peer group, not your own performance, often decides whether you come out ahead.
RAND's American School District Panel work by Melissa Kay Diliberti, Robin J. Lake and Steven R. Weiner, published on April 8, 2025, found that 48 percent of districts had provided AI training to teachers by fall 2024. Underneath that single number the spread is wide: 67 percent of low-poverty districts, 42 percent of middle-poverty districts and 39 percent of high-poverty districts. A district that had trained its teachers would be comfortably ahead of the high-poverty group and behind the low-poverty one, on the same day, with the same training, having made the same decision.
So the honest version of "are we behind?" is always "behind whom?" And the peer group is rarely handed to a leadership team by anyone neutral. It is chosen, usually late, usually by whoever is bringing the slide. Nobody has ever chosen a comparison group that made them look worse.
Why your own position on the scale is the least reliable number in the room
Self-placement flatters. This is the finding from the IDC benchmark that should give any leadership team pause: among the organizations that identified themselves as thrivers, only about one in six actually scored at the managed or optimized levels when IDC applied its own method. Roughly five out of six self-described leaders were somewhere less mature than they believed.
That gap is not dishonesty. It is what happens when the people scoring the organization are the people who made its decisions, and when the evidence they have is the evidence they collected. In the strategy sessions I run, the pattern is consistent: the senior team scores the organization on what it has decided, and the people further down score it on what has actually changed in their week. The two answers are usually several rungs apart, and the second one is the one that would show up in an inspection, a data incident or a parent complaint.
A benchmark built on self-assessment inherits that gap and then hides it behind a percentile.
The Comparison Test
Here is the filter I suggest leadership teams run before any AI benchmark is allowed to change what they do, or to reassure them that nothing needs to.
The Comparison Test is three questions I ask of any AI benchmark before a leadership team takes comfort from it: Who is in this comparison, and who decided they were our peers? If our position on it improved, what would have had to change in the work? And would the people it describes, our staff and our students, recognize the picture it paints? A benchmark that survives all three is information. A benchmark that fails any of them is a mirror, and leadership teams have never had trouble finding those.
The second question does most of the work. A benchmark position that could improve without anything changing in a classroom is measuring activity, which is the same trap as counting logins. That is the argument I made about internal numbers in how to measure AI maturity; this article is about other people's numbers, and they fail in a slightly different way. Your own weak metric misleads you. A benchmark misleads you and tells you that everyone agrees.
This is offered as a way of reading a benchmark, not a validated instrument. It is meant to be used out loud, in the meeting, while the slide is still on the screen.
The three comparisons that do change a decision
If peer comparison is the wrong first move, something has to take its place. Three comparisons earn their place on an agenda, in this order.
Compare yourself against your own record. Not the average school: your own last five decisions about AI. What did you decide in the spring? What has happened since? Who owned it, and would they say the same thing if asked separately? This is a harder comparison than any benchmark because you cannot lose it to a peer group; you can only lose it to yourselves.
Compare yourself against the school your strategy describes. Every leadership team that has written anything down about AI has described an institution. The useful gap is between that description and the building you walked into this morning. If your strategy promises that staff know what they may and may not put into a tool, go and ask four of them. The distance between the document and the answers is your real position, and it is measured in decisions rather than percentiles.
Compare yourself against the floor, not the field. Some things are not a matter of degree. Is there a named person who answers when AI gets something wrong? Can staff say what they are allowed to use it for? Has anyone told parents what is happening in their child's lessons? RAND's national survey work led by Christopher Joseph Doss, published on September 30, 2025, found that only 45 percent of principals reported having a school or district AI policy at all. That figure is often quoted as evidence that a school without one is in good company. It is better read the other way. Some questions do not have a percentile. "Most schools cannot answer that either" has never worked as a defense to anyone who was actually harmed.
What Good Looks Like
Leadership teams that use benchmarks well have one habit in common: they treat the benchmark as a prompt for an internal question rather than as a verdict.
In practice that looks unremarkable. The comparison goes on the agenda as information, not as a score. Somebody names which decision it might change before anybody discusses where the school sits on it. When the number is flattering, the team asks what it would look like if it were wrong. When it is unflattering, the team asks whether the thing being measured is one they had chosen to do. Then the benchmark comes off the agenda and the three comparisons above go on it.
The teams that use benchmarks badly do something quieter and more damaging. They find the comparison reassuring and stop. Nothing is decided, the slide is filed, and the organization has converted a genuine question into an afternoon of relief.
The most useful sentence I hear in these sessions is not "we are ahead" or "we are behind". It is somebody saying: I do not think we know, and here is how we would find out.
When peer comparison helps
There is a real case for benchmarking, and it deserves stating properly rather than being waved away.
Comparison is useful when it is specific enough to act on. Knowing that a neighboring trust settled on three approved tools rather than thirty tells you something you can use, because it describes a decision rather than a position. Comparison is useful for the political work of leadership too: a board that has seen what similar schools have done will approve a proposal it would otherwise defer, and leaders who work in isolation make slower decisions than leaders who can see the shape of the field.
The failure is not comparison itself. It is comparison used as a verdict, arriving at the end of the discussion instead of the beginning, and answering a question about adequacy that it was never able to answer.
That confusion has consequences beyond the meeting. Reporting on the sector this month found leaders frustrated by exactly this gap between information and instruction. In Tes on September 9, 2026, Jabed Ahmed reported Pip Sanderson of the National Institute of Teaching describing the response leaders give after AI training sessions: "This is all great and really interesting, but what do I do?" A benchmark is expert at producing that feeling and useless at resolving it.
Where to start instead
If your board has asked how you compare, do not answer with a number. Answer with a date.
Take your last five decisions about AI, write down when each was made, who owns it and what changed in the four weeks afterwards. Then pick the three floor questions above and get real answers to them, from people who are not in the leadership team. That exercise takes an afternoon, it cannot be gamed by choosing a friendlier peer group, and it produces something a benchmark never does: a next move with a name and a date attached to it.
You will still want to know what other schools are doing. Ask them what they decided, not how they scored.
If your leadership team is working through these questions, this is the kind of work I support through AI strategy sessions and leadership mentoring, and the School AI Readiness Scorecard is a reasonable place to start a conversation that is about your own decisions rather than someone else's percentile. It pairs well with the seven dimensions of AI readiness if you want a structure for the discussion, and with what it actually means to be AI ready if the definition itself is still contested in the room.
Sources and further reading
- New IBM Study Finds AI Adoption Is Outpacing K-12 Readiness, IBM and Morning Consult, September 2, 2026.
- Ambition Is Everywhere, Maturity Is Rare: Inside IDC's 2026 AI MaturityScape Benchmark, Xiao Liu and Andrea Siviero, IDC, August 11, 2026.
- More Districts Are Training Teachers on Artificial Intelligence: Findings from the American School District Panel, Melissa Kay Diliberti, Robin J. Lake and Steven R. Weiner, RAND, April 8, 2025.
- AI Use in Schools Is Quickly Increasing but Guidance Lags Behind: Findings from the RAND Survey Panels, Christopher Joseph Doss and colleagues, RAND, September 30, 2025.
- Schools 'desperate' for clearer AI guidance, Jabed Ahmed, Tes, September 9, 2026.
Dan Fitzpatrick is the founder of The AI Educator, a Forbes contributor and the author of four bestselling books on AI in education. He works with school and organizational leaders on AI strategy, readiness and leadership development. More about Dan.


