Ask a leadership team how they know their AI work is making a difference, and you will usually be shown a number that proves the tool was switched on.
Here is the answer nobody sells you: you cannot prove that AI caused an outcome in your school, and you should stop promising a board that you will. Proof in the strict sense means ruling out the other explanations, and in education there are always other explanations. What you can do, and what a serious board will accept, is lay out a chain of evidence and say plainly which link in it is weakest. That is a different promise, it is one you can keep, and leaders who make it are trusted further than leaders who arrive with a percentage and no provenance.
The question is live right now because the first evidence conversations of the year are landing. Governors and boards are asking about the AI work approved last year, and the honest answers are hard to find. Katelyn Schoenhofer of Wichita Public Schools put it to NPR's Lee V. Gaines on 28 September 2026 in the plainest terms anyone has managed: "We wrestle with what does success really mean when it comes to AI use?" She is not confused. She is describing a question that has no settled answer in the field, and pretending otherwise is how leadership teams end up defending figures they cannot source.
Two different questions get called "impact", and leaders keep answering the wrong one
"Did AI improve learning?" and "is this worth continuing?" are different questions with different evidence bars, and the second is the one your board is actually asking.
The first is a research question. It wants a causal effect, isolated from everything else happening in a school, and it is answered by trials with comparison groups. Stanford's SCALE initiative published the state of that research on 11 March 2026: Lily Fesler, JP Martinez Claeys, Chris Agnew and Susanna Loeb reviewed a repository that held "over 800 academic papers relevant to AI in K-12 education" as of October 2025 and found that "only a small subset (20 papers) produce strong causal evidence". More pointed still, for anyone preparing an American board paper: "None of these student-facing causal studies were conducted in U.S. K-12 school settings."
Read that twice. Twenty studies worldwide clear the causal bar, and none of the student-facing ones were run in a United States classroom. If the global research base cannot answer the first question yet, a single school with one term of data and no comparison group is not going to answer it either. Any leader who claims to has either run a trial they are not telling you about or is reporting something else and calling it impact.
The second question, the one a governing body is genuinely putting to you, is narrower and answerable. Is the work doing what we said it would do, is anyone worse off, and should it continue? That does not need a randomized trial. It needs an honest account of what changed and a clear statement of what you cannot yet show.
Most leadership teams answer the research question badly instead of the governance question well.
The Evidence Chain: four links, and the one that usually breaks
The Evidence Chain is the four links I ask leadership teams to lay out before they claim AI is having an impact: activity, that it happened; behavior, that the work is done differently now; outcome, that something the organization already cared about moved; and attribution, that the AI was the reason. Each link has to hold on its own, and a chain is only as strong as the link you cannot show. Most organizations can evidence the first link, assume the second, hope for the third and claim the fourth. The honest report names the link where the chain breaks and says what would close it.
This is my suggested way of organizing an evidence conversation, not a validated instrument. It is also not a maturity scale. The difference between a usage number and a maturity number is about whether your organization has changed; the chain is about whether a particular claim is safe to make out loud. You can be mature and still unable to close link four.
1. Activity: it happened
Activity evidence shows that the tool was available and that people opened it, and it is the only link most schools can currently produce.
Seat counts, login frequency, the number of staff who attended the training. This is the easiest evidence in the building and the least informative, because it rises whether you are leading this well or badly. Bellwether Education Partners put the problem precisely in "Measuring Artificial Intelligence in Education", published in October 2025 by Michelle Croft, Amy Chen Kulesa, Marisa Mission and Mary K. Wells: "most tools are measured by easy-to-track outputs (e.g., hours saved, number or frequency of logins, or features used), rather than whether they improve instruction, advance performance, or foster deeper student learning."
Collect it, report it in one line, and never let it travel under the word impact. A number that rises when the tool is switched on is evidence that the tool was switched on.
2. Behavior: the work is done differently now
Behavior evidence shows a named practice that is genuinely done differently from the way it was done a year ago, and it is the first link worth anything.
The test is whether someone who did not do the work could see the change. Not "staff feel more confident", which is a survey of feelings. Something a visitor could observe: the department's assessment feedback now goes out within two days instead of two weeks and has done for a term; the cover planning that used to take a middle leader a Sunday afternoon now happens in a shared document on Friday; every education, health and care plan review in this school now starts from a drafted summary that a named person checks.
Behavior evidence has a property that makes it unusually valuable to a board: it survives the tool. If the supplier disappears, the changed practice is still visible, and so is the decision that produced it. If you can evidence nothing else, evidence this.
3. Outcome: something you already cared about moved
Outcome evidence shows movement in a measure that existed before AI arrived, in the terms your organization already uses.
This is where leadership teams invent a metric, and inventing one is a confession. If the only measure that moved is one you created in order to show that the AI moved something, you have marked your own homework and published the grade. Use what was already on your improvement plan: attendance at the parent evening, the proportion of pupils who hand in the extended piece, staff retention in the department that was struggling, the number of complaints about feedback.
Two cautions. Measures move for many reasons, so write down the measure and its baseline before you start, not after, a discipline Bellwether frames as beginning with "a clear theory of change". And some outcomes are slower than your reporting cycle. A claim about reading comprehension at the end of one term is not patience, it is noise.
4. Attribution: the AI was the reason
Attribution evidence shows that the AI, rather than everything else you did that year, produced the movement, and this is the link a single school almost never closes.
The regulator has said so out loud. Ofsted visited 21 schools and further education colleges for "'The biggest risk is doing nothing': insights from early adopters of artificial intelligence in schools and further education colleges", published on 27 June 2025, and found that most leaders "relied on feedback from staff and students or tracked and monitored staff usage of AI (artificial intelligence) tools, rather than collecting data that could be used to measure the impact of AI (artificial intelligence) specifically on pupils". The report then concedes the difficulty instead of scolding anyone for it: "it is not always possible to evaluate the extent to which any impact was due to AI (artificial intelligence), rather than to any other factors such as pupils' prior knowledge, or the teaching approach."
That is an inspectorate saying the attribution link is hard in principle, not just hard for you. It is the single most useful sentence a school leader can take into a governors' meeting this term, and it is worth quoting to the governor who wants a causal number by Christmas.
So the honest position on link four is a sentence, not a figure: "We cannot attribute this to the AI, and here is the comparison that would let us try."
What I Tell Leadership Teams
What I tell leadership teams is that the chain is a reporting instrument before it is a measurement one, and that its value is in the link you admit you cannot show.
Across the leadership teams I work with, the mistake I see most often is not weak measurement. It is a strong claim resting on a weak link, made in a room where nobody has the standing to ask which link it rests on. A paper arrives saying the AI pilot "saved staff 200 hours and improved outcomes". The 200 hours is a supplier estimate from activity data. "Improved outcomes" is a hope attached to it with a conjunction. Nobody in the room can see the join, so nobody tests it, and the claim is now on the record and will be quoted back to that leader for years.
The teams that handle this best do something that feels, for about ten minutes, like weakness. They put the open link in the paper. They write the sentence that says what they cannot yet show. Then the conversation changes shape: the board stops interrogating the number and starts discussing whether closing that link is worth the effort. That is a governance conversation, which is what the meeting was for.
My own confidence in this comes from the other side of the table. Writing about AI and education for Forbes and in four books has meant reading a great deal of research and a great deal of material that quotes research it has not read, and the gap between the two is almost always an unexamined attribution claim. In a published study that claim has to be defended. In a board paper it rarely has to be, which is exactly why a leader should defend it unprompted.
The four lines an evidence note carries
An evidence note is four lines long, and a leadership team that can write all four is ready for the conversation.
- The claim. One sentence, in the school's own language, naming what changed and for whom.
- The link it rests on. Which of the four links your evidence actually reaches, named by its name.
- The link that is open. The thing you cannot show, stated before anyone asks.
- What would close it. The comparison, the time, or the measure that would move the claim up a link.
A hypothetical example, to show the shape. A trust reports that its secondary English departments now return written feedback within three days, down from roughly ten, and have done since January. That is behavior evidence, link two, drawn from the departments' own records. The open link is outcome: no change has been detected yet in the proportion of pupils acting on the feedback. What would close it is one term of the existing work-scrutiny data compared against the two departments that have not adopted the practice. Four lines, no invented figures, and nothing a governor could later discover was softer than it looked.
Compare that with "AI is saving our English teachers seven hours a week". The second sentence sounds stronger and is worth less, because the first one tells a reader where it stops.
When a stronger claim is warranted
A stronger claim is warranted when the tool was built for a specific instructional job, carries published evidence of its own, and you ran a real comparison, and those three conditions hold together more rarely than suppliers suggest.
The Stanford review is encouraging on the first condition: "Tools designed with pedagogical guardrails (such as AI chatbots for tutoring that provide step-by-step reasoning instead of direct answers) show more promise than general purpose AI tools." A tutoring system designed to withhold the answer and give a hint is a different proposition from a general chatbot, and the evidence treats it differently.
On the second, the base rate is sobering. Instructure, working with the nonprofit InnovateEDU, reviewed 150 widely used classroom technologies for "How to Choose Safe and Effective Classroom Technology", published on 10 March 2026, and found that "40% of purpose-built edtech tools have identifiable evidence aligned to the Every Student Succeeds Act (ESSA), compared to just 2% of consumer technologies used in classrooms". The Every Student Succeeds Act is the United States federal law that sets tiers of research evidence a school can rely on. Their warning is the one to carry into any demonstration: "A polished interface or time-saving feature does not replace verified impact on teaching and learning."
The third condition is yours alone. Two parallel groups, the same teacher or the same unit, one using the tool and one not, agreed before you start. That is less than a trial and far more than an anecdote, and a school with several forms of entry can usually arrange it.
One case calls for the opposite judgment, and chasing link four there wastes everyone's attention. Where the AI is doing administrative work and the outcome you care about is the released time itself, the chain is genuinely shorter: behavior evidence plus a named place the time went is the whole argument. The way that goes wrong is different, and it is the subject of why AI pilots go nowhere: not weak evidence, but nobody deciding what the released capacity is for.
What to do before your next governors' meeting
Take the single boldest AI claim in your most recent paper and work out which link it rests on.
Most teams find it rests on link one and was written as though it reached link three. Rewrite it as four lines. Then pick one practice you believe has genuinely changed, and gather the behavior evidence for it properly: the record, the date it changed, the person who would notice if it stopped. One well-evidenced link two beats four assumed links, and it is the foundation everything else is built on.
If you want a wider view of where evidence sits among the other things that have to be in place, it is the seventh dimension of AI readiness, and it is usually the weakest of the seven. It sits inside the larger question of what it actually means to be AI ready, which is a capacity to make and hold decisions rather than a stock of tools. And before anyone asks how you compare with other schools, it is worth knowing what a benchmark actually tells you, which is how common your position is rather than whether it is good enough.
If your leadership team is working out what it can honestly claim about its AI work and what it cannot, this is the kind of question I work through in AI readiness reviews with schools and trusts. I also write about this weekly for leaders who would rather see the evidence than the headline, in the newsletter.
Sources and further reading
- Lily Fesler, JP Martinez Claeys, Chris Agnew and Susanna Loeb, "The Evidence Base on AI in K-12: A 2026 Review", Stanford SCALE Initiative, 11 March 2026.
- Ofsted, "'The biggest risk is doing nothing': insights from early adopters of artificial intelligence in schools and further education colleges", 27 June 2025.
- Michelle Croft, Amy Chen Kulesa, Marisa Mission and Mary K. Wells, "Measuring Artificial Intelligence in Education", Bellwether Education Partners, October 2025.
- Instructure with InnovateEDU, "How to Choose Safe and Effective Classroom Technology", 10 March 2026.
- Lee V. Gaines, "Schools are experimenting with AI with little evidence or policy to guide them", NPR, 28 September 2026.
Dan Fitzpatrick is the founder of The AI Educator, a Forbes contributor and the author of four bestselling books on AI in education. He works with schools, trusts and organizations on leading AI well. More about Dan.


