A teacher puts a piece of work in front of you and says the sentence you have started to hear every September: "There is no way she wrote this." The vocabulary has jumped two years. A detector returned a number in the nineties. It does not sound like her. The student says she wrote it, and she is upset that anyone thinks otherwise. Now it is your decision, and whichever way you go, that family will remember it for years.
Most schools have a policy saying AI misuse is not allowed. Very few have decided what happens in the ten minutes after a teacher becomes suspicious. That gap is where the damage happens.
What should a school do when a student is accused of using AI?
Treat it as a question about the work, not a verdict about the student. Start from the rule you actually published, gather evidence about how the work was made rather than a score claiming what made it, give the student a genuine chance to explain, and decide on what you can show rather than on how certain the teacher feels. If you cannot show it, you mark what you can authenticate and you change the task next time. You do not convict on a hunch, and you do not let the hunch quietly lower a grade either.
The reason this goes wrong is that schools collapse three separate decisions into one. There is a question about the work: is this authentic enough to mark? There is a question about conduct: has a published rule been broken, and which one? And there is a question about the student: what does this young person need next, which may be teaching rather than sanction. Answer them in that order and most cases resolve without a disciplinary process. Answer them all at once and you get the case that ends up in a complaint, or in a lawyer's inbox.
Why you cannot prove a student used AI
You cannot prove it because no available tool can tell you, and the organizations responsible for qualifications have stopped pretending otherwise.
The Joint Council for Qualifications, which sets the rules for GCSEs and A levels in England, is explicit in its guidance AI Use in Assessments: Your role in protecting the integrity of qualifications, published on April 30, 2025. Detection tools, it notes, "will give lower scores for AI-generated content which has been subsequently amended by students, as they base their scores on the predictability of words." It recommends using more than one tool "to provide an additional source of evidence," and it puts the weight somewhere else entirely: "teachers will know their students best and so are best placed to assess the authenticity of work." The tool is a source of evidence. The judgment stays with a person.
The research is blunter. In a 2023 study, GPT detectors are biased against non-native English writers, Weixin Liang and colleagues at Stanford found that detectors "consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified." Their recommendation was to "caution against their use in evaluative or educational settings, particularly when they may inadvertently penalize or exclude non-native English speakers." Read that against your own roll. The students most likely to be wrongly flagged are often the ones with the least standing to argue back.
Institutions have started acting on this. Inside Higher Ed reported on August 5, 2026 that at least a dozen universities have disabled AI detection tools, Yale, Vanderbilt, Johns Hopkins, Northwestern, Georgetown and New York University among them, many after Turnitin's detector was found to produce far more false positives than originally claimed. They have moved the effort to oral defenses, in-class writing and asking students to show their reasoning. Chalkbeat's Lily Altavena found on September 2, 2026 that more than a third of states whose AI guidance mentions detection advise against relying on it alone, while most state guidance gives academic integrity a page or less and leaves schools to work out the rest. MIT's Justin Reich told her there is no evidence-based policy here yet.
So a detector score is an input to a professional judgment. It is never the judgment. This belongs on the short list of decisions you never let AI make in your school, alongside anything that determines a child's future.
The same logic cuts the other way, and leaders forget this half. A clean detector report is not evidence of innocence either. It is evidence of nothing.
How often does this actually happen?
Formally, far less often than staffroom conversation implies, and that gap is the thing to worry about.
Ofqual's official statistics for the summer 2025 exam series in England, published on December 11, 2025, record 100 penalties issued to students for plagiarism involving the misuse of artificial intelligence. That is 75.0 percent of all plagiarism penalties and 2.0 percent of all penalties issued to students, up from 85 cases the previous summer, when AI accounted for 55.4 percent of plagiarism penalties and 1.7 percent of the total. The share is rising fast. The absolute number is small.
Now set that against how often a teacher in your building has privately doubted a piece of work this term. The formal system sees a hundred cases a year nationally. Your staff see several a week. Almost none of those reach a process, which means they are being settled informally, by whoever happened to be holding the book, with no record, no consistency and no appeal. Two students who did the same thing in different classrooms get different outcomes, and nobody in leadership ever finds out.
That is not a detection problem. It is a governance problem, and it is yours.
The Burden Test: three questions before a suspicion becomes a case
This is the filter I suggest leadership teams put between a teacher's doubt and a school's action.
The Burden Test is three questions I ask before a school lets a suspicion about AI become a case against a student: Can we say what this student is alleged to have done in one sentence, without using the word AI? Would anything the student says change the outcome? And if we are wrong, what does this student lose, and can we give it back? A suspicion that survives all three is a case worth opening. A suspicion that fails any of them is a feeling, and a feeling should travel no further than the person who had it.
Take them one at a time, because each catches a different failure.
"Without using the word AI" forces you to name the rule. "She used AI" is not an offense. Submitting work as your own that you did not produce is an offense. Failing to acknowledge assistance where the task required it is an offense. Using a tool in a task that specified no tools is an offense. If you cannot complete the sentence, the problem is your task design or your published rule, not the student.
"Would anything the student says change the outcome?" tests whether you are holding a meeting or a sentencing. If you have already decided, do not stage a conversation to make the decision look fair. Staff can usually tell the difference, and students always can.
"If we are wrong, what does this student lose?" sets the standard of evidence. A conversation and a redraft is a low-stakes outcome, so a reasonable doubt is enough to justify it. A malpractice report to an awarding organization can end a qualification, so that demands evidence you would be willing to show the student's parents line by line. Match the burden to the consequence. Most schools apply one standard to everything, which means they are either too heavy for the small cases or far too light for the serious ones.
Treat it as a way of thinking rather than a validated instrument. It is a filter, not a finding.
What the process should be, step by step
Six steps, in this order, with a named owner for each. Copy them into your policy and change the job titles to match your structure.
- The teacher records the observation, not the conclusion. What in the work prompted the doubt, quoted. What the student's previous supervised work looks like. Any detector output, with the tool named and the date. No verdict.
- One named person triages within two school days. A designated leader, not the subject teacher, decides which of three routes this takes: no case, an academic conversation, or a formal investigation. Speed matters because memory of how the work was written fades within about a week.
- The student is told what is alleged, in writing, before any meeting. Including the specific rule and the possible outcomes. Nobody, adult or child, should walk into a meeting not knowing what it is about.
- The meeting asks about the work, not about the tool. Talk me through how you planned this. Which part took longest. Why did you choose that example. What does this word mean. A student who wrote it can nearly always do this. A student who did not usually cannot. It is the most reliable instrument you have, and it is free.
- The decision is written down with its reasoning and its evidence. Including the standard applied. One page. This is what protects the school if the decision is challenged, and it is what makes two classrooms consistent.
- Someone reviews the pattern each term. How many suspicions, in which subjects, about which groups of students, with what outcomes. If your flagged students cluster by ethnicity, by language background or by additional needs, you have found something more serious than cheating.
The pre-meeting checklist is shorter, and it is the one to put in front of staff:
- Which published rule is alleged to have been broken, in one sentence?
- Did the task state what AI use was permitted? Where, in writing?
- What does this student's supervised work look like?
- Is there process evidence: drafts, version history, notes, search history in a school account?
- What is the worst outcome available here, and does our evidence justify it?
- Who else has seen this, and who is deciding?
- Has the student been told what is alleged?
If more than two of those are blank, you are not ready to have the meeting.
What I tell leadership teams
The question I get asked most often about AI and cheating is the wrong question.
It arrives most weeks, through the newsletter I write for more than 44,000 educators and in the questions that follow the daily podcast, and it nearly always takes the same shape: which detector should we use? I understand the appeal. It is a procurement question, and procurement questions feel like progress. But it asks a machine to carry a decision about a child's integrity, and no machine can carry that. The better question is the one almost nobody asks: what will we do when we are not sure? You will not be sure. That is the permanent condition now. A school that has decided in advance how it behaves under uncertainty stays calm in March. A school that has not decided improvises, and improvisation under pressure favors whoever is most confident in the room.
Across the leadership teams I work with, the mistake I see most often is treating this as a discipline question when it is a task design question wearing a discipline costume. If a piece of work can be produced at home, alone, in one pass, by a machine, and you have no visibility of how it was made, then you have built an assessment you cannot authenticate and staffed it with a detector you cannot trust. The fix is upstream, in how the work is set and what counts as evidence of learning, not downstream in the meeting room.
The second pattern is quieter and worse. Where no process exists, a teacher who suspects something often does nothing, or marks the work down without saying why. The student never gets to answer. Nothing is recorded, nothing is taught, and the school's standards erode invisibly. An absent process does not produce fewer accusations. It produces unappealable ones.
What if the student really did it?
Then act, proportionately, and be clear that you are teaching first and sanctioning second.
The counterargument deserves taking seriously: some students are handing in work they did not write, and a school so anxious about false accusations that it challenges nothing teaches a cohort that the rules are decoration. That is a real failure, and it is not the safer option. The answer is not timidity. It is a process strong enough that staff trust it and use it.
Proportionality does most of the work. A 13-year-old who used a chatbot for a homework paragraph because they did not understand the task needs teaching, a redraft and a sentence in a record. The same behavior in a piece of assessed coursework, after signing an authentication declaration, is a different matter, and the JCQ guidance is clear that once that declaration is signed the case must be reported to the awarding organization rather than handled in school. Know which side of that line you are on before you start, because it determines who owns the decision.
One distinction leaders conflate constantly: undisclosed use is the offense, use is not. If your rule does not tell students what they are allowed to do, you cannot fairly punish them for guessing, and a blanket restriction you have not thought through will be enforced unevenly by the first staff member who has a bad week.
What good looks like before the first accusation
A school that handles these well has done five things before anyone is accused of anything.
The rule is written in the words a 14-year-old would use, task by task, not as a paragraph in a policy nobody has read. Every substantial task states what AI use is permitted, and the answer varies by task rather than by school. Work is set so that process evidence exists by default: a plan, a draft, version history, a two-minute conversation at the end, the three Ps of product, process and performance rather than a product alone. Staff have been trained on what an indicator is and is not, which matters more than it sounds: the JCQ list of possible indicators includes American spellings in a British script, and any leader can see how quickly that becomes a bad conversation with a bilingual student or a child who reads a lot of American fiction. And parents have been told how the school handles suspicion before it happens, which is far easier than explaining it to one angry family afterward.
None of that requires a tool. All of it requires a decision.
What to do in the next two weeks
Ask one question at your next leadership meeting: if a teacher came to you today convinced a student had used AI, who decides, on what evidence, and how would the student find out? If the answers are not the same from everyone in the room, write them down this term.
Then pick the three assessments in your school that carry the most weight and the least visibility, and add one piece of process evidence to each. That is a smaller job than a policy rewrite and it removes more risk.
Then look at last term's informal cases, the ones that never reached a process, and ask which students they involved. That is the audit nobody runs, and it is the one that tells you whether your school is fair.
Where this usually leads
Most leadership teams who work through this discover the real gap is not a cheating policy. It is that nobody has decided what the school's position on AI in student work actually is, which is why every case is argued from scratch. If that is where you are, this is the kind of work I support through AI strategy work with schools and trusts, turning a set of unresolved questions into a position your staff can apply on a Tuesday. You can also get the weekly thinking through my newsletter.
Sources and further reading
- Joint Council for Qualifications, AI Use in Assessments: Your role in protecting the integrity of qualifications, April 30, 2025.
- Ofqual, Malpractice in GCSE, AS and A level: summer 2025 exam series, official statistics, December 11, 2025.
- Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou, GPT detectors are biased against non-native English writers, 2023, published in Patterns.
- Kathryn Palmer, AI Detectors Are Out, New Assessments Are In, Inside Higher Ed, August 5, 2026.
- Lily Altavena, State AI guidance for schools skirts cheating, leaving teachers without solutions, Chalkbeat, September 2, 2026.
Dan Fitzpatrick is the founder of The AI Educator, a Forbes contributor and a bestselling author on AI in education who has advised the UK Department for Education, KHDA Dubai and the Ministry of Education in Kazakhstan. More about Dan.


