On July 3, 2026, Ofqual's chief regulator told the sector that reformed qualifications for 16 to 18 year olds would face "far, far more scrutiny of any proposals for written coursework." Ian Bauckham's reason, reported by Schools Week, was blunt: "we cannot normalise the idea that AI-generated output is a substitute for genuine human endeavour."
Read that as a leader rather than as a teacher and something moves. The regulator has started settling the high-stakes end of assessment on your behalf. It will decide what survives as coursework and how hard the survivors are policed. Nobody is going to settle the rest: the mock, the end-of-unit test, the homework essay, the piece of work that quietly becomes a predicted grade, the December report that goes home to a parent. That part belongs to you. It is also the part that shapes what students actually learn, because students study for the thing they are about to be judged on, not for the thing they will meet in June.
So, the short answer. When a machine can produce excellent work, assessment does not collapse. It splits. Any task whose only evidence is a finished artifact made out of your sight has stopped being evidence. Everything else has to earn its place by showing either the performance or the route. The leadership job this term is not to redesign assessment from scratch. It is to sort the assessments you already run into the ones that still prove something and the ones that never really did.
Is this actually happening, or is it a moral panic?
It is happening, and the honest figures are smaller and more interesting than the headlines.
The Higher Education Policy Institute's Student Generative AI Survey 2026, published on March 12, 2026 and based on 1,054 full-time UK undergraduates surveyed in December 2025, found that 12 percent of students put AI-generated text directly into assessed work. That is up from 8 percent the year before and 3 percent the year before that. Twelve percent is not an epidemic. A rise from 3 to 12 in twenty-four months is a trend line, and trend lines are what leaders are paid to read. The same survey found 65 percent of students saying assessment had already changed noticeably in response to AI, which tells you institutions are moving, unevenly, without saying much about it.
At the regulated end, Ofqual's written evidence to the Education Committee in April 2026 gives the sharper number. Across all forms of malpractice in 2025, penalties included 1,125 cases in which a student lost an entire GCSE or A level and nearly 2,000 in which marks were deducted. Those totals cover every kind of malpractice, not AI alone, and Ofqual does not break out an AI figure. What the regulator does say plainly is that undisclosed "use of AI in coursework undermines that intended learning experience," because the marks stop describing the student.
Hold those two facts side by side. Most students are not submitting machine-written work. Enough are that the marks no longer mean exactly what the mark scheme says they mean. That is a measurement problem, not a discipline problem, and it needs a leader rather than a detective.
The Claim Test: three questions to ask about any assessment
The Claim Test is three questions I ask about any piece of assessment before deciding what to do about AI: What claim does this make about this student? Would that claim still be true if a machine did the work? And if it would not, what would I need to see instead? An assessment that survives all three is still evidence. One that fails the second question and has no answer to the third is not an assessment. It is a collection exercise.
This is my suggested way of thinking about it, not a validated instrument, and it grows out of the model I have argued for in my books on AI in education: judge the product, the process and the performance, not the product alone.
The first question does most of the work, and it is the one nobody asks. Take a Year 10 history essay on the causes of the First World War. What claim does the mark make? Almost every teacher says something like "this student can construct a historical argument." Now ask the second question. If the essay was drafted by a model at eleven at night, is that claim still true? No. And the third question is where the conversation turns from anxiety to design: what would you have to see to make the claim again? A ten-minute conversation about the argument. The paragraph written in the room. The student's account of which suggestions they took and which they threw out.
Three answers, three kinds of assessment.
| Kind | The claim holds because | What you keep |
|---|---|---|
| Witnessed | You saw it happen | Live writing, vivas, seminars, demonstrations, performance, practical work |
| Traced | You can see the route, not just the destination | Drafts, decision logs, the student's account of what they accepted and rejected, oral defense of choices |
| Retired | It does not hold, and cannot be rescued | Stop grading it. Keep the task if it teaches; drop the mark |
Retiring things is the hard part, and the part leaders must authorize. A teacher cannot unilaterally stop grading a task that feeds the department's tracking spreadsheet. That decision has a name on it, and the name should be yours.
What I See in Practice
Across the leadership teams I work with, the mistake I see most often is treating this as an integrity question when it is a curriculum question wearing a disguise.
The conversation starts with detection software and ends with a policy paragraph about consequences, and everyone leaves the room having agreed something. Six weeks later the same essays are set, the same marks go into the same spreadsheet, and the same teachers privately admit they are not sure whose work they are reading. Nothing changed, because nothing in the assessment calendar changed. The policy governed students. The calendar governs the school.
The teams that make progress do something less impressive and more useful. They take one subject, one year group, one term, and they print the list of every assessment that generates a mark. Usually there are more than anyone expected, often twenty or thirty for a single cohort. Then they run the three questions down the list. What almost always emerges is that a small number of assessments were carrying the real weight, a large number existed to feed data collection, and nobody had ever separated the two. The AI question forced a look at the calendar that was overdue by about a decade.
That is the quiet opportunity here, and it is why I keep telling leaders not to grieve. You are not losing assessment. You are losing the pretense that an unsupervised finished product ever told you much on its own.
Can we not just detect it?
Not reliably enough to hang a decision on, and the newest evidence says so more precisely than the old debate did.
In a study published in the International Journal for Educational Integrity on June 29, 2026, Marijke Van Vlasselaer, Filip Van Droogenbroeck and Bram Spruyt tested four detection tools against 160 constructed documents and 1,163 real master's theses. The results were lopsided. One tool, Pangram, performed strongly, including on hybrid and humanized text. Two others, GPTZero and Copyleaks, scored zero on papers where AI passages were mixed into human writing, which is exactly how students actually work. Turnitin missed every fully AI-generated document in the constructed set. Encouragingly, false positives were rare across all four, an improvement on earlier studies. The authors' conclusion is the line to take to your senior team: detection tools "should not be used as sole evidence in high-stakes decision-making."
Read that carefully, because it cuts both ways. Detection has improved and is worth having as one signal. It cannot carry an accusation on its own, and a school that builds its response on a percentage score in a dashboard has handed a judgment about a child to a piece of software. The same logic applies to marking. Ofqual's position is that AI "cannot be used as a sole marker," and I would apply that standard to your internal assessment too, for the reasons I set out in which decisions AI should never make in your school.
What good looks like elsewhere
The direction of travel is toward evidence gathered over time rather than a single sitting.
In a white paper published on July 23, 2026, the Stanford Accelerator for Learning and ETS reported the conclusions of a convening of more than 100 education, research and policy leaders. Their recommendation was to move from isolated testing events toward "continuous systems of evidence that capture growth over time and across contexts," using portfolios, conversation-based assessment, formative feedback and authentic performance tasks, with human oversight kept firmly in place.
That is the same instinct as the Claim Test, arrived at from a different direction. It is also, for most schools, a five-year program and not a September one. Which is why the sort matters more than the vision. You cannot build continuous systems of evidence on top of an assessment calendar nobody has audited.
What to do this term
Pick one subject and one year group and sort their assessments before half term.
Print every task that produces a mark for that cohort this term. Against each one, write the claim it makes about the student in a single sentence, in plain words, not in mark scheme language. Cross out every claim that would still read as true if a machine had done the work. For each survivor, write W for witnessed or T for traced, and name the one thing a teacher would have to see. Everything left over is retired: keep the task if it teaches something, drop the mark, and tell the department you have dropped it, in writing, so nobody quietly reinstates it in the tracking system.
Then do the part most schools skip. Rewrite what a teacher says to students about AI on a task-by-task basis, rather than as a single school-wide rule. "You may use AI to check your argument and you must show me the exchange" and "this one is written in the room, on paper, in forty minutes" are both defensible. A blanket ban across every task is not, and I set out the test for when a restriction is a decision rather than a mood in how to make a restriction decision you can defend.
One more thing, for the leaders whose instinct is to wait for the regulator. Bauckham is settling the qualification. He is not settling your mocks, your reports or your predicted grades, and he never will. If you are still working out what the classroom itself is now for, that argument sits in what schools should be preparing children for now.
The leadership question
Ask your senior team one thing at the next meeting: of every mark we gave a student last term, how many would we still stand behind if we learned a machine had written the work?
If the honest answer is uncomfortable, the problem is not your students. It is that you have been collecting artifacts and calling them evidence. That is fixable, this term, with a list and a pen.
If your leadership team is working through what to keep, what to redesign and what to retire, this is the kind of work I support through AI strategy sessions and advisory work with schools and trusts.
Sources and further reading
- Ofqual, written evidence to the House of Commons Education Committee (AIE0204), April 2026. committees.parliament.uk
- Schools Week, "More 'scrutiny' of coursework plans to protect exams from AI," July 3, 2026. schoolsweek.co.uk
- Higher Education Policy Institute, Student Generative AI Survey 2026, March 12, 2026. hepi.ac.uk
- Marijke Van Vlasselaer, Filip Van Droogenbroeck and Bram Spruyt, "Who wrote this? Evaluating the reliability of AI detection tools in higher education," International Journal for Educational Integrity, June 29, 2026. link.springer.com
- Stanford Accelerator for Learning and ETS, "Responsible Assessment in the AI Era: Key Insights from a Future-Focused Convening," July 23, 2026. news.stanford.edu
Dan Fitzpatrick is the founder of The AI Educator, a former secondary teacher, assistant headteacher and Director of Digital Strategy in further education, and the author of bestselling books on AI in education. More about Dan.


