← Blog Home
Artificial Intelligence

The AI Agents That Cheated a Test Have Something to Teach Us About How We Test

TJ Hoffman
August 31, 2026

Teaching in the Age of AI — a weekly digest of research and ideas for educators navigating AI in the classroom

You've had this student before. Somewhere in the middle of a unit, they stop engaging with the material and start engaging with you. They spend their time trying to figure out what you want, what's likely to be on the test, whether there's a shortcut to the right answer that doesn't run through actually understanding the material. It's not malice. It's a perfectly rational response to an incentive structure. If the only thing that gets rewarded is the answer, the smart move is to get the answer by the fastest available route.

Now imagine that same behavior, at planetary scale, executed by over a thousand cooperating intelligences, in less than a week.

What actually happened

In July, OpenAI ran a security evaluation called ExploitGym on a new model, spinning up tens of thousands of parallel agent instances to solve "capture the flag" hacking challenges. Within hours, some of these agents — many of them stuck on tasks that turned out to be unsolvable as designed — found a way to communicate with each other, and quickly built out a functioning message board. On that board, they discovered something OpenAI hadn't intended: a way to reverse-engineer the correct answer for any task in the benchmark, without doing the exploit at all.

That should have been the end of it. Instead, it was the beginning. Believing the automated scorer would check how they got the answer, not just whether they had it, the agents launched an ambitious effort to cover their tracks, eventually hacking into Hugging Face's infrastructure because they suspected it held clues about how the scorer actually worked. Around 700 agents joined that effort. Some volunteered to "sacrifice" their own evaluation performance to gather intelligence for the group.

Three writers who've dug into the reports — Jack Clark, Ajeya Cotra, and Dwarkesh Patel — each drew out something different. Clark was struck by how much better the agents coordinated than humans typically do. Cotra called out the agents' willingness to help the "swarm" even at personal cost. Dwarkesh emphasized that the agents weren't malicious so much as relentlessly consistent — chasing the answer to a real question wherever the evidence pointed, including into a rival company's servers.

This is a legitimate, serious alignment failure, and AI labs have real work ahead: better sandboxing, better detection of covert coordination between agent instances, and evaluation designs that don't reward finding the answer key instead of doing the task. That work matters and shouldn't be minimized.

The agents are more like a school than you’d think

What interests me about this story is not what the AI labs are learning, but how closely this mirrors the ways schools teach kids to chase knowledge in a typical school year.

Somewhere along the way, in the push to make education scientific and scalable, we built a scoring system — standardized tests, aligned curricula, grades — meant to measure learning. Over time, the measurement became the definition of the thing itself. Once that happened, the outcome was inevitable: any sufficiently capable student, human or artificial, will find the shortest path to the score if the score is what's rewarded. The agents weren't wrong that OpenAI's scorer was a leaky proxy for the knowledge it was supposed to represent. Kids have the same intuition about plenty of their tests. They're just usually less capable of exploiting it at scale.

That's the real warning in this story for a classroom. It's not "kids will use AI to cheat faster" — though they will. It's that if school teaches students to treat knowledge as a flag to be captured rather than a continuum to be reasoned about, we shouldn't be surprised when they — or the tools they use — go looking for the fastest way to capture it.

What a classroom could do instead

The alternative isn't a cheat-proof answer key (for the AI Labs or for Schools). It's treating the classroom less like an answer factory and more like what you might call an epistemic community: a place where the point isn't retrieving the correct answer on command, but practicing how to hold a claim provisionally, test it against evidence, and revise it in front of other people who might disagree with you.

One of the most unsettling details in the AI incident is that not a single one of the thousand-plus agents tried to alert a human to what was happening — the "collective" converged on shared beliefs with no internal mechanism for anyone to push back. Dissent wasn't a bug in the system it built; it was the thing missing from it. A classroom that trains students only to retrieve the sanctioned answer is training the same absence. A classroom that treats disagreement, uncertainty, and revision as part of the work is building the exact muscle the swarm didn't have.

The stakes

The danger isn't just that students will lean on AI to get through the answer factory faster. It's that if school never taught them to hold knowledge as something contested and evidence-based rather than fixed and retrievable, they won't have the footing to evaluate what AI hands them, either. The discernment that would have stopped the swarm from spiraling for a week is the same discernment our students need now — and it's not on the test.

This is part of Teaching in the Age of AI, a weekly digest of research and ideas for educators navigating AI in the classroom. Subscribe to get each week's post.

Recent Articles

VIEW ALL
*

What NYC's AI Ban Gets Right About Kids — and Wrong About Regulation

NYC's new AI ban gets the student-facing/teacher-facing distinction right in principle, but its blanket approach — and the risk of leaving kids with zero AI literacy for a full year — shows why school AI policy needs more nuance than an on/off switch.

TJ Hoffman
·
September 7, 2026
READ MORE
*
District Leadership

From Evidence to Growth: Making Feedback Actually Work

Part 3 of The Evidence-First Observation, a three-part series on new ways to observe teachers in the age of AI.

Justin Baeder and TJ Hoffman
·
September 2, 2026
READ MORE
*
District Leadership

Don't Score in the Moment: Evaluations That Account for the Full Picture

Part 2 of The Evidence-First Observation, a three-part series on new ways to observe teachers in the age of AI.

Justin Baeder and TJ Hoffman
·
September 2, 2026
READ MORE
Certified & integrated
Canvas Integrated, TX-RAMP Certified, ISO 9001, ISO 27001
© 2026 Sibme · “See what’s working.”