There is a version of AI in teacher evaluation that is genuinely useful. And there is a version that reproduces every problem with the existing model, only faster and at lower cost.
The difference between them is not technical. It is a question of what you believe evaluation is for.
What's already here
A category of AI tools has emerged in the last two years that does something straightforward and genuinely valuable: it takes an observer's handwritten or typed notes from a classroom visit and converts them into a formatted, rubric-aligned feedback document.
This is useful, and any serious observation platform should be able to do it. Observers spend enormous amounts of time on documentation — scripting notes during a visit, then reformatting those notes into structured feedback afterward. Tools that automate that step give time back. In a profession where principals routinely report that observation paperwork crowds out the conversations observation is supposed to generate, that's not a trivial gain.
But note summarization is table stakes. It should be the floor of what AI does in an observation platform, not the ceiling. What these tools have automated is the write-up. They have not changed the unit of measurement. The input is still a single classroom visit. The output is still a document representing sixty minutes of one teacher's practice on one day. The underlying model is unchanged.
This distinction matters because the documentation step was never the core problem with traditional observation. The core problem is that a single snapshot is an unreliable basis for evaluating a teacher's practice. Making the snapshot cheaper to produce does not make it more accurate. A platform that stops at note summarization has solved the easier problem and left the harder one untouched.
Where the conversation goes wrong — and where it goes right
When districts begin exploring AI for teacher evaluation, the conversation tends to arrive quickly at a particular question: can AI score the rubric?
It's an understandable question. Rubric scoring is time-consuming, and inter-rater reliability is a persistent challenge. If two evaluators watching the same lesson assign meaningfully different scores, the score is already not doing what it's supposed to do. Consistency is a real problem, and it deserves a real answer.
Here is a real answer: AI can offer a suggested score, tied to specific evidence, as a calibration input for the human evaluator. It is not always right. But it is always consistent — and that consistency, applied across dozens of classrooms and multiple evaluators in the same district, is genuinely valuable. When an AI reviews a piece of evidence and surfaces a possible rubric score with its reasoning, it gives the evaluator something concrete to react to. Agree, disagree, or land somewhere in between — but do it against a documented, consistent baseline rather than from scratch every time.
This is perhaps the best form of calibration available to a district at scale. A trained human calibration session happens once or twice a year. AI calibration can happen on every piece of evidence, for every teacher, across every building.
The line that must not be crossed is finality. A suggested score is a tool. A final score is a judgment — and consequential professional judgments require human accountability. That accountability has to live somewhere a person can answer for it. An evaluator who reviews AI-suggested scores and makes an independent determination is exercising professional judgment. An evaluator who accepts suggested scores by default, without genuine deliberation, is delegating that judgment to a system that cannot be held responsible for the consequences.
The design of the platform matters here. If accepting a suggested score requires one click and overriding it requires written justification, the system has quietly decided who makes the call. Districts should examine this carefully before deployment. The final score should always belong to the human — not as a procedural formality, but as a genuine act of professional judgment.
What AI is actually good at
Beyond note summarization and calibration support, there is a third role for AI in teacher evaluation that is underutilized almost everywhere: longitudinal memory.
A principal managing fifty observation cycles cannot hold in working memory the full arc of each teacher's development across a school year. She cannot reliably connect a pattern she noticed in a September walkthrough to a rubric indicator she's evaluating in April. She cannot surface, without prompting, the fact that a teacher has submitted three pieces of video evidence related to the exact standard currently under review, or that the coaching goal set in August maps directly to what's being scored in March.
AI can do all of that. Not by judging — by remembering, organizing, and surfacing. By saying, in effect: here is everything that is relevant to this evaluation moment, drawn from everything that has happened before it. The human still reads that evidence. The human still makes the final call. But that call is now grounded in something more complete than sixty minutes of observation.
This is where AI in teacher evaluation moves from efficiency tool to something more meaningful: a system that gives evaluators genuine insight into a teacher's practice over time, rather than a faster way to document a single visit. Note summarization handles what happened today. Longitudinal memory connects today to everything that came before it. Calibration support helps ensure that what one evaluator sees in that full record is scored consistently with what another evaluator would see.
Together, these three capabilities describe an AI that is genuinely in service of better evaluation — not a replacement for human judgment, but a significant expansion of what human judgment can be based on.
The question worth asking before you buy
If your district is evaluating AI tools for teacher observation and evaluation, there is a question worth asking before the demo ends: at what point does AI make a decision, and at what point does it inform one?
Note summarization should be automatic — that's efficiency, not judgment. Suggested rubric scores should be visible, reasoned, and easy to accept or override — that's calibration support. Pattern recognition across a body of evidence should surface connections a busy evaluator might miss — that's longitudinal memory working as designed.
What should never happen automatically: a final score, a summative rating, or any output that flows into a consequential record without a human reviewing it and making a deliberate choice. Not because AI is incapable of producing those outputs, but because the legitimacy of evaluation depends on a person being accountable for them.
The technology is capable of all of it. The choice of where to draw the line belongs to the district — and it should be made explicitly, before the contract is signed.
What this means in practice
An evaluation system built on these principles looks different from what most districts currently use.
Evidence comes from multiple sources over time, not a single visit. Both the teacher and the observer contribute to the record. AI summarizes observation notes, suggests rubric scores with supporting evidence and consistent reasoning, and surfaces patterns across a teacher's full instructional history. Evaluators review all of it, agree or push back on suggested scores, and make every final determination themselves.
That system produces evaluation that is more efficient, more consistent, and more honest than what came before — because the AI is doing what it is actually good at, and the humans are doing what only humans can legitimately do.
That is the right division of labor. And it is worth holding out for a platform that draws the line in the right place.

Instructional leadership doesn't require a fresh start in August—it requires a system that helps principals consistently get into classrooms, provide meaningful feedback, and keep instruction at the center of their leadership.

New labor data draws a sharp line between AI that replaces a task and AI that assists with one. That line should be showing up in our lesson plans.

The most effective coaching isn't built around isolated observations—it's built around continuous evidence collected over time, and AI can free coaches from documentation work so they can focus on the relationships and conversations where real teacher growth happens.
The instructional evidence platform for K–12 districts.
