You walk into a classroom and you pay attention to everything you can—the tone, the transitions, the kid in the back row who's checked out, the question that almost landed but didn't. You document as much as you can.
And then, at the end of the process, almost none of it survives. It’s all converted into a number that ends up in a spreadsheet somewhere.
That's not a flaw in the system. That's what a score is—a purposeful simplification of something too complex to transmit whole.
Whoever receives the final score doesn't want your raw notes, and doesn’t want a nuanced, comprehensive understanding of the teacher’s practice; they want a number they can act on.
The question isn't whether scoring throws information away—it certainly does. It's whether we're honest about what's being thrown away, and whether we're doing it at the right time.
Most of the time, we’re scoring teacher practice too early.
Rubrics can be problematic things in teacher observations. It's not because rubrics are bad. It's that we keep using them like answer keys.
An answer key has the right answer. Two plus two, no ambiguity, no judgment required. Rubrics exist precisely because most of what matters in teaching isn't like that—it's too complex to reduce to a single correct response, so instead we build categories, levels, and look-fors. They’re a good and necessary tool.
But the moment you start treating a rubric like it hands you the correct answer—like there's one right way to fill every box, and your job is just to find the evidence that matches—you've turned a thoughtful instrument into a bureaucratic one. And that's where things go sideways.
This isn't just a classroom problem. Barry Schwartz and Kenneth Sharpe wrote an entire book about it—Practical Wisdom: The Right Way to Do the Right Thing—and their argument is blunt: rules and incentives are an insurance policy against disaster.
That's it. That's all they can promise you. They can keep the worst outcomes from happening, but they can't produce excellence, because excellence requires judgment, and judgment is exactly the thing rules are designed to remove.
Schwartz points directly at education as a casualty of this trade: when a school system stops trusting teachers' judgment, the fix is a teacher-proof script. These rigid systems create less room to think and more room to comply. They can prevent the worst teaching.
They also guarantee mediocre teaching, because the same script that keeps a weak teacher from doing real damage is the one preventing a great teacher from doing great teaching.
So ask yourself honestly: when you hand someone a twenty-page rubric and tell them every row needs a rating, are you building a shared language for talking about complex practice—or are you handing out an answer key and calling it professional judgment?
There's an implicit paradigm baked into a lot of "get into more classrooms" initiatives: the Olympic judge model. You observe, you score.

You do that enough times, the thinking goes, and the scores get more accurate—because you've probably caught someone's best performance somewhere in there, along with their average, and over enough reps you triangulate on the truth.
The problem is that teaching isn't a dive. A dive is over in three seconds. It's self-contained and discrete. And Olympic judges aren't interchangeable across events. A figure-skating judge can't score a high dive. They don't have the training.
Yet in education, we routinely hand one rubric to one observer and ask them to score everything from kindergarten reading groups to AP Physics using the same handful of domains, as if teaching were one uniform performance instead of dozens of different crafts bearing the same job title.
More observations help. Nobody's arguing against getting into classrooms more. But "more" doesn't solve the deeper issue, which is that a lot of what actually matters on a teacher evaluation rubric was never going to be observable in a 50-minute lesson, no matter how many times you showed up.
If the criteria include planning, or professional collaboration, or how a teacher communicates with families—you're not going to see more of that by walking through more often. You're going to see the same absence, more frequently.
This is where a lot of well-intentioned evaluation systems go wrong: they never ask what the right unit of evaluation actually is for a given criterion in a rubric. Our default unit is the lesson, because we've scheduled observations that way since forever, and it's convenient for data collection. But convenient for data collection and appropriate for judgment are two different things, and we tend to assume they're the same.
Some things really can be judged in seconds—a specific instructional technique, executed or not, done. Other things need a whole year to say anything true about. Take a criterion like "demonstrates leadership with parents and families." If I watch a lesson and see the teacher on the phone with a parent, that's not leadership — that's a phone call. Leadership is what that relationship adds up to over time, and no single lesson, however carefully observed, can show you that. Treating a 30-minute snapshot as sufficient evidence for a year-long criterion isn't rigor. It's a category error dressed up as data.
And there's a companion mistake here that I (Justin) think about a lot, which my co-author Heather Bell-Williams and I named observability bias in our book, Mapping Professional Practice: the tendency to focus on whatever is easiest to observe, rather than on what actually matters. If you're only counting what a camera or a checklist can catch, you'll systematically overweight the visible and underweight the invisible. And the invisible part—the professional judgment underneath the behavior—is usually what you actually care about.
Think about a pilot glimpsed through an open cockpit door. At any given moment, the pilot may not look busy. That doesn't mean nothing important is happening. And think about the rubrics that reward classrooms where "students are running the show"—the self-driving-vehicle version of teaching, where the teacher's visible footprint shrinks as the students' autonomy grows. That's often exactly what we want. It's also exactly the scenario where the things that made it possible—the planning, the routines built over months, the relationships—are least visible to whoever walks in that day. The less a good teacher seems to be doing, the harder it can be to give them credit for what actually got them there.
I (Justin) once watched an excellent teacher handle a strong-willed student in a way that looked, frankly, harsh. In the moment, my honest assessment was that she was being too hard on him. Afterward, she explained: they'd set specific goals together, she'd been in close contact with his family, and something he'd committed to hadn't happened. The observable, in-the-moment evidence that suggested unfairness was actually follow-through on a plan I had no way of seeing. The behavior was visible, but the judgment behind it wasn't. That gap is the whole ballgame.
Rather than a series of discrete performances that can be judged like Olympic sports, teaching is more like an iceberg: partially visible, but mostly hidden beneath the surface.

Above the surface sits visible behavior—the teacher actions we can directly observe. They’re important, because they’re what we can document and discuss in our feedback conversations. But they’re not the entirety of teacher practice.
Beneath the surface is where the heart of teaching can be found: in the thinking and professional judgment that determine a teacher’s effectiveness. Since we can’t directly see teacher judgment, we need another way to access it.
So if what we observe ourselves can’t give us the full picture of teacher practice, what can? Two things: conversation and artifacts.
Teacher-provided evidence is underused, and it's underused for a reason that doesn't hold up under scrutiny—the fear that teachers will only bring their best. Should we worry about that?
No, for two reasons. First, most people are bad liars. If someone hands you evidence, it's going to represent something real, even if it's their best foot forward—and there's real research, including work out of Harvard's Graduate School of Education, suggesting that "best foot forward" evidence actually facilitates more growth, not less.
Second, and more useful: when a teacher chooses what to bring you, they're showing you what they think good looks like. That's not a workaround to get real evidence. That's a richer, second-order kind of evidence—insight into their thinking that you'd never get any other way.
The conversation itself matters just as much, and the question you ask determines what you get back. Ask "why," and you'll usually get justification—a defense of a decision already made, because the stakes feel high and the instinct is to protect.
Ask "how," though, and something different happens. "How did you decide to group them that way?" "How does this compare to what you'd planned?" Questions like that invite explanation instead of defense, and they keep the conversation anchored to something you both actually saw, rather than devolving into a debate about whose opinion is right.
And this is where AI earns its place in the process—not by producing a rating, but by doing the unglamorous work of connecting evidence to criteria at a scale no human reviewing footage by hand could sustain. Run a lesson and its artifacts against a framework, and a good report will tell you not just where you have evidence, but where you don't—flagging "insufficient evidence to judge this criterion" instead of quietly forcing a rating you don't actually have grounds for.
Most forms don't give you permission to say "I don't know yet." A good tool should, so you can keep collecting the evidence you need.
If you want to have the greatest possible impact on teacher practice, separate evidence collection from scoring, as fully as you can. Use the rubric while you're in the room as a reminder of what to watch for—not as a form you're filling out live.
Talk to the teacher before you finalize anything. Watch for observability bias—keeping in mind that what you see is just the tip of the iceberg—and specifically hunt for the evidence of things that don't announce themselves on camera.
And when you don't have enough to judge a criterion fairly, say so, instead of manufacturing a rating to satisfy a form.
None of that gets you all the way there, though. You can gather the richest evidence in the world and score it as fairly as any process allows, and still walk away having changed nothing—because the evidence and the score were never the point. The conversation is the point. That's where we're headed next.

Part 3 of The Evidence-First Observation, a three-part series on new ways to observe teachers in the age of AI.

Part 2 of The Evidence-First Observation, a three-part series on new ways to observe teachers in the age of AI.

Part 1 of The Evidence-First Observation, a three-part series on new ways to observe teachers in the age of AI.
The instructional evidence platform for K–12 districts.
