Competence-Based Assessment in the Age of A.I.
I write about the things I'm thinking about because writing helps me understand them…
In fact, that's what I am doing right now. This piece is connected to ideas I'm developing for my current book project, Teacher Smart: How to Teach Like a Human in the Age of A.I. I don't yet have all the ideas fully arranged in my mind. It's a process. Writing gives me a way to work through that process. I can examine what I think I understand, notice where the connections are less clear, search for language precise enough to carry the ideas forward, and along the way, discover what becomes visible while I'm trying to explain it.
The understanding doesn't arrive fully formed and then make its way onto the page. Rather, I put the understanding together through writing. In some ways, I don't know what I think until I write about it.
But I also write to track my thinking—I want something I can return to. Writing creates a record of what I've learned. It preserves an understanding I may need to recover later. And because it gives my thinking a form that others can encounter, my writing becomes evidence of my understanding that I can share even as it continues to emerge. It's both an opportunity to learn and evidence of learning at the same time.
When the writing is going well, I feel empowered. I feel alive, engaged, confident, and enthusiastic. And I do mean feel… I feel it in my body. I feel a sense of ownership over what I am coming to understand. And once I recognize that I might've found something worthwhile, some gem of insight, I get a little charge of exuberance with it. I want to share it. But more than that, I want to share my thinking as clearly as possible. I become unwilling to hand over something that underrepresent
s the understanding or allows the gem to be corrupted by the insufficiency of the evidence I've produced.
When I reflect on this experience of writing for understanding, it helps me rethink what makes for a good assessment. A good assessment is both an opportunity to learn and the evidence of learning.
This idea isn’t unique to writing. Scientists develop their understanding through experimentation. Designers come to understand problems through prototyping. Musicians learn through rehearsal and performance. In meaningful human activity, the development and demonstration of competence frequently happen together.
School assessment, however, is often designed around a separation between the two. First, students learn. Then the learning stops so that assessment can begin. But that’s such a departure from how learning happens everywhere outside of traditional schooling. It’s as if the assessments are considered more rigorous if students aren’t able to learn from actually doing it and instead only provide evidence of what they know.
That separation is not incidental. It’s made possible by controlling the assessment conditions. Everyone is expected to complete the same kind of task, with access to the same resources, under the same restrictions. Variation is treated as a threat to rigor. But if understanding is naturally dynamic, contextual, and responsive to the person doing the learning, then sameness may give us cleaner comparisons while producing weaker evidence of competence.
This is where the claims about A.I. personalization become particularly interesting.
I’m cautious about many of the exuberant claims being made about A.I. personalization. Personalization can become one of those words with tremendous appeal, but it’s important to define what it means, determine what it requires, and establish specifically what value it produces. These are questions we should pursue—but embedded within the personalization argument is a foundational acknowledgment with which I can agree:
If personalization is preferred then sameness is insufficient.
If personalized learning and assessment represent an improvement over traditional methods for schooling, then we are already conceding that giving every student the same task, in the same format, under the same conditions is not necessarily the best way to measure what each student knows and can do.
That concession creates an opening for a much more consequential conversation. If sameness is insufficient, what should rigor mean?
For a long time, standardization has been treated as if it were synonymous with rigor. The sameness of the testing event supposedly guarantees its reliability and fairness. Students answer the same questions, or questions drawn from a common bank, under uniform conditions. Their performances can then be compared, sorted, ranked, and reported.
I agree with those who challenge the fairness of this manner of testing. I agree that they can marginalize students whose cultural and linguistic experiences are least represented in their assumptions and structures. But that’s not even my greatest concern: my strongest argument is that standardized tests are not valid measures of the quality and effectiveness of teaching and learning.
The problem begins with our definition of competence.
Standardized assessments are most efficient at measuring a student’s ability to recall information, reproduce an algorithm, recognize a familiar problem structure, or generate an expected answer within a narrow and controlled setting. But even before A.I. became mainstream, I have been one to claim this as insufficient evidence of competence.
Competence isn’t the reproduction of understanding under controlled conditions; it’s the capacity to preserve the essence of an understanding while adapting its use to the conditions of the moment.
Most people are mostly competent under perfect circumstances, but the real world is messy. It’s dynamic. The relevant variables don’t announce themselves. Problems don’t necessarily present themselves in the familiar form in which we learned to solve them. The stated outcome may not be the only outcome that must be considered. Conditions change, new information emerges, and other human beings bring needs and interpretations that we may not have anticipated. In other words, competence requires situational judgment.
A competent person must determine what a situation demands, decide which parts of an understanding are relevant, and adapt accordingly. They must recognize both the stated and unstated intentions of the task. They’ve got to respond to complications without losing sight of the essential understanding guiding the effort.
A.I. makes the insufficiency of our old assessments models more difficult to ignore. A student can use A.I. to reproduce information, follow recognizable steps, generate an explanation, or produce a polished final artifact—in seconds! The mere existence of the product can no longer tell us much about the human understanding behind it.
Some think the answer is to invest more energy in surveillance in order to preserve the testing structures we already have. But I say it’s time that we admit that the presence of A.I. has exposed a weakness that was there all along. We’ve relied too heavily on performances that can be reproduced without providing convincing evidence of integrated human understanding—the kind of head and heart understanding that changes you.
If the traditional products are no longer sufficient, we have to become more attentive to the performance in which it was created.
What judgments did the student make? What did the student notice? What resources did they select, and why? How did they respond when their first approach proved inadequate? What feedback did they use or reject? What changed between the first attempt and the last? What can they explain, defend, revise, and transfer into a new situation?
These aren’t secondary details surrounding the evidence. They are the evidence of human understanding.
Writing helps me understand because I’m making choices throughout the process. I’m testing relationships among ideas. I’m noticing when the language doesn’t yet match what I mean or capture the feeling I hope to convey. I’m revising not only the sentence but often, I’m interrogating my own underlying thought. If someone wants evidence of my competence as a writer and thinker, the final artifact reveals something important. But the decisions, revisions, explanations, and adaptations that produced it reveal something too.
A.I. is exposing that the strongest evidence of competence may not be a static product. It may be the learner’s movement through the work. That movement is richly human and culturally situated.
In Culturally Responsive Teaming, Lindsey Stevens and I refer to the volumes of scholarship that show how culture influences the expressions of intelligence. None of us learns to think in a vacuum. We inherit cultural fluencies within the social environments, or cognitive niches, where we learn how to human (v). Those fluencies shape the associations we make, the knowledge we recognize as relevant, the meaning we assign to events, and the ways we show what we know. This is why it’s so unwise to attempt to interpret a person’s assets outside of some understanding of their cultural context.
Standardized assessment doesn’t eliminate context. It installs an artificially controlled context and then treats that context as neutral; but students whose cultural fluencies align most closely with the language, assumptions, shared group references, and performance expectations embedded in the assessment are more likely to have their competence recognized. For others, the assessment may measure their familiarity with the testing context as much or more than it measures the competency it claims to assess.
In the book, we distinguish demographic data from psychographic data. Demographics gives us a high-level description of group characteristics, but psychographic data helps us develop person-level insight into values, interests, attitudes, motivations, behaviors, and ways of perceiving the world. Teachers have a person-level access to this knowledge through observation, relationships, conversation, and the daily experience of learning alongside students and colleagues.
Psychographic sources of information resist easy quantification, but that doesn’t make them irrelevant. Teachers work with humans, not data profiles.
This becomes particularly important when we circle back to consider competency-based assessment. Psychographic knowledge can help us design situations in which students’ competence is more likely to be visible. A student’s interests, experiences, cultural fluencies, and existing assets can provide on-ramps into a performance without determining or diminishing the intellectual expectations of the task. Put another way, the route that shows understanding can vary while the underlying competency remains coherent.
Assessment is SUPER important. I don't deny this. And reliability, validity, and replicability are necessary concerns. But I suggest that we update our thinking about what we mean by evidence. This is where the question of psychometric rigor becomes especially interesting. If students demonstrate competence through different tasks or modes of expression—and if we want schooling to be fair—we need a defensible basis for interpreting and comparing the evidence. And no, authenticity is not an excuse for vagueness nor should personalization mean that expectations dissolve into a collection of unrelated experiences.
But the idea that rigor is a function of sameness and evidence is contained in a single testing event is an outdated notion that needs to be set aside forever. We’re in the age of A.I. now, and it’s clear that our former methods for standardizing assessments are no longer valid.
A recent paper on competency-based assessment uses evidence-centered design to describe coherence among three interconnected models. The domain model defines the knowledge, skills, and attributes to be assessed. The task model specifies the situations that will elicit evidence of those competencies. The evidence model establishes how the resulting performance will be interpreted.
That framework helps us relocate rigor.
Psychometric rigor can be understood as a method achieved through coherence, intentional alignment, and structured design. The competency must be clearly defined. The task must genuinely induce the competency. The evidence must support reliable, valid, and justifiable conclusions about what a student understands and can do.
Those conditions don’t necessarily require every learner to complete precisely the same task in precisely the same way. Standardization is supposed to control for variation. Coherence makes variation interpretable.
This distinction allows us to imagine assessment experiences that are responsive without becoming incoherent. Students might work in different contexts, draw on different assets, or create different kinds of artifacts while demonstrating the same essential competency—kinda like what happens in the world outside of school. Shared performance indicators, transparent criteria, multiple sources of evidence, and opportunities for calibration can preserve the integrity of the assessment system.
In fact, I argue that allowing for variation strengthens validity. If competence is expressed differently according to experience and environment, then a single required mode of expression may obscure the very capacity we are attempting to measure. An assessment cannot make a credible claim about a student’s competence if features unrelated to the competency prevent that student from making the competence visible.
I can also make strong arguments for why we should be cautious in how we allow A.I. to manage the processes of assessment. A.I. personalization won’t solve for the issues of unfairness in assessment without our taking a careful look at the very nature of what we mean by evidence. Without that step, A.I. will merely reproduce the same reductive logic in a more technically sophisticated form. Standardized assessment can erase the person through sameness. Algorithmic personalization can flatten the person into a profile.
The goal isn’t to build a more precise mechanism for sorting students. It’s to create responsive conditions in which students can develop and demonstrate meaningful understanding.
And so I am back to thinking about the experience of writing this piece.
This writing doesn’t capture everything I think I understand about assessment. It certainly doesn’t resolve everything I want to say about competence, psychometric rigor, culture, or A.I.. This piece is a time- and space-specific representation of my understanding—something I’ll return to at some point to see where my mind was at some particular prior moment. But it’s also more than a snapshot because producing it has evolved what I understand.
I now have something I can interrogate, revise, and build upon. I have an artifact through which someone else can encounter my thinking. And I have evidence of the judgments, relationships, and emerging convictions that will inform what I compose next.
A.I. has created a genuine challenge for assessment, but it’s also given us an opportunity. It’s made it harder to mistake the production of an expected answer for evidence of human competence. It’s forced us to confront the weakness of assessments that ask students to reproduce what a machine can readily reproduce for them.
We don’t need to abandon rigor in response. We need to rescue rigor from its mistaken association with sameness.
A valid assessment gives students an opportunity to bring their knowledge, cultural fluencies, judgment, emotion, and experience into contact with a meaningful problem. The best assessments don’t interrupt learning to document it; they produce evidence while building momentum for more learning to come. The best assessments allow students to learn through the work, adapt as the conditions change, and produce evidence that makes their developing competence visible; and it sets up the teacher to leverage all they know about the learner and their craft to bring the student into possession of integrated understandings that change them forever.


