How to design assessments when students have AI
A polished final draft no longer tells you who did the thinking, and no rule at the top of the syllabus restores that. How to collect evidence that still holds.
A polished final draft no longer tells you who did the thinking, and no rule at the top of the syllabus restores that. What replaces it is a grading design. Decide which cognitive work each assignment protects, then collect evidence that the student did that work — in the header on the assignment sheet and in the independent check that goes with it.
What a finished product no longer proves
A history essay used to supply evidence about a dozen or more abilities at once: interpreting the prompt, evaluating sources, forming a thesis, organizing the reasoning, sustaining a voice to the end. That inference was never clean. Students have always had help from peers, parents, templates, and spellcheckers. What changed is that a generative tool now performs several of those operations quickly and at a decent level, so the same finished essay fits processes that share almost nothing.
One student pastes an almost unedited response produced in ninety seconds. Another spends the week developing ideas and an outline, then has AI write the prose. Another drafts alone and edits with AI for two hours. Another never opens a tool.
Call this construct migration: a task that once measured reasoning begins to measure a different combination of abilities — prompting, selection, editing, and access to a good model. The cost is invisible on the page in front of you, because all four of those students can hand in the same clean paragraph. Grade the product alone and your rank order can start tracking who had the better tool and who had help at home. It picks up editing ability, and willingness to break a rule the assignment never stated clearly. The content standard you meant to measure then survives in the grade only weakly.
The repair is not more surveillance. Gather more than one kind of affirmative evidence that the student understands the work.
What paper and blue books can prove
A handwritten in-class response gives you a time-bounded independent sample and far fewer competing screens. It shows the working: the annotation, the diagram, the crossings-out. And it gives you a piece of writing you can hold against the take-home draft that arrived last week.
What paper cannot do is establish that all the thinking happened independently. A student can generate material beforehand and memorize it, copy it in by hand, or get an answer from a device the room does not see. A handwritten answer also shows very little about how sources were chosen or how the reasoning developed, since both of those happened before the pen touched the page. Paper belongs in the design as one element of a larger evidence system, sampling what a student can do unaided in an hour.
Collecting product, process, and performance evidence
You already collect product evidence: the essay, the project, the document-based response, the policy brief. Process evidence is the trail that produced it — annotated sources, a dated outline, a claim–evidence chart, drafts and version history, a revision memo. Performance evidence is what the student can do in front of you: an oral defense, a seminar, an in-class response to an unseen source, a delayed retrieval check, a transfer task with no AI available. Between them they make an evidence triangle, and no corner of it has to carry a whole assessment by itself.
Weight the three per assignment rather than uniformly. An outlining unit puts most of its weight on process and one short independent performance. A research project can leave the polished artifact where it is and move the grade toward the source log and the defense.
A ten-minute in-class paragraph, an oral explanation of one source, a thesis and rough outline written in the room: collect these secure samples often enough and every student has a baseline. Their purpose is diagnosis before it is anything else. Knowing what a student can produce independently in ten minutes tells you how to read the polished draft, and tells you what to teach next.
Asking for a use note instead of a transcript
Don’t ask for a full chatbot transcript every time a student opens a tool. It buries you both, and it teaches nobody anything. Require documentation when AI materially shaped the intellectual product, and let the small uses go unreported.
Keep the disclosure short enough that students will write it. Six lines cover it: which tool, for what purpose, at which stage of the work. Then one suggestion the student accepted, one they rejected or corrected, and how they checked accuracy. Name it as required evidence on the assignment, and grade it as evidence.
The rejected suggestion is the interesting line. A student who can name what they turned down, and say why, has demonstrated judgment the essay may not show anywhere. Ask a class for their rejections and you are running a formative assessment you never had to design — plus you find the students who discover, out loud, that they rejected nothing.
Mapping the cognitive work in an assignment
Before you can decide what evidence to require, decide which operations in the assignment are the point. We call the result a cognitive work allocation map: the operations one assignment asks for, with a decision attached to each. Four steps, per assignment, in order.
Start with a precise learning claim. “Students will write an essay” gives you nothing to protect. “Students can build a hierarchical outline that organizes evidence around a defensible historical claim” tells you which operation is the target and which ones are incidental.
Break the assignment into its actual operations and give each one a status. Some are protected human work, which students perform independently. Some are human-led and AI-coached: students do the operation, and AI may question or critique it without producing it. Some are AI-permitted support, because that operation isn’t what you’re teaching this time. In a few assignments the collaboration itself is the target, and allocating work between a person and a tool is the thing being learned. Citation formatting is rarely the target, and forming the thesis usually is.
Attach visible evidence to every protected operation, and add one independent check. If evidence selection is protected, the claim–evidence chart gets submitted and graded. If outlining is protected, the dated outline carries real weight, and the student explains one structural decision in a two-minute conference. Then add a related task done without AI — a short unseen prompt, an in-class response — so the learning claim has assurance behind it.
Explain the allocation to students. Say what is being learned, why it is worth learning, which operations they have to practice, what AI may do, what evidence they owe you, and how the assignment will be assessed.
Writing the assignment header students read
Put the allocation on the assignment sheet itself, in a fixed block at the top, in the same shape every time. Students stop guessing at the rules once the rules sit in the same place on every handout. And nobody re-litigates them in the last week before a deadline.
The header holds seven fields:
- Learning target: what the student should be able to do afterwards.
- Why it matters: why this skill is worth developing.
- AI mode: the permission level for this assignment.
- AI role by phase: what AI may and may not do at each stage of the work, which is where a single assignment-wide label usually turns out to be too coarse.
- Required evidence: which artifacts get submitted, from notes and drafts to the use note and the source record.
- Independent check: conference, oral defense, in-class response, or transfer task.
- Criteria: what quality looks like here.
Filled in for a history outline, the AI role by phase field reads: select your sources, form your thesis, and build your outline without AI. Once the outline is approved, AI may flag unclear sentences in the draft, and may not add claims, evidence, or historical information. The independent check is a two-minute conference explaining one structural choice. The criteria are a defensible thesis, a logical hierarchy, relevant evidence, explicit reasoning, and accurate sourcing.
The AI mode field is the one that needs shared vocabulary across a department. The four permission levels we use, with sample policy language for each, are set out in The Four Modes of AI Assignments. For a single decision rather than a department vocabulary, the five tests for any proposed classroom AI use are the shorter route. Your district may already cite an external scale instead.
The AI Assessment Scale, the University of Kentucky scale, two-lane assessment, and TILT each cut at the same field from a different angle. The frameworks glossary defines each and says where it stops being useful.
Building assignments around the evidence you need
This design asks you to make more pieces than a single essay prompt did. The protected operation needs its own scaffold, and the independent check needs an unseen source nobody has read yet. Kindred K-12 builds and adapts those pieces out of the material you already teach, so the in-class writing prompt and the unseen-source check exist before the unit starts. It is a teacher’s tool: no student accounts, no student personal information, and the AI never chats with your students. If the question in front of you is a department-wide assessment stance rather than one assignment, how we support curriculum and instruction teams is the better starting point.
Rewriting assessments across a department is months of work nobody has time for. We help with that. Talk to us.