How to design assessments when students have AI
A polished final draft no longer tells you who did the thinking, and no rule at the top of the syllabus restores that. How to collect evidence that still holds.
A polished final draft no longer tells you who did the thinking, and no rule at the top of the syllabus restores that. What replaces it is a grading design. Decide which cognitive work each assignment protects, then collect evidence that the student did that work — in the header on the assignment sheet and in the independent check that goes with it.
What a finished product no longer proves
A history essay used to supply evidence about a dozen or more abilities at once: interpreting the prompt, evaluating sources, forming a thesis, organizing the reasoning, sustaining a voice to the end. That inference was never clean. Students have always had help from peers, parents, templates, and spellcheckers. What changed is that a generative tool now performs several of those operations quickly and at a decent level, so the same finished essay fits processes that share almost nothing.
One student pastes an almost unedited response produced in ninety seconds. Another spends the week developing ideas and an outline, then has AI write the prose. Another drafts alone and edits with AI for two hours. Another uses it only as a counterargument coach and writes every word. Another never opens a tool.
Call this construct migration: a task that once measured reasoning begins to measure a different combination of abilities — prompting, selection, editing, and access to a good model. The cost is invisible on the page in front of you, because all five of those students can hand in the same clean paragraph. Grade the product alone and your rank order can start tracking who had the better tool and who had help at home. It also picks up editing ability and willingness to break a rule the assignment never stated clearly. The content standard you meant to measure can end up surviving in the grade only weakly.
The repair is not more surveillance. Gather more than one kind of affirmative evidence that the student understands the work.
Paper is one element of the system, not the authorship answer
A handwritten in-class response — a blue book, a ten-minute paragraph — gives you a time-bounded independent sample and far fewer competing screens. It shows the working: the annotation, the diagram, the crossings-out. And it gives you a piece of writing you can hold against the take-home draft that arrived last week.
What paper cannot do is establish that all the thinking happened independently. A student can generate material beforehand and memorize it, copy it in by hand, or get an answer from a device the room does not see. A handwritten answer also shows very little about how sources were chosen or how the reasoning developed, since both of those happened before the pen touched the page. So the blue book earns a place in the design, and it does not close the question by itself.
Three kinds of evidence, and none carries the grade alone
You already collect product evidence, and it is the kind that stopped being conclusive on its own. The other two put back what the finished artifact stopped showing you. Together the three make an evidence triangle, and no corner of it has to carry a whole assessment by itself.
- Product evidence is the artifact you already grade: essay, project, presentation, lab report, document-based response, policy brief.
- Process evidence is the trail that produced it: notes, annotations, outline, claim–evidence chart, drafts, version history, the AI exchange, source log, revision memo, peer feedback.
- Performance evidence is what the student can do in front of you: oral defense, seminar, conference, in-class writing, an unseen source, live demonstration, delayed retrieval, a no-AI transfer task.
Weight the three per assignment. An outlining unit puts most of its weight on process and one short independent performance. A research project can leave the polished artifact where it is and move the grade toward the source log and the defense.
Secure samples are the ones taken under conditions you control: a ten-minute in-class paragraph, an oral explanation of one source, a thesis and rough outline written in the room. Collect them often enough and every student has a baseline. Their purpose is diagnostic. Knowing what a student can produce independently in ten minutes tells you how to read the polished draft, and tells you what to teach next.
Asking for a use note instead of a transcript
Don’t ask for a full chatbot transcript every time a student opens a tool. It buries you both in compliance work. Require documentation when AI materially shaped the intellectual product, and let the small uses go unreported.
Keep the disclosure short enough that students will write it. A student names the tool, the purpose, and the stage of the work, then one suggestion they accepted, one they rejected or corrected, and how they checked accuracy. Name it as required evidence on the assignment, and grade it as evidence.
The rejected suggestion is the interesting line. A student who can name what they turned down, and say why, has demonstrated judgment the essay may not show anywhere. Ask a class for their rejections and you are running a formative assessment you never had to design — plus you find the students who discover, out loud, that they rejected nothing.
Decide which operations are the point before you decide what to collect
Name the operations in the assignment you are actually teaching, before you decide what evidence to require. We call the result a cognitive work allocation map: the operations one assignment asks for, with a decision attached to each.
cognitive-work-allocation
- 1 Write a precise learning claim A learning claim names the operation: building a hierarchical outline that organizes evidence around a defensible historical claim. Writing an essay names the artifact and claims nothing.
- 2 Give every operation one of four statuses Protected, meaning students perform it alone. Coached, meaning AI may question or critique it without producing it. Permitted, where the operation is not what you are teaching this time. Or the target, where deciding what a person does and what a tool does is itself the thing being learned. the allocation decision
- 3 Attach evidence to every protected operation Protect evidence selection and the claim–evidence chart gets graded. Protect outlining and the dated outline carries real weight, with the student explaining one structural decision aloud.
- 4 Add one independent check, then explain the allocation A short unseen prompt or an in-class response, done with no AI. Then tell students what is being learned, what AI may do, what evidence they owe you, and how the work is assessed.
Most operations sort quickly once the learning claim is precise. Citation formatting is rarely the target, and forming the thesis usually is.
The same fields, in the same place, on every handout
Put the allocation on the assignment sheet itself, in a fixed block at the top, in the same shape every time. Students stop guessing at the rules once the rules sit in the same place on every handout.
Seven fields cover it:
assignment-header.txt
Learning target — what you should be able to do after this assignment.
Why it matters — why this skill is worth developing.
AI mode — the permission level for this assignment.
AI role by phase — what AI may and may not do at each stage of the work.
Required evidence — notes, outline, drafts, the AI use note, the source record, the final product.
Independent check — conference, oral defense, in-class response, or transfer task.
Criteria — what quality looks like here.
The phase field is the one that does the work a single label cannot. An assignment-wide permission level can hide the difference between a tool that may question a thesis and one that may write it. A student reading only the label cannot tell which they were given.
Filled in for a history outline, the header’s working fields read: select your sources, form your thesis, and build your outline without AI. Once the outline is approved, AI may flag unclear sentences in the draft, and may not add claims, evidence, or historical information. The independent check is a two-minute conference explaining one structural choice. The criteria are a defensible thesis, a logical hierarchy, relevant evidence, explicit reasoning, and accurate sourcing.
Departments need one vocabulary here, or the same label means four things down the hall. We use four permission levels, set out with sample policy language in The Four Modes of AI Assignments. When the question is one assignment rather than a shared vocabulary, the five tests for any proposed classroom AI use are the shorter route.
Your district may already cite an external scale instead. The AI Assessment Scale sorts assignments by how much AI is permitted; the University of Kentucky Student GAI Use Scale sorts them by the student’s intellectual role. Two-lane assessment works at the level of a whole course, and TILT covers an assignment’s purpose, task, and criteria without addressing AI. The frameworks glossary defines each and says where it stops being useful.
This design needs more pieces than one essay prompt did
The protected operation needs its own scaffold, and the independent check needs an unseen source nobody has read yet. Kindred K-12 builds and adapts those pieces out of the material you already teach, so the in-class writing prompt and the unseen-source check exist before the unit starts. It is a teacher’s tool: no student accounts, no student personal information, and the AI never chats with your students. If the question in front of you is a department-wide assessment stance rather than one assignment, how we support curriculum and instruction teams is the better starting point.
The assignment worth redesigning first is the one whose grade you would have the hardest time defending. The Kindred K-12 newsletter carries more of this kind of thing, and the other pages on AI in K-12 take up the rest of the argument.