Assess Learning When AI Is Available
Evaluate process, explanation, and judgment so assessment remains meaningful when learners can use generative tools.
Learning objectives
- Align assessment with the actual learning objective
- Ask learners to explain and critique AI-assisted work
- Set clear, proportionate AI-use expectations
ToolDix original visual
Frame
Name the outcome and constraints.
Build
Try one bounded workflow.
Review
Keep evidence, revise, and share.
The take-home essay was always a proxy. It measured writing to infer thinking, and the inference held because producing fluent prose required doing the thinking first. Generative tools broke that link, which is why the essay now feels unreliable — not because students changed, but because the proxy stopped proxying.
The productive response is not detection. It is fixing the proxy: assess the thing you actually care about, directly.
Detection is the wrong instrument
Before designing anything, it is worth being clear about why the obvious answer fails.
- Unreliable in both directions
- False positives fall on non-native writers
- Unappealable and corrosive to trust
- Teaches evasion, not judgement
- States what is permitted per task
- Asks for a short usage note
- Makes the conversation about method
- Models the professional norm
AI-text detectors are unreliable in both directions, and their errors are not randomly distributed. False positives fall disproportionately on non-native English writers, whose prose tends toward the regular patterns detectors flag, and on students who write in a plain style. A detector output is a probability, not evidence, and it is close to unappealable — the student cannot prove a negative, and you cannot show your working. Several institutions have disabled these tools for exactly this reason.
The alternative is not surveillance but declaration. State per assignment what assistance is permitted, ask for a short usage note where it is, and make the conversation about method rather than suspicion. This also happens to model the professional norm: in most workplaces AI use is expected and disclosed, not hidden.
A workable policy fits in one sentence per task. "You may use AI to generate practice questions and check your grammar; the analysis and the source selection must be yours, and attach a two-line note on what you used." Students follow rules they can understand.
Assess the capability you actually value
Ask what the assignment is for, then check whether the artifact still measures it.
If the goal is synthesis, the essay is no longer sufficient but the underlying skill is intact — so assess it through source choices, comparison of alternatives, and the defence of a conclusion. Ask for the sources considered and rejected, with reasons. That is where synthesis actually lives, and it is not what a general-purpose model produces unprompted.
If the goal is writing mechanics, you need a setting where the practice is visible: in-class writing, drafts with revision history, or a supervised component. This is a real constraint, and pretending otherwise is how a course ends up certifying something it did not teach.
If the goal is domain judgement, the strongest instrument is a flawed artifact. Supply an AI-generated answer containing a fabricated citation and a plausible but incorrect inference, and ask the student to find and explain both. Someone who understands the material finds them quickly. Someone who does not, cannot — and no tool helps, because they would have to know what is wrong to ask about it.
Reweight the rubric
A rubric is a statement of what you will pay for. If the fluent final answer carries most of the marks, the rubric now largely measures tool access.
Move weight down that table. Source selection quality still requires judgement, especially if you ask for the rejected options. Accuracy of critique discriminates strongly and grades quickly. Evidence of revision rewards the loop rather than the artifact, and it is visible in document history without any extra submission.
The bottom row — defending the work under questions — is the most reliable signal available and the most expensive to run. Two minutes per student, three questions, is enough to separate understanding from retrieval with high confidence. Because it is expensive, sample it: announce that a random subset will be asked to defend their submission. The announcement does most of the work, and the sampling makes it affordable.
One caution: oral defence advantages confident speakers and disadvantages anxious ones, so weight it as a check rather than the primary grade, and offer a written equivalent.
Practice: revise one assignment
Take an existing assignment and make four changes.
Add a process checkpoint. A plan, an outline, or an annotated source list submitted before the main work. Five minutes to write, and it establishes a baseline the final submission has to be consistent with.
Add an explanation prompt. "Explain why you chose this framing over the two alternatives you considered." Short, specific, and hard to answer without having considered them.
Add a critique component. Supply a flawed answer and ask for the errors. Reuse the same flawed answer across the cohort so grading is fast and comparison is easy.
Reweight the rubric so the polished artifact is at most a third of the marks.
Then state the AI policy for the task in one sentence at the top of the brief, and say what should be disclosed. Run it once and note the grading time — done well, process assessment is usually faster to grade than essays, because a wrong critique is obvious in ten seconds where a mediocre essay takes ten minutes.
Proportionality
The right level of restriction depends on learner age, subject, accessibility needs, institutional policy, and whether the task teaches a foundational skill or an applied one. A first-year student learning to construct an argument and a final-year student applying one need different rules, and applying the same policy to both is a mistake in one direction or the other.
Two constraints deserve particular attention when assessment is involved. First, if AI use is permitted but only paid tiers produce good output, the assessment partly measures spending. Second, for some learners generative assistance is a documented accommodation; an assessment rule that removes it is removing support, not restoring rigour. Both argue for assessing process, which is far less sensitive to which tool anyone had.
Common mistakes
Acting on detector scores. They are not evidence, the errors are biased, and using them damages the trust that makes disclosure policies work.
One rule for the whole course. Different tasks assess different skills and need different lines.
Adding process requirements without removing anything. If you add a plan, a critique, and a defence while keeping the full essay, you have tripled the workload for everyone. Shrink the artifact as you add the evidence.
Sources and license context
These references informed the lesson. ToolDix adds its own explanation, workflow, and practice rather than reproducing source material. Every link below leaves ToolDix and opens the publisher's own site in a new tab.
- UNESCO AI Competency Framework for Students (opens unesco.org in a new tab)External · unesco.org (Official framework; link only)
- European Commission Ethical AI Guidelines for Educators (opens education.ec.europa.eu in a new tab)External · education.ec.europa.eu (Official guidance; link only)
Keep going
Read these next on ToolDix.
Original lessons that build on what you just read.