Skip to main content
AI Education

Assess Learning When AI Is Available

Evaluate process, explanation, and judgment so assessment remains meaningful when learners can use generative tools.

Intermediate15 minBy ToolDix Editorial

Learning objectives

  • Align assessment with the actual learning objective
  • Ask learners to explain and critique AI-assisted work
  • Set clear, proportionate AI-use expectations

ToolDix original visual

AI Education practice loop
1

Frame

Name the outcome and constraints.

2

Build

Try one bounded workflow.

3

Review

Keep evidence, revise, and share.

The take-home essay was always a proxy. It measured writing to infer thinking, and the inference held because producing fluent prose required doing the thinking first. Generative tools broke that link, which is why the essay now feels unreliable — not because students changed, but because the proxy stopped proxying.

The productive response is not detection. It is fixing the proxy: assess the thing you actually care about, directly.

Detection is the wrong instrument

Before designing anything, it is worth being clear about why the obvious answer fails.

ToolDix original diagram
Replace guessing with a stated policy
Detection-based
  • Unreliable in both directions
  • False positives fall on non-native writers
  • Unappealable and corrosive to trust
  • Teaches evasion, not judgement
Disclosure-based
  • States what is permitted per task
  • Asks for a short usage note
  • Makes the conversation about method
  • Models the professional norm
Detector confidence scores are not evidence, and acting on them harms specific groups of students disproportionately. A stated policy is enforceable; a probability is not.

AI-text detectors are unreliable in both directions, and their errors are not randomly distributed. False positives fall disproportionately on non-native English writers, whose prose tends toward the regular patterns detectors flag, and on students who write in a plain style. A detector output is a probability, not evidence, and it is close to unappealable — the student cannot prove a negative, and you cannot show your working. Several institutions have disabled these tools for exactly this reason.

The alternative is not surveillance but declaration. State per assignment what assistance is permitted, ask for a short usage note where it is, and make the conversation about method rather than suspicion. This also happens to model the professional norm: in most workplaces AI use is expected and disclosed, not hidden.

A workable policy fits in one sentence per task. "You may use AI to generate practice questions and check your grammar; the analysis and the source selection must be yours, and attach a two-line note on what you used." Students follow rules they can understand.

Assess the capability you actually value

Ask what the assignment is for, then check whether the artifact still measures it.

If the goal is synthesis, the essay is no longer sufficient but the underlying skill is intact — so assess it through source choices, comparison of alternatives, and the defence of a conclusion. Ask for the sources considered and rejected, with reasons. That is where synthesis actually lives, and it is not what a general-purpose model produces unprompted.

If the goal is writing mechanics, you need a setting where the practice is visible: in-class writing, drafts with revision history, or a supervised component. This is a real constraint, and pretending otherwise is how a course ends up certifying something it did not teach.

If the goal is domain judgement, the strongest instrument is a flawed artifact. Supply an AI-generated answer containing a fabricated citation and a plausible but incorrect inference, and ask the student to find and explain both. Someone who understands the material finds them quickly. Someone who does not, cannot — and no tool helps, because they would have to know what is wrong to ask about it.

Reweight the rubric

A rubric is a statement of what you will pay for. If the fluent final answer carries most of the marks, the rubric now largely measures tool access.

ToolDix original diagram
What the rubric rewards is what you get
Fluent final answer
Now cheap to produce. Weight it low, or the rubric measures tool access.
Quality of source choices
Still requires judgement. Ask for the rejected options too.
Accuracy of critique
Can they find the flaw in a plausible wrong answer? Strongly discriminating.
Evidence of revision
What changed, and why. Rewards the loop rather than the artifact.
Defence under questions
The most reliable signal, and the most expensive. Sample it.
A rubric weighted toward the top row was already a weak instrument before generative tools. It is now close to measuring nothing.

Move weight down that table. Source selection quality still requires judgement, especially if you ask for the rejected options. Accuracy of critique discriminates strongly and grades quickly. Evidence of revision rewards the loop rather than the artifact, and it is visible in document history without any extra submission.

The bottom row — defending the work under questions — is the most reliable signal available and the most expensive to run. Two minutes per student, three questions, is enough to separate understanding from retrieval with high confidence. Because it is expensive, sample it: announce that a random subset will be asked to defend their submission. The announcement does most of the work, and the sampling makes it affordable.

One caution: oral defence advantages confident speakers and disadvantages anxious ones, so weight it as a check rather than the primary grade, and offer a written equivalent.

Practice: revise one assignment

Take an existing assignment and make four changes.

Add a process checkpoint. A plan, an outline, or an annotated source list submitted before the main work. Five minutes to write, and it establishes a baseline the final submission has to be consistent with.

Add an explanation prompt. "Explain why you chose this framing over the two alternatives you considered." Short, specific, and hard to answer without having considered them.

Add a critique component. Supply a flawed answer and ask for the errors. Reuse the same flawed answer across the cohort so grading is fast and comparison is easy.

Reweight the rubric so the polished artifact is at most a third of the marks.

Then state the AI policy for the task in one sentence at the top of the brief, and say what should be disclosed. Run it once and note the grading time — done well, process assessment is usually faster to grade than essays, because a wrong critique is obvious in ten seconds where a mediocre essay takes ten minutes.

Proportionality

ToolDix original diagram
Constraints that decide what you may require
Access is not uniform
Paid tiers, device availability, and home connectivity differ. Provide a no-AI path.
Accounts and personal data
Requiring a sign-up sends student data to a third party. Check policy and age limits first.
Availability changes
A tool free this term may not be next term. Do not build the objective around one product.
Accessibility
For some learners assistive generation is the accommodation, not the shortcut.
The fourth row cuts against blanket bans: a rule written to preserve rigour can remove an accommodation someone depends on.

The right level of restriction depends on learner age, subject, accessibility needs, institutional policy, and whether the task teaches a foundational skill or an applied one. A first-year student learning to construct an argument and a final-year student applying one need different rules, and applying the same policy to both is a mistake in one direction or the other.

Two constraints deserve particular attention when assessment is involved. First, if AI use is permitted but only paid tiers produce good output, the assessment partly measures spending. Second, for some learners generative assistance is a documented accommodation; an assessment rule that removes it is removing support, not restoring rigour. Both argue for assessing process, which is far less sensitive to which tool anyone had.

Common mistakes

Acting on detector scores. They are not evidence, the errors are biased, and using them damages the trust that makes disclosure policies work.

One rule for the whole course. Different tasks assess different skills and need different lines.

Adding process requirements without removing anything. If you add a plan, a critique, and a defence while keeping the full essay, you have tripled the workload for everyone. Shrink the artifact as you add the evidence.

Sources and license context

These references informed the lesson. ToolDix adds its own explanation, workflow, and practice rather than reproducing source material. Every link below leaves ToolDix and opens the publisher's own site in a new tab.

Keep going

Read these next on ToolDix.

Original lessons that build on what you just read.