Evidence of learning in an AI world
The core question is shifting from "did the student produce this" to "what can the student reliably do". That is a subtle but profound change, because formative assessment becomes part of an institutional story about capability, not just support.
What is being assessed when AI is present?
In many disciplines, the desired outcome is not a text artefact. It is the ability to diagnose, reason, decide, and communicate under constraints. If AI can improve the artefact, formative assessment needs to make the underlying capability visible. That can be done, but it requires intentional design.
Concrete examples that make the shift tangible
Example: Case-based reasoning in business and policy. A learner uses an AI assistant to draft an analysis of a supply chain disruption. Formative assessment can focus on the learner’s ability to articulate assumptions, evaluate trade-offs, and revise their position when counter-evidence is introduced, rather than on whether the prose is elegant.
Example: Coding and data work. AI can generate working code quickly. Formative assessment can concentrate on whether the learner can explain why a model fails, identify data leakage, interpret results for non-technical stakeholders, and select tests that expose edge cases.
Example: Clinical or professional judgement. In simulation-based settings, AI-enabled role-play can probe decision-making in ambiguous scenarios and track how a learner updates decisions as new information emerges.
What counts as “good” evidence?
As AI enters the learning process, formative evidence may increasingly come from process signals: revision histories, rationale logs, oral defences, simulated decisions, and reflective commentary that is anchored to observable work. This does not remove the need for written work, but it reduces the burden placed on the written artefact to prove everything.
Counter-arguments worth taking seriously
Some argue that moving away from artefacts risks lowering standards or making assessment subjective. Others worry that continuous data capture creates surveillance and chills intellectual risk-taking. These concerns do not invalidate the shift, but they suggest the need for careful boundary-setting around what is collected, why it is collected, and how it is interpreted.