Skip to main content
Data and AICorporate training

Generated items still need a review step

A model can draft assessment items far faster than a committee can. The bottleneck was never drafting, which is why the savings are smaller than they look.

2 min read

ST

Written by

Seratlas Team

Engineering and delivery

2 min read

We build and operate assessment software. We write here about the parts that are harder than they look.

Generating a hundred multiple-choice items from a syllabus takes minutes. The reason this has not collapsed the cost of building an assessment is that drafting was never the expensive step — reviewing was, and a hundred drafts is a hundred things to review.

What generation is good at

Volume, and coverage of the obvious. Given a well-specified topic, a model produces items that are grammatical, plausibly pitched, and reasonably well distributed across the content. For a low-stakes formative quiz that is genuinely most of the work.

It is also good at the narrow task of producing distractors for a stem someone else wrote, because a distractor is a bounded thing: wrong, but wrong in a way that corresponds to a real misconception.

What it is reliably bad at

Knowing which misconception matters

A good distractor encodes an error that learners actually make. A generated one often encodes an error that is merely adjacent — plausible to a reader, but not attractive to anyone who has misunderstood the topic. The item then discriminates poorly, which is only visible after it has been administered.

Staying inside the taught scope

Generated items drift toward what is common in the training data rather than what is in your curriculum. The result is an item that is defensible in general and unfair in context, because it tests material the candidate was never taught.

Not leaking the answer

Cueing is the characteristic failure: the correct option is longer, or more hedged, or the only one that is grammatically consistent with the stem. These are easy to spot once you are looking and easy to miss in a batch of a hundred.

A workflow that holds up

Treat generated items as drafts with provenance. Record which model, which prompt, and which source material produced each one, and keep that on the item for its whole life — when a fault pattern turns up later, provenance is what lets you find its siblings instead of auditing the whole bank.

Review in batches small enough to review properly, and pilot before scoring. Field-testing items unscored alongside live ones is the only step that substitutes for expert judgement, and it works regardless of how the item was written.

The honest summary

Generation moves effort from drafting to reviewing. That is a real gain where review is cheap and the stakes are low, and close to no gain at all on a certification exam, where every item needs a subject-matter expert's attention whether a person or a model produced the first version.

Common questions

It can catch mechanical faults — a duplicated option, a missing key, an answer that contradicts the stem. It cannot tell you whether the item measures the thing your curriculum claims to teach, because that is a fact about your curriculum.

Variants of a validated item, and distractors for a stem a human has already written. Both are bounded tasks with a reviewable output, which is the shape of task where this works.

Building something like this?

If any of the above is a problem you are currently having, we are happy to talk about it without a sales process attached.