Item exposure is the quiet failure of a small question bank
A bank that is too small stops measuring knowledge and starts measuring who has seen the questions. The arithmetic is unforgiving and easy to check.
A model can draft assessment items far faster than a committee can. The bottleneck was never drafting, which is why the savings are smaller than they look.
2 min read
Written by
Seratlas Team
Engineering and delivery
2 min read
We build and operate assessment software. We write here about the parts that are harder than they look.
Generating a hundred multiple-choice items from a syllabus takes minutes. The reason this has not collapsed the cost of building an assessment is that drafting was never the expensive step — reviewing was, and a hundred drafts is a hundred things to review.
Volume, and coverage of the obvious. Given a well-specified topic, a model produces items that are grammatical, plausibly pitched, and reasonably well distributed across the content. For a low-stakes formative quiz that is genuinely most of the work.
It is also good at the narrow task of producing distractors for a stem someone else wrote, because a distractor is a bounded thing: wrong, but wrong in a way that corresponds to a real misconception.
A good distractor encodes an error that learners actually make. A generated one often encodes an error that is merely adjacent — plausible to a reader, but not attractive to anyone who has misunderstood the topic. The item then discriminates poorly, which is only visible after it has been administered.
Generated items drift toward what is common in the training data rather than what is in your curriculum. The result is an item that is defensible in general and unfair in context, because it tests material the candidate was never taught.
Cueing is the characteristic failure: the correct option is longer, or more hedged, or the only one that is grammatically consistent with the stem. These are easy to spot once you are looking and easy to miss in a batch of a hundred.
Treat generated items as drafts with provenance. Record which model, which prompt, and which source material produced each one, and keep that on the item for its whole life — when a fault pattern turns up later, provenance is what lets you find its siblings instead of auditing the whole bank.
Review in batches small enough to review properly, and pilot before scoring. Field-testing items unscored alongside live ones is the only step that substitutes for expert judgement, and it works regardless of how the item was written.
Generation moves effort from drafting to reviewing. That is a real gain where review is cheap and the stakes are low, and close to no gain at all on a certification exam, where every item needs a subject-matter expert's attention whether a person or a model produced the first version.
A bank that is too small stops measuring knowledge and starts measuring who has seen the questions. The arithmetic is unforgiving and easy to check.
Proctoring produces evidence about a room, not a verdict about a person. Treating its output as a decision is where institutions get into trouble.
Static hosting removes a class of outage from the pages candidates see first, and costs less than the runtime it replaces.
Assessment traffic arrives as a spike with a deadline attached, and the usual scaling advice assumes a load shape this is not.
If any of the above is a problem you are currently having, we are happy to talk about it without a sales process attached.