Item exposure is the quiet failure of a small question bank
A bank that is too small stops measuring knowledge and starts measuring who has seen the questions. The arithmetic is unforgiving and easy to check.
Sixty percent is not a standard, it is a habit. Choosing a cut score is a judgement about consequences, and it should be recorded as one.
2 min read
Written by
Seratlas Team
Engineering and delivery
2 min read
We build and operate assessment software. We write here about the parts that are harder than they look.
Ask where a pass mark came from and the answer is usually that it has always been sixty percent. That is not a standard. It is a convention that has survived because nobody has had to defend it, and it stops surviving the first time someone does.
A cut score distributes two kinds of mistake. Set it low and some candidates pass who should not; set it high and some fail who should not. There is no setting that eliminates both, and the right balance depends entirely on which error costs more in this particular context.
For a pilot's licence, a false pass is catastrophic and a false fail is an inconvenience. For a placement test that routes students to a support class, it is close to reversed — a false fail wastes a term of someone's time.
Stating which way the asymmetry runs, before choosing a number, is most of the work. It is also the part that never makes it into the record.
A fixed percentage of a variable form is a moving standard. If this year's items are harder, sixty percent represents more knowledge than it did last year, and the cohorts are not comparable — but the score report presents them as if they were.
This is why standard-setting methods exist. They differ in how they gather expert judgement, and they share one property: the judgement is about the items, so the resulting cut score is anchored to content rather than to a round number.
A panel of people who know the domain, a definition of the borderline candidate specific enough to argue about, and item-by-item judgements aggregated into a threshold. The output is a number with a documented derivation, which is the difference that matters when it is challenged.
"Just barely qualified" is not specific enough to produce agreement. What can this person do, what do they still get wrong, and what would you let them do unsupervised? A panel that has not settled this is not making item judgements against a shared standard, and the aggregate then represents nothing in particular.
The asymmetry you chose and why. The method, the panel, and the definition of the borderline candidate. The resulting cut score, the form version it applies to, and the date it takes effect.
A cut score with that record behind it can be defended, adjusted, and explained to the person it failed. A cut score that is sixty percent because it has always been sixty percent can only be asserted, which works right up until somebody asks.
A bank that is too small stops measuring knowledge and starts measuring who has seen the questions. The arithmetic is unforgiving and easy to check.
Proctoring produces evidence about a room, not a verdict about a person. Treating its output as a decision is where institutions get into trouble.
Static hosting removes a class of outage from the pages candidates see first, and costs less than the runtime it replaces.
A model can draft assessment items far faster than a committee can. The bottleneck was never drafting, which is why the savings are smaller than they look.
If any of the above is a problem you are currently having, we are happy to talk about it without a sales process attached.