What a pass mark encodes
Sixty percent is not a standard, it is a habit. Choosing a cut score is a judgement about consequences, and it should be recorded as one.
A bank that is too small stops measuring knowledge and starts measuring who has seen the questions. The arithmetic is unforgiving and easy to check.
2 min read
Written by
Seratlas Team
Engineering and delivery
2 min read
We build and operate assessment software. We write here about the parts that are harder than they look.
A certification exam with two hundred items and four thousand candidates a year has a problem that does not appear in any score report. Every item is seen often enough, by enough people, that its content leaves the room. After a while the exam still discriminates — it just discriminates between candidates who encountered the circulating set and candidates who did not.
That is a measurement failure, and it is invisible from inside the results, because a compromised item often looks like an easy item.
Exposure is sittings multiplied by items per sitting, divided by bank size. Four thousand sittings of sixty items drawn from two hundred gives each item roughly twelve hundred viewings a year. No amount of candidate honesty survives that number; it only takes one person in twelve hundred to write the question down.
The reason this gets missed is that nobody computes it. Bank size is discussed as a content-production problem — how many items can we afford to write — rather than as a constraint derived from volume.
An item that has leaked gets easier over time, and only that item. Difficulty drift is normal across a whole form as the candidate population changes; drift concentrated in a handful of items, while their neighbours hold steady, is a different signal.
The stronger signal is that the item stops correlating with overall performance. A leaked item is answered correctly by people who prepared badly, which is precisely the pattern a discrimination index is designed to catch. Watching it per item, over time, catches exposure earlier than watching difficulty alone.
Retire on a schedule rather than on suspicion. An item with a planned service life is replaced before it has leaked; an item retired when someone notices something odd was compromised for however long the noticing took.
Draw from a larger pool than the form length, with content constraints applied at draw time so two candidates get statistically comparable forms from different items. That is more expensive to build than a fixed form, and it is the only structural answer — every other measure slows exposure without bounding it.
All of this needs the platform to know which items were served to whom and when, and to keep that history after the sitting is scored. Systems that treat a delivered form as transient throw away exactly the data the exposure analysis needs, and the gap is only discovered years later when someone asks how long an item has been in service.
Sixty percent is not a standard, it is a habit. Choosing a cut score is a judgement about consequences, and it should be recorded as one.
Assessment traffic arrives as a spike with a deadline attached, and the usual scaling advice assumes a load shape this is not.
Proctoring produces evidence about a room, not a verdict about a person. Treating its output as a decision is where institutions get into trouble.
Static hosting removes a class of outage from the pages candidates see first, and costs less than the runtime it replaces.
If any of the above is a problem you are currently having, we are happy to talk about it without a sales process attached.