Skip to main content
Assessment scienceK-12 education

What a pass mark encodes

Sixty percent is not a standard, it is a habit. Choosing a cut score is a judgement about consequences, and it should be recorded as one.

2 min read

ST

Written by

Seratlas Team

Engineering and delivery

2 min read

We build and operate assessment software. We write here about the parts that are harder than they look.

Ask where a pass mark came from and the answer is usually that it has always been sixty percent. That is not a standard. It is a convention that has survived because nobody has had to defend it, and it stops surviving the first time someone does.

The number is a judgement about two errors

A cut score distributes two kinds of mistake. Set it low and some candidates pass who should not; set it high and some fail who should not. There is no setting that eliminates both, and the right balance depends entirely on which error costs more in this particular context.

For a pilot's licence, a false pass is catastrophic and a false fail is an inconvenience. For a placement test that routes students to a support class, it is close to reversed — a false fail wastes a term of someone's time.

Stating which way the asymmetry runs, before choosing a number, is most of the work. It is also the part that never makes it into the record.

Why a percentage does not hold anything constant

A fixed percentage of a variable form is a moving standard. If this year's items are harder, sixty percent represents more knowledge than it did last year, and the cohorts are not comparable — but the score report presents them as if they were.

This is why standard-setting methods exist. They differ in how they gather expert judgement, and they share one property: the judgement is about the items, so the resulting cut score is anchored to content rather than to a round number.

What the methods have in common

A panel of people who know the domain, a definition of the borderline candidate specific enough to argue about, and item-by-item judgements aggregated into a threshold. The output is a number with a documented derivation, which is the difference that matters when it is challenged.

The borderline candidate is the load-bearing definition

"Just barely qualified" is not specific enough to produce agreement. What can this person do, what do they still get wrong, and what would you let them do unsupervised? A panel that has not settled this is not making item judgements against a shared standard, and the aggregate then represents nothing in particular.

What to record

The asymmetry you chose and why. The method, the panel, and the definition of the borderline candidate. The resulting cut score, the form version it applies to, and the date it takes effect.

A cut score with that record behind it can be defended, adjusted, and explained to the person it failed. A cut score that is sixty percent because it has always been sixty percent can only be asserted, which works right up until somebody asks.

Common questions

Because it ties the standard to the difficulty of the particular items on the form. Two forms covering the same content at different difficulties produce different candidates at the same 60%, which means the number is not holding anything constant.

Much less. The argument scales with consequence — where the outcome is feedback, an approximate threshold is fine, and where it is a licence or a grade, it is not.

Building something like this?

If any of the above is a problem you are currently having, we are happy to talk about it without a sales process attached.