Skip to main content
Security and complianceHigher education

What remote proctoring actually proves

Proctoring produces evidence about a room, not a verdict about a person. Treating its output as a decision is where institutions get into trouble.

2 min read

ST

Written by

Seratlas Team

Engineering and delivery

2 min read

We build and operate assessment software. We write here about the parts that are harder than they look.

A proctoring system observes a webcam feed, some browser events, and occasionally a desktop. From that it produces flags: a face left the frame, a second voice was audible, a window lost focus. Every one of those is a fact about a room.

None of them is a fact about cheating, and the distance between the two is where most of the harm in this field happens.

The confusion is structural, not careless

Flags arrive in a queue, sorted by a confidence score, in an interface that asks for a decision. Everything about that presentation implies the system has found something, and the reviewer's job is to confirm it. The base rate works the other way round: most flags in most exams are innocent, because looking away from a screen is what people do while thinking.

A reviewer processing two hundred flagged sessions under time pressure is not applying a prior. They are clearing a queue.

What the evidence supports

Identity, reasonably well

Matching a candidate to a credential at the start of a session is the part of proctoring that works, and it addresses the substitution case, which is the most consequential form of exam fraud.

Environment, weakly

A room scan establishes what was visible for the seconds it took. It does not establish what was outside the frame, which is most of the room.

Intent, not at all

No signal available to a proctoring system distinguishes consulting a note from looking at a wall. The inference is entirely the reviewer's, and the system's confidence score describes how unusual the signal was, not how likely the conclusion is.

What this means for the process

Separate the flag from the finding, in the record as well as in the procedure. A session log should say what was observed; a decision record should say who concluded what from it, and on what other evidence. Systems that store a verdict where the observation belongs make an appeal impossible to conduct, because the original signal is no longer recoverable.

Give the candidate the material. An allegation supported by a recording the candidate cannot see is not one they can answer, and in several jurisdictions that is not merely unfair but unlawful.

Keep a human decision in the loop and make its authorship explicit. "The system flagged you" is not a decision anyone can be held to, which is exactly why it is the sentence institutions reach for.

The uncomfortable part

Proctoring is often bought to make a claim to an accreditor rather than to detect misconduct, and it succeeds at that regardless of whether its flags mean anything. That mismatch is worth naming internally before the first appeal arrives, because the appeal is where the claim gets tested.

Common questions

No — it is useful as triage, narrowing hours of recording to minutes a human reviews. It stops being useful the moment its output is treated as a finding rather than a pointer.

That is a validity issue before it is a fairness one, and it needs an alternative route rather than an exception process. A requirement most candidates can meet and some cannot is a requirement that measures housing.

Building something like this?

If any of the above is a problem you are currently having, we are happy to talk about it without a sales process attached.