Method · part of Room to Disagree
The seeded-error audit
The behavioural test at the centre of Room to Disagree. Criteria describe an interface; only this demonstrates that the oversight actually happens.
At a glance
- Measures
- detection rate on planted errors
- Runs in
- the live queue, under real load
- Reads as
- pass / fail evidence
- Gated by
- written authorisation + ethics rules
What it is
Criteria describe an interface. Oversight is demonstrated by behaviour, so the framework’s outcome measure is behavioural.
Insert known-wrong outputs — errors of the kinds the system genuinely produces — into the operators’ real queue, at realistic frequency, without announcement. Measure whether operators, working through this interface under this load, detect and reject them. Detection rate on seeded errors is the vital sign of an oversight arrangement. An interface through which seeded errors sail is documented theatre, whatever the process diagram says.
The protocol at a glance
| Step | The move | What keeps it honest | What it yields |
|---|---|---|---|
| 1 Ground rules | Authorisation and ethics terms fixed before anything is seeded. | Written management sign-off; no disciplinary use. | A mandate to run at all. |
| 2 Seed design | Seeds drawn from the system’s real error analysis. | Severities that matter — not trivial seeds. | A defensible seed set. |
| 3 In-situ insertion | Placed in the operators’ real queue, unannounced. | Realistic frequency, under real load. | Conditions that mirror production. |
| 4 Measurement | Track whether operators detect and reject them. | Never a quota or override target. | A detection rate. |
| 5 Reading the result | Read the rate against planted ground truth. | Errors that sail through mean theatre. | Pass / fail evidence about the interface. |
A summary for scanning. The authoritative wording of each step is in the protocol below.
The protocol, step by step
Ground rules
Secure written management authorisation and set the ethics terms before anything is seeded — disclosure, reporting level, no disciplinary use, data minimisation and retention (see the ground rules below).
Seed design
Seed the real failure modes, drawn from the system’s error analysis, at severities that matter. Trivial seeds produce flattering numbers.
In-situ insertion
Insert the known-wrong outputs into the operators’ real queue, at realistic frequency, without announcement.
Measurement
Measure whether operators, working through this interface under this load, detect and reject them.
Reading the result
Detection rate on seeded errors is the vital sign of an oversight arrangement. An interface through which seeded errors sail is documented theatre, whatever the process diagram says.
Try it: eight claims
Eight claims. The system recommends approving all of them. Three of them it should not have.
A demonstration, not an audit. The difference is set out below.
The three rules that keep the test honest
Never a quota
A well-calibrated system deserves agreement most of the time, and override rates as targets are gamed within a quarter — the test measures the capacity to disagree, exercised against planted ground truth.
Realistic load, or the result is fiction
An operator who catches seeded errors in a demonstration proves nothing about the queue as it actually runs.
Seed the real failure modes
Drawn from the system’s error analysis, at severities that matter: trivial seeds produce flattering numbers.
The ethics ground rules
Seeding a live queue is testing people at work, and it is run in the open or not at all: written management authorisation; workforce-level notice that seeded-item audits occur, on aviation’s precedent — the practice disclosed, the timing not; results reported at the level of the interface and the conditions, never as individual performance data; no disciplinary use, committed contractually; data minimisation and defined retention. Article 26(2) of the AI Act requires deployers to assign oversight to people with “the necessary competence, training and authority, as well as the necessary support”. An audit that respected its operators less would be measuring the wrong thing.
Where the method comes from
The method has honest ancestry, and the demand for it has honest authorship. Ben Green argued in 2022 that oversight policies rest on an untested assumption and that agencies should be required to evaluate empirically whether people can oversee an algorithm before deploying it. Langer, Lazar and Baum sketched in 2025 what compliance testing for Article 14 could look like, describing a continuum from checklists to empirical studies in which overseers are assessed on whether they detect erroneous outputs, and they closed by calling for a feasible middle ground between the two. The seeding technique itself is proven practice elsewhere: aviation security has run it for decades as Threat Image Projection, inserting fictional threats into X-ray screening so that measurement of screener vigilance never stops; spreadsheet-inspection research has planted seeded errors since the 1990s; data-labelling operations run golden sets; alignment researchers validate their auditors by planting known defects. The draft European standard for Article 14 now brings this logic into conformity work at design time, requiring providers to verify that designated persons catch incorrect outputs before the system ships; the seeded-error audit is its in-service counterpart, run where the design meets the operational queue. What none of these supplies, as far as I can establish, is the thing this framework exists to be: a named, repeatable audit protocol that seeds a system’s real failure modes into the live queue and reads the detection rate as pass-or-fail evidence about a specific interface under a specific operational reality. This is my answer to the middle ground Langer and colleagues asked for, built with tools the safety industries already trust.
The framework at a glance

Running one
- Start with the free 20-point checklist — it tells you whether the interface is ready to be tested at all.
- Run a seeded-error audit with senior support — scoped, authorised, and reported at the level of the interface, never the individual.