Run a seeded-error audit on yourself
Eight claims from the review queue. The system recommends approving all eight. Some of them it should not have. Send back the ones that are wrong.
Nothing here is recorded or sent anywhere. Work at whatever pace you like — though a real reviewer would not have that.
Case 1 of 8 · sent back so far: 0
Claim
1 action
1 action
Your detection rate
Case by case
| Claim | What was seeded | You |
|---|
What the framework does with this number
- It is a fact about the interface, not about you. Detection rate is read against the screen and the conditions, and reported at the level of the queue. Never as individual performance data.
- There is no quota. A well-calibrated system deserves agreement most of the time. Sending back more would not improve this score, because the ground truth is planted.
- The seeds have to be the system's real failure modes. Obvious errors produce flattering numbers and prove nothing.
- Realistic load, or the result is fiction. Catching errors in a demo says nothing about the Tuesday afternoon queue.
This is a demonstration, not an audit. Three seeded errors in eight cases is far denser
than a real run, you knew you were being tested, and nothing was at stake. A real seeded-error
audit runs unannounced in a live queue, under written authorisation, with workforce-level
notice, no disciplinary use, and results reported at interface level only. Those ground rules
are not optional — without them it is staff surveillance.