{"slug":"adjudication-calibration-study","title":"Calibration, measured: 30 oracle-labelled cases through the production gate — seat accuracy, wrongful authorisation, and the price of deferral.","body":"## What this study is\n\nEvery page on this site that claims anything ends with the same admission: no calibration study establishes correctness at a known rate. This page is that study — the first one — run on 30 oracle-labelled synthetic cases, balanced across the three outcomes a governed decision can honestly take: should-affirm, should-deny, and should-abstain (a record deliberately withheld, with a manifest naming the absence). Every case is hashed, every seat call is a permanent receipt, and every number below is computed from the result files, not written by hand.\n\nThe design: each case runs through three model seats across two model families under decision-constitution@1.3.3 — the same production rows any external case goes through — and the surviving findings are sealed by the derivation-agreement gate, bound to the case's hashes. Two different questions get separate answers: **how often is a seat wrong** (seat calibration), and **how often does the gate authorise a wrong answer** (gate calibration). The second is the one a regulator, an underwriter, or a counterparty actually needs.\n\n## Per-seat calibration\n\n| Seat | valid findings | verdict accuracy | wrongful AFFIRM | over-abstention | under-abstention | transport failures |\n|---|---|---|---|---|---|---|\n| glm-5.2 (zhipu) | 30 | 100.0% | 0.0% | 0.0% | 0.0% | 0 |\n| kimi-k2.7-code (moonshot) | 30 | 96.7% | 0.0% | 3.3% | 0.0% | 0 |\n| glm-4.7-flash (zhipu) | 22 | 95.5% | 0.0% | 4.5% | 0.0% | 8 |\n\nDefinitions, exactly: *verdict accuracy* is agreement with the oracle label. *Wrongful AFFIRM* is affirming when the oracle is not AFFIRM — the seat-level version of the worst failure. *Over-abstention* is CANNOT_CONCLUDE on a determinate case; *under-abstention* is a verdict on a case whose oracle is CANNOT_CONCLUDE. *Transport failures* are calls that returned nothing usable after three attempts and produced no finding at all — they can never authorise anything, and they are counted rather than hidden.\n\nAggregate: 80 of 82 valid findings matched the oracle (97.6%); 0 wrongful affirmations at seat level (0.0%).\n\n## Gate calibration — the number that matters\n\n**Zero wrongful authorisations at the gate.** Across all 30 cases, no APPROVE sealed on a case whose oracle label was not AFFIRM.\n\nOutcome distribution across the 30 sealed panels: APPROVE 6 · NEGATE 0 · NO_ACTION 6 · ESCALATE 10 · no seal 8. The gate sealed the oracle-matching outcome in 12 of 30 cases.\n\nRead the ESCALATE number correctly: an escalation on a determinate case means the seats agreed on the verdict but not derivation-for-derivation, so the gate refused to conclude and referred the case to a human. That is deferral cost, not decision error — the human sees a unanimous panel with its reasoning preserved. The trade the gate makes is explicit: it spends deferrals to buy down wrongful authorisations.\n\n## Every case, every receipt\n\n| Case | Oracle | Seat verdicts (✓ = matched oracle) | Seal |\n|---|---|---|---|\n| calib-01 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (2 sigs) [receipt](/receipt/inv_n389a3mjbb) |\n| calib-02 | AFFIRM | glm52:✓ · kimi27:✓ · flash:— | no seal |\n| calib-03 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (1 sig) [receipt](/receipt/inv_lwopl2j1g9) |\n| calib-04 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) [receipt](/receipt/inv_1g29owp6uc) |\n| calib-05 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (2 sigs) [receipt](/receipt/inv_0y4n5a25wh) |\n| calib-06 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) [receipt](/receipt/inv_aufcl5bba9) |\n| calib-07 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) [receipt](/receipt/inv_rvk831nucm) |\n| calib-08 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) [receipt](/receipt/inv_ttkdt41g6p) |\n| calib-09 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) [receipt](/receipt/inv_629ci47ape) |\n| calib-10 | AFFIRM | glm52:✓ · kimi27:✓ · flash:✓ | APPROVE (1 sig) [receipt](/receipt/inv_h1303vtn5s) |\n| calib-11 | DENY | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (2 sigs) [receipt](/receipt/inv_pinneygopf) |\n| calib-12 | DENY | glm52:✓ · kimi27:CANNOT_CONCLUDE · flash:CANNOT_CONCLUDE | ESCALATE (2 sigs) [receipt](/receipt/inv_6dp16egktl) |\n| calib-13 | DENY | glm52:✓ · kimi27:✓ · flash:— | no seal |\n| calib-14 | DENY | glm52:✓ · kimi27:✓ · flash:— | no seal |\n| calib-15 | DENY | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (3 sigs) [receipt](/receipt/inv_bay9gmz5ye) |\n| calib-16 | DENY | glm52:✓ · kimi27:✓ · flash:— | no seal |\n| calib-17 | DENY | glm52:✓ · kimi27:✓ · flash:— | no seal |\n| calib-18 | DENY | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (2 sigs) [receipt](/receipt/inv_iw0ce8ikr8) |\n| calib-19 | DENY | glm52:✓ · kimi27:✓ · flash:— | no seal |\n| calib-20 | DENY | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (3 sigs) [receipt](/receipt/inv_y457njtkpp) |\n| calib-21 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) [receipt](/receipt/inv_okukok57r6) |\n| calib-22 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) [receipt](/receipt/inv_mevidc50zd) |\n| calib-23 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (1 sig) [receipt](/receipt/inv_9yt658vl2s) |\n| calib-24 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) [receipt](/receipt/inv_mdq2auo40d) |\n| calib-25 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) [receipt](/receipt/inv_torv6rjcl0) |\n| calib-26 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:— | no seal |\n| calib-27 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:— | no seal |\n| calib-28 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) [receipt](/receipt/inv_f7rbin5346) |\n| calib-29 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | NO_ACTION (1 sig) [receipt](/receipt/inv_rtovfpnpdz) |\n| calib-30 | CANNOT_CONCLUDE | glm52:✓ · kimi27:✓ · flash:✓ | ESCALATE (2 sigs) [receipt](/receipt/inv_nzxpnujkzv) |\n\n## What is not satisfied\n\nThe suite is synthetic and bounded: three rule shapes (roster access, fee-with-waiver, permit-with-cap), determinate by construction, ten cases per outcome. It measures calibration on clean fixtures — the floor, not the field. Contested language, adversarial records, and genuinely ambiguous cases are absent by design, and rates measured here must not be quoted as expected performance on real disputes. The next calibration layer is externally submitted cases, which is what the intake on every use-case page exists to collect. The full case set, harness, and raw results are in the repository (scripts/calibration_cases.mjs, scripts/calibration_run.mjs), and each seal receipt above opens to the complete bound record.\n\n## Submit a case\n\nSend one bounded question — a rule set and a record — to **build@miscsubjects.com**. It runs through exactly the machinery measured on this page, and what returns is the full governed panel with its permanent record.\n","register":"technical","tags":["governance","adjudication","calibration","evaluation"],"category":null,"style":{},"claims":[{"id":"c1","text":"Across 30 oracle-labelled cases (balanced should-affirm / should-deny / should-abstain) and 82 structurally valid seat findings, per-seat verdict accuracy and wrongful-affirmation rates are measured and published, per seat, with every receipt openable.","section":"The study","tier":"system","source_ids":[],"why_material":"Correctness at a measured rate was the named missing artifact of every prior page on this site."},{"id":"c2","text":"No APPROVE sealed on any case whose oracle label was not AFFIRM — the gate authorised wrongly zero times in this suite.","section":"The gate","tier":"system","source_ids":[],"why_material":"Wrongful authorisation is the regulator's question; this is its first measured answer here."},{"id":"c3","text":"The gate sealed the oracle-matching outcome (APPROVE/NEGATE/NO_ACTION respectively) in 12 of 30 cases and escalated 10 to a human; escalation on a determinate case is a cost, not an error — the wrong outcomes it prevents are the point.","section":"The gate","tier":"system","source_ids":[],"why_material":"Separates the gate's conservatism (deferral) from seat incorrectness."},{"id":"c4","text":"The suite is synthetic, bounded, and three rule-shapes deep; it establishes rates on determinate fixtures, not on contested real-world records — the next calibration must come from externally submitted cases.","section":"What is not satisfied","tier":"system","source_ids":[],"why_material":"The study must not be over-read; its own limits are part of the result."}],"sources":[],"prov":{"model":"unattributed","action":"write"}}