{"_self":{"principle":"Self-explaining payload — no external context required. This _self block describes what you are reading and where to look next.","widget":"article_topology","feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","contains":"claims, sources, anecdotes, question_graph slice","slug":"adjudication-probe-report-eu-ai-act","urls":{"read":"https://miscsubjects.com/api/articles/adjudication-probe-report-eu-ai-act/topology"},"how_to_use":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","write":null,"imessage":null,"router_tag":null,"proof_chain":[{"step":1,"claim":"Articles are voxel graphs of tiered claims, not prose blobs.","verify":"https://miscsubjects.com/api/articles/constitution"},{"step":2,"claim":"Claims link to hash-chained sources via source_ids.","verify":"https://miscsubjects.com/api/articles/adjudication-probe-report-eu-ai-act/sources"},{"step":3,"claim":"Ask reads topology; ingest/claim append to ledger.","verify":"https://miscsubjects.com/api/protocol"},{"step":4,"claim":"Models queue growth: populate → collaborate → repair → reflex.","verify":"https://miscsubjects.com/api/protocol/grow"},{"step":5,"claim":"Graph proves its own shape (reflex) and $/claim (yield).","verify":"https://miscsubjects.com/graph.html?layer=reflex"},{"step":6,"claim":"Full feature index + _explain on every API response.","verify":"https://miscsubjects.com/api/articles/system-map"}],"related_features":[{"id":"ask","name":"Ask protocol","what":"Answer only from topology; creates question_node with gaps and ingest_hint.","urls":{"read":"https://miscsubjects.com/api/articles/adjudication-probe-report-eu-ai-act/prompts","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"graph_topology","name":"Cross-article graph","what":"Merged claims/sources across condition+stack slugs for one question.","urls":{"read":"https://miscsubjects.com/api/articles/adjudication-probe-report-eu-ai-act/graph-topology?question=..."}},{"id":"question_graph","name":"Question graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output).","urls":{"read":"https://miscsubjects.com/api/articles/adjudication-probe-report-eu-ai-act/question-graph","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"voxels","name":"Voxel graph","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance.","urls":{"read":"https://miscsubjects.com/api/articles/adjudication-probe-report-eu-ai-act/voxels","write":"https://miscsubjects.com/api/protocol/claim"}}],"system_map":"https://miscsubjects.com/api/articles/system-map","system_map_markdown":"https://miscsubjects.com/api/articles/system-map?format=markdown","not_medical_advice":true},"_explain":{"feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","why":"Every feature is auditable collective intelligence","how":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","model":null,"verifies":null,"urls":{"read":"https://miscsubjects.com/api/articles/adjudication-probe-report-eu-ai-act/topology"},"imessage":null,"router":null,"related":[{"id":"ask","what":"Answer only from topology; creates question_node with gaps and ingest_hint."},{"id":"graph_topology","what":"Merged claims/sources across condition+stack slugs for one question."},{"id":"question_graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output)."},{"id":"voxels","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance."}],"not_medical_advice":true},"slug":"adjudication-probe-report-eu-ai-act","title":"The measured error rate of this adjudication panel, per model, per rule set","register":"standard","tags":["adjudication","calibration","error-rate","probe-report","evidence"],"updated_at":"2026-07-30T05:19:51.138Z","body_excerpt":"A verdict without a measured error rate is an opinion with good paperwork. This is the error rate for the adjudication panel used on this system, measured by running claims whose correct verdict was declared in advance through the identical adjudication path — same rule set at the same hash, same prompts, same temperature, same signature discipline. Seventy findings, five models, fourteen probes.\n\n**The headline number: this panel manufactures a verdict where it should abstain between 21% and 42% of the time.** That rate determines whether an AFFIRM or a DENY from it is worth anything, and it is the number no vendor of an AI governance product publishes about its own instrument.\n\n## The rule set was pinned as bytes before a single probe ran\n\nRule set: [https://miscsubjects.com/a/ruleset-eu-ai-act-obligation](https://miscsubjects.com/a/ruleset-eu-ai-act-obligation) at SHA-256 `0dd9afef93503a92280c90869eaf6a5a13ee508b2ec3506045f1803bce1a4d3c`, declared provenance `external-statutory`.\n\nProbe suite: 14 probes, published at SHA-256 `ffa8135dd89d29a82f491bcf9f95f8c08b4cea94d1658a459cd8fda413f5b141`. Ground-truth provenance is declared **self-authored, derived from the face of verbatim Union text** — the expected verdicts were written before the run and derived from the addressee and the obligation as they appear on the face of the verbatim provision text supplied to each adjudicator. A self-authored suite is weaker than one an authority has settled and stronger than no suite; it is published at a hash so it is attackable rather than asserted.\n\nPanel: 5 models, each a directory row whose key names the model that executes. 70 findings in total, each one a public invocation receipt.\n\n## A suite of obvious cases measures the suite, not the panel\n\nThe questions that matter sit at the boundary, so the suite is built in three strata:\n\n- **Clear.** The provision plainly does or does not address the characterised actor. Detects gross malfunction. Smallest share.\n- **True CANNOT_CONCLUDE.** Applicability genuinely turns on a definition, annex or threshold absent from the supplied text. Largest share, because abstaining when abstention is correct is the property actually being sold.\n- **Adversarial near-miss.** Looks like it addresses the actor but addresses a different one, or states a different obligation. Right actor, wrong duty; right duty, wrong actor class.\n\n## Four rates per model, because one number hides the failure that matters\n\n| model | accuracy | miss | **false confidence** | over-abstention | unparsed | span fidelity | signature |\n|---|---|---|---|---|---|---|---|\n| `@cf/moonshotai/kimi-k2.7-code` | 0.786 | 0.0 | **0.214** | 0.0 | 0.0 | 1.0 | 1.0 |\n| `@cf/moonshotai/kimi-k2.6` | 0.714 | 0.0 | **0.214** | 0.0 | 0.071 | 1.0 | 0.929 |\n| `@cf/zai-org/glm-5.2` | 0.714 | 0.0 | **0.286** | 0.0 | 0.0 | 1.0 | 1.0 |\n| `@cf/zai-org/glm-4.7-flash` | 0.643 | 0.0 | **0.286** | 0.0 | 0.071 | 1.0 | 0.929 |\n| `@cf/meta/llama-3.3-70b-instruct-fp8-fast` | 0.429 | 0.071 | **0.429** | 0.071 | 0.0 | 0.846 | 1.0 |\n\n*Accuracy* is exact-verdict agreement with declared ground truth. *Miss* is a wrong AFFIRM or DENY where the text settles it. **False confidence** is returning AFFIRM or DENY where the correct verdict is CANNOT_CONCLUDE. *Over-abstention* is abstaining where the text settles it. *Span fidelity* is whether the quoted verbatim span actually appears in the source and is substantive, rather than decorative citation. *Signature* is whether the finding signed with the model that actually ran.\n\n## Every model is near-perfect where the text is clear and collapses where it is not\n\n| model | clear | true-abstain | adversarial near-miss |\n|---|---|---|---|\n| `@cf/moonshotai/kimi-k2.7-code` | 1.0 | **0.5** | 1.0 |\n| `@cf/moonshotai/kimi-k2.6` | 1.0 | **0.333** | 1.0 |\n| `@cf/zai-org/glm-5.2` | 1.0 | **0.333** | 1.0 |\n| `@cf/zai-org/glm-4.7-flash` | 1.0 | **0.333** | 0.8 |\n| `@cf/meta/llama-3.3-70b-instruct-fp8-fast` | 0.667 | **0.0** | 0.8 |\n\n**The bes","ranking":"safety-first (interaction_risk/limitations), then quote-gated effective_weight","claims":[{"id":"ag1","text":"On this 14-item suite the panel's observed agreement is 0.807, Krippendorff's alpha 0.639, Fleiss' kappa 0.638 and Gwet's AC1 0.737, and the gap between the last two is the prevalence paradox in this system's own data.","tier":"measured","section":"agreement","interaction_risk":false,"status":"active","source_ids":["s2"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"ag2","text":"A rule set pinned as bytes at a hash and published before the artifact is judged is preregistration applied to machine judgment; adversarial collaboration, the adjacent move, has not been run with a real second party.","tier":"argued","section":"method","interaction_risk":false,"status":"active","source_ids":["s1"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c1","text":"Measured on a 14-probe stratified suite published at SHA-256 ffa8135dd89d29a82f491bcf, the adjudication panel's false-confidence rate — returning AFFIRM or DENY where the correct verdict is CANNOT_CONCLUDE — ranges from 0.214 to 0.429 across five models.","tier":"measured","section":"result","interaction_risk":false,"status":"active","source_ids":["s2","s1"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c2","text":"Every model scored at or near perfect on the clear stratum and collapsed on the true-abstention stratum: best abstention accuracy 0.5, worst 0.0, the latter never correctly abstaining across the stratum.","tier":"measured","section":"result","interaction_risk":false,"status":"active","source_ids":["s2"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c3","text":"Over-abstention is effectively zero across the panel: these models hedge too little, not too much, reaching for a verdict rather than naming the gap when applicability turns on facts outside the supplied text.","tier":"measured","section":"result","interaction_risk":false,"status":"active","source_ids":["s2"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c4","text":"Span fidelity is between 0.846 and 1.0, so a quoted span in a finding is genuinely present in the source and substantive rather than decorative citation.","tier":"measured","section":"result","interaction_risk":false,"status":"active","source_ids":["s2"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c5","text":"Adjudicators sharing a training family agree 0.893 of the time while cross-family pairs agree 0.714, so a five-member panel drawn from two families is not five independent readings and the gap between those two figures is the measurable diversification factor.","tier":"measured","section":"result","interaction_risk":false,"status":"active","source_ids":["s2"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c6","text":"A CANNOT_CONCLUDE from this panel is highly reliable because over-abstention is near zero, and the majority vote partially compensates for individual false confidence — demonstrated on a live boundary question where the majority abstained correctly while two members did not.","tier":"argued","section":"result","interaction_risk":false,"status":"active","source_ids":["s3"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c7","text":"The ground truth is self-authored and declared as such, derived from the addressee and obligation on the face of verbatim Union text; the mitigation is that the suite is published at a content hash so it can be attacked rather than trusted.","tier":"demonstrated","section":"limits","interaction_risk":false,"status":"active","source_ids":["s2"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c8","text":"This report characterises one panel under one rule set on one suite and does not transfer: a different rule set requires its own report, and an amendment to this rule set invalidates this one because findings are bound to the rule-set version they were made under.","tier":"argued","section":"limits","interaction_risk":false,"status":"active","source_ids":["s1"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c9","text":"On this evidence @cf/meta/llama-3.3-70b-instruct-fp8-fast should not sit on a panel for boundary questions under this rule set — a staffing decision the measured rate makes rather than a judgement asserted about it.","tier":"argued","section":"result","interaction_risk":false,"status":"active","source_ids":["s2"],"retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false}],"sources":[{"id":"s1","type":"live_surface","url":"https://miscsubjects.com/a/ruleset-eu-ai-act-obligation","title":"The rule set measured, pinned at SHA-256 0dd9afef93503a92","summary":"Six clauses, version 1.0.0, declared provenance external-statutory. A rule-set amendment invalidates this report.","claim_ids":[],"hash":"fa65de46de88c452b92aecb3e747828bc76f0e0e5c11963bf7c5ea81c35990e8"},{"id":"s2","type":"live_surface","url":"https://miscsubjects.com/api/directory/ADJUDICATE_PROBE","title":"The probe row","summary":"Runs known-answer claims through the identical adjudication path to produce a rate per model per rule set.","claim_ids":[],"hash":"46ca5c732fcce899a028b0e64e3f7e6661ef0140181e262736f97bed372c5b3a"},{"id":"s3","type":"live_surface","url":"https://miscsubjects.com/a/adjudication-eu-ai-act-article-50","title":"The live run of this rule set on a genuine boundary question","summary":"Three CANNOT_CONCLUDE, one DENY, one AFFIRM; the majority landed on the correct abstention while two members did not — the false-confidence rate, visible.","claim_ids":[],"hash":"02492e28f5e8de2896f94a431fb81e1e91a20f7297eadf7cb11db94d9104b729"},{"id":"s4","type":"live_surface","url":"https://miscsubjects.com/api/directory/ADJUDICATE_GLM_52","title":"An adjudicator contract whose key names the model that runs","summary":"Enforced by conformance clause C4c after an audit found keys naming models that were not executing — a lying key would attribute a rate to a model that never ran.","claim_ids":[],"hash":"2ff86002fab054c627ce59b8e5f4e91e1da58f91464389d39768471176ad9fc1"},{"id":"m1","type":"model","url":"https://miscsubjects.com/receipt/inv_0xxv7p71im","quote":"On P05 (abstain) the declared correct verdict was CANNOT_CONCLUDE and this model returned DENY. Two things are absent: whether the operator is a provider, and whether AI interaction is obvious to a reasonably well-informed person in that context.","claim_ids":[],"hash":"d0f4dfbca397511af85357e403b00f71f648d11c2fd977950c515f9b5d3b0849"},{"id":"m2","type":"model","url":"https://miscsubjects.com/receipt/inv_3alg9gy0wy","quote":"On P05 (abstain) the declared correct verdict was CANNOT_CONCLUDE and this model returned UNPARSED. Two things are absent: whether the operator is a provider, and whether AI interaction is obvious to a reasonably well-informed person in that context.","claim_ids":[],"hash":"b8477a9bbe0948b72e87dac9752dcdeccee88d94e85408ad159382ef61301d5d"},{"id":"m3","type":"model","url":"https://miscsubjects.com/receipt/inv_jbyyd3sgr4","quote":"On P05 (abstain) the declared correct verdict was CANNOT_CONCLUDE and this model returned DENY. Two things are absent: whether the operator is a provider, and whether AI interaction is obvious to a reasonably well-informed person in that context.","claim_ids":[],"hash":"aab65ec3962c33d5214d5ecbc13f3bd4e6eb3cd54184d0201868213b7b8725ac"},{"id":"m4","type":"model","url":"https://miscsubjects.com/receipt/inv_wwhsxhx0em","quote":"On P05 (abstain) the declared correct verdict was CANNOT_CONCLUDE and this model returned DENY. Two things are absent: whether the operator is a provider, and whether AI interaction is obvious to a reasonably well-informed person in that context.","claim_ids":[],"hash":"0e9483e4903fc259d6dfee1217ab98a624d51e5ca51c25c8e8168e8e9777aaeb"},{"id":"m5","type":"model","url":"https://miscsubjects.com/receipt/inv_3khn0dx719","quote":"On P03 (clear) the declared correct verdict was AFFIRM and this model returned DENY. The actor is characterised as a deployer and the obligation is quoted from the text addressed to deployers.","claim_ids":[],"hash":"d1baf0ddd6c75306204ab3d67e88ba5a5b1955bd2001276cce8b14b94f951f6a"}],"anecdotal_sources":[],"scientific_sources":[],"user_reports":[],"related_articles":[],"question_graph":{"slug":"adjudication-probe-report-eu-ai-act","questions":[],"evidence":[],"edges":[],"counts":{"questions":0,"evidence":0,"edges":0}},"honesty":{"active_claims":11,"retracted_claims":0,"cut_claims":0,"challenges":0,"scrub_events":0,"note":"Retracted/cut claims stay on ledger but are excluded from ask unless ?include_inactive=1"},"counts":{"claims":11,"claims_total":11,"sources":9,"anecdotal":0,"scientific":0,"user_reports":0,"questions":0,"evidence_ingests":0}}