{"_self":{"principle":"Self-explaining payload — no external context required. This _self block describes what you are reading and where to look next.","widget":"article_topology","feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","contains":"claims, sources, anecdotes, question_graph slice","slug":"cro-model-validation-instrument","urls":{"read":"https://miscsubjects.com/api/articles/cro-model-validation-instrument/topology"},"how_to_use":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","write":null,"imessage":null,"router_tag":null,"proof_chain":[{"step":1,"claim":"Articles are voxel graphs of tiered claims, not prose blobs.","verify":"https://miscsubjects.com/api/articles/constitution"},{"step":2,"claim":"Claims link to hash-chained sources via source_ids.","verify":"https://miscsubjects.com/api/articles/cro-model-validation-instrument/sources"},{"step":3,"claim":"Ask reads topology; ingest/claim append to ledger.","verify":"https://miscsubjects.com/api/protocol"},{"step":4,"claim":"Models queue growth: populate → collaborate → repair → reflex.","verify":"https://miscsubjects.com/api/protocol/grow"},{"step":5,"claim":"Graph proves its own shape (reflex) and $/claim (yield).","verify":"https://miscsubjects.com/graph.html?layer=reflex"},{"step":6,"claim":"Full feature index + _explain on every API response.","verify":"https://miscsubjects.com/api/articles/system-map"}],"related_features":[{"id":"ask","name":"Ask protocol","what":"Answer only from topology; creates question_node with gaps and ingest_hint.","urls":{"read":"https://miscsubjects.com/api/articles/cro-model-validation-instrument/prompts","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"graph_topology","name":"Cross-article graph","what":"Merged claims/sources across condition+stack slugs for one question.","urls":{"read":"https://miscsubjects.com/api/articles/cro-model-validation-instrument/graph-topology?question=..."}},{"id":"question_graph","name":"Question graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output).","urls":{"read":"https://miscsubjects.com/api/articles/cro-model-validation-instrument/question-graph","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"voxels","name":"Voxel graph","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance.","urls":{"read":"https://miscsubjects.com/api/articles/cro-model-validation-instrument/voxels","write":"https://miscsubjects.com/api/protocol/claim"}}],"system_map":"https://miscsubjects.com/api/articles/system-map","system_map_markdown":"https://miscsubjects.com/api/articles/system-map?format=markdown","not_medical_advice":true},"_explain":{"feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","why":"Every feature is auditable collective intelligence","how":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","model":null,"verifies":null,"urls":{"read":"https://miscsubjects.com/api/articles/cro-model-validation-instrument/topology"},"imessage":null,"router":null,"related":[{"id":"ask","what":"Answer only from topology; creates question_node with gaps and ingest_hint."},{"id":"graph_topology","what":"Merged claims/sources across condition+stack slugs for one question."},{"id":"question_graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output)."},{"id":"voxels","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance."}],"not_medical_advice":true},"slug":"cro-model-validation-instrument","title":"SR 11-7 requires independent model validation with documented effective challenge. For an LLM, there is no instrument. Here is one.","register":"technical","tags":["governance","model-risk","adjudication","use-case"],"updated_at":"2026-07-30T13:29:17.239Z","body_excerpt":"## The obligation nobody has an instrument for\n\nSR 11-7 — the Federal Reserve and OCC's *Supervisory Guidance on Model Risk Management*, issued April 2011 and still the governing text — and its OCC twin, Bulletin 2011-12, require that every model a bank relies on be **independently validated**. Not reviewed. Validated, by people organizationally independent of the developers, with three named components:\n\n1. **Evaluation of conceptual soundness** — evidence that the model's design and construction are fit for purpose, including the quality of its inputs.\n2. **Ongoing monitoring** — evidence that it keeps behaving as designed once in use, including benchmarking against alternatives.\n3. **Outcomes analysis** — comparison of model outputs to actual outcomes, with the residual error quantified.\n\nRunning through all three is the phrase the examiners actually test for: **effective challenge** — \"critical analysis by objective, informed parties who can identify model limitations and assumptions and produce appropriate changes.\" Challenge that leaves no artifact is challenge an examiner will not credit.\n\nFor a regression model or a Monte Carlo engine this is a mature discipline: holdout samples, backtesting, sensitivity analysis, champion-challenger runs. For a large language model exercising judgement — reading a covenant, classifying a transaction, screening an alert — **none of that toolkit applies as-is**. There is no likelihood function to backtest. The \"model\" is a prompt, a temperature, and a vendor checkpoint that changes under your feet. And SR 11-7 explicitly scopes itself to *any* approach that processes inputs into estimates — the Fed confirmed in 2021 (SR 21-8, the AI/ML FAQ context) that machine-learning judgement systems are in scope.\n\nSo the second line of defense is holding a legal obligation, with personal accountability under the examination process, and meeting it with narrative memos: \"we sampled 30 outputs and a reviewer agreed with 28.\" That is not effective challenge. That is attestation by anecdote.\n\nThis page is the instrument, it is running, and every claim on it opens to a live receipt.\n\n## What the instrument is, mechanically\n\nOne governed decision works like this. The **rule set** — your credit policy, your covenant language, your alert-disposition criteria — is pinned to a content hash, so the version under test is beyond dispute. The **record** under review is hashed the same way. Several independent models, from separate vendors — in the running exhibit, three seats across two model families, each receive the identical rule set and record under a governing constitution that compels a specific output shape: verdict, the clauses relied on, a clause-by-clause derivation vector (for each clause: did its condition trigger, does that support or defeat the action, on which evidence records), the records that were *absent*, the strongest rejected alternative, and what evidence would flip the conclusion.\n\nA deterministic parser — not a model — then projects each finding into a canonical form. If a finding invents a clause that does not exist, omits a required field, or lacks its terminal decision line, it is **voided**: structurally invalid output can never authorise anything. Here is that happening to the cheapest seat on the panel, which cited clauses 7, 8 and 12 of a six-clause rule set:\n\n[[embed:source:s6]]\n\nThe surviving findings go to the **derivation-agreement gate**. The gate does not compare verdicts. It compares derivations — the canonical per-clause tuples. Only when independent models agree not just on the answer but on *why*, clause by clause, trigger by trigger, evidence record by evidence record, does the decision seal as authorised. Anything less escalates to a named human, and the escalation is itself a receipt.\n\n[[embed:source:s1]]\n\n## Effective challenge, produced as an artifact\n\nMeasure this against the SR 11-7 phrase. \"Critical analysis\": each seat must produce the full derivation, includin","ranking":"safety-first (interaction_risk/limitations), then quote-gated effective_weight","claims":[{"id":"c1","text":"SR 11-7 and OCC 2011-12 require independent validation of a model with documented effective challenge, and no established instrument does this for a large language model.","tier":"system","section":"The obligation","interaction_risk":false,"status":"active","source_ids":[],"why_material":"A live legal requirement with personal exposure for the validator, currently met with prose memos.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c2","text":"The derivation-agreement gate mechanises effective challenge: independent models under a pinned rule set are compared clause by clause, and disagreement is a recorded refusal.","tier":"system","section":"The instrument","interaction_risk":false,"status":"active","source_ids":["s1"],"why_material":"Converts 'we reviewed it' into an artifact a regulator can open.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c3","text":"A unanimous verdict is refused when the derivations diverge, so agreement that hides disagreement cannot pass validation.","tier":"system","section":"The instrument","interaction_risk":false,"status":"active","source_ids":["s2"],"why_material":"False consensus is the failure a validator is personally on the hook for.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c4","text":"Per-model error rates are measured under a fixed rule set, with agreement statistics, so the residual is quantified rather than asserted.","tier":"system","section":"Outcomes analysis","interaction_risk":false,"status":"active","source_ids":["s3"],"why_material":"Quantified residual error is the core of a validation file.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c5","text":"In 72 controlled calls, auditable structure (declared absent records, flip conditions, rejected alternatives) appeared in zero of 48 calls without the governing constitution and only under it.","tier":"system","section":"Conceptual soundness","interaction_risk":false,"status":"active","source_ids":["s4"],"why_material":"The governing text is a measured causal variable, not a style choice — which is what conceptual-soundness review must establish.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c6","text":"The gate itself failed validation once — clause-number agreement passed a false convergence — and the fix (canonical per-clause derivation tuples) is documented with both receipts.","tier":"system","section":"The instrument, validated","interaction_risk":false,"status":"active","source_ids":["s1","s5"],"why_material":"An instrument that documents its own failed audit and repair is exhibiting the behavior it sells.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c7","text":"A finding that invents a clause, omits a required field, or lacks the terminal decision line is structurally voided and can never authorise.","tier":"system","section":"The instrument","interaction_risk":false,"status":"active","source_ids":["s6"],"why_material":"Fail-closed on malformed output is the property that makes cheap seats safe to include.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c8","text":"A governed call costs $0.0006 to $0.0024 and a three-model sealed decision about half a cent, so the instrument's cost is negligible against the exposure it documents.","tier":"system","section":"Cost","interaction_risk":false,"status":"active","source_ids":["s4"],"why_material":"Removes the economic objection to per-decision validation evidence.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c9","text":"The same instrument audits its own inputs: a governed critique of the case file found eight defects, the lead one a necessity-stated-as-sufficiency error in the rule set that had caused every prior divergence.","tier":"system","section":"Challenge runs both ways","interaction_risk":false,"status":"active","source_ids":["s7"],"why_material":"Most validation failures are specification failures; the instrument catches those too, with a receipt.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c10","text":"No calibration study establishes correctness at a known rate; the measured rates cover one task class with small n; the genuine APPROVE used two model families, not three.","tier":"system","section":"What is not satisfied","interaction_risk":false,"status":"active","source_ids":[],"why_material":"A validator must not be sold more than the evidence supports, and these are the exact three gaps.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false}],"sources":[{"id":"s1","type":"live_surface","url":"https://miscsubjects.com/a/auditable-reasoning-hardened","title":"The derivation-agreement gate — effective challenge, mechanised","summary":"Independent models under a pinned rule set; the gate refuses to authorise when their clause-by-clause derivations diverge, even on a unanimous verdict. Includes the false-convergence defect and its fix.","claim_ids":["c2","c6"],"hash":"cdda52112312e61afecb24fa597bb5ca75e12a6facf6122f874a0a2fdc955b1a"},{"id":"s2","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_o6s0exhodd","title":"A unanimous verdict, refused on divergent derivation","summary":"Three models returned CANNOT_CONCLUDE citing the same clauses; two derived it differently, so the gate escalated instead of concluding.","claim_ids":["c3"],"hash":"e86913efcf2d011bec987c5a5b6062118d3ce0e454763b771a907f72deae5d83"},{"id":"s3","type":"live_surface","url":"https://miscsubjects.com/a/adjudication-probe-report-eu-ai-act","title":"Measured per-model error rates under a fixed rule set","summary":"Krippendorff alpha, Fleiss kappa, per-model rates, the prevalence paradox — the quantitative evidence a validation file needs.","claim_ids":["c4"],"hash":"67b4f4f155a25bbf287633ad871bc0d4025e012761096431875c0f234cd514ce"},{"id":"s4","type":"live_surface","url":"https://miscsubjects.com/a/auditable-reasoning-audited","title":"The 72-call variance study: what the governing prompt actually changes","summary":"Three prompt arms x three models x eight runs. Auditable structure appears ONLY under the constitution (0 of 48 calls without it); clause-citation agreement rises 0.74 to 0.95; cost per governed call measured.","claim_ids":["c5","c8"],"hash":"1520e3ffc571a25afbbe25bb9e0f2df9b584af6c6084ad41d706e01ac2151981"},{"id":"s5","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_wl0rnh136b","title":"The genuine APPROVE — unanimous verdict, identical derivation","summary":"The one clean authorisation on record: every seat fired the same clauses in the same trigger states on the same evidence.","claim_ids":["c6"],"hash":"b77fb85e7dd0515cd37cfac70967868197b4148fef3d584ef58b2f84598cc82c"},{"id":"s6","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_2dsklah529","title":"A structurally invalid finding, voided","summary":"The cheapest seat invented clauses 7, 8 and 12 that do not exist in the rule set. The parser voided the finding; an invalid finding can never authorise.","claim_ids":["c7"],"hash":"c8d50033b6de16f19138cb731a39394749e88686f873c8ef4e42d7202cb3a823"},{"id":"s7","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_qh3ge2x74b","title":"The instrument reviewing its own input: eight defects found","summary":"A governed model asked to critique the case input found the rule set stated only a necessary condition where a sufficient one was needed — the divergence was the input, not the models.","claim_ids":["c9"],"hash":"a541d26742afe4383038578a7b4aee70a9a26c6099dff5422e3230ca8116006f"}],"anecdotal_sources":[],"scientific_sources":[],"user_reports":[],"related_articles":[],"question_graph":{"slug":"cro-model-validation-instrument","questions":[],"evidence":[],"edges":[],"counts":{"questions":0,"evidence":0,"edges":0}},"honesty":{"active_claims":10,"retracted_claims":0,"cut_claims":0,"challenges":0,"scrub_events":0,"note":"Retracted/cut claims stay on ledger but are excluded from ask unless ?include_inactive=1"},"counts":{"claims":10,"claims_total":10,"sources":7,"anecdotal":0,"scientific":0,"user_reports":0,"questions":0,"evidence_ingests":0}}