{"_self":{"principle":"Self-explaining payload — no external context required. This _self block describes what you are reading and where to look next.","widget":"article_topology","feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","contains":"claims, sources, anecdotes, question_graph slice","slug":"nist-ai-rmf-measure-reference","urls":{"read":"https://miscsubjects.com/api/articles/nist-ai-rmf-measure-reference/topology"},"how_to_use":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","write":null,"imessage":null,"router_tag":null,"proof_chain":[{"step":1,"claim":"Articles are voxel graphs of tiered claims, not prose blobs.","verify":"https://miscsubjects.com/api/articles/constitution"},{"step":2,"claim":"Claims link to hash-chained sources via source_ids.","verify":"https://miscsubjects.com/api/articles/nist-ai-rmf-measure-reference/sources"},{"step":3,"claim":"Ask reads topology; ingest/claim append to ledger.","verify":"https://miscsubjects.com/api/protocol"},{"step":4,"claim":"Models queue growth: populate → collaborate → repair → reflex.","verify":"https://miscsubjects.com/api/protocol/grow"},{"step":5,"claim":"Graph proves its own shape (reflex) and $/claim (yield).","verify":"https://miscsubjects.com/graph.html?layer=reflex"},{"step":6,"claim":"Full feature index + _explain on every API response.","verify":"https://miscsubjects.com/api/articles/system-map"}],"related_features":[{"id":"ask","name":"Ask protocol","what":"Answer only from topology; creates question_node with gaps and ingest_hint.","urls":{"read":"https://miscsubjects.com/api/articles/nist-ai-rmf-measure-reference/prompts","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"graph_topology","name":"Cross-article graph","what":"Merged claims/sources across condition+stack slugs for one question.","urls":{"read":"https://miscsubjects.com/api/articles/nist-ai-rmf-measure-reference/graph-topology?question=..."}},{"id":"question_graph","name":"Question graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output).","urls":{"read":"https://miscsubjects.com/api/articles/nist-ai-rmf-measure-reference/question-graph","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"voxels","name":"Voxel graph","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance.","urls":{"read":"https://miscsubjects.com/api/articles/nist-ai-rmf-measure-reference/voxels","write":"https://miscsubjects.com/api/protocol/claim"}}],"system_map":"https://miscsubjects.com/api/articles/system-map","system_map_markdown":"https://miscsubjects.com/api/articles/system-map?format=markdown","not_medical_advice":true},"_explain":{"feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","why":"Every feature is auditable collective intelligence","how":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","model":null,"verifies":null,"urls":{"read":"https://miscsubjects.com/api/articles/nist-ai-rmf-measure-reference/topology"},"imessage":null,"router":null,"related":[{"id":"ask","what":"Answer only from topology; creates question_node with gaps and ingest_hint."},{"id":"graph_topology","what":"Merged claims/sources across condition+stack slugs for one question."},{"id":"question_graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output)."},{"id":"voxels","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance."}],"not_medical_advice":true},"slug":"nist-ai-rmf-measure-reference","title":"NIST AI RMF's MEASURE function describes what to measure. Nothing runnable exists to point at. Here is a candidate reference implementation.","register":"technical","tags":["nist-ai-rmf","iso-42001","measure","reference-implementation","ai-governance","calibration"],"updated_at":"2026-07-30T14:44:20.008Z","body_excerpt":"## The gap between a framework and a mechanism\n\nNIST's *Artificial Intelligence Risk Management Framework* (AI RMF 1.0, NIST AI 100-1, January 2023) organises the discipline into four functions: **GOVERN**, **MAP**, **MEASURE**, **MANAGE**. It is voluntary by design, and its MEASURE function is the load-bearing one — \"quantitative, qualitative, or mixed-method tools, techniques, and methodologies to analyze, assess, benchmark, and monitor AI risk.\" The Generative AI Profile (NIST AI 600-1, July 2024) extends the same functions to generative systems. ISO/IEC 42001:2023 does the certifiable version of the same move: clause 9 requires an organization to determine what will be monitored and measured, the methods for monitoring, measurement, analysis and evaluation, and to *retain documented information as evidence of the results*.\n\nBoth documents are careful, considered, and correct. Both ship as prose. Neither ships a runnable mechanism. MEASURE tells you that AI systems should be evaluated for trustworthy characteristics with documented, repeatable methods; it cannot show you one executing. Clause 9 tells you to retain evidence of measurement results; it cannot show you what such evidence looks like when it is produced per decision rather than per audit cycle. So every implementer performs the same private translation — framework prose into bespoke internal process — and every certification audit reviews the translation, not a mechanism. There is no reference implementation to point at, diff against, or attack.\n\nThis page offers one. Not for the whole of MEASURE — for a specific slice: the measurement of model judgement under a governing rule set. It is running now, every element below opens to a live exhibit, and the closing section states exactly what it does not satisfy. The property being claimed is narrow and unusual: **a standards author can point at this rather than describe it.**\n\n## The candidate, element by element\n\n**Versioned governing law at a content hash.** The rule set a decision is judged under — and the constitution compelling the output shape — are pinned to content hashes, so the version under test is beyond dispute. This is MEASURE's precondition stated as an artifact: you cannot measure a system's behaviour against criteria unless the criteria are frozen. And the governing text is not asserted to matter — its effect is measured. A 72-call controlled study ran three prompt arms across three models, eight runs each: auditable structure (declared-absent records, flip conditions, rejected alternatives) appeared in **zero of 48 calls** without the constitution, and only under it; clause-citation agreement rose from 0.74 to 0.95.\n\n[[embed:source:s5]]\n\n**Machine-comparable per-seat reasoning.** Each model seat is compelled into a canonical form: verdict, the clauses relied on, a clause-by-clause derivation vector (did the clause trigger, does it support or defeat the action, on which evidence records), the records that were *absent*, the strongest rejected alternative, and the finding that would flip the conclusion. A deterministic parser voids anything malformed — a finding that invents a clause ([here is one citing clauses 7, 8 and 12 of a six-clause rule set](/receipt/inv_2dsklah529)) can never authorise. The point for a measurement regime: free-text rationales are not comparable units. Canonical derivation tuples are. Disagreement between independent evaluators becomes something you compute, not something a committee characterises.\n\n**A deterministic agreement gate with four sealed outcomes.** The surviving findings go to a gate that is code, not a model. It compares derivations — not verdicts — and seals exactly one of four outcomes: authorise, negate, abstain, or escalate to a named human. The finite vocabulary matters to a framework author because it makes the mechanism itself auditable: there is no fifth outcome, no silent pass. The sharpest exhibit is [a unanimous verdict the gate refused](/receipt/inv_o6","ranking":"safety-first (interaction_risk/limitations), then quote-gated effective_weight","claims":[{"id":"c1","text":"NIST AI RMF 1.0 is a voluntary framework whose MEASURE function calls for quantitative and qualitative methods to analyze, assess, benchmark, and monitor AI risks, but it ships as prose: it specifies what to measure, not a runnable mechanism that measures it.","tier":"system","section":"The gap","interaction_risk":false,"status":"active","source_ids":["s1"],"why_material":"The absence of a reference implementation is the premise; if MEASURE shipped one, this page would be redundant.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c2","text":"ISO/IEC 42001 clause 9 requires organizations to determine measurement methods and retain documented evidence of results, and certification audits accept process documentation because no executable reference exists to point at.","tier":"system","section":"The gap","interaction_risk":false,"status":"active","source_ids":["s2"],"why_material":"The same specification-without-mechanism gap, in the certifiable standard.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c3","text":"The governing law of each decision is a versioned text pinned to a content hash, and a 72-call controlled study measured its causal effect: auditable structure appeared in zero of 48 ungoverned calls and only under the constitution.","tier":"system","section":"The candidate, element by element","interaction_risk":false,"status":"active","source_ids":["s5"],"why_material":"Versioned, hash-pinned governing law with a measured effect is the first thing a MEASURE implementer needs and the first thing prose frameworks cannot supply.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c4","text":"Per-seat reasoning is compelled into a canonical machine-comparable form — verdict, clauses relied on, per-clause derivation tuples, declared-absent records, rejected alternative, flip condition — so disagreement is computable rather than narrated.","tier":"system","section":"The candidate, element by element","interaction_risk":false,"status":"active","source_ids":["s3"],"why_material":"Measurement requires comparable units; free-text rationales are not comparable units.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c5","text":"A deterministic gate — not a model — compares the canonical derivations and seals exactly one of four outcomes: authorise, negate, abstain, or escalate to a named human, and the refusals are receipts too.","tier":"system","section":"The candidate, element by element","interaction_risk":false,"status":"active","source_ids":["s4"],"why_material":"A finite outcome vocabulary with fail-closed refusal is what makes the mechanism auditable as a mechanism.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c6","text":"The gate's own failure is on the record: its first version passed a false convergence on clause numbers, sealed an APPROVE, was caught, retracted, and fixed to compare full derivation tuples — with both the defective and the genuine seal public.","tier":"system","section":"The candidate, element by element","interaction_risk":false,"status":"active","source_ids":["s4","s7"],"why_material":"A measurement instrument that documents its own failed audit exhibits the property the framework asks for.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c7","text":"Every decision — including refusals and abstentions — emits a permanent public receipt carrying the complete request and response payloads and the content hashes it was bound to.","tier":"system","section":"The candidate, element by element","interaction_risk":false,"status":"active","source_ids":["s7","s8"],"why_material":"Retained documented evidence of results, the ISO 42001 clause 9 requirement, produced per decision rather than per audit cycle.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c8","text":"A 30-case oracle-labelled calibration study ran the production gate end to end: glm-5.2 scored 30/30, kimi-k2.7-code 29/30, and the gate sealed zero wrongful authorisations in 30 cases, with escalation counted as deferral cost, not hidden.","tier":"system","section":"The candidate, element by element","interaction_risk":false,"status":"active","source_ids":["s3"],"why_material":"A wrongful-authorisation rate against oracle labels is the exact quantitative artifact MEASURE describes in prose.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c9","text":"The decision record is already mapped element by element against FRE 902, ISA 705, EU AI Act Articles 12 and 14, NIST, ISO 42001, IEC 61508, and Toulmin, with each mapping's failures stated alongside it.","tier":"system","section":"Offered for testing, not claimed as satisfied","interaction_risk":false,"status":"active","source_ids":["s6"],"why_material":"The mapping is the artifact a standards author would start from — and it names its own gaps.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c10","text":"This is a candidate reference implementation, not a conformant one: self-declared conformance is worthless, the calibration corpus is synthetic and single task class, and the panel spans two model families, not three.","tier":"system","section":"Offered for testing, not claimed as satisfied","interaction_risk":false,"status":"active","source_ids":[],"why_material":"A framework body must not be sold more than the evidence supports; these are the exact limits of what is on the record.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false}],"sources":[{"id":"s1","type":"standard","url":"https://www.nist.gov/itl/ai-risk-management-framework","title":"Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1","summary":"The voluntary framework, January 2023: four functions — GOVERN, MAP, MEASURE, MANAGE. MEASURE covers employing quantitative and qualitative methods to analyze, assess, benchmark, and monitor AI risk; the Generative AI Profile (NIST AI 600-1, July 2024) is its first cross-sectoral profile.","claim_ids":["c1"],"hash":"75df671bc20d5712256cccc4dd66a6f6f099b388eb726acccb48841d9cf2cbef"},{"id":"s2","type":"standard","url":"https://www.iso.org/standard/42001","title":"ISO/IEC 42001:2023 — Artificial intelligence management system","summary":"The certifiable AI management-system standard. Clause 9 requires the organization to determine what needs to be monitored and measured, the methods for monitoring, measurement, analysis and evaluation, and to retain documented information as evidence of the results.","claim_ids":["c2"],"hash":"36fd7f6399fef831bad48e72982a839c9f4a5a6cd3b65e47d9a805f188a19400"},{"id":"s3","type":"live_surface","url":"https://miscsubjects.com/a/adjudication-calibration-study","title":"Calibration, measured: 30 oracle-labelled cases through the production gate","summary":"Three seats across two model families under decision-constitution@1.3.3 on 30 hashed, oracle-labelled synthetic cases: glm-5.2 30/30, kimi-k2.7-code 29/30, zero wrongful authorisations at the gate. Seat calibration and gate calibration answered separately, every case a receipt.","claim_ids":["c4","c8"],"hash":"948bf81f45ffef0061e786620c02ea4d48ae21df39f5bfe1bba3236064090700"},{"id":"s4","type":"live_surface","url":"https://miscsubjects.com/a/auditable-reasoning-hardened","title":"The gate compares derivations, not citations","summary":"The derivation-agreement gate: independent seats under a pinned rule set, compared clause by clause; four sealed outcomes; the false-convergence defect it caught in itself, with both receipts.","claim_ids":["c5","c6"],"hash":"00fc9d6bd081337d28f498fe179611844d6156d22c7a3696a0ac16bb141fb4a5"},{"id":"s5","type":"live_surface","url":"https://miscsubjects.com/a/auditable-reasoning-audited","title":"The 72-call variance study: what the governing prompt actually changes","summary":"Three prompt arms x three models x eight runs. Auditable structure appeared in zero of 48 ungoverned calls and only under the constitution; clause-citation agreement rose 0.74 to 0.95. The governing text is a measured causal variable.","claim_ids":["c3"],"hash":"58feae6aceed965866d1f1d2519273f662c2a055da83be572c830317fe9975c1"},{"id":"s6","type":"live_surface","url":"https://miscsubjects.com/a/attested-finding-conformance-map","title":"Every primitive mapped to its frame","summary":"The attested finding mapped element by element against FRE 902, ISA 705, EU AI Act Articles 12 and 14, NIST, ISO 42001, IEC 61508, and Toulmin — including what each mapping fails.","claim_ids":["c9"],"hash":"c7f3ff6774f30a281587134d05aa01a1221006bb2c195e198093a2648e45361b"},{"id":"s7","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_wl0rnh136b","title":"A genuine APPROVE: unanimous verdict, identical derivation","summary":"The sealed authorisation: every seat fired the same clauses in the same trigger states on the same evidence, bound to the case hashes.","claim_ids":["c6","c7"],"hash":"3be45b11862d778fead744c78bcb11f98d8c7620a9d037276a5d832bd73d6f77"},{"id":"s8","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_7rqy8ywuls","title":"The first clean NO_ACTION: abstention as a sealed outcome","summary":"A record deliberately absent, a manifest naming the absence, and a panel sealing abstention rather than guessing — the outcome class most measurement regimes cannot even represent.","claim_ids":["c7"],"hash":"f5ad437300df719ff02c4157a2006337420407e79d71af44daa9e3b6e5037cc2"},{"id":"em_es_d83908a2604b492a86a9","type":"email","url":"https://miscsubjects.com/letter-nist-2026-07-30","title":"Letter to Elham Tabassi — 2026-07-30","claim_ids":[],"hash":"aa34070190579ef6ab83a4038c2aad495de45f98c37c43dafcb5f177d7641dcf"}],"anecdotal_sources":[],"scientific_sources":[],"user_reports":[],"related_articles":[],"question_graph":{"slug":"nist-ai-rmf-measure-reference","questions":[],"evidence":[],"edges":[],"counts":{"questions":0,"evidence":0,"edges":0}},"honesty":{"active_claims":10,"retracted_claims":0,"cut_claims":0,"challenges":0,"scrub_events":0,"note":"Retracted/cut claims stay on ledger but are excluded from ask unless ?include_inactive=1"},"counts":{"claims":10,"claims_total":10,"sources":9,"anecdotal":0,"scientific":0,"user_reports":0,"questions":0,"evidence_ingests":0}}