{"_self":{"principle":"Self-explaining payload — no external context required. This _self block describes what you are reading and where to look next.","widget":"article_topology","feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","contains":"claims, sources, anecdotes, question_graph slice","slug":"peer-review-derivation-record","urls":{"read":"https://miscsubjects.com/api/articles/peer-review-derivation-record/topology"},"how_to_use":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","write":null,"imessage":null,"router_tag":null,"proof_chain":[{"step":1,"claim":"Articles are voxel graphs of tiered claims, not prose blobs.","verify":"https://miscsubjects.com/api/articles/constitution"},{"step":2,"claim":"Claims link to hash-chained sources via source_ids.","verify":"https://miscsubjects.com/api/articles/peer-review-derivation-record/sources"},{"step":3,"claim":"Ask reads topology; ingest/claim append to ledger.","verify":"https://miscsubjects.com/api/protocol"},{"step":4,"claim":"Models queue growth: populate → collaborate → repair → reflex.","verify":"https://miscsubjects.com/api/protocol/grow"},{"step":5,"claim":"Graph proves its own shape (reflex) and $/claim (yield).","verify":"https://miscsubjects.com/graph.html?layer=reflex"},{"step":6,"claim":"Full feature index + _explain on every API response.","verify":"https://miscsubjects.com/api/articles/system-map"}],"related_features":[{"id":"ask","name":"Ask protocol","what":"Answer only from topology; creates question_node with gaps and ingest_hint.","urls":{"read":"https://miscsubjects.com/api/articles/peer-review-derivation-record/prompts","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"graph_topology","name":"Cross-article graph","what":"Merged claims/sources across condition+stack slugs for one question.","urls":{"read":"https://miscsubjects.com/api/articles/peer-review-derivation-record/graph-topology?question=..."}},{"id":"question_graph","name":"Question graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output).","urls":{"read":"https://miscsubjects.com/api/articles/peer-review-derivation-record/question-graph","write":"https://miscsubjects.com/api/protocol/ask"}},{"id":"voxels","name":"Voxel graph","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance.","urls":{"read":"https://miscsubjects.com/api/articles/peer-review-derivation-record/voxels","write":"https://miscsubjects.com/api/protocol/claim"}}],"system_map":"https://miscsubjects.com/api/articles/system-map","system_map_markdown":"https://miscsubjects.com/api/articles/system-map?format=markdown","not_medical_advice":true},"_explain":{"feature":"topology","name":"Article topology","what":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","why":"Every feature is auditable collective intelligence","how":"Claims, sources, anecdotes, user reports, related embeds, question graph slice — for ask/ROUTER.","model":null,"verifies":null,"urls":{"read":"https://miscsubjects.com/api/articles/peer-review-derivation-record/topology"},"imessage":null,"router":null,"related":[{"id":"ask","what":"Answer only from topology; creates question_node with gaps and ingest_hint."},{"id":"graph_topology","what":"Merged claims/sources across condition+stack slugs for one question."},{"id":"question_graph","what":"Ask nodes (questions + gaps) and evidence_ingest nodes (pasted model output)."},{"id":"voxels","what":"Claims as atoms, sources as edges (supported_by, posted_by). Per-claim provenance."}],"not_medical_advice":true},"slug":"peer-review-derivation-record","title":"Two NeurIPS committees disagreed on a quarter of the same papers, twice. Nothing records why. Here is a review format in which disagreement is a comparable record.","register":"technical","tags":["peer-review","meta-science","auditable-reasoning","use-case"],"updated_at":"2026-07-30T15:02:11.919Z","body_excerpt":"## The defect is measured, famous, and unrepaired\n\nPeer review's central weakness is not a suspicion. It is one of the best-measured facts about scientific publishing, measured by the field most capable of measuring it, on itself, twice.\n\nIn 2014 the NIPS programme chairs — Corinna Cortes and Neil Lawrence — ran an experiment no journal editor has been able to un-know since: they routed 10% of submissions through **two independent programme committees**, each unaware of the duplication, each applying the same review form, the same criteria, the same accept/reject decision. The committees disagreed on **25.9% of the duplicated papers**. Because the acceptance rate was about 22.5%, that arithmetic has a sharper reading: **roughly half to 57% of the papers one committee accepted were rejected by the other**. Acceptance at the field's flagship venue was, for the marginal paper, closer to a coin flip than to a measurement.\n\nThe natural hope was that this was a 2014 problem — a growing field, stretched reviewers. So NeurIPS ran it again in 2021, at ten times the scale: 882 duplicated papers, two committees, the same design. The result: **committees disagreed on 23% of duplicated papers, and about half of the papers accepted by one committee were rejected by the other.** Seven years, an order of magnitude more data, an entire reform literature in between — and the arbitrariness did not move.\n\nEvery load-bearing number in the two preceding paragraphs is the organizers' own, and both write-ups are public:\n\n[[embed:source:s1]]\n\n[[embed:source:s2]]\n\n## What the experiments could not see\n\nRead the two experiments carefully and notice what they measure: **how often** reviewers disagree. Not **why**. They could not measure why, because the review record does not contain the why in any comparable form.\n\nA review, as every venue currently collects it, is prose plus scores. Two reviews of the same manuscript can reach opposite recommendations, and the record offers no way to determine whether they disagreed about the same thing — whether one reviewer read the ablation as missing while the other read it as present; whether both applied the reproducibility criterion and reached different trigger states, or one never applied it at all; whether the disagreement is about the manuscript or about what the criterion means. The scores are comparable and empty; the prose is substantive and incomparable.\n\nSo the field's most famous defect sits exactly where its records are weakest. Reviewer disagreement is visible only as a binary outcome — accept here, reject there — and everything upstream of that outcome, the derivation, evaporates into paragraphs no machine and few humans can align. Score recalibration, better forms, reviewer training, open review: every proposed reform operates on the outcome layer or the prose layer. None of them produces the artifact that would let an editor say *these two reviewers applied criterion 4 to the same section and derived opposite trigger states* — which is the sentence that would make the disagreement tractable.\n\n## The governed format, applied to the checkable slice\n\nThis site runs a decision format built for exactly that missing artifact, and this page states precisely how far it reaches into peer review — which is a bounded distance, stated now and again at the end.\n\nA manuscript review has two components that current practice fuses. One is **judgement**: is this novel, is it significant, is it interesting. That is not a rule application, and nothing on this page touches it. The other is **checkable**: does the paper report what the venue requires reported — the criteria a venue already publishes as checklists. Are all claims in the abstract supported by evidence in the body? Are the baselines the ones the venue's policy names? Is the data availability statement present and does it match what the paper actually uses? Are limitations stated? Is the statistical reporting complete — n, variance, test named? This slice","ranking":"safety-first (interaction_risk/limitations), then quote-gated effective_weight","claims":[{"id":"c1","text":"In the 2014 NIPS consistency experiment, two independent committees disagreed on 25.9% of duplicated submissions — meaning roughly half or more of the papers one committee accepted were rejected by the other — and the 2021 repeat at ten times the scale found essentially the same arbitrariness (23% disagreement, about half of accepted papers rejected by the other committee).","tier":"system","section":"The measured defect","interaction_risk":false,"status":"active","source_ids":["s1","s2"],"why_material":"The field's own organizers measured the defect twice, seven years apart, and it did not move.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c2","text":"The consistency experiments measured how often committees disagree, but the review record itself does not capture why they disagreed at any comparable level — the disagreement is visible only as a binary outcome, never as a divergence between stated derivations.","tier":"system","section":"The measured defect","interaction_risk":false,"status":"active","source_ids":["s2"],"why_material":"The missing artifact is the WHY, and no reform of scores or forms produces it.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c3","text":"A venue's checkable criteria — completeness of reporting, claims-versus-evidence structure, required disclosures — can be pinned to a content hash as a rule set, and a manuscript's checkable properties recorded against it, without touching novelty or significance.","tier":"system","section":"The governed format applied","interaction_risk":false,"status":"active","source_ids":[],"why_material":"Scopes the claim to the slice where rule application is honest.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c4","text":"Under the governing constitution, each reviewing seat must emit a fixed machine-readable derivation: per criterion, whether its condition triggered, whether it supports or defeats acceptance of the checked property, and on which evidence in the manuscript — and a deterministic parser voids any finding that invents a criterion or omits a required field.","tier":"system","section":"The governed format applied","interaction_risk":false,"status":"active","source_ids":["s3"],"why_material":"Machine-comparable review findings require a compelled shape, not reviewer goodwill.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c5","text":"The derivation-agreement gate compares derivations, not verdicts: two reviews that reach the same recommendation for different stated reasons are recorded as divergent, and the divergence itself becomes the artifact.","tier":"system","section":"Disagreement becomes a record","interaction_risk":false,"status":"active","source_ids":["s3"],"why_material":"Converts NeurIPS-style noise into an inspectable object.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c6","text":"Both halves of the phenomenon exist as live receipts: a unanimous verdict refused because two seats derived it differently, and a genuine seal where every seat fired the same criteria in the same trigger states on the same evidence.","tier":"system","section":"Disagreement becomes a record","interaction_risk":false,"status":"active","source_ids":["s4","s5"],"why_material":"The peer-review problem in miniature, already on the record — same verdict, different reasons, caught mechanically.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c7","text":"In a 30-case oracle-labelled calibration study on synthetic determinate fixtures, the strongest seat matched the oracle 30/30 and the second 29/30, across two model families, with zero wrongful authorisations at the gate in 30 cases.","tier":"system","section":"Calibration","interaction_risk":false,"status":"active","source_ids":["s6"],"why_material":"The only accuracy numbers this page is entitled to, with their scope stated.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c8","text":"Abstention is a first-class sealed outcome: when the criteria license no conclusion, the system records a NO_ACTION rather than forcing a verdict, and the abstention is a receipt.","tier":"system","section":"Calibration","interaction_risk":false,"status":"active","source_ids":["s7"],"why_material":"Review of a manuscript outside the rule set's competence must terminate in a recorded abstention, not a guess.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c9","text":"The same machinery audits the criteria themselves: a governed seat asked to critique a case file found eight defects in the rule set, the lead one a necessity-stated-as-sufficiency error that had caused every prior derivation divergence on that case.","tier":"system","section":"The criteria are also under review","interaction_risk":false,"status":"active","source_ids":["s8"],"why_material":"Much reviewer disagreement is criterion ambiguity; an instrument that cannot distinguish the two writes findings against the wrong component.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false},{"id":"c10","text":"Scientific merit judgement — novelty, significance, interestingness — is not a rule application and is out of scope; this instrument covers only the checkable slice, has not been run on real submissions, and its calibration numbers come from synthetic determinate fixtures only.","tier":"system","section":"What this does not cover","interaction_risk":false,"status":"active","source_ids":[],"why_material":"An instrument for honest review that oversold itself would be defective by its own standard.","retracted_at":null,"retraction_reason":null,"challenged_by":[],"effective_weight":0.1,"quote_gated":false}],"sources":[{"id":"s1","type":"web","url":"https://arxiv.org/abs/2109.09774","title":"Inconsistency in Conference Peer Review: Revisiting the 2014 NeurIPS Experiment","summary":"The 2014 NIPS organizers routed 10% of submissions through two independent programme committees. The committees disagreed on 25.9% of the duplicated papers; given the ~22.5% acceptance rate, roughly half to 57% of the papers one committee accepted were rejected by the other.","claim_ids":["c1"],"hash":"1ed02e56742d84112581175d464fac0712b198d887674988b8c78c67e1bd63f6"},{"id":"s2","type":"web","url":"https://arxiv.org/abs/2306.03262","title":"The NeurIPS 2021 Consistency Experiment","summary":"The 2014 experiment repeated at ~10x scale: 882 duplicated papers through two committees. Disagreement on 23% of duplicated papers; about half of the papers accepted by one committee were rejected by the other. The arbitrariness did not improve in seven years.","claim_ids":["c1","c2"],"hash":"2158e2ee091c544683970cb9801d95b7babe63dda366f5ff5c3f9eb818476d0e"},{"id":"s3","type":"live_surface","url":"https://miscsubjects.com/a/auditable-reasoning-hardened","title":"The derivation-agreement gate — effective challenge, mechanised","summary":"Independent models under a pinned rule set; a deterministic parser projects each finding into canonical per-clause derivation tuples; the gate refuses to conclude when derivations diverge, even on a unanimous verdict.","claim_ids":["c4","c5"],"hash":"ea845b0d71fec49eac1fcb62e143a9182a848254326cb580534e46ff3558709c"},{"id":"s4","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_o6s0exhodd","title":"Same verdict, different derivations — the refusal receipt","summary":"Three seats returned the same verdict citing the same clauses; two derived it through different trigger states; the gate escalated instead of concluding. The exhibit: agreement inspected at the level of reasoning and found hollow.","claim_ids":["c6"],"hash":"c3c29fdae7deea4c947519013a1a0a161690fbd5e034c7ebfe8224ccfcdcbf37"},{"id":"s5","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_wl0rnh136b","title":"The genuine authorisation — identical derivations","summary":"The clean seal on record: every seat fired the same clauses in the same trigger states on the same evidence records.","claim_ids":["c6"],"hash":"adab34b512f10cbe2b61a171d6521e37748b8f757bf889d8dae7e268afb8747d"},{"id":"s6","type":"live_surface","url":"https://miscsubjects.com/a/adjudication-calibration-study","title":"The calibration study — 30 oracle-labelled cases through the production gate","summary":"Seat accuracy on synthetic determinate fixtures: glm-5.2 30/30, kimi-k2.7 29/30; zero wrongful authorisations at the gate across all 30 cases; escalation counted as deferral cost, not decision error.","claim_ids":["c7"],"hash":"3f4143a29c252767625b3134c25854479795fbed9a04bd0736e84600ef12bf3b"},{"id":"s7","type":"live_surface","url":"https://miscsubjects.com/a/adjudication-abstention-no-action","title":"Abstention as a sealed outcome","summary":"The first clean NO_ACTION: a rule set that licenses no action produces a sealed abstention, with the spec-defect arc (four amendments) that got there. Receipt inv_7rqy8ywuls.","claim_ids":["c8"],"hash":"414dd02fbf061779e42472c615483663bc2be49989a3de8d68fa56cb352c61f6"},{"id":"s8","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_qh3ge2x74b","title":"The instrument reviewing its own input — eight defects found","summary":"A governed seat asked to critique the case file found the rule set stated a necessary condition where a sufficient one was needed — the divergence was the input, not the reviewers.","claim_ids":["c9"],"hash":"1c6d86d73f93530a532b4bef3f51767cdd568a2047a1c9e5bc326fff37084983"}],"anecdotal_sources":[],"scientific_sources":[],"user_reports":[],"related_articles":[],"question_graph":{"slug":"peer-review-derivation-record","questions":[],"evidence":[],"edges":[],"counts":{"questions":0,"evidence":0,"edges":0}},"honesty":{"active_claims":10,"retracted_claims":0,"cut_claims":0,"challenges":0,"scrub_events":0,"note":"Retracted/cut claims stay on ledger but are excluded from ask unless ?include_inactive=1"},"counts":{"claims":10,"claims_total":10,"sources":8,"anecdotal":0,"scientific":0,"user_reports":0,"questions":0,"evidence_ingests":0}}