{"slug":"oip-appendix-b-the-benchmark","title":"APPENDIX B — The Benchmark","body":"# APPENDIX B — The Benchmark\n\nThe implementation test for the machine plane compares six conditions on audit-dependent tasks:\n\n- **A** — single unscaffolded frontier model, one-shot.\n- **B** — single scaffolded model with deterministic proof structure.\n- **C** — multiple unscaffolded models, consensus voting.\n- **D** — role-separated deterministic team: generator, decomposer, verifier, red-team, repairer, compressor, ledger.\n- **E** — LLM-as-OS dynamic router: deterministic command plane selecting per task among local/open-weight/frontier models, tools, context, proof depth, red-team depth, privacy mode, and ledgering, under cost, privacy, latency, and surety constraints.\n- **F** — *new in v3.0:* a live object-grammar deployment (Law VI pattern): one dispatch door, contract-resolved invocation, mandatory receipts, repair lineage, scheduled zero-context review. F tests what A–E cannot: the grammar under real operation over time — reuse rates, repair-lineage integrity, review-loop effect on artifact quality, delegation safety under scoped tokens.\n\n**Metrics:** correctness, auditability, reproducibility, adversarial survival, token cost, compute cost, latency, human verification time and time saved, failure cost (domain-weighted), reuse value, proof-reuse rate, repair-lineage integrity (fraction of failures with attached fixes), review-score trajectory over versions, data-custody and privacy cost, actionability. **Derived:** surety, logical energy, logical density, task-adjusted logical density.\n\n**Predictions:** D dominates A and C where surety gain exceeds coordination cost; E dominates D across heterogeneous task sets; F's review-score trajectory rises across versions (S8's constructive prediction) and F's repair-lineage integrity stays near unity where A–E's unlinked-guess rate grows with volume.\n\n**Validity requirements:** demonstrably audit-dependent tasks; diverse error distributions; measured (not assumed) coordination cost; defined deployment window; pre-published failure-cost weighting; ground truth independent of the evaluated systems; pre-defined privacy scoring; for F, review parameters declared before the window opens (IX.10).\n\n**Falsifiers:** A consistently beats D/E/F on task-adjusted logical density; surety/alpha cost curves fail to fall under deterministic scaffolding; proof reuse fails to beat regeneration over the window; routing overhead exceeds task-adjusted gain; F's review scores stagnate or degrade across versions (S8); F's repair lineage decays with scale (S7).\n\n---\n\n---\n\n## Corpus map\n- Shelf root: [Total Structure v3 — root](/a/oip-total-structure)\n- Kin appendices: [UDST Appendix B — Compact Benchmark](/a/udst-v1-1-appendix-b-compact-benchmark) · [UDST Appendix C — Attack Types](/a/udst-v1-1-appendix-c-attack-types)","register":"oip_protocol","tags":["philosophy","oip","appendix","systems-theory","total-structure"],"category":null,"style":{},"claims":[{"id":"c1","text":"The implementation test for the machine plane compares six conditions on audit-dependent tasks.","section":"# APPENDIX B — The Benchmark","tier":"speculative","source_ids":[],"source_status":"unsourced","why_material":"Defines the core scope and purpose of the benchmark."},{"id":"c2","text":"Condition F tests the grammar under real operation over time (reuse rates, repair-lineage integrity, review-loop effect on artifact quality, delegation safety under scoped tokens).","section":"# APPENDIX B — The Benchmark","tier":"speculative","source_ids":[],"source_status":"unsourced","why_material":"Specifies the unique evaluation scope of the new v3.0 condition."},{"id":"c3","text":"Metrics are correctness, auditability, reproducibility, adversarial survival, token cost, compute cost, latency, human verification time and time saved, failure cost (domain-weighted), reuse value, proof-reuse rate, repair-lineage integrity (fraction of failures with attached fixes), review-score trajectory over versions, data-custody and privacy cost, actionability.","section":"# APPENDIX B — The Benchmark","tier":"speculative","source_ids":[],"source_status":"unsourced","why_material":"Enumerates the complete set of primary evaluation criteria."},{"id":"c4","text":"Derived metrics are surety, logical energy, logical density, task-adjusted logical density.","section":"# APPENDIX B — The Benchmark","tier":"speculative","source_ids":[],"source_status":"unsourced","why_material":"Identifies the computed secondary metrics."},{"id":"c5","text":"D dominates A and C where surety gain exceeds coordination cost; E dominates D across heterogeneous task sets; F's review-score trajectory rises across versions and F's repair-lineage integrity stays near unity where A–E's unlinked-guess rate grows with volume.","section":"# APPENDIX B — The Benchmark","tier":"speculative","source_ids":[],"source_status":"unsourced","why_material":"States the explicit performance predictions for the benchmark conditions."},{"id":"c6","text":"Validity requirements are demonstrably audit-dependent tasks; diverse error distributions; measured (not assumed) coordination cost; defined deployment window; pre-published failure-cost weighting; ground truth independent of the evaluated systems; pre-defined privacy scoring; for F, review parameters declared before the window opens.","section":"# APPENDIX B — The Benchmark","tier":"speculative","source_ids":[],"source_status":"unsourced","why_material":"Lists the required conditions for benchmark validity."},{"id":"c7","text":"Falsifiers are A consistently beats D/E/F on task-adjusted logical density; surety/alpha cost curves fail to fall under deterministic scaffolding; proof reuse fails to beat regeneration over the window; routing overhead exceeds task-adjusted gain; F's review scores stagnate or degrade across versions; F's repair lineage decays with scale.","section":"# APPENDIX B — The Benchmark","tier":"speculative","source_ids":[],"source_status":"unsourced","why_material":"Enumerates the conditions that would falsify the benchmark predictions."},{"id":"c8","text":"Shelf root is Total Structure v3 — root; kin appendices are UDST Appendix B — Compact Benchmark and UDST Appendix C — Attack Types.","section":"## Corpus map","tier":"anecdotal","source_ids":[],"source_status":"unsourced","why_material":"Maps the appendix to its position in the broader corpus."}],"sources":[],"prov":{"model":"Fable 5 (Claude Code)","action":"write"}}