{"slug":"oip-v3-appendix-b","title":"Total Structure v3: Appendix B","body":"# APPENDIX B — The Benchmark\n\nThe implementation test for the machine plane compares six conditions on audit-dependent tasks:\n\n- **A** — single unscaffolded frontier model, one-shot.\n- **B** — single scaffolded model with deterministic proof structure.\n- **C** — multiple unscaffolded models, consensus voting.\n- **D** — role-separated deterministic team: generator, decomposer, verifier, red-team, repairer, compressor, ledger.\n- **E** — LLM-as-OS dynamic router: deterministic command plane selecting per task among local/open-weight/frontier models, tools, context, proof depth, red-team depth, privacy mode, and ledgering, under cost, privacy, latency, and surety constraints.\n- **F** — *new in v3.0:* a live object-grammar deployment (Law VI pattern): one dispatch door, contract-resolved invocation, mandatory receipts, repair lineage, scheduled zero-context review. F tests what A–E cannot: the grammar under real operation over time — reuse rates, repair-lineage integrity, review-loop effect on artifact quality, delegation safety under scoped tokens.\n\n**Metrics:** correctness, auditability, reproducibility, adversarial survival, token cost, compute cost, latency, human verification time and time saved, failure cost (domain-weighted), reuse value, proof-reuse rate, repair-lineage integrity (fraction of failures with attached fixes), review-score trajectory over versions, data-custody and privacy cost, actionability. **Derived:** surety, logical energy, logical density, task-adjusted logical density.\n\n**Predictions:** D dominates A and C where surety gain exceeds coordination cost; E dominates D across heterogeneous task sets; F's review-score trajectory rises across versions (S8's constructive prediction) and F's repair-lineage integrity stays near unity where A–E's unlinked-guess rate grows with volume.\n\n**Validity requirements:** demonstrably audit-dependent tasks; diverse error distributions; measured (not assumed) coordination cost; defined deployment window; pre-published failure-cost weighting; ground truth independent of the evaluated systems; pre-defined privacy scoring; for F, review parameters declared before the window opens (IX.10).\n\n**Falsifiers:** A consistently beats D/E/F on task-adjusted logical density; surety/alpha cost curves fail to fall under deterministic scaffolding; proof reuse fails to beat regeneration over the window; routing overhead exceeds task-adjusted gain; F's review scores stagnate or degrade across versions (S8); F's repair lineage decays with scale (S7).\n\n---\n\n---\n\n## Corpus map\n- Canonical shelf: [Total Structure root](/a/oip-total-structure)","register":"oip_protocol","tags":["philosophy","oip","appendix","total-structure","systems-theory"],"category":null,"style":{},"claims":[{"id":"c1","text":"Commercial vendors and clinics market this compound (1 commercial/clinic sources catalogued) — marketing material, not evidence.","section":"who_claims_what","tier":"speculative","source_ids":["s1"],"source_status":"sourced","why_material":"Collapsed 4 duplicate marketing claims"}],"sources":[{"id":"s1","type":"adjacent","url":"https://miscsubjects.com/a/oip-v3-appendix-b","title":"Total Structure v3: Appendix B","quote":"The implementation test for the machine plane compares six conditions on audit-dependent tasks:","summary":"Primary source for Total Structure v3: Appendix B.","claim_ids":["c1"]}],"prov":{"model":"Fable 5 (Claude Code)","action":"write"}}