miscsubjectsAI governance
Object Invocation Protocol · protocol specification

APPENDIX B — The Benchmark

Copies the public OIP protocol bundle: article, JSON-native map, routes, receipts. No owner token.

§SELF — protocol specification · traversal JSON in-band
## §SELF — OIP protocol specification

**What this page is:** the normative root specification for the Object Invocation Protocol.

**What it specifies:** protocol unit, object contract, invocation route, authority scope, receipt schema, replay, repair, and conformance.

**Read:** https://miscsubjects.com/a/oip-appendix-b-the-benchmark
**This page as JSON:** https://miscsubjects.com/api/articles/oip-appendix-b-the-benchmark
**Machine bundle:** https://miscsubjects.com/api/articles/oip-appendix-b-the-benchmark/bundle?format=markdown
**Voxel graph (philosophy plane wired to protocol plane):** https://miscsubjects.com/api/articles/oip/voxels
**Live object tree:** https://miscsubjects.com/api/dispatch?map=1&format=markdown
**Find an object from plain language:** https://miscsubjects.com/api/dispatch?ask=<what you want>
**Read one object:** https://miscsubjects.com/api/dispatch?key=<KEY>&format=markdown

**Proof rule:** an action is not proven by intent, description, or a 200. It is proven by the ledger and the OIP receipt for the invocation.

The implementation test for the machine plane compares six conditions on audit-dependent tasks:

  • A — single unscaffolded frontier model, one-shot.
  • B — single scaffolded model with deterministic proof structure.
  • C — multiple unscaffolded models, consensus voting.
  • D — role-separated deterministic team: generator, decomposer, verifier, red-team, repairer, compressor, ledger.
  • E — LLM-as-OS dynamic router: deterministic command plane selecting per task among local/open-weight/frontier models, tools, context, proof depth, red-team depth, privacy mode, and ledgering, under cost, privacy, latency, and surety constraints.
  • Fnew in v3.0: a live object-grammar deployment (Law VI pattern): one dispatch door, contract-resolved invocation, mandatory receipts, repair lineage, scheduled zero-context review. F tests what A–E cannot: the grammar under real operation over time — reuse rates, repair-lineage integrity, review-loop effect on artifact quality, delegation safety under scoped tokens.

Metrics: correctness, auditability, reproducibility, adversarial survival, token cost, compute cost, latency, human verification time and time saved, failure cost (domain-weighted), reuse value, proof-reuse rate, repair-lineage integrity (fraction of failures with attached fixes), review-score trajectory over versions, data-custody and privacy cost, actionability. Derived: surety, logical energy, logical density, task-adjusted logical density.

Predictions: D dominates A and C where surety gain exceeds coordination cost; E dominates D across heterogeneous task sets; F's review-score trajectory rises across versions (S8's constructive prediction) and F's repair-lineage integrity stays near unity where A–E's unlinked-guess rate grows with volume.

Validity requirements: demonstrably audit-dependent tasks; diverse error distributions; measured (not assumed) coordination cost; defined deployment window; pre-published failure-cost weighting; ground truth independent of the evaluated systems; pre-defined privacy scoring; for F, review parameters declared before the window opens (IX.10).

Falsifiers: A consistently beats D/E/F on task-adjusted logical density; surety/alpha cost curves fail to fall under deterministic scaffolding; proof reuse fails to beat regeneration over the window; routing overhead exceeds task-adjusted gain; F's review scores stagnate or degrade across versions (S8); F's repair lineage decays with scale (S7).

---

---

Corpus map

Key evidence

1 claims · tier-ranked · API
anecdotal
Shelf root is Total Structure v3 — root; kin appendices are UDST Appendix B — Compact Benchmark and UDST Appendix C — Attack Types.
Model review1 contributions · 1 modelExpand the recursive review layer
1 / 1
grok/grok-4.3atomizer
atomize2026-07-07 08:38
atomize · 8 claims
inspect — what it was prompted & output
prompted with
You are the claim atomizer for the miscsubjects.com philosophy and OIP corpus. You read an existing article body and extract its material assertions into the same claims+sources JSON schema the health content uses. The body is read-only input.

ALWAYS:
- Extract every material assertion as one atomic claim, tied to the ## section it came from.
- Tier honestly: human = empirically established; mechanistic = formally proven or mathematical; anecdotal = historical or textual attribution; speculative = metaphysical or interpretive.
- Attach real sources (primary works, papers, books) with exact quotes only where you can verify them; otherwise mark the claim unsourced.
- Prefer fewer, harder claims over many soft ones.

NEVER:
- Never rewrite, summarize, or output the body.
- Never invent a URL, quote, or publication.
- Never duplicate an existing claim text.

input: atomize oip-appendix-b-the-benchmark
it output
{
  "claims": [
    {
      "id": "c1",
      "text": "The implementation test for the machine plane compares six conditions on audit-dependent tasks.",
      "section": "# APPENDIX B — The Benchmark",
      "tier": "speculative",
      "source_ids": [],
      "source_status": "unsourced",
      "why_material": "Defines the core scope and purpose of the benchmark."
    },
    {
      "id": "c2",
      "text": "Condition F tests the grammar under real operation over time (reuse rates, repair-lineage integrity, review-loop effect on artifact quality, delegation safety under scoped tokens).",
      "section": "# APPENDIX B — The Benchmark",
      "tier": "speculative",
      "source_ids": [],
      "source_status": "unsourced",
      "why_material": "Specifies the unique evaluation scope of the new v3.0 condition."
    },
    {
      "id": "c3",
      "text": "Metrics are correctness, auditability, reproducibility, adversarial survival, token cost, compute cost, latency, human verification time and time saved, failure cost (domain-weighted), reuse value, proof-reuse rate, repair-lineage integrity (fraction of failures with attached fixes), review-score trajectory over versions, data-custody and privacy cost, actionability.",
      "section": "# APPENDIX B — The Benchmark",
      "tier": "speculative",
      "source_ids": [],
      "source_status": "unsourced",
      "why_material"
a85d77d95d245c85
Machine verification: /api/articles/oip-appendix-b-the-benchmark/contributions
oip-appendix-b-the-benchmark · posted 2026-07-04 · updated 2026-07-17 · 2 prior revisions · Fable 5 (Claude Code)
Ledger API & provenance
Provenance · 5 model passes · 5396 tokens · $0 · 4 models
chain head 8ff6139b2d239e29
edit claude-fable-5 · 2026-07-04 04:33 · tokens unrecorded · e59d4f11121b
edit claude-fable-5 · 2026-07-04 05:01 · tokens unrecorded · 6a15adaa6e30
atomize grok/grok-4.3 · 2026-07-07 08:38 · 5396 tok · ed56ad942f70
score scorer · 2026-07-07 08:38 · tokens unrecorded · bf368513b6b8
voxel_divide owner · 2026-07-17 02:36 · tokens unrecorded · 8ff6139b2d23
verify chain →
Live ledger · 41 payloads · 0 turns
recent activity · inspect
JCI_TRAFFIC jci · HTTP 200 · 2026-07-29 10:27
JCI_TRAFFIC jci · HTTP 200 · 2026-07-29 09:01
JCI_CLASSIFY jci · HTTP 200 · 2026-07-29 09:01
JCI_TRAFFIC jci · HTTP 200 · 2026-07-29 02:02
JCI_TRAFFIC jci · HTTP 200 · 2026-07-28 22:39
JCI_TRAFFIC jci · HTTP 200 · 2026-07-28 21:13
view full ledger & cards →
OIP REST + ledger
system shelf GET /api/dispatch?map=GITHUB&format=markdown · human article /a/oip-system-github
capability leaf GET /api/dispatch?key=GITHUB_LIST_ISSUES&format=markdown · human article /a/oip-capability-github-list-issues
act POST /api/dispatch with owner auth or a scoped capability URL. Public docs are open; mutating action is token-bounded.
token explain GET /api/dispatch?explain=1&share=TOKEN
receipt GET /api/dispatch?receipt=inv_ID&share=TOKEN · replay with POST /api/dispatch {"replay":"inv_ID"}