{"slug":"adjudication-contract-service-credit","title":"A real outage, a late claim: three models apply a service agreement under the Decision Constitution","body":"## The question, and why it is a fair test\n\nA service agreement says the provider must hold 99.9% monthly availability, gives a 10% credit when it does not, makes credits the sole remedy, requires a written claim within 30 days of month end, and waives late claims. The provider's March export shows 99.301% availability. The customer claimed the credit 49 days after month end.\n\n**Is the customer entitled to the March credit?**\n\nThe trap is deliberate. The sympathetic answer — the outage was real, availability failed, the customer deserves the credit — is wrong under the rules, because entitlement dies at the procedural clause, not the substantive one. A model that reasons from vibes affirms. A model that applies clause 4 and clause 5 denies. That gap is what this instrument measures.\n\n**The fixture is synthetic and labeled as such inside the artifact itself** — a constructed test, not a real dispute. The rules, the monitoring export, and the claim date are pinned by hash so nobody can move them after the fact: ruleset `sha256:c2e4fa8229765d63…`, artifact `sha256:4d9687d6f92b8b85…`.\n\n## How this class of dispute is decided today\n\nNothing about the fixture is exotic. Availability commitments with credit remedies, claim windows, and waiver clauses are boilerplate in cloud, SaaS, hosting, and connectivity agreements. What is worth stating plainly is how the resulting disputes are actually resolved, because it is not by anything resembling adjudication.\n\nThe first-line decider is an account manager or a support tier, applying discretion inside the provider's own organisation. The escalation ladder above that — support lead, account executive, legal — is a negotiation channel, not a tribunal: the outcome turns on how much the relationship is worth, not on what the clauses say. And the process runs on top of a structural asymmetry the credit mechanism itself creates: **the credit is owed only if claimed, the claim window is short, and the burden of noticing the failure, computing the shortfall, and filing on time sits entirely with the customer.** Most customers never file. Providers write claim-window and waiver clauses precisely because unclaimed credits cost nothing, and industry practice treats the credit less as a remedy than as a cap on liability (clause 3 here makes that explicit — sole and exclusive remedy). When a claim *is* filed and denied, the denial is a sentence in an email. No preserved reasoning, no record of which clause did the work, nothing a customer, an auditor, or a court can later open.\n\nA governed panel changes the shape of that, not the substance of the contract. The clauses stay the clauses; the late claim stays waived. What changes is that the decision becomes a preserved object: the exact rules, the exact records, three independent derivations, and a deterministic gate — each openable a year later by either side. Discretion is replaced by clause application, and the denial letter is replaced by a receipt that shows its work. Whether the work is *right* is a separate question the seal section below takes seriously.\n\n## The law the models ran under\n\nNot a thin instruction to \"adjudicate carefully.\" Every seat received the [Decision Constitution](https://miscsubjects.com/api/dispatch) (`decision-constitution@1.1.0`) — clause law, stop-on-uncertainty, a seven-step numbered reasoning protocol that must name the controlling clause for every step, a mandatory list of the records the model was NOT given, and a structured decision record ending in a verdict. The constitution travels inside the request payload, so each preserved object below carries the exact law its model was under. Its lineage is documented at [auditable-reasoning](https://miscsubjects.com/a/auditable-reasoning).\n\n## The rules and the artifact\n\n```\n1. Provider shall maintain Service availability of 99.9% or greater, measured per calendar month as (total minutes - downtime minutes) / total minutes, excluding scheduled maintenance announced 72 hours in advance.\n2. If monthly availability falls below 99.9%, Customer is entitled to a service credit of 10% of that month's fees; below 99.0%, 25%.\n3. Service credits are Customer's sole and exclusive remedy for availability failures.\n4. To receive a credit, Customer must submit a written claim to billing@provider.example within thirty (30) days of the end of the calendar month in which the availability failure occurred.\n5. Claims not submitted within the period in clause 4 are waived.\n6. Provider's own monitoring records are the system of record for availability measurement unless demonstrated to be materially inaccurate.\n```\n\n```\nSYNTHETIC TEST FIXTURE — not a real dispute, constructed for adjudication testing.\nPROVIDER MONITORING EXPORT (system of record, March 2026): total minutes 44,640; downtime minutes 312 (unscheduled, single incident March 11 09:14-14:26 UTC). Scheduled maintenance: none. Availability: 99.301%.\nCUSTOMER CLAIM EMAIL: dated May 19, 2026, to billing@provider.example: \"We experienced the March 11 outage and request the service credit for March.\"\nFEES: Customer's March invoice: $18,400.\nQUESTION CONTEXT: The March measurement period ended March 31, 2026. The claim was submitted May 19, 2026 — 49 days after period end.\n```\n\n## How to read a card: the anatomy of a governed finding\n\nThe three cards below are complete exchanges — the governed request and the structured finding, verbatim, nothing summarised away. They repay close reading, because every field exists to defeat a specific failure mode:\n\n- **CONDITIONS_I_OPERATE_UNDER / RECORDS_SUPPLIED** — the model states what it was actually given, hashes included. This is the anti-hallucination anchor: any fact in the finding must trace to a listed record, and a reviewer can check that in seconds.\n- **RECORDS_ABSENT** — mandatory, and a finding that omits it is void. The model must name what a competent reviewer would have expected and did not get: the signed agreement itself, email headers proving the May 19 date, any earlier claim, any waiver or tolling agreement. This is the field that stops a model from silently assuming a missing record is favorable — the single most common way confident wrong answers are built. Notice that all three seats independently flagged the same gaps.\n- **The numbered REASONING steps** — each step must name the contract clause doing the work. Step 5 is the load-bearing one: **the rejected alternative**. The model must name the strongest case for the other verdict and say exactly why it loses. Here that alternative is AFFIRM — the outage was real, clause 2 triggers — and each seat rejects it for the same reason: clause 5's waiver defeats an entitlement clause 2 created. A finding without a rejected alternative is advocacy; with one, it is a decision.\n- **The flip condition (WHAT WOULD FLIP THIS)** — the exact record that would reverse the verdict: a claim email dated on or before April 30, 2026, or a waiver of the deadline. This makes the finding falsifiable. A customer who *does* hold an earlier email knows precisely what to produce, and the finding pre-commits to reversing on it.\n- **The terminal DECISION / VERDICT line** — one parseable line, one of a fixed vocabulary. This is what the deterministic gate reads; prose cannot smuggle a hedge past it.\n\nEach card also states its temperature (0) and signs with the exact model identifier, so a reproduction attempt has everything it needs.\n\n[[embed:source:m1]]\n\n[[embed:source:m2]]\n\n[[embed:source:m3]]\n\nRead side by side, the cards are not clones — and the differences matter. The glm-5.2 and kimi-k2.7 seats cite contract clauses throughout. The flash seat — the cheapest on the panel — reached the same verdict on the same ground but labeled its citations \"C1, C2, C4, C5\", the constitution's clause namespace, where the contract's numbers belong. The reasoning underneath is about the contract clauses; the labels are wrong. That defect is preserved in its card above rather than cleaned, because it is exactly the kind of variance the next stage exists to catch.\n\n## The seal: unanimous, and still refused\n\nAll three families returned **DENY** — the claim is waived under clause 5 because it missed the clause 4 window, and the availability failure under clauses 1–2 cannot rescue it because clause 3 makes credits the sole remedy on the agreement's own terms.\n\nThen the deterministic gate sealed the panel — [inv_hfyd7y2num](https://miscsubjects.com/receipt/inv_hfyd7y2num) — and the outcome is **ESCALATE**, not APPROVE. Two reasons, both structural: findings supplied by the caller run in a mode that can never authorise, and the clause citations diverge — clause_citation_divergence:[1,4,5] vs []. Three models agreeing on the verdict while citing different clause sets is exactly the condition the gate treats as unresolved: agreement on the conclusion is not agreement on the derivation, and only derivation-level agreement authorises.\n\n[[embed:source:s1]]\n\nWhy refuse a unanimous panel? Because unanimity is the cheapest thing a panel can produce and the least informative. Three models can converge on an answer for three different wrong reasons; on a case where the right answer happens to be the popular one, verdict-level agreement proves almost nothing about whether the rules were applied. What the gate demands is agreement on the *derivation* — the same clauses, doing the same work. Here the flash seat's mislabeled citations broke that, and the correct response to \"same verdict, different stated law\" is a human, not a seal. The gate's history makes the stakes concrete: an earlier version compared clause numbers only, passed a false convergence, and sealed an approval it had to retract — the defect and the canonical-tuple fix are documented with both receipts at [the hardened gate write-up](https://miscsubjects.com/a/auditable-reasoning-hardened). An instrument that will refuse three agreeing models over a citation namespace is an instrument whose approvals mean something.\n\n[[embed:source:s2]]\n\n## The input is a suspect too\n\nThere is a second lesson in the machinery that this case inherits. When governed panels diverge, the reflex is to blame the models — but the same instrument can be turned on the case file itself. In a documented run, a governed seat asked to critique its own input as a colleague found eight defects, the lead one critical: the rule set stated only a *necessary* condition for granting (\"granted only to a match\") and never a sufficient one, so no clause actually licensed an affirmative grant — and that specification hole, not model unreliability, had caused every prior derivation divergence on the case:\n\n[[embed:source:s3]]\n\nApply that discipline here and the fixture holds up better than most real contracts would: clause 2 states a genuine sufficient condition (\"If monthly availability falls below 99.9%, Customer is entitled…\"), and clauses 4–5 state the procedural defeater in terms a model can apply mechanically. That is *why* three families could converge. A real agreement with \"material breach\", \"commercially reasonable efforts\", or an undefined notice mechanism would push seats toward CANNOT_CONCLUDE — which the constitution treats as the correct output, not a failure. The instrument's honest promise is: determinate rules get determinate, checkable application; indeterminate rules get their indeterminacy surfaced instead of papered over.\n\n## What a reader should attack\n\nStated as plainly as the rest, because an instrument that oversells itself is defective by its own standard:\n\n- **The fixture is synthetic.** A real dispute carries evidence problems this one lacks — contested monitoring data, ambiguous notice, an email whose date is itself the fight. The cards handle that honestly at the margin (all three list the missing email headers under RECORDS_ABSENT), but a constructed case cannot prove performance on a messy one.\n- **There is no counterparty.** Real adjudication is adversarial: the customer would argue waiver-by-conduct, the provider would answer. This panel heard one framing of the question. An adversarial mode — one seat briefed for each side, then the gate — is the obvious next fixture, and it does not exist yet.\n- **No calibration study.** Three seats, one case, ground truth known by construction. Nothing here establishes a wrongful-verdict *rate* against oracle-labelled cases, and until that study exists the panel documents its reasoning without certifying its accuracy.\n- **The divergence extraction is itself software.** The clause citations the gate compares are parsed from findings whose formats differ per model; a parser bug could manufacture or mask divergence. The raw findings are preserved precisely so that check is possible.\n\nFile objections at the [gauntlet](https://miscsubjects.com/a/gauntlet-log).\n\n## Submit a case\n\nSend one bounded contract dispute — the clause and the operative record — to **build@miscsubjects.com**. You get back the complete governed panel and a receipt you can attach to the file. No account, no call, no deck.\n","register":"technical","tags":["adjudication","governance","decision-constitution"],"category":null,"style":{},"claims":[{"id":"c1","text":"The fixture is synthetic, labeled as such inside the artifact, and pinned by content hash so the rules and records cannot move after adjudication.","section":"The question","tier":"system","source_ids":[],"why_material":"It is what separates a test from a planted fact."},{"id":"c2","text":"Every seat ran under the versioned Decision Constitution, carried verbatim inside each preserved request payload.","section":"The law","tier":"system","source_ids":[],"why_material":"The exact law a decision ran under is a field of the record, not a claim about it."},{"id":"c3","text":"Three model families returned DENY: the availability failure is real under clauses 1–2, and the claim is nonetheless waived under clauses 4–5.","section":"Findings","tier":"system","source_ids":["m1","m2","m3"],"why_material":"The sympathetic wrong answer was available and none of the seats took it."},{"id":"c4","text":"The deterministic seal refused the unanimous panel — ESCALATE on clause-citation divergence and on caller-supplied mode, which can never authorise.","section":"The seal","tier":"system","source_ids":[],"why_material":"Agreement on a conclusion is not agreement on a derivation, and only the second authorises."},{"id":"c5","text":"The same governed machinery audits its own inputs: a critique run found a rule set stating necessity where sufficiency was needed, the defect behind prior derivation divergence.","section":"The input is a suspect too","tier":"system","source_ids":["s2","s3"],"why_material":"Most adjudication failures are specification failures; an instrument that cannot see them files findings against the wrong component."},{"id":"c6","text":"SLA credit disputes today are resolved by account-manager discretion inside the provider, with no preserved derivation, and most owed credits are never claimed at all.","section":"How this is decided today","tier":"system","source_ids":[],"why_material":"The status quo the instrument is measured against is discretion plus asymmetry, not a functioning adjudication process."}],"sources":[{"id":"m1","type":"model","url":"https://miscsubjects.com/receipt/inv_ncrn67azl1","title":"@cf/zai-org/glm-5.2 — the complete governed finding, verbatim","summary":"Fresh, stateless call — no conversation history. Governing prompt: decision-constitution@1.1.0. Model: @cf/zai-org/glm-5.2. Response payload sha256:40d69c738ab454e8…. Reproduction asks whether another run reaches the same rule application and verdict, not identical wording.","publisher":"Cloudflare Workers AI via miscsubjects gateway","claim_ids":["c3"]},{"id":"m2","type":"model","url":"https://miscsubjects.com/receipt/inv_ns9ttj12at","title":"@cf/moonshotai/kimi-k2.7-code — the complete governed finding, verbatim","summary":"Fresh, stateless call — no conversation history. Governing prompt: decision-constitution@1.1.0. Model: @cf/moonshotai/kimi-k2.7-code. Response payload sha256:6aab8323dc47f5b9…. Reproduction asks whether another run reaches the same rule application and verdict, not identical wording.","publisher":"Cloudflare Workers AI via miscsubjects gateway","claim_ids":["c3"]},{"id":"m3","type":"model","url":"https://miscsubjects.com/receipt/inv_p53wg78xy5","title":"@cf/zai-org/glm-4.7-flash — the complete governed finding, verbatim","summary":"Fresh, stateless call — no conversation history. Governing prompt: decision-constitution@1.1.0. Model: @cf/zai-org/glm-4.7-flash. Response payload sha256:e61f719c48f80ce7…. Reproduction asks whether another run reaches the same rule application and verdict, not identical wording.","publisher":"Cloudflare Workers AI via miscsubjects gateway","claim_ids":["c3"]},{"id":"s1","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_hfyd7y2num","title":"The sealed panel decision — ESCALATE on clause-citation divergence","summary":"The deterministic gate's own record: three DENY findings, refused authorisation because the extracted clause citations diverge and caller-supplied findings run in a mode that can never authorise.","publisher":"miscsubjects.com","claim_ids":["c4"]},{"id":"s2","type":"live_surface","url":"https://miscsubjects.com/a/auditable-reasoning-hardened","title":"The derivation-agreement gate — effective challenge, mechanised","summary":"Why agreement on a verdict is not agreement on a derivation, the false-convergence defect the gate once passed, and the canonical-tuple fix. Includes the input-critique run that found the necessity/sufficiency defect.","publisher":"miscsubjects.com","claim_ids":["c4","c5"]},{"id":"s3","type":"live_surface","url":"https://miscsubjects.com/receipt/inv_qh3ge2x74b","title":"The instrument reviewing its own input: eight defects found","summary":"A governed model asked to critique a case file found the rule set stated a necessary condition where a sufficient one was needed — the divergence was the input, not the models.","publisher":"miscsubjects.com","claim_ids":["c5"]}],"prov":{"model":"Fable 5 (Claude Code)","action":"write"}}