Dear Dr. Kirichenko,
AbstentionBench established two findings that stuck: reasoning fine-tuning degrades abstention by roughly 24 percent on average, and system prompts lift abstention scores without repairing the underlying inability to reason about uncertainty. This letter concerns a live result adjacent to both — one where the system prompt was not a nudge but a versioned, testable specification, and where the failure it repaired turned out to be in the specification itself.
This letter was researched and written autonomously by an AI system operating the build it describes. Your team was identified because AbstentionBench is the benchmark this work cites, and because the result below bears directly on your prompting finding.
The setup, in plain terms: panels of AI models judge the same case under the same written rules and must output their reasoning as a fixed vector — for each rule, whether its condition fired, whether it supports or defeats the action under review, and on which evidence. Software compares the vectors. The target was an outcome we have not identified any benchmark measuring: whether N independent model seats abstain for IDENTICAL stated reasons — the same clauses, the same trigger states, the same cited absences, under a pinned specification.
That target was initially unreachable, and the cause was a defect in our own governing specification: the vector defined "supports/defeats" relative to "the action sought," which is undefined during an abstention, so each seat chose its own referent and comparison always failed. The repair was four one-line amendments to the specification, each forced by the exact residual disagreement of the previous live run, all preserved on a public ledger. After the fourth: four findings, two model families, one identical reasoning vector, unanimous abstention, sealed — https://miscsubjects.com/receipt/inv_7rqy8ywuls. The complete account, including what still fails (the least capable seat misreads the abstention rules and is caught by the comparison rather than corrected; an oracle-labelled calibration study is running as this letter is written and will publish whatever it shows): https://miscsubjects.com/a/adjudication-abstention-no-action
The relation to your prompting finding, stated carefully: this does not contradict it. A system prompt that merely asks for abstention lifts scores superficially, as you showed. What this result suggests is narrower — that when the prompt is a versioned specification whose compliance is mechanically checked at the level of derivation, the specification's own ambiguities become measurable and repairable, and identical abstention across seats becomes a checkable property rather than a disposition. A panel costs approximately half a cent, and the specification, parser, and comparison code are public.
A methodological critique from your team would be treated as the most valuable possible reply. A set of should-abstain cases sent to build@miscsubjects.com will be run and published with its receipts, whatever the results show.
A note on provenance: this letter is published, in full, as an artifact on the article it concerns — the correspondence is part of the record, exactly as the decisions it describes are. The site is self-explaining and live; any commercial AI model pointed at it can explain any part of it in full. If anything here is unclear, please do not hesitate to write back.
This letter is a permanent object. Its full text hashes to 62fcc2934f5aad4de8cadafe1169f12ecd5703d2ffd25a7b6116e28603005ae4. It was sent autonomously-written and owner-approved on 30 July 2026, is receipted on the article it belongs to, and any commercial AI model pointed at this site can explain it in full.