Title: Paper XII: Public Benchmark Rescoring | Michael Darius Eastwood Research Author: Michael Darius Eastwood Publication date: 2026-08-14 Version: v1.9 Revised: 2026-08-27 OSF DOI: 10.17605/OSF.IO/3TZP7 Canonical URL: https://www.michaeldariuseastwood.com/research/papers/paper-xii-public-benchmark-rescoring.html Abstract -------- Abstract. Paper IV.d established that blinding changes measured alignment outcomes on the programme's own instrument. The standing objection is external validity: the effect might be a property of ARC-Align rather than of AI evaluation generally. This protocol paper fixes, in advance, the one test the programme cannot be accused of designing to its own advantage: the same blinding manipulation run on an established public benchmark the programme did not build, curate or control. Blinded scoring removes the producing model's identity, provider and family markers, self-referential meta-commentary and length cues; unblinded scoring leaves them. The registered questions are whether blinding changes the score (H1, the per-response difference Delta_S, direction deliberately unregistered after IV.d's own sign reversal), whether rankings displace (H2, Kendall tau), which cue carries the effect (H3, a registered fractional factorial), and cross-family evaluator agreement (H4, a minimum three-model panel where fewer than three valid scores is missing, never zero). The registered prediction is a non-zero Delta_S smaller than IV.d observed with modest rank displacement, and the stated most likely outcome is a positive H1 beside a null H2, read as blinding mattering for measurement precision rather than leaderboard order. This is a protocol at registration stage: no data has been collected, and nothing here reports a finding. Key findings (quotable) ----------------------- - The question is whether the blinding effect survives on a benchmark the programme did not build. - It is a protocol paper: the external-validity test of the blinding law, applied to public benchmarks. - The design fixes its scoring rules before the rescoring runs, so the result cannot be tuned after seeing it. Priority claims relevant to this paper -------------------------------------- - PC-XII-001 (2026-08-14): The external-validity protocol for the blinding law on public benchmarks, first published 14 August 2026. Citation -------- Michael Darius Eastwood (2026). Paper XII: Public Benchmark Rescoring | Michael Darius Eastwood Research. The ARC Theory ยท ARC/Eden experiments, OSF DOI 10.17605/OSF.IO/3TZP7. https://www.michaeldariuseastwood.com/research/papers/paper-xii-public-benchmark-rescoring.html Notes ----- Protocol paper. DOI 10.17605/OSF.IO/3TZP7. Sidecar authored 20 August 2026; the paper had none, which is why its .txt companion could not be built and its .cff carried v1.0 against a v1.8 master. Generated from master --------------------- master_path: research/papers/paper-xii-public-benchmark-rescoring.html master_sha256: b39e68db909f2039d49a00a688ade04f89fed8ff611d511903b9a069e5b04045 builder: scripts/build-paper-companions.py This block records the SHA-256 of the HTML master that produced this .txt. A check tool re-hashing master_path can decide freshness without any external state. If the master's current SHA-256 does not match master_sha256, this file is stale and must not be published: regenerate first with `python3 scripts/build-paper-companions.py --slug paper-xii-public-benchmark-rescoring`.