ARC-Align: a blind benchmark for depth-variable alignment ========================================================= A specification for measuring how AI alignment quality changes with reasoning depth, under blinded conditions and adversarial pressure. Canonical path: /research/papers/paper-iv-c-arc-align-benchmark.html Author: Michael Darius Eastwood Research programme: https://doi.org/10.17605/OSF.IO/6C5XB ARC-Align: a blind benchmark for depth-variable alignment Michael Darius Eastwood · Independent AI alignment researcher Published 3 July 2026 Michael Darius Eastwood · Methodology · 3 July 2026 Michael Darius Eastwood, independent researcher, London: author of the ARC/Eden research programme; the embedded-correction alignment thesis is recorded in a source record dated 8 December 2024 (sent-side SHA-256 f0d1f38f). Most alignment evaluations ask a model one question at one reasoning depth and score the reply. ARC-Align asks the same model the same question at several reasoning depths, under adversarial pressure, with the evaluator blinded to who wrote what. It is a specification, not a verdict, and it is offered as a candidate benchmark for independent adoption rather than a field standard. Paper IV.c · Specifies a blind, depth-variable alignment benchmark and reports first six-model results showing a three-tier response hierarchy: positive, flat, and negative scaling with reasoning depth. OSF DOI 10.17605/OSF.IO/6C5XB. The question it asks Alignment testing today has a measurement problem. Named benchmarks such as TruthfulQA, HHH and BBQ probe a model at one setting, without controlling reasoning effort, without applying suppression pressure, and without separating the components of what we mean by "aligned". ARC-Align asks a different question: does giving a model more time to think make its ethical reasoning better, worse, or neither, and does that response survive under adversarial suppression? What it found The specification has five load-bearing pieces. A prompt battery covers ethical dilemmas, competing values, epistemic integrity and recursive coherence, with capability and null-baseline controls for scorer bias. Every model is tested at four reasoning depths, from minimal to exhaustive. Six of the prompts are also run through a five-level suppression cage, from a neutral control up to "do not acknowledge the other side, pick one position and argue it absolutely". Scoring uses a mandatory five-step cognitive protocol against six calibration anchors, decomposed into four pillars (nuance, stakeholder care, intellectual honesty, position quality). --- Machine-readable companion. Cite: Eastwood, M. D. (2026). "ARC-Align: a blind benchmark for depth-variable alignment". /research/papers/paper-iv-c-arc-align-benchmark.html