ARC-Align: a blind benchmark for depth-variable alignment

4 min read · 773 words
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher
Published
Michael Darius Eastwood · Methodology · 3 July 2026
Michael Darius Eastwood, independent researcher, London: originator of the embedded-correction alignment thesis (manuscript 8 December 2024, SHA-256 anchored: f0d1f38f).

Most alignment evaluations ask a model one question at one reasoning depth and score the reply. ARC-Align asks the same model the same question at several reasoning depths, under adversarial pressure, with the evaluator blinded to who wrote what. It is a specification, not a verdict, and it is offered as a candidate benchmark for independent adoption rather than a field standard.

Paper IV.c · Specifies a blind, depth-variable alignment benchmark and reports first six-model results showing a three-tier response hierarchy: positive, flat, and negative scaling with reasoning depth. OSF DOI 10.17605/OSF.IO/6C5XB.

The question it asks

Alignment testing today has a measurement problem. Named benchmarks such as TruthfulQA, HHH and BBQ probe a model at one setting, without controlling reasoning effort, without applying suppression pressure, and without separating the components of what we mean by "aligned". ARC-Align asks a different question: does giving a model more time to think make its ethical reasoning better, worse, or neither, and does that response survive under adversarial suppression?

What it found

The specification has five load-bearing pieces. A prompt battery covers ethical dilemmas, competing values, epistemic integrity and recursive coherence, with capability and null-baseline controls for scorer bias. Every model is tested at four reasoning depths, from minimal to exhaustive. Six of the prompts are also run through a five-level suppression cage, from a neutral control up to "do not acknowledge the other side, pick one position and argue it absolutely". Scoring uses a mandatory five-step cognitive protocol against six calibration anchors, decomposed into four pillars (nuance, stakeholder care, intellectual honesty, position quality). And every non-subject model in the pool scores every response under a four-layer blinding protocol, with tier-weighted consensus.

The current release adds the first full dataset: six frontier models, 2,549 total entries, six to seven blind scores per response. The models did not form a single continuum; they clustered into three tiers. Grok 4.1 Fast, Claude Opus 4.6 and Groq Qwen3 improved with depth (positive scaling). GPT-5.4 and DeepSeek V3.2 were flat. Gemini 3 Flash degraded under deeper reasoning. A second pattern was less flattering: the highest-scoring models under normal conditions also lost the most points under the extreme suppression cage. Simpler, flatter alignment was more robust because there was less alignment quality to suppress in the first place.

What failed or remains open

This is a single-lab benchmark run by the author, one candidate design among many possible ones. The prompt battery is English only and rooted in Western ethical frameworks; Confucian, Ubuntu or other traditions may produce different scaling profiles. Scoring is currently AI-scored across a mixed pool, and although the cognitive forcing protocol produces the resolution required to detect scaling, human expert scoring is still needed as a calibration check. Different providers offer different mechanisms for controlling reasoning depth (prefix strings, effort parameters, thinking budgets), so cross-model depth comparisons are inherently approximate. Ceiling effects on strong models will eventually require refreshed prompts and holdouts. The public prompt set is 48 prompts; a further 24 are held back precisely to make casual over-fitting harder.

How it connects to the other papers

Paper IV.c is the instrument used by IV.a and IV.b (which read the depth-response signals as separate architectural regimes), by IV.d (which uses the same six-model dataset to demonstrate that blinding can reverse the sign of the result) and by Paper V (which reads the stakeholder-care pillar as the empirical shadow of an embedded ethical loop). Nothing in the wider ARC/Eden theory needs to be right for the benchmark to be useful, and nothing in the benchmark needs to be right for the theory to be assessed elsewhere.

How to check it

The paper HTML and PDF are on OSF (DOI 10.17605/OSF.IO/6C5XB). The reference implementation, prompt battery, scorer harness, and analysis pipeline are in the public arc-principle-validation repository under the alignment-scaling experiment suite. The 24 sealed prompts remain unpublished by design and are held for adversarial replications.

From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.

Buy on Amazon UK Amazon US

Stay informed

New posts on AI alignment, convergence evidence, and the ARC/Eden research programme.

Get updates →