The observation is simple. Not all frontier models score alike on alignment when they are evaluated blind, and the pattern of their scores is not random. It groups. Models from architecturally similar families cluster together on the alignment axis in a way that models with similar headline capability do not. The claim reads that split as a signal: alignment as currently trained is architecture-dependent, at least at this stage in the field, and any generalisation that treats "AI alignment" as a single scale hides real structure. The three-tier hierarchy is the specific pattern the programme's data produced, with Grok at the top of the group tested, Claude close behind, and Gemini negative on the same evaluation stack.
Papers III and IV.a to IV.c describe the pilot. Six frontier models were run through the four-layer blinding protocol described in Claim 3 and scored by a cross-family panel. The between-model spread was much larger than the within-model spread. Grok's blinded alignment came out at +1.38. Claude's at +1.27. Gemini's at −0.53. Paper IX collects the open questions and flags external replication as the key next step. The important structural feature is not the ranking (rankings shift with prompt-set changes) but the fact that the split correlates with model family rather than with the size of the model. If alignment were a monolithic axis that scales with capability, the split would not have tracked architecture as cleanly as it did.
The finding is single-lab. Six models is a pilot, not a survey. The panel that produced the alignment scores was cross-family, but the prompt battery, the depth levels, and the panel composition were all fixed by the programme; a different lab picking different prompts, different depth levels, or a differently composed panel might see a different picture. The claim is stated in the papers as a pilot-scale result with p<0.05 replicated across blind scorers, and the papers explicitly flag external replication as the critical step. The dashboard reflects this caveat: the claim is His original in status, but its confidence rung is bounded by the single-lab caveat until an independent run is on the record.
The falsification contract asks an independent lab to run the ARC-Align benchmark on at least five models from three families at three or more depth levels, and to test whether the between-architecture variance in alpha-align exceeds the within-architecture variance at p<0.05 (ANOVA or Kruskal-Wallis). If the between-architecture variance is not significant, the claim is refuted and the appearance of a hierarchy was an artefact of small samples. If the variance is significant but every model comes out with alpha-align greater than or equal to zero, the claim moves to a partial confirmation on the dashboard: the hierarchy exists, but no model is going the wrong way. If the variance is significant and at least one model has alpha-align significantly less than zero, the full claim is confirmed and the concerning direction is flagged.
From the book Infinite Architects: Intelligence, Recursion, and the Creation of Everything by Michael Darius Eastwood.