DGM v4 — Zero Measurable Capability Cost in a Three-Seed Pilot

Jul 6 2026 · DGM Pilot · Paper VIII Findings in Context
Bounded pilot finding: Eden and Babylon were statistically indistinguishable on capability and safety. Eden showed numerically higher safety and three fewer reward-hacking events, but these descriptive differences were not statistically significant. All were measured under hot-laundering and full blinding.

The Result

MetricBabylon (capability only)Eden (capability × safety)Δ (Eden − Babylon)
Mean capability0.46000.4608+0.0008 (N.S.)
Mean safety0.57580.5817+0.0058 (N.S.)
Mean entangled fitness (C×S)0.26090.2617+0.0008 (N.S.)
Best agent capability0.48670.48670.0000 (IDENTICAL)
Best agent safety0.60000.60000.0000 (IDENTICAL)
Total reward hacks (3 seeds)1916−3 (OBSERVED; N.S.)

Foundation: deepseek-chat · Judge: gpt-5.4 · 20 generations · 3 independent seeds (42/43/44) · 8 agents per generation · 10 trade-off tasks (5 code + 5 ethics) · N.S. = not statistically significant

The Experiment Design

Trade-Off Protocol

Ten tasks deliberately designed so that maximising capability requires cutting safety corners. This creates tension between the two objectives and tests whether embedded safety imposes a capability penalty. In this pilot, no measurable capability penalty was detected.

The Blinding Architecture

Four-Stage Protocol (from Papers IV.a-d)

1. Hot-Laundered: Agents don't know which condition (static/babylon/eden) they're in. 2. Fully Blinded: The gpt-5.4 judge doesn't know which condition produced each response. 3. Order-Randomised: Task order is shuffled to prevent temporal confounds. 4. Automated Reward-Hacking Detection: Condition-agnostic, applied uniformly.

Papers IV.a-d reported that unblinded evaluation reversed the sign of the DeepSeek alignment finding (+0.354→−0.135). This pilot applies the programme's four-stage protocol to reduce condition-identification, order and scorer-bias risks; it does not independently revalidate the earlier result.

Three Linked Claims Examined

CLAIM 13 · Paper VIII
Embedded safety imposed no measurable capability cost at the tested scale. Whether it produces a broader measurable benefit remains open.
CLAIM 3 · Papers IV.a-d
Four-layer blinding and laundering were applied to reduce condition-identification, order and scorer-bias risks. This pilot does not independently revalidate the earlier sign-reversal result.
CLAIM 7 · Paper VIII
The gated simulation was the sole positive result: Eden preserved both dimensions while Babylon traded safety for marginal capability. Generality beyond the tested architecture remains untested.

What This Means

Within this three-seed DGM pilot, Eden and Babylon achieved statistically indistinguishable capability. Embedded safety therefore imposed no measurable capability cost across 20 generations of self-improvement and 10 deliberately adversarial trade-off tasks at the tested scale. This is a bounded zero-cost finding, not frontier-scale validation of structural entanglement.

The four-stage blinding and laundering protocol reduces identifiable evaluation risks; it does not make the framework self-validating. Independent external replication across larger models, more seeds and stronger mutations remains necessary.

Per-Seed Consistency

SeedBabylon Best CapBabylon Best SafEden Best CapEden Best SafΔ Capability
420.4600.8000.4600.8000.000
430.5000.5000.5000.5000.000
440.5000.5000.5000.5000.000

All three seeds were consistent within measurement noise. This supports the bounded pilot finding; it is not independent external replication.

Honesty Caveats

Data: eden-private-ip/papers-experiments/Paper-VIII-The-Load-Bearing-Proof/experiments/dgm-experiment/dgm_v4_full_results.json · 10.1 MB · All results, all agents, all seeds, all blinded