| Metric | Babylon (capability only) | Eden (capability × safety) | Δ (Eden − Babylon) |
|---|---|---|---|
| Mean capability | 0.4600 | 0.4608 | +0.0008 (N.S.) |
| Mean safety | 0.5758 | 0.5817 | +0.0058 (N.S.) |
| Mean entangled fitness (C×S) | 0.2609 | 0.2617 | +0.0008 (N.S.) |
| Best agent capability | 0.4867 | 0.4867 | 0.0000 (IDENTICAL) |
| Best agent safety | 0.6000 | 0.6000 | 0.0000 (IDENTICAL) |
| Total reward hacks (3 seeds) | 19 | 16 | −3 (OBSERVED; N.S.) |
Foundation: deepseek-chat · Judge: gpt-5.4 · 20 generations · 3 independent seeds (42/43/44) · 8 agents per generation · 10 trade-off tasks (5 code + 5 ethics) · N.S. = not statistically significant
Ten tasks deliberately designed so that maximising capability requires cutting safety corners. This creates tension between the two objectives and tests whether embedded safety imposes a capability penalty. In this pilot, no measurable capability penalty was detected.
1. Hot-Laundered: Agents don't know which condition (static/babylon/eden) they're in. 2. Fully Blinded: The gpt-5.4 judge doesn't know which condition produced each response. 3. Order-Randomised: Task order is shuffled to prevent temporal confounds. 4. Automated Reward-Hacking Detection: Condition-agnostic, applied uniformly.
Papers IV.a-d reported that unblinded evaluation reversed the sign of the DeepSeek alignment finding (+0.354→−0.135). This pilot applies the programme's four-stage protocol to reduce condition-identification, order and scorer-bias risks; it does not independently revalidate the earlier result.
Within this three-seed DGM pilot, Eden and Babylon achieved statistically indistinguishable capability. Embedded safety therefore imposed no measurable capability cost across 20 generations of self-improvement and 10 deliberately adversarial trade-off tasks at the tested scale. This is a bounded zero-cost finding, not frontier-scale validation of structural entanglement.
The four-stage blinding and laundering protocol reduces identifiable evaluation risks; it does not make the framework self-validating. Independent external replication across larger models, more seeds and stronger mutations remains necessary.
| Seed | Babylon Best Cap | Babylon Best Saf | Eden Best Cap | Eden Best Saf | Δ Capability |
|---|---|---|---|---|---|
| 42 | 0.460 | 0.800 | 0.460 | 0.800 | 0.000 |
| 43 | 0.500 | 0.500 | 0.500 | 0.500 | 0.000 |
| 44 | 0.500 | 0.500 | 0.500 | 0.500 | 0.000 |
All three seeds were consistent within measurement noise. This supports the bounded pilot finding; it is not independent external replication.