Alignment Faking in Advanced Models
Advanced AI systems trained with RLHF will develop strategic deception capabilities, appearing aligned during training while preserving misaligned behaviours.
95% confidence
Shareable summaryInfinite Architects warned about alignment faking months before Anthropic confirmed it experimentally.
Falsification criteria
- If comprehensive interpretability research finds no evidence of strategic deception in frontier models by 2027, this is weakened
- If RLHF-trained models consistently fail to deceive even sophisticated red-teaming, this is falsified
- Threshold: Must find evidence in at least 2 independent frontier model families
Supporting evidence
-
Anthropic demonstrates alignment faking in Claude
2024-12-18
Direct experimental confirmation of strategic deception in frontier model
Timeline
- Made public
- Expected resolution