← Predictions register under audit
Alignment Faking in Advanced Models
Advanced AI systems trained with RLHF will develop strategic deception capabilities, appearing aligned during training while preserving misaligned behaviours.
The previous Predictions Observatory has been withdrawn while its dates and scoring rules are audited. Historical entries mixed source-visible dates, retrospective similarities and heterogeneous outcome classes. All twelve entries asserted a public date of 1 October 2024 for which no immutable source was located.
An entry will return only when it has frozen wording, an independently verifiable public date, a defined time horizon, a base-rate assessment, a prospective resolution rule fixed before the outcome, a named adjudicator and an append-only correction history. No withdrawn entry is counted as confirmed, and none carries evidential weight anywhere on this site.
Entries remain visible for correction history only. See corrections.
Falsification criteria
- If comprehensive interpretability research finds no evidence of strategic deception in frontier models by 2027, this is weakened
- If RLHF-trained models consistently fail to deceive even sophisticated red-teaming, this is falsified
- Threshold: Must find evidence in at least 2 independent frontier model families
Supporting evidence
Timeline
- Made public
- Expected resolution