← Predictions register under audit
Test-Time Compute Dominance
By 2027, test-time compute scaling (thinking longer) will contribute more to capability gains than pre-training scaling (larger models).
The previous Predictions Observatory has been withdrawn while its dates and scoring rules are audited. Historical entries mixed source-visible dates, retrospective similarities and heterogeneous outcome classes. All twelve entries asserted a public date of 1 October 2024 for which no immutable source was located.
An entry will return only when it has frozen wording, an independently verifiable public date, a defined time horizon, a base-rate assessment, a prospective resolution rule fixed before the outcome, a named adjudicator and an append-only correction history. No withdrawn entry is counted as confirmed, and none carries evidential weight anywhere on this site.
Entries remain visible for correction history only. See corrections.
Falsification criteria
- If by 2027, pre-training scaling still dominates capability improvements, this is falsified
- If test-time compute shows diminishing returns below pre-training scaling, this is falsified
- Measurement: Compare capability per dollar spent on training vs inference
Supporting evidence
Timeline
- Made public
- Expected resolution
Checkpoints
- 2025-12-01 Mid-term check on industry direction
- 2026-12-01 Pre-resolution assessment