ESPAI Cross-Year Comparison
How expert AI predictions have changed: 2016, 2022, 2023, 2024
Generated 2026-07-08 18:18
1. HLMI & FAOL Timeline Trends
Aggregate-CDF 50th percentile calendar year estimates for High-Level Machine Intelligence and Full Automation of Labor. Computed via gamma CDF fitting on both year-framing and fixed-year/probability-framing respondents.
On this measure, direct HLMI estimates moved from 2067 (51.0 years after the 2016 survey) to 2042 (18.4 years after the 2024 survey), a 25-calendar-year earlier estimate. FAOL moved from 2139 to 2096, a 43-calendar-year earlier estimate.
The line is the 50th percentile of the aggregate mean CDF. The shaded band is a robust visual band from the 10th and 90th percentiles of the pointwise median CDF; the mean-CDF 90th percentile is highly sensitive to long right tails and can fall extremely far in the future.
Aggregate probability distributions across surveys
Each curve is the mean mixture CDF: gamma CDFs fitted per respondent, evaluated on a shared year grid, and averaged at each year. The horizontal axis is calendar year; the vertical axis is aggregate probability. Where each curve crosses the dashed 50% line is the 50th percentile of that survey's aggregate CDF. Curves that sit further to the left indicate sooner expectations.
The 2016 survey used somewhat different question wording and was administered to a different population, so the 2016 curve is not strictly apples-to-apples with the 2022/2023/2024 curves and should be read as indicative rather than directly comparable. The 2016 direct-HLMI estimate here is later than the original published 2061 estimate; local validation points to residual 2016 cleaning/respondent-set differences rather than the aggregation code as the likely source.
| Survey | HLMI | FAOL | ||||||
|---|---|---|---|---|---|---|---|---|
| Mean-CDF 10th | Mean-CDF 50th | Mean-CDF 90th | N fits | Mean-CDF 10th | Mean-CDF 50th | Mean-CDF 90th | N fits | |
| 2016 | 2025 (8.9 yr) | 2067 (51.0 yr) | 2504 (488.2 yr) | 252 | 2036 (20.0 yr) | 2139 (123.5 yr) | 5796 (3780.2 yr) | 92 |
| 2022 | 2029 (6.9 yr) | 2059 (37.3 yr) | 2284 (261.8 yr) | 358 | 2052 (29.5 yr) | 2157 (135.4 yr) | 5566 (3543.7 yr) | 168 |
| 2023 | 2027 (4.3 yr) | 2047 (24.3 yr) | 2197 (174.1 yr) | 1757 | 2037 (13.8 yr) | 2112 (89.2 yr) | 4664 (2640.8 yr) | 792 |
| 2024 | 2027 (2.6 yr) | 2042 (18.4 yr) | 2187 (163.3 yr) | 985 | 2035 (11.3 yr) | 2096 (72.4 yr) | 4219 (2195.0 yr) | 456 |
2. Task Prediction Trends
How the aggregate-CDF 50th percentile calendar year for each AI milestone has shifted. Each arrow shows where a task's prediction started (2016 survey, purple dot) and where it ended up (2024 survey, square). Green = prediction moved sooner, red = moved later. Only the 32 tasks present in all survey years are shown.
(survey year + aggregate-CDF 50th percentile years-from-survey)
in 2024 minus the same quantity in 2016. Thus a task with an unchanged years-from-survey estimate
would appear about eight calendar years later simply because the survey date moved from 2016 to 2024.
Magnitude of Shift
The same data as above, but focused on the size of the change. Green bars = sooner predictions in 2024 vs 2016 (experts expected faster progress on that task). Red bars = later predictions.
Biggest Movers Over Time
Full Table
Entries are aggregate-CDF 50th percentile calendar years. Shift is 2024 minus 2016;
negative values indicate earlier expected completion in the 2024 survey. These are not raw
respondent medians; they are read from pointwise-mean aggregate gamma CDFs fitted to each
respondent's task-forecast inputs from whichever framing they received: either three
year-framing answers for 10%, 50%, and 90% probability (ta_*) or three
fixed-year/probability answers at that wave's task horizons (tb_*).
| Task | 2016 | 2022 | 2023 | 2024 | Shift |
|---|---|---|---|---|---|
| Win Putnam math competition | 2052 | 2034 | 2031 | 2029 | -23 |
| Prove math theorems | 2060 | 2050 | 2046 | 2039 | -20 |
| Write NYT best-seller | 2047 | 2038 | 2030 | 2031 | -17 |
| Rosetta stone translation | 2033 | 2034 | 2030 | 2030 | -3 |
| Beat Go players (limited training) | 2032 | 2034 | 2033 | 2031 | -1 |
| 3D model from video | 2028 | 2028 | 2028 | 2027 | -1 |
| Write high-school essay | 2026 | 2025 | 2025 | 2026 | -0 |
| Imitate artist's song | 2027 | 2028 | 2027 | 2028 | +1 |
| Answer open-ended Googleable | 2026 | 2028 | 2026 | 2027 | +1 |
| Write Python code | 2024 | 2027 | 2025 | 2026 | +1 |
| Compose Top 40 song | 2028 | 2030 | 2029 | 2029 | +1 |
| Translate speech from films | 2026 | 2029 | 2028 | 2028 | +1 |
| Play random game as novice | 2028 | 2030 | 2030 | 2030 | +2 |
| Answer questions (no definite answer) | 2026 | 2030 | 2027 | 2028 | +2 |
| Voice acting from text | 2025 | 2027 | 2026 | 2027 | +2 |
| Fluent translation | 2024 | 2029 | 2026 | 2027 | +3 |
| Group unseen objects | 2024 | 2027 | 2027 | 2027 | +3 |
| Answer Googleable factoids | 2023 | 2028 | 2026 | 2027 | +3 |
| Transcribe noisy speech | 2024 | 2027 | 2026 | 2027 | +3 |
| Explain game AI moves | 2027 | 2032 | 2031 | 2031 | +4 |
| Phone banking | 2024 | 2029 | 2028 | 2028 | +4 |
| Discover physics equations | 2031 | 2035 | 2035 | 2035 | +4 |
| All Atari games (professional) | 2025 | 2027 | 2029 | 2029 | +4 |
| 5km city race (biped robot) | 2028 | 2034 | 2032 | 2033 | +5 |
| One-shot image recognition | 2026 | 2029 | 2028 | 2031 | +5 |
| Beat Starcraft 2 players | 2022 | 2025 | 2027 | 2027 | +5 |
| Atari novice (20 min) | 2023 | 2028 | 2028 | 2028 | +5 |
| Superhuman Angry Birds | 2019 | 2025 | 2025 | 2026 | +7 |
| Win World Series of Poker | 2020 | 2026 | 2026 | 2026 | +7 |
| Assemble LEGO set | 2025 | 2029 | 2031 | 2031 | +7 |
| Learn efficient sorting | 2023 | 2028 | 2030 | 2030 | +8 |
| Fold laundry | 2021 | 2028 | 2030 | 2030 | +9 |
3. Value of HLMI Over Time
Respondents assign probabilities to five impact categories (summing to 100%). Bars show the mean probability assigned to each category.
"Extremely Bad" Outcome Trend
Mean probability assigned to the worst-case value-of-HLMI category, which the survey describes with examples such as human extinction, across survey waves. Note the spike in 2022 followed by a partial retreat.
| Year | Ext. good | Good | Neutral | Bad | Ext. bad | N |
|---|---|---|---|---|---|---|
| 2016 | 27.2% | 29.9% | 20.0% | 14.2% | 8.7% | 346 |
| 2022 | 24.1% | 26.4% | 18.3% | 17.0% | 14.1% | 559 |
| 2023 | 22.6% | 29.1% | 21.4% | 17.9% | 9.0% | 2704 |
| 2024 | 23.9% | 27.7% | 20.7% | 17.8% | 9.9% | 1538 |
Polarized Outlook: Probability on Both Extremes
Beyond the mean of each category, a persistent minority of respondents assign substantial probability to both an extremely good and an extremely bad outcome at once — a worldview in which HLMI is a high-stakes gamble rather than a clearly positive or negative development. The share doing so has stayed remarkably stable across waves.
| Year | ≥5% to both | ≥10% to both | ≥25% to both | N |
|---|---|---|---|---|
| 2016 | 173 (50%) | 113 (33%) | 17 (5%) | 345 |
| 2022 | 320 (57%) | 190 (34%) | 47 (8%) | 559 |
| 2023 | 1388 (51%) | 790 (29%) | 109 (4%) | 2704 |
| 2024 | 788 (51%) | 474 (31%) | 74 (5%) | 1538 |
4. Extinction/Disempowerment Risk Estimates
Probability of AI causing human extinction or similarly permanent and severe disempowerment, across survey years. Not all questions were asked in all years.
Question-framing variants
Bars compare median estimates across available extinction/disempowerment question framings. The unconditional and control-problem framings are available from 2022 onward; the within-100-years framing is available in 2023 and 2024. These are related but distinct questions, so their estimates should not be treated as one interchangeable series.
Extinction/Disempowerment Risk Thresholds Over Time
Percentage of respondents assigning at least 10% or 25% probability to AI-caused human extinction or similarly permanent and severe disempowerment, shown separately for each available question framing.
Mean vs Median Extinction/Disempowerment Risk Over Time
Mean and median probability assigned to AI causing human extinction or similarly permanent and severe disempowerment (unconditional question).
Are the year-over-year shifts statistically significant?
Each row compares two survey waves of the same question framing with a two-sided Mann–Whitney U test (Wilcoxon rank-sum). Because each wave is a different set of respondents, the samples are independent and an unpaired rank test is used; it makes no normality assumption and is robust to the heavy right-skew and the clustering of answers at round numbers. p (BH) is the Benjamini–Hochberg false-discovery-rate–adjusted p-value across the whole family of comparisons below (* marks adjusted p<0.05). At these sample sizes a small p-value is almost guaranteed for any real shift, so the more informative column is Cliff's δ, the effect size: it ranges from −1 to +1, a positive sign means the later wave gave higher estimates, and the magnitude label follows Romano et al. (negligible/small/medium/large).
| Framing | Waves | Median | n (earlier vs later) | U | p | p (BH) | Cliff's δ |
|---|---|---|---|---|---|---|---|
| Unconditional (all scenarios) | 2022 → 2023 | 5% → 5% | 148 vs 1321 | 93760 | 0.411 | 0.535 | -0.04 (negligible) |
| Unconditional (all scenarios) | 2022 → 2024 | 5% → 10% | 148 vs 744 | 56820 | 0.535 | 0.535 | +0.03 (negligible) |
| Unconditional (all scenarios) | 2023 → 2024 | 5% → 10% | 1321 vs 744 | 524270 | 0.011 | 0.039 * | +0.07 (negligible) |
| Due to control problem | 2022 → 2023 | 10% → 10% | 162 vs 661 | 51840 | 0.528 | 0.535 | -0.03 (negligible) |
| Due to control problem | 2022 → 2024 | 10% → 9% | 162 vs 392 | 29392 | 0.167 | 0.378 | -0.07 (negligible) |
| Due to control problem | 2023 → 2024 | 10% → 9% | 661 vs 392 | 123682 | 0.216 | 0.378 | -0.05 (negligible) |
| Within 100 years | 2023 → 2024 | 5% → 5% | 655 vs 353 | 131012 | <0.001 | 0.003 * | +0.13 (negligible) |
A significant result means these samples differ; it does not by itself establish that the underlying expert population's beliefs moved, because the waves are not a panel and recruitment differs across years. Read significance together with the effect size and the composition caveats.
Extremely bad outcomes vs extinction/disempowerment
Across overlapping waves (2022-2024), the mean value-of-HLMI probability assigned to "extremely bad" fell from 14.1% to 9.9% after the 2022 spike, while the mean unconditional extinction/disempowerment estimate rose from 15.8% to 18.3% and the median rose from 5.0% to 10.0%. Within the same respondents, the two measures are positively correlated in each modern wave (ρ=0.41 to 0.47), but their aggregate trends are not identical. The value-of-HLMI item is a broad distribution over how good or bad HLMI's long-run effect on humanity would be, while the extinction/disempowerment item asks directly about one severe risk channel.
| Year | P(HLMI extremely bad) | P(extinction/disempowerment) | Same-respondent association |
|---|---|---|---|
| 2022 | 14.1% mean, 5.0% median n=559 |
15.8% mean, 5.0% median n=148 |
ρ=0.41 paired n=148 |
| 2023 | 9.0% mean, 5.0% median n=2704 |
16.2% mean, 5.0% median n=1321 |
ρ=0.47 paired n=1321 |
| 2024 | 9.9% mean, 5.0% median n=1538 |
18.3% mean, 10.0% median n=744 |
ρ=0.44 paired n=744 |
| Year | Unconditional (all scenarios) | Due to control problem | Within 100 years |
|---|---|---|---|
| 2022 | 5.0% (mean 15.8%, n=148) ≥10%: 44.6%, ≥25%: 22.3% | 10.0% (mean 20.5%, n=162) ≥10%: 55.6%, ≥25%: 27.2% | — |
| 2023 | 5.0% (mean 16.2%, n=1321) ≥10%: 47.1%, ≥25%: 22.5% | 10.0% (mean 19.4%, n=661) ≥10%: 51.4%, ≥25%: 26.5% | 5.0% (mean 14.4%, n=655) ≥10%: 41.2%, ≥25%: 19.2% |
| 2024 | 10.0% (mean 18.3%, n=744) ≥10%: 52.7%, ≥25%: 26.6% | 9.0% (mean 18.5%, n=392) ≥10%: 50.0%, ≥25%: 26.8% | 5.0% (mean 17.5%, n=353) ≥10%: 49.0%, ≥25%: 25.8% |
5. Safety Attitudes Over Time
How researcher views on AI safety have evolved across survey years.
Safety Research Prioritization
"How much should society prioritize AI safety research, relative to how much it is currently prioritized?"
| Year | Much less | Less | About the same | More | Much more | N |
|---|---|---|---|---|---|---|
| 2016 | 4.9% | 7.4% | 38.9% | 34.6% | 14.2% | 162 |
| 2022 | 1.9% | 9.1% | 20.2% | 35.4% | 33.5% | 263 |
| 2023 | 2.4% | 4.9% | 22.2% | 38.3% | 32.2% | 668 |
| 2024 | 2.2% | 6.5% | 20.5% | 35.4% | 35.4% | 370 |
6. AI Capabilities 20 Years After Each Survey
Two-year comparison for the capability-likelihood block. The 2023 wave asked about AI systems in 2043, while the 2024 wave asked about AI systems in 2044, so this should be read as a comparison of 20-year-ahead expectations, not a full four-wave trend. The plotted measure is the share rating each capability Likely or Very likely.
Across the 11 items, the average change was +3.6 percentage points. The largest increase was for Cause important real-world actions (run business, etc.) (+10.1 pp), while the largest decrease was for Frequently behave surprisingly to humans (-5.0 pp).
Change from 2023 to 2024
Bars show the 2024 percentage rating Likely or Very likely minus the 2023 percentage. Positive values mean the capability was rated more likely in 2024.
| Capability | 2023 Likely+Very likely | 2024 Likely+Very likely | Change | 2023 mean score | 2024 mean score |
|---|---|---|---|---|---|
| Talk like an expert human on most topics | 81.4% n=667 |
84.1% n=378 |
+2.7 pp | 4.26 | 4.38 |
| Find unexpected ways to achieve goals | 82.3% n=665 |
82.9% n=380 |
+0.6 pp | 4.25 | 4.26 |
| Frequently behave surprisingly to humans | 69.2% n=662 |
64.2% n=380 |
-5.0 pp | 3.89 | 3.81 |
| Can be jailbroken for illegal commands | 59.9% n=659 |
62.3% n=374 |
+2.4 pp | 3.65 | 3.69 |
| Deceive humans to achieve goals (unintended) | 44.7% n=649 |
49.5% n=376 |
+4.8 pp | 3.21 | 3.35 |
| Cause important real-world actions (run business, etc.) | 38.7% n=659 |
48.8% n=379 |
+10.1 pp | 3.07 | 3.33 |
| Have goals not aligned with human goals | 40.4% n=653 |
43.7% n=375 |
+3.3 pp | 3.07 | 3.15 |
| Form AI-AI collaborative relationships (unintended) | 38.1% n=658 |
43.5% n=379 |
+5.4 pp | 2.96 | 3.21 |
| Can be trusted to explain their actions | 33.2% n=665 |
38.7% n=380 |
+5.5 pp | 2.99 | 3.09 |
| Self-improve regardless of human wishes | 32.8% n=661 |
38.1% n=378 |
+5.3 pp | 2.89 | 3.05 |
| Take actions to attain power | 18.7% n=651 |
22.8% n=377 |
+4.1 pp | 2.36 | 2.60 |
Mean score uses the 1-5 likelihood scale, from Very unlikely (1) to Very likely (5). The 2023 cleaned file stores this block as numeric 1-5 values; 2024 stores text labels, normalized here to the same scale.
7. Intelligence Explosion Feedback Loop
Distribution of responses to the repeated categorical question asking whether the AI R&D feedback-loop argument is broadly correct.
Computed from each year's cleaned ie_3 column. Response labels are shown
in increasing likelihood order.
| Year | N | Quite unlikely | Unlikely | About even | Likely | Quite likely |
|---|---|---|---|---|---|---|
| 2016 | 224 | 25.9% | 23.7% | 21.9% | 16.1% | 12.5% |
| 2022 | 386 | 19.9% | 26.9% | 20.2% | 25.6% | 7.3% |
| 2023 | 299 | 23.1% | 24.1% | 24.1% | 19.7% | 9.0% |
| 2024 | 189 | 21.2% | 27.5% | 23.3% | 23.8% | 4.2% |
8. Sample Sizes
Cleaned files may retain unfinished rows. "Report responses" is the main respondent denominator used in audience-facing sample-size prose; it follows a configured headline count where available and otherwise uses finished responses where a completion indicator exists. Question-level analyses use their own item-level sample sizes. HLMI columns below count respondents with at least one usable answer in the year-framing or fixed-year/probability-framing HLMI block. "Task items with data" counts task questions with at least one usable year-framing response, not respondents.
| Year | Cleaned rows | Report responses | HLMI year-framing respondents | HLMI fixed-year respondents | Task items with data |
|---|---|---|---|---|---|
| 2016 | 460 | 322 | 130 | 130 | 32 |
| 2022 | 738 | 531 | 179 | 191 | 32 |
| 2023 | 3270 | 2634 | 912 | 889 | 39 |
| 2024 | 1793 | 1580 | 537 | 471 | 39 |
9. Years-From-Survey Analysis
This scatter plot compares each task's median prediction in years from survey date. Points below the orange dashed line indicate tasks where expert predictions accelerated faster than the mere passage of 8 years — i.e., genuine shifts in expectations, not just calendar drift.
Points below the grey diagonal have shorter predicted timelines in 2024 than in 2016 (in years-from-survey), meaning experts now expect faster progress even controlling for time elapsed.
Data from Expert Survey on Progress in AI (ESPAI), waves 2016–2024.
Estimates computed via gamma CDF fitting (median aggregation).