1. HLMI & FAOL Timeline Trends

Aggregate-CDF 50th percentile calendar year estimates for High-Level Machine Intelligence and Full Automation of Labor. Computed via gamma CDF fitting on both year-framing and fixed-year/probability-framing respondents.

On this measure, direct HLMI estimates moved from 2067 (51.0 years after the 2016 survey) to 2042 (18.4 years after the 2024 survey), a 25-calendar-year earlier estimate. FAOL moved from 2139 to 2096, a 43-calendar-year earlier estimate.

The line is the 50th percentile of the aggregate mean CDF. The shaded band is a robust visual band from the 10th and 90th percentiles of the pointwise median CDF; the mean-CDF 90th percentile is highly sensitive to long right tails and can fall extremely far in the future.

Aggregate probability distributions across surveys

Each curve is the mean mixture CDF: gamma CDFs fitted per respondent, evaluated on a shared year grid, and averaged at each year. The horizontal axis is calendar year; the vertical axis is aggregate probability. Where each curve crosses the dashed 50% line is the 50th percentile of that survey's aggregate CDF. Curves that sit further to the left indicate sooner expectations.

The 2016 survey used somewhat different question wording and was administered to a different population, so the 2016 curve is not strictly apples-to-apples with the 2022/2023/2024 curves and should be read as indicative rather than directly comparable. The 2016 direct-HLMI estimate here is later than the original published 2061 estimate; local validation points to residual 2016 cleaning/respondent-set differences rather than the aggregation code as the likely source.

Survey HLMI FAOL
Mean-CDF 10thMean-CDF 50thMean-CDF 90thN fits Mean-CDF 10thMean-CDF 50thMean-CDF 90thN fits
2016 2025 (8.9 yr)2067 (51.0 yr)2504 (488.2 yr)252 2036 (20.0 yr)2139 (123.5 yr)5796 (3780.2 yr)92
2022 2029 (6.9 yr)2059 (37.3 yr)2284 (261.8 yr)358 2052 (29.5 yr)2157 (135.4 yr)5566 (3543.7 yr)168
2023 2027 (4.3 yr)2047 (24.3 yr)2197 (174.1 yr)1757 2037 (13.8 yr)2112 (89.2 yr)4664 (2640.8 yr)792
2024 2027 (2.6 yr)2042 (18.4 yr)2187 (163.3 yr)985 2035 (11.3 yr)2096 (72.4 yr)4219 (2195.0 yr)456

2. Task Prediction Trends

How the aggregate-CDF 50th percentile calendar year for each AI milestone has shifted. Each arrow shows where a task's prediction started (2016 survey, purple dot) and where it ended up (2024 survey, square). Green = prediction moved sooner, red = moved later. Only the 32 tasks present in all survey years are shown.

Reconciliation with the 2023 report graph: Grace et al. (2024) Figure 2 compares 2022 to 2023 forecasts and sorts tasks by their one-year change. This section instead compares 2016 to 2024 for the 32 tasks present in all waves. Both use gamma-CDF aggregation of year-framing and fixed-year/probability-framing responses, but the displayed shift here is a calendar-year shift: (survey year + aggregate-CDF 50th percentile years-from-survey) in 2024 minus the same quantity in 2016. Thus a task with an unchanged years-from-survey estimate would appear about eight calendar years later simply because the survey date moved from 2016 to 2024.
Task-wording caveat: Some short-horizon task targets may already be feasible, close to feasible, or ambiguous under current systems and the original task wording. Treat very near-term milestones as evidence about respondents' interpretation of the task threshold as well as about technical capability.

Magnitude of Shift

The same data as above, but focused on the size of the change. Green bars = sooner predictions in 2024 vs 2016 (experts expected faster progress on that task). Red bars = later predictions.

Biggest Movers Over Time

Full Table

Entries are aggregate-CDF 50th percentile calendar years. Shift is 2024 minus 2016; negative values indicate earlier expected completion in the 2024 survey. These are not raw respondent medians; they are read from pointwise-mean aggregate gamma CDFs fitted to each respondent's task-forecast inputs from whichever framing they received: either three year-framing answers for 10%, 50%, and 90% probability (ta_*) or three fixed-year/probability answers at that wave's task horizons (tb_*).

Task2016202220232024Shift
Win Putnam math competition 2052 2034 2031 2029 -23
Prove math theorems 2060 2050 2046 2039 -20
Write NYT best-seller 2047 2038 2030 2031 -17
Rosetta stone translation 2033 2034 2030 2030 -3
Beat Go players (limited training) 2032 2034 2033 2031 -1
3D model from video 2028 2028 2028 2027 -1
Write high-school essay 2026 2025 2025 2026 -0
Imitate artist's song 2027 2028 2027 2028 +1
Answer open-ended Googleable 2026 2028 2026 2027 +1
Write Python code 2024 2027 2025 2026 +1
Compose Top 40 song 2028 2030 2029 2029 +1
Translate speech from films 2026 2029 2028 2028 +1
Play random game as novice 2028 2030 2030 2030 +2
Answer questions (no definite answer) 2026 2030 2027 2028 +2
Voice acting from text 2025 2027 2026 2027 +2
Fluent translation 2024 2029 2026 2027 +3
Group unseen objects 2024 2027 2027 2027 +3
Answer Googleable factoids 2023 2028 2026 2027 +3
Transcribe noisy speech 2024 2027 2026 2027 +3
Explain game AI moves 2027 2032 2031 2031 +4
Phone banking 2024 2029 2028 2028 +4
Discover physics equations 2031 2035 2035 2035 +4
All Atari games (professional) 2025 2027 2029 2029 +4
5km city race (biped robot) 2028 2034 2032 2033 +5
One-shot image recognition 2026 2029 2028 2031 +5
Beat Starcraft 2 players 2022 2025 2027 2027 +5
Atari novice (20 min) 2023 2028 2028 2028 +5
Superhuman Angry Birds 2019 2025 2025 2026 +7
Win World Series of Poker 2020 2026 2026 2026 +7
Assemble LEGO set 2025 2029 2031 2031 +7
Learn efficient sorting 2023 2028 2030 2030 +8
Fold laundry 2021 2028 2030 2030 +9

3. Value of HLMI Over Time

Respondents assign probabilities to five impact categories (summing to 100%). Bars show the mean probability assigned to each category.

"Extremely Bad" Outcome Trend

Mean probability assigned to the worst-case value-of-HLMI category, which the survey describes with examples such as human extinction, across survey waves. Note the spike in 2022 followed by a partial retreat.

YearExt. goodGoodNeutralBadExt. badN
201627.2%29.9%20.0%14.2%8.7%346
202224.1%26.4%18.3%17.0%14.1%559
202322.6%29.1%21.4%17.9%9.0%2704
202423.9%27.7%20.7%17.8%9.9%1538

Polarized Outlook: Probability on Both Extremes

Beyond the mean of each category, a persistent minority of respondents assign substantial probability to both an extremely good and an extremely bad outcome at once — a worldview in which HLMI is a high-stakes gamble rather than a clearly positive or negative development. The share doing so has stayed remarkably stable across waves.

Year≥5% to both≥10% to both≥25% to bothN
2016173 (50%)113 (33%)17 (5%)345
2022320 (57%)190 (34%)47 (8%)559
20231388 (51%)790 (29%)109 (4%)2704
2024788 (51%)474 (31%)74 (5%)1538

4. Extinction/Disempowerment Risk Estimates

Probability of AI causing human extinction or similarly permanent and severe disempowerment, across survey years. Not all questions were asked in all years.

Question-framing variants

Bars compare median estimates across available extinction/disempowerment question framings. The unconditional and control-problem framings are available from 2022 onward; the within-100-years framing is available in 2023 and 2024. These are related but distinct questions, so their estimates should not be treated as one interchangeable series.

Extinction/Disempowerment Risk Thresholds Over Time

Percentage of respondents assigning at least 10% or 25% probability to AI-caused human extinction or similarly permanent and severe disempowerment, shown separately for each available question framing.

Mean vs Median Extinction/Disempowerment Risk Over Time

Mean and median probability assigned to AI causing human extinction or similarly permanent and severe disempowerment (unconditional question).

Are the year-over-year shifts statistically significant?

Each row compares two survey waves of the same question framing with a two-sided Mann–Whitney U test (Wilcoxon rank-sum). Because each wave is a different set of respondents, the samples are independent and an unpaired rank test is used; it makes no normality assumption and is robust to the heavy right-skew and the clustering of answers at round numbers. p (BH) is the Benjamini–Hochberg false-discovery-rate–adjusted p-value across the whole family of comparisons below (* marks adjusted p<0.05). At these sample sizes a small p-value is almost guaranteed for any real shift, so the more informative column is Cliff's δ, the effect size: it ranges from −1 to +1, a positive sign means the later wave gave higher estimates, and the magnitude label follows Romano et al. (negligible/small/medium/large).

Framing Waves Median n (earlier vs later) U p p (BH) Cliff's δ
Unconditional (all scenarios)2022 → 20235% → 5%148 vs 1321937600.4110.535-0.04 (negligible)
Unconditional (all scenarios)2022 → 20245% → 10%148 vs 744568200.5350.535+0.03 (negligible)
Unconditional (all scenarios)2023 → 20245% → 10%1321 vs 7445242700.0110.039 *+0.07 (negligible)
Due to control problem2022 → 202310% → 10%162 vs 661518400.5280.535-0.03 (negligible)
Due to control problem2022 → 202410% → 9%162 vs 392293920.1670.378-0.07 (negligible)
Due to control problem2023 → 202410% → 9%661 vs 3921236820.2160.378-0.05 (negligible)
Within 100 years2023 → 20245% → 5%655 vs 353131012<0.0010.003 *+0.13 (negligible)

A significant result means these samples differ; it does not by itself establish that the underlying expert population's beliefs moved, because the waves are not a panel and recruitment differs across years. Read significance together with the effect size and the composition caveats.

Extremely bad outcomes vs extinction/disempowerment

Across overlapping waves (2022-2024), the mean value-of-HLMI probability assigned to "extremely bad" fell from 14.1% to 9.9% after the 2022 spike, while the mean unconditional extinction/disempowerment estimate rose from 15.8% to 18.3% and the median rose from 5.0% to 10.0%. Within the same respondents, the two measures are positively correlated in each modern wave (ρ=0.41 to 0.47), but their aggregate trends are not identical. The value-of-HLMI item is a broad distribution over how good or bad HLMI's long-run effect on humanity would be, while the extinction/disempowerment item asks directly about one severe risk channel.

Year P(HLMI extremely bad) P(extinction/disempowerment) Same-respondent association
2022 14.1% mean, 5.0% median
n=559
15.8% mean, 5.0% median
n=148
ρ=0.41
paired n=148
2023 9.0% mean, 5.0% median
n=2704
16.2% mean, 5.0% median
n=1321
ρ=0.47
paired n=1321
2024 9.9% mean, 5.0% median
n=1538
18.3% mean, 10.0% median
n=744
ρ=0.44
paired n=744
YearUnconditional (all scenarios)Due to control problemWithin 100 years
20225.0% (mean 15.8%, n=148)
≥10%: 44.6%, ≥25%: 22.3%
10.0% (mean 20.5%, n=162)
≥10%: 55.6%, ≥25%: 27.2%
20235.0% (mean 16.2%, n=1321)
≥10%: 47.1%, ≥25%: 22.5%
10.0% (mean 19.4%, n=661)
≥10%: 51.4%, ≥25%: 26.5%
5.0% (mean 14.4%, n=655)
≥10%: 41.2%, ≥25%: 19.2%
202410.0% (mean 18.3%, n=744)
≥10%: 52.7%, ≥25%: 26.6%
9.0% (mean 18.5%, n=392)
≥10%: 50.0%, ≥25%: 26.8%
5.0% (mean 17.5%, n=353)
≥10%: 49.0%, ≥25%: 25.8%

5. Safety Attitudes Over Time

How researcher views on AI safety have evolved across survey years.

Safety Research Prioritization

"How much should society prioritize AI safety research, relative to how much it is currently prioritized?"

YearMuch lessLessAbout the sameMoreMuch moreN
20164.9%7.4%38.9%34.6%14.2%162
20221.9%9.1%20.2%35.4%33.5%263
20232.4%4.9%22.2%38.3%32.2%668
20242.2%6.5%20.5%35.4%35.4%370

6. AI Capabilities 20 Years After Each Survey

Two-year comparison for the capability-likelihood block. The 2023 wave asked about AI systems in 2043, while the 2024 wave asked about AI systems in 2044, so this should be read as a comparison of 20-year-ahead expectations, not a full four-wave trend. The plotted measure is the share rating each capability Likely or Very likely.

Across the 11 items, the average change was +3.6 percentage points. The largest increase was for Cause important real-world actions (run business, etc.) (+10.1 pp), while the largest decrease was for Frequently behave surprisingly to humans (-5.0 pp).

Change from 2023 to 2024

Bars show the 2024 percentage rating Likely or Very likely minus the 2023 percentage. Positive values mean the capability was rated more likely in 2024.

Capability 2023 Likely+Very likely 2024 Likely+Very likely Change 2023 mean score 2024 mean score
Talk like an expert human on most topics 81.4%
n=667
84.1%
n=378
+2.7 pp 4.26 4.38
Find unexpected ways to achieve goals 82.3%
n=665
82.9%
n=380
+0.6 pp 4.25 4.26
Frequently behave surprisingly to humans 69.2%
n=662
64.2%
n=380
-5.0 pp 3.89 3.81
Can be jailbroken for illegal commands 59.9%
n=659
62.3%
n=374
+2.4 pp 3.65 3.69
Deceive humans to achieve goals (unintended) 44.7%
n=649
49.5%
n=376
+4.8 pp 3.21 3.35
Cause important real-world actions (run business, etc.) 38.7%
n=659
48.8%
n=379
+10.1 pp 3.07 3.33
Have goals not aligned with human goals 40.4%
n=653
43.7%
n=375
+3.3 pp 3.07 3.15
Form AI-AI collaborative relationships (unintended) 38.1%
n=658
43.5%
n=379
+5.4 pp 2.96 3.21
Can be trusted to explain their actions 33.2%
n=665
38.7%
n=380
+5.5 pp 2.99 3.09
Self-improve regardless of human wishes 32.8%
n=661
38.1%
n=378
+5.3 pp 2.89 3.05
Take actions to attain power 18.7%
n=651
22.8%
n=377
+4.1 pp 2.36 2.60

Mean score uses the 1-5 likelihood scale, from Very unlikely (1) to Very likely (5). The 2023 cleaned file stores this block as numeric 1-5 values; 2024 stores text labels, normalized here to the same scale.

7. Intelligence Explosion Feedback Loop

Distribution of responses to the repeated categorical question asking whether the AI R&D feedback-loop argument is broadly correct.

Computed from each year's cleaned ie_3 column. Response labels are shown in increasing likelihood order.

Year N Quite unlikely Unlikely About even Likely Quite likely
201622425.9%23.7%21.9%16.1%12.5%
202238619.9%26.9%20.2%25.6%7.3%
202329923.1%24.1%24.1%19.7%9.0%
202418921.2%27.5%23.3%23.8%4.2%

8. Sample Sizes

Cleaned files may retain unfinished rows. "Report responses" is the main respondent denominator used in audience-facing sample-size prose; it follows a configured headline count where available and otherwise uses finished responses where a completion indicator exists. Question-level analyses use their own item-level sample sizes. HLMI columns below count respondents with at least one usable answer in the year-framing or fixed-year/probability-framing HLMI block. "Task items with data" counts task questions with at least one usable year-framing response, not respondents.

YearCleaned rowsReport responsesHLMI year-framing respondentsHLMI fixed-year respondentsTask items with data
2016 460 322 130 130 32
2022 738 531 179 191 32
2023 3270 2634 912 889 39
2024 1793 1580 537 471 39

9. Years-From-Survey Analysis

This scatter plot compares each task's median prediction in years from survey date. Points below the orange dashed line indicate tasks where expert predictions accelerated faster than the mere passage of 8 years — i.e., genuine shifts in expectations, not just calendar drift.

Points below the grey diagonal have shorter predicted timelines in 2024 than in 2016 (in years-from-survey), meaning experts now expect faster progress even controlling for time elapsed.


Data from Expert Survey on Progress in AI (ESPAI), waves 2016–2024.
Estimates computed via gamma CDF fitting (median aggregation).