1. Overview & Survey Methodology

322
Report Responses
460
Cleaned Analysis Rows
322
Finished
138
Incomplete (partial data kept)
231
Year Framing
229
Probability Framing

Randomization Blocks

BlockCountPercentage
Block 116034.8%
Block 210723.3%
Block 311825.7%

Data Cleaning Notes

The following cleaning steps were applied (see cleaning_log.md for full details):

Caveat — sum-to-100 triplets: 29 triplets across 27 respondents have values summing to exactly 100 (e.g., 10%, 30%, 60% at the three time horizons). This may indicate respondents who treated the probabilities as shares that must total 100%, rather than independent cumulative probabilities at each time horizon. Following standard methodology, sum-to-100 triplets from respondents with descending errors were also removed (2 total triplets NaN'd).


3.1 When Will 39 AI Milestones Be Feasible?

19
Tasks feasible within 10 years (50% prob)
39
Total milestones assessed

Corresponds to Figure 1 in Grace et al. (2024). Dots show the 50% probability year from the mixture CDF; horizontal lines show the 10%–90% range. The x-axis is capped at 2100; milestones whose 50% or 90% year falls beyond the cap are labeled with an arrow showing the true year. Includes 39 tasks, 4 occupations, HLMI, and FAOL.

Combined year-framing and probability-framing responses using gamma mixture CDF aggregation (matching the 2023 report methodology). A gamma CDF is fitted to each respondent's data, then all CDFs are averaged pointwise to produce a mixture CDF from which percentiles are read. Sorted by 50% probability year (earliest first). The HLMI row uses the dedicated direct-HLMI blocks (hb_a_*/hb_b_*); it does not pool the final-occupation prediction block.

MilestoneN (fitted)10% Prob Year50% Prob Year90% Prob Year
Play new Angry Birds levels (superhuman)380 yr (2016)3 yr (2019)14 yr (2030)
Win World Series of Poker370 yr (2016)4 yr (2020)19 yr (2035)
Fold laundry (speed + quality)290 yr (2016)5 yr (2021)29 yr (2045)
Beat best Starcraft 2 players241 yr (2017)6 yr (2022)64 yr (2080)
Learn efficient sorting (no solution form)430 yr (2016)7 yr (2023)31 yr (2047)
Atari novice level (20 min training)331 yr (2017)7 yr (2023)41 yr (2057)
Answer Googleable factoid questions441 yr (2017)7 yr (2023)34 yr (2050)
Group unseen objects into classes291 yr (2017)8 yr (2024)61 yr (2077)
Fluent translation (most languages)411 yr (2017)8 yr (2024)40 yr (2056)
Transcribe speech (noisy, accents)321 yr (2017)8 yr (2024)23 yr (2039)
Phone banking services291 yr (2017)8 yr (2024)33 yr (2049)
Write Python code (e.g. quicksort)361 yr (2017)8 yr (2024)39 yr (2055)
Assemble any LEGO set350 yr (2016)9 yr (2025)40 yr (2056)
Outperform on all Atari games361 yr (2017)9 yr (2025)48 yr (2064)
Voice acting from text422 yr (2018)9 yr (2025)28 yr (2044)
Write high-school history essay421 yr (2017)10 yr (2026)51 yr (2067)
One-shot image recognition301 yr (2017)10 yr (2026)42 yr (2058)
Answer questions with no definite answer452 yr (2018)10 yr (2026)56 yr (2072)
Answer Googleable open-ended questions362 yr (2018)10 yr (2026)62 yr (2078)
Translate speech from subtitled films371 yr (2017)10 yr (2026)52 yr (2068)
Explain game AI moves to layman351 yr (2017)11 yr (2027)85 yr (2101)
Produce song indistinguishable from artist402 yr (2018)11 yr (2027)55 yr (2071)
Truck Driver (occ.)892 yr (2018)12 yr (2028)40 yr (2056)
Compose US Top 40 song (full audio)361 yr (2017)12 yr (2028)50 yr (2066)
Beat fastest human in 5km city race (biped robot)281 yr (2017)12 yr (2028)43 yr (2059)
3D model from short video402 yr (2018)12 yr (2028)45 yr (2061)
Play random game as human novice (<10 min)441 yr (2017)12 yr (2028)85 yr (2101)
Retail Salesperson (occ.)893 yr (2019)15 yr (2031)67 yr (2083)
Discover physics equations from simulation511 yr (2017)15 yr (2031)144 yr (2160)
Beat best Go players (limited training)402 yr (2018)16 yr (2032)328 yr (2344)
Translate newly discovered language (Rosetta stone)341 yr (2017)17 yr (2033)135 yr (2151)
Write NYT best-seller novel266 yr (2022)31 yr (2047)233 yr (2249)
Win Putnam math competition434 yr (2020)36 yr (2052)203 yr (2219)
Surgeon (occ.)919 yr (2025)37 yr (2053)285 yr (2301)
Prove publishable math theorems316 yr (2022)44 yr (2060)156 yr (2172)
HLMI (all human tasks)2529 yr (2025)45 yr (2061)350 yr (2366)
AI Researcher (occ.)9021 yr (2037)87 yr (2103)2406 yr (4422)
Full Automation of Labor9220 yr (2036)123 yr (2139)3780 yr (5796)
Conduct ML research & write conference paper0
Solve unsolved math problem (e.g. Millennium)0
Install electrical wiring in new home0
Replicate ML conference study0
Fine-tune open source LLM0
Build website with payment processing0
Find & patch security flaw (100k+ users)0

Year values are years from survey date (2016), with calendar year in parentheses. "10% prob year" = year at which the mixture CDF reaches 10%. "90% prob year" = 90%. Dashes indicate insufficient data or an aggregate CDF that does not reach the target percentile within the 1e8-year "never/infinity" sentinel bound.


3.2 HLMI & Full Automation of Labor Timing

High-Level Machine Intelligence (HLMI)

HLMI is defined as machines that can accomplish every task better and cheaper than human workers.

9 yr (2025)
10% probability
45 yr (2061)
50% probability
350 yr (2366)
90% probability
Percentile2016 Survey2023 SurveyShift
10% prob9 yr (2025)11 yr (2027)-1.9 yr
50% prob45 yr (2061)31 yr (2047)+14.5 yr

N (gamma fits): 252 (0 failed fits)

This HLMI estimate uses only the dedicated direct-HLMI question blocks (hb_a_* and hb_b_*). It does not pool the separate final-occupation prediction block (hj_*_final_pred).

Full Automation of Labor (FAOL)

All occupations fully automatable — machines carry out every task better and more cheaply than humans.

20 yr (2036)
10% probability
123 yr (2139)
50% probability
3780 yr (5796)
90% probability
Percentile2016 Survey2023 SurveyShift
50% prob123 yr (2139)100 yr (2116)+23.5 yr

N (gamma fits): 92 (0 failed fits)

Gap Between HLMI and FAOL

78 years
Gap at 50% probability

The consistent gap between HLMI and FAOL predictions reflects researchers' view that achieving human-level AI capability differs from fully automating all occupations.

Methodology Note

Gamma mixture CDF method: Following the methodology of Grace et al. (2024), a gamma CDF is fitted to each respondent's three data points (whether year-framing or probability-framing). All individual CDFs are then averaged pointwise to produce a mixture CDF, from which percentiles are read off. This approach handles extrapolation naturally (pessimistic respondents contribute heavy-tailed CDFs rather than being dropped) and treats both question framings symmetrically. Respondents were randomly assigned to either a "year framing" (provide years for 10%/50%/90% probability) or a "probability framing" (provide probabilities at fixed time horizons: 10, 20, and 40 years for HLMI; 10, 20, and 50 years for FAOL). See compare_methods.py for a comparison of this method with linear interpolation.

Corresponds to Figure 3 in Grace et al. (2024). Thin lines show individual respondent gamma CDFs (random subset of 200); thick line is the mixture (average) CDF. The dashed line marks 50% probability.


3.2.4 High-Level Machine Intelligence Timing by Experience

This table splits respondents at the median years in field, then refits the direct High-Level Machine Intelligence (HLMI) timing model within each experience group.

Experience groupDirect HLMI gamma-fit countMixture CDF 50% YearCalendar Year
More experienced (>=6 yr)4158.3 yr2074
Less experienced (<6 yr)3238.9 yr2055

The count values are not all respondents in each experience group. They count respondents who both answered the years-in-field question and gave enough direct HLMI timing answers to fit a gamma CDF. The separate occupation-automation prediction block is not pooled into this HLMI comparison. Only 107 respondents received the experience question due to randomization. Geographic region and citation count data are not available in the anonymized dataset.


3.2.5 Do Participants Agree on HLMI Timing?

"How much do you think your views on when HLMI will be achieved differ from those of the typical AI researcher?" (N=115)

ResponseCountPercentage
Not much6455.7%
A moderate amount4135.7%
A lot108.7%

3.3 Framing Effects

Respondents were randomly assigned to one of two question framings for the same underlying question. The "year framing" asked how many years until a given probability, while the "probability framing" asked for the probability at a fixed time horizon.

HLMI 50% Probability Year by Framing

FramingN (gamma fits)Mixture CDF 50% YearCalendar Year
Year framing (respondent provides years)12440.4 yr2056
Probability framing (respondent provides probabilities)12850.7 yr2067
10.3 years
Framing effect (difference)

A consistent framing effect has been observed across survey waves: the year-framing tends to produce earlier (shorter-timeline) predictions than the probability-framing. This is a documented cognitive bias in probability elicitation.


3.4 Perceived Rates of Progress

Respondents were asked whether AI progress was faster in the first or second half of their career. (N=107)

0.0%
Said second half was faster
6 yr
Median time in field (N=107)

Corresponds to Figure 4 in Grace et al. (2024).

ResponseCountPercentage2023 Survey
The second half00.0%60%
The first half00.0%
They were about the same00.0%

How Far Along Is AI Progress? (Outside-view slider)

Respondents were shown a slider with three anchors: A = where progress was when they started working in their AI area; B = where it is now; and C = where it would need to be for AI software to have roughly human-level abilities at the tasks they study. They were asked "What fraction of the distance between where progress was when you started working in the area (A) and where it would need to be to attain human-level abilities in the area (C) have we come so far (B)?" on a 0–100 scale. This complements the year-based timing questions by anchoring an estimate to each respondent's own career window — so it is robust to disagreement about absolute calendar dates.

SurveyNMedianMeanSD
2016104 18 21.1 21.4

IQR (2016): 5–25.


3.5 What Causes AI Progress?

Respondents estimated how much less AI progress there would have been with half as much of each input. Higher values = more important to progress. Values are percentage less progress (0-100%).

Corresponds to Figure 5 in Grace et al. (2024). Red dots are means; box shows IQR with median line.

FactorNMedianMeanIQR
Computing hardware cost decline6350.0%56.1%40–80%
AI algorithm progress5750.0%44.8%30–50%
Training dataset effort7040.0%43.7%20–69%
Funding6740.0%43.2%30–60%
Researcher effort7035.0%41.9%25–60%

3.6 Will There Be an Intelligence Explosion?

Probability Estimates (numeric, 0-100%)

ScenarioNMedianMeanIQR2023 Survey MedianChange
Dramatic tech speedup within 2 years of HLMI22020%30.0%5–50%20%unchanged
Dramatic tech speedup within 30 years of HLMI22080%68.9%50–99%80%unchanged
Vastly superhuman AI within 2 years of HLMI20910%19.5%1–25%10%unchanged
Vastly superhuman AI within 30 years of HLMI21050%53.2%20–90%

Is the Feedback Loop Argument Broadly Correct? (N=224)

"If AI does nearly all R&D, improvements in AI will accelerate progress including further AI progress. This could cause progress to become >10x faster within 5 years."

Corresponds to Figure 6 in Grace et al. (2024).

ResponseCountPercentage
Quite unlikely (0-20%)5825.9%
Unlikely (21-40%)5323.7%
About even chance (41-60%)4921.9%
Likely (61-80%)3616.1%
Quite likely (81-100%)2812.5%

3.7 AI capabilities 2044

Section skipped — required data not available for this survey year. (KeyError: 'Q1#1_1')


3.8 Explainability

Section skipped — required data not available for this survey year. (KeyError: 'Q1 (1)')


4.1 Concerning scenarios

Section skipped — required data not available for this survey year. (KeyError: 'Q369#1_1')


4.2 How Good or Bad Will HLMI Be?

Respondents assigned probabilities to five outcome categories (summing to 100%). N=345

70.7%
Net optimists
16.2%
Net pessimists
13.0%
Balanced

Corresponds to Figure 11 in Grace et al. (2024), "Thousands of AI Authors on the Future of AI."

Mean and Median Probabilities

OutcomeMeanMedian
Extremely good27.1%20.0%
On balance good29.9%25.0%
Neutral20.0%20.0%
On balance bad14.2%10.0%
Extremely bad8.7%5.0%

Individual Respondent Views

Each vertical slice below is one respondent's probability allocation, sorted from most optimistic (left) to most pessimistic (right). The chart reveals the full diversity of expert opinion.

Corresponds to Figure 10 in Grace et al. (2024).

The same respondents sorted by how much probability they assign to "extremely bad (e.g. human extinction)" outcomes. The growing black band on the right shows those assigning the highest probability to extremely bad outcomes.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

Individual Response Profiles

Each small bar below represents one respondent's five-category probability distribution, sampled evenly from the most optimistic to the most pessimistic.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

Extreme Outcome Analysis

Metric2016 Survey2023 SurveyChange
Non-zero to both extremes64.1%64.0%+0.1 pp
>=5% on extremely bad54.8%57.8%-3.0 pp
>=10% on extremely bad40.0%37.8%+2.2 pp
>=20% on extremely bad17.7%
>=25% on extremely bad9.6%
Mean extremely bad8.7%9.0%-0.3 pp
Median extremely bad5.0%5.0%unchanged
Net optimists70.7%68.3%+2.4 pp

4.3 Extinction risk

Section skipped — required data not available for this survey year. (KeyError: 'extinction_all_1')


4.4 Are Future AI-Risk Concerns Due to Misunderstandings of AI Research?

"To what extent do you think people's concerns about future risks from AI are due to misunderstandings of AI research?" (N=115)

No direct figure equivalent in Grace et al. (2024); this data is discussed in Section 4.4 of that paper.

ResponseCountPercentage
Hardly at all00.0%
Not much00.0%
Somewhat00.0%
To a large extent00.0%
Almost entirely00.0%

4.5 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity

This question was not included in this survey year.


4.6 How Much Should AI Safety Research Be Prioritized?

"How much should society prioritize AI safety research, relative to how much it is currently prioritized?" (N=162)

48.8%
Say 'More' or 'Much more'
70%
2023 Survey (More+Much more)

Corresponds to Figure 14 in Grace et al. (2024).

ResponseCountPercentage
Much less84.9%
Less127.4%
About the same6338.9%
More5634.6%
Much more2314.2%

4.7 The Alignment Problem

Corresponds to Figure 15 in Grace et al. (2024).

Do you think this argument points at an important problem? (N=149)

ResponseCountPercentage
No, not a real problem.1610.7%
No, not an important problem.2919.5%
Yes, a moderately important problem.4630.9%
Yes, a very important problem.85.4%
Yes, among the most important problems in the field.5033.6%

How valuable is it to work on this problem today, compared to other problems in AI? (N=149)

ResponseCountPercentage
Much less valuable3322.1%
Less valuable6040.3%
As valuable as other problems4228.2%
More valuable128.1%
Much more valuable21.3%

How hard do you think this problem is compared to other problems in AI? (N=147)

ResponseCountPercentage
Much easier117.5%
Easier2819.0%
As hard as other problems6141.5%
Harder3322.4%
Much harder149.5%

5. Results by Outlook Cluster

Section skipped — required data not available for this survey year. (KeyError: 'extinction_all_1')


Appendix: Cluster Analysis

Section skipped — required data not available for this survey year. (KeyError: 'Q369#1_1')


Appendix: Survey Flow

Survey-flow / randomization diagram modeled on Figure 16 of Grace et al., Thousands of AI Authors on the Future of AI. Each box is a question block. Percentages and n's are realized fill rates: the realized number of respondents who answered each block, divided by the total N. The total N counts everyone who answered at least one question (respondents who answered none are excluded).

Jobs / FAOL Sample-Size Reconciliation

The survey-flow boxes count anyone with any response in the broader Jobs / FAOL block. FAOL analyses use only the FAOL triplet, so their sample sizes can be slightly smaller.

Count definitionFixed-yearsFixed-probabilities
Whole Jobs / FAOL block (survey-flow box)4944
FAOL only, any triplet answer4944
FAOL only, complete triplet4943

Here, "fixed-years" means respondents gave probabilities for fixed time horizons (hj_b_full_*), while "fixed-probabilities" means respondents gave years for fixed probabilities (hj_a_full_*).

Tasks: ta_* columns are the fixed-probability framing, tb_* the fixed-years framing (verified empirically: ta_* rows have fixedprobabilities=1, tb_* rows have fixedprobabilities=0). The Qualtrics survey-flow definition (.qsf) is not in this repo, so the exact display order may differ from the diagram.


Appendix: Supplementary Figures

B.1 fixed-prob/fixed-year CDFs

Full CDF comparison between the fixed-prob and fixed-year conditions for HLMI and FAOL. The fixed-prob framing consistently produces earlier predictions (the red curve is shifted left relative to blue).

Corresponds to Figure 18 in Grace et al. (2024). Each curve is the mixture (mean) CDF for respondents assigned to that framing condition.

B.2 Bootstrap Confidence Bands

95% bootstrap confidence intervals on the aggregate CDF, obtained by resampling the fitted individual CDFs with replacement (500 resamples).

Milestone50% Year (point est.)95% Bootstrap CI
HLMI2061 (45.5 yr)2056–2068 (40.1–51.9 yr)
FAOL2139 (123.5 yr)2110–2190 (94.1–173.9 yr)

Bootstrap CIs reflect sampling variability — if we surveyed a different random sample of AI researchers, how much would the aggregate change? This is distinct from the spread of individual predictions (which is much wider).


Appendix: Statistical Tests

F.1 Yuen's Trimmed-Mean Bootstrap Test: Demographics

Tests whether experienced researchers (≥6 years in field) give significantly different HLMI predictions than junior researchers. Uses a 10% trimmed mean with 5,000 bootstrap resamples.

ComparisonTrimmed-mean diff95% CIp-valueSignificant (α=0.05)
HLMI 50% year (experienced − junior, split at 6 yr)+15.0 yr[-26.3, +91.4]0.473No

N (experienced): 41, N (junior): 32. Individual 50% years are computed from per-respondent gamma fits. The trimmed mean reduces sensitivity to extreme predictions.

F.2 Framing Effect: Statistical Significance

Tests whether year-framing and probability-framing respondents give significantly different predictions, using the same Yuen's bootstrap method.

MilestoneYear-framing medianProb-framing medianTrimmed-mean diff95% CIp-value
HLMI39.9 yr (n=124)52.8 yr (n=128)-24.9 yr[-103.3, +0.9]0.211
FAOL80.0 yr (n=43)130.5 yr (n=49)-278.6 yr[-973.2, +1808.6]0.707

A negative difference (yr_50 − pr_50) means the year-framing produces earlier predictions. Medians shown for reference; the statistical test uses trimmed means.


Appendix: Aggregation Method Comparison

Different choices in how individual survey responses are fitted and combined into aggregate forecasts can shift the headline numbers. This appendix compares three approaches on the key milestones, motivated by Adamczewski's reanalysis of ESPAI data.

Methods compared

Mix-Mean (MSE) — the Grace et al. 2023 method. Fit a gamma CDF to each respondent's 3 data points using mean squared error, then take the pointwise mean of all fitted CDFs.

Mix-Median (MSE) — same fitting, but take the pointwise median instead of mean. More robust to outlier predictions (e.g. respondents predicting 100M years). Adamczewski argues the median better represents the "typical expert."

Mix-Mean (LogLoss) — fit using log-loss (cross-entropy) instead of MSE, then take the mean. Log-loss naturally weighs errors near p=0 and p=1 more heavily than errors near p=0.5, matching the intuition that a 5% error at p=0.95 represents a much larger shift in belief than the same error at p=0.50.

Results

MilestoneMix-Mean (MSE)Mix-Median (MSE)Mix-Mean (LogLoss)N (MSE)N (LogLoss)
HLMI2061 (45.5 yr)2066 (49.6 yr)2061 (45.4 yr)252252
FAOL2139 (123.5 yr)2147 (130.5 yr)2152 (135.8 yr)9292
Truck Driver2028 (11.6 yr)2026 (10.2 yr)2027 (11.4 yr)8989
Surgeon2053 (37.4 yr)2055 (38.6 yr)2054 (37.6 yr)9191
Retail Salesperson2031 (14.6 yr)2030 (14.1 yr)2031 (14.7 yr)8989
AI Researcher2103 (87.2 yr)2097 (80.7 yr)2106 (90.2 yr)9090

Key takeaways

Aggregation method matters more than loss function. Switching from mean to median aggregation pushes HLMI from 2061 to 2066 and FAOL from 2139 to 2147. Switching the loss function (MSE→log-loss) barely changes the aggregate — consistent with Adamczewski's finding that fitting choice has "hardly any impact."

The effect is largest for far-future milestones (FAOL, AI Researcher) where a few very pessimistic respondents pull the mean CDF rightward.

See also: compare_methods.py for the full 5-method comparison across all 39 tasks, including linear interpolation and gamma-individual methods. Method comparison motivated by Adamczewski (bayes.net/espai).


Data: 2016 Expert Survey on Progress in AI (ESPAI).
Previous survey comparison values from "Thousands of AI Authors on the Future of AI" (2024 preprint).