1. Overview & Survey Methodology

531
Report Responses
738
Cleaned Analysis Rows
531
Finished
207
Incomplete (partial data kept)
361
Year Framing
377
Probability Framing

Randomization Blocks

BlockCountPercentage
Block 124332.9%
Block 223231.4%
Block 326335.6%

Data Cleaning Notes

The following cleaning steps were applied (see cleaning_log.md for full details):

Caveat — sum-to-100 triplets: 118 triplets across 65 respondents have values summing to exactly 100 (e.g., 10%, 30%, 60% at the three time horizons). This may indicate respondents who treated the probabilities as shares that must total 100%, rather than independent cumulative probabilities at each time horizon. Following standard methodology, sum-to-100 triplets from respondents with descending errors were also removed (77 total triplets NaN'd).


3.1 When Will 39 AI Milestones Be Feasible?

24
Tasks feasible within 10 years (50% prob)
39
Total milestones assessed

Corresponds to Figure 1 in Grace et al. (2024). Dots show the 50% probability year from the mixture CDF; horizontal lines show the 10%–90% range. The x-axis is capped at 2100; milestones whose 50% or 90% year falls beyond the cap are labeled with an arrow showing the true year. Includes 39 tasks, 4 occupations, HLMI, and FAOL.

Combined year-framing and probability-framing responses using gamma mixture CDF aggregation (matching the 2023 report methodology). A gamma CDF is fitted to each respondent's data, then all CDFs are averaged pointwise to produce a mixture CDF from which percentiles are read. Sorted by 50% probability year (earliest first). The HLMI row uses the dedicated direct-HLMI blocks (hb_a_*/hb_b_*); it does not pool the final-occupation prediction block.

MilestoneN (fitted)10% Prob Year50% Prob Year90% Prob Year
Play new Angry Birds levels (superhuman)790 yr (2022)3 yr (2025)14 yr (2036)
Beat best Starcraft 2 players670 yr (2022)3 yr (2025)16 yr (2038)
Write high-school history essay680 yr (2022)3 yr (2025)27 yr (2049)
Win World Series of Poker640 yr (2022)4 yr (2026)28 yr (2050)
Group unseen objects into classes580 yr (2022)5 yr (2027)59 yr (2081)
Voice acting from text620 yr (2022)5 yr (2027)26 yr (2048)
Write Python code (e.g. quicksort)560 yr (2022)5 yr (2027)34 yr (2056)
Outperform on all Atari games390 yr (2022)5 yr (2027)34 yr (2056)
Transcribe speech (noisy, accents)660 yr (2022)5 yr (2027)28 yr (2050)
Answer Googleable factoid questions671 yr (2023)6 yr (2028)32 yr (2054)
Answer Googleable open-ended questions580 yr (2022)6 yr (2028)27 yr (2049)
Produce song indistinguishable from artist550 yr (2022)6 yr (2028)29 yr (2051)
3D model from short video771 yr (2023)6 yr (2028)25 yr (2047)
Fold laundry (speed + quality)530 yr (2022)6 yr (2028)31 yr (2053)
Learn efficient sorting (no solution form)700 yr (2022)6 yr (2028)79 yr (2101)
Atari novice level (20 min training)571 yr (2023)6 yr (2028)33 yr (2055)
Fluent translation (most languages)551 yr (2023)7 yr (2029)45 yr (2067)
Phone banking services701 yr (2023)7 yr (2029)41 yr (2063)
Translate speech from subtitled films641 yr (2023)7 yr (2029)43 yr (2065)
Assemble any LEGO set731 yr (2023)7 yr (2029)35 yr (2057)
One-shot image recognition561 yr (2023)7 yr (2029)52 yr (2074)
Compose US Top 40 song (full audio)631 yr (2023)8 yr (2030)59 yr (2081)
Answer questions with no definite answer671 yr (2023)8 yr (2030)60 yr (2082)
Play random game as human novice (<10 min)681 yr (2023)8 yr (2030)37 yr (2059)
Explain game AI moves to layman592 yr (2024)10 yr (2032)97 yr (2119)
Win Putnam math competition641 yr (2023)12 yr (2034)102 yr (2124)
Truck Driver (occ.)1692 yr (2024)12 yr (2034)59 yr (2081)
Beat best Go players (limited training)702 yr (2024)12 yr (2034)137 yr (2159)
Beat fastest human in 5km city race (biped robot)622 yr (2024)12 yr (2034)77 yr (2099)
Translate newly discovered language (Rosetta stone)641 yr (2023)12 yr (2034)108 yr (2130)
Discover physics equations from simulation691 yr (2023)13 yr (2035)129 yr (2151)
Retail Salesperson (occ.)1691 yr (2023)14 yr (2036)89 yr (2111)
Write NYT best-seller novel562 yr (2024)16 yr (2038)207 yr (2229)
Prove publishable math theorems723 yr (2025)28 yr (2050)502 yr (2524)
Surgeon (occ.)1688 yr (2030)37 yr (2059)276 yr (2298)
HLMI (all human tasks)3587 yr (2029)37 yr (2059)262 yr (2284)
AI Researcher (occ.)16912 yr (2034)69 yr (2091)2268 yr (4290)
Full Automation of Labor16830 yr (2052)135 yr (2157)3544 yr (5566)
Conduct ML research & write conference paper0
Solve unsolved math problem (e.g. Millennium)0
Install electrical wiring in new home0
Replicate ML conference study0
Fine-tune open source LLM0
Build website with payment processing0
Find & patch security flaw (100k+ users)0

Year values are years from survey date (2022), with calendar year in parentheses. "10% prob year" = year at which the mixture CDF reaches 10%. "90% prob year" = 90%. Dashes indicate insufficient data or an aggregate CDF that does not reach the target percentile within the 1e8-year "never/infinity" sentinel bound.


3.2 HLMI & Full Automation of Labor Timing

High-Level Machine Intelligence (HLMI)

HLMI is defined as machines that can accomplish every task better and cheaper than human workers.

7 yr (2029)
10% probability
37 yr (2059)
50% probability
262 yr (2284)
90% probability
Percentile2022 Survey2023 SurveyShift
10% prob7 yr (2029)5 yr (2027)+1.9 yr
50% prob37 yr (2059)25 yr (2047)+12.3 yr

N (gamma fits): 358 (0 failed fits)

This HLMI estimate uses only the dedicated direct-HLMI question blocks (hb_a_* and hb_b_*). It does not pool the separate final-occupation prediction block (hj_*_final_pred).

Full Automation of Labor (FAOL)

All occupations fully automatable — machines carry out every task better and more cheaply than humans.

30 yr (2052)
10% probability
135 yr (2157)
50% probability
3544 yr (5566)
90% probability
Percentile2022 Survey2023 SurveyShift
50% prob135 yr (2157)94 yr (2116)+41.4 yr

N (gamma fits): 168 (0 failed fits)

Gap Between HLMI and FAOL

98 years
Gap at 50% probability

The consistent gap between HLMI and FAOL predictions reflects researchers' view that achieving human-level AI capability differs from fully automating all occupations.

Methodology Note

Gamma mixture CDF method: Following the methodology of Grace et al. (2024), a gamma CDF is fitted to each respondent's three data points (whether year-framing or probability-framing). All individual CDFs are then averaged pointwise to produce a mixture CDF, from which percentiles are read off. This approach handles extrapolation naturally (pessimistic respondents contribute heavy-tailed CDFs rather than being dropped) and treats both question framings symmetrically. Respondents were randomly assigned to either a "year framing" (provide years for 10%/50%/90% probability) or a "probability framing" (provide probabilities at fixed time horizons: 10, 20, and 40 years for HLMI; 10, 20, and 50 years for FAOL). See compare_methods.py for a comparison of this method with linear interpolation.

Corresponds to Figure 3 in Grace et al. (2024). Thin lines show individual respondent gamma CDFs (random subset of 200); thick line is the mixture (average) CDF. The dashed line marks 50% probability.


3.2.4 High-Level Machine Intelligence Timing by Experience

This table splits respondents at the median years in field, then refits the direct High-Level Machine Intelligence (HLMI) timing model within each experience group.

Experience groupDirect HLMI gamma-fit countMixture CDF 50% YearCalendar Year
More experienced (>=5 yr)7635.1 yr2057
Less experienced (<5 yr)4532.8 yr2055

The count values are not all respondents in each experience group. They count respondents who both answered the years-in-field question and gave enough direct HLMI timing answers to fit a gamma CDF. The separate occupation-automation prediction block is not pooled into this HLMI comparison. Only 180 respondents received the experience question due to randomization. Geographic region and citation count data are not available in the anonymized dataset.


3.2.5 Do Participants Agree on HLMI Timing?

"How much do you think your views on when HLMI will be achieved differ from those of the typical AI researcher?" (N=184)

ResponseCountPercentage
Not much7942.9%
A moderate amount9350.5%
A lot126.5%

3.3 Framing Effects

Respondents were randomly assigned to one of two question framings for the same underlying question. The "year framing" asked how many years until a given probability, while the "probability framing" asked for the probability at a fixed time horizon.

HLMI 50% Probability Year by Framing

FramingN (gamma fits)Mixture CDF 50% YearCalendar Year
Year framing (respondent provides years)17129.5 yr2052
Probability framing (respondent provides probabilities)18746.0 yr2068
16.5 years
Framing effect (difference)

A consistent framing effect has been observed across survey waves: the year-framing tends to produce earlier (shorter-timeline) predictions than the probability-framing. This is a documented cognitive bias in probability elicitation.


3.4 Perceived Rates of Progress

Respondents were asked whether AI progress was faster in the first or second half of their career. (N=180)

56.1%
Said second half was faster
5 yr
Median time in field (N=180)

Corresponds to Figure 4 in Grace et al. (2024).

ResponseCountPercentage2023 Survey
The second half10156.1%60%
The first half3217.8%
They were about the same4726.1%

How Far Along Is AI Progress? (Outside-view slider)

Respondents were shown a slider with three anchors: A = where progress was when they started working in their AI area; B = where it is now; and C = where it would need to be for AI software to have roughly human-level abilities at the tasks they study. They were asked "What fraction of the distance between where progress was when you started working in the area (A) and where it would need to be to attain human-level abilities in the area (C) have we come so far (B)?" on a 0–100 scale. This complements the year-based timing questions by anchoring an estimate to each respondent's own career window — so it is robust to disagreement about absolute calendar dates.

SurveyNMedianMeanSD
2022172 20 25.9 25.0

IQR (2022): 5–40.


3.5 What Causes AI Progress?

Respondents estimated how much less AI progress there would have been with half as much of each input. Higher values = more important to progress. Values are percentage less progress (0-100%).

Corresponds to Figure 5 in Grace et al. (2024). Red dots are means; box shows IQR with median line.

FactorNMedianMeanIQR
Computing hardware cost decline10760.0%61.6%50–80%
Training dataset effort10850.0%48.4%30–70%
AI algorithm progress10050.0%44.3%30–60%
Funding10935.0%42.3%25–50%
Researcher effort11125.0%34.1%20–40%

3.6 Will There Be an Intelligence Explosion?

Probability Estimates (numeric, 0-100%)

ScenarioNMedianMeanIQR2023 Survey MedianChange
Dramatic tech speedup within 2 years of HLMI33920%29.5%5–50%20%unchanged
Dramatic tech speedup within 30 years of HLMI34080%74.0%50–95%80%unchanged
Vastly superhuman AI within 2 years of HLMI36510%19.1%1–25%10%unchanged
Vastly superhuman AI within 30 years of HLMI36560%55.9%20–90%

Is the Feedback Loop Argument Broadly Correct? (N=386)

"If AI does nearly all R&D, improvements in AI will accelerate progress including further AI progress. This could cause progress to become >10x faster within 5 years."

Corresponds to Figure 6 in Grace et al. (2024).

ResponseCountPercentage
Quite unlikely (0-20%)7719.9%
Unlikely (21-40%)10426.9%
About even chance (41-60%)7820.2%
Likely (61-80%)9925.6%
Quite likely (81-100%)287.3%

3.7 AI capabilities 2044

Section skipped — required data not available for this survey year. (KeyError: 'Q1#1_1')


3.8 Explainability

Section skipped — required data not available for this survey year. (KeyError: 'Q1 (1)')


4.1 Concerning scenarios

Section skipped — required data not available for this survey year. (KeyError: 'Q369#1_1')


4.2 How Good or Bad Will HLMI Be?

Respondents assigned probabilities to five outcome categories (summing to 100%). N=559

62.4%
Net optimists
24.0%
Net pessimists
13.6%
Balanced

Corresponds to Figure 11 in Grace et al. (2024), "Thousands of AI Authors on the Future of AI."

Mean and Median Probabilities

OutcomeMeanMedian
Extremely good24.1%10.0%
On balance good26.4%20.0%
Neutral18.3%15.0%
On balance bad17.0%10.0%
Extremely bad14.1%5.0%

Individual Respondent Views

Each vertical slice below is one respondent's probability allocation, sorted from most optimistic (left) to most pessimistic (right). The chart reveals the full diversity of expert opinion.

Corresponds to Figure 10 in Grace et al. (2024).

The same respondents sorted by how much probability they assign to "extremely bad (e.g. human extinction)" outcomes. The growing black band on the right shows those assigning the highest probability to extremely bad outcomes.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

Individual Response Profiles

Each small bar below represents one respondent's five-category probability distribution, sampled evenly from the most optimistic to the most pessimistic.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

Extreme Outcome Analysis

Metric2022 Survey2023 SurveyChange
Non-zero to both extremes68.3%64.0%+4.3 pp
>=5% on extremely bad66.7%57.8%+8.9 pp
>=10% on extremely bad47.6%37.8%+9.8 pp
>=20% on extremely bad27.0%
>=25% on extremely bad17.7%
Mean extremely bad14.1%9.0%+5.1 pp
Median extremely bad5.0%5.0%unchanged
Net optimists62.4%68.3%-5.9 pp

4.3 Extinction risk

Section skipped — required data not available for this survey year. (KeyError: 'extinction_100_1')


4.4 Are Future AI-Risk Concerns Due to Misunderstandings of AI Research?

"To what extent do you think people's concerns about future risks from AI are due to misunderstandings of AI research?" (N=185)

No direct figure equivalent in Grace et al. (2024); this data is discussed in Section 4.4 of that paper.

ResponseCountPercentage
Hardly at all52.7%
Not much2513.5%
Somewhat4624.9%
To a large extent8948.1%
Almost entirely2010.8%

4.5 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity

This question was not included in this survey year.


4.6 How Much Should AI Safety Research Be Prioritized?

"How much should society prioritize AI safety research, relative to how much it is currently prioritized?" (N=263)

68.8%
Say 'More' or 'Much more'
70%
2023 Survey (More+Much more)

Corresponds to Figure 14 in Grace et al. (2024).

ResponseCountPercentage
Much less51.9%
Less249.1%
About the same5320.2%
More9335.4%
Much more8833.5%

4.7 The Alignment Problem

Corresponds to Figure 15 in Grace et al. (2024).

Do you think this argument points at an important problem? (N=272)

ResponseCountPercentage
No, not a real problem.114.0%
No, not an important problem.3814.0%
Yes, a moderately important problem.6423.5%
Yes, a very important problem.10137.1%
Yes, among the most important problems in the field.5821.3%

How valuable is it to work on this problem today, compared to other problems in AI? (N=272)

ResponseCountPercentage
Much less valuable279.9%
Less valuable8230.1%
As valuable as other problems8932.7%
More valuable5319.5%
Much more valuable217.7%

How hard do you think this problem is compared to other problems in AI? (N=272)

ResponseCountPercentage
Much easier134.8%
Easier248.8%
As hard as other problems7828.7%
Harder8531.2%
Much harder7226.5%

5. Results by Outlook Cluster

Section skipped — required data not available for this survey year. (KeyError: 'Q369#1_1')


Appendix: Cluster Analysis

Section skipped — required data not available for this survey year. (KeyError: 'Q369#1_1')


Appendix: Survey Flow

Survey-flow / randomization diagram modeled on Figure 16 of Grace et al., Thousands of AI Authors on the Future of AI. Each box is a question block. Percentages and n's are realized fill rates: the realized number of respondents who answered each block, divided by the total N. The total N counts everyone who answered at least one question (respondents who answered none are excluded).

Jobs / FAOL Sample-Size Reconciliation

The survey-flow boxes count anyone with any response in the broader Jobs / FAOL block. FAOL analyses use only the FAOL triplet, so their sample sizes can be slightly smaller.

Count definitionFixed-yearsFixed-probabilities
Whole Jobs / FAOL block (survey-flow box)8981
FAOL only, any triplet answer8881
FAOL only, complete triplet8880

Here, "fixed-years" means respondents gave probabilities for fixed time horizons (hj_b_full_*), while "fixed-probabilities" means respondents gave years for fixed probabilities (hj_a_full_*).

Tasks: ta_* columns are the fixed-probability framing, tb_* the fixed-years framing (verified empirically: ta_* rows have fixedprobabilities=1, tb_* rows have fixedprobabilities=0). The Qualtrics survey-flow definition (.qsf) is not in this repo, so the exact display order may differ from the diagram.


Appendix: Supplementary Figures

B.1 fixed-prob/fixed-year CDFs

Full CDF comparison between the fixed-prob and fixed-year conditions for HLMI and FAOL. The fixed-prob framing consistently produces earlier predictions (the red curve is shifted left relative to blue).

Corresponds to Figure 18 in Grace et al. (2024). Each curve is the mixture (mean) CDF for respondents assigned to that framing condition.

B.2 Bootstrap Confidence Bands

95% bootstrap confidence intervals on the aggregate CDF, obtained by resampling the fitted individual CDFs with replacement (500 resamples).

Milestone50% Year (point est.)95% Bootstrap CI
HLMI2059 (37.3 yr)2056–2063 (33.7–41.4 yr)
FAOL2157 (135.4 yr)2130–2199 (108.1–176.8 yr)

Bootstrap CIs reflect sampling variability — if we surveyed a different random sample of AI researchers, how much would the aggregate change? This is distinct from the spread of individual predictions (which is much wider).


Appendix: Statistical Tests

F.1 Yuen's Trimmed-Mean Bootstrap Test: Demographics

Tests whether experienced researchers (≥5 years in field) give significantly different HLMI predictions than junior researchers. Uses a 10% trimmed mean with 5,000 bootstrap resamples.

ComparisonTrimmed-mean diff95% CIp-valueSignificant (α=0.05)
HLMI 50% year (experienced − junior, split at 5 yr)+6.1 yr[-57.7, +28.9]0.630No

N (experienced): 76, N (junior): 45. Individual 50% years are computed from per-respondent gamma fits. The trimmed mean reduces sensitivity to extreme predictions.

F.2 Framing Effect: Statistical Significance

Tests whether year-framing and probability-framing respondents give significantly different predictions, using the same Yuen's bootstrap method.

MilestoneYear-framing medianProb-framing medianTrimmed-mean diff95% CIp-value
HLMI30.7 yr (n=171)44.1 yr (n=187)-31.8 yr[-108.0, -15.6]0.105
FAOL102.2 yr (n=80)308.4 yr (n=88)-917.4 yr[-1249.6, -593.5]0.002

A negative difference (yr_50 − pr_50) means the year-framing produces earlier predictions. Medians shown for reference; the statistical test uses trimmed means.


Appendix: Aggregation Method Comparison

Different choices in how individual survey responses are fitted and combined into aggregate forecasts can shift the headline numbers. This appendix compares three approaches on the key milestones, motivated by Adamczewski's reanalysis of ESPAI data.

Methods compared

Mix-Mean (MSE) — the Grace et al. 2023 method. Fit a gamma CDF to each respondent's 3 data points using mean squared error, then take the pointwise mean of all fitted CDFs.

Mix-Median (MSE) — same fitting, but take the pointwise median instead of mean. More robust to outlier predictions (e.g. respondents predicting 100M years). Adamczewski argues the median better represents the "typical expert."

Mix-Mean (LogLoss) — fit using log-loss (cross-entropy) instead of MSE, then take the mean. Log-loss naturally weighs errors near p=0 and p=1 more heavily than errors near p=0.5, matching the intuition that a 5% error at p=0.95 represents a much larger shift in belief than the same error at p=0.50.

Results

MilestoneMix-Mean (MSE)Mix-Median (MSE)Mix-Mean (LogLoss)N (MSE)N (LogLoss)
HLMI2059 (37.3 yr)2062 (39.9 yr)2060 (38.1 yr)358358
FAOL2157 (135.4 yr)2168 (146.1 yr)2166 (144.5 yr)168168
Truck Driver2034 (11.7 yr)2032 (10.3 yr)2034 (11.6 yr)169169
Surgeon2059 (36.8 yr)2061 (38.8 yr)2059 (37.3 yr)168168
Retail Salesperson2036 (13.7 yr)2037 (14.9 yr)2036 (13.7 yr)169169
AI Researcher2091 (68.5 yr)2091 (68.9 yr)2091 (69.3 yr)169169

Key takeaways

Aggregation method matters more than loss function. Switching from mean to median aggregation pushes HLMI from 2059 to 2062 and FAOL from 2157 to 2168. Switching the loss function (MSE→log-loss) barely changes the aggregate — consistent with Adamczewski's finding that fitting choice has "hardly any impact."

The effect is largest for far-future milestones (FAOL, AI Researcher) where a few very pessimistic respondents pull the mean CDF rightward.

See also: compare_methods.py for the full 5-method comparison across all 39 tasks, including linear interpolation and gamma-individual methods. Method comparison motivated by Adamczewski (bayes.net/espai).


Data: 2022 Expert Survey on Progress in AI (ESPAI).
Previous survey comparison values from "Thousands of AI Authors on the Future of AI" (2024 preprint).