1. Overview & Survey Methodology

1,580
Report Responses
1,793
Cleaned Analysis Rows
1,502
Finished
291
Incomplete (partial data kept)
917
Year Framing
876
Probability Framing

Randomization Blocks

BlockCountPercentage
Block 159633.2%
Block 259133.0%
Block 360633.8%

Data Cleaning Notes

The following cleaning steps were applied (see cleaning_log.md for full details):

Caveat — sum-to-100 triplets: 393 triplets across 159 respondents have values summing to exactly 100 (e.g., 10%, 30%, 60% at the three time horizons). This may indicate respondents who treated the probabilities as shares that must total 100%, rather than independent cumulative probabilities at each time horizon. Following standard methodology, sum-to-100 triplets from respondents with descending errors were also removed (368 total triplets NaN'd).


3.1 When Will 39 AI Milestones Be Feasible?

34
Tasks feasible within 10 years (50% prob)
39
Total milestones assessed

Corresponds to Figure 1 in Grace et al. (2024). Dots show the 50% probability year from the mixture CDF; horizontal lines show the 10%–90% range. The x-axis is capped at 2100; milestones whose 50% or 90% year falls beyond the cap are labeled with an arrow showing the true year. Includes 39 tasks, 4 occupations, HLMI, and FAOL.

Combined year-framing and probability-framing responses using gamma mixture CDF aggregation (matching the 2023 report methodology). A gamma CDF is fitted to each respondent's data, then all CDFs are averaged pointwise to produce a mixture CDF from which percentiles are read. Sorted by 50% probability year (earliest first). The HLMI row uses the dedicated direct-HLMI blocks (hb_a_*/hb_b_*); it does not pool the final-occupation prediction block.

MilestoneN (2024, fitted)2023 10%2024 10%Δ10 (cal yr)2023 50%2024 50%Δ50 (cal yr)2023 90%2024 90%Δ90 (cal yr)
Write Python code (e.g. quicksort)1110 yr (2023)0 yr (2024)+12 yr (2025)2 yr (2026)+114 yr (2037)11 yr (2035)-2
Write high-school history essay1290 yr (2023)0 yr (2024)+12 yr (2025)2 yr (2026)+113 yr (2036)10 yr (2034)-2
Play new Angry Birds levels (superhuman)1670 yr (2023)0 yr (2024)+12 yr (2025)2 yr (2026)+110 yr (2033)10 yr (2034)+1
Win World Series of Poker1390 yr (2023)0 yr (2024)+13 yr (2026)2 yr (2026)±020 yr (2043)15 yr (2039)-3
Answer Googleable factoid questions1440 yr (2023)0 yr (2024)+13 yr (2026)3 yr (2027)+121 yr (2044)22 yr (2046)+2
Group unseen objects into classes1460 yr (2023)0 yr (2024)+14 yr (2027)3 yr (2027)±023 yr (2046)19 yr (2043)-3
Voice acting from text1410 yr (2023)0 yr (2024)+13 yr (2026)3 yr (2027)+117 yr (2040)14 yr (2038)-2
Fluent translation (most languages)1580 yr (2023)0 yr (2024)+13 yr (2026)3 yr (2027)±016 yr (2039)18 yr (2042)+3
Answer Googleable open-ended questions1570 yr (2023)0 yr (2024)+13 yr (2026)3 yr (2027)+123 yr (2046)23 yr (2047)+1
Transcribe speech (noisy, accents)1350 yr (2023)0 yr (2024)+13 yr (2026)3 yr (2027)+117 yr (2040)16 yr (2040)-1
Beat best Starcraft 2 players1470 yr (2023)0 yr (2024)+14 yr (2027)3 yr (2027)±025 yr (2048)18 yr (2042)-7
3D model from short video1510 yr (2023)0 yr (2024)+15 yr (2028)3 yr (2027)-127 yr (2050)18 yr (2042)-8
Answer questions with no definite answer1460 yr (2023)0 yr (2024)+14 yr (2027)4 yr (2028)±035 yr (2058)20 yr (2044)-14
Produce song indistinguishable from artist1490 yr (2023)0 yr (2024)+14 yr (2027)4 yr (2028)+126 yr (2049)20 yr (2044)-5
Build website with payment processing1340 yr (2023)0 yr (2024)+15 yr (2028)4 yr (2028)±031 yr (2054)20 yr (2044)-9
Translate speech from subtitled films1200 yr (2023)0 yr (2024)+15 yr (2028)4 yr (2028)±041 yr (2064)24 yr (2048)-16
Atari novice level (20 min training)1480 yr (2023)0 yr (2024)+15 yr (2028)4 yr (2028)±036 yr (2059)32 yr (2056)-3
Phone banking services1430 yr (2023)0 yr (2024)+15 yr (2028)4 yr (2028)±030 yr (2053)26 yr (2050)-4
Outperform on all Atari games1500 yr (2023)0 yr (2024)+16 yr (2029)5 yr (2029)±035 yr (2058)25 yr (2049)-10
Fine-tune open source LLM1391 yr (2024)0 yr (2024)+15 yr (2028)5 yr (2029)+147 yr (2070)29 yr (2053)-16
Compose US Top 40 song (full audio)1480 yr (2023)0 yr (2024)+16 yr (2029)5 yr (2029)±046 yr (2069)47 yr (2071)+2
Win Putnam math competition1581 yr (2024)0 yr (2024)±08 yr (2031)5 yr (2029)-271 yr (2094)31 yr (2055)-39
Translate newly discovered language (Rosetta stone)1221 yr (2024)0 yr (2024)±07 yr (2030)6 yr (2030)±052 yr (2075)55 yr (2079)+4
Play random game as human novice (<10 min)1401 yr (2024)1 yr (2025)+17 yr (2030)6 yr (2030)±038 yr (2061)38 yr (2062)±0
Fold laundry (speed + quality)1621 yr (2024)1 yr (2025)+17 yr (2030)6 yr (2030)±040 yr (2063)35 yr (2059)-5
Learn efficient sorting (no solution form)1351 yr (2024)0 yr (2024)±07 yr (2030)6 yr (2030)+155 yr (2078)51 yr (2075)-3
Write NYT best-seller novel1551 yr (2024)1 yr (2025)+17 yr (2030)7 yr (2031)±062 yr (2085)65 yr (2089)+4
One-shot image recognition1340 yr (2023)1 yr (2025)+15 yr (2028)7 yr (2031)+234 yr (2057)37 yr (2061)+5
Explain game AI moves to layman1501 yr (2024)1 yr (2025)+18 yr (2031)7 yr (2031)±081 yr (2104)48 yr (2072)-32
Beat best Go players (limited training)1481 yr (2024)1 yr (2025)+110 yr (2033)7 yr (2031)-1100 yr (2123)88 yr (2112)-11
Assemble any LEGO set1411 yr (2024)1 yr (2025)+18 yr (2031)7 yr (2031)+148 yr (2071)32 yr (2056)-15
Retail Salesperson (occ.)4601 yr (2024)1 yr (2025)±010 yr (2033)7 yr (2031)-175 yr (2098)68 yr (2092)-6
Find & patch security flaw (100k+ users)1352 yr (2025)1 yr (2025)+110 yr (2033)7 yr (2031)-187 yr (2110)63 yr (2087)-23
Replicate ML conference study1402 yr (2025)2 yr (2026)+112 yr (2035)8 yr (2032)-3109 yr (2132)53 yr (2077)-56
Truck Driver (occ.)4592 yr (2025)1 yr (2025)±012 yr (2035)9 yr (2033)-258 yr (2081)46 yr (2070)-11
Beat fastest human in 5km city race (biped robot)1251 yr (2024)1 yr (2025)+19 yr (2032)9 yr (2033)+144 yr (2067)40 yr (2064)-3
Discover physics equations from simulation1321 yr (2024)1 yr (2025)+112 yr (2035)11 yr (2035)±0123 yr (2146)127 yr (2151)+5
Conduct ML research & write conference paper1642 yr (2025)2 yr (2026)±020 yr (2043)12 yr (2036)-7246 yr (2269)159 yr (2183)-86
Prove publishable math theorems1493 yr (2026)2 yr (2026)±023 yr (2046)15 yr (2039)-7270 yr (2293)258 yr (2282)-11
Install electrical wiring in new home1303 yr (2026)3 yr (2027)+117 yr (2040)16 yr (2040)±0104 yr (2127)96 yr (2120)-7
HLMI (all human tasks)9854 yr (2027)3 yr (2027)-124 yr (2047)18 yr (2042)-5174 yr (2197)163 yr (2187)-10
Surgeon (occ.)4587 yr (2030)5 yr (2029)±033 yr (2056)26 yr (2050)-6310 yr (2333)299 yr (2323)-9
AI Researcher (occ.)4586 yr (2029)4 yr (2028)-140 yr (2063)28 yr (2052)-111321 yr (3344)583 yr (2607)-737
Solve unsolved math problem (e.g. Millennium)1304 yr (2027)4 yr (2028)±027 yr (2050)30 yr (2054)+3329 yr (2352)766 yr (2790)+437
Full Automation of Labor45614 yr (2037)11 yr (2035)-189 yr (2112)72 yr (2096)-162641 yr (4664)2195 yr (4219)-445

Year values are years from survey date (2024), with calendar year in parentheses. "10% prob year" = year at which the mixture CDF reaches 10%. "90% prob year" = 90%. Dashes indicate insufficient data or an aggregate CDF that does not reach the target percentile within the 1e8-year "never/infinity" sentinel bound. Δ columns show shift in calendar prediction year between 2024 and 2023 surveys (positive = pushed later; red ≥ +3 yr, green ≤ −3 yr).


3.2 HLMI & Full Automation of Labor Timing

High-Level Machine Intelligence (HLMI)

HLMI is defined as machines that can accomplish every task better and cheaper than human workers.

3 yr (2027)
10% probability
18 yr (2042)
50% probability
163 yr (2187)
90% probability
Percentile2024 Survey2023 SurveyShift
10% prob3 yr (2027)3 yr (2027)-0.4 yr
50% prob18 yr (2042)23 yr (2047)-4.6 yr

N (gamma fits): 985 (0 failed fits)

This HLMI estimate uses only the dedicated direct-HLMI question blocks (hb_a_* and hb_b_*). It does not pool the separate final-occupation prediction block (hj_*_final_pred).

Full Automation of Labor (FAOL)

All occupations fully automatable — machines carry out every task better and more cheaply than humans.

11 yr (2035)
10% probability
72 yr (2096)
50% probability
2195 yr (4219)
90% probability
Percentile2024 Survey2023 SurveyShift
50% prob72 yr (2096)92 yr (2116)-19.6 yr

N (gamma fits): 456 (0 failed fits)

Gap Between HLMI and FAOL

54 years
Gap at 50% probability

The consistent gap between HLMI and FAOL predictions reflects researchers' view that achieving human-level AI capability differs from fully automating all occupations.

Methodology Note

Gamma mixture CDF method: Following the methodology of Grace et al. (2024), a gamma CDF is fitted to each respondent's three data points (whether year-framing or probability-framing). All individual CDFs are then averaged pointwise to produce a mixture CDF, from which percentiles are read off. This approach handles extrapolation naturally (pessimistic respondents contribute heavy-tailed CDFs rather than being dropped) and treats both question framings symmetrically. Respondents were randomly assigned to either a "year framing" (provide years for 10%/50%/90% probability) or a "probability framing" (provide probabilities at fixed time horizons: 10, 20, and 40 years for HLMI; 10, 20, and 50 years for FAOL). See compare_methods.py for a comparison of this method with linear interpolation.

Corresponds to Figure 3 in Grace et al. (2024). Thin lines show individual respondent gamma CDFs (random subset of 200); thick line is the mixture (average) CDF. The dashed line marks 50% probability.


3.2.4 High-Level Machine Intelligence Timing by Experience

This table splits respondents at the median years in field, then refits the direct High-Level Machine Intelligence (HLMI) timing model within each experience group.

Experience groupDirect HLMI gamma-fit countMixture CDF 50% YearCalendar Year
More experienced (>=6 yr)5818.2 yr2042
Less experienced (<6 yr)5513.4 yr2037

The count values are not all respondents in each experience group. They count respondents who both answered the years-in-field question and gave enough direct HLMI timing answers to fit a gamma CDF. The separate occupation-automation prediction block is not pooled into this HLMI comparison. Only 189 respondents received the experience question due to randomization. Geographic region and citation count data are not available in the anonymized dataset.


3.2.5 Do Participants Agree on HLMI Timing?

"How much do you think your views on when HLMI will be achieved differ from those of the typical AI researcher?" (N=385)

ResponseCountPercentage
Not much16943.9%
A moderate amount17445.2%
A lot4210.9%

3.3 Framing Effects

Respondents were randomly assigned to one of two question framings for the same underlying question. The "year framing" asked how many years until a given probability, while the "probability framing" asked for the probability at a fixed time horizon.

HLMI 50% Probability Year by Framing

FramingN (gamma fits)Mixture CDF 50% YearCalendar Year
Year framing (respondent provides years)52513.8 yr2038
Probability framing (respondent provides probabilities)46026.1 yr2050
12.3 years
Framing effect (difference)

A consistent framing effect has been observed across survey waves: the year-framing tends to produce earlier (shorter-timeline) predictions than the probability-framing. This is a documented cognitive bias in probability elicitation.


3.4 Perceived Rates of Progress

Respondents were asked whether AI progress was faster in the first or second half of their career. (N=191)

58.6%
Said second half was faster
6 yr
Median time in field (N=189)

Corresponds to Figure 4 in Grace et al. (2024).

ResponseCountPercentage2023 Survey
The second half11258.6%60%
The first half3618.8%
They were about the same4322.5%

How Far Along Is AI Progress? (Outside-view slider)

Respondents were shown a slider with three anchors: A = where progress was when they started working in their AI area; B = where it is now; and C = where it would need to be for AI software to have roughly human-level abilities at the tasks they study. They were asked "What fraction of the distance between where progress was when you started working in the area (A) and where it would need to be to attain human-level abilities in the area (C) have we come so far (B)?" on a 0–100 scale. This complements the year-based timing questions by anchoring an estimate to each respondent's own career window — so it is robust to disagreement about absolute calendar dates.

SurveyNMedianMeanSD
2024188 30 38.0 28.4
20233193037.626.0

IQR (2024): 14–60. The question was also asked in 2023 (raw values clamped to 0–100; one 2023 respondent entered an out-of-range value). Median and mean barely moved year-on-year, suggesting the field's self-assessment of remaining distance to HLMI is stable.


3.5 What Causes AI Progress?

Respondents estimated how much less AI progress there would have been with half as much of each input. Higher values = more important to progress. Values are percentage less progress (0-100%).

Corresponds to Figure 5 in Grace et al. (2024). Red dots are means; box shows IQR with median line.

FactorNMedianMeanIQR
Computing hardware cost decline11960.0%58.3%45–80%
Training dataset effort10950.0%51.9%30–75%
AI algorithm progress10350.0%46.1%30–60%
Funding11640.0%44.1%25–60%
Researcher effort9525.0%32.7%20–50%

3.6 Will There Be an Intelligence Explosion?

Probability Estimates (numeric, 0-100%)

ScenarioNMedianMeanIQR2023 Survey MedianChange
Dramatic tech speedup within 2 years of HLMI17320%30.3%10–50%20%unchanged
Dramatic tech speedup within 30 years of HLMI17280%71.6%50–95%80%unchanged
Vastly superhuman AI within 2 years of HLMI17110%20.9%1–30%10%unchanged
Vastly superhuman AI within 30 years of HLMI17260%59.6%30–95%

Is the Feedback Loop Argument Broadly Correct? (N=189)

"If AI does nearly all R&D, improvements in AI will accelerate progress including further AI progress. This could cause progress to become >10x faster within 5 years."

Corresponds to Figure 6 in Grace et al. (2024).

ResponseCountPercentage
Quite unlikely (0-20%)4021.2%
Unlikely (21-40%)5227.5%
About even chance (41-60%)4423.3%
Likely (61-80%)4523.8%
Quite likely (81-100%)84.2%

3.7 AI Capabilities in 2044

"In 2044, how likely do you think the following will be for at least some state-of-the-art AI systems?" Sorted by percentage rating Likely or Very Likely.

Corresponds to Figure 7 in Grace et al. (2024).

CapabilityNV.UnlikelyUnlikelyEvenLikelyV.LikelyLikely+V.Likely
Talk like an expert human on most topics3781%5%10%23%61%84%
Find unexpected ways to achieve goals3802%4%11%32%51%83%
Frequently behave surprisingly to humans3802%11%23%32%32%64%
Can be jailbroken for illegal commands3745%10%23%36%27%62%
Deceive humans to achieve goals (unintended)3769%16%26%31%19%49%
Cause important real-world actions (run business, etc.)3798%19%24%30%18%49%
Have goals not aligned with human goals37510%22%24%31%13%44%
Form AI-AI collaborative relationships (unintended)3799%21%26%26%17%44%
Can be trusted to explain their actions3808%24%29%28%11%39%
Self-improve regardless of human wishes37812%22%28%25%13%38%
Take actions to attain power37721%28%29%16%7%23%

3.8 Will AI Explain Its Decisions? (2029)

"For typical state-of-the-art AI systems in 2029, will users be able to know the true reasons for decisions?" (N=520)

21.7%
Rate it Likely or Very Likely

Corresponds to Figure 8 in Grace et al. (2024).

ResponseCountPercentage
Very unlikely (<10%)13826.5%
Unlikely (10-40%)15730.2%
Even odds (40-60%)11221.5%
Likely (60-90%)8015.4%
Very likely (>90%)336.3%

4.1 How Concerning Are Future AI Scenarios?

"How would you rate the level of concern these scenarios deserve over the next thirty years?" Sorted by percentage rating Substantial or Extreme concern.

Corresponds to Figure 9 in Grace et al. (2024).

ScenarioNNoneA LittleSubstantialExtreme2024 Sub.+Ext.2023 Sub.+Ext.Δ Sub.+Ext. (pp)
AI makes it easy to spread false info (deepfakes)7592%15%32%51%83%86%-2
AI manipulates large-scale public opinion7574%18%36%41%78%78%-1
AI lets dangerous groups make powerful tools (bioweapons)7564%24%40%32%72%73%-1
Authoritarian rulers use AI for control7547%23%33%36%69%73%-3
AI worsens economic inequality7537%25%35%34%69%71%-3
AI bias worsens unjust situations (hiring, etc.)7578%35%38%19%57%61%-4
Less human interaction (more time with AI)75914%37%31%17%49%45%+3
Automation leaves most people economically powerless75316%37%30%18%48%46%+2
Wrong goals: AI reduces human decision-making role75616%38%30%16%46%44%+2
Misaligned powerful AI causes catastrophe (weapons)75519%38%25%18%43%43%+1
Automation makes people struggle to find meaning76026%38%25%11%37%35%+1

4.2 How Good or Bad Will HLMI Be?

Respondents assigned probabilities to five outcome categories (summing to 100%). N=1538

67.0%
Net optimists
22.2%
Net pessimists
10.7%
Balanced

Corresponds to Figure 11 in Grace et al. (2024), "Thousands of AI Authors on the Future of AI."

Mean and Median Probabilities

OutcomeMeanMedian
Extremely good23.9%15.0%
On balance good27.7%25.0%
Neutral20.7%20.0%
On balance bad17.8%15.0%
Extremely bad9.9%5.0%

Individual Respondent Views

Each vertical slice below is one respondent's probability allocation, sorted from most optimistic (left) to most pessimistic (right). The chart reveals the full diversity of expert opinion.

Corresponds to Figure 10 in Grace et al. (2024).

The same respondents sorted by how much probability they assign to "extremely bad (e.g. human extinction)" outcomes. The growing black band on the right shows those assigning the highest probability to extremely bad outcomes.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

Individual Response Profiles

Each small bar below represents one respondent's five-category probability distribution, sampled evenly from the most optimistic to the most pessimistic.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

Extreme Outcome Analysis

Metric2024 Survey2023 SurveyChange
Non-zero to both extremes64.4%64.0%+0.4 pp
>=5% on extremely bad58.3%57.8%+0.5 pp
>=10% on extremely bad40.5%37.8%+2.7 pp
>=20% on extremely bad20.0%
>=25% on extremely bad11.9%
Mean extremely bad9.9%9.0%+0.9 pp
Median extremely bad5.0%5.0%unchanged
Net optimists67.0%68.3%-1.3 pp

4.3 How Likely Is AI to Cause Extinction or Severe Disempowerment?

Each respondent was randomly assigned ONE of three questions about the probability that future AI advances cause human extinction or similarly permanent and severe disempowerment. The chart below also includes "extremely bad" HLMI outcome probabilities from Section 4.2 for comparison.

Corresponds to Figure 13 in Grace et al. (2024).

Note: The "Extremely bad HLMI outcome" question was asked to all respondents (N is larger), while each extinction/disempowerment question was asked to a random subset of respondents.

Distribution of Individual Estimates

The 2024 estimates were higher for the unconditional and within-100-years questions, while estimates for the loss-of-control question were similar to 2023.

The filled curves show 2024 responses; dashed lines show 2023. Respondents are ranked independently within each year and question framing, so the horizontal axis is percentile rather than respondent count. Each question was asked to a different random subset of respondents.

Corresponds to Figure 12 in Grace et al. (2024).

QuestionNMean2023 Survey MeanChangeMedian2023 Survey MedianChange>=10%>=25%
Future AI causes extinction or severe disempowerment74418.3%16.2%+2.1 pp10.0%5.0%+5.0 pp52.7%26.6%
Inability to control advanced AI causes extinction/disempowerment39218.5%19.4%-0.9 pp9.0%10.0%-1.0 pp50.0%26.8%
AI-caused extinction/disempowerment within 100 years35317.5%14.4%+3.1 pp5.0%5.0%unchanged49.0%25.8%
Pooled (all variants combined)148918.2%10.0%51.1%26.5%

The "Pooled" row combines the raw responses from all variants above (744 + 392 + 353 = 1489 responses) and takes the median of the combined data. Because each respondent was randomly assigned exactly one variant, no respondent is counted twice. And because the other variants each add a constraint (a specific cause, a time limit) to the basic question, every response is a lower bound on that respondent's unconstrained probability — so the pooled median is a conservative basis for "the median researcher put at least 10%..." claims.


4.4 Are Future AI-Risk Concerns Due to Misunderstandings of AI Research?

"To what extent do you think people's concerns about future risks from AI are due to misunderstandings of AI research?" (N=386)

No direct figure equivalent in Grace et al. (2024); this data is discussed in Section 4.4 of that paper.

ResponseCountPercentage
Hardly at all102.6%
Not much6115.8%
Somewhat12532.4%
To a large extent16342.2%
Almost entirely277.0%

4.5 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity

Question wording: "What rate of global AI progress over the next five years would make you feel most optimistic for humanity's future? Assume any change in speed affects all projects equally." (N=382)

Corresponds to Table 3 in Grace et al. (2024); no figure equivalent in that paper.

ResponseCountPercentage2023 SurveyChange
Much slower348.9%4.8%+4.1 pp
Somewhat slower9524.9%29.9%-5.0 pp
Current speed11229.3%26.9%+2.4 pp
Somewhat faster7519.6%22.8%-3.2 pp
Much faster5614.7%15.6%-0.9 pp
Other102.6%

4.6 How Much Should AI Safety Research Be Prioritized?

"How much should society prioritize AI safety research, relative to how much it is currently prioritized?" (N=370)

70.8%
Say 'More' or 'Much more'
70%
2023 Survey (More+Much more)

Corresponds to Figure 14 in Grace et al. (2024).

ResponseCountPercentage
Much less82.2%
Less246.5%
About the same7620.5%
More13135.4%
Much more13135.4%

4.7 The Alignment Problem

Corresponds to Figure 15 in Grace et al. (2024).

Do you think this argument points at an important problem? (N=755)

ResponseCountPercentage
No, not a real problem.314.1%
No, not an important problem.9312.3%
Yes, a moderately important problem.26635.2%
Yes, a very important problem.27836.8%
Yes, among the most important problems in the field.8711.5%

How valuable is it to work on this problem today, compared to other problems in AI? (N=753)

ResponseCountPercentage
Much less valuable628.2%
Less valuable17823.6%
As valuable as other problems27035.9%
More valuable18925.1%
Much more valuable547.2%

How hard do you think this problem is compared to other problems in AI? (N=752)

ResponseCountPercentage
Much easier152.0%
Easier749.8%
As hard as other problems24132.0%
Harder26635.4%
Much harder15620.7%

5. Results by Value-Outlook Cluster

5.0 Methodology: Identifying Outlook Groups

One of the survey's core questions asks respondents to distribute 100 probability points across five possible outcomes of high-level machine intelligence: extremely good, on balance good, more or less neutral, on balance bad, and extremely bad. These five numbers form a compact signature of each respondent's overall outlook on advanced AI.

We apply a Gaussian Mixture Model (GMM) with K=4 components to these five-dimensional signatures (standardized to zero mean and unit variance). GMM is a soft-clustering method that models the data as a mixture of multivariate normal distributions; each respondent is assigned to the component with the highest posterior probability. We use 20 random initializations to avoid local optima.

The four resulting clusters are then named by inspecting each cluster's mean probability profile:

This four-way split captures meaningful variation in worldview that cuts across traditional demographic variables like experience level. 1,538 of 1,793 cleaned analysis rows had complete value-outlook data and were assigned to a cluster. The remaining subsections re-examine key survey results through this lens.

The bimodal outlook is not merely an artifact of the mixture model. Among the 1,538 respondents who answered both endpoints of the value question, 474 (31%) assigned at least 10% probability to both an extremely good and an extremely bad outcome, and 74 (5%) assigned at least 25% to both. This model-free count confirms that a substantial minority genuinely hedges across both extremes rather than settling on a single direction.

5.0.1 Cluster Profiles

ClusterN% of totalExt. goodGoodNeutralBadExt. bad
Strong Optimists45730%33%34%22%11%0%
Mild Optimists46030%17%38%22%15%8%
Concerned32921%6%17%29%35%12%
Polarized/Bimodal29219%40%14%6%13%27%

5.1 HLMI Timeline Predictions

How soon do different outlook groups expect human-level machine intelligence?

ClusterNMedian HLMI yearIQR
Strong Optimists16020392032–2074
Mild Optimists15620392031–2054
Concerned9720442034–2074
Polarized/Bimodal10520342031–2054

5.2 Extinction/Disempowerment Estimates

Probability that future AI causes human extinction or permanent severe disempowerment, broken down by outlook cluster.

ClusterNMeanMedian≥10%≥25%
Strong Optimists22311.7%1%37%15%
Mild Optimists21014.2%8%50%18%
Concerned16621.3%10%63%35%
Polarized/Bimodal14531.1%20%69%47%

5.3 Concerning Scenarios

Mean concern level (0 = no concern, 3 = extreme concern) for each of 11 AI risk scenarios. The table below highlights which scenarios show the largest divergence across clusters.

ScenarioHighest clusterScoreLowest clusterScoreGap
Wrong goals: AI reduces human decision-making roleConcerned1.72Strong Optimists1.110.61
Automation leaves most people economically powerlessConcerned1.75Strong Optimists1.140.61
Authoritarian rulers use AI for controlConcerned2.22Strong Optimists1.640.58
Misaligned powerful AI causes catastrophe (weapons)Polarized/Bimodal1.65Strong Optimists1.090.56
AI worsens economic inequalityConcerned2.23Strong Optimists1.680.55
Automation makes people struggle to find meaningConcerned1.50Strong Optimists1.010.49
AI manipulates large-scale public opinionConcerned2.41Strong Optimists1.950.46
Less human interaction (more time with AI)Concerned1.77Strong Optimists1.350.42
AI bias worsens unjust situations (hiring, etc.)Concerned1.91Polarized/Bimodal1.540.37
AI lets dangerous groups make powerful tools (bioweapons)Concerned2.11Strong Optimists1.800.31
AI makes it easy to spread false info (deepfakes)Concerned2.41Polarized/Bimodal2.220.19

5.4 Expected AI Capabilities in 2044

Mean rated likelihood (0 = very unlikely, 4 = very likely) that AI systems will exhibit each capability by 2044.

5.5 Intelligence Explosion

Median probability estimates for dramatic AI capability speedup and superhuman AI emergence, by outlook cluster.

ClusterDramatic speedup within 2yr (median %)Dramatic speedup within 30yr (median %)Superhuman within 2yr (median %)N
Strong Optimists20%80%8%50
Mild Optimists20%80%10%55
Concerned15%78%10%37
Polarized/Bimodal32%95%10%28

5.6 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity

What rate of global AI progress over the next five years would make each cluster feel most optimistic for humanity's future?

5.7 Safety and Alignment-Problem Views

Mean scores on whether the alignment argument points at an important problem, how much society should prioritize AI safety research, and how valuable it is to work on the alignment problem today compared with other AI problems (all 0–4 scales).

ClusterAlignment-problem importance (0–4)Safety research priority (0–4)Value of working on alignment problem today vs other AI problems (0–4)N
Strong Optimists2.172.691.78228
Mild Optimists2.442.821.99220
Concerned2.463.252.12162
Polarized/Bimodal2.613.292.19145

5.8 Counterfactual Progress Reduction if Factors Were Halved

Median estimated progress reduction if each factor were halved, by cluster.

Cluster assignments are based on GMM (K=4, 20 random initializations) applied to the five HLMI value-outcome probabilities (vb_1_1–vb_1_5). Respondents missing all five values are excluded (1,538 of 1,793 cleaned analysis rows assigned). Sample sizes vary across subsections due to block randomization.


Appendix: Cluster Analysis

A data-driven exploration of the natural groupings, the safety divide, and hardware vs. software beliefs among 1,793 ESPAI 2024 cleaned analysis rows.

Executive Summary

1. We tested clustering solutions from K=2 through K=6 and found that four clusters yielded the most informative groupings (n=1,538). Beyond the expected optimist and pessimist groups, a Polarized/Bimodal group emerges that assigns high probability to both extremely good AND extremely bad outcomes. This group is not simply pessimistic or technophobic -- their P(extremely good) is comparable to the most optimistic cluster. They are better understood as "high-impact believers" who are convinced AI will be transformative but genuinely uncertain whether that transformation will be positive or catastrophic.
2. We observe a meaningful distinction between more and less safety/severe-risk-concerned researchers, though it falls along a spectrum rather than a clean divide. Among 754 respondents with safety data, those in the high-concern group assign 12% mean probability to extremely bad outcomes (vs. 7% for low-concern), and give higher extinction/disempowerment probability estimates (32% vs. 11% mean), with moderate-to-large effect sizes on key measures.
3. Those who give higher extinction/disempowerment probability estimates tend to predict EARLIER HLMI (rho=-0.22, p=0.0000, n=446), suggesting that concern is driven by the belief that powerful AI is coming soon, not that AI is inherently dangerous regardless of timeline.
4. Hardware vs. software beliefs exist on a spectrum, leaning slightly hardware. Among respondents who rated both (n=58), 35 lean hardware and 18 lean software. Computing hardware is rated as the most impactful factor overall (median 60% progress reduction if halved), ahead of algorithms (50%), data (50%), and funding (40%). Notably, people who rate hardware as more important tend to rate safety as LESS important (rho=-0.32, p<0.01), hinting that hardware-focused people may have a more gradualist worldview.
Important caveat: Block randomization means respondents answered different question subsets. The cleaned analysis file contains 1,793 rows; the audience-facing 2024 response count is 1,580. Cross-cutting analyses use smaller item-level samples and should be treated accordingly. Sample sizes are reported throughout.

1. How Many Camps? The Natural Clusters

1.1 Clustering on Value Outlook (n=1,538)

The value-of-HLMI question asked respondents to assign probabilities (summing to 100%) across five outcomes: extremely good, on balance good, neutral, on balance bad, and extremely bad. It is the highest-coverage feature in this analysis. We cluster on these 5 dimensions using Gaussian Mixture Models.

K selection

Figure 1: Silhouette scores for K=2 through K=6 clusters on value outlook. K=2 has the highest silhouette, but K=3 and K=4 offer more interpretable structure.

Silhouette scores are modest (0.10-0.12), indicating the clusters are not sharply separated -- this is a continuous landscape of opinion, not discrete tribes. Still, the structure is meaningful.

1.2 Four-Camp Solution

Value outlook clusters K=4

Figure 2: Four natural camps in value outlook. Left: PCA projection. Center: mean probability profiles. Right: cluster sizes.

The most striking finding here is the Polarized/Bimodal group. These respondents assign high probability to both extremely good and extremely bad outcomes -- their P(extremely good) is comparable to the Strong Optimists, yet they simultaneously assign substantial probability to catastrophic outcomes. This is not a group of doomers or technophobes. They are researchers who believe AI will be hugely impactful, but who are genuinely torn on whether that impact will be positive or negative. The key axis for this group is magnitude of impact, not direction -- they have rejected the possibility that AI will be a modest or neutral development.

1.3 Three-Camp Solution (Simpler View)

Collapsing to 3 clusters (which has a higher silhouette score of 0.18 vs 0.07) merges the finer distinctions into a simpler optimist/moderate/pessimist framing. This loses the polarized group but provides a cleaner summary:

Value outlook clusters K=3

Figure 3: Three-camp simplification. The polarized group is absorbed into the moderate/pessimist clusters.

1.4 Intuitive View: How Soon vs. How Good

The PCA axes above lack intuitive meaning. Below, we plot each respondent on two directly interpretable dimensions: their predicted HLMI arrival year (x-axis) and their net optimism score (y-axis), defined as P(good + extremely good) minus P(bad + extremely bad). This reveals where the four camps sit in the space of "how soon" vs. "how beneficial."

HLMI year vs net optimism scatter

Figure 3b: Respondents (n=935) binned by predicted HLMI year and net optimism. Each bubble's size shows how many researchers from that cluster fall in the bin (labeled when ≥5). The vertical dashed line marks the median predicted year; the horizontal line separates net optimists from net pessimists.

1.5 Combined Clustering: Values + Concerns + Safety + Extinction

For the 191 respondents who answered all of: value outlook, concern scenarios, alignment-problem importance, and extinction/disempowerment probability, we ran a richer clustering incorporating 8 features.

Combined clusters

Figure 4: Combined clustering incorporating worldview, concerns, safety/alignment-problem views, and extinction/disempowerment probability. Mean concern averages 11 scenario concern ratings on a 0-3 scale (0=no concern, 3=extreme concern); alignment-problem importance is a 0-4 ordinal score (0=not a real problem, 4=among the most important problems in the field); P(extinction/disempowerment) is a percentage.

ClusterSizeP(Good / Ext good)P(Bad / Ext bad)Mean concernAlignment-problem impP(extinction/disempowerment)HLMI Year median [IQR]; mean
Optimistic / Low x-risk 103 (54%) 24% / 35% 12% / 3% 1.6/3 2.3/4 4% 2049 [2034-2074]; mean 16981 (n=67)
Pessimistic / High x-risk 88 (46%) 25% / 18% 23% / 18% 1.9/3 2.5/4 34% 2043 [2032-2064]; mean 2056 (n=56)

HLMI year summaries include the median, interquartile range, mean, and item-level n because several groups share the same median year while their wider distributions differ.

2. Higher vs Lower Safety/Severe-Risk Concern

2.1 Defining the Groups

We built a safety composite score from (normalized 0-1 and averaged):

Requiring at least 2 of 3 components, we obtained scores for 754 respondents and split at the median (0.500) into High (n=330) and Low (n=424) safety concern groups.

2.2 The Full Comparison

Safety comparison

Figure 5: Comprehensive comparison of higher vs lower safety/severe-risk-concern respondents across six dimensions. Alignment-problem importance uses the 0-4 ordinal scale above; mean concern averages 11 scenario concern ratings on a 0-3 scale; the optimism-maximizing AI progress-rate answer is coded 0=much slower, 2=current speed, 4=much faster.

2.3 Statistical Tests

Figure 5 panelTested variablen (High)n (Low)Median (High)Median (Low)Mean (High)Mean (Low)p-valueEffect r
HLMI timeline HLMI Year 210258 2044.02044.0 478730.025329.1 0.0491 * 0.105
Extinction/disempowerment estimate P(extinction/disempowerment) 130234 30.05.0 31.710.8 0.0000 *** -0.466
Value outlook P(Extremely bad) 330424 5.05.0 12.47.3 0.0000 *** -0.182
Value outlook P(Extremely good) 330424 10.020.0 22.324.8 0.0568 n.s. 0.080
Concern levels Mean concern 163220 1.91.6 1.91.6 0.0000 *** -0.348
AI progress rate for optimism Optimism-maximizing AI progress rate 81103 2.02.0 1.92.2 0.0501 n.s. 0.164

Selected Mann-Whitney U tests for scalar summaries from Figure 5. The value-outlook panel is summarized by P(extremely good) and P(extremely bad), the concern panel by mean concern, and the AI-capabilities panel is tested item-by-item in the next table. Effect size r is rank-biserial correlation (|r| > 0.3 = medium, |r| > 0.5 = large). Scale notes: mean concern 0-3, optimism-maximizing AI progress-rate answer 0-4, probabilities in percentage points. *** p<0.001, ** p<0.01, * p<0.05

2.4 AI Capabilities by 2044: What Do Safety People Expect?

Safety-concerned people don't just worry more -- they have systematically different expectations for what AI will be able to do by 2044:

CapabilityHypothesized directionMean (High Safety)Mean (Low Safety)High - LowObserved directionp-valuen
Talk like expert Exploratory 3.413.34 +0.07 High > Low 0.7903 n.s. 199
Self-improve regardless Exploratory 2.052.00 +0.05 High > Low 0.7481 n.s. 199
AI-AI collaborations High > Low 2.372.11 +0.26 High > Low 0.1911 n.s. 199
Deceive humans High > Low 2.522.23 +0.29 High > Low 0.0820 n.s. 197
Unexpected strategies Exploratory 3.263.34 -0.07 High < Low 0.3043 n.s. 200
Seek power High > Low 1.741.35 +0.38 High > Low 0.0104 * 200
Explain actions (trustworthy) Exploratory 2.082.10 -0.02 High < Low 0.7394 n.s. 200
Can be jailbroken Exploratory 2.862.68 +0.18 High > Low 0.3029 n.s. 194
Surprising behavior Exploratory 2.762.91 -0.15 High < Low 0.3611 n.s. 200
Real-world actions Exploratory 2.372.19 +0.18 High > Low 0.2788 n.s. 199
Misaligned goals High > Low 2.461.94 +0.53 High > Low 0.0016 ** 199

Scale: 0=Very unlikely, 1=Unlikely, 2=Even chance, 3=Likely, 4=Very likely. High - Low is the high-safety mean minus the low-safety mean, so positive values mean high-safety respondents rated the capability as more likely. Hypothesized direction is marked for risk-relevant capabilities where we expected High > Low; other rows are exploratory.

Overall risk-relevant capability pattern: Averaging the 4 hypothesized High > Low items (AI-AI collaboration, deception, power-seeking, and misaligned goals), high-safety respondents rated these capabilities as more likely by +0.37 points on the 0-4 scale (means 2.27 vs. 1.90; medians 2.38 vs. 2.00; n=84 high, 115 low; Mann-Whitney p=0.0017 **, |r|=0.26). Respondents are included if they answered at least 2 of the 4 items.
Key pattern: Safety-concerned respondents do rate certain dangerous capabilities (power-seeking, AI-AI collaboration, misaligned goals, deception) as more likely, while agreeing with others on benign capabilities (expert conversation, explaining actions). However, the magnitude of these differences is notably small -- typically 0.3-0.6 points on a 0-4 scale. Both groups largely agree on what AI will be able to do by 2044. This suggests the safety divide is driven less by disagreements about AI's technical capabilities and more by differences in values, risk tolerance, or beliefs about how those capabilities will be managed.

2.5 Intelligence Explosion Beliefs

Intelligence explosion

Figure 6: Probability estimates for intelligence explosion scenarios, by safety concern level.

3. Hardware vs. Software Progress Beliefs

3.1 Counterfactual Progress Reduction if Factors Were Halved

Respondents rated how much AI progress would decrease if each of 5 factors were cut in half (0-100% scale). Higher values mean the factor is more important.

Progress attribution

Figure 7: Counterfactual progress-factor analysis. Top-left: estimated progress reduction distributions. Top-right: hardware vs algorithm progress-reduction estimates. Bottom-left: HW-SW index distribution. Bottom-right: factor correlations with other variables.

Rankings (by median):
  1. Computing hardware: 60% (n=119) -- largest median estimated progress reduction if halved
  2. Training data: 50% (n=109)
  3. Algorithm progress: 50% (n=103)
  4. Funding: 40% (n=116)
  5. Researcher effort: 25% (n=95) -- surprisingly lowest

3.2 Hardware vs Software Index

We computed (Hardware - Algorithms) / (Hardware + Algorithms) for each respondent who rated both. Positive = hardware-leaning, negative = software-leaning.

  • Hardware-leaning: 35 (60%)
  • Software-leaning: 18 (31%)
  • Balanced: 5
  • Median index: 0.15 (slight hardware lean)

3.3 Do Progress Beliefs Predict Other Views?

Sample size note: Only ~60-120 respondents answered the progress cause questions. Cross-tabulations with other variables yield modest samples (n=15-60). Treat these as suggestive.

Significant cause-factor correlations (p < 0.10)

Cause FactorTargetSpearman rhop-valuen
Researcher effortHLMI Year -0.238 0.0647 n.s. 61
Computing hardwareAlignment-problem imp -0.316 0.0086 ** 68
Training dataMean concern 0.337 0.0166 * 50

Hardware-Software Index Correlations

Variable 1Variable 2Spearman rhop-valuen
HW-vs-SW IndexP(extinction/disempowerment) 0.089 0.6212 n.s. 33
HW-vs-SW IndexHLMI Year 0.073 0.6758 n.s. 35
HW-vs-SW IndexP(Ext bad) -0.165 0.2152 n.s. 58
HW-vs-SW IndexAlignment-problem imp -0.369 0.0344 * 33
HW-vs-SW IndexMean concern 0.106 0.5847 n.s. 29
HW-vs-SW IndexP(Ext good) 0.048 0.7207 n.s. 58
The hardware-safety link: People who attribute more importance to computing hardware tend to rate safety as less important. This suggests a meaningful (if modest-sample) divide: "hardware people" may see AI progress as more incremental and predictable, while "algorithm people" may see it as more unpredictable and discontinuous -- leading to greater safety concern.

3.4 Progress Beliefs by Outlook and Safety Groups

Do optimists, pessimists, and safety-concerned researchers differ in what they think drives AI progress? Below we break down cause factor importance by outlook cluster (left) and safety group (right).

Cause factors by group

Figure 7b: Mean importance ratings for each progress factor, split by outlook group (left) and safety group (right). Scale: 0-100% estimated decrease in progress if factor halved.

Pessimists emphasize algorithms more. Pessimists rate algorithm progress as notably more important than optimists do, while both groups rate hardware similarly. This aligns with the hardware-safety correlation: those who see AI progress as driven by algorithmic breakthroughs (rather than predictable hardware scaling) tend to be more concerned about safety -- perhaps because algorithmic advances feel less predictable and harder to control. By contrast, the safety-group split shows little difference in cause factor ratings, suggesting the connection between progress beliefs and concern runs primarily through overall outlook rather than safety concern per se.

4. How Outlook Predicts Everything Else

The three value-outlook clusters (from Section 1) predict views across many other dimensions:

VariableConcernedMild OptimistsStrong OptimistsPolarized/Bimodal
HLMI Year2046 [2034-2074]; mean 533995 (n=188)2044 [2032-2064]; mean 19307 (n=290)2044 [2034-2064]; mean 1101046 (n=274)2044 [2033-2069]; mean 2123 (n=183)
P(extinction/disempowerment)10.0 (n=166)7.5 (n=210)1.0 (n=223)20.0 (n=145)
Alignment-problem importance3.0 (n=162)3.0 (n=220)2.0 (n=228)3.0 (n=145)
Mean concern2.0 (n=165)1.7 (n=243)1.5 (n=198)1.8 (n=148)
Optimism-maximizing AI progress rate1.0 (n=75)2.0 (n=111)2.0 (n=109)2.0 (n=77)

Values are medians with sample sizes in parentheses, except HLMI Year, which shows median [IQR], mean, and item-level n because several groups share the same median year. Scale notes: alignment-problem importance 0-4, mean concern 0-3, optimism-maximizing AI progress-rate answer 0=much slower to 4=much faster, probabilities in percentage points.

Outlook group comparisons

Figure 8: How the four value-outlook camps compare across timelines, extinction/disempowerment probability, safety/alignment-problem views, concern levels, optimism-maximizing AI progress-rate answers, and P(extremely bad). Mean concern averages 11 scenario concern ratings on a 0-3 scale; alignment-problem importance and optimism-maximizing AI progress-rate answers use 0-4 ordinal scales.

5. What Dimensions Structure AI Researcher Beliefs?

Principal Component Analysis on respondents with values + concerns + safety + extinction/disempowerment data (n=180) reveals the latent axes of disagreement.

PCA loadings

Figure 9: PCA factor loadings. Green bars = positive loading, red = negative. Each panel shows one principal component.

PC1 (21.4% of variance) is the "concern axis." It loads positively on all 11 concern items, alignment-problem importance, value of working on the alignment problem today, and P(extinction/disempowerment), and negatively on optimistic outlook. This single dimension captures most of the disagreement: how worried are you about AI?
PC2 (10.5%) is the "extremity axis." It loads positively on both P(Extremely good) AND P(Extremely bad)/P(extinction/disempowerment), and negatively on P(Neutral). This separates people who hold extreme views in either direction from those with moderate views. Some respondents genuinely believe AI could be transformatively good AND pose extinction/disempowerment risk.
PC3 (7.9%) separates "institutional safety" from "pessimism." It loads positively on P(On balance bad) and P(extinction/disempowerment), but negatively on alignment-problem importance and value of working on the alignment problem today. This captures people who think bad outcomes are likely but DON'T think the alignment problem is important -- perhaps because they think the problems aren't technical, or aren't solvable.

6. The Full Correlation Structure

Correlation matrix

Figure 10: Spearman rank correlations between key variables. Stars indicate significance. Sample sizes shown in each cell. Alignment-problem importance uses a 0-4 ordinal scale from not a real problem to among the field's most important problems; value of working on the alignment problem today uses a 0-4 ordinal scale from much less valuable to much more valuable than other AI problems; mean concern uses a 0-3 scale; HLMI Year is a calendar-year estimate; probabilities are 0-100 percentages.

Key Pairwise Correlations

RelationshipSpearman rhop-valuen
P(extinction/disempowerment) x HLMI Year -0.216 0.0000 *** 446
P(extinction/disempowerment) x P(Ext bad) 0.436 0.0000 *** 744
Alignment-problem imp x P(extinction/disempowerment) 0.141 0.0072 ** 364
Alignment-problem imp x HLMI Year -0.057 0.2211 n.s. 468
Alignment-problem imp x Mean concern 0.276 0.0000 *** 383
P(Ext good) x P(Ext bad) -0.113 0.0000 *** 1538
HLMI Year x P(Ext bad) -0.025 0.4421 n.s. 935
HLMI Year x P(Ext good) -0.074 0.0228 * 935
Mean concern x P(extinction/disempowerment) 0.321 0.0000 *** 381
Mean concern x HLMI Year -0.102 0.0293 * 461
Optimism-max progress rate x P(extinction/disempowerment) -0.158 0.0327 * 182
Optimism-max progress rate x HLMI Year -0.210 0.0010 *** 242
Optimism-max progress rate x alignment-problem imp -0.016 0.8289 n.s. 184
Strongest relationships:
  • Alignment-problem importance and value of working on the alignment problem today are highly correlated (these aren't independent beliefs -- they form a coherent safety/alignment-problem view)
  • P(extinction/disempowerment) and P(Extremely bad) are strongly linked -- people who give high extinction/disempowerment probability estimates also see HLMI as likely to be extremely bad
  • P(extinction/disempowerment) and HLMI timeline are negatively correlated -- people with the highest extinction/disempowerment estimates tend to predict earlier arrival of HLMI, not later
  • Optimism-maximizing AI progress rate and HLMI timeline are negatively correlated -- people with shorter timelines more often choose slower progress as the rate that would make them most optimistic for humanity's future

7. Methods & Caveats

7.1 Approach

7.2 Caveats

Block randomization: Not all respondents answered all questions. The "causes of progress" questions were only seen by ~6% of respondents. Safety questions by ~42%. This creates inherently unequal coverage for cross-cutting analyses.
Soft boundaries: The silhouette scores (0.08-0.12) indicate that clusters are not crisply separated. These are regions of higher density in a continuous opinion landscape, not discrete tribes. Median splits are even more arbitrary. The true distribution of views is continuous.
Causality: Correlations between safety concern, timeline predictions, and value outlook don't establish causal direction. A "safety worldview" (short timelines + high concern + safety-focused) may be a coherent package of beliefs adopted together, not a chain where one causes another.
No demographic controls: We don't control for region, career stage, subfield, or other demographics that might confound the clustering. Some of the "camps" might partially reflect institutional or geographic cultures rather than independent belief formation.

7.3 Technical Details


Appendix: Survey Flow

Survey-flow / randomization diagram modeled on Figure 16 of Grace et al., Thousands of AI Authors on the Future of AI. Each box is a question block. Percentages and n's are realized fill rates: the realized number of respondents who answered each block, divided by the total N. The total N counts everyone who answered at least one question (respondents who answered none are excluded).

Jobs / FAOL Sample-Size Reconciliation

The survey-flow boxes count anyone with any response in the broader Jobs / FAOL block. FAOL analyses use only the FAOL triplet, so their sample sizes can be slightly smaller.

Count definitionFixed-yearsFixed-probabilities
Whole Jobs / FAOL block (survey-flow box)234241
FAOL only, any triplet answer233234
FAOL only, complete triplet233223

Here, "fixed-years" means respondents gave probabilities for fixed time horizons (hj_b_full_*), while "fixed-probabilities" means respondents gave years for fixed probabilities (hj_a_full_*).

Tasks: ta_* columns are the fixed-probability framing, tb_* the fixed-years framing (verified empirically: ta_* rows have fixedprobabilities=1, tb_* rows have fixedprobabilities=0). The Qualtrics survey-flow definition (.qsf) is not in this repo, so the exact display order may differ from the diagram.


Appendix: Supplementary Figures

B.1 fixed-prob/fixed-year CDFs

Full CDF comparison between the fixed-prob and fixed-year conditions for HLMI and FAOL. The fixed-prob framing consistently produces earlier predictions (the red curve is shifted left relative to blue).

Corresponds to Figure 18 in Grace et al. (2024). Each curve is the mixture (mean) CDF for respondents assigned to that framing condition.

B.2 Bootstrap Confidence Bands

95% bootstrap confidence intervals on the aggregate CDF, obtained by resampling the fitted individual CDFs with replacement (500 resamples).

Milestone50% Year (point est.)95% Bootstrap CI
HLMI2042 (18.4 yr)2041–2044 (16.9–20.0 yr)
FAOL2096 (72.4 yr)2088–2105 (64.0–81.5 yr)

Bootstrap CIs reflect sampling variability — if we surveyed a different random sample of AI researchers, how much would the aggregate change? This is distinct from the spread of individual predictions (which is much wider).


Appendix: Statistical Tests

F.1 Yuen's Trimmed-Mean Bootstrap Test: Demographics

Tests whether experienced researchers (≥6 years in field) give significantly different HLMI predictions than junior researchers. Uses a 10% trimmed mean with 5,000 bootstrap resamples.

ComparisonTrimmed-mean diff95% CIp-valueSignificant (α=0.05)
HLMI 50% year (experienced − junior, split at 6 yr)+10.4 yr[-3.6, +43.4]1.000No

N (experienced): 58, N (junior): 55. Individual 50% years are computed from per-respondent gamma fits. The trimmed mean reduces sensitivity to extreme predictions.

F.2 Framing Effect: Statistical Significance

Tests whether year-framing and probability-framing respondents give significantly different predictions, using the same Yuen's bootstrap method.

MilestoneYear-framing medianProb-framing medianTrimmed-mean diff95% CIp-value
HLMI14.8 yr (n=525)28.6 yr (n=459)-17.5 yr[-28.5, -10.4]0.005
FAOL49.9 yr (n=223)123.2 yr (n=233)-477.7 yr[-641.6, -316.3]0.000

A negative difference (yr_50 − pr_50) means the year-framing produces earlier predictions. Medians shown for reference; the statistical test uses trimmed means.


Appendix: Aggregation Method Comparison

Different choices in how individual survey responses are fitted and combined into aggregate forecasts can shift the headline numbers. This appendix compares three approaches on the key milestones, motivated by Adamczewski's reanalysis of ESPAI data.

Methods compared

Mix-Mean (MSE) — the Grace et al. 2023 method. Fit a gamma CDF to each respondent's 3 data points using mean squared error, then take the pointwise mean of all fitted CDFs.

Mix-Median (MSE) — same fitting, but take the pointwise median instead of mean. More robust to outlier predictions (e.g. respondents predicting 100M years). Adamczewski argues the median better represents the "typical expert."

Mix-Mean (LogLoss) — fit using log-loss (cross-entropy) instead of MSE, then take the mean. Log-loss naturally weighs errors near p=0 and p=1 more heavily than errors near p=0.5, matching the intuition that a 5% error at p=0.95 represents a much larger shift in belief than the same error at p=0.50.

Results

MilestoneMix-Mean (MSE)Mix-Median (MSE)Mix-Mean (LogLoss)N (MSE)N (LogLoss)
HLMI2042 (18.4 yr)2044 (20.0 yr)2043 (19.0 yr)985985
FAOL2096 (72.4 yr)2101 (77.1 yr)2096 (72.0 yr)456456
Truck Driver2033 (8.5 yr)2034 (9.7 yr)2033 (8.5 yr)459459
Surgeon2050 (26.2 yr)2053 (29.3 yr)2051 (26.6 yr)458458
Retail Salesperson2031 (7.4 yr)2033 (8.7 yr)2031 (7.4 yr)460460
AI Researcher2052 (27.7 yr)2055 (30.6 yr)2052 (27.6 yr)458458

Key takeaways

Aggregation method matters more than loss function. Switching from mean to median aggregation pushes HLMI from 2042 to 2044 and FAOL from 2096 to 2101. Switching the loss function (MSE→log-loss) barely changes the aggregate — consistent with Adamczewski's finding that fitting choice has "hardly any impact."

The effect is largest for far-future milestones (FAOL, AI Researcher) where a few very pessimistic respondents pull the mean CDF rightward.

See also: compare_methods.py for the full 5-method comparison across all 39 tasks, including linear interpolation and gamma-individual methods. Method comparison motivated by Adamczewski (bayes.net/espai).


Data: 2024 Expert Survey on Progress in AI (ESPAI).
Previous survey comparison values from "Thousands of AI Authors on the Future of AI" (2024 preprint).