1. Overview & Survey Methodology

2,634
Report Responses
3,270
Cleaned Analysis Rows
2,634
Finished
636
Incomplete (partial data kept)
1,676
Year Framing
1,594
Probability Framing

Randomization Blocks

BlockCountPercentage
Block 11,12234.3%
Block 21,08533.2%
Block 31,06332.5%

Data Cleaning Notes

The following cleaning steps were applied (see cleaning_log.md for full details):

This dataset was pre-cleaned upstream. Column renaming was applied; no further numeric cleaning was needed.


3.1 When Will 39 AI Milestones Be Feasible?

33
Tasks feasible within 10 years (50% prob)
39
Total milestones assessed

Corresponds to Figure 1 in Grace et al. (2024). Dots show the 50% probability year from the mixture CDF; horizontal lines show the 10%–90% range. The x-axis is capped at 2100; milestones whose 50% or 90% year falls beyond the cap are labeled with an arrow showing the true year. Includes 39 tasks, 4 occupations, HLMI, and FAOL.

Combined year-framing and probability-framing responses using gamma mixture CDF aggregation (matching the 2023 report methodology). A gamma CDF is fitted to each respondent's data, then all CDFs are averaged pointwise to produce a mixture CDF from which percentiles are read. Sorted by 50% probability year (earliest first). The HLMI row uses the dedicated direct-HLMI blocks (hb_a_*/hb_b_*); it does not pool the final-occupation prediction block.

MilestoneN (fitted)10% Prob Year50% Prob Year90% Prob Year
Write Python code (e.g. quicksort)2570 yr (2023)2 yr (2025)14 yr (2037)
Write high-school history essay2550 yr (2023)2 yr (2025)13 yr (2036)
Play new Angry Birds levels (superhuman)2240 yr (2023)2 yr (2025)10 yr (2033)
Answer Googleable factoid questions2230 yr (2023)3 yr (2026)21 yr (2044)
Voice acting from text2560 yr (2023)3 yr (2026)17 yr (2040)
Win World Series of Poker2540 yr (2023)3 yr (2026)20 yr (2043)
Transcribe speech (noisy, accents)2660 yr (2023)3 yr (2026)17 yr (2040)
Fluent translation (most languages)2410 yr (2023)3 yr (2026)16 yr (2039)
Answer Googleable open-ended questions2520 yr (2023)3 yr (2026)23 yr (2046)
Group unseen objects into classes2400 yr (2023)4 yr (2027)23 yr (2046)
Produce song indistinguishable from artist2540 yr (2023)4 yr (2027)26 yr (2049)
Answer questions with no definite answer2600 yr (2023)4 yr (2027)35 yr (2058)
Beat best Starcraft 2 players2470 yr (2023)4 yr (2027)25 yr (2048)
Translate speech from subtitled films2400 yr (2023)5 yr (2028)41 yr (2064)
Build website with payment processing2630 yr (2023)5 yr (2028)31 yr (2054)
Phone banking services2600 yr (2023)5 yr (2028)30 yr (2053)
3D model from short video2330 yr (2023)5 yr (2028)27 yr (2050)
Atari novice level (20 min training)2420 yr (2023)5 yr (2028)36 yr (2059)
Fine-tune open source LLM2201 yr (2024)5 yr (2028)47 yr (2070)
One-shot image recognition2600 yr (2023)5 yr (2028)34 yr (2057)
Compose US Top 40 song (full audio)2610 yr (2023)6 yr (2029)46 yr (2069)
Outperform on all Atari games2790 yr (2023)6 yr (2029)35 yr (2058)
Learn efficient sorting (no solution form)2431 yr (2024)7 yr (2030)55 yr (2078)
Fold laundry (speed + quality)2361 yr (2024)7 yr (2030)40 yr (2063)
Play random game as human novice (<10 min)2261 yr (2024)7 yr (2030)38 yr (2061)
Translate newly discovered language (Rosetta stone)2471 yr (2024)7 yr (2030)52 yr (2075)
Write NYT best-seller novel2321 yr (2024)7 yr (2030)62 yr (2085)
Explain game AI moves to layman2471 yr (2024)8 yr (2031)81 yr (2104)
Assemble any LEGO set2521 yr (2024)8 yr (2031)48 yr (2071)
Win Putnam math competition2341 yr (2024)8 yr (2031)71 yr (2094)
Beat fastest human in 5km city race (biped robot)2811 yr (2024)9 yr (2032)44 yr (2067)
Beat best Go players (limited training)2351 yr (2024)10 yr (2033)100 yr (2123)
Find & patch security flaw (100k+ users)2442 yr (2025)10 yr (2033)87 yr (2110)
Retail Salesperson (occ.)7881 yr (2024)10 yr (2033)75 yr (2098)
Discover physics equations from simulation2711 yr (2024)12 yr (2035)123 yr (2146)
Truck Driver (occ.)7842 yr (2025)12 yr (2035)58 yr (2081)
Replicate ML conference study2532 yr (2025)12 yr (2035)109 yr (2132)
Install electrical wiring in new home2443 yr (2026)17 yr (2040)104 yr (2127)
Conduct ML research & write conference paper2692 yr (2025)20 yr (2043)246 yr (2269)
Prove publishable math theorems2573 yr (2026)23 yr (2046)270 yr (2293)
HLMI (all human tasks)17574 yr (2027)24 yr (2047)174 yr (2197)
Solve unsolved math problem (e.g. Millennium)2494 yr (2027)27 yr (2050)329 yr (2352)
Surgeon (occ.)7887 yr (2030)33 yr (2056)310 yr (2333)
AI Researcher (occ.)7866 yr (2029)40 yr (2063)1321 yr (3344)
Full Automation of Labor79214 yr (2037)89 yr (2112)2641 yr (4664)

Year values are years from survey date (2023), with calendar year in parentheses. "10% prob year" = year at which the mixture CDF reaches 10%. "90% prob year" = 90%. Dashes indicate insufficient data or an aggregate CDF that does not reach the target percentile within the 1e8-year "never/infinity" sentinel bound.


3.2 HLMI & Full Automation of Labor Timing

High-Level Machine Intelligence (HLMI)

HLMI is defined as machines that can accomplish every task better and cheaper than human workers.

4 yr (2027)
10% probability
24 yr (2047)
50% probability
174 yr (2197)
90% probability
PercentileThis analysisPublished 2023Diff
10% prob4 yr (2027)4 yr (2027)+0.3 yr
50% prob24 yr (2047)24 yr (2047)+0.3 yr

The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.

N (gamma fits): 1757 (0 failed fits)

This HLMI estimate uses only the dedicated direct-HLMI question blocks (hb_a_* and hb_b_*). It does not pool the separate final-occupation prediction block (hj_*_final_pred).

Full Automation of Labor (FAOL)

All occupations fully automatable — machines carry out every task better and more cheaply than humans.

14 yr (2037)
10% probability
89 yr (2112)
50% probability
2641 yr (4664)
90% probability
PercentileThis analysisPublished 2023Diff
50% prob89 yr (2112)93 yr (2116)-3.8 yr

N (gamma fits): 792 (0 failed fits)

Gap Between HLMI and FAOL

65 years
Gap at 50% probability

The consistent gap between HLMI and FAOL predictions reflects researchers' view that achieving human-level AI capability differs from fully automating all occupations.

Methodology Note

Gamma mixture CDF method: Following the methodology of Grace et al. (2024), a gamma CDF is fitted to each respondent's three data points (whether year-framing or probability-framing). All individual CDFs are then averaged pointwise to produce a mixture CDF, from which percentiles are read off. This approach handles extrapolation naturally (pessimistic respondents contribute heavy-tailed CDFs rather than being dropped) and treats both question framings symmetrically. Respondents were randomly assigned to either a "year framing" (provide years for 10%/50%/90% probability) or a "probability framing" (provide probabilities at fixed time horizons: 10, 20, and 40 years for HLMI; 10, 20, and 50 years for FAOL). See compare_methods.py for a comparison of this method with linear interpolation.

Corresponds to Figure 3 in Grace et al. (2024). Thin lines show individual respondent gamma CDFs (random subset of 200); thick line is the mixture (average) CDF. The dashed line marks 50% probability.


3.2.4 High-Level Machine Intelligence Timing by Experience

Insufficient demographic data: the years-in-field question (hh_howlong) was redacted or absent in this year's anonymized dataset, so the experience split cannot be computed.


3.2.5 Do Participants Agree on HLMI Timing?

"How much do you think your views on when HLMI will be achieved differ from those of the typical AI researcher?" (N=670)

ResponseCountPercentage
Not much29443.9%
A moderate amount30645.7%
A lot7010.4%

3.3 Framing Effects

Respondents were randomly assigned to one of two question framings for the same underlying question. The "year framing" asked how many years until a given probability, while the "probability framing" asked for the probability at a fixed time horizon.

HLMI 50% Probability Year by Framing

FramingN (gamma fits)Mixture CDF 50% YearCalendar Year
Year framing (respondent provides years)89017.5 yr2041
Probability framing (respondent provides probabilities)86734.0 yr2057
16.5 years
Framing effect (difference)

A consistent framing effect has been observed across survey waves: the year-framing tends to produce earlier (shorter-timeline) predictions than the probability-framing. This is a documented cognitive bias in probability elicitation.


3.4 Perceived Rates of Progress

Respondents were asked whether AI progress was faster in the first or second half of their career. (N=333)

59.5%
Said second half was faster
Median time in field (N=0)

Corresponds to Figure 4 in Grace et al. (2024).

ResponseCountPercentagePublished 2023
The second half19859.5%60%
The first half5717.1%
They were about the same7823.4%

The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.

How Far Along Is AI Progress? (Outside-view slider)

Respondents were shown a slider with three anchors: A = where progress was when they started working in their AI area; B = where it is now; and C = where it would need to be for AI software to have roughly human-level abilities at the tasks they study. They were asked "What fraction of the distance between where progress was when you started working in the area (A) and where it would need to be to attain human-level abilities in the area (C) have we come so far (B)?" on a 0–100 scale. This complements the year-based timing questions by anchoring an estimate to each respondent's own career window — so it is robust to disagreement about absolute calendar dates.

SurveyNMedianMeanSD
2023319 30 37.6 26.0

IQR (2023): 18–60.


3.5 What Causes AI Progress?

Respondents estimated how much less AI progress there would have been with half as much of each input. Higher values = more important to progress. Values are percentage less progress (0-100%).

Corresponds to Figure 5 in Grace et al. (2024). Red dots are means; box shows IQR with median line.

FactorNMedianMeanIQR
Computing hardware cost decline19550.0%56.6%40–75%
Training dataset effort19650.0%53.7%30–80%
Funding19150.0%46.0%25–70%
AI algorithm progress19050.0%46.1%25–70%
Researcher effort20030.0%35.4%20–50%

3.6 Will There Be an Intelligence Explosion?

Probability Estimates (numeric, 0-100%)

ScenarioNMedianMeanIQRPublished 2023 MedianDiff
Dramatic tech speedup within 2 years of HLMI29820%28.2%5–50%20%matches published
Dramatic tech speedup within 30 years of HLMI29780%66.7%50–95%80%matches published
Vastly superhuman AI within 2 years of HLMI28110%19.7%1–30%10%matches published
Vastly superhuman AI within 30 years of HLMI28260%56.3%20–90%

The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.

Is the Feedback Loop Argument Broadly Correct? (N=299)

"If AI does nearly all R&D, improvements in AI will accelerate progress including further AI progress. This could cause progress to become >10x faster within 5 years."

Corresponds to Figure 6 in Grace et al. (2024).

ResponseCountPercentage
Quite unlikely (0-20%)6923.1%
Unlikely (21-40%)7224.1%
About even chance (41-60%)7224.1%
Likely (61-80%)5919.7%
Quite likely (81-100%)279.0%

3.7 AI Capabilities in 2043

"In 2043, how likely do you think the following will be for at least some state-of-the-art AI systems?" Sorted by percentage rating Likely or Very Likely.

Corresponds to Figure 7 in Grace et al. (2024).

CapabilityNV.UnlikelyUnlikelyEvenLikelyV.LikelyLikely+V.Likely
Find unexpected ways to achieve goals6652%5%11%31%51%82%
Talk like an expert human on most topics6672%4%13%28%53%81%
Frequently behave surprisingly to humans6622%11%18%35%34%69%
Can be jailbroken for illegal commands6596%12%22%31%29%60%
Deceive humans to achieve goals (unintended)64910%22%23%27%18%45%
Have goals not aligned with human goals65313%23%23%25%16%40%
Cause important real-world actions (run business, etc.)65911%24%26%24%15%39%
Form AI-AI collaborative relationships (unintended)65816%24%22%25%13%38%
Can be trusted to explain their actions66510%24%33%23%10%33%
Self-improve regardless of human wishes66115%24%28%23%10%33%
Take actions to attain power65128%32%21%13%6%19%

3.8 Will AI Explain Its Decisions? (2028)

"For typical state-of-the-art AI systems in 2028, will users be able to know the true reasons for decisions?" (N=912)

19.2%
Rate it Likely or Very Likely

Corresponds to Figure 8 in Grace et al. (2024).

ResponseCountPercentage
Very unlikely (<10%)23726.0%
Unlikely (10-40%)32135.2%
Even odds (40-60%)17919.6%
Likely (60-90%)13414.7%
Very likely (>90%)414.5%

4.1 How Concerning Are Future AI Scenarios?

"How would you rate the level of concern these scenarios deserve over the next thirty years?" Sorted by percentage rating Substantial or Extreme concern.

Corresponds to Figure 9 in Grace et al. (2024).

ScenarioNNoneA LittleSubstantialExtremeSub.+Ext.
AI makes it easy to spread false info (deepfakes)13452%13%34%52%86%
AI manipulates large-scale public opinion13424%18%37%41%78%
AI lets dangerous groups make powerful tools (bioweapons)13404%23%41%32%73%
Authoritarian rulers use AI for control13207%20%35%38%73%
AI worsens economic inequality13396%23%38%33%71%
AI bias worsens unjust situations (hiring, etc.)13448%31%38%23%61%
Automation leaves most people economically powerless133517%37%32%15%46%
Less human interaction (more time with AI)134017%38%30%15%45%
Wrong goals: AI reduces human decision-making role133317%38%31%13%44%
Misaligned powerful AI causes catastrophe (weapons)134120%38%24%18%43%
Automation makes people struggle to find meaning133626%39%25%11%35%

4.2 How Good or Bad Will HLMI Be?

Respondents assigned probabilities to five outcome categories (summing to 100%). N=2704

68.3%
Net optimists
20.6%
Net pessimists
11.1%
Balanced

Corresponds to Figure 11 in Grace et al. (2024), "Thousands of AI Authors on the Future of AI."

Mean and Median Probabilities

OutcomeMeanMedian
Extremely good22.6%10.0%
On balance good29.1%25.0%
Neutral21.4%20.0%
On balance bad17.9%15.0%
Extremely bad9.0%5.0%

Individual Respondent Views

Each vertical slice below is one respondent's probability allocation, sorted from most optimistic (left) to most pessimistic (right). The chart reveals the full diversity of expert opinion.

Corresponds to Figure 10 in Grace et al. (2024).

The same respondents sorted by how much probability they assign to "extremely bad (e.g. human extinction)" outcomes. The growing black band on the right shows those assigning the highest probability to extremely bad outcomes.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

Individual Response Profiles

Each small bar below represents one respondent's five-category probability distribution, sampled evenly from the most optimistic to the most pessimistic.

No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.

Extreme Outcome Analysis

MetricThis analysisPublished 2023Diff
Non-zero to both extremes63.5%64.0%-0.5 pp
>=5% on extremely bad57.8%57.8%matches published
>=10% on extremely bad37.8%37.8%matches published
>=20% on extremely bad17.0%
>=25% on extremely bad9.8%
Mean extremely bad9.0%9.0%matches published
Median extremely bad5.0%5.0%matches published
Net optimists68.3%68.3%matches published

The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.


4.3 How Likely Is AI to Cause Extinction or Severe Disempowerment?

Each respondent was randomly assigned ONE of three questions about the probability that future AI advances cause human extinction or similarly permanent and severe disempowerment. The chart below also includes "extremely bad" HLMI outcome probabilities from Section 4.2 for comparison.

Corresponds to Figure 13 in Grace et al. (2024).

Note: The "Extremely bad HLMI outcome" question was asked to all respondents (N is larger), while each extinction/disempowerment question was asked to a random subset of respondents.

Corresponds to Figure 12 in Grace et al. (2024).

QuestionNMeanPublished 2023 MeanDiffMedianPublished 2023 MedianDiff>=10%>=25%
Future AI causes extinction or severe disempowerment132116.2%16.2%matches published5.0%5.0%matches published47.1%22.5%
Inability to control advanced AI causes extinction/disempowerment66119.4%19.4%matches published10.0%10.0%matches published51.4%26.5%
AI-caused extinction/disempowerment within 100 years65514.4%14.4%matches published5.0%5.0%matches published41.2%19.2%
Pooled (all variants combined)263716.5%5.0%46.7%22.7%

The "Pooled" row combines the raw responses from all variants above (1321 + 661 + 655 = 2637 responses) and takes the median of the combined data. Because each respondent was randomly assigned exactly one variant, no respondent is counted twice. And because the other variants each add a constraint (a specific cause, a time limit) to the basic question, every response is a lower bound on that respondent's unconstrained probability — so the pooled median is a conservative basis for "the median researcher put at least 5%..." claims.

The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.


4.4 Are Future AI-Risk Concerns Due to Misunderstandings of AI Research?

"To what extent do you think people's concerns about future risks from AI are due to misunderstandings of AI research?" (N=671)

No direct figure equivalent in Grace et al. (2024); this data is discussed in Section 4.4 of that paper.

ResponseCountPercentage
Hardly at all111.6%
Not much9814.6%
Somewhat19529.1%
To a large extent29544.0%
Almost entirely7210.7%

4.5 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity

Question wording: "What rate of global AI progress over the next five years would make you feel most optimistic for humanity's future? Assume any change in speed affects all projects equally." (N=646)

Corresponds to Table 3 in Grace et al. (2024); no figure equivalent in that paper.

ResponseCountPercentagePublished 2023Diff
Much slower314.8%4.8%matches published
Somewhat slower19329.9%29.9%matches published
Current speed17426.9%26.9%matches published
Somewhat faster14722.8%22.8%matches published
Much faster10115.6%15.6%matches published
Other00.0%

The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.


4.6 How Much Should AI Safety Research Be Prioritized?

"How much should society prioritize AI safety research, relative to how much it is currently prioritized?" (N=668)

70.5%
Say 'More' or 'Much more'
70%
Published 2023 (More+Much more)

Corresponds to Figure 14 in Grace et al. (2024).

ResponseCountPercentage
Much less162.4%
Less334.9%
About the same14822.2%
More25638.3%
Much more21532.2%

4.7 The Alignment Problem

Corresponds to Figure 15 in Grace et al. (2024).

Do you think this argument points at an important problem? (N=1322)

ResponseCountPercentage
No, not a real problem.675.1%
No, not an important problem.1219.2%
Yes, a moderately important problem.42532.1%
Yes, a very important problem.53940.8%
Yes, among the most important problems in the field.17012.9%

How valuable is it to work on this problem today, compared to other problems in AI? (N=1322)

ResponseCountPercentage
Much less valuable1118.4%
Less valuable29322.2%
As valuable as other problems51238.7%
More valuable29422.2%
Much more valuable1128.5%

How hard do you think this problem is compared to other problems in AI? (N=1318)

ResponseCountPercentage
Much easier342.6%
Easier13710.4%
As hard as other problems39630.0%
Harder47035.7%
Much harder28121.3%

5. Results by Value-Outlook Cluster

5.0 Methodology: Identifying Outlook Groups

One of the survey's core questions asks respondents to distribute 100 probability points across five possible outcomes of high-level machine intelligence: extremely good, on balance good, more or less neutral, on balance bad, and extremely bad. These five numbers form a compact signature of each respondent's overall outlook on advanced AI.

We apply a Gaussian Mixture Model (GMM) with K=4 components to these five-dimensional signatures (standardized to zero mean and unit variance). GMM is a soft-clustering method that models the data as a mixture of multivariate normal distributions; each respondent is assigned to the component with the highest posterior probability. We use 20 random initializations to avoid local optima.

The four resulting clusters are then named by inspecting each cluster's mean probability profile:

This four-way split captures meaningful variation in worldview that cuts across traditional demographic variables like experience level. 2,704 of 3,270 cleaned analysis rows had complete value-outlook data and were assigned to a cluster. The remaining subsections re-examine key survey results through this lens.

The bimodal outlook is not merely an artifact of the mixture model. Among the 2,704 respondents who answered both endpoints of the value question, 790 (29%) assigned at least 10% probability to both an extremely good and an extremely bad outcome, and 109 (4%) assigned at least 25% to both. This model-free count confirms that a substantial minority genuinely hedges across both extremes rather than settling on a single direction.

5.0.1 Cluster Profiles

ClusterN% of totalExt. goodGoodNeutralBadExt. bad
Strong Optimists43616%38%36%12%8%5%
Mild Optimists83231%29%36%23%12%0%
Alarmed49318%22%12%9%30%28%
Uncertain/Neutral94335%11%29%31%21%9%

5.1 HLMI Timeline Predictions

How soon do different outlook groups expect human-level machine intelligence?

ClusterNMedian HLMI yearIQR
Strong Optimists15420382033–2053
Mild Optimists26620432033–2063
Alarmed13520432033–2073
Uncertain/Neutral32020432033–2063

5.2 Extinction/Disempowerment Estimates

Probability that future AI causes human extinction or permanent severe disempowerment, broken down by outlook cluster.

ClusterNMeanMedian≥10%≥25%
Strong Optimists2069.6%5%33%13%
Mild Optimists3799.8%1%32%12%
Alarmed27131.1%20%72%48%
Uncertain/Neutral46515.6%10%51%20%

5.3 Concerning Scenarios

Mean concern level (0 = no concern, 3 = extreme concern) for each of 11 AI risk scenarios. The table below highlights which scenarios show the largest divergence across clusters.

ScenarioHighest clusterScoreLowest clusterScoreGap
Misaligned powerful AI causes catastrophe (weapons)Alarmed1.81Mild Optimists1.120.69
Wrong goals: AI reduces human decision-making roleAlarmed1.80Mild Optimists1.140.66
Automation leaves most people economically powerlessAlarmed1.83Mild Optimists1.220.61
AI lets dangerous groups make powerful tools (bioweapons)Alarmed2.25Mild Optimists1.790.46
Automation makes people struggle to find meaningAlarmed1.46Mild Optimists1.050.42
Less human interaction (more time with AI)Alarmed1.69Strong Optimists1.300.39
AI worsens economic inequalityAlarmed2.18Strong Optimists1.830.36
Authoritarian rulers use AI for controlAlarmed2.22Mild Optimists1.860.36
AI manipulates large-scale public opinionAlarmed2.30Mild Optimists2.000.29
AI makes it easy to spread false info (deepfakes)Uncertain/Neutral2.52Mild Optimists2.240.27
AI bias worsens unjust situations (hiring, etc.)Uncertain/Neutral1.90Mild Optimists1.640.26

5.4 Expected AI Capabilities in 2043

Mean rated likelihood (0 = very unlikely, 4 = very likely) that AI systems will exhibit each capability by 2043.

5.5 Intelligence Explosion

Median probability estimates for dramatic AI capability speedup and superhuman AI emergence, by outlook cluster.

ClusterDramatic speedup within 2yr (median %)Dramatic speedup within 30yr (median %)Superhuman within 2yr (median %)N
Strong Optimists20%90%15%47
Mild Optimists15%80%5%75
Alarmed20%90%10%54
Uncertain/Neutral10%70%10%118

5.6 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity

What rate of global AI progress over the next five years would make each cluster feel most optimistic for humanity's future?

5.7 Safety and Alignment-Problem Views

Mean scores on whether the alignment argument points at an important problem, how much society should prioritize AI safety research, and how valuable it is to work on the alignment problem today compared with other AI problems (all 0–4 scales).

ClusterAlignment-problem importance (0–4)Safety research priority (0–4)Value of working on alignment problem today vs other AI problems (0–4)N
Strong Optimists2.382.821.88224
Mild Optimists2.232.721.82386
Alarmed2.773.322.27259
Uncertain/Neutral2.553.022.06453

5.8 Counterfactual Progress Reduction if Factors Were Halved

Median estimated progress reduction if each factor were halved, by cluster.

Cluster assignments are based on GMM (K=4, 20 random initializations) applied to the five HLMI value-outcome probabilities (vb_1_1–vb_1_5). Respondents missing all five values are excluded (2,704 of 3,270 cleaned analysis rows assigned). Sample sizes vary across subsections due to block randomization.


Appendix: Cluster Analysis

A data-driven exploration of the natural groupings, the safety divide, and hardware vs. software beliefs among 3,270 ESPAI 2023 cleaned analysis rows.

Executive Summary

1. We tested clustering solutions from K=2 through K=6 and found that four clusters yielded the most informative groupings (n=1,538). Beyond the expected optimist and pessimist groups, a Polarized/Bimodal group emerges that assigns high probability to both extremely good AND extremely bad outcomes. This group is not simply pessimistic or technophobic -- their P(extremely good) is comparable to the most optimistic cluster. They are better understood as "high-impact believers" who are convinced AI will be transformative but genuinely uncertain whether that transformation will be positive or catastrophic.
2. We observe a meaningful distinction between more and less safety/severe-risk-concerned researchers, though it falls along a spectrum rather than a clean divide. Among 1322 respondents with safety data, those in the high-concern group assign 12% mean probability to extremely bad outcomes (vs. 7% for low-concern), and give higher extinction/disempowerment probability estimates (33% vs. 8% mean), with moderate-to-large effect sizes on key measures.
3. Those who give higher extinction/disempowerment probability estimates tend to predict EARLIER HLMI (rho=-0.14, p=0.0001, n=809), suggesting that concern is driven by the belief that powerful AI is coming soon, not that AI is inherently dangerous regardless of timeline.
4. Hardware vs. software beliefs exist on a spectrum, leaning slightly hardware. Among respondents who rated both (n=92), 59 lean hardware and 25 lean software. Computing hardware is rated as the most impactful factor overall (median 60% progress reduction if halved), ahead of algorithms (50%), data (50%), and funding (40%). Notably, people who rate hardware as more important tend to rate safety as LESS important (rho=-0.32, p<0.01), hinting that hardware-focused people may have a more gradualist worldview.
Important caveat: Block randomization means respondents answered different question subsets. The cleaned analysis file contains 3,270 rows. Cross-cutting analyses use smaller item-level samples and should be treated accordingly. Sample sizes are reported throughout.

1. How Many Camps? The Natural Clusters

1.1 Clustering on Value Outlook (n=1,538)

The value-of-HLMI question asked respondents to assign probabilities (summing to 100%) across five outcomes: extremely good, on balance good, neutral, on balance bad, and extremely bad. It is the highest-coverage feature in this analysis. We cluster on these 5 dimensions using Gaussian Mixture Models.

K selection

Figure 1: Silhouette scores for K=2 through K=6 clusters on value outlook. K=2 has the highest silhouette, but K=3 and K=4 offer more interpretable structure.

Silhouette scores are modest (0.10-0.12), indicating the clusters are not sharply separated -- this is a continuous landscape of opinion, not discrete tribes. Still, the structure is meaningful.

1.2 Four-Camp Solution

Value outlook clusters K=4

Figure 2: Four natural camps in value outlook. Left: PCA projection. Center: mean probability profiles. Right: cluster sizes.

The most striking finding here is the Polarized/Bimodal group. These respondents assign high probability to both extremely good and extremely bad outcomes -- their P(extremely good) is comparable to the Strong Optimists, yet they simultaneously assign substantial probability to catastrophic outcomes. This is not a group of doomers or technophobes. They are researchers who believe AI will be hugely impactful, but who are genuinely torn on whether that impact will be positive or negative. The key axis for this group is magnitude of impact, not direction -- they have rejected the possibility that AI will be a modest or neutral development.

1.3 Three-Camp Solution (Simpler View)

Collapsing to 3 clusters (which has a higher silhouette score of 0.18 vs 0.07) merges the finer distinctions into a simpler optimist/moderate/pessimist framing. This loses the polarized group but provides a cleaner summary:

Value outlook clusters K=3

Figure 3: Three-camp simplification. The polarized group is absorbed into the moderate/pessimist clusters.

1.4 Intuitive View: How Soon vs. How Good

The PCA axes above lack intuitive meaning. Below, we plot each respondent on two directly interpretable dimensions: their predicted HLMI arrival year (x-axis) and their net optimism score (y-axis), defined as P(good + extremely good) minus P(bad + extremely bad). This reveals where the four camps sit in the space of "how soon" vs. "how beneficial."

HLMI year vs net optimism scatter

Figure 3b: Respondents (n=1663) binned by predicted HLMI year and net optimism. Each bubble's size shows how many researchers from that cluster fall in the bin (labeled when ≥5). The vertical dashed line marks the median predicted year; the horizontal line separates net optimists from net pessimists.

1.5 Combined Clustering: Values + Concerns + Safety + Extinction

For the 323 respondents who answered all of: value outlook, concern scenarios, alignment-problem importance, and extinction/disempowerment probability, we ran a richer clustering incorporating 8 features.

Combined clusters

Figure 4: Combined clustering incorporating worldview, concerns, safety/alignment-problem views, and extinction/disempowerment probability. Mean concern averages 11 scenario concern ratings on a 0-3 scale (0=no concern, 3=extreme concern); alignment-problem importance is a 0-4 ordinal score (0=not a real problem, 4=among the most important problems in the field); P(extinction/disempowerment) is a percentage.

ClusterSizeP(Good / Ext good)P(Bad / Ext bad)Mean concernAlignment-problem impP(extinction/disempowerment)HLMI Year median [IQR]; mean
Optimistic / High x-risk 72 (22%) 13% / 46% 11% / 4% 1.8/3 2.4/4 21% 2053 [2043-2073]; mean 2080 (n=53)
Optimistic / Low x-risk 60 (19%) 36% / 32% 13% / 8% 1.8/3 2.5/4 5% 2043 [2033-2073]; mean 2068 (n=43)
Optimistic / Low x-risk 96 (30%) 26% / 37% 14% / 4% 1.6/3 2.2/4 4% 2043 [2033-2072]; mean 2112 (n=58)
Pessimistic / Safety-focused / High x-risk / Broadly worried 42 (13%) 22% / 13% 19% / 36% 2.1/3 3.1/4 56% 2046 [2038-2063]; mean 2060 (n=21)
Pessimistic 38 (12%) 6% / 13% 50% / 11% 2.0/3 2.0/4 13% 2063 [2047-2085]; mean 45562 (n=23)
Moderate / Low x-risk 15 (5%) 2% / 15% 15% / 2% 1.6/3 2.4/4 4% 2063 [2040-2296]; mean 2186 (n=8)

HLMI year summaries include the median, interquartile range, mean, and item-level n because several groups share the same median year while their wider distributions differ.

2. Higher vs Lower Safety/Severe-Risk Concern

2.1 Defining the Groups

We built a safety composite score from (normalized 0-1 and averaged):

Requiring at least 2 of 3 components, we obtained scores for 1322 respondents and split at the median (0.500) into High (n=595) and Low (n=727) safety concern groups.

2.2 The Full Comparison

Safety comparison

Figure 5: Comprehensive comparison of higher vs lower safety/severe-risk-concern respondents across six dimensions. Alignment-problem importance uses the 0-4 ordinal scale above; mean concern averages 11 scenario concern ratings on a 0-3 scale; the optimism-maximizing AI progress-rate answer is coded 0=much slower, 2=current speed, 4=much faster.

2.3 Statistical Tests

Figure 5 panelTested variablen (High)n (Low)Median (High)Median (Low)Mean (High)Mean (Low)p-valueEffect r
HLMI timeline HLMI Year 357453 2043.02053.0 2229.04529.2 0.0074 ** 0.109
Extinction/disempowerment estimate P(extinction/disempowerment) 237406 25.02.0 32.57.9 0.0000 *** -0.586
Value outlook P(Extremely bad) 595727 7.05.0 12.46.7 0.0000 *** -0.223
Value outlook P(Extremely good) 595727 10.010.0 20.423.3 0.0724 n.s. 0.057
Concern levels Mean concern 274381 1.91.7 1.91.7 0.0000 *** -0.296
AI progress rate for optimism Optimism-maximizing AI progress rate 144204 2.02.0 1.92.2 0.0128 * 0.152

Selected Mann-Whitney U tests for scalar summaries from Figure 5. The value-outlook panel is summarized by P(extremely good) and P(extremely bad), the concern panel by mean concern, and the AI-capabilities panel is tested item-by-item in the next table. Effect size r is rank-biserial correlation (|r| > 0.3 = medium, |r| > 0.5 = large). Scale notes: mean concern 0-3, optimism-maximizing AI progress-rate answer 0-4, probabilities in percentage points. *** p<0.001, ** p<0.01, * p<0.05

2.4 AI Capabilities by 2043: What Do Safety People Expect?

Safety-concerned people don't just worry more -- they have systematically different expectations for what AI will be able to do by 2043:

CapabilityHypothesized directionMean (High Safety)Mean (Low Safety)High - LowObserved directionp-valuen
Talk like expert Exploratory 3.343.11 +0.23 High > Low 0.1252 n.s. 322
Self-improve regardless Exploratory 2.081.74 +0.34 High > Low 0.0192 * 317
AI-AI collaborations High > Low 2.131.86 +0.27 High > Low 0.0719 n.s. 313
Deceive humans High > Low 2.421.98 +0.44 High > Low 0.0021 ** 313
Unexpected strategies Exploratory 3.433.01 +0.43 High > Low 0.0000 *** 322
Seek power High > Low 1.561.25 +0.31 High > Low 0.0219 * 314
Explain actions (trustworthy) Exploratory 1.902.02 -0.11 High < Low 0.3351 n.s. 322
Can be jailbroken Exploratory 2.842.49 +0.36 High > Low 0.0097 ** 318
Surprising behavior Exploratory 3.102.69 +0.41 High > Low 0.0002 *** 321
Real-world actions Exploratory 2.431.85 +0.57 High > Low 0.0001 *** 317
Misaligned goals High > Low 2.341.87 +0.47 High > Low 0.0011 ** 316

Scale: 0=Very unlikely, 1=Unlikely, 2=Even chance, 3=Likely, 4=Very likely. High - Low is the high-safety mean minus the low-safety mean, so positive values mean high-safety respondents rated the capability as more likely. Hypothesized direction is marked for risk-relevant capabilities where we expected High > Low; other rows are exploratory.

Overall risk-relevant capability pattern: Averaging the 4 hypothesized High > Low items (AI-AI collaboration, deception, power-seeking, and misaligned goals), high-safety respondents rated these capabilities as more likely by +0.38 points on the 0-4 scale (means 2.12 vs. 1.74; medians 2.00 vs. 1.50; n=157 high, 162 low; Mann-Whitney p=0.0008 ***, |r|=0.22). Respondents are included if they answered at least 2 of the 4 items.
Key pattern: Safety-concerned respondents do rate certain dangerous capabilities (power-seeking, AI-AI collaboration, misaligned goals, deception) as more likely, while agreeing with others on benign capabilities (expert conversation, explaining actions). However, the magnitude of these differences is notably small -- typically 0.3-0.6 points on a 0-4 scale. Both groups largely agree on what AI will be able to do by 2043. This suggests the safety divide is driven less by disagreements about AI's technical capabilities and more by differences in values, risk tolerance, or beliefs about how those capabilities will be managed.

2.5 Intelligence Explosion Beliefs

Intelligence explosion

Figure 6: Probability estimates for intelligence explosion scenarios, by safety concern level.

3. Hardware vs. Software Progress Beliefs

3.1 Counterfactual Progress Reduction if Factors Were Halved

Respondents rated how much AI progress would decrease if each of 5 factors were cut in half (0-100% scale). Higher values mean the factor is more important.

Progress attribution

Figure 7: Counterfactual progress-factor analysis. Top-left: estimated progress reduction distributions. Top-right: hardware vs algorithm progress-reduction estimates. Bottom-left: HW-SW index distribution. Bottom-right: factor correlations with other variables.

Rankings (by median):
  1. Computing hardware: 50% (n=195) -- largest median estimated progress reduction if halved
  2. Training data: 50% (n=196)
  3. Algorithm progress: 50% (n=190)
  4. Funding: 50% (n=191)
  5. Researcher effort: 30% (n=200) -- surprisingly lowest

3.2 Hardware vs Software Index

We computed (Hardware - Algorithms) / (Hardware + Algorithms) for each respondent who rated both. Positive = hardware-leaning, negative = software-leaning.

  • Hardware-leaning: 59 (64%)
  • Software-leaning: 25 (27%)
  • Balanced: 8
  • Median index: 0.19 (slight hardware lean)

3.3 Do Progress Beliefs Predict Other Views?

Sample size note: Only ~60-120 respondents answered the progress cause questions. Cross-tabulations with other variables yield modest samples (n=15-60). Treat these as suggestive.

Significant cause-factor correlations (p < 0.10)

Cause FactorTargetSpearman rhop-valuen
Researcher effortAlignment-problem imp 0.296 0.0044 ** 91
Training dataHLMI Year -0.155 0.0848 n.s. 124
FundingAlignment-problem imp 0.223 0.0275 * 98

Hardware-Software Index Correlations

Variable 1Variable 2Spearman rhop-valuen
HW-vs-SW IndexP(extinction/disempowerment) 0.151 0.3601 n.s. 39
HW-vs-SW IndexHLMI Year 0.276 0.0430 * 54
HW-vs-SW IndexP(Ext bad) 0.176 0.0939 n.s. 92
HW-vs-SW IndexAlignment-problem imp -0.128 0.3871 n.s. 48
HW-vs-SW IndexMean concern 0.189 0.2368 n.s. 41
HW-vs-SW IndexP(Ext good) -0.154 0.1423 n.s. 92
The hardware-safety link: People who attribute more importance to computing hardware tend to rate safety as less important. This suggests a meaningful (if modest-sample) divide: "hardware people" may see AI progress as more incremental and predictable, while "algorithm people" may see it as more unpredictable and discontinuous -- leading to greater safety concern.

3.4 Progress Beliefs by Outlook and Safety Groups

Do optimists, pessimists, and safety-concerned researchers differ in what they think drives AI progress? Below we break down cause factor importance by outlook cluster (left) and safety group (right).

Cause factors by group

Figure 7b: Mean importance ratings for each progress factor, split by outlook group (left) and safety group (right). Scale: 0-100% estimated decrease in progress if factor halved.

Pessimists emphasize algorithms more. Pessimists rate algorithm progress as notably more important than optimists do, while both groups rate hardware similarly. This aligns with the hardware-safety correlation: those who see AI progress as driven by algorithmic breakthroughs (rather than predictable hardware scaling) tend to be more concerned about safety -- perhaps because algorithmic advances feel less predictable and harder to control. By contrast, the safety-group split shows little difference in cause factor ratings, suggesting the connection between progress beliefs and concern runs primarily through overall outlook rather than safety concern per se.

4. How Outlook Predicts Everything Else

The three value-outlook clusters (from Section 1) predict views across many other dimensions:

VariableMild OptimistsUncertain/NeutralStrong OptimistsAlarmed
HLMI Year2043 [2033-2073]; mean 2393 (n=515)2048 [2033-2073]; mean 2183 (n=595)2043 [2033-2063]; mean 2080 (n=280)2053 [2038-2073]; mean 6157 (n=273)
P(extinction/disempowerment)1.0 (n=379)10.0 (n=465)5.0 (n=206)20.0 (n=271)
Alignment-problem importance2.0 (n=386)3.0 (n=453)2.0 (n=224)3.0 (n=259)
Mean concern1.6 (n=435)1.8 (n=444)1.6 (n=237)2.0 (n=229)
Optimism-maximizing AI progress rate2.0 (n=188)2.0 (n=234)3.0 (n=93)1.0 (n=131)

Values are medians with sample sizes in parentheses, except HLMI Year, which shows median [IQR], mean, and item-level n because several groups share the same median year. Scale notes: alignment-problem importance 0-4, mean concern 0-3, optimism-maximizing AI progress-rate answer 0=much slower to 4=much faster, probabilities in percentage points.

Outlook group comparisons

Figure 8: How the four value-outlook camps compare across timelines, extinction/disempowerment probability, safety/alignment-problem views, concern levels, optimism-maximizing AI progress-rate answers, and P(extremely bad). Mean concern averages 11 scenario concern ratings on a 0-3 scale; alignment-problem importance and optimism-maximizing AI progress-rate answers use 0-4 ordinal scales.

5. What Dimensions Structure AI Researcher Beliefs?

Principal Component Analysis on respondents with values + concerns + safety + extinction/disempowerment data (n=301) reveals the latent axes of disagreement.

PCA loadings

Figure 9: PCA factor loadings. Green bars = positive loading, red = negative. Each panel shows one principal component.

PC1 (22.9% of variance) is the "concern axis." It loads positively on all 11 concern items, alignment-problem importance, value of working on the alignment problem today, and P(extinction/disempowerment), and negatively on optimistic outlook. This single dimension captures most of the disagreement: how worried are you about AI?
PC2 (11.9%) is the "extremity axis." It loads positively on both P(Extremely good) AND P(Extremely bad)/P(extinction/disempowerment), and negatively on P(Neutral). This separates people who hold extreme views in either direction from those with moderate views. Some respondents genuinely believe AI could be transformatively good AND pose extinction/disempowerment risk.
PC3 (8.6%) separates "institutional safety" from "pessimism." It loads positively on P(On balance bad) and P(extinction/disempowerment), but negatively on alignment-problem importance and value of working on the alignment problem today. This captures people who think bad outcomes are likely but DON'T think the alignment problem is important -- perhaps because they think the problems aren't technical, or aren't solvable.

6. The Full Correlation Structure

Correlation matrix

Figure 10: Spearman rank correlations between key variables. Stars indicate significance. Sample sizes shown in each cell. Alignment-problem importance uses a 0-4 ordinal scale from not a real problem to among the field's most important problems; value of working on the alignment problem today uses a 0-4 ordinal scale from much less valuable to much more valuable than other AI problems; mean concern uses a 0-3 scale; HLMI Year is a calendar-year estimate; probabilities are 0-100 percentages.

Key Pairwise Correlations

RelationshipSpearman rhop-valuen
P(extinction/disempowerment) x HLMI Year -0.137 0.0001 *** 809
P(extinction/disempowerment) x P(Ext bad) 0.469 0.0000 *** 1321
Alignment-problem imp x P(extinction/disempowerment) 0.312 0.0000 *** 643
Alignment-problem imp x HLMI Year -0.090 0.0102 * 810
Alignment-problem imp x Mean concern 0.301 0.0000 *** 655
P(Ext good) x P(Ext bad) -0.010 0.6123 n.s. 2704
HLMI Year x P(Ext bad) -0.023 0.3468 n.s. 1663
HLMI Year x P(Ext good) -0.072 0.0034 ** 1663
Mean concern x P(extinction/disempowerment) 0.378 0.0000 *** 652
Mean concern x HLMI Year -0.109 0.0015 ** 843
Optimism-max progress rate x P(extinction/disempowerment) -0.189 0.0004 *** 354
Optimism-max progress rate x HLMI Year 0.008 0.8744 n.s. 383
Optimism-max progress rate x alignment-problem imp -0.168 0.0016 ** 348
Strongest relationships:
  • Alignment-problem importance and value of working on the alignment problem today are highly correlated (these aren't independent beliefs -- they form a coherent safety/alignment-problem view)
  • P(extinction/disempowerment) and P(Extremely bad) are strongly linked -- people who give high extinction/disempowerment probability estimates also see HLMI as likely to be extremely bad
  • P(extinction/disempowerment) and HLMI timeline are negatively correlated -- people with the highest extinction/disempowerment estimates tend to predict earlier arrival of HLMI, not later
  • Optimism-maximizing AI progress rate and HLMI timeline are negatively correlated -- people with shorter timelines more often choose slower progress as the rate that would make them most optimistic for humanity's future

7. Methods & Caveats

7.1 Approach

7.2 Caveats

Block randomization: Not all respondents answered all questions. The "causes of progress" questions were only seen by ~6% of respondents. Safety questions by ~42%. This creates inherently unequal coverage for cross-cutting analyses.
Soft boundaries: The silhouette scores (0.08-0.12) indicate that clusters are not crisply separated. These are regions of higher density in a continuous opinion landscape, not discrete tribes. Median splits are even more arbitrary. The true distribution of views is continuous.
Causality: Correlations between safety concern, timeline predictions, and value outlook don't establish causal direction. A "safety worldview" (short timelines + high concern + safety-focused) may be a coherent package of beliefs adopted together, not a chain where one causes another.
No demographic controls: We don't control for region, career stage, subfield, or other demographics that might confound the clustering. Some of the "camps" might partially reflect institutional or geographic cultures rather than independent belief formation.

7.3 Technical Details


Appendix: Survey Flow

Survey-flow / randomization diagram modeled on Figure 16 of Grace et al., Thousands of AI Authors on the Future of AI. Each box is a question block. Percentages and n's are realized fill rates: the realized number of respondents who answered each block, divided by the total N. The total N counts everyone who answered at least one question (respondents who answered none are excluded).

Jobs / FAOL Sample-Size Reconciliation

The survey-flow boxes count anyone with any response in the broader Jobs / FAOL block. FAOL analyses use only the FAOL triplet, so their sample sizes can be slightly smaller.

Count definitionFixed-yearsFixed-probabilities
Whole Jobs / FAOL block (survey-flow box)384426
FAOL only, any triplet answer380418
FAOL only, complete triplet379413

Here, "fixed-years" means respondents gave probabilities for fixed time horizons (hj_b_full_*), while "fixed-probabilities" means respondents gave years for fixed probabilities (hj_a_full_*).

Tasks: ta_* columns are the fixed-probability framing, tb_* the fixed-years framing (verified empirically: ta_* rows have fixedprobabilities=1, tb_* rows have fixedprobabilities=0). The Qualtrics survey-flow definition (.qsf) is not in this repo, so the exact display order may differ from the diagram.


Appendix: Supplementary Figures

B.1 fixed-prob/fixed-year CDFs

Full CDF comparison between the fixed-prob and fixed-year conditions for HLMI and FAOL. The fixed-prob framing consistently produces earlier predictions (the red curve is shifted left relative to blue).

Corresponds to Figure 18 in Grace et al. (2024). Each curve is the mixture (mean) CDF for respondents assigned to that framing condition.

B.2 Bootstrap Confidence Bands

95% bootstrap confidence intervals on the aggregate CDF, obtained by resampling the fitted individual CDFs with replacement (500 resamples).

Milestone50% Year (point est.)95% Bootstrap CI
HLMI2047 (24.3 yr)2046–2049 (23.0–25.6 yr)
FAOL2112 (89.2 yr)2105–2121 (81.8–98.4 yr)

Bootstrap CIs reflect sampling variability — if we surveyed a different random sample of AI researchers, how much would the aggregate change? This is distinct from the spread of individual predictions (which is much wider).


Appendix: Statistical Tests

F.1 Yuen's Trimmed-Mean Bootstrap Test: Demographics

Tests whether experienced researchers (split at the median years in field) give significantly different HLMI predictions than junior researchers. Uses a 10% trimmed mean with 5,000 bootstrap resamples.

Insufficient data for demographic comparison (years-in-field was redacted in the anonymized dataset).

N (experienced): 0, N (junior): 0. Individual 50% years are computed from per-respondent gamma fits. The trimmed mean reduces sensitivity to extreme predictions.

F.2 Framing Effect: Statistical Significance

Tests whether year-framing and probability-framing respondents give significantly different predictions, using the same Yuen's bootstrap method.

MilestoneYear-framing medianProb-framing medianTrimmed-mean diff95% CIp-value
HLMI19.8 yr (n=890)36.4 yr (n=867)-26.7 yr[-34.7, -20.6]0.000
FAOL74.2 yr (n=413)130.5 yr (n=379)-672.3 yr[-815.8, -533.0]0.000

A negative difference (yr_50 − pr_50) means the year-framing produces earlier predictions. Medians shown for reference; the statistical test uses trimmed means.


Appendix: Aggregation Method Comparison

Different choices in how individual survey responses are fitted and combined into aggregate forecasts can shift the headline numbers. This appendix compares three approaches on the key milestones, motivated by Adamczewski's reanalysis of ESPAI data.

Methods compared

Mix-Mean (MSE) — the Grace et al. 2023 method. Fit a gamma CDF to each respondent's 3 data points using mean squared error, then take the pointwise mean of all fitted CDFs.

Mix-Median (MSE) — same fitting, but take the pointwise median instead of mean. More robust to outlier predictions (e.g. respondents predicting 100M years). Adamczewski argues the median better represents the "typical expert."

Mix-Mean (LogLoss) — fit using log-loss (cross-entropy) instead of MSE, then take the mean. Log-loss naturally weighs errors near p=0 and p=1 more heavily than errors near p=0.5, matching the intuition that a 5% error at p=0.95 represents a much larger shift in belief than the same error at p=0.50.

Results

MilestoneMix-Mean (MSE)Mix-Median (MSE)Mix-Mean (LogLoss)N (MSE)N (LogLoss)
HLMI2047 (24.3 yr)2049 (25.5 yr)2048 (24.7 yr)17571757
FAOL2112 (89.2 yr)2121 (97.5 yr)2112 (89.0 yr)792792
Truck Driver2035 (11.6 yr)2034 (10.6 yr)2035 (11.6 yr)784784
Surgeon2056 (32.7 yr)2057 (34.1 yr)2056 (33.5 yr)788788
Retail Salesperson2033 (9.8 yr)2033 (10.0 yr)2033 (9.8 yr)788788
AI Researcher2063 (40.0 yr)2071 (48.2 yr)2063 (40.2 yr)786786

Key takeaways

Aggregation method matters more than loss function. Switching from mean to median aggregation pushes HLMI from 2047 to 2049 and FAOL from 2112 to 2121. Switching the loss function (MSE→log-loss) barely changes the aggregate — consistent with Adamczewski's finding that fitting choice has "hardly any impact."

The effect is largest for far-future milestones (FAOL, AI Researcher) where a few very pessimistic respondents pull the mean CDF rightward.

See also: compare_methods.py for the full 5-method comparison across all 39 tasks, including linear interpolation and gamma-individual methods. Method comparison motivated by Adamczewski (bayes.net/espai).


Data: 2023 Expert Survey on Progress in AI (ESPAI).
Previous survey comparison values from "Thousands of AI Authors on the Future of AI" (2024 preprint).