2024 ESPAI Survey Analysis
Expert Survey on Progress in AI — Comprehensive Data Report
Generated 2026-07-16 19:37
1. Overview & Survey Methodology
Randomization Blocks
| Block | Count | Percentage |
|---|---|---|
| Block 1 | 596 | 33.2% |
| Block 2 | 591 | 33.0% |
| Block 3 | 606 | 33.8% |
Data Cleaning Notes
The following cleaning steps were applied (see cleaning_log.md for full details):
- Literal year values (2000–3000) converted to years-from-survey. 183 values converted.
- Extreme year values (>1e8) capped at 1e8 (1 values).
- Monotonicity violations found in 320 complete triplets.
- 331 descending triplets found across 111 respondents in year and probability framings (values set to NaN).
Caveat — sum-to-100 triplets: 393 triplets across 159 respondents have values summing to exactly 100 (e.g., 10%, 30%, 60% at the three time horizons). This may indicate respondents who treated the probabilities as shares that must total 100%, rather than independent cumulative probabilities at each time horizon. Following standard methodology, sum-to-100 triplets from respondents with descending errors were also removed (368 total triplets NaN'd).
3.1 When Will 39 AI Milestones Be Feasible?
Corresponds to Figure 1 in Grace et al. (2024). Dots show the 50% probability year from the mixture CDF; horizontal lines show the 10%–90% range. The x-axis is capped at 2100; milestones whose 50% or 90% year falls beyond the cap are labeled with an arrow showing the true year. Includes 39 tasks, 4 occupations, HLMI, and FAOL.
Combined year-framing and probability-framing responses using gamma mixture CDF aggregation
(matching the 2023 report methodology). A gamma CDF is fitted to each respondent's data, then
all CDFs are averaged pointwise to produce a mixture CDF from which percentiles are read.
Sorted by 50% probability year (earliest first). The HLMI row uses the dedicated direct-HLMI
blocks (hb_a_*/hb_b_*); it does not pool the final-occupation
prediction block.
| Milestone | N (2024, fitted) | 2023 10% | 2024 10% | Δ10 (cal yr) | 2023 50% | 2024 50% | Δ50 (cal yr) | 2023 90% | 2024 90% | Δ90 (cal yr) |
|---|---|---|---|---|---|---|---|---|---|---|
| Write Python code (e.g. quicksort) | 111 | 0 yr (2023) | 0 yr (2024) | +1 | 2 yr (2025) | 2 yr (2026) | +1 | 14 yr (2037) | 11 yr (2035) | -2 |
| Write high-school history essay | 129 | 0 yr (2023) | 0 yr (2024) | +1 | 2 yr (2025) | 2 yr (2026) | +1 | 13 yr (2036) | 10 yr (2034) | -2 |
| Play new Angry Birds levels (superhuman) | 167 | 0 yr (2023) | 0 yr (2024) | +1 | 2 yr (2025) | 2 yr (2026) | +1 | 10 yr (2033) | 10 yr (2034) | +1 |
| Win World Series of Poker | 139 | 0 yr (2023) | 0 yr (2024) | +1 | 3 yr (2026) | 2 yr (2026) | ±0 | 20 yr (2043) | 15 yr (2039) | -3 |
| Answer Googleable factoid questions | 144 | 0 yr (2023) | 0 yr (2024) | +1 | 3 yr (2026) | 3 yr (2027) | +1 | 21 yr (2044) | 22 yr (2046) | +2 |
| Group unseen objects into classes | 146 | 0 yr (2023) | 0 yr (2024) | +1 | 4 yr (2027) | 3 yr (2027) | ±0 | 23 yr (2046) | 19 yr (2043) | -3 |
| Voice acting from text | 141 | 0 yr (2023) | 0 yr (2024) | +1 | 3 yr (2026) | 3 yr (2027) | +1 | 17 yr (2040) | 14 yr (2038) | -2 |
| Fluent translation (most languages) | 158 | 0 yr (2023) | 0 yr (2024) | +1 | 3 yr (2026) | 3 yr (2027) | ±0 | 16 yr (2039) | 18 yr (2042) | +3 |
| Answer Googleable open-ended questions | 157 | 0 yr (2023) | 0 yr (2024) | +1 | 3 yr (2026) | 3 yr (2027) | +1 | 23 yr (2046) | 23 yr (2047) | +1 |
| Transcribe speech (noisy, accents) | 135 | 0 yr (2023) | 0 yr (2024) | +1 | 3 yr (2026) | 3 yr (2027) | +1 | 17 yr (2040) | 16 yr (2040) | -1 |
| Beat best Starcraft 2 players | 147 | 0 yr (2023) | 0 yr (2024) | +1 | 4 yr (2027) | 3 yr (2027) | ±0 | 25 yr (2048) | 18 yr (2042) | -7 |
| 3D model from short video | 151 | 0 yr (2023) | 0 yr (2024) | +1 | 5 yr (2028) | 3 yr (2027) | -1 | 27 yr (2050) | 18 yr (2042) | -8 |
| Answer questions with no definite answer | 146 | 0 yr (2023) | 0 yr (2024) | +1 | 4 yr (2027) | 4 yr (2028) | ±0 | 35 yr (2058) | 20 yr (2044) | -14 |
| Produce song indistinguishable from artist | 149 | 0 yr (2023) | 0 yr (2024) | +1 | 4 yr (2027) | 4 yr (2028) | +1 | 26 yr (2049) | 20 yr (2044) | -5 |
| Build website with payment processing | 134 | 0 yr (2023) | 0 yr (2024) | +1 | 5 yr (2028) | 4 yr (2028) | ±0 | 31 yr (2054) | 20 yr (2044) | -9 |
| Translate speech from subtitled films | 120 | 0 yr (2023) | 0 yr (2024) | +1 | 5 yr (2028) | 4 yr (2028) | ±0 | 41 yr (2064) | 24 yr (2048) | -16 |
| Atari novice level (20 min training) | 148 | 0 yr (2023) | 0 yr (2024) | +1 | 5 yr (2028) | 4 yr (2028) | ±0 | 36 yr (2059) | 32 yr (2056) | -3 |
| Phone banking services | 143 | 0 yr (2023) | 0 yr (2024) | +1 | 5 yr (2028) | 4 yr (2028) | ±0 | 30 yr (2053) | 26 yr (2050) | -4 |
| Outperform on all Atari games | 150 | 0 yr (2023) | 0 yr (2024) | +1 | 6 yr (2029) | 5 yr (2029) | ±0 | 35 yr (2058) | 25 yr (2049) | -10 |
| Fine-tune open source LLM | 139 | 1 yr (2024) | 0 yr (2024) | +1 | 5 yr (2028) | 5 yr (2029) | +1 | 47 yr (2070) | 29 yr (2053) | -16 |
| Compose US Top 40 song (full audio) | 148 | 0 yr (2023) | 0 yr (2024) | +1 | 6 yr (2029) | 5 yr (2029) | ±0 | 46 yr (2069) | 47 yr (2071) | +2 |
| Win Putnam math competition | 158 | 1 yr (2024) | 0 yr (2024) | ±0 | 8 yr (2031) | 5 yr (2029) | -2 | 71 yr (2094) | 31 yr (2055) | -39 |
| Translate newly discovered language (Rosetta stone) | 122 | 1 yr (2024) | 0 yr (2024) | ±0 | 7 yr (2030) | 6 yr (2030) | ±0 | 52 yr (2075) | 55 yr (2079) | +4 |
| Play random game as human novice (<10 min) | 140 | 1 yr (2024) | 1 yr (2025) | +1 | 7 yr (2030) | 6 yr (2030) | ±0 | 38 yr (2061) | 38 yr (2062) | ±0 |
| Fold laundry (speed + quality) | 162 | 1 yr (2024) | 1 yr (2025) | +1 | 7 yr (2030) | 6 yr (2030) | ±0 | 40 yr (2063) | 35 yr (2059) | -5 |
| Learn efficient sorting (no solution form) | 135 | 1 yr (2024) | 0 yr (2024) | ±0 | 7 yr (2030) | 6 yr (2030) | +1 | 55 yr (2078) | 51 yr (2075) | -3 |
| Write NYT best-seller novel | 155 | 1 yr (2024) | 1 yr (2025) | +1 | 7 yr (2030) | 7 yr (2031) | ±0 | 62 yr (2085) | 65 yr (2089) | +4 |
| One-shot image recognition | 134 | 0 yr (2023) | 1 yr (2025) | +1 | 5 yr (2028) | 7 yr (2031) | +2 | 34 yr (2057) | 37 yr (2061) | +5 |
| Explain game AI moves to layman | 150 | 1 yr (2024) | 1 yr (2025) | +1 | 8 yr (2031) | 7 yr (2031) | ±0 | 81 yr (2104) | 48 yr (2072) | -32 |
| Beat best Go players (limited training) | 148 | 1 yr (2024) | 1 yr (2025) | +1 | 10 yr (2033) | 7 yr (2031) | -1 | 100 yr (2123) | 88 yr (2112) | -11 |
| Assemble any LEGO set | 141 | 1 yr (2024) | 1 yr (2025) | +1 | 8 yr (2031) | 7 yr (2031) | +1 | 48 yr (2071) | 32 yr (2056) | -15 |
| Retail Salesperson (occ.) | 460 | 1 yr (2024) | 1 yr (2025) | ±0 | 10 yr (2033) | 7 yr (2031) | -1 | 75 yr (2098) | 68 yr (2092) | -6 |
| Find & patch security flaw (100k+ users) | 135 | 2 yr (2025) | 1 yr (2025) | +1 | 10 yr (2033) | 7 yr (2031) | -1 | 87 yr (2110) | 63 yr (2087) | -23 |
| Replicate ML conference study | 140 | 2 yr (2025) | 2 yr (2026) | +1 | 12 yr (2035) | 8 yr (2032) | -3 | 109 yr (2132) | 53 yr (2077) | -56 |
| Truck Driver (occ.) | 459 | 2 yr (2025) | 1 yr (2025) | ±0 | 12 yr (2035) | 9 yr (2033) | -2 | 58 yr (2081) | 46 yr (2070) | -11 |
| Beat fastest human in 5km city race (biped robot) | 125 | 1 yr (2024) | 1 yr (2025) | +1 | 9 yr (2032) | 9 yr (2033) | +1 | 44 yr (2067) | 40 yr (2064) | -3 |
| Discover physics equations from simulation | 132 | 1 yr (2024) | 1 yr (2025) | +1 | 12 yr (2035) | 11 yr (2035) | ±0 | 123 yr (2146) | 127 yr (2151) | +5 |
| Conduct ML research & write conference paper | 164 | 2 yr (2025) | 2 yr (2026) | ±0 | 20 yr (2043) | 12 yr (2036) | -7 | 246 yr (2269) | 159 yr (2183) | -86 |
| Prove publishable math theorems | 149 | 3 yr (2026) | 2 yr (2026) | ±0 | 23 yr (2046) | 15 yr (2039) | -7 | 270 yr (2293) | 258 yr (2282) | -11 |
| Install electrical wiring in new home | 130 | 3 yr (2026) | 3 yr (2027) | +1 | 17 yr (2040) | 16 yr (2040) | ±0 | 104 yr (2127) | 96 yr (2120) | -7 |
| HLMI (all human tasks) | 985 | 4 yr (2027) | 3 yr (2027) | -1 | 24 yr (2047) | 18 yr (2042) | -5 | 174 yr (2197) | 163 yr (2187) | -10 |
| Surgeon (occ.) | 458 | 7 yr (2030) | 5 yr (2029) | ±0 | 33 yr (2056) | 26 yr (2050) | -6 | 310 yr (2333) | 299 yr (2323) | -9 |
| AI Researcher (occ.) | 458 | 6 yr (2029) | 4 yr (2028) | -1 | 40 yr (2063) | 28 yr (2052) | -11 | 1321 yr (3344) | 583 yr (2607) | -737 |
| Solve unsolved math problem (e.g. Millennium) | 130 | 4 yr (2027) | 4 yr (2028) | ±0 | 27 yr (2050) | 30 yr (2054) | +3 | 329 yr (2352) | 766 yr (2790) | +437 |
| Full Automation of Labor | 456 | 14 yr (2037) | 11 yr (2035) | -1 | 89 yr (2112) | 72 yr (2096) | -16 | 2641 yr (4664) | 2195 yr (4219) | -445 |
Year values are years from survey date (2024), with calendar year in parentheses. "10% prob year" = year at which the mixture CDF reaches 10%. "90% prob year" = 90%. Dashes indicate insufficient data or an aggregate CDF that does not reach the target percentile within the 1e8-year "never/infinity" sentinel bound. Δ columns show shift in calendar prediction year between 2024 and 2023 surveys (positive = pushed later; red ≥ +3 yr, green ≤ −3 yr).
3.2 HLMI & Full Automation of Labor Timing
High-Level Machine Intelligence (HLMI)
HLMI is defined as machines that can accomplish every task better and cheaper than human workers.
| Percentile | 2024 Survey | 2023 Survey | Shift |
|---|---|---|---|
| 10% prob | 3 yr (2027) | 3 yr (2027) | -0.4 yr |
| 50% prob | 18 yr (2042) | 23 yr (2047) | -4.6 yr |
N (gamma fits): 985 (0 failed fits)
This HLMI estimate uses only the dedicated direct-HLMI
question blocks (hb_a_* and hb_b_*). It does not
pool the separate final-occupation prediction block
(hj_*_final_pred).
Full Automation of Labor (FAOL)
All occupations fully automatable — machines carry out every task better and more cheaply than humans.
| Percentile | 2024 Survey | 2023 Survey | Shift |
|---|---|---|---|
| 50% prob | 72 yr (2096) | 92 yr (2116) | -19.6 yr |
N (gamma fits): 456 (0 failed fits)
Gap Between HLMI and FAOL
The consistent gap between HLMI and FAOL predictions reflects researchers' view that achieving human-level AI capability differs from fully automating all occupations.
Methodology Note
Gamma mixture CDF method: Following the methodology of Grace et al. (2024),
a gamma CDF is fitted to each respondent's three data points (whether year-framing or probability-framing).
All individual CDFs are then averaged pointwise to produce a mixture CDF, from which percentiles are read off.
This approach handles extrapolation naturally (pessimistic respondents contribute heavy-tailed CDFs rather
than being dropped) and treats both question framings symmetrically.
Respondents were randomly assigned to either a "year framing" (provide years for 10%/50%/90% probability)
or a "probability framing" (provide probabilities at fixed time horizons: 10, 20, and 40 years for
HLMI; 10, 20, and 50 years for FAOL).
See compare_methods.py for a comparison of this method with linear interpolation.
Corresponds to Figure 3 in Grace et al. (2024). Thin lines show individual respondent gamma CDFs (random subset of 200); thick line is the mixture (average) CDF. The dashed line marks 50% probability.
3.2.4 High-Level Machine Intelligence Timing by Experience
This table splits respondents at the median years in field, then refits the direct High-Level Machine Intelligence (HLMI) timing model within each experience group.
| Experience group | Direct HLMI gamma-fit count | Mixture CDF 50% Year | Calendar Year |
|---|---|---|---|
| More experienced (>=6 yr) | 58 | 18.2 yr | 2042 |
| Less experienced (<6 yr) | 55 | 13.4 yr | 2037 |
The count values are not all respondents in each experience group. They count respondents who both answered the years-in-field question and gave enough direct HLMI timing answers to fit a gamma CDF. The separate occupation-automation prediction block is not pooled into this HLMI comparison. Only 189 respondents received the experience question due to randomization. Geographic region and citation count data are not available in the anonymized dataset.
3.2.5 Do Participants Agree on HLMI Timing?
"How much do you think your views on when HLMI will be achieved differ from those of the typical AI researcher?" (N=385)
| Response | Count | Percentage |
|---|---|---|
| Not much | 169 | 43.9% |
| A moderate amount | 174 | 45.2% |
| A lot | 42 | 10.9% |
3.3 Framing Effects
Respondents were randomly assigned to one of two question framings for the same underlying question. The "year framing" asked how many years until a given probability, while the "probability framing" asked for the probability at a fixed time horizon.
HLMI 50% Probability Year by Framing
| Framing | N (gamma fits) | Mixture CDF 50% Year | Calendar Year |
|---|---|---|---|
| Year framing (respondent provides years) | 525 | 13.8 yr | 2038 |
| Probability framing (respondent provides probabilities) | 460 | 26.1 yr | 2050 |
A consistent framing effect has been observed across survey waves: the year-framing tends to produce earlier (shorter-timeline) predictions than the probability-framing. This is a documented cognitive bias in probability elicitation.
3.4 Perceived Rates of Progress
Respondents were asked whether AI progress was faster in the first or second half of their career. (N=191)
Corresponds to Figure 4 in Grace et al. (2024).
| Response | Count | Percentage | 2023 Survey |
|---|---|---|---|
| The second half | 112 | 58.6% | 60% |
| The first half | 36 | 18.8% | — |
| They were about the same | 43 | 22.5% | — |
How Far Along Is AI Progress? (Outside-view slider)
Respondents were shown a slider with three anchors: A = where progress was when they started working in their AI area; B = where it is now; and C = where it would need to be for AI software to have roughly human-level abilities at the tasks they study. They were asked "What fraction of the distance between where progress was when you started working in the area (A) and where it would need to be to attain human-level abilities in the area (C) have we come so far (B)?" on a 0–100 scale. This complements the year-based timing questions by anchoring an estimate to each respondent's own career window — so it is robust to disagreement about absolute calendar dates.
| Survey | N | Median | Mean | SD |
|---|---|---|---|---|
| 2024 | 188 | 30 | 38.0 | 28.4 |
| 2023 | 319 | 30 | 37.6 | 26.0 |
IQR (2024): 14–60. The question was also asked in 2023 (raw values clamped to 0–100; one 2023 respondent entered an out-of-range value). Median and mean barely moved year-on-year, suggesting the field's self-assessment of remaining distance to HLMI is stable.
3.5 What Causes AI Progress?
Respondents estimated how much less AI progress there would have been with half as much of each input. Higher values = more important to progress. Values are percentage less progress (0-100%).
Corresponds to Figure 5 in Grace et al. (2024). Red dots are means; box shows IQR with median line.
| Factor | N | Median | Mean | IQR |
|---|---|---|---|---|
| Computing hardware cost decline | 119 | 60.0% | 58.3% | 45–80% |
| Training dataset effort | 109 | 50.0% | 51.9% | 30–75% |
| AI algorithm progress | 103 | 50.0% | 46.1% | 30–60% |
| Funding | 116 | 40.0% | 44.1% | 25–60% |
| Researcher effort | 95 | 25.0% | 32.7% | 20–50% |
3.6 Will There Be an Intelligence Explosion?
Probability Estimates (numeric, 0-100%)
| Scenario | N | Median | Mean | IQR | 2023 Survey Median | Change |
|---|---|---|---|---|---|---|
| Dramatic tech speedup within 2 years of HLMI | 173 | 20% | 30.3% | 10–50% | 20% | unchanged |
| Dramatic tech speedup within 30 years of HLMI | 172 | 80% | 71.6% | 50–95% | 80% | unchanged |
| Vastly superhuman AI within 2 years of HLMI | 171 | 10% | 20.9% | 1–30% | 10% | unchanged |
| Vastly superhuman AI within 30 years of HLMI | 172 | 60% | 59.6% | 30–95% | — | — |
Is the Feedback Loop Argument Broadly Correct? (N=189)
"If AI does nearly all R&D, improvements in AI will accelerate progress including further AI progress. This could cause progress to become >10x faster within 5 years."
Corresponds to Figure 6 in Grace et al. (2024).
| Response | Count | Percentage |
|---|---|---|
| Quite unlikely (0-20%) | 40 | 21.2% |
| Unlikely (21-40%) | 52 | 27.5% |
| About even chance (41-60%) | 44 | 23.3% |
| Likely (61-80%) | 45 | 23.8% |
| Quite likely (81-100%) | 8 | 4.2% |
3.7 AI Capabilities in 2044
"In 2044, how likely do you think the following will be for at least some state-of-the-art AI systems?" Sorted by percentage rating Likely or Very Likely.
Corresponds to Figure 7 in Grace et al. (2024).
| Capability | N | V.Unlikely | Unlikely | Even | Likely | V.Likely | Likely+V.Likely |
|---|---|---|---|---|---|---|---|
| Talk like an expert human on most topics | 378 | 1% | 5% | 10% | 23% | 61% | 84% |
| Find unexpected ways to achieve goals | 380 | 2% | 4% | 11% | 32% | 51% | 83% |
| Frequently behave surprisingly to humans | 380 | 2% | 11% | 23% | 32% | 32% | 64% |
| Can be jailbroken for illegal commands | 374 | 5% | 10% | 23% | 36% | 27% | 62% |
| Deceive humans to achieve goals (unintended) | 376 | 9% | 16% | 26% | 31% | 19% | 49% |
| Cause important real-world actions (run business, etc.) | 379 | 8% | 19% | 24% | 30% | 18% | 49% |
| Have goals not aligned with human goals | 375 | 10% | 22% | 24% | 31% | 13% | 44% |
| Form AI-AI collaborative relationships (unintended) | 379 | 9% | 21% | 26% | 26% | 17% | 44% |
| Can be trusted to explain their actions | 380 | 8% | 24% | 29% | 28% | 11% | 39% |
| Self-improve regardless of human wishes | 378 | 12% | 22% | 28% | 25% | 13% | 38% |
| Take actions to attain power | 377 | 21% | 28% | 29% | 16% | 7% | 23% |
3.8 Will AI Explain Its Decisions? (2029)
"For typical state-of-the-art AI systems in 2029, will users be able to know the true reasons for decisions?" (N=520)
Corresponds to Figure 8 in Grace et al. (2024).
| Response | Count | Percentage |
|---|---|---|
| Very unlikely (<10%) | 138 | 26.5% |
| Unlikely (10-40%) | 157 | 30.2% |
| Even odds (40-60%) | 112 | 21.5% |
| Likely (60-90%) | 80 | 15.4% |
| Very likely (>90%) | 33 | 6.3% |
4.1 How Concerning Are Future AI Scenarios?
"How would you rate the level of concern these scenarios deserve over the next thirty years?" Sorted by percentage rating Substantial or Extreme concern.
Corresponds to Figure 9 in Grace et al. (2024).
| Scenario | N | None | A Little | Substantial | Extreme | 2024 Sub.+Ext. | 2023 Sub.+Ext. | Δ Sub.+Ext. (pp) |
|---|---|---|---|---|---|---|---|---|
| AI makes it easy to spread false info (deepfakes) | 759 | 2% | 15% | 32% | 51% | 83% | 86% | -2 |
| AI manipulates large-scale public opinion | 757 | 4% | 18% | 36% | 41% | 78% | 78% | -1 |
| AI lets dangerous groups make powerful tools (bioweapons) | 756 | 4% | 24% | 40% | 32% | 72% | 73% | -1 |
| Authoritarian rulers use AI for control | 754 | 7% | 23% | 33% | 36% | 69% | 73% | -3 |
| AI worsens economic inequality | 753 | 7% | 25% | 35% | 34% | 69% | 71% | -3 |
| AI bias worsens unjust situations (hiring, etc.) | 757 | 8% | 35% | 38% | 19% | 57% | 61% | -4 |
| Less human interaction (more time with AI) | 759 | 14% | 37% | 31% | 17% | 49% | 45% | +3 |
| Automation leaves most people economically powerless | 753 | 16% | 37% | 30% | 18% | 48% | 46% | +2 |
| Wrong goals: AI reduces human decision-making role | 756 | 16% | 38% | 30% | 16% | 46% | 44% | +2 |
| Misaligned powerful AI causes catastrophe (weapons) | 755 | 19% | 38% | 25% | 18% | 43% | 43% | +1 |
| Automation makes people struggle to find meaning | 760 | 26% | 38% | 25% | 11% | 37% | 35% | +1 |
4.2 How Good or Bad Will HLMI Be?
Respondents assigned probabilities to five outcome categories (summing to 100%). N=1538
Corresponds to Figure 11 in Grace et al. (2024), "Thousands of AI Authors on the Future of AI."
Mean and Median Probabilities
| Outcome | Mean | Median |
|---|---|---|
| Extremely good | 23.9% | 15.0% |
| On balance good | 27.7% | 25.0% |
| Neutral | 20.7% | 20.0% |
| On balance bad | 17.8% | 15.0% |
| Extremely bad | 9.9% | 5.0% |
Individual Respondent Views
Each vertical slice below is one respondent's probability allocation, sorted from most optimistic (left) to most pessimistic (right). The chart reveals the full diversity of expert opinion.
Corresponds to Figure 10 in Grace et al. (2024).
The same respondents sorted by how much probability they assign to "extremely bad (e.g. human extinction)" outcomes. The growing black band on the right shows those assigning the highest probability to extremely bad outcomes.
No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.
No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.
Individual Response Profiles
Each small bar below represents one respondent's five-category probability distribution, sampled evenly from the most optimistic to the most pessimistic.
No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.
Extreme Outcome Analysis
| Metric | 2024 Survey | 2023 Survey | Change |
|---|---|---|---|
| Non-zero to both extremes | 64.4% | 64.0% | +0.4 pp |
| >=5% on extremely bad | 58.3% | 57.8% | +0.5 pp |
| >=10% on extremely bad | 40.5% | 37.8% | +2.7 pp |
| >=20% on extremely bad | 20.0% | — | — |
| >=25% on extremely bad | 11.9% | — | — |
| Mean extremely bad | 9.9% | 9.0% | +0.9 pp |
| Median extremely bad | 5.0% | 5.0% | unchanged |
| Net optimists | 67.0% | 68.3% | -1.3 pp |
4.3 How Likely Is AI to Cause Extinction or Severe Disempowerment?
Each respondent was randomly assigned ONE of three questions about the probability that future AI advances cause human extinction or similarly permanent and severe disempowerment. The chart below also includes "extremely bad" HLMI outcome probabilities from Section 4.2 for comparison.
Corresponds to Figure 13 in Grace et al. (2024).
Note: The "Extremely bad HLMI outcome" question was asked to all respondents (N is larger), while each extinction/disempowerment question was asked to a random subset of respondents.
Distribution of Individual Estimates
The 2024 estimates were higher for the unconditional and within-100-years questions, while estimates for the loss-of-control question were similar to 2023.
The filled curves show 2024 responses; dashed lines show 2023. Respondents are ranked independently within each year and question framing, so the horizontal axis is percentile rather than respondent count. Each question was asked to a different random subset of respondents.
Corresponds to Figure 12 in Grace et al. (2024).
| Question | N | Mean | 2023 Survey Mean | Change | Median | 2023 Survey Median | Change | >=10% | >=25% |
|---|---|---|---|---|---|---|---|---|---|
| Future AI causes extinction or severe disempowerment | 744 | 18.3% | 16.2% | +2.1 pp | 10.0% | 5.0% | +5.0 pp | 52.7% | 26.6% |
| Inability to control advanced AI causes extinction/disempowerment | 392 | 18.5% | 19.4% | -0.9 pp | 9.0% | 10.0% | -1.0 pp | 50.0% | 26.8% |
| AI-caused extinction/disempowerment within 100 years | 353 | 17.5% | 14.4% | +3.1 pp | 5.0% | 5.0% | unchanged | 49.0% | 25.8% |
| Pooled (all variants combined) | 1489 | 18.2% | — | — | 10.0% | — | — | 51.1% | 26.5% |
The "Pooled" row combines the raw responses from all variants above (744 + 392 + 353 = 1489 responses) and takes the median of the combined data. Because each respondent was randomly assigned exactly one variant, no respondent is counted twice. And because the other variants each add a constraint (a specific cause, a time limit) to the basic question, every response is a lower bound on that respondent's unconstrained probability — so the pooled median is a conservative basis for "the median researcher put at least 10%..." claims.
4.4 Are Future AI-Risk Concerns Due to Misunderstandings of AI Research?
"To what extent do you think people's concerns about future risks from AI are due to misunderstandings of AI research?" (N=386)
No direct figure equivalent in Grace et al. (2024); this data is discussed in Section 4.4 of that paper.
| Response | Count | Percentage |
|---|---|---|
| Hardly at all | 10 | 2.6% |
| Not much | 61 | 15.8% |
| Somewhat | 125 | 32.4% |
| To a large extent | 163 | 42.2% |
| Almost entirely | 27 | 7.0% |
4.5 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity
Question wording: "What rate of global AI progress over the next five years would make you feel most optimistic for humanity's future? Assume any change in speed affects all projects equally." (N=382)
Corresponds to Table 3 in Grace et al. (2024); no figure equivalent in that paper.
| Response | Count | Percentage | 2023 Survey | Change |
|---|---|---|---|---|
| Much slower | 34 | 8.9% | 4.8% | +4.1 pp |
| Somewhat slower | 95 | 24.9% | 29.9% | -5.0 pp |
| Current speed | 112 | 29.3% | 26.9% | +2.4 pp |
| Somewhat faster | 75 | 19.6% | 22.8% | -3.2 pp |
| Much faster | 56 | 14.7% | 15.6% | -0.9 pp |
| Other | 10 | 2.6% | — | — |
4.6 How Much Should AI Safety Research Be Prioritized?
"How much should society prioritize AI safety research, relative to how much it is currently prioritized?" (N=370)
Corresponds to Figure 14 in Grace et al. (2024).
| Response | Count | Percentage |
|---|---|---|
| Much less | 8 | 2.2% |
| Less | 24 | 6.5% |
| About the same | 76 | 20.5% |
| More | 131 | 35.4% |
| Much more | 131 | 35.4% |
4.7 The Alignment Problem
Corresponds to Figure 15 in Grace et al. (2024).
Do you think this argument points at an important problem? (N=755)
| Response | Count | Percentage |
|---|---|---|
| No, not a real problem. | 31 | 4.1% |
| No, not an important problem. | 93 | 12.3% |
| Yes, a moderately important problem. | 266 | 35.2% |
| Yes, a very important problem. | 278 | 36.8% |
| Yes, among the most important problems in the field. | 87 | 11.5% |
How valuable is it to work on this problem today, compared to other problems in AI? (N=753)
| Response | Count | Percentage |
|---|---|---|
| Much less valuable | 62 | 8.2% |
| Less valuable | 178 | 23.6% |
| As valuable as other problems | 270 | 35.9% |
| More valuable | 189 | 25.1% |
| Much more valuable | 54 | 7.2% |
How hard do you think this problem is compared to other problems in AI? (N=752)
| Response | Count | Percentage |
|---|---|---|
| Much easier | 15 | 2.0% |
| Easier | 74 | 9.8% |
| As hard as other problems | 241 | 32.0% |
| Harder | 266 | 35.4% |
| Much harder | 156 | 20.7% |
5. Results by Value-Outlook Cluster
5.0 Methodology: Identifying Outlook Groups
One of the survey's core questions asks respondents to distribute 100 probability points across five possible outcomes of high-level machine intelligence: extremely good, on balance good, more or less neutral, on balance bad, and extremely bad. These five numbers form a compact signature of each respondent's overall outlook on advanced AI.
We apply a Gaussian Mixture Model (GMM) with K=4 components to these five-dimensional signatures (standardized to zero mean and unit variance). GMM is a soft-clustering method that models the data as a mixture of multivariate normal distributions; each respondent is assigned to the component with the highest posterior probability. We use 20 random initializations to avoid local optima.
The four resulting clusters are then named by inspecting each cluster's mean probability profile:
- Strong Optimists — assign high probability to extremely good outcomes and very little to bad outcomes.
- Mild Optimists — lean positive overall (combined good > 55%) but with more hedging and moderate expectations.
- Concerned — assign elevated probability to bad or extremely bad outcomes (combined pessimism > 30%).
- Polarized/Bimodal — the most distinctive group: they assign substantial probability to both extremely good and extremely bad outcomes, reflecting a worldview where advanced AI is seen as a high-stakes gamble rather than a clearly positive or negative development.
This four-way split captures meaningful variation in worldview that cuts across traditional demographic variables like experience level. 1,538 of 1,793 cleaned analysis rows had complete value-outlook data and were assigned to a cluster. The remaining subsections re-examine key survey results through this lens.
The bimodal outlook is not merely an artifact of the mixture model. Among the 1,538 respondents who answered both endpoints of the value question, 474 (31%) assigned at least 10% probability to both an extremely good and an extremely bad outcome, and 74 (5%) assigned at least 25% to both. This model-free count confirms that a substantial minority genuinely hedges across both extremes rather than settling on a single direction.
5.0.1 Cluster Profiles
| Cluster | N | % of total | Ext. good | Good | Neutral | Bad | Ext. bad |
|---|---|---|---|---|---|---|---|
| Strong Optimists | 457 | 30% | 33% | 34% | 22% | 11% | 0% |
| Mild Optimists | 460 | 30% | 17% | 38% | 22% | 15% | 8% |
| Concerned | 329 | 21% | 6% | 17% | 29% | 35% | 12% |
| Polarized/Bimodal | 292 | 19% | 40% | 14% | 6% | 13% | 27% |
5.1 HLMI Timeline Predictions
How soon do different outlook groups expect human-level machine intelligence?
| Cluster | N | Median HLMI year | IQR |
|---|---|---|---|
| Strong Optimists | 160 | 2039 | 2032–2074 |
| Mild Optimists | 156 | 2039 | 2031–2054 |
| Concerned | 97 | 2044 | 2034–2074 |
| Polarized/Bimodal | 105 | 2034 | 2031–2054 |
5.2 Extinction/Disempowerment Estimates
Probability that future AI causes human extinction or permanent severe disempowerment, broken down by outlook cluster.
| Cluster | N | Mean | Median | ≥10% | ≥25% |
|---|---|---|---|---|---|
| Strong Optimists | 223 | 11.7% | 1% | 37% | 15% |
| Mild Optimists | 210 | 14.2% | 8% | 50% | 18% |
| Concerned | 166 | 21.3% | 10% | 63% | 35% |
| Polarized/Bimodal | 145 | 31.1% | 20% | 69% | 47% |
5.3 Concerning Scenarios
Mean concern level (0 = no concern, 3 = extreme concern) for each of 11 AI risk scenarios. The table below highlights which scenarios show the largest divergence across clusters.
| Scenario | Highest cluster | Score | Lowest cluster | Score | Gap |
|---|---|---|---|---|---|
| Wrong goals: AI reduces human decision-making role | Concerned | 1.72 | Strong Optimists | 1.11 | 0.61 |
| Automation leaves most people economically powerless | Concerned | 1.75 | Strong Optimists | 1.14 | 0.61 |
| Authoritarian rulers use AI for control | Concerned | 2.22 | Strong Optimists | 1.64 | 0.58 |
| Misaligned powerful AI causes catastrophe (weapons) | Polarized/Bimodal | 1.65 | Strong Optimists | 1.09 | 0.56 |
| AI worsens economic inequality | Concerned | 2.23 | Strong Optimists | 1.68 | 0.55 |
| Automation makes people struggle to find meaning | Concerned | 1.50 | Strong Optimists | 1.01 | 0.49 |
| AI manipulates large-scale public opinion | Concerned | 2.41 | Strong Optimists | 1.95 | 0.46 |
| Less human interaction (more time with AI) | Concerned | 1.77 | Strong Optimists | 1.35 | 0.42 |
| AI bias worsens unjust situations (hiring, etc.) | Concerned | 1.91 | Polarized/Bimodal | 1.54 | 0.37 |
| AI lets dangerous groups make powerful tools (bioweapons) | Concerned | 2.11 | Strong Optimists | 1.80 | 0.31 |
| AI makes it easy to spread false info (deepfakes) | Concerned | 2.41 | Polarized/Bimodal | 2.22 | 0.19 |
5.4 Expected AI Capabilities in 2044
Mean rated likelihood (0 = very unlikely, 4 = very likely) that AI systems will exhibit each capability by 2044.
5.5 Intelligence Explosion
Median probability estimates for dramatic AI capability speedup and superhuman AI emergence, by outlook cluster.
| Cluster | Dramatic speedup within 2yr (median %) | Dramatic speedup within 30yr (median %) | Superhuman within 2yr (median %) | N |
|---|---|---|---|---|
| Strong Optimists | 20% | 80% | 8% | 50 |
| Mild Optimists | 20% | 80% | 10% | 55 |
| Concerned | 15% | 78% | 10% | 37 |
| Polarized/Bimodal | 32% | 95% | 10% | 28 |
5.6 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity
What rate of global AI progress over the next five years would make each cluster feel most optimistic for humanity's future?
5.7 Safety and Alignment-Problem Views
Mean scores on whether the alignment argument points at an important problem, how much society should prioritize AI safety research, and how valuable it is to work on the alignment problem today compared with other AI problems (all 0–4 scales).
| Cluster | Alignment-problem importance (0–4) | Safety research priority (0–4) | Value of working on alignment problem today vs other AI problems (0–4) | N |
|---|---|---|---|---|
| Strong Optimists | 2.17 | 2.69 | 1.78 | 228 |
| Mild Optimists | 2.44 | 2.82 | 1.99 | 220 |
| Concerned | 2.46 | 3.25 | 2.12 | 162 |
| Polarized/Bimodal | 2.61 | 3.29 | 2.19 | 145 |
5.8 Counterfactual Progress Reduction if Factors Were Halved
Median estimated progress reduction if each factor were halved, by cluster.
Cluster assignments are based on GMM (K=4, 20 random initializations) applied to the five HLMI value-outcome probabilities (vb_1_1–vb_1_5). Respondents missing all five values are excluded (1,538 of 1,793 cleaned analysis rows assigned). Sample sizes vary across subsections due to block randomization.
Appendix: Cluster Analysis
A data-driven exploration of the natural groupings, the safety divide, and hardware vs. software beliefs among 1,793 ESPAI 2024 cleaned analysis rows.
Executive Summary
1. How Many Camps? The Natural Clusters
1.1 Clustering on Value Outlook (n=1,538)
The value-of-HLMI question asked respondents to assign probabilities (summing to 100%) across five outcomes: extremely good, on balance good, neutral, on balance bad, and extremely bad. It is the highest-coverage feature in this analysis. We cluster on these 5 dimensions using Gaussian Mixture Models.
Figure 1: Silhouette scores for K=2 through K=6 clusters on value outlook. K=2 has the highest silhouette, but K=3 and K=4 offer more interpretable structure.
Silhouette scores are modest (0.10-0.12), indicating the clusters are not sharply separated -- this is a continuous landscape of opinion, not discrete tribes. Still, the structure is meaningful.
1.2 Four-Camp Solution
Figure 2: Four natural camps in value outlook. Left: PCA projection. Center: mean probability profiles. Right: cluster sizes.
The most striking finding here is the Polarized/Bimodal group. These respondents assign high probability to both extremely good and extremely bad outcomes -- their P(extremely good) is comparable to the Strong Optimists, yet they simultaneously assign substantial probability to catastrophic outcomes. This is not a group of doomers or technophobes. They are researchers who believe AI will be hugely impactful, but who are genuinely torn on whether that impact will be positive or negative. The key axis for this group is magnitude of impact, not direction -- they have rejected the possibility that AI will be a modest or neutral development.
1.3 Three-Camp Solution (Simpler View)
Collapsing to 3 clusters (which has a higher silhouette score of 0.18 vs 0.07) merges the finer distinctions into a simpler optimist/moderate/pessimist framing. This loses the polarized group but provides a cleaner summary:
Figure 3: Three-camp simplification. The polarized group is absorbed into the moderate/pessimist clusters.
1.4 Intuitive View: How Soon vs. How Good
The PCA axes above lack intuitive meaning. Below, we plot each respondent on two directly interpretable dimensions: their predicted HLMI arrival year (x-axis) and their net optimism score (y-axis), defined as P(good + extremely good) minus P(bad + extremely bad). This reveals where the four camps sit in the space of "how soon" vs. "how beneficial."
Figure 3b: Respondents (n=935) binned by predicted HLMI year and net optimism. Each bubble's size shows how many researchers from that cluster fall in the bin (labeled when ≥5). The vertical dashed line marks the median predicted year; the horizontal line separates net optimists from net pessimists.
1.5 Combined Clustering: Values + Concerns + Safety + Extinction
For the 191 respondents who answered all of: value outlook, concern scenarios, alignment-problem importance, and extinction/disempowerment probability, we ran a richer clustering incorporating 8 features.
Figure 4: Combined clustering incorporating worldview, concerns, safety/alignment-problem views, and extinction/disempowerment probability. Mean concern averages 11 scenario concern ratings on a 0-3 scale (0=no concern, 3=extreme concern); alignment-problem importance is a 0-4 ordinal score (0=not a real problem, 4=among the most important problems in the field); P(extinction/disempowerment) is a percentage.
| Cluster | Size | P(Good / Ext good) | P(Bad / Ext bad) | Mean concern | Alignment-problem imp | P(extinction/disempowerment) | HLMI Year median [IQR]; mean |
|---|---|---|---|---|---|---|---|
| Optimistic / Low x-risk | 103 (54%) | 24% / 35% | 12% / 3% | 1.6/3 | 2.3/4 | 4% | 2049 [2034-2074]; mean 16981 (n=67) |
| Pessimistic / High x-risk | 88 (46%) | 25% / 18% | 23% / 18% | 1.9/3 | 2.5/4 | 34% | 2043 [2032-2064]; mean 2056 (n=56) |
HLMI year summaries include the median, interquartile range, mean, and item-level n because several groups share the same median year while their wider distributions differ.
2. Higher vs Lower Safety/Severe-Risk Concern
2.1 Defining the Groups
We built a safety composite score from (normalized 0-1 and averaged):
- Alignment-problem importance rating (0-4 ordinal scale: not a real problem to among the most important problems in the field)
- Value of working on the alignment problem today compared with other AI problems (0-4 ordinal scale)
- P(extinction/disempowerment from AI) (0-100%)
Requiring at least 2 of 3 components, we obtained scores for 754 respondents and split at the median (0.500) into High (n=330) and Low (n=424) safety concern groups.
2.2 The Full Comparison
Figure 5: Comprehensive comparison of higher vs lower safety/severe-risk-concern respondents across six dimensions. Alignment-problem importance uses the 0-4 ordinal scale above; mean concern averages 11 scenario concern ratings on a 0-3 scale; the optimism-maximizing AI progress-rate answer is coded 0=much slower, 2=current speed, 4=much faster.
2.3 Statistical Tests
| Figure 5 panel | Tested variable | n (High) | n (Low) | Median (High) | Median (Low) | Mean (High) | Mean (Low) | p-value | Effect r |
|---|---|---|---|---|---|---|---|---|---|
| HLMI timeline | HLMI Year | 210 | 258 | 2044.0 | 2044.0 | 478730.0 | 25329.1 | 0.0491 * | 0.105 |
| Extinction/disempowerment estimate | P(extinction/disempowerment) | 130 | 234 | 30.0 | 5.0 | 31.7 | 10.8 | 0.0000 *** | -0.466 |
| Value outlook | P(Extremely bad) | 330 | 424 | 5.0 | 5.0 | 12.4 | 7.3 | 0.0000 *** | -0.182 |
| Value outlook | P(Extremely good) | 330 | 424 | 10.0 | 20.0 | 22.3 | 24.8 | 0.0568 n.s. | 0.080 |
| Concern levels | Mean concern | 163 | 220 | 1.9 | 1.6 | 1.9 | 1.6 | 0.0000 *** | -0.348 |
| AI progress rate for optimism | Optimism-maximizing AI progress rate | 81 | 103 | 2.0 | 2.0 | 1.9 | 2.2 | 0.0501 n.s. | 0.164 |
Selected Mann-Whitney U tests for scalar summaries from Figure 5. The value-outlook panel is summarized by P(extremely good) and P(extremely bad), the concern panel by mean concern, and the AI-capabilities panel is tested item-by-item in the next table. Effect size r is rank-biserial correlation (|r| > 0.3 = medium, |r| > 0.5 = large). Scale notes: mean concern 0-3, optimism-maximizing AI progress-rate answer 0-4, probabilities in percentage points. *** p<0.001, ** p<0.01, * p<0.05
2.4 AI Capabilities by 2044: What Do Safety People Expect?
Safety-concerned people don't just worry more -- they have systematically different expectations for what AI will be able to do by 2044:
| Capability | Hypothesized direction | Mean (High Safety) | Mean (Low Safety) | High - Low | Observed direction | p-value | n |
|---|---|---|---|---|---|---|---|
| Talk like expert | Exploratory | 3.41 | 3.34 | +0.07 | High > Low | 0.7903 n.s. | 199 |
| Self-improve regardless | Exploratory | 2.05 | 2.00 | +0.05 | High > Low | 0.7481 n.s. | 199 |
| AI-AI collaborations | High > Low | 2.37 | 2.11 | +0.26 | High > Low | 0.1911 n.s. | 199 |
| Deceive humans | High > Low | 2.52 | 2.23 | +0.29 | High > Low | 0.0820 n.s. | 197 |
| Unexpected strategies | Exploratory | 3.26 | 3.34 | -0.07 | High < Low | 0.3043 n.s. | 200 |
| Seek power | High > Low | 1.74 | 1.35 | +0.38 | High > Low | 0.0104 * | 200 |
| Explain actions (trustworthy) | Exploratory | 2.08 | 2.10 | -0.02 | High < Low | 0.7394 n.s. | 200 |
| Can be jailbroken | Exploratory | 2.86 | 2.68 | +0.18 | High > Low | 0.3029 n.s. | 194 |
| Surprising behavior | Exploratory | 2.76 | 2.91 | -0.15 | High < Low | 0.3611 n.s. | 200 |
| Real-world actions | Exploratory | 2.37 | 2.19 | +0.18 | High > Low | 0.2788 n.s. | 199 |
| Misaligned goals | High > Low | 2.46 | 1.94 | +0.53 | High > Low | 0.0016 ** | 199 |
Scale: 0=Very unlikely, 1=Unlikely, 2=Even chance, 3=Likely, 4=Very likely. High - Low is the high-safety mean minus the low-safety mean, so positive values mean high-safety respondents rated the capability as more likely. Hypothesized direction is marked for risk-relevant capabilities where we expected High > Low; other rows are exploratory.
2.5 Intelligence Explosion Beliefs
Figure 6: Probability estimates for intelligence explosion scenarios, by safety concern level.
3. Hardware vs. Software Progress Beliefs
3.1 Counterfactual Progress Reduction if Factors Were Halved
Respondents rated how much AI progress would decrease if each of 5 factors were cut in half (0-100% scale). Higher values mean the factor is more important.
Figure 7: Counterfactual progress-factor analysis. Top-left: estimated progress reduction distributions. Top-right: hardware vs algorithm progress-reduction estimates. Bottom-left: HW-SW index distribution. Bottom-right: factor correlations with other variables.
- Computing hardware: 60% (n=119) -- largest median estimated progress reduction if halved
- Training data: 50% (n=109)
- Algorithm progress: 50% (n=103)
- Funding: 40% (n=116)
- Researcher effort: 25% (n=95) -- surprisingly lowest
3.2 Hardware vs Software Index
We computed (Hardware - Algorithms) / (Hardware + Algorithms) for each respondent who rated both. Positive = hardware-leaning, negative = software-leaning.
- Hardware-leaning: 35 (60%)
- Software-leaning: 18 (31%)
- Balanced: 5
- Median index: 0.15 (slight hardware lean)
3.3 Do Progress Beliefs Predict Other Views?
Significant cause-factor correlations (p < 0.10)
| Cause Factor | Target | Spearman rho | p-value | n |
|---|---|---|---|---|
| Researcher effort | HLMI Year | -0.238 | 0.0647 n.s. | 61 |
| Computing hardware | Alignment-problem imp | -0.316 | 0.0086 ** | 68 |
| Training data | Mean concern | 0.337 | 0.0166 * | 50 |
Hardware-Software Index Correlations
| Variable 1 | Variable 2 | Spearman rho | p-value | n |
|---|---|---|---|---|
| HW-vs-SW Index | P(extinction/disempowerment) | 0.089 | 0.6212 n.s. | 33 |
| HW-vs-SW Index | HLMI Year | 0.073 | 0.6758 n.s. | 35 |
| HW-vs-SW Index | P(Ext bad) | -0.165 | 0.2152 n.s. | 58 |
| HW-vs-SW Index | Alignment-problem imp | -0.369 | 0.0344 * | 33 |
| HW-vs-SW Index | Mean concern | 0.106 | 0.5847 n.s. | 29 |
| HW-vs-SW Index | P(Ext good) | 0.048 | 0.7207 n.s. | 58 |
3.4 Progress Beliefs by Outlook and Safety Groups
Do optimists, pessimists, and safety-concerned researchers differ in what they think drives AI progress? Below we break down cause factor importance by outlook cluster (left) and safety group (right).
Figure 7b: Mean importance ratings for each progress factor, split by outlook group (left) and safety group (right). Scale: 0-100% estimated decrease in progress if factor halved.
4. How Outlook Predicts Everything Else
The three value-outlook clusters (from Section 1) predict views across many other dimensions:
| Variable | Concerned | Mild Optimists | Strong Optimists | Polarized/Bimodal |
|---|---|---|---|---|
| HLMI Year | 2046 [2034-2074]; mean 533995 (n=188) | 2044 [2032-2064]; mean 19307 (n=290) | 2044 [2034-2064]; mean 1101046 (n=274) | 2044 [2033-2069]; mean 2123 (n=183) |
| P(extinction/disempowerment) | 10.0 (n=166) | 7.5 (n=210) | 1.0 (n=223) | 20.0 (n=145) |
| Alignment-problem importance | 3.0 (n=162) | 3.0 (n=220) | 2.0 (n=228) | 3.0 (n=145) |
| Mean concern | 2.0 (n=165) | 1.7 (n=243) | 1.5 (n=198) | 1.8 (n=148) |
| Optimism-maximizing AI progress rate | 1.0 (n=75) | 2.0 (n=111) | 2.0 (n=109) | 2.0 (n=77) |
Values are medians with sample sizes in parentheses, except HLMI Year, which shows median [IQR], mean, and item-level n because several groups share the same median year. Scale notes: alignment-problem importance 0-4, mean concern 0-3, optimism-maximizing AI progress-rate answer 0=much slower to 4=much faster, probabilities in percentage points.
Figure 8: How the four value-outlook camps compare across timelines, extinction/disempowerment probability, safety/alignment-problem views, concern levels, optimism-maximizing AI progress-rate answers, and P(extremely bad). Mean concern averages 11 scenario concern ratings on a 0-3 scale; alignment-problem importance and optimism-maximizing AI progress-rate answers use 0-4 ordinal scales.
5. What Dimensions Structure AI Researcher Beliefs?
Principal Component Analysis on respondents with values + concerns + safety + extinction/disempowerment data (n=180) reveals the latent axes of disagreement.
Figure 9: PCA factor loadings. Green bars = positive loading, red = negative. Each panel shows one principal component.
6. The Full Correlation Structure
Figure 10: Spearman rank correlations between key variables. Stars indicate significance. Sample sizes shown in each cell. Alignment-problem importance uses a 0-4 ordinal scale from not a real problem to among the field's most important problems; value of working on the alignment problem today uses a 0-4 ordinal scale from much less valuable to much more valuable than other AI problems; mean concern uses a 0-3 scale; HLMI Year is a calendar-year estimate; probabilities are 0-100 percentages.
Key Pairwise Correlations
| Relationship | Spearman rho | p-value | n |
|---|---|---|---|
| P(extinction/disempowerment) x HLMI Year | -0.216 | 0.0000 *** | 446 |
| P(extinction/disempowerment) x P(Ext bad) | 0.436 | 0.0000 *** | 744 |
| Alignment-problem imp x P(extinction/disempowerment) | 0.141 | 0.0072 ** | 364 |
| Alignment-problem imp x HLMI Year | -0.057 | 0.2211 n.s. | 468 |
| Alignment-problem imp x Mean concern | 0.276 | 0.0000 *** | 383 |
| P(Ext good) x P(Ext bad) | -0.113 | 0.0000 *** | 1538 |
| HLMI Year x P(Ext bad) | -0.025 | 0.4421 n.s. | 935 |
| HLMI Year x P(Ext good) | -0.074 | 0.0228 * | 935 |
| Mean concern x P(extinction/disempowerment) | 0.321 | 0.0000 *** | 381 |
| Mean concern x HLMI Year | -0.102 | 0.0293 * | 461 |
| Optimism-max progress rate x P(extinction/disempowerment) | -0.158 | 0.0327 * | 182 |
| Optimism-max progress rate x HLMI Year | -0.210 | 0.0010 *** | 242 |
| Optimism-max progress rate x alignment-problem imp | -0.016 | 0.8289 n.s. | 184 |
- Alignment-problem importance and value of working on the alignment problem today are highly correlated (these aren't independent beliefs -- they form a coherent safety/alignment-problem view)
- P(extinction/disempowerment) and P(Extremely bad) are strongly linked -- people who give high extinction/disempowerment probability estimates also see HLMI as likely to be extremely bad
- P(extinction/disempowerment) and HLMI timeline are negatively correlated -- people with the highest extinction/disempowerment estimates tend to predict earlier arrival of HLMI, not later
- Optimism-maximizing AI progress rate and HLMI timeline are negatively correlated -- people with shorter timelines more often choose slower progress as the rate that would make them most optimistic for humanity's future
7. Methods & Caveats
7.1 Approach
- Strategy: Multiple focused analyses on overlapping subsets (rather than one clustering requiring all features, which drops to n~180). This maximizes statistical power while giving complementary views.
- Clustering: Gaussian Mixture Models (GMM) with full covariance, 10-20 random initializations, model selection by BIC with parsimony preference
- Statistical tests: Non-parametric throughout (Mann-Whitney U for two groups, Kruskal-Wallis for three+, Spearman rank correlations)
- Visualization: PCA for dimensionality reduction (visualization only -- clustering done in full feature space)
7.2 Caveats
7.3 Technical Details
- Random seed: 42
- All tests two-sided; no multiple comparison correction (exploratory analysis)
- HLMI year capped at 2150 for visualization; uncapped for statistics where noted
Appendix: Survey Flow
Survey-flow / randomization diagram modeled on Figure 16 of Grace et al., Thousands of AI Authors on the Future of AI. Each box is a question block. Percentages and n's are realized fill rates: the realized number of respondents who answered each block, divided by the total N. The total N counts everyone who answered at least one question (respondents who answered none are excluded).
Jobs / FAOL Sample-Size Reconciliation
The survey-flow boxes count anyone with any response in the broader Jobs / FAOL block. FAOL analyses use only the FAOL triplet, so their sample sizes can be slightly smaller.
| Count definition | Fixed-years | Fixed-probabilities |
|---|---|---|
| Whole Jobs / FAOL block (survey-flow box) | 234 | 241 |
| FAOL only, any triplet answer | 233 | 234 |
| FAOL only, complete triplet | 233 | 223 |
Here, "fixed-years" means respondents gave probabilities for fixed
time horizons (hj_b_full_*), while "fixed-probabilities" means respondents
gave years for fixed probabilities (hj_a_full_*).
Tasks: ta_* columns are the fixed-probability framing,
tb_* the fixed-years framing (verified empirically:
ta_* rows have fixedprobabilities=1,
tb_* rows have fixedprobabilities=0).
The Qualtrics survey-flow definition (.qsf) is not in
this repo, so the exact display order may differ from the diagram.
Appendix: Supplementary Figures
B.1 fixed-prob/fixed-year CDFs
Full CDF comparison between the fixed-prob and fixed-year conditions for HLMI and FAOL. The fixed-prob framing consistently produces earlier predictions (the red curve is shifted left relative to blue).
Corresponds to Figure 18 in Grace et al. (2024). Each curve is the mixture (mean) CDF for respondents assigned to that framing condition.
B.2 Bootstrap Confidence Bands
95% bootstrap confidence intervals on the aggregate CDF, obtained by resampling the fitted individual CDFs with replacement (500 resamples).
| Milestone | 50% Year (point est.) | 95% Bootstrap CI |
|---|---|---|
| HLMI | 2042 (18.4 yr) | 2041–2044 (16.9–20.0 yr) |
| FAOL | 2096 (72.4 yr) | 2088–2105 (64.0–81.5 yr) |
Bootstrap CIs reflect sampling variability — if we surveyed a different random sample of AI researchers, how much would the aggregate change? This is distinct from the spread of individual predictions (which is much wider).
Appendix: Statistical Tests
F.1 Yuen's Trimmed-Mean Bootstrap Test: Demographics
Tests whether experienced researchers (≥6 years in field) give significantly different HLMI predictions than junior researchers. Uses a 10% trimmed mean with 5,000 bootstrap resamples.
| Comparison | Trimmed-mean diff | 95% CI | p-value | Significant (α=0.05) |
|---|---|---|---|---|
| HLMI 50% year (experienced − junior, split at 6 yr) | +10.4 yr | [-3.6, +43.4] | 1.000 | No |
N (experienced): 58, N (junior): 55. Individual 50% years are computed from per-respondent gamma fits. The trimmed mean reduces sensitivity to extreme predictions.
F.2 Framing Effect: Statistical Significance
Tests whether year-framing and probability-framing respondents give significantly different predictions, using the same Yuen's bootstrap method.
| Milestone | Year-framing median | Prob-framing median | Trimmed-mean diff | 95% CI | p-value |
|---|---|---|---|---|---|
| HLMI | 14.8 yr (n=525) | 28.6 yr (n=459) | -17.5 yr | [-28.5, -10.4] | 0.005 |
| FAOL | 49.9 yr (n=223) | 123.2 yr (n=233) | -477.7 yr | [-641.6, -316.3] | 0.000 |
A negative difference (yr_50 − pr_50) means the year-framing produces earlier predictions. Medians shown for reference; the statistical test uses trimmed means.
Appendix: Aggregation Method Comparison
Different choices in how individual survey responses are fitted and combined into aggregate forecasts can shift the headline numbers. This appendix compares three approaches on the key milestones, motivated by Adamczewski's reanalysis of ESPAI data.
Methods compared
Mix-Mean (MSE) — the Grace et al. 2023 method. Fit a gamma CDF to each respondent's 3 data points using mean squared error, then take the pointwise mean of all fitted CDFs.
Mix-Median (MSE) — same fitting, but take the pointwise median instead of mean. More robust to outlier predictions (e.g. respondents predicting 100M years). Adamczewski argues the median better represents the "typical expert."
Mix-Mean (LogLoss) — fit using log-loss (cross-entropy) instead of MSE, then take the mean. Log-loss naturally weighs errors near p=0 and p=1 more heavily than errors near p=0.5, matching the intuition that a 5% error at p=0.95 represents a much larger shift in belief than the same error at p=0.50.
Results
| Milestone | Mix-Mean (MSE) | Mix-Median (MSE) | Mix-Mean (LogLoss) | N (MSE) | N (LogLoss) |
|---|---|---|---|---|---|
| HLMI | 2042 (18.4 yr) | 2044 (20.0 yr) | 2043 (19.0 yr) | 985 | 985 |
| FAOL | 2096 (72.4 yr) | 2101 (77.1 yr) | 2096 (72.0 yr) | 456 | 456 |
| Truck Driver | 2033 (8.5 yr) | 2034 (9.7 yr) | 2033 (8.5 yr) | 459 | 459 |
| Surgeon | 2050 (26.2 yr) | 2053 (29.3 yr) | 2051 (26.6 yr) | 458 | 458 |
| Retail Salesperson | 2031 (7.4 yr) | 2033 (8.7 yr) | 2031 (7.4 yr) | 460 | 460 |
| AI Researcher | 2052 (27.7 yr) | 2055 (30.6 yr) | 2052 (27.6 yr) | 458 | 458 |
Key takeaways
Aggregation method matters more than loss function. Switching from mean to median aggregation pushes HLMI from 2042 to 2044 and FAOL from 2096 to 2101. Switching the loss function (MSE→log-loss) barely changes the aggregate — consistent with Adamczewski's finding that fitting choice has "hardly any impact."
The effect is largest for far-future milestones (FAOL, AI Researcher) where a few very pessimistic respondents pull the mean CDF rightward.
See also: compare_methods.py for the full 5-method comparison
across all 39 tasks, including linear interpolation and gamma-individual methods.
Method comparison motivated by
Adamczewski (bayes.net/espai).
Data: 2024 Expert Survey on Progress in AI (ESPAI).
Previous survey comparison values from "Thousands of AI Authors on the Future of AI" (2024 preprint).