2023 ESPAI Survey Analysis
Expert Survey on Progress in AI — Comprehensive Data Report
Generated 2026-07-09 00:38
1. Overview & Survey Methodology
Randomization Blocks
| Block | Count | Percentage |
|---|---|---|
| Block 1 | 1,122 | 34.3% |
| Block 2 | 1,085 | 33.2% |
| Block 3 | 1,063 | 32.5% |
Data Cleaning Notes
The following cleaning steps were applied (see cleaning_log.md for full details):
This dataset was pre-cleaned upstream. Column renaming was applied; no further numeric cleaning was needed.
3.1 When Will 39 AI Milestones Be Feasible?
Corresponds to Figure 1 in Grace et al. (2024). Dots show the 50% probability year from the mixture CDF; horizontal lines show the 10%–90% range. The x-axis is capped at 2100; milestones whose 50% or 90% year falls beyond the cap are labeled with an arrow showing the true year. Includes 39 tasks, 4 occupations, HLMI, and FAOL.
Combined year-framing and probability-framing responses using gamma mixture CDF aggregation
(matching the 2023 report methodology). A gamma CDF is fitted to each respondent's data, then
all CDFs are averaged pointwise to produce a mixture CDF from which percentiles are read.
Sorted by 50% probability year (earliest first). The HLMI row uses the dedicated direct-HLMI
blocks (hb_a_*/hb_b_*); it does not pool the final-occupation
prediction block.
| Milestone | N (fitted) | 10% Prob Year | 50% Prob Year | 90% Prob Year |
|---|---|---|---|---|
| Write Python code (e.g. quicksort) | 257 | 0 yr (2023) | 2 yr (2025) | 14 yr (2037) |
| Write high-school history essay | 255 | 0 yr (2023) | 2 yr (2025) | 13 yr (2036) |
| Play new Angry Birds levels (superhuman) | 224 | 0 yr (2023) | 2 yr (2025) | 10 yr (2033) |
| Answer Googleable factoid questions | 223 | 0 yr (2023) | 3 yr (2026) | 21 yr (2044) |
| Voice acting from text | 256 | 0 yr (2023) | 3 yr (2026) | 17 yr (2040) |
| Win World Series of Poker | 254 | 0 yr (2023) | 3 yr (2026) | 20 yr (2043) |
| Transcribe speech (noisy, accents) | 266 | 0 yr (2023) | 3 yr (2026) | 17 yr (2040) |
| Fluent translation (most languages) | 241 | 0 yr (2023) | 3 yr (2026) | 16 yr (2039) |
| Answer Googleable open-ended questions | 252 | 0 yr (2023) | 3 yr (2026) | 23 yr (2046) |
| Group unseen objects into classes | 240 | 0 yr (2023) | 4 yr (2027) | 23 yr (2046) |
| Produce song indistinguishable from artist | 254 | 0 yr (2023) | 4 yr (2027) | 26 yr (2049) |
| Answer questions with no definite answer | 260 | 0 yr (2023) | 4 yr (2027) | 35 yr (2058) |
| Beat best Starcraft 2 players | 247 | 0 yr (2023) | 4 yr (2027) | 25 yr (2048) |
| Translate speech from subtitled films | 240 | 0 yr (2023) | 5 yr (2028) | 41 yr (2064) |
| Build website with payment processing | 263 | 0 yr (2023) | 5 yr (2028) | 31 yr (2054) |
| Phone banking services | 260 | 0 yr (2023) | 5 yr (2028) | 30 yr (2053) |
| 3D model from short video | 233 | 0 yr (2023) | 5 yr (2028) | 27 yr (2050) |
| Atari novice level (20 min training) | 242 | 0 yr (2023) | 5 yr (2028) | 36 yr (2059) |
| Fine-tune open source LLM | 220 | 1 yr (2024) | 5 yr (2028) | 47 yr (2070) |
| One-shot image recognition | 260 | 0 yr (2023) | 5 yr (2028) | 34 yr (2057) |
| Compose US Top 40 song (full audio) | 261 | 0 yr (2023) | 6 yr (2029) | 46 yr (2069) |
| Outperform on all Atari games | 279 | 0 yr (2023) | 6 yr (2029) | 35 yr (2058) |
| Learn efficient sorting (no solution form) | 243 | 1 yr (2024) | 7 yr (2030) | 55 yr (2078) |
| Fold laundry (speed + quality) | 236 | 1 yr (2024) | 7 yr (2030) | 40 yr (2063) |
| Play random game as human novice (<10 min) | 226 | 1 yr (2024) | 7 yr (2030) | 38 yr (2061) |
| Translate newly discovered language (Rosetta stone) | 247 | 1 yr (2024) | 7 yr (2030) | 52 yr (2075) |
| Write NYT best-seller novel | 232 | 1 yr (2024) | 7 yr (2030) | 62 yr (2085) |
| Explain game AI moves to layman | 247 | 1 yr (2024) | 8 yr (2031) | 81 yr (2104) |
| Assemble any LEGO set | 252 | 1 yr (2024) | 8 yr (2031) | 48 yr (2071) |
| Win Putnam math competition | 234 | 1 yr (2024) | 8 yr (2031) | 71 yr (2094) |
| Beat fastest human in 5km city race (biped robot) | 281 | 1 yr (2024) | 9 yr (2032) | 44 yr (2067) |
| Beat best Go players (limited training) | 235 | 1 yr (2024) | 10 yr (2033) | 100 yr (2123) |
| Find & patch security flaw (100k+ users) | 244 | 2 yr (2025) | 10 yr (2033) | 87 yr (2110) |
| Retail Salesperson (occ.) | 788 | 1 yr (2024) | 10 yr (2033) | 75 yr (2098) |
| Discover physics equations from simulation | 271 | 1 yr (2024) | 12 yr (2035) | 123 yr (2146) |
| Truck Driver (occ.) | 784 | 2 yr (2025) | 12 yr (2035) | 58 yr (2081) |
| Replicate ML conference study | 253 | 2 yr (2025) | 12 yr (2035) | 109 yr (2132) |
| Install electrical wiring in new home | 244 | 3 yr (2026) | 17 yr (2040) | 104 yr (2127) |
| Conduct ML research & write conference paper | 269 | 2 yr (2025) | 20 yr (2043) | 246 yr (2269) |
| Prove publishable math theorems | 257 | 3 yr (2026) | 23 yr (2046) | 270 yr (2293) |
| HLMI (all human tasks) | 1757 | 4 yr (2027) | 24 yr (2047) | 174 yr (2197) |
| Solve unsolved math problem (e.g. Millennium) | 249 | 4 yr (2027) | 27 yr (2050) | 329 yr (2352) |
| Surgeon (occ.) | 788 | 7 yr (2030) | 33 yr (2056) | 310 yr (2333) |
| AI Researcher (occ.) | 786 | 6 yr (2029) | 40 yr (2063) | 1321 yr (3344) |
| Full Automation of Labor | 792 | 14 yr (2037) | 89 yr (2112) | 2641 yr (4664) |
Year values are years from survey date (2023), with calendar year in parentheses. "10% prob year" = year at which the mixture CDF reaches 10%. "90% prob year" = 90%. Dashes indicate insufficient data or an aggregate CDF that does not reach the target percentile within the 1e8-year "never/infinity" sentinel bound.
3.2 HLMI & Full Automation of Labor Timing
High-Level Machine Intelligence (HLMI)
HLMI is defined as machines that can accomplish every task better and cheaper than human workers.
| Percentile | This analysis | Published 2023 | Diff |
|---|---|---|---|
| 10% prob | 4 yr (2027) | 4 yr (2027) | +0.3 yr |
| 50% prob | 24 yr (2047) | 24 yr (2047) | +0.3 yr |
The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.
N (gamma fits): 1757 (0 failed fits)
This HLMI estimate uses only the dedicated direct-HLMI
question blocks (hb_a_* and hb_b_*). It does not
pool the separate final-occupation prediction block
(hj_*_final_pred).
Full Automation of Labor (FAOL)
All occupations fully automatable — machines carry out every task better and more cheaply than humans.
| Percentile | This analysis | Published 2023 | Diff |
|---|---|---|---|
| 50% prob | 89 yr (2112) | 93 yr (2116) | -3.8 yr |
N (gamma fits): 792 (0 failed fits)
Gap Between HLMI and FAOL
The consistent gap between HLMI and FAOL predictions reflects researchers' view that achieving human-level AI capability differs from fully automating all occupations.
Methodology Note
Gamma mixture CDF method: Following the methodology of Grace et al. (2024),
a gamma CDF is fitted to each respondent's three data points (whether year-framing or probability-framing).
All individual CDFs are then averaged pointwise to produce a mixture CDF, from which percentiles are read off.
This approach handles extrapolation naturally (pessimistic respondents contribute heavy-tailed CDFs rather
than being dropped) and treats both question framings symmetrically.
Respondents were randomly assigned to either a "year framing" (provide years for 10%/50%/90% probability)
or a "probability framing" (provide probabilities at fixed time horizons: 10, 20, and 40 years for
HLMI; 10, 20, and 50 years for FAOL).
See compare_methods.py for a comparison of this method with linear interpolation.
Corresponds to Figure 3 in Grace et al. (2024). Thin lines show individual respondent gamma CDFs (random subset of 200); thick line is the mixture (average) CDF. The dashed line marks 50% probability.
3.2.4 High-Level Machine Intelligence Timing by Experience
Insufficient demographic data: the years-in-field question (hh_howlong) was redacted or absent in this year's anonymized dataset, so the experience split cannot be computed.
3.2.5 Do Participants Agree on HLMI Timing?
"How much do you think your views on when HLMI will be achieved differ from those of the typical AI researcher?" (N=670)
| Response | Count | Percentage |
|---|---|---|
| Not much | 294 | 43.9% |
| A moderate amount | 306 | 45.7% |
| A lot | 70 | 10.4% |
3.3 Framing Effects
Respondents were randomly assigned to one of two question framings for the same underlying question. The "year framing" asked how many years until a given probability, while the "probability framing" asked for the probability at a fixed time horizon.
HLMI 50% Probability Year by Framing
| Framing | N (gamma fits) | Mixture CDF 50% Year | Calendar Year |
|---|---|---|---|
| Year framing (respondent provides years) | 890 | 17.5 yr | 2041 |
| Probability framing (respondent provides probabilities) | 867 | 34.0 yr | 2057 |
A consistent framing effect has been observed across survey waves: the year-framing tends to produce earlier (shorter-timeline) predictions than the probability-framing. This is a documented cognitive bias in probability elicitation.
3.4 Perceived Rates of Progress
Respondents were asked whether AI progress was faster in the first or second half of their career. (N=333)
Corresponds to Figure 4 in Grace et al. (2024).
| Response | Count | Percentage | Published 2023 |
|---|---|---|---|
| The second half | 198 | 59.5% | 60% |
| The first half | 57 | 17.1% | — |
| They were about the same | 78 | 23.4% | — |
The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.
How Far Along Is AI Progress? (Outside-view slider)
Respondents were shown a slider with three anchors: A = where progress was when they started working in their AI area; B = where it is now; and C = where it would need to be for AI software to have roughly human-level abilities at the tasks they study. They were asked "What fraction of the distance between where progress was when you started working in the area (A) and where it would need to be to attain human-level abilities in the area (C) have we come so far (B)?" on a 0–100 scale. This complements the year-based timing questions by anchoring an estimate to each respondent's own career window — so it is robust to disagreement about absolute calendar dates.
| Survey | N | Median | Mean | SD |
|---|---|---|---|---|
| 2023 | 319 | 30 | 37.6 | 26.0 |
IQR (2023): 18–60.
3.5 What Causes AI Progress?
Respondents estimated how much less AI progress there would have been with half as much of each input. Higher values = more important to progress. Values are percentage less progress (0-100%).
Corresponds to Figure 5 in Grace et al. (2024). Red dots are means; box shows IQR with median line.
| Factor | N | Median | Mean | IQR |
|---|---|---|---|---|
| Computing hardware cost decline | 195 | 50.0% | 56.6% | 40–75% |
| Training dataset effort | 196 | 50.0% | 53.7% | 30–80% |
| Funding | 191 | 50.0% | 46.0% | 25–70% |
| AI algorithm progress | 190 | 50.0% | 46.1% | 25–70% |
| Researcher effort | 200 | 30.0% | 35.4% | 20–50% |
3.6 Will There Be an Intelligence Explosion?
Probability Estimates (numeric, 0-100%)
| Scenario | N | Median | Mean | IQR | Published 2023 Median | Diff |
|---|---|---|---|---|---|---|
| Dramatic tech speedup within 2 years of HLMI | 298 | 20% | 28.2% | 5–50% | 20% | matches published |
| Dramatic tech speedup within 30 years of HLMI | 297 | 80% | 66.7% | 50–95% | 80% | matches published |
| Vastly superhuman AI within 2 years of HLMI | 281 | 10% | 19.7% | 1–30% | 10% | matches published |
| Vastly superhuman AI within 30 years of HLMI | 282 | 60% | 56.3% | 20–90% | — | — |
The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.
Is the Feedback Loop Argument Broadly Correct? (N=299)
"If AI does nearly all R&D, improvements in AI will accelerate progress including further AI progress. This could cause progress to become >10x faster within 5 years."
Corresponds to Figure 6 in Grace et al. (2024).
| Response | Count | Percentage |
|---|---|---|
| Quite unlikely (0-20%) | 69 | 23.1% |
| Unlikely (21-40%) | 72 | 24.1% |
| About even chance (41-60%) | 72 | 24.1% |
| Likely (61-80%) | 59 | 19.7% |
| Quite likely (81-100%) | 27 | 9.0% |
3.7 AI Capabilities in 2043
"In 2043, how likely do you think the following will be for at least some state-of-the-art AI systems?" Sorted by percentage rating Likely or Very Likely.
Corresponds to Figure 7 in Grace et al. (2024).
| Capability | N | V.Unlikely | Unlikely | Even | Likely | V.Likely | Likely+V.Likely |
|---|---|---|---|---|---|---|---|
| Find unexpected ways to achieve goals | 665 | 2% | 5% | 11% | 31% | 51% | 82% |
| Talk like an expert human on most topics | 667 | 2% | 4% | 13% | 28% | 53% | 81% |
| Frequently behave surprisingly to humans | 662 | 2% | 11% | 18% | 35% | 34% | 69% |
| Can be jailbroken for illegal commands | 659 | 6% | 12% | 22% | 31% | 29% | 60% |
| Deceive humans to achieve goals (unintended) | 649 | 10% | 22% | 23% | 27% | 18% | 45% |
| Have goals not aligned with human goals | 653 | 13% | 23% | 23% | 25% | 16% | 40% |
| Cause important real-world actions (run business, etc.) | 659 | 11% | 24% | 26% | 24% | 15% | 39% |
| Form AI-AI collaborative relationships (unintended) | 658 | 16% | 24% | 22% | 25% | 13% | 38% |
| Can be trusted to explain their actions | 665 | 10% | 24% | 33% | 23% | 10% | 33% |
| Self-improve regardless of human wishes | 661 | 15% | 24% | 28% | 23% | 10% | 33% |
| Take actions to attain power | 651 | 28% | 32% | 21% | 13% | 6% | 19% |
3.8 Will AI Explain Its Decisions? (2028)
"For typical state-of-the-art AI systems in 2028, will users be able to know the true reasons for decisions?" (N=912)
Corresponds to Figure 8 in Grace et al. (2024).
| Response | Count | Percentage |
|---|---|---|
| Very unlikely (<10%) | 237 | 26.0% |
| Unlikely (10-40%) | 321 | 35.2% |
| Even odds (40-60%) | 179 | 19.6% |
| Likely (60-90%) | 134 | 14.7% |
| Very likely (>90%) | 41 | 4.5% |
4.1 How Concerning Are Future AI Scenarios?
"How would you rate the level of concern these scenarios deserve over the next thirty years?" Sorted by percentage rating Substantial or Extreme concern.
Corresponds to Figure 9 in Grace et al. (2024).
| Scenario | N | None | A Little | Substantial | Extreme | Sub.+Ext. |
|---|---|---|---|---|---|---|
| AI makes it easy to spread false info (deepfakes) | 1345 | 2% | 13% | 34% | 52% | 86% |
| AI manipulates large-scale public opinion | 1342 | 4% | 18% | 37% | 41% | 78% |
| AI lets dangerous groups make powerful tools (bioweapons) | 1340 | 4% | 23% | 41% | 32% | 73% |
| Authoritarian rulers use AI for control | 1320 | 7% | 20% | 35% | 38% | 73% |
| AI worsens economic inequality | 1339 | 6% | 23% | 38% | 33% | 71% |
| AI bias worsens unjust situations (hiring, etc.) | 1344 | 8% | 31% | 38% | 23% | 61% |
| Automation leaves most people economically powerless | 1335 | 17% | 37% | 32% | 15% | 46% |
| Less human interaction (more time with AI) | 1340 | 17% | 38% | 30% | 15% | 45% |
| Wrong goals: AI reduces human decision-making role | 1333 | 17% | 38% | 31% | 13% | 44% |
| Misaligned powerful AI causes catastrophe (weapons) | 1341 | 20% | 38% | 24% | 18% | 43% |
| Automation makes people struggle to find meaning | 1336 | 26% | 39% | 25% | 11% | 35% |
4.2 How Good or Bad Will HLMI Be?
Respondents assigned probabilities to five outcome categories (summing to 100%). N=2704
Corresponds to Figure 11 in Grace et al. (2024), "Thousands of AI Authors on the Future of AI."
Mean and Median Probabilities
| Outcome | Mean | Median |
|---|---|---|
| Extremely good | 22.6% | 10.0% |
| On balance good | 29.1% | 25.0% |
| Neutral | 21.4% | 20.0% |
| On balance bad | 17.9% | 15.0% |
| Extremely bad | 9.0% | 5.0% |
Individual Respondent Views
Each vertical slice below is one respondent's probability allocation, sorted from most optimistic (left) to most pessimistic (right). The chart reveals the full diversity of expert opinion.
Corresponds to Figure 10 in Grace et al. (2024).
The same respondents sorted by how much probability they assign to "extremely bad (e.g. human extinction)" outcomes. The growing black band on the right shows those assigning the highest probability to extremely bad outcomes.
No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.
No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.
Individual Response Profiles
Each small bar below represents one respondent's five-category probability distribution, sampled evenly from the most optimistic to the most pessimistic.
No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.
Extreme Outcome Analysis
| Metric | This analysis | Published 2023 | Diff |
|---|---|---|---|
| Non-zero to both extremes | 63.5% | 64.0% | -0.5 pp |
| >=5% on extremely bad | 57.8% | 57.8% | matches published |
| >=10% on extremely bad | 37.8% | 37.8% | matches published |
| >=20% on extremely bad | 17.0% | — | — |
| >=25% on extremely bad | 9.8% | — | — |
| Mean extremely bad | 9.0% | 9.0% | matches published |
| Median extremely bad | 5.0% | 5.0% | matches published |
| Net optimists | 68.3% | 68.3% | matches published |
The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.
4.3 How Likely Is AI to Cause Extinction or Severe Disempowerment?
Each respondent was randomly assigned ONE of three questions about the probability that future AI advances cause human extinction or similarly permanent and severe disempowerment. The chart below also includes "extremely bad" HLMI outcome probabilities from Section 4.2 for comparison.
Corresponds to Figure 13 in Grace et al. (2024).
Note: The "Extremely bad HLMI outcome" question was asked to all respondents (N is larger), while each extinction/disempowerment question was asked to a random subset of respondents.
Corresponds to Figure 12 in Grace et al. (2024).
| Question | N | Mean | Published 2023 Mean | Diff | Median | Published 2023 Median | Diff | >=10% | >=25% |
|---|---|---|---|---|---|---|---|---|---|
| Future AI causes extinction or severe disempowerment | 1321 | 16.2% | 16.2% | matches published | 5.0% | 5.0% | matches published | 47.1% | 22.5% |
| Inability to control advanced AI causes extinction/disempowerment | 661 | 19.4% | 19.4% | matches published | 10.0% | 10.0% | matches published | 51.4% | 26.5% |
| AI-caused extinction/disempowerment within 100 years | 655 | 14.4% | 14.4% | matches published | 5.0% | 5.0% | matches published | 41.2% | 19.2% |
| Pooled (all variants combined) | 2637 | 16.5% | — | — | 5.0% | — | — | 46.7% | 22.7% |
The "Pooled" row combines the raw responses from all variants above (1321 + 661 + 655 = 2637 responses) and takes the median of the combined data. Because each respondent was randomly assigned exactly one variant, no respondent is counted twice. And because the other variants each add a constraint (a specific cause, a time limit) to the basic question, every response is a lower bound on that respondent's unconstrained probability — so the pooled median is a conservative basis for "the median researcher put at least 5%..." claims.
The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.
4.4 Are Future AI-Risk Concerns Due to Misunderstandings of AI Research?
"To what extent do you think people's concerns about future risks from AI are due to misunderstandings of AI research?" (N=671)
No direct figure equivalent in Grace et al. (2024); this data is discussed in Section 4.4 of that paper.
| Response | Count | Percentage |
|---|---|---|
| Hardly at all | 11 | 1.6% |
| Not much | 98 | 14.6% |
| Somewhat | 195 | 29.1% |
| To a large extent | 295 | 44.0% |
| Almost entirely | 72 | 10.7% |
4.5 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity
Question wording: "What rate of global AI progress over the next five years would make you feel most optimistic for humanity's future? Assume any change in speed affects all projects equally." (N=646)
Corresponds to Table 3 in Grace et al. (2024); no figure equivalent in that paper.
| Response | Count | Percentage | Published 2023 | Diff |
|---|---|---|---|---|
| Much slower | 31 | 4.8% | 4.8% | matches published |
| Somewhat slower | 193 | 29.9% | 29.9% | matches published |
| Current speed | 174 | 26.9% | 26.9% | matches published |
| Somewhat faster | 147 | 22.8% | 22.8% | matches published |
| Much faster | 101 | 15.6% | 15.6% | matches published |
| Other | 0 | 0.0% | — | — |
The comparison column shows the published value for this same 2023 survey (Grace et al. 2024); any difference reflects cleaning and aggregation choices in this re-analysis, not a change in expert opinion.
4.6 How Much Should AI Safety Research Be Prioritized?
"How much should society prioritize AI safety research, relative to how much it is currently prioritized?" (N=668)
Corresponds to Figure 14 in Grace et al. (2024).
| Response | Count | Percentage |
|---|---|---|
| Much less | 16 | 2.4% |
| Less | 33 | 4.9% |
| About the same | 148 | 22.2% |
| More | 256 | 38.3% |
| Much more | 215 | 32.2% |
4.7 The Alignment Problem
Corresponds to Figure 15 in Grace et al. (2024).
Do you think this argument points at an important problem? (N=1322)
| Response | Count | Percentage |
|---|---|---|
| No, not a real problem. | 67 | 5.1% |
| No, not an important problem. | 121 | 9.2% |
| Yes, a moderately important problem. | 425 | 32.1% |
| Yes, a very important problem. | 539 | 40.8% |
| Yes, among the most important problems in the field. | 170 | 12.9% |
How valuable is it to work on this problem today, compared to other problems in AI? (N=1322)
| Response | Count | Percentage |
|---|---|---|
| Much less valuable | 111 | 8.4% |
| Less valuable | 293 | 22.2% |
| As valuable as other problems | 512 | 38.7% |
| More valuable | 294 | 22.2% |
| Much more valuable | 112 | 8.5% |
How hard do you think this problem is compared to other problems in AI? (N=1318)
| Response | Count | Percentage |
|---|---|---|
| Much easier | 34 | 2.6% |
| Easier | 137 | 10.4% |
| As hard as other problems | 396 | 30.0% |
| Harder | 470 | 35.7% |
| Much harder | 281 | 21.3% |
5. Results by Value-Outlook Cluster
5.0 Methodology: Identifying Outlook Groups
One of the survey's core questions asks respondents to distribute 100 probability points across five possible outcomes of high-level machine intelligence: extremely good, on balance good, more or less neutral, on balance bad, and extremely bad. These five numbers form a compact signature of each respondent's overall outlook on advanced AI.
We apply a Gaussian Mixture Model (GMM) with K=4 components to these five-dimensional signatures (standardized to zero mean and unit variance). GMM is a soft-clustering method that models the data as a mixture of multivariate normal distributions; each respondent is assigned to the component with the highest posterior probability. We use 20 random initializations to avoid local optima.
The four resulting clusters are then named by inspecting each cluster's mean probability profile:
- Strong Optimists — assign high probability to extremely good outcomes and very little to bad outcomes.
- Mild Optimists — lean positive overall (combined good > 55%) but with more hedging and moderate expectations.
- Concerned — assign elevated probability to bad or extremely bad outcomes (combined pessimism > 30%).
- Polarized/Bimodal — the most distinctive group: they assign substantial probability to both extremely good and extremely bad outcomes, reflecting a worldview where advanced AI is seen as a high-stakes gamble rather than a clearly positive or negative development.
This four-way split captures meaningful variation in worldview that cuts across traditional demographic variables like experience level. 2,704 of 3,270 cleaned analysis rows had complete value-outlook data and were assigned to a cluster. The remaining subsections re-examine key survey results through this lens.
The bimodal outlook is not merely an artifact of the mixture model. Among the 2,704 respondents who answered both endpoints of the value question, 790 (29%) assigned at least 10% probability to both an extremely good and an extremely bad outcome, and 109 (4%) assigned at least 25% to both. This model-free count confirms that a substantial minority genuinely hedges across both extremes rather than settling on a single direction.
5.0.1 Cluster Profiles
| Cluster | N | % of total | Ext. good | Good | Neutral | Bad | Ext. bad |
|---|---|---|---|---|---|---|---|
| Strong Optimists | 436 | 16% | 38% | 36% | 12% | 8% | 5% |
| Mild Optimists | 832 | 31% | 29% | 36% | 23% | 12% | 0% |
| Alarmed | 493 | 18% | 22% | 12% | 9% | 30% | 28% |
| Uncertain/Neutral | 943 | 35% | 11% | 29% | 31% | 21% | 9% |
5.1 HLMI Timeline Predictions
How soon do different outlook groups expect human-level machine intelligence?
| Cluster | N | Median HLMI year | IQR |
|---|---|---|---|
| Strong Optimists | 154 | 2038 | 2033–2053 |
| Mild Optimists | 266 | 2043 | 2033–2063 |
| Alarmed | 135 | 2043 | 2033–2073 |
| Uncertain/Neutral | 320 | 2043 | 2033–2063 |
5.2 Extinction/Disempowerment Estimates
Probability that future AI causes human extinction or permanent severe disempowerment, broken down by outlook cluster.
| Cluster | N | Mean | Median | ≥10% | ≥25% |
|---|---|---|---|---|---|
| Strong Optimists | 206 | 9.6% | 5% | 33% | 13% |
| Mild Optimists | 379 | 9.8% | 1% | 32% | 12% |
| Alarmed | 271 | 31.1% | 20% | 72% | 48% |
| Uncertain/Neutral | 465 | 15.6% | 10% | 51% | 20% |
5.3 Concerning Scenarios
Mean concern level (0 = no concern, 3 = extreme concern) for each of 11 AI risk scenarios. The table below highlights which scenarios show the largest divergence across clusters.
| Scenario | Highest cluster | Score | Lowest cluster | Score | Gap |
|---|---|---|---|---|---|
| Misaligned powerful AI causes catastrophe (weapons) | Alarmed | 1.81 | Mild Optimists | 1.12 | 0.69 |
| Wrong goals: AI reduces human decision-making role | Alarmed | 1.80 | Mild Optimists | 1.14 | 0.66 |
| Automation leaves most people economically powerless | Alarmed | 1.83 | Mild Optimists | 1.22 | 0.61 |
| AI lets dangerous groups make powerful tools (bioweapons) | Alarmed | 2.25 | Mild Optimists | 1.79 | 0.46 |
| Automation makes people struggle to find meaning | Alarmed | 1.46 | Mild Optimists | 1.05 | 0.42 |
| Less human interaction (more time with AI) | Alarmed | 1.69 | Strong Optimists | 1.30 | 0.39 |
| AI worsens economic inequality | Alarmed | 2.18 | Strong Optimists | 1.83 | 0.36 |
| Authoritarian rulers use AI for control | Alarmed | 2.22 | Mild Optimists | 1.86 | 0.36 |
| AI manipulates large-scale public opinion | Alarmed | 2.30 | Mild Optimists | 2.00 | 0.29 |
| AI makes it easy to spread false info (deepfakes) | Uncertain/Neutral | 2.52 | Mild Optimists | 2.24 | 0.27 |
| AI bias worsens unjust situations (hiring, etc.) | Uncertain/Neutral | 1.90 | Mild Optimists | 1.64 | 0.26 |
5.4 Expected AI Capabilities in 2043
Mean rated likelihood (0 = very unlikely, 4 = very likely) that AI systems will exhibit each capability by 2043.
5.5 Intelligence Explosion
Median probability estimates for dramatic AI capability speedup and superhuman AI emergence, by outlook cluster.
| Cluster | Dramatic speedup within 2yr (median %) | Dramatic speedup within 30yr (median %) | Superhuman within 2yr (median %) | N |
|---|---|---|---|---|
| Strong Optimists | 20% | 90% | 15% | 47 |
| Mild Optimists | 15% | 80% | 5% | 75 |
| Alarmed | 20% | 90% | 10% | 54 |
| Uncertain/Neutral | 10% | 70% | 10% | 118 |
5.6 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity
What rate of global AI progress over the next five years would make each cluster feel most optimistic for humanity's future?
5.7 Safety and Alignment-Problem Views
Mean scores on whether the alignment argument points at an important problem, how much society should prioritize AI safety research, and how valuable it is to work on the alignment problem today compared with other AI problems (all 0–4 scales).
| Cluster | Alignment-problem importance (0–4) | Safety research priority (0–4) | Value of working on alignment problem today vs other AI problems (0–4) | N |
|---|---|---|---|---|
| Strong Optimists | 2.38 | 2.82 | 1.88 | 224 |
| Mild Optimists | 2.23 | 2.72 | 1.82 | 386 |
| Alarmed | 2.77 | 3.32 | 2.27 | 259 |
| Uncertain/Neutral | 2.55 | 3.02 | 2.06 | 453 |
5.8 Counterfactual Progress Reduction if Factors Were Halved
Median estimated progress reduction if each factor were halved, by cluster.
Cluster assignments are based on GMM (K=4, 20 random initializations) applied to the five HLMI value-outcome probabilities (vb_1_1–vb_1_5). Respondents missing all five values are excluded (2,704 of 3,270 cleaned analysis rows assigned). Sample sizes vary across subsections due to block randomization.
Appendix: Cluster Analysis
A data-driven exploration of the natural groupings, the safety divide, and hardware vs. software beliefs among 3,270 ESPAI 2023 cleaned analysis rows.
Executive Summary
1. How Many Camps? The Natural Clusters
1.1 Clustering on Value Outlook (n=1,538)
The value-of-HLMI question asked respondents to assign probabilities (summing to 100%) across five outcomes: extremely good, on balance good, neutral, on balance bad, and extremely bad. It is the highest-coverage feature in this analysis. We cluster on these 5 dimensions using Gaussian Mixture Models.
Figure 1: Silhouette scores for K=2 through K=6 clusters on value outlook. K=2 has the highest silhouette, but K=3 and K=4 offer more interpretable structure.
Silhouette scores are modest (0.10-0.12), indicating the clusters are not sharply separated -- this is a continuous landscape of opinion, not discrete tribes. Still, the structure is meaningful.
1.2 Four-Camp Solution
Figure 2: Four natural camps in value outlook. Left: PCA projection. Center: mean probability profiles. Right: cluster sizes.
The most striking finding here is the Polarized/Bimodal group. These respondents assign high probability to both extremely good and extremely bad outcomes -- their P(extremely good) is comparable to the Strong Optimists, yet they simultaneously assign substantial probability to catastrophic outcomes. This is not a group of doomers or technophobes. They are researchers who believe AI will be hugely impactful, but who are genuinely torn on whether that impact will be positive or negative. The key axis for this group is magnitude of impact, not direction -- they have rejected the possibility that AI will be a modest or neutral development.
1.3 Three-Camp Solution (Simpler View)
Collapsing to 3 clusters (which has a higher silhouette score of 0.18 vs 0.07) merges the finer distinctions into a simpler optimist/moderate/pessimist framing. This loses the polarized group but provides a cleaner summary:
Figure 3: Three-camp simplification. The polarized group is absorbed into the moderate/pessimist clusters.
1.4 Intuitive View: How Soon vs. How Good
The PCA axes above lack intuitive meaning. Below, we plot each respondent on two directly interpretable dimensions: their predicted HLMI arrival year (x-axis) and their net optimism score (y-axis), defined as P(good + extremely good) minus P(bad + extremely bad). This reveals where the four camps sit in the space of "how soon" vs. "how beneficial."
Figure 3b: Respondents (n=1663) binned by predicted HLMI year and net optimism. Each bubble's size shows how many researchers from that cluster fall in the bin (labeled when ≥5). The vertical dashed line marks the median predicted year; the horizontal line separates net optimists from net pessimists.
1.5 Combined Clustering: Values + Concerns + Safety + Extinction
For the 323 respondents who answered all of: value outlook, concern scenarios, alignment-problem importance, and extinction/disempowerment probability, we ran a richer clustering incorporating 8 features.
Figure 4: Combined clustering incorporating worldview, concerns, safety/alignment-problem views, and extinction/disempowerment probability. Mean concern averages 11 scenario concern ratings on a 0-3 scale (0=no concern, 3=extreme concern); alignment-problem importance is a 0-4 ordinal score (0=not a real problem, 4=among the most important problems in the field); P(extinction/disempowerment) is a percentage.
| Cluster | Size | P(Good / Ext good) | P(Bad / Ext bad) | Mean concern | Alignment-problem imp | P(extinction/disempowerment) | HLMI Year median [IQR]; mean |
|---|---|---|---|---|---|---|---|
| Optimistic / High x-risk | 72 (22%) | 13% / 46% | 11% / 4% | 1.8/3 | 2.4/4 | 21% | 2053 [2043-2073]; mean 2080 (n=53) |
| Optimistic / Low x-risk | 60 (19%) | 36% / 32% | 13% / 8% | 1.8/3 | 2.5/4 | 5% | 2043 [2033-2073]; mean 2068 (n=43) |
| Optimistic / Low x-risk | 96 (30%) | 26% / 37% | 14% / 4% | 1.6/3 | 2.2/4 | 4% | 2043 [2033-2072]; mean 2112 (n=58) |
| Pessimistic / Safety-focused / High x-risk / Broadly worried | 42 (13%) | 22% / 13% | 19% / 36% | 2.1/3 | 3.1/4 | 56% | 2046 [2038-2063]; mean 2060 (n=21) |
| Pessimistic | 38 (12%) | 6% / 13% | 50% / 11% | 2.0/3 | 2.0/4 | 13% | 2063 [2047-2085]; mean 45562 (n=23) |
| Moderate / Low x-risk | 15 (5%) | 2% / 15% | 15% / 2% | 1.6/3 | 2.4/4 | 4% | 2063 [2040-2296]; mean 2186 (n=8) |
HLMI year summaries include the median, interquartile range, mean, and item-level n because several groups share the same median year while their wider distributions differ.
2. Higher vs Lower Safety/Severe-Risk Concern
2.1 Defining the Groups
We built a safety composite score from (normalized 0-1 and averaged):
- Alignment-problem importance rating (0-4 ordinal scale: not a real problem to among the most important problems in the field)
- Value of working on the alignment problem today compared with other AI problems (0-4 ordinal scale)
- P(extinction/disempowerment from AI) (0-100%)
Requiring at least 2 of 3 components, we obtained scores for 1322 respondents and split at the median (0.500) into High (n=595) and Low (n=727) safety concern groups.
2.2 The Full Comparison
Figure 5: Comprehensive comparison of higher vs lower safety/severe-risk-concern respondents across six dimensions. Alignment-problem importance uses the 0-4 ordinal scale above; mean concern averages 11 scenario concern ratings on a 0-3 scale; the optimism-maximizing AI progress-rate answer is coded 0=much slower, 2=current speed, 4=much faster.
2.3 Statistical Tests
| Figure 5 panel | Tested variable | n (High) | n (Low) | Median (High) | Median (Low) | Mean (High) | Mean (Low) | p-value | Effect r |
|---|---|---|---|---|---|---|---|---|---|
| HLMI timeline | HLMI Year | 357 | 453 | 2043.0 | 2053.0 | 2229.0 | 4529.2 | 0.0074 ** | 0.109 |
| Extinction/disempowerment estimate | P(extinction/disempowerment) | 237 | 406 | 25.0 | 2.0 | 32.5 | 7.9 | 0.0000 *** | -0.586 |
| Value outlook | P(Extremely bad) | 595 | 727 | 7.0 | 5.0 | 12.4 | 6.7 | 0.0000 *** | -0.223 |
| Value outlook | P(Extremely good) | 595 | 727 | 10.0 | 10.0 | 20.4 | 23.3 | 0.0724 n.s. | 0.057 |
| Concern levels | Mean concern | 274 | 381 | 1.9 | 1.7 | 1.9 | 1.7 | 0.0000 *** | -0.296 |
| AI progress rate for optimism | Optimism-maximizing AI progress rate | 144 | 204 | 2.0 | 2.0 | 1.9 | 2.2 | 0.0128 * | 0.152 |
Selected Mann-Whitney U tests for scalar summaries from Figure 5. The value-outlook panel is summarized by P(extremely good) and P(extremely bad), the concern panel by mean concern, and the AI-capabilities panel is tested item-by-item in the next table. Effect size r is rank-biserial correlation (|r| > 0.3 = medium, |r| > 0.5 = large). Scale notes: mean concern 0-3, optimism-maximizing AI progress-rate answer 0-4, probabilities in percentage points. *** p<0.001, ** p<0.01, * p<0.05
2.4 AI Capabilities by 2043: What Do Safety People Expect?
Safety-concerned people don't just worry more -- they have systematically different expectations for what AI will be able to do by 2043:
| Capability | Hypothesized direction | Mean (High Safety) | Mean (Low Safety) | High - Low | Observed direction | p-value | n |
|---|---|---|---|---|---|---|---|
| Talk like expert | Exploratory | 3.34 | 3.11 | +0.23 | High > Low | 0.1252 n.s. | 322 |
| Self-improve regardless | Exploratory | 2.08 | 1.74 | +0.34 | High > Low | 0.0192 * | 317 |
| AI-AI collaborations | High > Low | 2.13 | 1.86 | +0.27 | High > Low | 0.0719 n.s. | 313 |
| Deceive humans | High > Low | 2.42 | 1.98 | +0.44 | High > Low | 0.0021 ** | 313 |
| Unexpected strategies | Exploratory | 3.43 | 3.01 | +0.43 | High > Low | 0.0000 *** | 322 |
| Seek power | High > Low | 1.56 | 1.25 | +0.31 | High > Low | 0.0219 * | 314 |
| Explain actions (trustworthy) | Exploratory | 1.90 | 2.02 | -0.11 | High < Low | 0.3351 n.s. | 322 |
| Can be jailbroken | Exploratory | 2.84 | 2.49 | +0.36 | High > Low | 0.0097 ** | 318 |
| Surprising behavior | Exploratory | 3.10 | 2.69 | +0.41 | High > Low | 0.0002 *** | 321 |
| Real-world actions | Exploratory | 2.43 | 1.85 | +0.57 | High > Low | 0.0001 *** | 317 |
| Misaligned goals | High > Low | 2.34 | 1.87 | +0.47 | High > Low | 0.0011 ** | 316 |
Scale: 0=Very unlikely, 1=Unlikely, 2=Even chance, 3=Likely, 4=Very likely. High - Low is the high-safety mean minus the low-safety mean, so positive values mean high-safety respondents rated the capability as more likely. Hypothesized direction is marked for risk-relevant capabilities where we expected High > Low; other rows are exploratory.
2.5 Intelligence Explosion Beliefs
Figure 6: Probability estimates for intelligence explosion scenarios, by safety concern level.
3. Hardware vs. Software Progress Beliefs
3.1 Counterfactual Progress Reduction if Factors Were Halved
Respondents rated how much AI progress would decrease if each of 5 factors were cut in half (0-100% scale). Higher values mean the factor is more important.
Figure 7: Counterfactual progress-factor analysis. Top-left: estimated progress reduction distributions. Top-right: hardware vs algorithm progress-reduction estimates. Bottom-left: HW-SW index distribution. Bottom-right: factor correlations with other variables.
- Computing hardware: 50% (n=195) -- largest median estimated progress reduction if halved
- Training data: 50% (n=196)
- Algorithm progress: 50% (n=190)
- Funding: 50% (n=191)
- Researcher effort: 30% (n=200) -- surprisingly lowest
3.2 Hardware vs Software Index
We computed (Hardware - Algorithms) / (Hardware + Algorithms) for each respondent who rated both. Positive = hardware-leaning, negative = software-leaning.
- Hardware-leaning: 59 (64%)
- Software-leaning: 25 (27%)
- Balanced: 8
- Median index: 0.19 (slight hardware lean)
3.3 Do Progress Beliefs Predict Other Views?
Significant cause-factor correlations (p < 0.10)
| Cause Factor | Target | Spearman rho | p-value | n |
|---|---|---|---|---|
| Researcher effort | Alignment-problem imp | 0.296 | 0.0044 ** | 91 |
| Training data | HLMI Year | -0.155 | 0.0848 n.s. | 124 |
| Funding | Alignment-problem imp | 0.223 | 0.0275 * | 98 |
Hardware-Software Index Correlations
| Variable 1 | Variable 2 | Spearman rho | p-value | n |
|---|---|---|---|---|
| HW-vs-SW Index | P(extinction/disempowerment) | 0.151 | 0.3601 n.s. | 39 |
| HW-vs-SW Index | HLMI Year | 0.276 | 0.0430 * | 54 |
| HW-vs-SW Index | P(Ext bad) | 0.176 | 0.0939 n.s. | 92 |
| HW-vs-SW Index | Alignment-problem imp | -0.128 | 0.3871 n.s. | 48 |
| HW-vs-SW Index | Mean concern | 0.189 | 0.2368 n.s. | 41 |
| HW-vs-SW Index | P(Ext good) | -0.154 | 0.1423 n.s. | 92 |
3.4 Progress Beliefs by Outlook and Safety Groups
Do optimists, pessimists, and safety-concerned researchers differ in what they think drives AI progress? Below we break down cause factor importance by outlook cluster (left) and safety group (right).
Figure 7b: Mean importance ratings for each progress factor, split by outlook group (left) and safety group (right). Scale: 0-100% estimated decrease in progress if factor halved.
4. How Outlook Predicts Everything Else
The three value-outlook clusters (from Section 1) predict views across many other dimensions:
| Variable | Mild Optimists | Uncertain/Neutral | Strong Optimists | Alarmed |
|---|---|---|---|---|
| HLMI Year | 2043 [2033-2073]; mean 2393 (n=515) | 2048 [2033-2073]; mean 2183 (n=595) | 2043 [2033-2063]; mean 2080 (n=280) | 2053 [2038-2073]; mean 6157 (n=273) |
| P(extinction/disempowerment) | 1.0 (n=379) | 10.0 (n=465) | 5.0 (n=206) | 20.0 (n=271) |
| Alignment-problem importance | 2.0 (n=386) | 3.0 (n=453) | 2.0 (n=224) | 3.0 (n=259) |
| Mean concern | 1.6 (n=435) | 1.8 (n=444) | 1.6 (n=237) | 2.0 (n=229) |
| Optimism-maximizing AI progress rate | 2.0 (n=188) | 2.0 (n=234) | 3.0 (n=93) | 1.0 (n=131) |
Values are medians with sample sizes in parentheses, except HLMI Year, which shows median [IQR], mean, and item-level n because several groups share the same median year. Scale notes: alignment-problem importance 0-4, mean concern 0-3, optimism-maximizing AI progress-rate answer 0=much slower to 4=much faster, probabilities in percentage points.
Figure 8: How the four value-outlook camps compare across timelines, extinction/disempowerment probability, safety/alignment-problem views, concern levels, optimism-maximizing AI progress-rate answers, and P(extremely bad). Mean concern averages 11 scenario concern ratings on a 0-3 scale; alignment-problem importance and optimism-maximizing AI progress-rate answers use 0-4 ordinal scales.
5. What Dimensions Structure AI Researcher Beliefs?
Principal Component Analysis on respondents with values + concerns + safety + extinction/disempowerment data (n=301) reveals the latent axes of disagreement.
Figure 9: PCA factor loadings. Green bars = positive loading, red = negative. Each panel shows one principal component.
6. The Full Correlation Structure
Figure 10: Spearman rank correlations between key variables. Stars indicate significance. Sample sizes shown in each cell. Alignment-problem importance uses a 0-4 ordinal scale from not a real problem to among the field's most important problems; value of working on the alignment problem today uses a 0-4 ordinal scale from much less valuable to much more valuable than other AI problems; mean concern uses a 0-3 scale; HLMI Year is a calendar-year estimate; probabilities are 0-100 percentages.
Key Pairwise Correlations
| Relationship | Spearman rho | p-value | n |
|---|---|---|---|
| P(extinction/disempowerment) x HLMI Year | -0.137 | 0.0001 *** | 809 |
| P(extinction/disempowerment) x P(Ext bad) | 0.469 | 0.0000 *** | 1321 |
| Alignment-problem imp x P(extinction/disempowerment) | 0.312 | 0.0000 *** | 643 |
| Alignment-problem imp x HLMI Year | -0.090 | 0.0102 * | 810 |
| Alignment-problem imp x Mean concern | 0.301 | 0.0000 *** | 655 |
| P(Ext good) x P(Ext bad) | -0.010 | 0.6123 n.s. | 2704 |
| HLMI Year x P(Ext bad) | -0.023 | 0.3468 n.s. | 1663 |
| HLMI Year x P(Ext good) | -0.072 | 0.0034 ** | 1663 |
| Mean concern x P(extinction/disempowerment) | 0.378 | 0.0000 *** | 652 |
| Mean concern x HLMI Year | -0.109 | 0.0015 ** | 843 |
| Optimism-max progress rate x P(extinction/disempowerment) | -0.189 | 0.0004 *** | 354 |
| Optimism-max progress rate x HLMI Year | 0.008 | 0.8744 n.s. | 383 |
| Optimism-max progress rate x alignment-problem imp | -0.168 | 0.0016 ** | 348 |
- Alignment-problem importance and value of working on the alignment problem today are highly correlated (these aren't independent beliefs -- they form a coherent safety/alignment-problem view)
- P(extinction/disempowerment) and P(Extremely bad) are strongly linked -- people who give high extinction/disempowerment probability estimates also see HLMI as likely to be extremely bad
- P(extinction/disempowerment) and HLMI timeline are negatively correlated -- people with the highest extinction/disempowerment estimates tend to predict earlier arrival of HLMI, not later
- Optimism-maximizing AI progress rate and HLMI timeline are negatively correlated -- people with shorter timelines more often choose slower progress as the rate that would make them most optimistic for humanity's future
7. Methods & Caveats
7.1 Approach
- Strategy: Multiple focused analyses on overlapping subsets (rather than one clustering requiring all features, which drops to n~180). This maximizes statistical power while giving complementary views.
- Clustering: Gaussian Mixture Models (GMM) with full covariance, 10-20 random initializations, model selection by BIC with parsimony preference
- Statistical tests: Non-parametric throughout (Mann-Whitney U for two groups, Kruskal-Wallis for three+, Spearman rank correlations)
- Visualization: PCA for dimensionality reduction (visualization only -- clustering done in full feature space)
7.2 Caveats
7.3 Technical Details
- Random seed: 42
- All tests two-sided; no multiple comparison correction (exploratory analysis)
- HLMI year capped at 2150 for visualization; uncapped for statistics where noted
Appendix: Survey Flow
Survey-flow / randomization diagram modeled on Figure 16 of Grace et al., Thousands of AI Authors on the Future of AI. Each box is a question block. Percentages and n's are realized fill rates: the realized number of respondents who answered each block, divided by the total N. The total N counts everyone who answered at least one question (respondents who answered none are excluded).
Jobs / FAOL Sample-Size Reconciliation
The survey-flow boxes count anyone with any response in the broader Jobs / FAOL block. FAOL analyses use only the FAOL triplet, so their sample sizes can be slightly smaller.
| Count definition | Fixed-years | Fixed-probabilities |
|---|---|---|
| Whole Jobs / FAOL block (survey-flow box) | 384 | 426 |
| FAOL only, any triplet answer | 380 | 418 |
| FAOL only, complete triplet | 379 | 413 |
Here, "fixed-years" means respondents gave probabilities for fixed
time horizons (hj_b_full_*), while "fixed-probabilities" means respondents
gave years for fixed probabilities (hj_a_full_*).
Tasks: ta_* columns are the fixed-probability framing,
tb_* the fixed-years framing (verified empirically:
ta_* rows have fixedprobabilities=1,
tb_* rows have fixedprobabilities=0).
The Qualtrics survey-flow definition (.qsf) is not in
this repo, so the exact display order may differ from the diagram.
Appendix: Supplementary Figures
B.1 fixed-prob/fixed-year CDFs
Full CDF comparison between the fixed-prob and fixed-year conditions for HLMI and FAOL. The fixed-prob framing consistently produces earlier predictions (the red curve is shifted left relative to blue).
Corresponds to Figure 18 in Grace et al. (2024). Each curve is the mixture (mean) CDF for respondents assigned to that framing condition.
B.2 Bootstrap Confidence Bands
95% bootstrap confidence intervals on the aggregate CDF, obtained by resampling the fitted individual CDFs with replacement (500 resamples).
| Milestone | 50% Year (point est.) | 95% Bootstrap CI |
|---|---|---|
| HLMI | 2047 (24.3 yr) | 2046–2049 (23.0–25.6 yr) |
| FAOL | 2112 (89.2 yr) | 2105–2121 (81.8–98.4 yr) |
Bootstrap CIs reflect sampling variability — if we surveyed a different random sample of AI researchers, how much would the aggregate change? This is distinct from the spread of individual predictions (which is much wider).
Appendix: Statistical Tests
F.1 Yuen's Trimmed-Mean Bootstrap Test: Demographics
Tests whether experienced researchers (split at the median years in field) give significantly different HLMI predictions than junior researchers. Uses a 10% trimmed mean with 5,000 bootstrap resamples.
Insufficient data for demographic comparison (years-in-field was redacted in the anonymized dataset).
N (experienced): 0, N (junior): 0. Individual 50% years are computed from per-respondent gamma fits. The trimmed mean reduces sensitivity to extreme predictions.
F.2 Framing Effect: Statistical Significance
Tests whether year-framing and probability-framing respondents give significantly different predictions, using the same Yuen's bootstrap method.
| Milestone | Year-framing median | Prob-framing median | Trimmed-mean diff | 95% CI | p-value |
|---|---|---|---|---|---|
| HLMI | 19.8 yr (n=890) | 36.4 yr (n=867) | -26.7 yr | [-34.7, -20.6] | 0.000 |
| FAOL | 74.2 yr (n=413) | 130.5 yr (n=379) | -672.3 yr | [-815.8, -533.0] | 0.000 |
A negative difference (yr_50 − pr_50) means the year-framing produces earlier predictions. Medians shown for reference; the statistical test uses trimmed means.
Appendix: Aggregation Method Comparison
Different choices in how individual survey responses are fitted and combined into aggregate forecasts can shift the headline numbers. This appendix compares three approaches on the key milestones, motivated by Adamczewski's reanalysis of ESPAI data.
Methods compared
Mix-Mean (MSE) — the Grace et al. 2023 method. Fit a gamma CDF to each respondent's 3 data points using mean squared error, then take the pointwise mean of all fitted CDFs.
Mix-Median (MSE) — same fitting, but take the pointwise median instead of mean. More robust to outlier predictions (e.g. respondents predicting 100M years). Adamczewski argues the median better represents the "typical expert."
Mix-Mean (LogLoss) — fit using log-loss (cross-entropy) instead of MSE, then take the mean. Log-loss naturally weighs errors near p=0 and p=1 more heavily than errors near p=0.5, matching the intuition that a 5% error at p=0.95 represents a much larger shift in belief than the same error at p=0.50.
Results
| Milestone | Mix-Mean (MSE) | Mix-Median (MSE) | Mix-Mean (LogLoss) | N (MSE) | N (LogLoss) |
|---|---|---|---|---|---|
| HLMI | 2047 (24.3 yr) | 2049 (25.5 yr) | 2048 (24.7 yr) | 1757 | 1757 |
| FAOL | 2112 (89.2 yr) | 2121 (97.5 yr) | 2112 (89.0 yr) | 792 | 792 |
| Truck Driver | 2035 (11.6 yr) | 2034 (10.6 yr) | 2035 (11.6 yr) | 784 | 784 |
| Surgeon | 2056 (32.7 yr) | 2057 (34.1 yr) | 2056 (33.5 yr) | 788 | 788 |
| Retail Salesperson | 2033 (9.8 yr) | 2033 (10.0 yr) | 2033 (9.8 yr) | 788 | 788 |
| AI Researcher | 2063 (40.0 yr) | 2071 (48.2 yr) | 2063 (40.2 yr) | 786 | 786 |
Key takeaways
Aggregation method matters more than loss function. Switching from mean to median aggregation pushes HLMI from 2047 to 2049 and FAOL from 2112 to 2121. Switching the loss function (MSE→log-loss) barely changes the aggregate — consistent with Adamczewski's finding that fitting choice has "hardly any impact."
The effect is largest for far-future milestones (FAOL, AI Researcher) where a few very pessimistic respondents pull the mean CDF rightward.
See also: compare_methods.py for the full 5-method comparison
across all 39 tasks, including linear interpolation and gamma-individual methods.
Method comparison motivated by
Adamczewski (bayes.net/espai).
Data: 2023 Expert Survey on Progress in AI (ESPAI).
Previous survey comparison values from "Thousands of AI Authors on the Future of AI" (2024 preprint).