2016 ESPAI Survey Analysis
Expert Survey on Progress in AI — Comprehensive Data Report
Generated 2026-07-09 00:34
1. Overview & Survey Methodology
Randomization Blocks
| Block | Count | Percentage |
|---|---|---|
| Block 1 | 160 | 34.8% |
| Block 2 | 107 | 23.3% |
| Block 3 | 118 | 25.7% |
Data Cleaning Notes
The following cleaning steps were applied (see cleaning_log.md for full details):
- No literal year values found — all values already in years-from-survey format.
- No year values above the 1e8 cap.
- Monotonicity violations found in 2 complete triplets.
- 2 descending triplets found across 2 respondents in year and probability framings (values set to NaN).
Caveat — sum-to-100 triplets: 29 triplets across 27 respondents have values summing to exactly 100 (e.g., 10%, 30%, 60% at the three time horizons). This may indicate respondents who treated the probabilities as shares that must total 100%, rather than independent cumulative probabilities at each time horizon. Following standard methodology, sum-to-100 triplets from respondents with descending errors were also removed (2 total triplets NaN'd).
3.1 When Will 39 AI Milestones Be Feasible?
Corresponds to Figure 1 in Grace et al. (2024). Dots show the 50% probability year from the mixture CDF; horizontal lines show the 10%–90% range. The x-axis is capped at 2100; milestones whose 50% or 90% year falls beyond the cap are labeled with an arrow showing the true year. Includes 39 tasks, 4 occupations, HLMI, and FAOL.
Combined year-framing and probability-framing responses using gamma mixture CDF aggregation
(matching the 2023 report methodology). A gamma CDF is fitted to each respondent's data, then
all CDFs are averaged pointwise to produce a mixture CDF from which percentiles are read.
Sorted by 50% probability year (earliest first). The HLMI row uses the dedicated direct-HLMI
blocks (hb_a_*/hb_b_*); it does not pool the final-occupation
prediction block.
| Milestone | N (fitted) | 10% Prob Year | 50% Prob Year | 90% Prob Year |
|---|---|---|---|---|
| Play new Angry Birds levels (superhuman) | 38 | 0 yr (2016) | 3 yr (2019) | 14 yr (2030) |
| Win World Series of Poker | 37 | 0 yr (2016) | 4 yr (2020) | 19 yr (2035) |
| Fold laundry (speed + quality) | 29 | 0 yr (2016) | 5 yr (2021) | 29 yr (2045) |
| Beat best Starcraft 2 players | 24 | 1 yr (2017) | 6 yr (2022) | 64 yr (2080) |
| Learn efficient sorting (no solution form) | 43 | 0 yr (2016) | 7 yr (2023) | 31 yr (2047) |
| Atari novice level (20 min training) | 33 | 1 yr (2017) | 7 yr (2023) | 41 yr (2057) |
| Answer Googleable factoid questions | 44 | 1 yr (2017) | 7 yr (2023) | 34 yr (2050) |
| Group unseen objects into classes | 29 | 1 yr (2017) | 8 yr (2024) | 61 yr (2077) |
| Fluent translation (most languages) | 41 | 1 yr (2017) | 8 yr (2024) | 40 yr (2056) |
| Transcribe speech (noisy, accents) | 32 | 1 yr (2017) | 8 yr (2024) | 23 yr (2039) |
| Phone banking services | 29 | 1 yr (2017) | 8 yr (2024) | 33 yr (2049) |
| Write Python code (e.g. quicksort) | 36 | 1 yr (2017) | 8 yr (2024) | 39 yr (2055) |
| Assemble any LEGO set | 35 | 0 yr (2016) | 9 yr (2025) | 40 yr (2056) |
| Outperform on all Atari games | 36 | 1 yr (2017) | 9 yr (2025) | 48 yr (2064) |
| Voice acting from text | 42 | 2 yr (2018) | 9 yr (2025) | 28 yr (2044) |
| Write high-school history essay | 42 | 1 yr (2017) | 10 yr (2026) | 51 yr (2067) |
| One-shot image recognition | 30 | 1 yr (2017) | 10 yr (2026) | 42 yr (2058) |
| Answer questions with no definite answer | 45 | 2 yr (2018) | 10 yr (2026) | 56 yr (2072) |
| Answer Googleable open-ended questions | 36 | 2 yr (2018) | 10 yr (2026) | 62 yr (2078) |
| Translate speech from subtitled films | 37 | 1 yr (2017) | 10 yr (2026) | 52 yr (2068) |
| Explain game AI moves to layman | 35 | 1 yr (2017) | 11 yr (2027) | 85 yr (2101) |
| Produce song indistinguishable from artist | 40 | 2 yr (2018) | 11 yr (2027) | 55 yr (2071) |
| Truck Driver (occ.) | 89 | 2 yr (2018) | 12 yr (2028) | 40 yr (2056) |
| Compose US Top 40 song (full audio) | 36 | 1 yr (2017) | 12 yr (2028) | 50 yr (2066) |
| Beat fastest human in 5km city race (biped robot) | 28 | 1 yr (2017) | 12 yr (2028) | 43 yr (2059) |
| 3D model from short video | 40 | 2 yr (2018) | 12 yr (2028) | 45 yr (2061) |
| Play random game as human novice (<10 min) | 44 | 1 yr (2017) | 12 yr (2028) | 85 yr (2101) |
| Retail Salesperson (occ.) | 89 | 3 yr (2019) | 15 yr (2031) | 67 yr (2083) |
| Discover physics equations from simulation | 51 | 1 yr (2017) | 15 yr (2031) | 144 yr (2160) |
| Beat best Go players (limited training) | 40 | 2 yr (2018) | 16 yr (2032) | 328 yr (2344) |
| Translate newly discovered language (Rosetta stone) | 34 | 1 yr (2017) | 17 yr (2033) | 135 yr (2151) |
| Write NYT best-seller novel | 26 | 6 yr (2022) | 31 yr (2047) | 233 yr (2249) |
| Win Putnam math competition | 43 | 4 yr (2020) | 36 yr (2052) | 203 yr (2219) |
| Surgeon (occ.) | 91 | 9 yr (2025) | 37 yr (2053) | 285 yr (2301) |
| Prove publishable math theorems | 31 | 6 yr (2022) | 44 yr (2060) | 156 yr (2172) |
| HLMI (all human tasks) | 252 | 9 yr (2025) | 45 yr (2061) | 350 yr (2366) |
| AI Researcher (occ.) | 90 | 21 yr (2037) | 87 yr (2103) | 2406 yr (4422) |
| Full Automation of Labor | 92 | 20 yr (2036) | 123 yr (2139) | 3780 yr (5796) |
| Conduct ML research & write conference paper | 0 | — | — | — |
| Solve unsolved math problem (e.g. Millennium) | 0 | — | — | — |
| Install electrical wiring in new home | 0 | — | — | — |
| Replicate ML conference study | 0 | — | — | — |
| Fine-tune open source LLM | 0 | — | — | — |
| Build website with payment processing | 0 | — | — | — |
| Find & patch security flaw (100k+ users) | 0 | — | — | — |
Year values are years from survey date (2016), with calendar year in parentheses. "10% prob year" = year at which the mixture CDF reaches 10%. "90% prob year" = 90%. Dashes indicate insufficient data or an aggregate CDF that does not reach the target percentile within the 1e8-year "never/infinity" sentinel bound.
3.2 HLMI & Full Automation of Labor Timing
High-Level Machine Intelligence (HLMI)
HLMI is defined as machines that can accomplish every task better and cheaper than human workers.
| Percentile | 2016 Survey | 2023 Survey | Shift |
|---|---|---|---|
| 10% prob | 9 yr (2025) | 11 yr (2027) | -1.9 yr |
| 50% prob | 45 yr (2061) | 31 yr (2047) | +14.5 yr |
N (gamma fits): 252 (0 failed fits)
This HLMI estimate uses only the dedicated direct-HLMI
question blocks (hb_a_* and hb_b_*). It does not
pool the separate final-occupation prediction block
(hj_*_final_pred).
Full Automation of Labor (FAOL)
All occupations fully automatable — machines carry out every task better and more cheaply than humans.
| Percentile | 2016 Survey | 2023 Survey | Shift |
|---|---|---|---|
| 50% prob | 123 yr (2139) | 100 yr (2116) | +23.5 yr |
N (gamma fits): 92 (0 failed fits)
Gap Between HLMI and FAOL
The consistent gap between HLMI and FAOL predictions reflects researchers' view that achieving human-level AI capability differs from fully automating all occupations.
Methodology Note
Gamma mixture CDF method: Following the methodology of Grace et al. (2024),
a gamma CDF is fitted to each respondent's three data points (whether year-framing or probability-framing).
All individual CDFs are then averaged pointwise to produce a mixture CDF, from which percentiles are read off.
This approach handles extrapolation naturally (pessimistic respondents contribute heavy-tailed CDFs rather
than being dropped) and treats both question framings symmetrically.
Respondents were randomly assigned to either a "year framing" (provide years for 10%/50%/90% probability)
or a "probability framing" (provide probabilities at fixed time horizons: 10, 20, and 40 years for
HLMI; 10, 20, and 50 years for FAOL).
See compare_methods.py for a comparison of this method with linear interpolation.
Corresponds to Figure 3 in Grace et al. (2024). Thin lines show individual respondent gamma CDFs (random subset of 200); thick line is the mixture (average) CDF. The dashed line marks 50% probability.
3.2.4 High-Level Machine Intelligence Timing by Experience
This table splits respondents at the median years in field, then refits the direct High-Level Machine Intelligence (HLMI) timing model within each experience group.
| Experience group | Direct HLMI gamma-fit count | Mixture CDF 50% Year | Calendar Year |
|---|---|---|---|
| More experienced (>=6 yr) | 41 | 58.3 yr | 2074 |
| Less experienced (<6 yr) | 32 | 38.9 yr | 2055 |
The count values are not all respondents in each experience group. They count respondents who both answered the years-in-field question and gave enough direct HLMI timing answers to fit a gamma CDF. The separate occupation-automation prediction block is not pooled into this HLMI comparison. Only 107 respondents received the experience question due to randomization. Geographic region and citation count data are not available in the anonymized dataset.
3.2.5 Do Participants Agree on HLMI Timing?
"How much do you think your views on when HLMI will be achieved differ from those of the typical AI researcher?" (N=115)
| Response | Count | Percentage |
|---|---|---|
| Not much | 64 | 55.7% |
| A moderate amount | 41 | 35.7% |
| A lot | 10 | 8.7% |
3.3 Framing Effects
Respondents were randomly assigned to one of two question framings for the same underlying question. The "year framing" asked how many years until a given probability, while the "probability framing" asked for the probability at a fixed time horizon.
HLMI 50% Probability Year by Framing
| Framing | N (gamma fits) | Mixture CDF 50% Year | Calendar Year |
|---|---|---|---|
| Year framing (respondent provides years) | 124 | 40.4 yr | 2056 |
| Probability framing (respondent provides probabilities) | 128 | 50.7 yr | 2067 |
A consistent framing effect has been observed across survey waves: the year-framing tends to produce earlier (shorter-timeline) predictions than the probability-framing. This is a documented cognitive bias in probability elicitation.
3.4 Perceived Rates of Progress
Respondents were asked whether AI progress was faster in the first or second half of their career. (N=107)
Corresponds to Figure 4 in Grace et al. (2024).
| Response | Count | Percentage | 2023 Survey |
|---|---|---|---|
| The second half | 0 | 0.0% | 60% |
| The first half | 0 | 0.0% | — |
| They were about the same | 0 | 0.0% | — |
How Far Along Is AI Progress? (Outside-view slider)
Respondents were shown a slider with three anchors: A = where progress was when they started working in their AI area; B = where it is now; and C = where it would need to be for AI software to have roughly human-level abilities at the tasks they study. They were asked "What fraction of the distance between where progress was when you started working in the area (A) and where it would need to be to attain human-level abilities in the area (C) have we come so far (B)?" on a 0–100 scale. This complements the year-based timing questions by anchoring an estimate to each respondent's own career window — so it is robust to disagreement about absolute calendar dates.
| Survey | N | Median | Mean | SD |
|---|---|---|---|---|
| 2016 | 104 | 18 | 21.1 | 21.4 |
IQR (2016): 5–25.
3.5 What Causes AI Progress?
Respondents estimated how much less AI progress there would have been with half as much of each input. Higher values = more important to progress. Values are percentage less progress (0-100%).
Corresponds to Figure 5 in Grace et al. (2024). Red dots are means; box shows IQR with median line.
| Factor | N | Median | Mean | IQR |
|---|---|---|---|---|
| Computing hardware cost decline | 63 | 50.0% | 56.1% | 40–80% |
| AI algorithm progress | 57 | 50.0% | 44.8% | 30–50% |
| Training dataset effort | 70 | 40.0% | 43.7% | 20–69% |
| Funding | 67 | 40.0% | 43.2% | 30–60% |
| Researcher effort | 70 | 35.0% | 41.9% | 25–60% |
3.6 Will There Be an Intelligence Explosion?
Probability Estimates (numeric, 0-100%)
| Scenario | N | Median | Mean | IQR | 2023 Survey Median | Change |
|---|---|---|---|---|---|---|
| Dramatic tech speedup within 2 years of HLMI | 220 | 20% | 30.0% | 5–50% | 20% | unchanged |
| Dramatic tech speedup within 30 years of HLMI | 220 | 80% | 68.9% | 50–99% | 80% | unchanged |
| Vastly superhuman AI within 2 years of HLMI | 209 | 10% | 19.5% | 1–25% | 10% | unchanged |
| Vastly superhuman AI within 30 years of HLMI | 210 | 50% | 53.2% | 20–90% | — | — |
Is the Feedback Loop Argument Broadly Correct? (N=224)
"If AI does nearly all R&D, improvements in AI will accelerate progress including further AI progress. This could cause progress to become >10x faster within 5 years."
Corresponds to Figure 6 in Grace et al. (2024).
| Response | Count | Percentage |
|---|---|---|
| Quite unlikely (0-20%) | 58 | 25.9% |
| Unlikely (21-40%) | 53 | 23.7% |
| About even chance (41-60%) | 49 | 21.9% |
| Likely (61-80%) | 36 | 16.1% |
| Quite likely (81-100%) | 28 | 12.5% |
3.7 AI capabilities 2044
Section skipped — required data not available for this survey year. (KeyError: 'Q1#1_1')
3.8 Explainability
Section skipped — required data not available for this survey year. (KeyError: 'Q1 (1)')
4.1 Concerning scenarios
Section skipped — required data not available for this survey year. (KeyError: 'Q369#1_1')
4.2 How Good or Bad Will HLMI Be?
Respondents assigned probabilities to five outcome categories (summing to 100%). N=345
Corresponds to Figure 11 in Grace et al. (2024), "Thousands of AI Authors on the Future of AI."
Mean and Median Probabilities
| Outcome | Mean | Median |
|---|---|---|
| Extremely good | 27.1% | 20.0% |
| On balance good | 29.9% | 25.0% |
| Neutral | 20.0% | 20.0% |
| On balance bad | 14.2% | 10.0% |
| Extremely bad | 8.7% | 5.0% |
Individual Respondent Views
Each vertical slice below is one respondent's probability allocation, sorted from most optimistic (left) to most pessimistic (right). The chart reveals the full diversity of expert opinion.
Corresponds to Figure 10 in Grace et al. (2024).
The same respondents sorted by how much probability they assign to "extremely bad (e.g. human extinction)" outcomes. The growing black band on the right shows those assigning the highest probability to extremely bad outcomes.
No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.
No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.
Individual Response Profiles
Each small bar below represents one respondent's five-category probability distribution, sampled evenly from the most optimistic to the most pessimistic.
No direct equivalent in Grace et al. (2024); novel visualization of the same Section 4.2 data.
Extreme Outcome Analysis
| Metric | 2016 Survey | 2023 Survey | Change |
|---|---|---|---|
| Non-zero to both extremes | 64.1% | 64.0% | +0.1 pp |
| >=5% on extremely bad | 54.8% | 57.8% | -3.0 pp |
| >=10% on extremely bad | 40.0% | 37.8% | +2.2 pp |
| >=20% on extremely bad | 17.7% | — | — |
| >=25% on extremely bad | 9.6% | — | — |
| Mean extremely bad | 8.7% | 9.0% | -0.3 pp |
| Median extremely bad | 5.0% | 5.0% | unchanged |
| Net optimists | 70.7% | 68.3% | +2.4 pp |
4.3 Extinction risk
Section skipped — required data not available for this survey year. (KeyError: 'extinction_all_1')
4.4 Are Future AI-Risk Concerns Due to Misunderstandings of AI Research?
"To what extent do you think people's concerns about future risks from AI are due to misunderstandings of AI research?" (N=115)
No direct figure equivalent in Grace et al. (2024); this data is discussed in Section 4.4 of that paper.
| Response | Count | Percentage |
|---|---|---|
| Hardly at all | 0 | 0.0% |
| Not much | 0 | 0.0% |
| Somewhat | 0 | 0.0% |
| To a large extent | 0 | 0.0% |
| Almost entirely | 0 | 0.0% |
4.5 Rates of 5-Year Global AI Progress Prompting Most Optimism for Humanity
This question was not included in this survey year.
4.6 How Much Should AI Safety Research Be Prioritized?
"How much should society prioritize AI safety research, relative to how much it is currently prioritized?" (N=162)
Corresponds to Figure 14 in Grace et al. (2024).
| Response | Count | Percentage |
|---|---|---|
| Much less | 8 | 4.9% |
| Less | 12 | 7.4% |
| About the same | 63 | 38.9% |
| More | 56 | 34.6% |
| Much more | 23 | 14.2% |
4.7 The Alignment Problem
Corresponds to Figure 15 in Grace et al. (2024).
Do you think this argument points at an important problem? (N=149)
| Response | Count | Percentage |
|---|---|---|
| No, not a real problem. | 16 | 10.7% |
| No, not an important problem. | 29 | 19.5% |
| Yes, a moderately important problem. | 46 | 30.9% |
| Yes, a very important problem. | 8 | 5.4% |
| Yes, among the most important problems in the field. | 50 | 33.6% |
How valuable is it to work on this problem today, compared to other problems in AI? (N=149)
| Response | Count | Percentage |
|---|---|---|
| Much less valuable | 33 | 22.1% |
| Less valuable | 60 | 40.3% |
| As valuable as other problems | 42 | 28.2% |
| More valuable | 12 | 8.1% |
| Much more valuable | 2 | 1.3% |
How hard do you think this problem is compared to other problems in AI? (N=147)
| Response | Count | Percentage |
|---|---|---|
| Much easier | 11 | 7.5% |
| Easier | 28 | 19.0% |
| As hard as other problems | 61 | 41.5% |
| Harder | 33 | 22.4% |
| Much harder | 14 | 9.5% |
5. Results by Outlook Cluster
Section skipped — required data not available for this survey year. (KeyError: 'extinction_all_1')
Appendix: Cluster Analysis
Section skipped — required data not available for this survey year. (KeyError: 'Q369#1_1')
Appendix: Survey Flow
Survey-flow / randomization diagram modeled on Figure 16 of Grace et al., Thousands of AI Authors on the Future of AI. Each box is a question block. Percentages and n's are realized fill rates: the realized number of respondents who answered each block, divided by the total N. The total N counts everyone who answered at least one question (respondents who answered none are excluded).
Jobs / FAOL Sample-Size Reconciliation
The survey-flow boxes count anyone with any response in the broader Jobs / FAOL block. FAOL analyses use only the FAOL triplet, so their sample sizes can be slightly smaller.
| Count definition | Fixed-years | Fixed-probabilities |
|---|---|---|
| Whole Jobs / FAOL block (survey-flow box) | 49 | 44 |
| FAOL only, any triplet answer | 49 | 44 |
| FAOL only, complete triplet | 49 | 43 |
Here, "fixed-years" means respondents gave probabilities for fixed
time horizons (hj_b_full_*), while "fixed-probabilities" means respondents
gave years for fixed probabilities (hj_a_full_*).
Tasks: ta_* columns are the fixed-probability framing,
tb_* the fixed-years framing (verified empirically:
ta_* rows have fixedprobabilities=1,
tb_* rows have fixedprobabilities=0).
The Qualtrics survey-flow definition (.qsf) is not in
this repo, so the exact display order may differ from the diagram.
Appendix: Supplementary Figures
B.1 fixed-prob/fixed-year CDFs
Full CDF comparison between the fixed-prob and fixed-year conditions for HLMI and FAOL. The fixed-prob framing consistently produces earlier predictions (the red curve is shifted left relative to blue).
Corresponds to Figure 18 in Grace et al. (2024). Each curve is the mixture (mean) CDF for respondents assigned to that framing condition.
B.2 Bootstrap Confidence Bands
95% bootstrap confidence intervals on the aggregate CDF, obtained by resampling the fitted individual CDFs with replacement (500 resamples).
| Milestone | 50% Year (point est.) | 95% Bootstrap CI |
|---|---|---|
| HLMI | 2061 (45.5 yr) | 2056–2068 (40.1–51.9 yr) |
| FAOL | 2139 (123.5 yr) | 2110–2190 (94.1–173.9 yr) |
Bootstrap CIs reflect sampling variability — if we surveyed a different random sample of AI researchers, how much would the aggregate change? This is distinct from the spread of individual predictions (which is much wider).
Appendix: Statistical Tests
F.1 Yuen's Trimmed-Mean Bootstrap Test: Demographics
Tests whether experienced researchers (≥6 years in field) give significantly different HLMI predictions than junior researchers. Uses a 10% trimmed mean with 5,000 bootstrap resamples.
| Comparison | Trimmed-mean diff | 95% CI | p-value | Significant (α=0.05) |
|---|---|---|---|---|
| HLMI 50% year (experienced − junior, split at 6 yr) | +15.0 yr | [-26.3, +91.4] | 0.473 | No |
N (experienced): 41, N (junior): 32. Individual 50% years are computed from per-respondent gamma fits. The trimmed mean reduces sensitivity to extreme predictions.
F.2 Framing Effect: Statistical Significance
Tests whether year-framing and probability-framing respondents give significantly different predictions, using the same Yuen's bootstrap method.
| Milestone | Year-framing median | Prob-framing median | Trimmed-mean diff | 95% CI | p-value |
|---|---|---|---|---|---|
| HLMI | 39.9 yr (n=124) | 52.8 yr (n=128) | -24.9 yr | [-103.3, +0.9] | 0.211 |
| FAOL | 80.0 yr (n=43) | 130.5 yr (n=49) | -278.6 yr | [-973.2, +1808.6] | 0.707 |
A negative difference (yr_50 − pr_50) means the year-framing produces earlier predictions. Medians shown for reference; the statistical test uses trimmed means.
Appendix: Aggregation Method Comparison
Different choices in how individual survey responses are fitted and combined into aggregate forecasts can shift the headline numbers. This appendix compares three approaches on the key milestones, motivated by Adamczewski's reanalysis of ESPAI data.
Methods compared
Mix-Mean (MSE) — the Grace et al. 2023 method. Fit a gamma CDF to each respondent's 3 data points using mean squared error, then take the pointwise mean of all fitted CDFs.
Mix-Median (MSE) — same fitting, but take the pointwise median instead of mean. More robust to outlier predictions (e.g. respondents predicting 100M years). Adamczewski argues the median better represents the "typical expert."
Mix-Mean (LogLoss) — fit using log-loss (cross-entropy) instead of MSE, then take the mean. Log-loss naturally weighs errors near p=0 and p=1 more heavily than errors near p=0.5, matching the intuition that a 5% error at p=0.95 represents a much larger shift in belief than the same error at p=0.50.
Results
| Milestone | Mix-Mean (MSE) | Mix-Median (MSE) | Mix-Mean (LogLoss) | N (MSE) | N (LogLoss) |
|---|---|---|---|---|---|
| HLMI | 2061 (45.5 yr) | 2066 (49.6 yr) | 2061 (45.4 yr) | 252 | 252 |
| FAOL | 2139 (123.5 yr) | 2147 (130.5 yr) | 2152 (135.8 yr) | 92 | 92 |
| Truck Driver | 2028 (11.6 yr) | 2026 (10.2 yr) | 2027 (11.4 yr) | 89 | 89 |
| Surgeon | 2053 (37.4 yr) | 2055 (38.6 yr) | 2054 (37.6 yr) | 91 | 91 |
| Retail Salesperson | 2031 (14.6 yr) | 2030 (14.1 yr) | 2031 (14.7 yr) | 89 | 89 |
| AI Researcher | 2103 (87.2 yr) | 2097 (80.7 yr) | 2106 (90.2 yr) | 90 | 90 |
Key takeaways
Aggregation method matters more than loss function. Switching from mean to median aggregation pushes HLMI from 2061 to 2066 and FAOL from 2139 to 2147. Switching the loss function (MSE→log-loss) barely changes the aggregate — consistent with Adamczewski's finding that fitting choice has "hardly any impact."
The effect is largest for far-future milestones (FAOL, AI Researcher) where a few very pessimistic respondents pull the mean CDF rightward.
See also: compare_methods.py for the full 5-method comparison
across all 39 tasks, including linear interpolation and gamma-individual methods.
Method comparison motivated by
Adamczewski (bayes.net/espai).
Data: 2016 Expert Survey on Progress in AI (ESPAI).
Previous survey comparison values from "Thousands of AI Authors on the Future of AI" (2024 preprint).