Appendix: Cluster Analysis

A data-driven exploration of the natural groupings, the safety divide, and hardware vs. software beliefs among 1,793 ESPAI 2024 cleaned analysis rows.

Executive Summary

1. We tested clustering solutions from K=2 through K=6 and found that four clusters yielded the most informative groupings (n=1,538). Beyond the expected optimist and pessimist groups, a Polarized/Bimodal group emerges that assigns high probability to both extremely good AND extremely bad outcomes. This group is not simply pessimistic or technophobic -- their P(extremely good) is comparable to the most optimistic cluster. They are better understood as "high-impact believers" who are convinced AI will be transformative but genuinely uncertain whether that transformation will be positive or catastrophic.
2. We observe a meaningful distinction between more and less safety/severe-risk-concerned researchers, though it falls along a spectrum rather than a clean divide. Among 754 respondents with safety data, those in the high-concern group assign 12% mean probability to extremely bad outcomes (vs. 7% for low-concern), and give higher extinction/disempowerment probability estimates (32% vs. 11% mean), with moderate-to-large effect sizes on key measures.
3. Those who give higher extinction/disempowerment probability estimates tend to predict EARLIER HLMI (rho=-0.22, p=0.0000, n=446), suggesting that concern is driven by the belief that powerful AI is coming soon, not that AI is inherently dangerous regardless of timeline.
4. Hardware vs. software beliefs exist on a spectrum, leaning slightly hardware. Among respondents who rated both (n=58), 35 lean hardware and 18 lean software. Computing hardware is rated as the most impactful factor overall (median 60% progress reduction if halved), ahead of algorithms (50%), data (50%), and funding (40%). Notably, people who rate hardware as more important tend to rate safety as LESS important (rho=-0.32, p<0.01), hinting that hardware-focused people may have a more gradualist worldview.
Important caveat: Block randomization means respondents answered different question subsets. The cleaned analysis file contains 1,793 rows; the audience-facing 2024 response count is 1,580. Cross-cutting analyses use smaller item-level samples and should be treated accordingly. Sample sizes are reported throughout.

1. How Many Camps? The Natural Clusters

1.1 Clustering on Value Outlook (n=1,538)

The value-of-HLMI question asked respondents to assign probabilities (summing to 100%) across five outcomes: extremely good, on balance good, neutral, on balance bad, and extremely bad. It is the highest-coverage feature in this analysis. We cluster on these 5 dimensions using Gaussian Mixture Models.

K selection

Figure 1: Silhouette scores for K=2 through K=6 clusters on value outlook. K=2 has the highest silhouette, but K=3 and K=4 offer more interpretable structure.

Silhouette scores are modest (0.10-0.12), indicating the clusters are not sharply separated -- this is a continuous landscape of opinion, not discrete tribes. Still, the structure is meaningful.

1.2 Four-Camp Solution

Value outlook clusters K=4

Figure 2: Four natural camps in value outlook. Left: PCA projection. Center: mean probability profiles. Right: cluster sizes.

The most striking finding here is the Polarized/Bimodal group. These respondents assign high probability to both extremely good and extremely bad outcomes -- their P(extremely good) is comparable to the Strong Optimists, yet they simultaneously assign substantial probability to catastrophic outcomes. This is not a group of doomers or technophobes. They are researchers who believe AI will be hugely impactful, but who are genuinely torn on whether that impact will be positive or negative. The key axis for this group is magnitude of impact, not direction -- they have rejected the possibility that AI will be a modest or neutral development.

1.3 Three-Camp Solution (Simpler View)

Collapsing to 3 clusters (which has a higher silhouette score of 0.18 vs 0.07) merges the finer distinctions into a simpler optimist/moderate/pessimist framing. This loses the polarized group but provides a cleaner summary:

Value outlook clusters K=3

Figure 3: Three-camp simplification. The polarized group is absorbed into the moderate/pessimist clusters.

1.4 Intuitive View: How Soon vs. How Good

The PCA axes above lack intuitive meaning. Below, we plot each respondent on two directly interpretable dimensions: their predicted HLMI arrival year (x-axis) and their net optimism score (y-axis), defined as P(good + extremely good) minus P(bad + extremely bad). This reveals where the four camps sit in the space of "how soon" vs. "how beneficial."

HLMI year vs net optimism scatter

Figure 3b: Respondents (n=935) binned by predicted HLMI year and net optimism. Each bubble's size shows how many researchers from that cluster fall in the bin (labeled when ≥5). The vertical dashed line marks the median predicted year; the horizontal line separates net optimists from net pessimists.

1.5 Combined Clustering: Values + Concerns + Safety + Extinction

For the 191 respondents who answered all of: value outlook, concern scenarios, alignment-problem importance, and extinction/disempowerment probability, we ran a richer clustering incorporating 8 features.

Combined clusters

Figure 4: Combined clustering incorporating worldview, concerns, safety/alignment-problem views, and extinction/disempowerment probability. Mean concern averages 11 scenario concern ratings on a 0-3 scale (0=no concern, 3=extreme concern); alignment-problem importance is a 0-4 ordinal score (0=not a real problem, 4=among the most important problems in the field); P(extinction/disempowerment) is a percentage.

ClusterSizeP(Good / Ext good)P(Bad / Ext bad)Mean concernAlignment-problem impP(extinction/disempowerment)HLMI Year median [IQR]; mean
Optimistic / Low x-risk 103 (54%) 24% / 35% 12% / 3% 1.6/3 2.3/4 4% 2049 [2034-2074]; mean 16981 (n=67)
Pessimistic / High x-risk 88 (46%) 25% / 18% 23% / 18% 1.9/3 2.5/4 34% 2043 [2032-2064]; mean 2056 (n=56)

HLMI year summaries include the median, interquartile range, mean, and item-level n because several groups share the same median year while their wider distributions differ.

2. Higher vs Lower Safety/Severe-Risk Concern

2.1 Defining the Groups

We built a safety composite score from (normalized 0-1 and averaged):

Requiring at least 2 of 3 components, we obtained scores for 754 respondents and split at the median (0.500) into High (n=330) and Low (n=424) safety concern groups.

2.2 The Full Comparison

Safety comparison

Figure 5: Comprehensive comparison of higher vs lower safety/severe-risk-concern respondents across six dimensions. Alignment-problem importance uses the 0-4 ordinal scale above; mean concern averages 11 scenario concern ratings on a 0-3 scale; the optimism-maximizing AI progress-rate answer is coded 0=much slower, 2=current speed, 4=much faster.

2.3 Statistical Tests

Figure 5 panelTested variablen (High)n (Low)Median (High)Median (Low)Mean (High)Mean (Low)p-valueEffect r
HLMI timeline HLMI Year 210258 2044.02044.0 478730.025329.1 0.0491 * 0.105
Extinction/disempowerment estimate P(extinction/disempowerment) 130234 30.05.0 31.710.8 0.0000 *** -0.466
Value outlook P(Extremely bad) 330424 5.05.0 12.47.3 0.0000 *** -0.182
Value outlook P(Extremely good) 330424 10.020.0 22.324.8 0.0568 n.s. 0.080
Concern levels Mean concern 163220 1.91.6 1.91.6 0.0000 *** -0.348
AI progress rate for optimism Optimism-maximizing AI progress rate 81103 2.02.0 1.92.2 0.0501 n.s. 0.164

Selected Mann-Whitney U tests for scalar summaries from Figure 5. The value-outlook panel is summarized by P(extremely good) and P(extremely bad), the concern panel by mean concern, and the AI-capabilities panel is tested item-by-item in the next table. Effect size r is rank-biserial correlation (|r| > 0.3 = medium, |r| > 0.5 = large). Scale notes: mean concern 0-3, optimism-maximizing AI progress-rate answer 0-4, probabilities in percentage points. *** p<0.001, ** p<0.01, * p<0.05

2.4 AI Capabilities by 2044: What Do Safety People Expect?

Safety-concerned people don't just worry more -- they have systematically different expectations for what AI will be able to do by 2044:

CapabilityHypothesized directionMean (High Safety)Mean (Low Safety)High - LowObserved directionp-valuen
Talk like expert Exploratory 3.413.34 +0.07 High > Low 0.7903 n.s. 199
Self-improve regardless Exploratory 2.052.00 +0.05 High > Low 0.7481 n.s. 199
AI-AI collaborations High > Low 2.372.11 +0.26 High > Low 0.1911 n.s. 199
Deceive humans High > Low 2.522.23 +0.29 High > Low 0.0820 n.s. 197
Unexpected strategies Exploratory 3.263.34 -0.07 High < Low 0.3043 n.s. 200
Seek power High > Low 1.741.35 +0.38 High > Low 0.0104 * 200
Explain actions (trustworthy) Exploratory 2.082.10 -0.02 High < Low 0.7394 n.s. 200
Can be jailbroken Exploratory 2.862.68 +0.18 High > Low 0.3029 n.s. 194
Surprising behavior Exploratory 2.762.91 -0.15 High < Low 0.3611 n.s. 200
Real-world actions Exploratory 2.372.19 +0.18 High > Low 0.2788 n.s. 199
Misaligned goals High > Low 2.461.94 +0.53 High > Low 0.0016 ** 199

Scale: 0=Very unlikely, 1=Unlikely, 2=Even chance, 3=Likely, 4=Very likely. High - Low is the high-safety mean minus the low-safety mean, so positive values mean high-safety respondents rated the capability as more likely. Hypothesized direction is marked for risk-relevant capabilities where we expected High > Low; other rows are exploratory.

Overall risk-relevant capability pattern: Averaging the 4 hypothesized High > Low items (AI-AI collaboration, deception, power-seeking, and misaligned goals), high-safety respondents rated these capabilities as more likely by +0.37 points on the 0-4 scale (means 2.27 vs. 1.90; medians 2.38 vs. 2.00; n=84 high, 115 low; Mann-Whitney p=0.0017 **, |r|=0.26). Respondents are included if they answered at least 2 of the 4 items.
Key pattern: Safety-concerned respondents do rate certain dangerous capabilities (power-seeking, AI-AI collaboration, misaligned goals, deception) as more likely, while agreeing with others on benign capabilities (expert conversation, explaining actions). However, the magnitude of these differences is notably small -- typically 0.3-0.6 points on a 0-4 scale. Both groups largely agree on what AI will be able to do by 2044. This suggests the safety divide is driven less by disagreements about AI's technical capabilities and more by differences in values, risk tolerance, or beliefs about how those capabilities will be managed.

2.5 Intelligence Explosion Beliefs

Intelligence explosion

Figure 6: Probability estimates for intelligence explosion scenarios, by safety concern level.

3. Hardware vs. Software Progress Beliefs

3.1 Counterfactual Progress Reduction if Factors Were Halved

Respondents rated how much AI progress would decrease if each of 5 factors were cut in half (0-100% scale). Higher values mean the factor is more important.

Progress attribution

Figure 7: Counterfactual progress-factor analysis. Top-left: estimated progress reduction distributions. Top-right: hardware vs algorithm progress-reduction estimates. Bottom-left: HW-SW index distribution. Bottom-right: factor correlations with other variables.

Rankings (by median):
  1. Computing hardware: 60% (n=119) -- largest median estimated progress reduction if halved
  2. Training data: 50% (n=109)
  3. Algorithm progress: 50% (n=103)
  4. Funding: 40% (n=116)
  5. Researcher effort: 25% (n=95) -- surprisingly lowest

3.2 Hardware vs Software Index

We computed (Hardware - Algorithms) / (Hardware + Algorithms) for each respondent who rated both. Positive = hardware-leaning, negative = software-leaning.

3.3 Do Progress Beliefs Predict Other Views?

Sample size note: Only ~60-120 respondents answered the progress cause questions. Cross-tabulations with other variables yield modest samples (n=15-60). Treat these as suggestive.

Significant cause-factor correlations (p < 0.10)

Cause FactorTargetSpearman rhop-valuen
Researcher effortHLMI Year -0.238 0.0647 n.s. 61
Computing hardwareAlignment-problem imp -0.316 0.0086 ** 68
Training dataMean concern 0.337 0.0166 * 50

Hardware-Software Index Correlations

Variable 1Variable 2Spearman rhop-valuen
HW-vs-SW IndexP(extinction/disempowerment) 0.089 0.6212 n.s. 33
HW-vs-SW IndexHLMI Year 0.073 0.6758 n.s. 35
HW-vs-SW IndexP(Ext bad) -0.165 0.2152 n.s. 58
HW-vs-SW IndexAlignment-problem imp -0.369 0.0344 * 33
HW-vs-SW IndexMean concern 0.106 0.5847 n.s. 29
HW-vs-SW IndexP(Ext good) 0.048 0.7207 n.s. 58
The hardware-safety link: People who attribute more importance to computing hardware tend to rate safety as less important. This suggests a meaningful (if modest-sample) divide: "hardware people" may see AI progress as more incremental and predictable, while "algorithm people" may see it as more unpredictable and discontinuous -- leading to greater safety concern.

3.4 Progress Beliefs by Outlook and Safety Groups

Do optimists, pessimists, and safety-concerned researchers differ in what they think drives AI progress? Below we break down cause factor importance by outlook cluster (left) and safety group (right).

Cause factors by group

Figure 7b: Mean importance ratings for each progress factor, split by outlook group (left) and safety group (right). Scale: 0-100% estimated decrease in progress if factor halved.

Pessimists emphasize algorithms more. Pessimists rate algorithm progress as notably more important than optimists do, while both groups rate hardware similarly. This aligns with the hardware-safety correlation: those who see AI progress as driven by algorithmic breakthroughs (rather than predictable hardware scaling) tend to be more concerned about safety -- perhaps because algorithmic advances feel less predictable and harder to control. By contrast, the safety-group split shows little difference in cause factor ratings, suggesting the connection between progress beliefs and concern runs primarily through overall outlook rather than safety concern per se.

4. How Outlook Predicts Everything Else

The three value-outlook clusters (from Section 1) predict views across many other dimensions:

VariableConcernedMild OptimistsStrong OptimistsPolarized/Bimodal
HLMI Year2046 [2034-2074]; mean 533995 (n=188)2044 [2032-2064]; mean 19307 (n=290)2044 [2034-2064]; mean 1101046 (n=274)2044 [2033-2069]; mean 2123 (n=183)
P(extinction/disempowerment)10.0 (n=166)7.5 (n=210)1.0 (n=223)20.0 (n=145)
Alignment-problem importance3.0 (n=162)3.0 (n=220)2.0 (n=228)3.0 (n=145)
Mean concern2.0 (n=165)1.7 (n=243)1.5 (n=198)1.8 (n=148)
Optimism-maximizing AI progress rate1.0 (n=75)2.0 (n=111)2.0 (n=109)2.0 (n=77)

Values are medians with sample sizes in parentheses, except HLMI Year, which shows median [IQR], mean, and item-level n because several groups share the same median year. Scale notes: alignment-problem importance 0-4, mean concern 0-3, optimism-maximizing AI progress-rate answer 0=much slower to 4=much faster, probabilities in percentage points.

Outlook group comparisons

Figure 8: How the four value-outlook camps compare across timelines, extinction/disempowerment probability, safety/alignment-problem views, concern levels, optimism-maximizing AI progress-rate answers, and P(extremely bad). Mean concern averages 11 scenario concern ratings on a 0-3 scale; alignment-problem importance and optimism-maximizing AI progress-rate answers use 0-4 ordinal scales.

5. What Dimensions Structure AI Researcher Beliefs?

Principal Component Analysis on respondents with values + concerns + safety + extinction/disempowerment data (n=180) reveals the latent axes of disagreement.

PCA loadings

Figure 9: PCA factor loadings. Green bars = positive loading, red = negative. Each panel shows one principal component.

PC1 (21.4% of variance) is the "concern axis." It loads positively on all 11 concern items, alignment-problem importance, value of working on the alignment problem today, and P(extinction/disempowerment), and negatively on optimistic outlook. This single dimension captures most of the disagreement: how worried are you about AI?
PC2 (10.5%) is the "extremity axis." It loads positively on both P(Extremely good) AND P(Extremely bad)/P(extinction/disempowerment), and negatively on P(Neutral). This separates people who hold extreme views in either direction from those with moderate views. Some respondents genuinely believe AI could be transformatively good AND pose extinction/disempowerment risk.
PC3 (7.9%) separates "institutional safety" from "pessimism." It loads positively on P(On balance bad) and P(extinction/disempowerment), but negatively on alignment-problem importance and value of working on the alignment problem today. This captures people who think bad outcomes are likely but DON'T think the alignment problem is important -- perhaps because they think the problems aren't technical, or aren't solvable.

6. The Full Correlation Structure

Correlation matrix

Figure 10: Spearman rank correlations between key variables. Stars indicate significance. Sample sizes shown in each cell. Alignment-problem importance uses a 0-4 ordinal scale from not a real problem to among the field's most important problems; value of working on the alignment problem today uses a 0-4 ordinal scale from much less valuable to much more valuable than other AI problems; mean concern uses a 0-3 scale; HLMI Year is a calendar-year estimate; probabilities are 0-100 percentages.

Key Pairwise Correlations

RelationshipSpearman rhop-valuen
P(extinction/disempowerment) x HLMI Year -0.216 0.0000 *** 446
P(extinction/disempowerment) x P(Ext bad) 0.436 0.0000 *** 744
Alignment-problem imp x P(extinction/disempowerment) 0.141 0.0072 ** 364
Alignment-problem imp x HLMI Year -0.057 0.2211 n.s. 468
Alignment-problem imp x Mean concern 0.276 0.0000 *** 383
P(Ext good) x P(Ext bad) -0.113 0.0000 *** 1538
HLMI Year x P(Ext bad) -0.025 0.4421 n.s. 935
HLMI Year x P(Ext good) -0.074 0.0228 * 935
Mean concern x P(extinction/disempowerment) 0.321 0.0000 *** 381
Mean concern x HLMI Year -0.102 0.0293 * 461
Optimism-max progress rate x P(extinction/disempowerment) -0.158 0.0327 * 182
Optimism-max progress rate x HLMI Year -0.210 0.0010 *** 242
Optimism-max progress rate x alignment-problem imp -0.016 0.8289 n.s. 184
Strongest relationships:

7. Methods & Caveats

7.1 Approach

7.2 Caveats

Block randomization: Not all respondents answered all questions. The "causes of progress" questions were only seen by ~6% of respondents. Safety questions by ~42%. This creates inherently unequal coverage for cross-cutting analyses.
Soft boundaries: The silhouette scores (0.08-0.12) indicate that clusters are not crisply separated. These are regions of higher density in a continuous opinion landscape, not discrete tribes. Median splits are even more arbitrary. The true distribution of views is continuous.
Causality: Correlations between safety concern, timeline predictions, and value outlook don't establish causal direction. A "safety worldview" (short timelines + high concern + safety-focused) may be a coherent package of beliefs adopted together, not a chain where one causes another.
No demographic controls: We don't control for region, career stage, subfield, or other demographics that might confound the clustering. Some of the "camps" might partially reflect institutional or geographic cultures rather than independent belief formation.

7.3 Technical Details