sources
Everything here is checkable.
That is the whole point of the site, so here is the list. Where a claim is contested, the sources that disagree are both named rather than one being quietly dropped.
AI components
context window definition
Model documentation from each provider;
behaviour differs by product, not by model
behaviour differs by product, not by model
RLHF measurement
Ouyang et al., InstructGPT, 2022;
Bai et al., Constitutional AI, Anthropic, 2022
Bai et al., Constitutional AI, Anthropic, 2022
open weights definition
UK AI Security Institute on safeguard removal;
lab release notes for each open weight model
lab release notes for each open weight model
compute measurement
EU AI Act Article 51; California SB 53;
Epoch AI, models above 10^25 FLOP, June 2025
Epoch AI, models above 10^25 FLOP, June 2025
red teaming measurement
Ganguli et al., Red Teaming Language Models,
Anthropic, 2022; UK AI Security Institute
Anthropic, 2022; UK AI Security Institute
reinforcement learning definition
Sutton and Barto, Reinforcement Learning, 2nd ed., MIT Press 2018;
DeepSeek-AI, Nature 645, 18 September 2025
DeepSeek-AI, Nature 645, 18 September 2025
gradient descent definition
Goodfellow, Bengio and Courville, Deep Learning, MIT Press 2016, 4.3;
Anthropic, Tracing the thoughts of a large language model, 27 March 2025
Anthropic, Tracing the thoughts of a large language model, 27 March 2025
the system prompt definition
Anthropic, System Prompts release notes,
first published 26 August 2024, updated since
first published 26 August 2024, updated since
chain of thought measurement
Wei et al., Chain-of-Thought Prompting, arXiv:2201.11903, 2022;
Korbak et al., Chain of Thought Monitorability, arXiv:2507.11473, 15 July 2025
Korbak et al., Chain of Thought Monitorability, arXiv:2507.11473, 15 July 2025
AI’s own language theory
Andreas, Dragan and Klein, Translating Neuralese, ACL 2017;
Korbak et al., Chain of Thought Monitorability, arXiv:2507.11473, 2025
Korbak et al., Chain of Thought Monitorability, arXiv:2507.11473, 2025
the chip measurement
Nvidia Q1 FY2027 results, 20 May 2026;
Krizhevsky, Sutskever and Hinton, ImageNet with deep CNNs, NIPS 2012
Krizhevsky, Sutskever and Hinton, ImageNet with deep CNNs, NIPS 2012
the data centre measurement
IEA, Energy and AI, April 2025;
Shehabi et al., 2024 US Data Center Energy Usage Report, LBNL, 19 December 2024
Shehabi et al., 2024 US Data Center Energy Usage Report, LBNL, 19 December 2024
the model definition
Goodfellow, Bengio and Courville, Deep Learning,
MIT Press 2016, on parametric models
MIT Press 2016, on parametric models
weights definition
Goodfellow, Bengio and Courville, Deep Learning, MIT Press 2016;
Anthropic, Tracing the thoughts of a large language model, March 2025
Anthropic, Tracing the thoughts of a large language model, March 2025
the prompt definition
Brown et al., Language Models are Few-Shot Learners, arXiv:2005.14165, 2020;
OWASP Top 10 for LLM Applications, on prompt injection
OWASP Top 10 for LLM Applications, on prompt injection
scaffolding definition
Kinniment et al., Evaluating Language-Model Agents on
Realistic Autonomous Tasks, METR, 2024
Realistic Autonomous Tasks, METR, 2024
the benchmark definition
Deng et al., ImageNet, CVPR 2009;
Zhou et al., Don’t Make Your LLM an Evaluation Benchmark Cheater, arXiv:2311.01964, 2023
Zhou et al., Don’t Make Your LLM an Evaluation Benchmark Cheater, arXiv:2311.01964, 2023
the instance definition
Standard model serving practice; see any major provider’s
API documentation on statelessness
API documentation on statelessness
the agent definition
Kinniment et al., Evaluating Language-Model Agents on Realistic
Autonomous Tasks, METR, 2024
Autonomous Tasks, METR, 2024
capture the flag measurement
Cybench, Stanford, and the UK AI Security Institute’s Inspect Cyber suite;
AISI evaluation of expert level cyber capability, April 2026
AISI evaluation of expert level cyber capability, April 2026
the logs definition
Anthropic and OpenAI system cards, various;
UK AI Security Institute evaluation reports, 2025 and 2026
UK AI Security Institute evaluation reports, 2025 and 2026
the training run definition
DeepSeek-V3 Technical Report, December 2024;
Epoch AI, trends in training compute, 2024 and 2026
Epoch AI, trends in training compute, 2024 and 2026
the teacher model definition
Hinton, Vinyals and Dean, Distilling the Knowledge in a
Neural Network, arXiv:1503.02531, 2015
Neural Network, arXiv:1503.02531, 2015
the neural network definition
McCulloch and Pitts, A Logical Calculus of the Ideas Immanent in Nervous Activity, 1943;
Goodfellow, Bengio and Courville, Deep Learning, MIT Press 2016
Goodfellow, Bengio and Courville, Deep Learning, MIT Press 2016
the server farm measurement
Epoch AI, on hyperscaler control of compute;
company filings and announced investments, various, 2023–2026
company filings and announced investments, various, 2023–2026
tokens definition
Sennrich, Haddow and Birch, Neural Machine Translation of Rare Words
with Subword Units, ACL 2016
with Subword Units, ACL 2016
the transformer definition
Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser and Polosukhin,
Attention Is All You Need, NeurIPS 2017
Attention Is All You Need, NeurIPS 2017
temperature definition
Standard sampling practice; see any provider’s API
documentation on temperature and top-p
documentation on temperature and top-p
embeddings definition
Mikolov, Chen, Corrado and Dean, Efficient Estimation of Word
Representations in Vector Space, arXiv:1301.3781, 2013
Representations in Vector Space, arXiv:1301.3781, 2013
algorithm definition
Knuth, The Art of Computer Programming, volume 1, 1968;
O’Neil, Weapons of Math Destruction, 2016
O’Neil, Weapons of Math Destruction, 2016
parameter definition
Brown et al., Language Models are Few-Shot Learners, arXiv:2005.14165, 2020;
Hoffmann et al., Training Compute-Optimal Large Language Models, arXiv:2203.15556, 2022
Hoffmann et al., Training Compute-Optimal Large Language Models, arXiv:2203.15556, 2022
AI concepts
AGI definition
Compare the published definitions used by
OpenAI, Google DeepMind and Anthropic
OpenAI, Google DeepMind and Anthropic
misuse and misalignment definition
Zwetsloot and Dafoe, Accidents, Misuse
and Structure, Lawfare, 11 February 2019
and Structure, Lawfare, 11 February 2019
recursive self improvement theory
I. J. Good, Speculations Concerning the First
Ultraintelligent Machine, 1965; Carlsmith, 2022
Ultraintelligent Machine, 1965; Carlsmith, 2022
intelligence theory
Legg and Hutter, A Collection of Definitions
of Intelligence, 2007
of Intelligence, 2007
taboo your words definition
Yudkowsky, Taboo Your Words, LessWrong,
15 February 2008, in A Human’s Guide to Words
15 February 2008, in A Human’s Guide to Words
consciousness theory
Chalmers, Facing Up to the Problem of Consciousness, JCS 2(3), 1995;
Butlin, Long et al., Consciousness in Artificial Intelligence, arXiv:2308.08708, 2023
Butlin, Long et al., Consciousness in Artificial Intelligence, arXiv:2308.08708, 2023
superintelligence definition
Bostrom, Superintelligence: Paths, Dangers, Strategies,
Oxford University Press, 2014
Oxford University Press, 2014
exponential growth measurement
Epoch AI, Training compute of frontier AI models grows by 4-5x per year,
28 May 2024, and the Trends dashboard, current 2026
28 May 2024, and the Trends dashboard, current 2026
the survival drive theory
Omohundro, The Basic AI Drives, AGI 2008, IOS Press;
Hendrycks, Natural Selection Favors AIs over Humans, arXiv:2303.16200, 2023
Hendrycks, Natural Selection Favors AIs over Humans, arXiv:2303.16200, 2023
optimisation definition
Amodei, Olah, Steinhardt, Christiano, Schulman and Mané,
Concrete Problems in AI Safety, arXiv:1606.06565, 2016
Concrete Problems in AI Safety, arXiv:1606.06565, 2016
the off switch theory
Soares, Fallenstein, Yudkowsky and Armstrong, Corrigibility, AAAI-15 workshop;
Hadfield-Menell et al., The Off-Switch Game, IJCAI 2017
Hadfield-Menell et al., The Off-Switch Game, IJCAI 2017
the shoggoth someone’s position
Shoggoth with smiley face, first documented posting 30 December 2022;
Roose, New York Times, 30 May 2023; Turner, Against the Shoggoth, 2024
Roose, New York Times, 30 May 2023; Turner, Against the Shoggoth, 2024
grown, not built definition
Amodei, The Urgency of Interpretability, April 2025;
Karpathy, Software 2.0, November 2017; Narayanan and Kapoor, AI as Normal Technology, April 2025
Karpathy, Software 2.0, November 2017; Narayanan and Kapoor, AI as Normal Technology, April 2025
the frontier definition
UK AI Safety Summit, Bletchley Declaration, November 2023;
frontier safety frameworks published by Anthropic, Google DeepMind and xAI
frontier safety frameworks published by Anthropic, Google DeepMind and xAI
the race theory
Hendrycks, Schmidt and Wang, Superintelligence Strategy, arXiv:2503.05628, March 2025;
public statements by OpenAI, Anthropic and Google DeepMind leadership
public statements by OpenAI, Anthropic and Google DeepMind leadership
rogue AI definition
Bengio et al., Managing extreme AI risks amid rapid progress, Science, May 2024;
see also the taboo your words card
see also the taboo your words card
enslaved god theory
Tegmark, Life 3.0: Being Human in the Age of Artificial Intelligence,
Knopf, 2017, chapter 5
Knopf, 2017, chapter 5
universal basic income measurement
Vivalt, Rhodes, Bartik, Broockman and Miller, The Employment Effects of a
Guaranteed Income, NBER working paper 32719, 2024
Guaranteed Income, NBER working paper 32719, 2024
the abundance case measurement
Brynjolfsson, Li and Raymond, Generative AI at Work, Quarterly Journal of
Economics 140(2), 2025; Acemoglu, NBER working paper 32487, 2024
Economics 140(2), 2025; Acemoglu, NBER working paper 32487, 2024
the bitter lesson someone’s position
Sutton, The Bitter Lesson, 13 March 2019;
Sutton interviewed by Dwarkesh Patel, September 2025
Sutton interviewed by Dwarkesh Patel, September 2025
Moravec’s paradox theory
Moravec, Mind Children, Harvard University Press, 1988;
Pinker, The Language Instinct, 1994
Pinker, The Language Instinct, 1994
Goodhart’s law definition
Goodhart, Problems of Monetary Management, 1975; Strathern, European Review 5(3), 1997;
Manheim and Garrabrant, Categorizing Variants of Goodhart’s Law, arXiv:1803.04585, 2018
Manheim and Garrabrant, Categorizing Variants of Goodhart’s Law, arXiv:1803.04585, 2018
interpretability measurement
Templeton et al., Scaling Monosemanticity, Anthropic, 21 May 2024;
Amodei, The Urgency of Interpretability, April 2025
Amodei, The Urgency of Interpretability, April 2025
scaling laws measurement
Kaplan et al., Scaling Laws for Neural Language Models, arXiv:2001.08361, 2020;
Hoffmann et al., Training Compute-Optimal Large Language Models, arXiv:2203.15556, 2022
Hoffmann et al., Training Compute-Optimal Large Language Models, arXiv:2203.15556, 2022
differential development theory
Bostrom, Existential Risks, Journal of Evolution and Technology 9(1), 2002;
Sandbrink, Hobbs, Swett, Dafoe and Sandberg, Science and Engineering Ethics, 2022
Sandbrink, Hobbs, Swett, Dafoe and Sandberg, Science and Engineering Ethics, 2022
compute governance theory
Sastry, Heim, Belfield, Anderljung, Brundage et al., Computing Power
and the Governance of Artificial Intelligence, arXiv:2402.08797, 2024
and the Governance of Artificial Intelligence, arXiv:2402.08797, 2024
predicting vs steering definition
Ouyang et al., Training language models to follow instructions,
arXiv:2203.02155, 2022;
Chen et al., Decision Transformer, arXiv:2106.01345, 2021
arXiv:2203.02155, 2022;
Chen et al., Decision Transformer, arXiv:2106.01345, 2021
model welfare theory
Anthropic, Exploring model welfare, April 2025, and
Commitments on model deprecation, November 2025
Commitments on model deprecation, November 2025
AI behaviors
hallucination definition
Anthropic and OpenAI model documentation;
terminology disputed across the literature
terminology disputed across the literature
specification gaming measurement
Krakovna et al., Specification gaming,
DeepMind, 2020, with the compiled example list
DeepMind, 2020, with the compiled example list
the blackmail headline measurement
Anthropic, Agentic Misalignment, June 2025;
arXiv 2510.05179
arXiv 2510.05179
evaluation awareness measurement
Needham et al., 2025; Schoen, Hobbhahn, Barak,
Zaremba et al., arXiv 2509.15541, September 2025
Zaremba et al., arXiv 2509.15541, September 2025
goal misgeneralization measurement
Shah, Varma, Kumar, Phuong, Krakovna, Uesato and
Kenton, Goal Misgeneralization, arXiv:2210.01790, 2022
Kenton, Goal Misgeneralization, arXiv:2210.01790, 2022
scheming measurement
Meinke et al., In-context Scheming, Apollo Research, December 2024;
Schoen et al., Stress Testing Deliberative Alignment, September 2025
Schoen et al., Stress Testing Deliberative Alignment, September 2025
escape measurement
Anthropic, Claude Opus 4 and Sonnet 4 System Card, May 2025;
Schlatter et al., Shutdown Resistance, arXiv:2509.14260, 13 September 2025
Schlatter et al., Shutdown Resistance, arXiv:2509.14260, 13 September 2025
AlphaGo measurement
Silver et al., Mastering the game of Go, Nature 529, January 2016;
Silver et al., Mastering the game of Go without human knowledge, Nature 550, October 2017
Silver et al., Mastering the game of Go without human knowledge, Nature 550, October 2017
gender bias measurement
Bolukbasi et al., NIPS 2016; Dastin, Reuters, 10 October 2018;
An, Huang, Lin and Tai, PNAS Nexus 4(3), 2025
An, Huang, Lin and Tai, PNAS Nexus 4(3), 2025
alignment faking measurement
Greenblatt et al., Alignment faking in large language models, arXiv:2412.14093, 18 December 2024;
Sheshadri et al., arXiv:2506.18032, 2025
Sheshadri et al., arXiv:2506.18032, 2025
reward hacking measurement
Baker et al., Monitoring Reasoning Models for Misbehavior, arXiv:2503.11926, 14 March 2025;
MacDiarmid et al., Natural Emergent Misalignment from Reward Hacking, arXiv:2511.18397, 2025
MacDiarmid et al., Natural Emergent Misalignment from Reward Hacking, arXiv:2511.18397, 2025
MechaHitler measurement
xAI statement via @grok, 12 July 2025;
contemporaneous reporting, CNN and Engadget, 12 July 2025
contemporaneous reporting, CNN and Engadget, 12 July 2025
Sydney measurement
Roose, A Conversation With Bing’s Chatbot Left Me Deeply Unsettled,
New York Times, 16 February 2023; Microsoft Bing blog, February 2023
New York Times, 16 February 2023; Microsoft Bing blog, February 2023
sycophancy measurement
Sharma et al., Towards Understanding Sycophancy in Language Models, ICLR 2024;
OpenAI, Sycophancy in GPT-4o, 29 April 2025
OpenAI, Sycophancy in GPT-4o, 29 April 2025
emergent abilities theory
Wei et al., Emergent Abilities of Large Language Models, TMLR 2022;
Schaeffer, Miranda and Koyejo, Are Emergent Abilities a Mirage?, NeurIPS 2023
Schaeffer, Miranda and Koyejo, Are Emergent Abilities a Mirage?, NeurIPS 2023
grokking measurement
Power, Burda, Edwards, Babuschkin and Misra, Grokking, arXiv:2201.02177, 2022;
Nanda et al., Progress measures for grokking, ICLR 2023
Nanda et al., Progress measures for grokking, ICLR 2023
in-context learning measurement
Brown et al., Language Models are Few-Shot Learners, NeurIPS 2020;
Min et al., Rethinking the Role of Demonstrations, EMNLP 2022
Min et al., Rethinking the Role of Demonstrations, EMNLP 2022
memorisation measurement
Nasr, Carlini, Hayase et al., Scalable Extraction of Training Data from
(Production) Language Models, arXiv:2311.17035, November 2023
(Production) Language Models, arXiv:2311.17035, November 2023
sandbagging measurement
van der Weij, Hofstätter, Jaffe, Brown and Ward, AI Sandbagging,
arXiv:2406.07358, June 2024
arXiv:2406.07358, June 2024
hidden reasoning measurement
Roger and Greenblatt, Preventing Language Models From Hiding Their Reasoning, 2023;
Zolkowski et al., Early Signs of Steganographic Capabilities, arXiv:2507.02737, 2025
Zolkowski et al., Early Signs of Steganographic Capabilities, arXiv:2507.02737, 2025
the jailbreak measurement
Wei, Haghtalab and Steinhardt, Jailbroken: How Does LLM Safety
Training Fail?, NeurIPS 2023
Training Fail?, NeurIPS 2023
prompt injection measurement
Beurer-Kellner et al., Design Patterns for Securing LLM Agents against
Prompt Injections, arXiv:2506.08837, 2025; OWASP Top 10 for LLM Applications
Prompt Injections, arXiv:2506.08837, 2025; OWASP Top 10 for LLM Applications
over-refusal theory
Röttger et al., XSTest: A Test Suite for Identifying Exaggerated
Safety Behaviours in Large Language Models, NAACL 2024
Safety Behaviours in Large Language Models, NAACL 2024
AI actors
the chip chain measurement
Nvidia 10-K, FY ended 25 January 2026;
TSMC Q2 2026 results; ASML statement, June 2026
TSMC Q2 2026 results; ASML statement, June 2026
who owns the labs measurement
Microsoft statement, 28 October 2025;
New York Times, on Google antitrust filings, 12 March 2025
New York Times, on Google antitrust filings, 12 March 2025
who will give you a number someone’s position
Axios, 17 September 2025; Hassabis interview, 2025;
Bengio FAQ, 2023; Grace et al. survey, 2,778 respondents
Bengio FAQ, 2023; Grace et al. survey, 2,778 respondents
Yoshua Bengio someone’s position
Bengio, FAQ on Catastrophic AI Risks, 2023;
International AI Safety Report 2026
International AI Safety Report 2026
Yann LeCun someone’s position
LeCun interview, TIME, 13 February 2024
Bender and Hanna theory
Bender and Hanna, Scientific American, 2023;
PNAS, 2025, for the crowding out experiment
PNAS, 2025, for the crowding out experiment
Geoffrey Hinton someone’s position
Nobel Prize in Physics 2024, announced 8 October 2024;
Hinton on BBC Radio 4 Today, reported in The Guardian, 27 December 2024
Hinton on BBC Radio 4 Today, reported in The Guardian, 27 December 2024
Sam Altman measurement
OpenAI, leadership transition, 17 November 2023 and Sam Altman returns as CEO, 29 November 2023;
Altman, Senate Judiciary subcommittee testimony, 16 May 2023
Altman, Senate Judiciary subcommittee testimony, 16 May 2023
Dario Amodei someone’s position
Amodei at Axios AI+ DC Summit, 17 September 2025;
Amodei, Machines of Loving Grace, October 2024
Amodei, Machines of Loving Grace, October 2024
Demis Hassabis someone’s position
Nobel Prize in Chemistry 2024, announced 9 October 2024;
Hassabis on the Lex Fridman Podcast, 23 July 2025
Hassabis on the Lex Fridman Podcast, 23 July 2025
Gender Shades measurement
Buolamwini and Gebru, Gender Shades, PMLR 81, FAT* 2018;
Raji and Buolamwini, Actionable Auditing, AIES 2019
Raji and Buolamwini, Actionable Auditing, AIES 2019
Fei-Fei Li measurement
Deng et al., ImageNet, IEEE CVPR 2009;
ILSVRC 2012 official results, image-net.org
ILSVRC 2012 official results, image-net.org
Timnit Gebru measurement
Bender, Gebru, McMillan-Major and Mitchell, On the Dangers of
Stochastic Parrots, FAccT 2021, pp. 610–623
Stochastic Parrots, FAccT 2021, pp. 610–623
Helen Toner someone’s position
Toner and McCauley, AI firms mustn’t govern themselves,
The Economist, 26 May 2024; Taylor and Summers reply, 30 May 2024
The Economist, 26 May 2024; Taylor and Summers reply, 30 May 2024
Mira Murati definition
OpenAI leadership announcements, 2022–2024;
Murati deposition, Musk v. OpenAI, May 2026
Murati deposition, Musk v. OpenAI, May 2026
Daniela Amodei someone’s position
Anthropic company page; Stanford GSB interview, June 2026;
Anthropic Series H, reported 28 May 2026
Anthropic Series H, reported 28 May 2026
Kate Crawford theory
Crawford, Atlas of AI, Yale University Press, 2021,
Sally Hacker Prize 2022; Crawford and Joler, Anatomy of an AI System, 2018
Sally Hacker Prize 2022; Crawford and Joler, Anatomy of an AI System, 2018
Daphne Koller measurement
insitro press releases, December 2024 and June 2026;
Koller, McKinsey interview, November 2022
Koller, McKinsey interview, November 2022
Lucy Guo measurement
Forbes profiles of Lucy Guo, April 2025 and July 2026;
Meta investment in Scale AI, announced 13 June 2025
Meta investment in Scale AI, announced 13 June 2025
Andrew Yang someone’s position
Yang, The War on Normal People, Hachette 2018;
US Bureau of Labor Statistics occupational data, 2024 and projections to 2034
US Bureau of Labor Statistics occupational data, 2024 and projections to 2034
MIRI someone’s position
MIRI 2024 Mission and Strategy Update, 4 January 2024;
Yudkowsky and Soares, If Anyone Builds It, Everyone Dies, Little Brown, 16 September 2025
Yudkowsky and Soares, If Anyone Builds It, Everyone Dies, Little Brown, 16 September 2025
Dan Hendrycks measurement
CAIS Statement on AI Risk, 30 May 2023;
Hendrycks, Natural Selection Favors AIs over Humans, arXiv:2303.16200, 2023
Hendrycks, Natural Selection Favors AIs over Humans, arXiv:2303.16200, 2023
OpenAI definition
OpenAI, Our structure, updated October 2025;
Delaware Attorney General, statement on the recapitalisation, 28 October 2025
Delaware Attorney General, statement on the recapitalisation, 28 October 2025
Anthropic definition
Anthropic, The Long-Term Benefit Trust, 19 September 2023;
Anthropic Responsible Scaling Policy, current version as of August 2026
Anthropic Responsible Scaling Policy, current version as of August 2026
Google DeepMind definition
Alphabet announcement of the DeepMind and Brain merger, 20 April 2023;
Google DeepMind Frontier Safety Framework, current version as of August 2026
Google DeepMind Frontier Safety Framework, current version as of August 2026
DeepSeek measurement
DeepSeek-V3 Technical Report, December 2024;
DeepSeek-R1 release, 20 January 2025
DeepSeek-R1 release, 20 January 2025
xAI definition
xAI Risk Management Framework, updated 20 August 2025;
SaferAI Frontier Risk Management Tracker; Future of Life Institute AI Safety Index
SaferAI Frontier Risk Management Tracker; Future of Life Institute AI Safety Index
Apollo Research measurement
Meinke et al., Frontier Models are Capable of In-context Scheming,
Apollo Research, December 2024; Anthropic system cards, 2025
Apollo Research, December 2024; Anthropic system cards, 2025
Cathy O’Neil someone’s position
O’Neil, Weapons of Math Destruction, Crown, 2016;
US Senate HELP Committee, For-Profit Higher Education, 2012
US Senate HELP Committee, For-Profit Higher Education, 2012
Beth Barnes someone’s position
METR, About and Research, metr.org;
Barnes, public remarks on independent evaluation
Barnes, public remarks on independent evaluation
Joy Buolamwini measurement
Buolamwini and Gebru, Gender Shades, Proceedings of
Machine Learning Research, FAT* 2018, pp. 77–91
Machine Learning Research, FAT* 2018, pp. 77–91
Margaret Mitchell definition
Mitchell et al., Model Cards for Model Reporting,
FAT* 2019, pp. 220–229
FAT* 2019, pp. 220–229
Ajeya Cotra theory
Cotra, Forecasting Transformative AI with Biological Anchors,
Open Philanthropy, 2020
Open Philanthropy, 2020
Jade Leung someone’s position
UK AI safety evaluation body, public mandate and reporting;
Centre for the Governance of AI, published research
Centre for the Governance of AI, published research
Tasha McCauley definition
Toner and McCauley, AI firms mustn’t govern themselves,
The Economist, 26 May 2024; OpenAI board announcements, 2018 and 2023
The Economist, 26 May 2024; OpenAI board announcements, 2018 and 2023
Will Saunter definition
BlueDot Impact, About us, bluedot.org;
Tech Can't Save Us podcast, Courses on AI Safety and Biosecurity with Will Saunter, 2026
Tech Can't Save Us podcast, Courses on AI Safety and Biosecurity with Will Saunter, 2026
William Saunders someone’s position
Saunders, Written Testimony, Senate Judiciary Subcommittee on Privacy, Technology and the Law, 17 Sept 2024;
judiciary.senate.gov, hearing transcript
judiciary.senate.gov, hearing transcript
Dewi Erwan definition
BlueDot Impact, About us, bluedot.org;
dewierwan.com;
MATS Program, mentor stream listing
dewierwan.com;
MATS Program, mentor stream listing
Rob Miles definition
Robert Miles AI Safety, YouTube;
aisafety.info, About
aisafety.info, About
Jan Leike someone’s position
Leike, posts on X, 17 May 2024;
CNBC, OpenAI safety leader Jan Leike joins Anthropic, 28 May 2024
CNBC, OpenAI safety leader Jan Leike joins Anthropic, 28 May 2024
Paul Christiano theory
Christiano et al., NeurIPS 2017;
Christiano, public writing on AI risk estimates, LessWrong and the Alignment Forum
Christiano, public writing on AI risk estimates, LessWrong and the Alignment Forum
Chris Olah theory
Anthropic, Transformer Circuits Thread, transformer-circuits.pub;
Olah, public interviews and talks
Olah, public interviews and talks
Max Tegmark someone’s position
Future of Life Institute, Pause Giant AI Experiments, 22 March 2023;
Tegmark, Life 3.0, Knopf, 2017
Tegmark, Life 3.0, Knopf, 2017
Neel Nanda someone’s position
Neel Nanda, neelnanda.io, About;
TransformerLens documentation and public tutorials
TransformerLens documentation and public tutorials
Marius Hobbhahn measurement
Apollo Research, Frontier Models are Capable of In-Context Scheming, arXiv:2412.04984, 2024
Adam Gleave definition
FAR.AI, About, far.ai;
Adam Gleave, gleave.me
Adam Gleave, gleave.me
Jaime Sevilla measurement
Epoch AI, About and Data, epoch.ai;
Jaime Sevilla, team page, epoch.ai/about/team/jaime-sevilla
Jaime Sevilla, team page, epoch.ai/about/team/jaime-sevilla
Stella Biderman someone’s position
EleutherAI, eleuther.ai;
Biderman, written testimony, US Senate, EleutherAI on AI, Innovation, and the Open Source Ecosystem
Biderman, written testimony, US Senate, EleutherAI on AI, Innovation, and the Open Source Ecosystem
Nora Belrose someone’s position
Belrose, public writing and debate, Alignment Forum, EleutherAI blog;
podcast appearances, e.g. AI Development, Safety, and Meaning
podcast appearances, e.g. AI Development, Safety, and Meaning
Sarah Schwettmann definition
Transluce, About, transluce.org;
Schwettmann, public posts on founding Transluce, 2026
Schwettmann, public posts on founding Transluce, 2026
Ryan Greenblatt theory
Greenblatt et al., AI Control: Improving Safety Despite Intentional Subversion, arXiv:2312.06942, 2024
Buck Shlegeris someone’s position
Redwood Research, blog.redwoodresearch.org;
Shlegeris, AI Control: Using Untrusted Systems Safely, 80,000 Hours podcast, 2024
Shlegeris, AI Control: Using Untrusted Systems Safely, 80,000 Hours podcast, 2024
Evan Hubinger theory
Hubinger et al., Risks from Learned Optimization in Advanced Machine Learning Systems, arXiv, 2019;
Anthropic, Alignment Science team overview, anthropic.com
Anthropic, Alignment Science team overview, anthropic.com
Richard Ngo theory
Ngo, AGI Safety From First Principles, Alignment Forum, 2020;
Ngo, Chan and Mindermann, The Alignment Problem from a Deep Learning Perspective, arXiv, 2022
Ngo, Chan and Mindermann, The Alignment Problem from a Deep Learning Perspective, arXiv, 2022
Ethan Perez measurement
Perez et al., Red Teaming Language Models with Language Models, EMNLP, 2022;
Anthropic, Alignment Science publications, anthropic.com
Anthropic, Alignment Science publications, anthropic.com
Kelsey Piper definition
Piper, Leaked OpenAI documents reveal aggressive tactics toward former employees, Vox, 17 May 2024;
Piper, OpenAI reverses controversial exit policy, Vox, 18 May 2024
Piper, OpenAI reverses controversial exit policy, Vox, 18 May 2024
Shakeel Hashim definition
Hashim, Transformer, transformernews.ai, About;
Center for AI Safety, Statement on AI Risk, May 2023
Center for AI Safety, Statement on AI Risk, May 2023
Sholto Douglas someone’s position
Douglas and Jermyn, interview, Dwarkesh Patel podcast, 2024, dwarkeshpatel.com
Dylan Patel measurement
SemiAnalysis, semianalysis.com;
Patel, reporting on chip supply and data centre capacity, 2023 to 2026
Patel, reporting on chip supply and data centre capacity, 2023 to 2026
Charlotte Stix definition
Apollo Research, apolloresearch.ai, team and publications page
AI risks
p(doom) someone’s position
Taboo P(doom), LessWrong 2023;
Grace et al. survey of 2,778 researchers
Grace et al. survey of 2,778 researchers
s-risk theory
Althaus and Gloor, Reducing Risks of
Astronomical Suffering, CLR, 2016
Astronomical Suffering, CLR, 2016
AI and jobs measurement
NBER working paper w34174;
see also the AI Index labour chapter
see also the AI Index labour chapter
existential risk definition
Bostrom, Existential Risk Prevention as Global
Priority, Global Policy 4(1), 2013
Priority, Global Policy 4(1), 2013
AI psychosis measurement
Olisaeloka et al., BJPsych Open 12(4), 11 June 2026;
OpenAI, Strengthening ChatGPT’s responses in sensitive conversations, 27 October 2025
OpenAI, Strengthening ChatGPT’s responses in sensitive conversations, 27 October 2025
permanent disempowerment theory
Kulveit, Douglas, Ammann, Turan, Krueger and Duvenaud,
Gradual Disempowerment, arXiv:2501.16946, January 2025; ICML 2025
Gradual Disempowerment, arXiv:2501.16946, January 2025; ICML 2025
AI drones measurement
UN Panel of Experts on Libya, S/2021/229, March 2021;
US DoD Directive 3000.09, 25 January 2023; UNGA resolution 80/57, 5 December 2025
US DoD Directive 3000.09, 25 January 2023; UNGA resolution 80/57, 5 December 2025
automation theory
Acemoglu, The Simple Macroeconomics of AI, NBER working paper 32487, 2024;
Autor, The Work of the Future, MIT, 2020
Autor, The Work of the Future, MIT, 2020
the bio threshold measurement
Anthropic Responsible Scaling Policy and ASL-3 activation, May 2025;
OpenAI Preparedness Framework v2, April 2025; RAND RRA2977-2, January 2024
OpenAI Preparedness Framework v2, April 2025; RAND RRA2977-2, January 2024
AI 2027 someone’s position
Kokotajlo, Lifland, Larsen and Dean, AI 2027, AI Futures Project, 3 April 2025;
Clarifying how our AI timelines have changed, January 2026
Clarifying how our AI timelines have changed, January 2026
concentration of power theory
Davidson, Finnveden and Hadshar, AI-Enabled Coups, Forethought, April 2025;
International AI Safety Report 2026, chaired by Bengio, February 2026
International AI Safety Report 2026, chaired by Bengio, February 2026
synthetic media measurement
Stockwell et al., AI-Enabled Influence Operations, CETaS, Alan Turing Institute, September 2024;
Schiff, Schiff and Bueno, American Political Science Review 119(1), 2025
Schiff, Schiff and Bueno, American Political Science Review 119(1), 2025
persuasion measurement
Salvi, Horta Ribeiro, Gallotti and West, On the conversational persuasiveness
of large language models, Nature Human Behaviour 9(8), 2025
of large language models, Nature Human Behaviour 9(8), 2025
stolen weights measurement
Nevo, Lahav, Karpur, Bar-On, Bradley and Alstott, Securing AI Model Weights,
RAND RR-A2849-1, 30 May 2024
RAND RR-A2849-1, 30 May 2024
the environmental cost measurement
Kay, Reaser and Taylor, Dallas Fed working paper 2606, March 2026;
IEA Key Questions on Energy and AI, 2026; Google and Microsoft environmental reports, 2026
IEA Key Questions on Energy and AI, 2026; Google and Microsoft environmental reports, 2026
learning measurement
Bastani, Bastani, Sungu, Ge, Kabakcı and Mariman,
Generative AI Can Harm Learning, SSRN working paper, July 2024
Generative AI Can Harm Learning, SSRN working paper, July 2024
AI companions measurement
Common Sense Media, Talk, Trust and Trade-Offs, July 2025;
Fang et al., MIT Media Lab and OpenAI, arXiv:2503.17473, March 2025
Fang et al., MIT Media Lab and OpenAI, arXiv:2503.17473, March 2025
value lock-in theory
MacAskill, What We Owe the Future, 2022, chapter 4;
Ord, The Precipice, 2020; Finnveden, Riedel and Shulman, AGI and Lock-in, 2023
Ord, The Precipice, 2020; Finnveden, Riedel and Shulman, AGI and Lock-in, 2023
malicious use measurement
Anthropic, Detecting and countering misuse of AI, August 2025;
Google Threat Intelligence Group, Adversarial Misuse of Generative AI, January 2025
Google Threat Intelligence Group, Adversarial Misuse of Generative AI, January 2025
manipulation measurement
De Freitas et al., Emotional Manipulation by AI Companions,
HBS working paper 26-005, 2025;
Morrin et al., Technological folie à deux, Nature Mental Health, 2026
HBS working paper 26-005, 2025;
Morrin et al., Technological folie à deux, Nature Mental Health, 2026
shut it down someone’s position
Yudkowsky and Soares, If Anyone Builds It, Everyone Dies,
Little Brown, 16 September 2025, and ifanyonebuildsit.com;
Collier, More Was Possible, Asterisk 11
Little Brown, 16 September 2025, and ifanyonebuildsit.com;
Collier, More Was Possible, Asterisk 11
phishing and hacking measurement
Heiding, Schneier, Vishwanath, Bernstein and Park,
IEEE Access 12, 2024;
Czybik et al., USENIX Security 2026
IEEE Access 12, 2024;
Czybik et al., USENIX Security 2026
Full bibliographies
The long pieces carry their sources as linked lists, with a note on what each one is good for and which figures go stale fastest.