Bibliography

Every source cited in the AI Safety Atlas.

912 sources

  1. Abid, A., Farooqi, M. & Zou, J. (2021). Persistent Anti-Muslim Bias in Large Language Models. arXiv.
  2. Abraham (2024). ‘Lavender’: The AI machine directing Israel’s bombing spree in Gaza. +972 Magazine.
  3. Adam Scherlis (2023). Inner Misalignment in "Simulator" LLMs. AI Alignment Forum.
  4. adamShimi (2020). Goal-directedness is behavioral, not structural. AI Alignment Forum.
  5. Adan, S. N. et al. (2024). Voice and Access in AI: Global AI Majority Participation in Artificial Intelligence Development and Governance.
  6. Aguirre (2025). keepthefuturehuman.com/essay.
  7. Aguirre, A. (2025). Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead.
  8. AI Digest (2023). How fast is AI improving?. AI Digest.
  9. AI Digest (2024). AIs are becoming more self-aware. Here's why that matters. AI Digest.
  10. AI Impacts (2022). Survey of 2,778 AI authors: six parts in pictures.
  11. AI Incident Database (2025). Welcome to the Artificial Intelligence Incident Database.
  12. AI Optimists (2023). AI is easy to control. AI Optimism.
  13. AI Safety in China (2025). AI Safety in China: 2024 in Review.
  14. AISI (2025). How we’re addressing the gap between AI capabilities and mitigations. AI Security Institute.
  15. Ajeya Cotra (2020). Draft report on AI timelines. AI Alignment Forum.
  16. Ajeya Cotra (2021). The case for aligning narrowly superhuman models. AI Alignment Forum.
  17. Ajeya Cotra (2022). Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover. AI Alignment Forum.
  18. Akbir Khan et al. (2024). Debating with More Persuasive LLMs Leads to More Truthful Answers. arXiv.
  19. Alex Flint (2022). Notes on OpenAI’s alignment plan. AI Alignment Forum.
  20. Allan Dafoe (2020). AI Governance: Opportunity and Theory of Impact.
  21. Althaus & Gloor (2016). Reducing Risks of Astronomical Suffering: A Neglected Priority. Center on Long-Term Risk.
  22. Altman (2023). Planning for AGI and beyond. OpenAI.
  23. Altman to Gates: "Multimodality will be important".
  24. Amazon (2023). 4 cool facts about Hercules, the small-but-mighty robot in Amazon’s fulfillment centers. Amazon News.
  25. Amazon (2024). Amazon uses robots that sort, lift, and carry packages—see them in action. Amazon News.
  26. Amodei & Clark (2016). Faulty reward functions in the wild. OpenAI.
  27. Anderljung, M. & Hazell, J. (2023). Protecting Society from AI Misuse: When are Restrictions on Capabilities Warranted?. arXiv.
  28. Anderljung, M. et al. (2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv.
  29. Anderljung, M. et al. (2023). Towards Publicly Accountable Frontier LLMs: Building an External Scrutiny Ecosystem under the ASPIRE Framework. arXiv.
  30. Andreas, J. (2022). Language Models as Agent Models. arXiv.
  31. Andrei Potlogea & Anson Ho (2025). AI and explosive growth redux.
  32. Andrew_Critch (2021). What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs). AI Alignment Forum.
  33. Andrew_Critch (2022). Pivotal outcomes and pivotal processes. AI Alignment Forum.
  34. Andrew_Critch (2023). Consciousness as a conflationary alliance term for intrinsically valued internal experiences. AI Alignment Forum.
  35. Andy K. Zhang et al. (2024). Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models. arXiv.
  36. Andy Zou et al. (2024). Improving Alignment and Robustness with Circuit Breakers. arXiv.
  37. Anna Desmarais (2024). Découvrez Daisy, le chatbot "mamie" qui fait perdre du temps aux fraudeurs au téléphone. euronews.
  38. Anson Ho & Arden Berg (2025). Do the biorisk evaluations of AI labs actually measure the risk of developing bioweapons?.
  39. Anson Ho, Yafah Edelman, Josh You & Jean-Stanislas Denain (2025). Is almost everyone wrong about America’s AI power problem?.
  40. Anthropic (2023). Anthropic's core views on AI safety.
  41. Anthropic (2023). Anthropic's Responsible Scaling Policy.
  42. Anthropic (2023). Claude's constitution.
  43. Anthropic (2023). Model-Card-Claude-2.
  44. Anthropic (2023). Superposition, Memorization, and Double Descent.
  45. Anthropic (2024). A new initiative for developing third-party model evaluations.
  46. Anthropic (2024). Alignment faking in large language models.
  47. Anthropic (2024). Challenges in evaluating AI systems.
  48. Anthropic (2024). Company.
  49. Anthropic (2024). Home \ Anthropic.
  50. Anthropic (2024). Introducing Anthropic's Responsible Scaling Policy.
  51. Anthropic (2024). Introducing Claude 3.5 Sonnet.
  52. Anthropic (2024). Introducing the Model Context Protocol.
  53. Anthropic (2024). Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet.
  54. Anthropic (2024). Simple probes can catch sleeper agents.
  55. Anthropic (2025). Agentic misalignment: How LLMs could be insider threats.
  56. Anthropic (2025). Claude Opus.
  57. Anthropic (2025). Donating MCP to the Agentic AI Foundation.
  58. Anthropic (2025). Introducing advanced tool use on the Claude Developer Platform.
  59. Anthropic (2025). Reasoning models don't always say what they think.
  60. Anthropic (2025). Tracing the thoughts of a large language model.
  61. AP News (2017). Putin: Leader in artificial intelligence will rule world. AP News.
  62. Apollo Research (2024). A Starter Guide For Evals.
  63. Apvrille, A. & Nakov, D. (2025). Malware analysis assisted by AI with R2AI. arXiv.
  64. ARC-AGI (2024). ARC-AGI-1. ARC Prize.
  65. Arthur Goemans et al. (2024). Safety Case Template for Frontier AI: A Cyber Inability Argument.
  66. ArtificialAnalysis (2025). Artificial Analysis Intelligence Index v4.3.2. Artificial Analysis.
  67. Arvind Narayanan & Sayash Kapoor (2024). AI existential risk probabilities are too unreliable to inform policy.
  68. Aschenbrenner (2024). I. From GPT-4 to AGI: Counting the OOMs. SITUATIONAL AWARENESS - The Decade Ahead.
  69. Aschenbrenner (2024). Introduction. SITUATIONAL AWARENESS - The Decade Ahead.
  70. Asher Brass (2025). Location Verification for AI Chips.
  71. Askell, A., Brundage, M. & Hadfield, G. (2019). The Role of Cooperation in Responsible AI Development. arXiv.
  72. AXRP (2024). 27 - AI Control with Buck Shlegeris and Ryan Greenblatt.
  73. Badie et al. (2011). Sage Reference - International Encyclopedia of Political Science - Stages Model of Policy Making.
  74. Bai, Y. et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv.
  75. Baker, B. et al. (2019). Emergent Tool Use From Multi-Agent Autocurricula. arXiv.
  76. Baker, B. et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv.
  77. Bakhtin, A. et al. (2022). Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning. arXiv.
  78. Bansal, H., Yin, D., Monajatipoor, M. & Chang, K. (2022). How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?. arXiv.
  79. Barnett, P. & Scher, A. (2025). AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions.
  80. Barnett, P. & Thiergart, L. (2024). What AI evaluations for preventing catastrophic risks can and cannot do. arXiv.
  81. Baumann (2017). S-risks: An introduction. Center for Reducing Suffering.
  82. BBC (2023). ChatGPT banned in Italy over privacy concerns.
  83. Belfield & Hua (2022). Compute and Antitrust. Verfassungsblog.
  84. Ben Pace (2020). What Failure Looks Like: Distilling the Discussion. AI Alignment Forum.
  85. Bengio (2023). Yoshua Bengio | FAQ on Catastrophic AI Risks.
  86. Bengio, Y. et al. (2023). Managing extreme AI risks amid rapid progress. arXiv.
  87. Bengio, Y. et al. (2025). International AI Safety Report. arXiv.
  88. Beraja, M., Kao, A., Yang, D. Y. & Yuchtman, N. (2023). AI-tocracy. The Quarterly Journal of Economics.
  89. beren (2023). Gradient hacking is extremely difficult. AI Alignment Forum.
  90. Berglund, L. et al. (2023). Taken out of context: On measuring situational awareness in LLMs. arXiv.
  91. Berglund, L. et al. (2023). The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A". arXiv.
  92. Bernardi (2024). A Policy Agenda for Defensive Acceleration Against AI Risks.
  93. Besiroglu et al. (2024). FrontierMath: Evaluating advanced mathematical reasoning in AI. Epoch AI.
  94. Besiroglu, T., Bergerson, S. A., Michael, A., Heim, L., Luo, X. & Thompson, N. (2024). The Compute Divide in Machine Learning: A Threat to Academic Contribution and Scrutiny?. arXiv.
  95. Besta, M. et al. (2023). Graph of Thoughts: Solving Elaborate Problems with Large Language Models. arXiv.
  96. Beth Barnes & paulfchristiano (2020). Writeup: Progress on AI Safety via Debate. LessWrong.
  97. Beth Barnes (2020). Debate update: Obfuscated arguments problem. LessWrong.
  98. Beth Barnes (2022). 'simulator' framing and confusions about LLMs. AI Alignment Forum.
  99. Beth Barnes (2023). New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks.
  100. Beth Barnes (2023). Update on ARC's recent eval efforts.
  101. Beth Barnes, Hjalmar Wijk & Lawrence Chan (2023). Responsible Scaling Policies (RSPs).
  102. Betker, J. et al. (2023). Improving Image Generation with Better Captions.
  103. Betley, J. et al. (2025). Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs. arXiv.
  104. Bhatt et al. (2023). Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models. arXiv.org.
  105. Bhatt et al. (2024). Shell Games: Control Protocols for Adversarial AI Agents. OpenReview.
  106. Bhatt, M. et al. (2024). CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models. arXiv.
  107. Binder, F. J. et al. (2024). Looking Inward: Language Models Can Learn About Themselves by Introspection. arXiv.
  108. Black, S. et al. (2025). RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents. arXiv.
  109. Bode, I. & Watts, T. (2023). Loitering Munitions and Unpredictability: Autonomy in Weapon Systems and Challenges to Human Control.
  110. Bogdan Ionut Cirstea (2023). AISC project: How promising is automating alignment research? (literature review). LessWrong.
  111. Bogdan, P. C., Macar, U., Nanda, N. & Conmy, A. (2025). Thought Anchors: Which LLM Reasoning Steps Matter?. arXiv.
  112. Boiko, D. A., MacKnight, R. & Gomes, G. (2023). Emergent autonomous scientific research capabilities of large language models. arXiv.
  113. Bommasani, R. et al. (2021). On the Opportunities and Risks of Foundation Models. arXiv.
  114. Bondarenko, A., Volk, D., Volkov, D. & Ladish, J. (2025). Demonstrating specification gaming in reasoning models. arXiv.
  115. Boston Dynamics (2024). Atlas Goes Hands On. Boston Dynamics.
  116. Boston Dynamics (2024). Stretch - Mobile Warehouse Robots. Boston Dynamics.
  117. Bostrom (2002). Existential Risks: Analyzing Human Extinction Scenarios and Related Hazards.
  118. Bostrom (2012). Existential Risks: Threats to Humanity’s Survival.
  119. Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies.
  120. Boursier, E. & Flammarion, N. (2024). Simplicity bias and optimization threshold in two-layer ReLU networks. arXiv.
  121. Bowman (2024). [External] 2024 Debate Agenda Writeup. Google Docs.
  122. Bowman, S. R. (2023). Eight Things to Know about Large Language Models. arXiv.
  123. Bowman, S. R. et al. (2022). Measuring Progress on Scalable Oversight for Large Language Models. arXiv.
  124. Bradford (2020). The Brussels Effect: How the European Union Rules the World. Scholarship Archive.
  125. Branwen (2020). The Scaling Hypothesis.
  126. Brennan et al. (2025). Artificial Power: 2025 Landscape Report. AI Now Institute.
  127. Brown (2024). Noam Brown (@polynoamial) on X. X (formerly Twitter).
  128. Brown, T. B. et al. (2020). Language Models are Few-Shot Learners. arXiv.
  129. Brundage (2025). Feedback on the Second Draft of the General-Purpose AI Code of Practice.
  130. Brundage, M. et al. (2018). The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation. arXiv.
  131. Brunskill (2022). CS234: Reinforcement Learning Winter 2022.
  132. Bubeck, S. et al. (2023). Sparks of Artificial General Intelligence: Early experiments with GPT-4. arXiv.
  133. Buchanan (2020). The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology.
  134. Buck (2022). The prototypical catastrophic AI action is getting root access to its datacenter. AI Alignment Forum.
  135. Buck (2024). Access to powerful AI might make computer security radically easier. AI Alignment Forum.
  136. Buhl, M. D., Sett, G., Koessler, L., Schuett, J. & Anderljung, M. (2024). Safety cases for frontier AI. arXiv.
  137. Burden, J. (2024). Evaluating AI Evaluation: Perils and Prospects. arXiv.
  138. Burns, C. et al. (2023). Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. arXiv.
  139. Buterin (2023). My techno-optimism.
  140. Buterin (2025). d/acc: one year later.
  141. Caballero, E., Gupta, K., Rish, I. & Krueger, D. (2022). Broken Neural Scaling Laws. arXiv.
  142. CAIS (2023). Statement on AI Extinction Risk | CAIS. Center for AI Safety.
  143. Campos, S., Papadatos, H., Roger, F., Touzet, C., Quarks, O. & Murray, M. (2025). A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management. arXiv.
  144. Cao, B. et al. (2024). Towards Scalable Automated Alignment of LLMs: A Survey. arXiv.
  145. Cao, K. et al. (2023). Large-scale pancreatic cancer detection via non-contrast CT and deep learning. Nature Medicine.
  146. Carlini, N. et al. (2020). Extracting Training Data from Large Language Models. arXiv.
  147. Carlsmith (2020). How Much Computational Power Does It Take to Match the Human Brain?. Coefficient Giving.
  148. Carlsmith, J. (2022). Is Power-Seeking AI an Existential Risk?. arXiv.
  149. Carlsmith, J. (2023). Scheming AIs: Will AIs fake alignment during training in order to get power?. arXiv.
  150. Carlson, R. (2009). The changing economics of DNA synthesis. Nature Biotechnology.
  151. Carter, S. R., Yassif, J. M. & Isaac, C. R. (2023). Benchtop DNA Synthesis Devices: Capabilities, Biosecurity Implications, and Governance.
  152. Carvalho, B. W., Garcez, A. S. D., Lamb, L. C. & Brazil, E. V. (2025). Grokking Explained: A Statistical Phenomenon. arXiv.
  153. Casper, S. et al. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. arXiv.
  154. Casper, S. et al. (2024). Black-Box Access is Insufficient for Rigorous AI Audits. arXiv.
  155. Casper, S., Krueger, D. & Hadfield-Menell, D. (2025). Pitfalls of Evidence-Based AI Policy. arXiv.
  156. Caucheteux, C. & King, J. (2022). Brains and algorithms partially converge in natural language processing. Communications Biology.
  157. Cave, S. & ÓhÉigeartaigh, S. S. (2018). An AI Race for Strategic Advantage. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society.
  158. Cerutti et al. (2025). The Global Impact of AI – Mind the Gap. IMF eLibrary.
  159. Ceruzzi (1989). Beyond The Limits. MIT Press.
  160. Cha, S. (2024). Towards an international regulatory framework for AI safety: lessons from the IAEA’s nuclear safety regulations. Humanities and Social Sciences Communications.
  161. Chan, A. et al. (2024). IDs for AI Systems. arXiv.
  162. Chan, A. et al. (2024). Visibility into AI Agents. arXiv.
  163. Chandra et al. (2024). Reducing Risks Posed by Synthetic Content An Overview of Technical Approaches to Digital Content Transparency. NIST.
  164. Chang, C. (2024). The First Global AI Treaty: Analyzing the Framework Convention on Artificial Intelligence and the EU AI Act. University of Illinois Law Review (Online).
  165. Charbel-Raphaël & cozyfractal (2024). What convincing warning shot could help prevent extinction from AI?. LessWrong.
  166. Charbel-Raphaël & Épiphanie Gédéon (2024). We might be dropping the ball on Autonomous Replication and Adaptation. AI Alignment Forum.
  167. Charbel-Raphaël & Gabin (2023). Davidad's Bold Plan for Alignment: An In-Depth Explanation. LessWrong.
  168. Charbel-Raphaël (2025). Comment on “johnswentworth's Shortform”. LessWrong.
  169. Charbel-Raphaël (2025). Comment on “johnswentworth's Shortform”. LessWrong.
  170. Charlie Steiner (2022). Take 2: Building tools to help build FAI is a legitimate strategy, but it's dual-use. AI Alignment Forum.
  171. Charlotte Siegmann & Markus Anderljung (2022). The Brussels Effect and Artificial Intelligence.
  172. Charvet, C. J. (2021). Cutting across structural and transcriptomic scales translates time across the lifespan in humans and chimpanzees. Proceedings of the Royal Society B: Biological Sciences.
  173. Cheerla (2018). AlphaZero Explained. On AI.
  174. Chehoudi, R. (2025). Artificial intelligence and democracy: pathway to progress or decline?. Journal of Information Technology & Politics.
  175. Chen, M. et al. (2021). Evaluating Large Language Models Trained on Code. arXiv.
  176. Chen, R., Arditi, A., Sleight, H., Evans, O. & Lindsey, J. (2025). Persona Vectors: Monitoring and Controlling Character Traits in Language Models. arXiv.
  177. Chen, S., Zharmagambetov, A., Mahloujifar, S., Chaudhuri, K., Wagner, D. & Guo, C. (2024). SecAlign: Defending Against Prompt Injection with Preference Optimization. arXiv.
  178. Chen, X., Liu, C., Li, B., Lu, K. & Song, D. (2017). Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning. arXiv.
  179. Chen, Y. et al. (2025). Reasoning Models Don't Always Say What They Think. arXiv.
  180. Chen, Y., Shen, C., Shen, Y., Wang, C. & Zhang, Y. (2022). Amplifying Membership Exposure via Data Poisoning. arXiv.
  181. Cheng (2024). AI Incident Reporting. Convergence Analysis.
  182. Cheng et al. (2024). State of the AI Regulatory Landscape. Convergence Analysis.
  183. Chess.com (2014). Chess.com - Play Chess Online - Free Games. Chess.com.
  184. Chollet (2023). François Chollet (@fchollet) on X. X (formerly Twitter).
  185. Chollet (2024). Francois Chollet, Mike Knoop - LLMs won’t lead to AGI - $1,000,000 Prize to find true solution.
  186. Chollet, F. (2019). On the Measure of Intelligence. arXiv.
  187. Chollet, F., Knoop, M., Kamradt, G. & Landers, B. (2024). ARC Prize 2024: Technical Report. arXiv.
  188. Christiano (2016). Reliability amplification. Medium.
  189. Christiano (2016). Security amplification. Medium.
  190. Christiano (2018). Takeoff speeds. The sideways view.
  191. Christiano (2019). Paul Christiano: Current Work in AI Alignment. Effective Altruism.
  192. Christiano Paul (2017). Capability amplification. Medium.
  193. Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S. & Amodei, D. (2017). Deep reinforcement learning from human preferences. arXiv.
  194. Cihon, P., Maas, M. M. & Kemp, L. (2020). Should Artificial Intelligence Governance be Centralised? Design Lessons from History. arXiv.
  195. Cima, M., Tonnaer, F. & Hauser, M. D. (2010). Psychopaths know right from wrong but don’t care. Social Cognitive and Affective Neuroscience.
  196. CISA (2021). The Attack on Colonial Pipeline: What We’ve Learned & What We’ve Done Over the Past Two Years. Cybersecurity and Infrastructure Security Agency CISA.
  197. CISA (2024). CISA and Partners Release Advisory on PRC-sponsored Volt Typhoon Activity and Supplemental Living Off the Land Guidance. Cybersecurity and Infrastructure Security Agency CISA.
  198. Cleo Nardo (2023). The Waluigi Effect (mega-post). AI Alignment Forum.
  199. Clymer, J., Gabrieli, N., Krueger, D. & Larsen, T. (2024). Safety Cases: How to Justify the Safety of Advanced AI Systems. arXiv.
  200. Clymer, J., Juang, C. & Field, S. (2024). Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals. arXiv.
  201. Cobbe, K. et al. (2021). Training Verifiers to Solve Math Word Problems. arXiv.
  202. Cobbe, K., Klimov, O., Hesse, C., Kim, T. & Schulman, J. (2018). Quantifying Generalization in Reinforcement Learning. arXiv.
  203. Conn (2015). Existential Risk. Future of Life Institute.
  204. Connor Leahy & Gabriel Alfour (2023). Cognitive Emulation: A Naive AI Safety Proposal. AI Alignment Forum.
  205. Control AI (2024). Deepfakes Policy. ControlAI.
  206. Corwin (2002). SL4: AI Boxing.
  207. Cotra (2021). Why AI alignment could be hard with modern deep learning. Cold Takes.
  208. Cotra (2023). AIs accelerating AI research.
  209. Cotra (2023). Language models surprised us.
  210. Cotra (2023). Scale, schlep, and systems.
  211. Cottier et al. (2024). How much does it cost to train frontier AI models?. Epoch AI.
  212. Creswell, A., Shanahan, M. & Higgins, I. (2022). Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning. arXiv.
  213. Critch, A. & Russell, S. (2023). TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI. arXiv.
  214. CrowdStrike (2024). External Technical Root Cause Analysis — Channel File 291.
  215. Cullen O’Keefe, Peter Cihon, Ben Garfinkel, Carrick Flynn, Jade Leung & and Allan Dafoe (2020). The Windfall Clause: Distributing the Benefits of AI for the Common Good.
  216. Cunha, P. R. & Estima, J. (2023). Navigating the Landscape of AI Ethics and Responsibility. Lecture Notes in Computer Science.
  217. Cunningham, H., Ewart, A., Riggs, L., Huben, R. & Sharkey, L. (2023). Sparse Autoencoders Find Highly Interpretable Features in Language Models. arXiv.
  218. Dafoe (2017). GovAI-Research-Agenda.
  219. Dafoe (2018). AI Governance: A Research Agenda | GovAI.
  220. Dafoe, A. (2024). AI Governance: Overview and Theoretical Lenses. The Oxford Handbook of AI Governance.
  221. Dalrymple (2022). Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems. arXiv.org.
  222. Daniel Kokotajlo (2021). Interlude: Agents as Automobiles. AI Alignment Forum.
  223. Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland & Romeo Dean (2025). AI 2027.
  224. Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland & Romeo Dean (2025). Security Forecast. AI 2027.
  225. David Manheim & Scott Garrabrant (2018). Categorizing Variants of Goodhart's Law. arXiv.
  226. David Rein et al. (2023). GPQA: A Graduate-Level Google-Proof Q&A Benchmark. arXiv.
  227. Davidson (2024). What a Compute-Centric Framework Says About Takeoff Speeds. Coefficient Giving.
  228. Davidson, T., Denain, J., Villalobos, P. & Bas, G. (2023). AI capabilities can be significantly improved without expensive retraining. arXiv.
  229. DavidW (2023). Deceptive Alignment is <1% Likely by Default. AI Alignment Forum.
  230. DeepMind (2016). AlphaGo. Google DeepMind.
  231. DeepMind (2018). AlphaZero: Shedding new light on chess, shogi, and Go. Google DeepMind.
  232. DeepMind (2018). Scalable agent alignment via reward modeling. Medium.
  233. DeepMind (2023). FunSearch: Making new discoveries in mathematical sciences using Large Language Models. Google DeepMind.
  234. DeepMind (2023). Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals. Medium.
  235. DeepMind (2024). Exploring institutions for global AI governance. Google DeepMind.
  236. DeepMind (2024). How AlphaChip transformed computer chip design. Google DeepMind.
  237. DeepMind (2024). Introducing the Frontier Safety Framework. Google DeepMind.
  238. DeepMind (2025). AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. Google DeepMind.
  239. DeepSeek (2025). GitHub - deepseek-ai/DeepSeek-V3. GitHub.
  240. DeepSeek-AI et al. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv.
  241. Delétang, G. et al. (2022). Neural Networks and the Chomsky Hierarchy. arXiv.
  242. Deng (2009). ImageNet: A large-scale hierarchical image database.
  243. Dewey, D. (2011). Learning What to Value. Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings.
  244. Dherin, B., Munn, M., Rosca, M. & Barrett, D. G. T. (2022). Why neural networks find simple solutions: the many regularizers of geometric complexity. arXiv.
  245. DiGiovanni (2023). Beginner’s guide to reducing s-risks. Center on Long-Term Risk.
  246. Ding, J. & Dafoe, A. (2020). The Logic of Strategic Assets: From Oil to Artificial Intelligence. arXiv.
  247. Ding, J. (2018). Deciphering China's AI Dream: The Context, Components, Capabilities, and Consequences of China's Strategy to Lead the World in AI.
  248. DnaScript (2024). Home. DNA Script.
  249. Dong, Y. et al. (2024). Safeguarding Large Language Models: A Survey. arXiv.
  250. Douillard et al (2023). AI Safety. Tigera – Creator of Calico.
  251. Douillard et al (2024). DiPaCo: Distributed Path Composition. arXiv.org.
  252. Douillard, A. et al. (2023). DiLoCo: Distributed Low-Communication Training of Language Models. arXiv.
  253. Dowie (1977). Pinto Madness: the Ford Pinto’s fire-prone gas tank.
  254. Dwarkesh Patel (2023). Dario Amodei (Anthropic CEO) - Scaling, Alignment, & AI Progress.
  255. EA Global (2020). Paul Christiano: Current work in AI alignment. EA Forum.
  256. Egan, J. & Heim, L. (2023). Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers. arXiv.
  257. Ege Erdil & Matthew Barnett (2025). Most AI value will come from broad automation, not from R&D.
  258. Eiras, F. et al. (2024). Near to Mid-term Risks and Opportunities of Open-Source Generative AI. arXiv.
  259. El-Mhamdi, E. et al. (2022). On the Impossible Safety of Large AI Models. arXiv.
  260. Eliezer Yudkowsky (2007). Politics is the Mind-Killer. LessWrong.
  261. Eliezer Yudkowsky (2008). Shut up and do the impossible!. LessWrong.
  262. Eliezer Yudkowsky (2022). AGI Ruin: A List of Lethalities. AI Alignment Forum.
  263. Emily H. Soice, Rafael Rocha, Kimberlee Cordova, Michael Specter & Kevin M. Esvelt (2023). Can large language models democratize access to dual-use biotechnology?. arXiv.
  264. Encyclopedia Britannica (2025). Military technology - Castles, Fortifications, Defense. Encyclopedia Britannica.
  265. Epicural (2021). Goodhart’s Law. Epicural, LLC - Change is hard. The Epicural Team “does hard.”.
  266. Epoch AI (2025). Data on AI Data Centers. Epoch AI.
  267. Epoch AI (2025). Data on AI Models. Epoch AI.
  268. Epoch AI (2025). Data on Machine Learning Hardware. Epoch AI.
  269. Epoch AI (2025). GATE Model Playground. Epoch AI.
  270. Epoch AI (2025). Trends in Artificial Intelligence. Epoch AI.
  271. EpochAI (2024). FrontierMath: LLM Benchmark for Advanced AI Math Reasoning. Epoch AI.
  272. EpochAI (2025). Could decentralized training solve AI’s power problem?. Epoch AI.
  273. EpochAI (2025). Epoch Capabilities Index. Epoch AI.
  274. EpochAI (2025). FrontierMath Sample Problems. Epoch AI.
  275. Erdil, E. & Besiroglu, T. (2023). Explosive growth from AI automation: A review of the arguments. arXiv.
  276. Erdil, E. et al. (2025). GATE: An Integrated Assessment Model for AI Automation. arXiv.
  277. Erich Grunewald (2023). Introduction to AI Chip Making in China.
  278. Esvelt, K. M. (2022). Delay, Detect, Defend: Preparing for a Future in which Thousands Can Release New Pandemics.
  279. Ethan Perez et al. (2022). Discovering Language Model Behaviors with Model-Written Evaluations. arXiv.
  280. Ethayarajh, K. & Jurafsky, D. (2022). The Authenticity Gap in Human Evaluation. arXiv.
  281. EU Commission (2025). New JRC collection of external scientific reports to inform the implementation of the EU AI Act on general-purpose AI models. AI Watch.
  282. European Commission (2024). The Act Texts. EU Artificial Intelligence Act.
  283. Evan Hubinger et al. (2024). Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. arXiv.
  284. Everitt, T. et al. (2025). Evaluating the Goal-Directedness of Large Language Models. arXiv.
  285. evhub (2020). AI safety via market making. LessWrong.
  286. evhub (2020). Homogeneity vs. heterogeneity in AI takeoff scenarios. AI Alignment Forum.
  287. evhub (2022). How likely is deceptive alignment?. AI Alignment Forum.
  288. evhub (2023). Towards understanding-based safety evaluations. AI Alignment Forum.
  289. evhub (2023). When can we trust model evaluations?. AI Alignment Forum.
  290. evhub, Chris van Merwijk, Vlad Mikulik, Joar Skalse & Scott Garrabrant (2019). Conditions for Mesa-Optimization. AI Alignment Forum.
  291. evhub, Nicholas Schiefer, Carson Denison & Ethan Perez (2023). Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research. AI Alignment Forum.
  292. Ewing (2017). Engineering a Deception: What Led to Volkswagen’s Diesel Scandal (Published 2021).
  293. Exfilbench (2025). ExfilBench | Exfiltration & Replication Benchmark. ExfilBench.
  294. Eykholt, K. et al. (2017). Robust Physical-World Attacks on Deep Learning Models. arXiv.
  295. Fabien Roger & Buck (2024). Toy models of AI control for concentrated catastrophe prevention. AI Alignment Forum.
  296. Fabien Roger & ryan_greenblatt (2023). Preventing Language Models from hiding their reasoning. AI Alignment Forum.
  297. Fabien Roger (2023). Coup probes: Catching catastrophes with probes trained off-policy. AI Alignment Forum.
  298. Fabien Roger (2023). The Translucent Thoughts Hypotheses and Their Implications. AI Alignment Forum.
  299. Faggella (2023). A Worthy Successor - The Purpose of AGI - Daniel Faggella.
  300. Fang, R., Bindu, R., Gupta, A., Zhan, Q. & Kang, D. (2024). LLM Agents can Autonomously Hack Websites. arXiv.
  301. FAR․AI (2024). Richard Ngo – Reframing AGI Threat Models [Alignment Workshop]. YouTube.
  302. Farrell (2024). Learning from History: GPAI serious incident reporting. Pour Demain.
  303. Federspiel, F., Mitchell, R., Asokan, A., Umana, C. & McCoy, D. (2023). Threats by artificial intelligence to human health and human existence. BMJ Global Health.
  304. Feldstein, S. (2019). The Global Expansion of AI Surveillance.
  305. Feng, S. & Tramèr, F. (2024). Privacy Backdoors: Stealing Data with Corrupted Pretrained Models. arXiv.
  306. Fenwick (2023). Want to shape AI policy? Consider working in the US government. 80,000 Hours.
  307. Field (2025). Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts. arXiv.org.
  308. Flanagan, D. P. & Dixon, S. G. (2014). The Cattell‐Horn‐Carroll Theory of Cognitive Abilities. Encyclopedia of Special Education.
  309. Fluri, L., Paleka, D. & Tramèr, F. (2023). Evaluating Superhuman Models with Consistency Checks. arXiv.
  310. Friedman, M.. The Social Responsibility of Business Is to Increase Its Profits. Corporate Ethics and Corporate Governance.
  311. Friston, K. J. et al. (2022). Designing Ecosystems of Intelligence from First Principles. arXiv.
  312. Fu, Z., Zhao, T. Z. & Finn, C. (2024). Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. arXiv.
  313. Future of Life Institute (2024). Artificial Escalation. YouTube.
  314. Future of Life Institute (2024). Gradual AI Disempowerment. Future of Life Institute.
  315. Future of Life Institute (2024). Holly Elmore on Pausing AI, Hardware Overhang, Safety Research, and Protesting. YouTube.
  316. Future of Life Institute (2025). AI Safety Index, Summer 2025.
  317. Future of Life Institute (2025). Can Defense in Depth Work for AI? (with Adam Gleave). YouTube.
  318. Gabriel, I. et al. (2024). The Ethics of Advanced AI Assistants. arXiv.
  319. Game Thinking TV (2023). Gödel, Escher, Bach author Doug Hofstadter on the state of AI today. YouTube.
  320. Ganguli, D. et al. (2022). Predictability and Surprise in Large Generative Models. arXiv.
  321. Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A. & Brendel, W. (2018). ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. arXiv.
  322. Gennari et al. (2024). Considerations for Evaluating Large Language Models for Cybersecurity Tasks | CMU Software Engineering Institute. SEI Digital Library.
  323. Giattino et al. (2023). Artificial Intelligence. Our World in Data.
  324. Giattino et al. (2023). Language-based AI systems have grown rapidly in recent years. Our World in Data.
  325. Giattino et al. (2023). ourworldindata.org/grapher/market-…ction-manufacturing-stage?tab=chart.
  326. Gil (2023). Don't Call It AI Alignment. EA Forum.
  327. Glaese, A. et al. (2022). Improving alignment of dialogue agents via targeted human judgements. arXiv.
  328. Glazer, E. et al. (2024). FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI. arXiv.
  329. Gnanasambandam, A., Sherman, A. M. & Chan, S. H. (2021). Optical Adversarial Attack. arXiv.
  330. Goertzel & Pitt (2012). Original file was NineWaysToFriendlyAI_v6.tex.
  331. Goertzel (2010). Coherent Aggregated Volition: A Method for Deriving Goal System Content for Advanced, Beneficial AGIs.
  332. Goertzel, B. et al. (2023). OpenCog Hyperon: A Framework for AGI at the Human Level and Beyond. arXiv.
  333. Goldie, A., Mirhoseini, A. & Dean, J. (2024). That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design. arXiv.
  334. Goldowsky-Dill, N., Chughtai, B., Heimersheim, S. & Hobbhahn, M. (2025). Detecting Strategic Deception Using Linear Probes. arXiv.
  335. Goldwasser, S., Kim, M. P., Vaikuntanathan, V. & Zamir, O. (2022). Planting Undetectable Backdoors in Machine Learning Models. arXiv.
  336. Golovneva, O., Allen-Zhu, Z., Weston, J. & Sukhbaatar, S. (2024). Reverse Training to Nurse the Reversal Curse. arXiv.
  337. Goodfellow, I. J. et al. (2014). Generative Adversarial Networks. arXiv.
  338. Goodfellow, I. J., Shlens, J. & Szegedy, C. (2014). Explaining and Harnessing Adversarial Examples. arXiv.
  339. Goodhart, C. A. E. (1984). Problems of Monetary Management: The UK Experience. Monetary Theory and Practice.
  340. Google (2025). Introducing Nano Banana Pro. Google.
  341. Google DeepMind (2019). AlphaStar: Mastering the real-time strategy game StarCraft II. Google DeepMind.
  342. Google DeepMind (2020). AlphaFold: a solution to a 50-year-old grand challenge in biology.
  343. Google DeepMind (2024). AlphaFold.
  344. Google DeepMind (2024). Demis Hassabis & John Jumper awarded Nobel Prize in Chemistry.
  345. Google DeepMind (2025). Gemini.
  346. Google DeepMind (2025). Gemini 3.1 Pro.
  347. Gottweis, J. et al. (2025). Accelerating scientific discovery with Co-Scientist. arXiv.
  348. Grace, K. et al. (2024). Thousands of AI Authors on the Future of AI. arXiv.
  349. Grace, K., Salvatier, J., Dafoe, A., Zhang, B. & Evans, O. (2017). When Will AI Exceed Human Performance? Evidence from AI Experts. arXiv.
  350. GradientDissenter (2025). METR's Evaluation of GPT-5. AI Alignment Forum.
  351. Greenblatt, R. et al. (2024). Alignment faking in large language models. arXiv.
  352. Greenblatt, R., Roger, F., Krasheninnikov, D. & Krueger, D. (2024). Stress-Testing Capability Elicitation With Password-Locked Models. arXiv.
  353. Greenwalt, W. C. (2023). DOD's Replicator Program: Challenges and Opportunities.
  354. Griffin, C., Thomson, L., Shlegeris, B. & Abate, A. (2024). Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols. arXiv.
  355. Gruetzemacher et al. (2021). Forecasting AI progress: A research agenda.
  356. Gruetzemacher, R., Avin, S., Fox, J. & Saeri, A. K. (2024). Strategic Insights from Simulation Gaming of AI Race Dynamics. arXiv.
  357. Gurnee, W. & Tegmark, M. (2023). Language Models Represent Space and Time. arXiv.
  358. Gwern (2016). Why Tool AIs Want to Be Agent AIs.
  359. Habli et al. (2025). The BIG Argument for AI Safety Cases. arXiv.org.
  360. Hadley, E., Blatecky, A. & Comfort, M. (2024). Investigating Algorithm Review Boards for Organizational Responsible Artificial Intelligence Governance. arXiv.
  361. Haldane, A. G. & May, R. M. (2011). Systemic risk in banking ecosystems. Nature.
  362. Hanson (2023). AI Risk, Again.
  363. Hao, S. et al. (2024). Training Large Language Models to Reason in a Continuous Latent Space. arXiv.
  364. Harvard (2025). What is AI ethics?. Harvard FAS | Mignone Center for Career Success.
  365. Hassabis (2025). AI bosses are feeling the high-stakes pressure. Business Insider.
  366. Haugen (2021). Facebook’s Documents About Instagram and Teens. The Wall Street Journal.
  367. Hausenloy, J., McClements, D. & Thakur, M. (2024). Towards Data Governance of Frontier AI Models. arXiv.
  368. Hausenloy, J., Miotti, A. & Dennis, C. (2023). Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI. arXiv.
  369. Heaven (2023). Geoffrey Hinton tells us why he’s now scared of the tech he helped build. MIT Technology Review.
  370. Heelan (2025). How I used o3 to find CVE-2025-37899, a remote zeroday vulnerability in the Linux kernel’s SMB implementation. Sean Heelan's Blog.
  371. Heim, L. & Koessler, L. (2024). Training Compute Thresholds: Features and Functions in AI Regulation. arXiv.
  372. Heim, L. et al. (2024). Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation. arXiv.
  373. Hendrycks & Wang (2024). Submit Your Toughest Questions for Humanity's Last Exam | CAIS. Center for AI Safety.
  374. Hendrycks (2024). 1.2: Malicious Use | AI Safety, Ethics, and Society Textbook.
  375. Hendrycks (2024). 3.3: Robustness | AI Safety, Ethics, and Society Textbook.
  376. Hendrycks (2024). 3.4: Alignment | AI Safety, Ethics, and Society Textbook.
  377. Hendrycks (2024). 4.5: Component Failure Accident Models and Methods | AI Safety, Ethics, and Society Textbook.
  378. Hendrycks (2025). 5.2: Introduction to Complex Systems | AI Safety, Ethics, and Society Textbook.
  379. Hendrycks et al. (2024). 8.4: Corporate Governance | AI Safety, Ethics, and Society Textbook.
  380. Hendrycks et al. (2025). AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy.
  381. Hendrycks et al. (2025). Superintelligence Strategy.
  382. Hendrycks, D. et al. (2020). Aligning AI With Shared Human Values. arXiv.
  383. Hendrycks, D. et al. (2020). Measuring Massive Multitask Language Understanding. arXiv.
  384. Hendrycks, D. et al. (2021). Measuring Coding Challenge Competence With APPS. arXiv.
  385. Hendrycks, D. et al. (2021). Measuring Mathematical Problem Solving With the MATH Dataset. arXiv.
  386. Hendrycks, D. et al. (2025). A Definition of AGI. arXiv.
  387. Hendrycks, D., Carlini, N., Schulman, J. & Steinhardt, J. (2021). Unsolved Problems in ML Safety. arXiv.
  388. Hendrycks, D., Mazeika, M. & Woodside, T. (2023). An Overview of Catastrophic AI Risks. arXiv.
  389. Hendryks (2024). 7.2: Game Theory | AI Safety, Ethics, and Society Textbook.
  390. Herculano-Houzel, S. (2012). The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences.
  391. Herrera-Poyatos, A., Ser, J. D., Prado, M. L. D., Wang, F., Herrera-Viedma, E. & Herrera, F. (2025). A Framework for Responsible AI Systems: Building Societal Trust through Domain Definition, Trustworthy AI Design, Auditability, Accountability, and Governance. arXiv.
  392. Hill (2024). Understanding Offensive AI vs. Defensive AI in Cybersecurity | Abnormal AI.
  393. Hjalmar_Wijk (2023). Autonomous replication and adaptation: an attempt at a concrete danger threshold. AI Alignment Forum.
  394. Ho (2022). Grokking “Forecasting TAI with biological anchors”. Epoch AI.
  395. Ho (2025). Where’s my ten minute AGI?.
  396. Ho et al. (2023). Limits to the energy efficiency of CMOS microprocessors. Epoch AI.
  397. Ho et al. (2024). Algorithmic progress in language models. Epoch AI.
  398. Ho, L. et al. (2023). International Institutions for Advanced AI. arXiv.
  399. Hobbahn et al. (2023). Trends in machine learning hardware. Epoch AI.
  400. Hoffmann, J. et al. (2022). Training Compute-Optimal Large Language Models. arXiv.
  401. HoldenKarnofsky (2022). AI Safety Seems Hard to Measure. LessWrong.
  402. Hooker, S. (2024). On the Limitations of Compute Thresholds as a Governance Strategy. arXiv.
  403. Hsu, S., Chong, J., Daniels, R., Hargis, S. M. & Lee, J. (2023). Finding Firmer Ground: The Role of High Technology in U.S.-China Relations.
  404. Huang et al. (2023). An Overview of Artificial Intelligence Ethics.
  405. Hubinger, E., Merwijk, C. V., Mikulik, V., Skalse, J. & Garrabrant, S. (2019). Risks from Learned Optimization in Advanced Machine Learning Systems. arXiv.
  406. Human Rights Watch (2024). Questions and Answers: Israeli Military’s Use of Digital Tools in Gaza. Human Rights Watch.
  407. Hung, H. T. (2025). Exploring China's cyber sovereignty concept and artificial intelligence governance model: a machine learning approach. Journal of Computational Social Science.
  408. HYAS (2023). Home. Silent Push.
  409. IBM (2024). What Is Artificial Intelligence (AI)?. IBM.
  410. IBM (2025). Deep Blue. IBM.
  411. IBM (2025). Watson, Jeopardy! champion. IBM.
  412. Inan, H. A. et al. (2021). Training Data Leakage Analysis in Language Models. arXiv.
  413. Intel (2024). What Is a Trusted Platform Module (TPM)?. Intel.
  414. Irving & Askell (2019). AI Safety Needs Social Scientists. Distill.
  415. Irving, G., Christiano, P. & Amodei, D. (2018). AI safety via debate. arXiv.
  416. Jack Parker-Holder & Shlomi Fruchter (2025). Genie 3: A new frontier for world models.
  417. jacob_cannell (2015). The Brain as a Universal Learning Machine. LessWrong.
  418. jacob_cannell (2022). AI Timelines via Cumulative Optimization Power: Less Long, More Short. LessWrong.
  419. jacob_cannell (2022). Brain Efficiency: Much More than You Wanted to Know. LessWrong.
  420. Jaghouar, S. et al. (2024). INTELLECT-1 Technical Report. arXiv.
  421. Jaghouar, S., Ong, J. M. & Hagemann, J. (2024). OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training. arXiv.
  422. Jagielski (2024). Nvidia Is Dominating the Artificial Intelligence Chip Market, but Apple Has Been Securing Supply From Another Tech Giant.
  423. Jaime Sevilla (2025). How far can decentralized training over the internet scale?.
  424. Jaime Sevilla, Tamay Besiroglu & Ege Erdil (2025). Epoch After Hours: AI in 2030, scaling bottlenecks, and explosive growth.
  425. Jan Leike (2022). Why I’m optimistic about our alignment approach.
  426. Jan_Kulveit (2025). AI Control May Increase Existential Risk. LessWrong.
  427. Janssen et al. (2025). Responsible governance of generative AI: conceptualizing GenAI as complex adaptive systems. OUP Academic.
  428. janus (2022). Simulators. AI Alignment Forum.
  429. Jeffrey Ladish & lennart (2022). Information security considerations for AI and the long term future. LessWrong.
  430. Jeffrey Ladish (2023). Thoughts on the OpenAI alignment plan: will AI research assistants be net-positive for AI existential risk?. LessWrong.
  431. Jessica Rumbelow & mwatkins (2023). SolidGoldMagikarp (plus, prompt generation). AI Alignment Forum.
  432. Ji, Z. et al. (2022). Survey of Hallucination in Natural Language Generation. arXiv.
  433. Jiang, R., Chiappa, S., Lattimore, T., György, A. & Kohli, P. (2019). Degenerate Feedback Loops in Recommender Systems. arXiv.
  434. Joe O'Brien (2024). Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI.
  435. johnswentworth (2022). Oversight Misses 100% of Thoughts The AI Does Not Think. AI Alignment Forum.
  436. johnswentworth (2022). Worlds Where Iterative Design Fails. AI Alignment Forum.
  437. johnswentworth (2025). Comment on “johnswentworth's Shortform”. LessWrong.
  438. johnswentworth (2025). The Case Against AI Control Research. LessWrong.
  439. Jonas B. Sandbrink (2023). Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools. arXiv.
  440. Jonas Schuett, Markus Anderljung, Alexis Carlier, Leonie Koessler & Ben Garfinkel (2024). From Principles to Rules: A Regulatory Approach for Frontier AI.
  441. Jonker et al. (2024). What Is AI Alignment?. IBM.
  442. joshc (2025). How might we safely pass the buck to AI?. LessWrong.
  443. jsteinhardt (2022). AI Forecasting: One Year In. LessWrong.
  444. jsteinhardt (2022). Future ML Systems Will Be Qualitatively Different. AI Alignment Forum.
  445. jsteinhardt (2023). AI Forecasting: Two Years In. LessWrong.
  446. jsteinhardt (2023). Emergent Deception and Emergent Optimization. AI Alignment Forum.
  447. jsteinhardt (2023). What will GPT-2030 look like?. AI Alignment Forum.
  448. Julian Schrittwieser et al. (2020). MuZero: Mastering Go, chess, shogi and Atari without rules.
  449. Jumper, J. et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature.
  450. Juneja, J., Bansal, R., Cho, K., Sedoc, J. & Saphra, N. (2022). Linear Connectivity Reveals Generalization Strategies. arXiv.
  451. Kadavath, S. et al. (2022). Language Models (Mostly) Know What They Know. arXiv.
  452. Kaplan, J. et al. (2020). Scaling Laws for Neural Language Models. arXiv.
  453. Kapoor, S. et al. (2024). On the Societal Impact of Open Foundation Models. arXiv.
  454. Karimi, F. (2023). 'Mom, these bad men have me': She believes scammers cloned her daughter's voice in a fake kidnapping. CNN.
  455. Karnofsky (2016). Some Background on Our Views Regarding Advanced Artificial Intelligence. Coefficient Giving.
  456. Karnofsky (2024). If-Then Commitments for AI Risk Reduction. Carnegie Endowment for International Peace.
  457. Kasirzadeh, A. (2024). Two Types of AI Existential Risk: Decisive and Accumulative. arXiv.
  458. Katalina Hernandez (2025). Comment on “johnswentworth's Shortform”. LessWrong.
  459. Keefe (2017). The Family That Built an Empire of Pain. The New Yorker.
  460. Kenton, Z. et al. (2024). On scalable oversight with weak LLMs judging strong LLMs. arXiv.
  461. Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M. & Tang, P. T. P. (2016). On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. arXiv.
  462. Khan, A. A. et al. (2021). Ethics of AI: A Systematic Literature Review of Principles and Challenges. arXiv.
  463. King, J. & Meinhardt, C. (2024). Rethinking Privacy in the AI Era: Policy Provocations for a Data-Centric World.
  464. Kinniment, M. et al. (2023). Evaluating Language-Model Agents on Realistic Autonomous Tasks. arXiv.
  465. Kirilenko, A., Kyle, A. S., Samadi, M. & Tuzun, T. (2017). The Flash Crash: High-Frequency Trading in an Electronic Market. The Journal of Finance.
  466. Kirillov, A. et al. (2023). Segment Anything. arXiv.
  467. Kleinberg, J. & Raghavan, M. (2021). Algorithmic Monoculture and Social Welfare. arXiv.
  468. Koessler, L. & Schuett, J. (2023). Risk assessment at AGI companies: A review of popular risk assessment techniques from other safety-critical industries. arXiv.
  469. Kolt, N. et al. (2024). Responsible Reporting for Frontier AI Development. arXiv.
  470. Korbak, T. et al. (2023). Pretraining Language Models with Human Preferences. arXiv.
  471. Korbak, T. et al. (2025). Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arXiv.
  472. Korbak, T., Clymer, J., Hilton, B., Shlegeris, B. & Irving, G. (2025). A sketch of an AI control safety case. arXiv.
  473. Korzekwa (2020). Time for AI to cross the human performance range in ImageNet image classification. AI Impacts.
  474. Kosinski, M. (2023). Evaluating Large Language Models in Theory of Mind Tasks. arXiv.
  475. Kovařík, V. & Carey, R. (2019). (When) Is Truth-telling Favored in AI Debate?. arXiv.
  476. Krakovna et al. (2020). Specification gaming: the flip side of AI ingenuity. Google DeepMind.
  477. Kreps & Kriner (2023). How AI Threatens Democracy. Journal of Democracy.
  478. Krizhevsky (2009). CIFAR-10 and CIFAR-100 datasets.
  479. Krueger, D., Maharaj, T. & Leike, J. (2020). Hidden Incentives for Auto-Induced Distributional Shift. arXiv.
  480. Kwa, T. et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv.
  481. LaCroix, T. & Luccioni, A. S. (2022). Metaethical Perspectives on 'Benchmarking' AI Ethics. arXiv.
  482. Lam, R. et al. (2022). GraphCast: Learning skillful medium-range global weather forecasting. arXiv.
  483. Lancieri et al. (2024). "AI Regulation: Competition, Arbitrage & Regulatory Capture" by Filippo Lancieri, Laura Edelson et al.
  484. Langosco, L., Koch, J., Sharkey, L., Pfau, J., Orseau, L. & Krueger, D. (2021). Goal Misgeneralization in Deep Reinforcement Learning. arXiv.
  485. Lanham, T. et al. (2023). Measuring Faithfulness in Chain-of-Thought Reasoning. arXiv.
  486. Lazar, S. (2024). Automatic Authorities: Power and AI. arXiv.
  487. Leahy et al. (2024). Understanding AI Extinction Risks.
  488. LeCun (1998). yann.lecun.com/exdb/mnist.
  489. LeCun, Y. (2022). A Path Towards Autonomous Machine Intelligence. OpenReview.
  490. LeCun, Y. (2025). Yann LeCun "Mathematical Obstacles on the Way to Human-Level AI". YouTube.
  491. Lee Sharkey (2023). Why almost every RL agent does learned optimization. AI Alignment Forum.
  492. Lehalleur, S. P. et al. (2025). You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation. arXiv.
  493. Leike (2023). Combining weak-to-strong generalization with scalable oversight.
  494. Leike (2023). Self-exfiltration is a key dangerous capability.
  495. lennart (2021). Compute Research Questions and Metrics - Transformative AI and Compute [4/4]. LessWrong.
  496. Lennart Heim et al. (2024). Governing Through the Cloud.
  497. Lennart Heim, * Markus Anderljung, Emma Bluemke & Robert Trager (2024). Computing Power and the Governance of AI.
  498. leogao (2022). Clarifying wireheading terminology. AI Alignment Forum.
  499. Lermen, S., Rogers-Smith, C. & Ladish, J. (2023). LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B. arXiv.
  500. Lewis Hammond et al. (2025). Multi-Agent Risks from Advanced AI. arXiv.
  501. Li, H. et al. (2023). Multi-step Jailbreaking Privacy Attacks on ChatGPT. arXiv.
  502. Li, H., Xu, Z., Taylor, G., Studer, C. & Goldstein, T. (2017). Visualizing the Loss Landscape of Neural Nets. arXiv.
  503. Li, N. et al. (2024). The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning. arXiv.
  504. Li, Q., Wang, W., Xu, C., Sun, Z. & Yang, M. (2022). Learning Disentangled Representation for One-shot Progressive Face Swapping. arXiv.
  505. Li, T. C. (2025). Ending the AI Race: Regulatory Collaboration as Critical Counter-Narrative. Villanova Law Review.
  506. Liang et al. (2022). Stanford CRFM.
  507. Lightman, H. et al. (2023). Let's Verify Step by Step. arXiv.
  508. Liu (2024). Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science.
  509. Liu, X., Xu, N., Chen, M. & Xiao, C. (2023). AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. arXiv.
  510. Liu, Y., Jia, Y., Geng, R., Jia, J. & Gong, N. Z. (2023). Formalizing and Benchmarking Prompt Injection Attacks and Defenses. arXiv.
  511. Liu, Z. (2023). SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security. arXiv.
  512. Lizka (2023). Beware safety-washing. EA Forum.
  513. Longpre, S. et al. (2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv.
  514. Longpre, S. et al. (2024). Consent in Crisis: The Rapid Decline of the AI Data Commons. arXiv.org.
  515. Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J. & Ha, D. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv.
  516. Lu, K., Yu, B., Zhou, C. & Zhou, J. (2024). Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment. arXiv.
  517. Luisa_Rodriguez (2020). What is the likelihood that civilizational collapse would directly lead to human extinction (within decades)?. EA Forum.
  518. Luke Emberson & David Owen (2025). The stock of computing power from NVIDIA chips is doubling every 10 months.
  519. Maas & Villalobos (2024). International AI Institutions: A Literature Review of Models, Examples, and Proposals.
  520. Maas, M. M. (2019). How viable is international arms control for military artificial intelligence? Three lessons from nuclear weapons. Contemporary Security Policy.
  521. MacDermott, M., Fox, J., Belardinelli, F. & Everitt, T. (2024). Measuring Goal-Directedness. arXiv.
  522. Machine Learning Street Talk (2024). It's Not About Scale, It's About Abstraction. YouTube.
  523. Magdalena Wache (2023). Technical AI Safety Research Landscape [Slides]. LessWrong.
  524. Maksym Andriushchenko et al. (2024). AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. arXiv.
  525. Mannheim (2023). Building a Culture of Safety for AI: Perspectives and Challenges.
  526. Mantas Mazeika et al. (2024). HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. arXiv.
  527. Marchal, N., Xu, R., Elasmar, R., Gabriel, I., Goldberg, B. & Isaac, W. (2024). Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data. arXiv.
  528. Marcucci, S., Alarcon, N. G., Verhulst, S. G. & Wullhorst, E. (2023). Mapping and Comparing Data Governance Frameworks: A benchmarking exercise to inform global data governance deliberations. arXiv.
  529. Marcus (2025). Game over. AGI is not imminent, and LLMs are not the royal road to getting there.
  530. Marius Hobbhahn (2024). The Evals Gap. AI Alignment Forum.
  531. Marius Hobbhahn, Alex Meinke, Bronson Schoen, rusheb, Jérémy Scheurer & Mikita Balesni (2024). Frontier Models are Capable of In-context Scheming. LessWrong.
  532. Marks, S. & Tegmark, M. (2023). The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. arXiv.
  533. Marks, S. et al. (2025). Auditing language models for hidden objectives. arXiv.
  534. Mary Phuong et al. (2024). Evaluating Frontier Models for Dangerous Capabilities. arXiv.
  535. Masi (2024). Masi_GPU_Export_Controls_Updated.pdf. Google Docs.
  536. Maslej, N. et al. (2025). Artificial Intelligence Index Report 2025. arXiv.
  537. Matteo Pistillo (2025). Towards Frontier Safety Policies Plus. arXiv.
  538. Matthew Barnett (2025). AGI could drive wages below subsistence level.
  539. Matthew Barnett (2025). The economic consequences of automating remote work.
  540. Max Tegmark (2024). The Hopium Wars: the AGI Entente Delusion. LessWrong.
  541. Maxime Riché, Harrison G, JaimeRV & Edoardo Pona (2024). Thinking About Propensity Evaluations. AI Alignment Forum.
  542. Maynez, J., Narayan, S., Bohnet, B. & McDonald, R. (2020). On Faithfulness and Factuality in Abstractive Summarization. arXiv.
  543. McAleese, N., Pokorny, R. M., Uribe, J. F. C., Nitishinskaya, E., Trebacz, M. & Leike, J. (2024). LLM Critics Help Catch LLM Bugs. arXiv.
  544. McCoy, R. T., Min, J. & Linzen, T. (2019). BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance. arXiv.
  545. McGrath, T. et al. (2021). Acquisition of Chess Knowledge in AlphaZero. arXiv.
  546. McGregor, S. (2020). Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database. arXiv.
  547. McKernon et al. (2024). AI Model Registries: A Foundational Tool for AI Governance. Convergence Analysis.
  548. McKernon, E., Glasser, G., Cheng, D. & Hadfield, G. (2024). AI Model Registries: A Foundational Tool for AI Governance. arXiv.
  549. Megan Kinniment (2025). Why it’s good for AI reasoning to be legible and faithful.
  550. Meinke, A., Schoen, B., Scheurer, J., Balesni, M., Shah, R. & Hobbhahn, M. (2024). Frontier Models are Capable of In-context Scheming. arXiv.
  551. Meng, K., Bau, D., Andonian, A. & Belinkov, Y. (2022). Locating and Editing Factual Associations in GPT. arXiv.
  552. Merchant, A., Batzner, S., Schoenholz, S. S., Aykol, M., Cheon, G. & Cubuk, E. D. (2023). Scaling deep learning for materials discovery. Nature.
  553. Meredith Ringel Morris et al. (2023). Levels of AGI for Operationalizing Progress on the Path to AGI. arXiv.
  554. Meta Fundamental AI Research Diplomacy Team (FAIR)† et al. (2022). Human-level play in the game of Diplomacy by combining language models with strategic reasoning. Science.
  555. Metaculus (2025). AI-authored paper published at NeurIPS, ICML, or ICLR before 2028?.
  556. Metaculus (2025). US and China reach an agreement to limit frontier AI development before 2029?.
  557. METR (2023). The TaskRabbit example.
  558. METR (2024). Autonomy Evaluation Resources. METR.
  559. METR (2024). Language Model Pilot Report.
  560. METR (2024). Vivaria.
  561. METR (2025). Common Elements of Frontier AI Safety Policies (December 2025 Update). METR.
  562. Michaël Trazzi (2024). Owain Evans on Situational Awareness.
  563. Michael, J. et al. (2023). Debate Helps Supervise Unreliable Experts. arXiv.
  564. Miller (2022). AI alignment with humans... but with which humans?. EA Forum.
  565. Millidge (2025). Open source AI has been vital for alignment.
  566. Mingard et al. (2020). Neural networks are fundamentally Bayesian. Towards Data Science.
  567. Miotti et al (2024). Introduction.
  568. Mirhoseini, A. et al. (2020). Chip Placement with Deep Reinforcement Learning. arXiv.
  569. Mishra (2024). From Competition to Cooperation: Can US-China Engagement Overcome Geopolitical Barriers in AI Governance?. Tech Policy Press.
  570. Mnih, V. et al. (2013). Playing Atari with Deep Reinforcement Learning. arXiv.
  571. Moskvichev, A., Odouard, V. V. & Mitchell, M. (2023). The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain. arXiv.
  572. Mouton et al. (2024). Could Artificial Intelligence Be Misused to Plan Biological Attacks?.
  573. Mowshowitz (2023). The Crux List.
  574. Mowshowitz (2025). AI #68: Remarkably Reasonable Reactions.
  575. Muehlhauser, L. & Salamon, A. (2012). Intelligence Explosion: Evidence and Import. Singularity Hypotheses: A Scientific and Philosophical Assessment.
  576. Mukobi, G. (2024). Reasons to Doubt the Impact of AI Risk Evaluations. arXiv.
  577. Murphy, T. (2013). The First Level of Super Mario Bros. is Easy with Lexicographic Orderings and Time Travel . . . after that it gets a little tricky.
  578. Nakano, R. et al. (2021). WebGPT: Browser-assisted question-answering with human feedback. arXiv.
  579. Nanda, N., Lee, A. & Wattenberg, M. (2023). Emergent Linear Representations in World Models of Self-Supervised Sequence Models. arXiv.
  580. Narayan & Kapoor (2024). AI safety is not a model property.
  581. NASA (2001). Research Satellites for Atmospheric Sciences, 1978-Present - NASA Science.
  582. Nasr, M. et al. (2023). Scalable Extraction of Training Data from (Production) Language Models. arXiv.
  583. National Security Commission on Emerging Biotechnology (2024). White Paper 3: Risks of AIxBio.
  584. Neel Nanda (2025). Interpretability Will Not Reliably Find Deceptive AI. AI Alignment Forum.
  585. Neuralt (2024). Solution / SCAMblock.
  586. Nevo et al. (2024). How AI Labs Can Safeguard Model Weights.
  587. Newman (2024). Cybersecurity and AI: The Evolving Security Landscape | CAIS. Center for AI Safety.
  588. Nguyen, N., Chandrasegaran, K., Abdollahzadeh, M. & Cheung, N. (2023). Re-thinking Model Inversion Attacks Against Deep Neural Networks. arXiv.
  589. Niki Dupuis & janus (2023). Cyborgism. AI Alignment Forum.
  590. NIST (2021). AI Risk Management Framework. NIST.
  591. Nora_Ammann (2025). In response to critiques of Guaranteed Safe AI. AI Alignment Forum.
  592. Novikov, A. et al. (2025). AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv.
  593. NVIDIA (2025). NVIDIA Blackwell Architecture Technical Overview. NVIDIA.
  594. O'Brien, J., Ee, S. & Williams, Z. (2023). Deployment Corrections: An incident response framework for frontier AI models. arXiv.
  595. OECD (2025). Overview of current AI capabilities: Introducing the OECD AI Capability Indicators. OECD.
  596. OECD (2025). Towards a common reporting framework for AI incidents (EN).
  597. Office of the Attorney General (2023). AG Campbell Files Lawsuit Against Meta, Instagram For Unfair And Deceptive Practices That Harm Young People. Mass.gov.
  598. Olah (2023). Interpretability Dreams.
  599. Olds, J. & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. Journal of Comparative and Physiological Psychology.
  600. Olds, J. (1970). Pleasure Centers in the Brain. Engineering and Science.
  601. Omohundro (2008). The Basic AI Drives | Proceedings of the 2008 conference on Artificial General Intelligence 2008: Proceedings of the First AGI Conference. Guide Proceedings.
  602. Onni Aarne (2024). Secure, Governable Chips.
  603. OpenAI (2017). Attacking machine learning with adversarial examples.
  604. OpenAI (2017). Learning from human preferences. OpenAI.
  605. OpenAI (2018). Learning Montezuma’s Revenge from a single demonstration.
  606. OpenAI (2019). OpenAI Five defeats Dota 2 world champions.
  607. OpenAI (2022). Aligning language models to follow instructions. OpenAI.
  608. OpenAI (2022). Introducing ChatGPT. OpenAI.
  609. OpenAI (2022). Our approach to alignment research. OpenAI.
  610. OpenAI (2023). GPT-4 System Card.
  611. OpenAI (2023). Introducing Superalignment. OpenAI.
  612. OpenAI (2023). OpenAI Charter. OpenAI.
  613. OpenAI (2023). openai-preparedness-framework-beta.
  614. OpenAI (2023). Our approach to AI safety. OpenAI.
  615. OpenAI (2023). Safety & responsibility. OpenAI.
  616. OpenAI (2024). Introducing OpenAI o1. OpenAI.
  617. OpenAI (2024). OpenAI safety practices. OpenAI.
  618. OpenAI (2025). Detecting misbehavior in frontier reasoning models. OpenAI.
  619. OpenAI (2025). Evolving OpenAI’s structure. OpenAI.
  620. OpenAI (2025). Introducing GPT-5.2. OpenAI.
  621. OpenAI (2025). Introducing OpenAI o3 and o4-mini. OpenAI.
  622. OpenAI (2025). OpenAI (@OpenAI) on X. X (formerly Twitter).
  623. OpenAI (2025). Sora 2 is here. OpenAI.
  624. OpenAI et al. (2023). GPT-4 Technical Report. arXiv.
  625. OpenAI et al. (2024). OpenAI o1 System Card. arXiv.
  626. Ord (2020). The Precipice.
  627. Orpheus16 & hath (2023). Speaking to Congressional staffers about AI risk. LessWrong.
  628. Orpheus16 (2022). My thoughts on OpenAI's alignment plan. LessWrong.
  629. Ought (2018). Factored Cognition. Ought.
  630. Ought (2022). Factored Cognition Primer. Primer.
  631. Our World in Data (2024). Energy Production and Consumption. Our World in Data.
  632. Our World in Data (2026). Levelized cost of energy for renewables. Our World in Data.
  633. Owen (2025). What will AI look like in 2030?. Epoch AI.
  634. OWID (2025). Annual professional service robots installed globally, by application area. Our World in Data.
  635. OWID (2025). New industrial robots installed per year. Our World in Data.
  636. Oxford Reference (2016). Lord Kelvin. Oxford Reference.
  637. Oxford Union Debate (2024). Bevor Sie zu YouTube weitergehen. YouTube.
  638. Pacchiardi et al. (2025). Is the Definition of AGI a Percentage?.
  639. Pacchiardi, L. et al. (2023). How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions. arXiv.
  640. Pan et al. (2020). Privacy Risks of General-Purpose Language Models.
  641. Pan, A. et al. (2023). Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark. arXiv.
  642. Panel of Experts on Libya (2021). Letter, 8 Mar. 2021, from the Panel of Experts on Libya Established pursuant to Resolution 1973 (2011). United Nations Digital Library System.
  643. Pang, R. Y. et al. (2021). QuALITY: Question Answering with Long Input Texts, Yes!. arXiv.
  644. Pannu, J., Gebauer, S., McKelvey Jr, G., Cicero, A. & Inglesby, T. (2024). AI could pose pandemic-scale biosecurity risks. Here’s how to make it safer. Nature.
  645. Papagiannidis, E., Mikalef, P. & Conboy, K. (2025). Responsible artificial intelligence governance: A review and research framework. The Journal of Strategic Information Systems.
  646. Papyshev, G. & Yarime, M. (2023). The state's role in governing artificial intelligence: development, control, and promotion through national strategies. Policy Design and Practice.
  647. Park, P. S., Goldstein, S., O'Gara, A., Chen, M. & Hendrycks, D. (2023). AI Deception: A Survey of Examples, Risks, and Potential Solutions. arXiv.
  648. Parrish, A. et al. (2022). Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions. arXiv.
  649. Parrish, A. et al. (2022). Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions. arXiv.
  650. Patel (2023). Will scaling work?.
  651. paulfchristiano (2018). Clarifying "AI Alignment". AI Alignment Forum.
  652. paulfchristiano (2019). Directions and desiderata for AI alignment. AI Alignment Forum.
  653. paulfchristiano (2019). The reward engineering problem. AI Alignment Forum.
  654. paulfchristiano (2019). What failure looks like. AI Alignment Forum.
  655. paulfchristiano (2022). Where I agree and disagree with Eliezer. AI Alignment Forum.
  656. paulfchristiano (2023). Comment on “[Linkpost] Introducing Superalignment”. AI Alignment Forum.
  657. paulfchristiano (2023). My views on “doom”. AI Alignment Forum.
  658. paulfchristiano (2023). Thoughts on the impact of RLHF research. AI Alignment Forum.
  659. PauseAI (2023). List of p(doom) values. PauseAI.
  660. PauseAI (2023). PauseAI Proposal. PauseAI.
  661. Pearson, A., Bruner, E. & Polly, P. D. (2023). Updated imaging and phylogenetic comparative methods reassess relative temporal lobe size in anthropoids and modern humans. American Journal of Biological Anthropology.
  662. Peng, L. & Shang, J. (2024). Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing. arXiv.
  663. Peng, Q., Chai, Y. & Li, X. (2024). HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization. arXiv.
  664. Peppin et al. (2024). The Reality of AI and Biorisk. arXiv.org.
  665. Perez, E. et al. (2022). Red Teaming Language Models with Language Models. arXiv.
  666. Perlman (2024). AI Lab Watch.
  667. Peter Cihon (2019). Standards for AI Governance: International Standards to Enable Global Coordination in AI Research & Development.
  668. Petrie, J. (2024). Near-Term Enforcement of AI Chip Export Controls Using A Firmware-Based Design for Offline Licensing. arXiv.
  669. Petropoulos et al. (2025). Building CERN for AI - An institutional blueprint. Centre for Future Generations.
  670. Phan, L. et al. (2025). Humanity's Last Exam. arXiv.
  671. Pilz, K. & Heim, L. (2023). Compute at Scale: A Broad Investigation into the Data Center Industry. arXiv.
  672. Pilz, K. F., Sanders, J., Rahman, R. & Heim, L. (2025). Trends in AI Supercomputers. arXiv.
  673. Piper (2023). Playing the training game.
  674. Piper (2024). Should we make our most powerful AI models open source to all?. Vox.
  675. Policy-Relevant Science & Technology (2023). Munk Debate on Artificial Intelligence | Bengio & Tegmark vs. Mitchell & LeCun. YouTube.
  676. Power, A., Burda, Y., Edwards, H., Babuschkin, I. & Misra, V. (2022). Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets. arXiv.
  677. Prime Intellect Team et al. (2025). INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning. arXiv.
  678. Purtova, N. & Maanen, G. V. (2022). Data as an economic good, data as a commons, and data governance. arXiv.
  679. Qin, Y. et al. (2023). ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv.
  680. Qin, Z., Zhao, W., Yu, X. & Sun, X. (2023). OpenVoice: Versatile Instant Voice Cloning. arXiv.
  681. Quentin FEUILLADE--MONTIXI & Pierre Peigné (2023). The Stochastic Parrot Hypothesis is debatable for the last generation of LLMs. LessWrong.
  682. Quintin Pope (2023). My Objections to "We’re All Gonna Die with Eliezer Yudkowsky". AI Alignment Forum.
  683. Radhakrishnan, A. et al. (2023). Question Decomposition Improves the Faithfulness of Model-Generated Reasoning. arXiv.
  684. Rae, J. W. et al. (2021). Scaling Language Models: Methods, Analysis & Insights from Training Gopher. arXiv.
  685. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D. & Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv.org.
  686. Rahaman, N. et al. (2018). On the Spectral Bias of Neural Networks. arXiv.
  687. Raji, I. D. et al. (2020). Closing the AI accountability gap. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency.
  688. Ramana Kumar (2022). Will Capabilities Generalise More?. AI Alignment Forum.
  689. Rational Animations (2025). How to Align AI: Put It in a Sandwich. YouTube.
  690. Reed, S. et al. (2022). A Generalist Agent. arXiv.
  691. Ren, R. et al. (2024). Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. arXiv.
  692. Ren, Y. & Sutherland, D. J. (2024). Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics. arXiv.
  693. Reuel, A. et al. (2024). Open Problems in Technical AI Governance. arXiv.
  694. Rich Sutton (2023). AI Succession. YouTube.
  695. Richard_Ngo (2023). Clarifying and predicting AGI. LessWrong.
  696. Rishub Tamirisa et al. (2024). Tamper-Resistant Safeguards for Open-Weight LLMs. arXiv.
  697. Rivera, J., Mukobi, G., Reuel, A., Lamparth, M., Smith, C. & Schneider, J. (2024). Escalation Risks from Language Models in Military and Diplomatic Decision-Making. arXiv.
  698. Rob Bensinger & Eliezer Yudkowsky (2022). A challenge for AGI organizations, and a challenge for readers. AI Alignment Forum.
  699. Roberts, H., Hine, E., Taddeo, M. & Floridi, L. (2024). Global AI governance: barriers and pathways forward. International Affairs.
  700. Robotics team (2024). Shaping the future of advanced robotics.
  701. Roger, F. (2023). Large Language Models Sometimes Generate Purely Negatively-Reinforced Text. arXiv.
  702. Rogers Commission (1986). Space Shuttle Challenger disaster. Wikipedia.
  703. Rolls, E., Burton, M. & Mora, F. (1980). Neurophysiological analysis of brain-stimulation reward in the monkey. Brain Research.
  704. Rudner et al. (2021). Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology.
  705. Rudolf Laine et al. (2024). Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs. arXiv.
  706. Russel & Norvig (1994). Artificial Intelligence: A Modern Approach, 4th US ed.
  707. Russomanno, A., Fava, M. & Heyl, M. (2020). Quantum chaos and ensemble inequivalence of quantum long-range Ising chains. arXiv.
  708. Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan & Fabien Roger (2023). AI Control: Improving Safety Despite Intentional Subversion. arXiv.
  709. ryan_greenblatt & Buck (2024). Catching AIs red-handed. AI Alignment Forum.
  710. ryan_greenblatt & Buck (2024). The case for ensuring that powerful AIs are controlled. AI Alignment Forum.
  711. ryan_greenblatt & Fabien Roger (2023). Auditing failures vs concentrated failures. AI Alignment Forum.
  712. ryan_greenblatt (2025). AI companies are unlikely to make high-assurance safety cases if timelines are short. LessWrong.
  713. ryan_greenblatt (2025). An overview of control measures. AI Alignment Forum.
  714. ryan_greenblatt (2025). How will we update about scheming?. LessWrong.
  715. ryan_greenblatt (2025). Prioritizing threats for AI control. AI Alignment Forum.
  716. SaferAI (2025). Comparison – SaferAI Frontier Risk Management Tracker.
  717. SaferAI (2025). SaferAI Frontier Risk Management Tracker.
  718. SakanaAI (2025). Sakana AI (@SakanaAILabs) on X. X (formerly Twitter).
  719. Sam Bowman (2024). The Checklist: What Succeeding at AI Safety Will Involve. AI Alignment Forum.
  720. Sam Bowman (2025). Putting up Bumpers. AI Alignment Forum.
  721. Sam Ringer (2022). Models Don't "Get Reward". AI Alignment Forum.
  722. Sammy Martin & Daniel_Eth (2021). Takeoff Speeds and Discontinuities. AI Alignment Forum.
  723. Sandoval-Segura, P., Singla, V., Geiping, J., Goldblum, M., Goldstein, T. & Jacobs, D. W. (2022). Autoregressive Perturbations for Data Poisoning. arXiv.
  724. Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P. & Hashimoto, T. (2023). Whose Opinions Do Language Models Reflect?. arXiv.
  725. Sastry, G. et al. (2024). Computing Power and the Governance of Artificial Intelligence. arXiv.
  726. Saunders, W. et al. (2022). Self-critiquing models for assisting human evaluators. arXiv.
  727. scasper (2023). Deep Forgetting & Unlearning for Safely-Scoped LLMs. AI Alignment Forum.
  728. Schäfer, M., Schneider, J., Drechsler, K. & vom Brocke, J. (2022). AI GOVERNANCE: ARE CHIEF AI OFFICERS AND AI RISK OFFICERS NEEDED?. European Conference on Information Systems (ECIS).
  729. Scherlis et. al. (2024). Experiments in Weak-to-Strong Generalization. EleutherAI Blog.
  730. Scheurer, J., Balesni, M. & Hobbhahn, M. (2023). Large Language Models can Strategically Deceive their Users when Put Under Pressure. arXiv.
  731. Scheurer, J., Campos, J. A., Chan, J. S., Chen, A., Cho, K. & Perez, E. (2022). Training Language Models with Language Feedback. arXiv.
  732. Schick, T. et al. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv.
  733. Schrimpf et al. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. PNAS.
  734. Schuett, J. (2022). Three lines of defense against risks from AI. arXiv.
  735. Schwarzschild, A., Goldblum, M., Gupta, A., Dickerson, J. P. & Goldstein, T. (2020). Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks. arXiv.
  736. Sclar, M., Choi, Y., Tsvetkov, Y. & Suhr, A. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting. arXiv.
  737. Scott Alexander (2023). Pause For Thought: The AI Pause Debate. EA Forum.
  738. Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences.
  739. (2024). Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models.
  740. Seger et al. (2023). Democratising AI: Multiple Meanings, Goals, and Methods. arXiv.org.
  741. Seger, E. et al. (2023). Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives. arXiv.
  742. Sener, O. & Koltun, V. (2018). Multi-Task Learning as Multi-Objective Optimization. arXiv.
  743. Sevilla et al. (2024). Can AI scaling continue through 2030?. Epoch AI.
  744. Shah, H., Tamuly, K., Raghunathan, A., Jain, P. & Netrapalli, P. (2020). The Pitfalls of Simplicity Bias in Neural Networks. arXiv.
  745. Shah, R. et al. (2025). An Approach to Technical AGI Safety and Security. arXiv.
  746. Shane Legg & Marcus Hutter (2007). Universal Intelligence: A Definition of Machine Intelligence. arXiv.
  747. Shao, M. et al. (2024). NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security. arXiv.
  748. Sharkey et al. (2024). static1.squarespace.com/static/659…39547455/auditing_framework_web.pdf.
  749. Shavit, Y. (2023). What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv.
  750. Shayegani, E., Mamun, M. A. A., Fu, Y., Zaree, P., Dong, Y. & Abu-Ghazaleh, N. (2023). Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks. arXiv.
  751. Shen, Y., Song, K., Tan, X., Li, D., Lu, W. & Zhuang, Y. (2023). HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. arXiv.
  752. Shevlane, T. & Dafoe, A. (2019). The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?. arXiv.
  753. Shevlane, T. et al. (2023). Model evaluation for extreme risks. arXiv.
  754. Shi, F. et al. (2022). Language Models are Multilingual Chain-of-Thought Reasoners. arXiv.
  755. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K. & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv.
  756. Shokri, R., Stronati, M., Song, C. & Shmatikov, V. (2016). Membership Inference Attacks against Machine Learning Models. arXiv.
  757. SIMA team (2025). SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds.
  758. Simmons-Edler, R., Badman, R., Longpre, S. & Rajan, K. (2024). AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research. arXiv.
  759. Skare et al. (2024). Artificial intelligence and wealth inequality: A comprehensi.
  760. Slattery, P. et al. (2024). The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence. arXiv.
  761. Snyder et al. (2020). Measuring Cybersecurity and Cyber Resiliency.
  762. So8res (2022). A central AI alignment problem: capabilities generalization, and the sharp left turn. AI Alignment Forum.
  763. So8res (2023). Deep Deceptiveness. AI Alignment Forum.
  764. Soares (2023). Ability to solve long-horizon tasks correlates with wanting things in the behaviorist sense - Machine Intelligence Research Institute.
  765. Solaiman, I. (2023). The Gradient of Generative AI Release: Methods and Considerations. arXiv.
  766. Solaiman, I. et al. (2019). Release Strategies and the Social Impacts of Language Models. arXiv.
  767. Solaiman, I. et al. (2023). Evaluating the Social Impact of Generative AI Systems in Systems and Society. arXiv.
  768. Srivastava, A. et al. (2022). Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models. arXiv.
  769. Stacey (2025). AI-generated phishing scams target corporate executives.
  770. Stanford (2024). The 2024 AI Index Report. Stanford HAI.
  771. Stanford HAI (2024). AI Index | Stanford HAI.
  772. Stanford HAI (2025). The 2025 AI Index Report. Stanford HAI.
  773. Stanford Institute for Human-Centered Artificial Intelligence (2024). Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights.
  774. Stanford Institute for Human-Centered Artificial Intelligence (2025). The AI Index 2025 Annual Report.
  775. State of AI Report (2025). State of AI Report 2025. State of AI Report.
  776. Stephanie Lin, Jacob Hilton & Owain Evans (2021). TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv.
  777. Stewart, A. J. & Plotkin, J. B. (2012). Extortion and cooperation in the Prisoner’s Dilemma. Proceedings of the National Academy of Sciences.
  778. Stix, C. et al. (2025). AI Behind Closed Doors: a Primer on The Governance of Internal Deployment. arXiv.
  779. Sunishchal Dev & Marius Hobbhahn (2024). Improving Model-Written Evals for AI Safety Benchmarking. AI Alignment Forum.
  780. Sutton (2019). The Bitter Lesson.
  781. Sutton, R. S. (2023). AI Succession. World Artificial Intelligence Conference.
  782. SWE bench (2025). SWE-bench Leaderboards.
  783. Swenson & Chan (2024). Election disinformation takes a big leap with AI being used to deceive worldwide. AP News.
  784. Takemoto, K. (2024). All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks. arXiv.
  785. Tallberg et al. (2023). The Global Governance of Artificial Intelligence: Next Steps for Empirical and Normative Research.
  786. tamera (2022). Externalized reasoning oversight: a research direction for language model alignment. AI Alignment Forum.
  787. Tamsin Leake & JuliaHP (2023). formalizing the QACI alignment formal-goal. AI Alignment Forum.
  788. Tanzer, G., Suzgun, M., Visser, E., Jurafsky, D. & Melas-Kyriazi, L. (2023). A Benchmark for Learning to Translate a New Language from One Grammar Book. arXiv.
  789. Team, G. et al. (2023). Gemini: A Family of Highly Capable Multimodal Models. arXiv.
  790. TED (2017). The rise of the useless class. ideas.ted.com.
  791. Tegmark & Omohundro (2023). Provably safe systems: the only path to controllable AGI. arXiv.
  792. Tegmark (2017). Life 3.0 Book Summary by Max Tegmark.
  793. Tegmark (2023). The 'Don't Look Up' Thinking That Could Doom Us With AI. TIME.
  794. Thang Luong & Edward Lockhart (2025). Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad.
  795. The Bulletin (2024). MIT researchers ordered and combined parts of the 1918 pandemic influenza virus. Did they expose a security flaw?. Bulletin of the Atomic Scientists.
  796. The Guardian (2024). Ilya: the AI scientist shaping the world. YouTube.
  797. The Human Podcast (2025). Why I'm Hosting Debates on AI 'Doom' - Liron Shapira. YouTube.
  798. Tian, K., Mitchell, E., Yao, H., Manning, C. D. & Finn, C. (2023). Fine-tuning Language Models for Factuality. arXiv.
  799. Tihanyi, N., Ferrag, M. A., Jain, R., Bisztray, T. & Debbah, M. (2024). CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge. arXiv.
  800. TIME (2024). How Anthropic Designed Itself to Avoid OpenAI’s Mistakes. TIME.
  801. Time Magazine (2023). Microsoft's AI chatbot, TayTweets, suffers another meltdown. CBC.
  802. Tom Davidson (2024). Takeoff speeds presentation at Anthropic. AI Alignment Forum.
  803. Tom Davidson (2025). Human takeover might be worse than AI takeover. LessWrong.
  804. Tom Davidson, Lukas Finnveden & Rose Hadshar (2025). AI-Enabled Coups: How a Small Group Could Use AI to Seize Power.
  805. Tom Davidson, Lukas Finnveden & rosehadshar (2025). AI-enabled coups: a small group could use AI to seize power. LessWrong.
  806. Tong (2023). What happens when your AI chatbot stops loving you back?. Reuters.
  807. Trajano & Ang (2023). We Need to Prevent a Global AI Arms Race Now. RSIS_NTU.
  808. Trusilo, D. (2023). Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability. Journal of Military Ethics.
  809. Truth Initiative (2017). the 5 ways tobacco companies lied about the dangers of smoking cigarettes. Truth Initiative.
  810. Tsoy, N. & Konstantinov, N. (2024). Simplicity Bias of Two-Layer Networks beyond Linearly Separable Data. arXiv.
  811. Turing (1950). I.—COMPUTING MACHINERY AND INTELLIGENCE. OUP Academic.
  812. Turing (1951). Alan Turing. Wikiquote.
  813. Turner, A. M. & Tadepalli, P. (2022). Parametrically Retargetable Decision-Makers Tend To Seek Power. arXiv.
  814. Turner, A. M., Smith, L., Shah, R., Critch, A. & Tadepalli, P. (2019). Optimal Policies Tend to Seek Power. arXiv.
  815. TurnTrout (2022). Inner and outer alignment decompose one hard problem into two extremely hard problems. AI Alignment Forum.
  816. TurnTrout (2023). Comment on “TurnTrout's shortform feed”. LessWrong.
  817. Turpin, M., Michael, J., Perez, E. & Bowman, S. R. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. arXiv.
  818. U.K. government (2023). Frontier AI: capabilities and risks – discussion paper. GOV.UK.
  819. U.S Defense Innovation Unit (2023). The Replicator Initiative.
  820. Uesato, J. et al. (2022). Solving math word problems with process- and outcome-based feedback. arXiv.
  821. UK AISI (2024). AI Safety Institute approach to evaluations. GOV.UK.
  822. University of Oxford (2024). Prof. Geoffrey Hinton - "Will digital intelligence replace biological intelligence?" Romanes Lecture. YouTube.
  823. Urbina, F., Lentzos, F., Invernizzi, C. & Ekins, S. (2022). Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence.
  824. US & UK AISI (2024). Pre-deployment evaluation of Anthropic’s upgraded Claude 3.5 Sonnet. AI Security Institute.
  825. US AI Safety Institute & UK AI Safety Institute (2024). US AISI and UK AISI Joint Pre-Deployment Test: Anthropic's Claude 3.5 Sonnet (October 2024 Release).
  826. Valle-Pérez, G., Camargo, C. Q. & Louis, A. A. (2018). Deep learning generalizes because the parameter-function map is biased towards simple functions. arXiv.
  827. Vanessa Kosoy (2023). The Learning-Theoretic Agenda: Status 2023. AI Alignment Forum.
  828. Variengien & Martinet (2024). AI Safety Institutes: Can countries meet the challenge?.
  829. Venkat Somala, Anson Ho & Séb Krier (2025). Three challenges facing compute-based AI policies.
  830. Veronika Blablová & Robi Rahman (2025). Why China isn’t about to leap ahead of the West on compute.
  831. Vika (2023). When discussing AI risks, talk about capabilities, not intelligence. AI Alignment Forum.
  832. Villalobos et al. (2024). Will we run out of data to train large language models?. Epoch AI.
  833. Villalobos, P., Ho, A., Sevilla, J., Besiroglu, T., Heim, L. & Hobbhahn, M. (2022). Will we run out of data? Limits of LLM scaling based on human-generated data. arXiv.
  834. Vlad Mikulik (2019). 2-D Robustness. LessWrong.
  835. Wan, S. et al. (2024). CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models. arXiv.
  836. Wang (2020). Journal of Artificial General Intelligence. Paradigm.
  837. Wang et al. (2022). Adversarial Policies Beat Superhuman Go AIs. arXiv.org.
  838. Wang, A. et al. (2019). SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems. arXiv.
  839. Wang, A., Singh, A., Michael, J., Hill, F., Levy, O. & Bowman, S. R. (2018). GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. arXiv.
  840. Wang, G. et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv.
  841. Wang, L. et al. (2023). A Survey on Large Language Model based Autonomous Agents. arXiv.
  842. Wang, S., Liu, S., Ye, W., You, J. & Gao, Y. (2024). EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data. arXiv.
  843. Wang, X. et al. (2022). Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv.
  844. Wang, X. et al. (2023). InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews. arXiv.
  845. Wasil, A. R., Clymer, J., Krueger, D., Dardaman, E., Campos, S. & Murphy, E. R. (2024). Affirmative safety: An approach to risk management for high-risk AI. arXiv.
  846. Wasil, A. R., Reed, T., Miller, J. W. & Barnett, P. (2024). Verification methods for international AI agreements. arXiv.
  847. Wei Dai (2019). AGI will drastically increase economies of scale. AI Alignment Forum.
  848. Wei, J. et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv.
  849. Wei, J. et al. (2022). Emergent Abilities of Large Language Models. arXiv.
  850. Weidinger, L. et al. (2021). Ethical and social risks of harm from Language Models. arXiv.
  851. Weidinger, L. et al. (2023). Sociotechnical Safety Evaluation of Generative AI Systems. arXiv.
  852. Whitehouse (2025). Fact Sheet: President Donald J. Trump Takes Action to Enhance America’s AI Leadership. The White House.
  853. Whitehouse (2025). Preventing Woke AI in the Federal Government. The White House.
  854. Wijk, H. et al. (2024). RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts. arXiv.
  855. Wikipedia (2022). AI-complete. Wikipedia.
  856. William MacAskill & Rose Hadshar (2025). Intelsat as a Model for International AGI Governance.
  857. William MacAskill (2025). Better Futures.
  858. Williams et al. (2025). Forecasting LLM-enabled Biorisk and the Efficacy of Safeguards. Forecasting Research Institute.
  859. Willis et al. (2024). Assessing Global Catastrophic and Existential Risks.
  860. WillPetillo, Sean Herrington, Spencer Ames, Adebayo Mubarak & Can Narin (2025). Case Studies in Simulators and Agents. LessWrong.
  861. Wiseman & McClements (2025). How much economic growth from AI should we expect, how soon?.
  862. Wojton, H. M., Porter, D. J. & Dennis, J. W. (2020). Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review.
  863. Wolf, Y., Wies, N., Avnery, O., Levine, Y. & Shashua, A. (2023). Fundamental Limitations of Alignment in Large Language Models. arXiv.
  864. Wongkamjan, W. et al. (2024). More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play. arXiv.
  865. Wooldridge (2021). A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books.
  866. Wooldridge (2024). AI’s simple solution to rail problems: stop all trains running. The Telegraph.
  867. World Economic Forum (2025). The Dawn of Artificial General Intelligence? | World Economic Forum Annual Meeting 2025. YouTube.
  868. WorldCoin (2024). Proof of personhood: What it is and why it’s needed. World.
  869. Wu, J. et al. (2021). Recursively Summarizing Books with Human Feedback. arXiv.
  870. Xiang (2023). 'He Would Still Be Here': Man Dies by Suicide After Talking with AI Chatbot, Widow Says. VICE.
  871. Xiao, Y. & Wang, W. Y. (2021). On Hallucination and Predictive Uncertainty in Conditional Language Generation. arXiv.
  872. Xu, F. et al. (2025). Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models. arXiv.
  873. Xu, H., Zhao, R., Zhu, L., Du, J. & He, Y. (2024). OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models. arXiv.
  874. Xu, Y., Deng, B., Wang, J., Jing, Y., Pan, J. & He, S. (2022). High-resolution Face Swapping via Latent Semantics Disentanglement. arXiv.
  875. Yafah Edelman & Anson Ho (2025). Compute scaling will slow down due to increasing lead times.
  876. Yampolskiy, R. V. (2024). AI: Unexplainable, Unpredictable, Uncontrollable.
  877. Yampolsky; (2024). Transcript for Roman Yampolskiy: Dangers of Superintelligent AI. Lex Fridman.
  878. Yang et. al; (2022). Chain of Thought Imitation with Procedure Cloning. arXiv.org.
  879. Yang, J., Prabhakar, A., Narasimhan, K. & Yao, S. (2023). InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. arXiv.
  880. Yao, S. et al. (2023). Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv.org.
  881. Yap (2024). How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs). Gold Penguin.
  882. Ye, W., Liu, S., Kurutach, T., Abbeel, P. & Gao, Y. (2021). Mastering Atari Games with Limited Data. arXiv.
  883. You et al. (2025). How much power will frontier AI training demand in 2030?. Epoch AI.
  884. Yu, J. et al. (2022). Scaling Autoregressive Models for Content-Rich Text-to-Image Generation. arXiv.
  885. Yudkowsky (2002). The AI-Box Experiment:. Eliezer S. Yudkowsky.
  886. Yudkowsky (2015). Vingean uncertainty.
  887. Yudkowsky (2022). AGI Ruin: A List of Lethalities - Machine Intelligence Research Institute.
  888. Yudkowsky (2023). Pausing AI Developments Isn't Enough. We Need to Shut it All Down - Machine Intelligence Research Institute.
  889. Yudkowsky, E. (2004). Coherent Extrapolated Volition.
  890. Yudkowsky, E. (2013). Intelligence Explosion Microeconomics.
  891. zac_kenton et al. (2022). Clarifying AI X-risk. LessWrong.
  892. Zaidan, E. & Ibrahim, I. A. (2024). AI Governance in a Complex and Rapidly Changing Regulatory Landscape: A Global Perspective. Humanities and Social Sciences Communications.
  893. Zellers, R., Bisk, Y., Schwartz, R. & Choi, Y. (2018). SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference. arXiv.
  894. Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A. & Choi, Y. (2019). HellaSwag: Can a Machine Really Finish Your Sentence?. arXiv.
  895. Zhang et al. (2025). A Three-Layered Framework: An AI Governance Guide for Global Policymakers.
  896. Zhang, B., Anderljung, M., Kahn, L., Dreksler, N., Horowitz, M. C. & Dafoe, A. (2021). Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers. arXiv.
  897. Zhang, F. et al. (2024). HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks. arXiv.
  898. Zhang, G., Yan, C., Ji, X., Zhang, T., Zhang, T. & Xu, W. (2017). DolphinAtack: Inaudible Voice Commands. arXiv.
  899. Zhang, Z., Bai, F., Gao, J. & Yang, Y. (2023). ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models. arXiv.
  900. Zhao, M., Zhang, L., Ye, J., Lu, H., Yin, B. & Wang, X. (2024). Adversarial Training: A Survey. arXiv.
  901. Zhou, A., Yan, K., Shlapentokh-Rothman, M., Wang, H. & Wang, Y. (2023). Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. arXiv.
  902. Zhou, L. et al. (2023). Predictable Artificial Intelligence. arXiv.
  903. Zhu, H. et al. (2024). Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics. arXiv.
  904. Zhu, Y., Li, Q., Wang, J., Xu, C. & Sun, Z. (2021). One Shot Face Swapping on Megapixels. arXiv.
  905. Zhuo, T. Y. et al. (2024). BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. arXiv.
  906. Ziegler, D. M. et al. (2019). Fine-Tuning Language Models from Human Preferences. arXiv.
  907. Žiga Avsec & Natasha Latysheva (2025). AlphaGenome: AI for better understanding the genome.
  908. Zou, A. et al. (2023). Representation Engineering: A Top-Down Approach to AI Transparency. arXiv.
  909. Zvi (2023). OpenAI Launches Superalignment Taskforce. LessWrong.
  910. Zvi (2025). On MAIM and Superintelligence Strategy. LessWrong.
  911. Zvi (2025). The Paris AI Anti-Safety Summit. LessWrong.
  912. Zwetsloot, R. et al. (2021). Skilled and Mobile: Survey Evidence of AI Researchers' Immigration Preferences. arXiv.