Developing strategies to ensure the safety of increasingly capable AI systems presents unique and significant challenges. These difficulties stem from the nature of AI itself, the current state of the research field, and the complexity of the risks involved.
We do not know how to train systems to robustly behave well.
The Nature of the Problem #
Several intrinsic properties make AI safety a particularly hard problem :
AI risk is an emerging problem that is still poorly understood. AI risk is a relatively new field dealing with rapidly evolving technology. Our understanding of the full spectrum of potential failure modes and long-term consequences is incomplete. Devising robust safeguards for technologies that do not yet exist, but which could have profoundly negative outcomes, is inherently difficult.
The field is still pre-paradigmatic. There is currently no single, universally accepted paradigm for AI safety. Researchers disagree on fundamental aspects, including the most likely threat models (e.g., sudden takeover (Yudkowsky, 2022) vs. gradual loss of control (Critch, 2021)), and the most promising solution paths. The research agendas of some researchers seem scarcely useful to others, and one of the favorite activities of alignment researchers is to criticize each other’s plan constructively.
AIs are black boxes that are trained, not built. Modern deep learning models are "black boxes." While we know how to train them, the specific algorithms they learn and their internal decision-making processes remain largely opaque. These models lack the apparent modularity common in traditional software engineering, making it difficult to decompose, analyze, or verify their behavior. Progress in interpretability has yet to fully overcome this challenge.
Complexity is the source of many blind spots. The sheer complexity of large AI models means that unexpected and potentially harmful behaviors can emerge without warning. Issues like "glitch tokens", e.g., "SolidGoldMagikarp", causing erratic behavior in GPT models (Rumbelow & Watkins, 2023), demonstrate how unforeseen interactions between components (like tokenizers and training data ) can lead to failures. When GPT encounters this infrequent word, it behaves unpredictably and erratically. This phenomenon occurs because GPT uses a tokenizer to break down sentences into tokens (sets of letters such as words or combinations of letters and numbers), and the token "SolidGoldMagikarp" was present in the tokenizer's dataset but not in the GPT model's dataset. This blind spot is not an isolated incident.
Creating an exhaustive risk framework is difficult. There are many, many different classifications of risk scenarios that focus on various types of harm (Critch & Russel, 2023; Hendrycks et al., 2023; Slattery et al., 2024). Proposing a solid single-risk model beyond criticism is extremely difficult, and the risk scenarios often contain a degree of vagueness. No scenario captures most of the probability mass, and there is a wide diversity of potentially catastrophic scenarios (Pace, 2020).
Some arguments that seem initially appealing may be misleading. For example, the principal author of the paper (Turner et al., 2023) presenting a mathematical result on instrumental convergence, Alex Turner, now believes his theorem is a poor way to think about the problem (Turner, 2023). Some other classical arguments have been criticized recently, like the counting argument (AI Optimists, 2023) or the utility maximization frameworks, which will be discussed in the chapter "Goal Misgeneralization".
We may not have time. Many experts in the field believe that AGI, and shortly after ASI, could arrive before 2030. We need to solve these massive problems, or at least set the strategy for the launch, before it happens. For example, the scenario AI-2027, is based on a detailed forecasting of timelines of AGI arrival, and argues for a strong likelihood of ASI before the end of the decade.
Many essential terms in AI safety are complicated to define. They often require knowledge in philosophy (epistemology, theory of mind) and AI. For instance, to determine if an AI is an agent, one must clarify "what does agency mean?" which, as we'll see in later chapters, requires nuance and may be an intrinsically ill-defined and fuzzy term. Some topics in AI safety are so challenging to grasp and are thought to be non-scientific in the machine learning community, such as discussing situational awareness (Hinton, 2024) or why AI might be able to "really understand". These concepts are far from consensus among philosophers and AI researchers and require a lot of caution.
A simple solution probably doesn’t exist. For instance, the response to climate change is not just one measure, like saving electricity in winter at home. A whole range of potentially different solutions must be applied. Just as there are various problems to consider when building an airplane, similarly, when training and deploying an AI, a range of issues could arise, requiring precautions and various security measures.
AI safety is hard to measure. Working on the problem can lead to an illusion of understanding, thereby creating the illusion of control. AI safety lacks clear feedback loops. Progress in AI capability advancement is relatively easy to measure and benchmark, while progress in safety is comparatively harder to measure. For example, it’s much easier to monitor the inference speed than to monitor the truthfulness of a system or monitor its safety properties.
Uncertainty and Disagreement #
The pre-paradigmatic nature of AI safety leads to significant disagreements among experts. These differences in perspective are crucial to understanding when evaluating these proposed strategies.
The consequences of failures in AI alignment are steeped in uncertainty. New insights could challenge many high-level considerations discussed in this textbook. For instance, Zvi Mowshowitz has compiled a list of central questions marked by significant uncertainty (Mowshowitz, 2023). For example, what worlds count as catastrophic versus non-catastrophic? What would count as a non-catastrophic outcome? What is valuable? What do we care about? If answered differently, these questions could significantly alter one's estimate of the likelihood and severity of catastrophes stemming from unaligned AGI.
Divergent Worldviews. These disagreements often stem from fundamentally different worldviews. Some experts, like Robin Hanson, may approach AI risk through economic or evolutionary lenses, potentially leading to different conclusions about takeoff speeds and the likelihood of stable control compared to those focusing on agent foundations or technical alignment failures (Hanson, 2023). Others, like Richard Sutton, have expressed views suggesting an acceptance or even embrace of AI potentially succeeding humanity, framing it as a natural evolutionary step rather than an existential catastrophe (Sutton, 2023). These differing philosophical stances influence strategic priorities.
Safety Washing #
The combination of high stakes, public concern, and lack of consensus creates fertile ground for "safety washing"—the practice of misleadingly portraying AI products, research, or practices as safer or more aligned with safety goals than they actually are (Vaintrob, 2023).
Safety washing can create a false sense of security. Companies developing powerful AI face incentives to appear safety-conscious to appease the public, regulators, and potential employees. Safetywashing can involve overstating the safety benefits of certain features, focusing on less critical aspects of safety while downplaying existential risks, or funding/conducting research that primarily advances capabilities under the guise of safety. This can lead to insufficient risk mitigation efforts (Lizka, 2023). It can misdirect funding and talent towards less impactful work and make it harder to build a genuine scientific consensus on the true state of AI safety.
Assessing progress in safety is tricky. Even with the intention to help, actions might have a net negative impact (e.g., from second-order effects, like accelerating deployment of dangerous technologies), and determining the contribution's impact is far from trivial. For example, the impact of reinforcement learning from human feedback (RLHF), currently used to instruction-tune and make ChatGPT safer, is still debated in the community (Christiano, 2023). One reason the impact of RLHF may be negative is that this technique may create an illusion of alignment that would make spotting deceptive alignment even more challenging. The alignment of the systems trained through RLHF is shallow (Casper et al., 2023), and the alignment properties might break with future, more situationally aware models. Similarly, certain interpretability work faces dual-use concerns (Wache, 2023). Some argue that much current "AI safety" research solves easy problems that primarily benefit developers economically, potentially speeding up capabilities rather than meaningfully reducing existential risk (catubc, 2024). As a consequence, even well-intentioned research might inadvertently accelerate risks.
References
- AI Optimists (2023). AI is easy to control. AI Optimism.AI Optimists. (2023, November 29). AI is easy to control. AI Optimism. https://optimists.ai/2023/11/28/ai-is-easy-to-controlAI Optimists. 2023. “AI Is Easy to Control”. AI Optimism, November 29. https://optimists.ai/2023/11/28/ai-is-easy-to-control.AI Optimists. “AI Is Easy to Control”. AI Optimism, 29 Nov. 2023, https://optimists.ai/2023/11/28/ai-is-easy-to-control.AI Optimists. AI is easy to control. AI Optimism https://optimists.ai/2023/11/28/ai-is-easy-to-control (2023).AI Optimists, “AI is easy to control”, AI Optimism. [Online]. Available: https://optimists.ai/2023/11/28/ai-is-easy-to-control
- Andrew_Critch (2021). What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs). AI Alignment Forum.Andrew_Critch. (2021, March 31). What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs). AI Alignment Forum. https://alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnosticAndrew_Critch. 2021. “What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs)”. AI Alignment Forum, March 31. https://alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic.Andrew_Critch. “What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs)”. AI Alignment Forum, 31 Mar. 2021, https://alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic.Andrew_Critch. What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs). AI Alignment Forum https://alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic (2021).Andrew_Critch, “What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs)”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic
- Anthropic (2023). Anthropic's core views on AI safety.Anthropic. (2023). Anthropic's core views on AI safety. https://anthropic.com/news/core-views-on-ai-safetyAnthropic. 2023. “Anthropic's Core Views on AI Safety”. https://anthropic.com/news/core-views-on-ai-safety.Anthropic. Anthropic's Core Views on AI Safety. 2023, https://anthropic.com/news/core-views-on-ai-safety.Anthropic. Anthropic's core views on AI safety. https://anthropic.com/news/core-views-on-ai-safety (2023).Anthropic, “Anthropic's core views on AI safety”. [Online]. Available: https://anthropic.com/news/core-views-on-ai-safety
- Ben Pace (2020). What Failure Looks Like: Distilling the Discussion. AI Alignment Forum.Ben Pace. (2020, July 29). What Failure Looks Like: Distilling the Discussion. AI Alignment Forum. https://alignmentforum.org/posts/6jkGf5WEKMpMFXZp2/what-failure-looks-like-distilling-the-discussionBen Pace. 2020. “What Failure Looks Like: Distilling the Discussion”. AI Alignment Forum, July 29. https://alignmentforum.org/posts/6jkGf5WEKMpMFXZp2/what-failure-looks-like-distilling-the-discussion.Ben Pace. “What Failure Looks Like: Distilling the Discussion”. AI Alignment Forum, 29 July 2020, https://alignmentforum.org/posts/6jkGf5WEKMpMFXZp2/what-failure-looks-like-distilling-the-discussion.Ben Pace. What Failure Looks Like: Distilling the Discussion. AI Alignment Forum https://alignmentforum.org/posts/6jkGf5WEKMpMFXZp2/what-failure-looks-like-distilling-the-discussion (2020).Ben Pace, “What Failure Looks Like: Distilling the Discussion”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/6jkGf5WEKMpMFXZp2/what-failure-looks-like-distilling-the-discussion
- Casper, S. et al. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. arXiv.Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., … Hadfield-Menell, D. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. In arXiv. https://arxiv.org/abs/2307.15217Casper, S., X. Davies, C. Shi, et al. 2023. “Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback”. In arXiv. Preprint, July 27. https://arxiv.org/abs/2307.15217.Casper, S., et al. “Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback”. arXiv, 27 July 2023, https://arxiv.org/abs/2307.15217.Casper, S. et al. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. arXiv Preprint at https://arxiv.org/abs/2307.15217 (2023).S. Casper et al., “Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback”, Jul. 27, 2023. [Online]. Available: https://arxiv.org/abs/2307.15217
- Critch, A. & Russell, S. (2023). TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI. arXiv.Critch, A., & Russell, S. (2023). TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI. In arXiv. https://arxiv.org/abs/2306.06924Critch, A., and S. Russell. 2023. “TASRA: A Taxonomy and Analysis of Societal-Scale Risks from AI”. In arXiv. Preprint, June 12. https://arxiv.org/abs/2306.06924.Critch, A., and S. Russell. “TASRA: A Taxonomy and Analysis of Societal-Scale Risks from AI”. arXiv, 12 June 2023, https://arxiv.org/abs/2306.06924.Critch, A. & Russell, S. TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI. arXiv Preprint at https://arxiv.org/abs/2306.06924 (2023).A. Critch and S. Russell, “TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI”, Jun. 12, 2023. [Online]. Available: https://arxiv.org/abs/2306.06924
- Eliezer Yudkowsky (2022). AGI Ruin: A List of Lethalities. AI Alignment Forum.Eliezer Yudkowsky. (2022, June 5). AGI Ruin: A List of Lethalities. AI Alignment Forum. https://alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalitiesEliezer Yudkowsky. 2022. “AGI Ruin: A List of Lethalities”. AI Alignment Forum, June 5. https://alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities.Eliezer Yudkowsky. “AGI Ruin: A List of Lethalities”. AI Alignment Forum, 5 June 2022, https://alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities.Eliezer Yudkowsky. AGI Ruin: A List of Lethalities. AI Alignment Forum https://alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities (2022).Eliezer Yudkowsky, “AGI Ruin: A List of Lethalities”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities
- Hanson (2023). AI Risk, Again.Hanson. (2023). AI Risk, Again. https://www.overcomingbias.com/p/ai-risk-againHanson. 2023. “AI Risk, Again”. https://www.overcomingbias.com/p/ai-risk-again.Hanson. AI Risk, Again. 2023, https://www.overcomingbias.com/p/ai-risk-again.Hanson. AI Risk, Again. https://www.overcomingbias.com/p/ai-risk-again (2023).Hanson, “AI Risk, Again”. [Online]. Available: https://www.overcomingbias.com/p/ai-risk-again
- Hendrycks, D., Mazeika, M. & Woodside, T. (2023). An Overview of Catastrophic AI Risks. arXiv.Hendrycks, D., Mazeika, M., & Woodside, T. (2023). An Overview of Catastrophic AI Risks. In arXiv. https://arxiv.org/abs/2306.12001Hendrycks, D., M. Mazeika, and T. Woodside. 2023. “An Overview of Catastrophic AI Risks”. In arXiv. Preprint, June 21. https://arxiv.org/abs/2306.12001.Hendrycks, D., et al. “An Overview of Catastrophic AI Risks”. arXiv, 21 June 2023, https://arxiv.org/abs/2306.12001.Hendrycks, D., Mazeika, M. & Woodside, T. An Overview of Catastrophic AI Risks. arXiv Preprint at https://arxiv.org/abs/2306.12001 (2023).D. Hendrycks, M. Mazeika, and T. Woodside, “An Overview of Catastrophic AI Risks”, Jun. 21, 2023. [Online]. Available: https://arxiv.org/abs/2306.12001
- Jessica Rumbelow & mwatkins (2023). SolidGoldMagikarp (plus, prompt generation). AI Alignment Forum.Jessica Rumbelow, & mwatkins. (2023, February 5). SolidGoldMagikarp (plus, prompt generation). AI Alignment Forum. https://alignmentforum.org/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generationJessica Rumbelow, and mwatkins. 2023. “SolidGoldMagikarp (plus, Prompt Generation)”. AI Alignment Forum, February 5. https://alignmentforum.org/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation.Jessica Rumbelow, and mwatkins. “SolidGoldMagikarp (plus, Prompt Generation)”. AI Alignment Forum, 5 Feb. 2023, https://alignmentforum.org/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation.Jessica Rumbelow & mwatkins. SolidGoldMagikarp (plus, prompt generation). AI Alignment Forum https://alignmentforum.org/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation (2023).Jessica Rumbelow and mwatkins, “SolidGoldMagikarp (plus, prompt generation)”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation
- Lizka (2023). Beware safety-washing. EA Forum.Lizka. (2023, January 13). Beware safety-washing. EA Forum. Internet Archive (https://web.archive.org/web/20260519093736/https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing). https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washingLizka. 2023. “Beware Safety-washing”. EA Forum, January 13. Https://web.archive.org/web/20260519093736/https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing. Internet Archive. https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing.Lizka. “Beware Safety-washing”. EA Forum, 13 Jan. 2023, Internet Archive, https://web.archive.org/web/20260519093736/https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing, https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing.Lizka. Beware safety-washing. EA Forum https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing (2023).Lizka, “Beware safety-washing”, EA Forum. Accessed: May 19, 2026. [Online]. Available: https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing
- Magdalena Wache (2023). Technical AI Safety Research Landscape [Slides]. LessWrong.Magdalena Wache. (2023, September 18). Technical AI Safety Research Landscape [Slides]. LessWrong. https://lesswrong.com/posts/x2n7mBLryDXuLwGhx/technical-ai-safety-research-landscape-slidesMagdalena Wache. 2023. “Technical AI Safety Research Landscape [Slides]”. LessWrong, September 18. https://lesswrong.com/posts/x2n7mBLryDXuLwGhx/technical-ai-safety-research-landscape-slides.Magdalena Wache. “Technical AI Safety Research Landscape [Slides]”. LessWrong, 18 Sept. 2023, https://lesswrong.com/posts/x2n7mBLryDXuLwGhx/technical-ai-safety-research-landscape-slides.Magdalena Wache. Technical AI Safety Research Landscape [Slides]. LessWrong https://lesswrong.com/posts/x2n7mBLryDXuLwGhx/technical-ai-safety-research-landscape-slides (2023).Magdalena Wache, “Technical AI Safety Research Landscape [Slides]”, LessWrong. [Online]. Available: https://lesswrong.com/posts/x2n7mBLryDXuLwGhx/technical-ai-safety-research-landscape-slides
- Mowshowitz (2023). The Crux List.Mowshowitz. (2023). The Crux List. https://thezvi.substack.com/p/the-crux-listMowshowitz. 2023. The Crux List. Edition. https://thezvi.substack.com/p/the-crux-list.Mowshowitz. The Crux List. 2023, https://thezvi.substack.com/p/the-crux-list.Mowshowitz. The Crux List. https://thezvi.substack.com/p/the-crux-list (2023).Mowshowitz, “The Crux List”. [Online]. Available: https://thezvi.substack.com/p/the-crux-list
- paulfchristiano (2023). Thoughts on the impact of RLHF research. AI Alignment Forum.paulfchristiano. (2023, January 25). Thoughts on the impact of RLHF research. AI Alignment Forum. https://alignmentforum.org/posts/vwu4kegAEZTBtpT6p/thoughts-on-the-impact-of-rlhf-researchpaulfchristiano. 2023. “Thoughts on the Impact of RLHF Research”. AI Alignment Forum, January 25. https://alignmentforum.org/posts/vwu4kegAEZTBtpT6p/thoughts-on-the-impact-of-rlhf-research.paulfchristiano. “Thoughts on the Impact of RLHF Research”. AI Alignment Forum, 25 Jan. 2023, https://alignmentforum.org/posts/vwu4kegAEZTBtpT6p/thoughts-on-the-impact-of-rlhf-research.paulfchristiano. Thoughts on the impact of RLHF research. AI Alignment Forum https://alignmentforum.org/posts/vwu4kegAEZTBtpT6p/thoughts-on-the-impact-of-rlhf-research (2023).paulfchristiano, “Thoughts on the impact of RLHF research”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/vwu4kegAEZTBtpT6p/thoughts-on-the-impact-of-rlhf-research
- Rich Sutton (2023). AI Succession. YouTube.Rich Sutton. (2023). AI Succession [Video recording]. In YouTube. https://www.youtube.com/watch?v=NgHFMolXs3URich Sutton. 2023. “AI Succession”. YouTube. https://www.youtube.com/watch?v=NgHFMolXs3U.Rich Sutton. “AI Succession”. YouTube, 2023, https://www.youtube.com/watch?v=NgHFMolXs3U.Rich Sutton. AI Succession. YouTube (2023).Rich Sutton, AI Succession, (2023). [Online Video]. Available: https://www.youtube.com/watch?v=NgHFMolXs3U
- Slattery, P. et al. (2024). The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence. arXiv.Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2024). The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence. In arXiv. https://doi.org/10.1016/j.patter.2026.101517Slattery, P., A. K. Saeri, E. A. C. Grundy, et al. 2024. “The AI Risk Repository: A Meta-review, Database, and Taxonomy of Risks from Artificial Intelligence”. In arXiv. Preprint, August 14. https://doi.org/10.1016/j.patter.2026.101517.Slattery, P., et al. “The AI Risk Repository: A Meta-review, Database, and Taxonomy of Risks from Artificial Intelligence”. arXiv, 14 Aug. 2024, https://doi.org/10.1016/j.patter.2026.101517.Slattery, P. et al. The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence. arXiv Preprint at https://doi.org/10.1016/j.patter.2026.101517 (2024).P. Slattery et al., “The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence”, Aug. 14, 2024. doi: 10.1016/j.patter.2026.101517.
- Turner, A. M., Smith, L., Shah, R., Critch, A. & Tadepalli, P. (2019). Optimal Policies Tend to Seek Power. arXiv.Turner, A. M., Smith, L., Shah, R., Critch, A., & Tadepalli, P. (2019). Optimal Policies Tend to Seek Power. In arXiv. https://arxiv.org/abs/1912.01683Turner, A. M., L. Smith, R. Shah, A. Critch, and P. Tadepalli. 2019. “Optimal Policies Tend to Seek Power”. In arXiv. Preprint, December 3. https://arxiv.org/abs/1912.01683.Turner, A. M., et al. “Optimal Policies Tend to Seek Power”. arXiv, 3 Dec. 2019, https://arxiv.org/abs/1912.01683.Turner, A. M., Smith, L., Shah, R., Critch, A. & Tadepalli, P. Optimal Policies Tend to Seek Power. arXiv Preprint at https://arxiv.org/abs/1912.01683 (2019).A. M. Turner, L. Smith, R. Shah, A. Critch, and P. Tadepalli, “Optimal Policies Tend to Seek Power”, Dec. 03, 2019. [Online]. Available: https://arxiv.org/abs/1912.01683
- TurnTrout (2023). Comment on “TurnTrout's shortform feed”. LessWrong.TurnTrout. (2023, June 1). Comment on “TurnTrout's shortform feed”. LessWrong. https://lesswrong.com/posts/dqSwccGTWyBgxrR58/turntrout-s-shortform-feed?commentId=Sw89AxHGJ5j7E7ETfTurnTrout. 2023. “Comment on “TurnTrout's Shortform Feed””. LessWrong, June 1. https://lesswrong.com/posts/dqSwccGTWyBgxrR58/turntrout-s-shortform-feed?commentId=Sw89AxHGJ5j7E7ETf.TurnTrout. “Comment on “TurnTrout's Shortform Feed””. LessWrong, 1 June 2023, https://lesswrong.com/posts/dqSwccGTWyBgxrR58/turntrout-s-shortform-feed?commentId=Sw89AxHGJ5j7E7ETf.TurnTrout. Comment on “TurnTrout's shortform feed”. LessWrong https://lesswrong.com/posts/dqSwccGTWyBgxrR58/turntrout-s-shortform-feed?commentId=Sw89AxHGJ5j7E7ETf (2023).TurnTrout, “Comment on “TurnTrout's shortform feed””, LessWrong. [Online]. Available: https://lesswrong.com/posts/dqSwccGTWyBgxrR58/turntrout-s-shortform-feed?commentId=Sw89AxHGJ5j7E7ETf
- University of Oxford (2024). Prof. Geoffrey Hinton - "Will digital intelligence replace biological intelligence?" Romanes Lecture. YouTube.University of Oxford. (2024). Prof. Geoffrey Hinton - "Will digital intelligence replace biological intelligence?" Romanes Lecture [Video recording]. In YouTube. https://www.youtube.com/watch?v=N1TEjTeQeg0University of Oxford. 2024. “Prof. Geoffrey Hinton - "Will Digital Intelligence Replace Biological Intelligence?" Romanes Lecture”. YouTube. https://www.youtube.com/watch?v=N1TEjTeQeg0.University of Oxford. “Prof. Geoffrey Hinton - "Will Digital Intelligence Replace Biological Intelligence?" Romanes Lecture”. YouTube, 2024, https://www.youtube.com/watch?v=N1TEjTeQeg0.University of Oxford. Prof. Geoffrey Hinton - "Will Digital Intelligence Replace Biological Intelligence?" Romanes Lecture. YouTube (2024).University of Oxford, Prof. Geoffrey Hinton - "Will digital intelligence replace biological intelligence?" Romanes Lecture, (2024). [Online Video]. Available: https://www.youtube.com/watch?v=N1TEjTeQeg0
Was this section useful?
Thank you for your feedback
Your input helps improve the Atlas.