How we define problems directly impacts which strategies we pursue in solving that problem. In a new and evolving field like AI safety, clearly defined terms are essential for effective communication and research. Ambiguity leads to miscommunication, hinders collaboration, obscures disagreements, and facilitates safety washing (Ren et al., 2024; Lizka, 2023). The terms we use reflect our assumptions about the nature of the problems we're trying to solve and shape the solutions we develop. Terms like "alignment" and "safety" are used with varying meanings, reflecting different underlying assumptions about the nature of the problem and the research goals. The goal of this section is to explain different perspectives on these words, what specific safety strategies aim to achieve, and establish how our text will utilize them.
AI Safety #
AI safety ensures AI systems do not cause harm to humans or the environment. It encompasses the broadest range of research and engineering practices focused on preventing harmful outcomes from AI systems. While alignment focuses on aspects such as an AI's goals and intentions, safety addresses a broader range of concerns (Rudner et al., 2021). It is concerned with ensuring that AI systems do not inadvertently or deliberately cause harm or danger to humans or the environment. AI safety research seeks to identify the causes of unintended AI behavior and develop tools for ensuring safe and reliable operation. It can include technical subfields like robustness (ensuring reliable performance, including against adversarial attacks), monitoring (observing AI behavior), and capability control (limiting potentially dangerous abilities).
AI Alignment #
AI alignment aims to ensure AI systems act in accordance with human intentions and values. Alignment is a subset of safety that focuses specifically on the technical problem of ensuring AI objectives align with human intentions and values. Theoretically, a system could be aligned but unsafe (e.g., competently pursuing the wrong goal due to misspecification) or safe but unaligned (e.g., constrained by control mechanisms despite misaligned objectives). While this sounds straightforward, the precise scope varies significantly across research communities. We already saw a brief definition of alignment in the previous chapter, but this section offers a more nuanced perspective on the various definitions that we could potentially work with.
AI Ethics #
AI ethics is the field that examines the moral principles and societal implications of AI systems. It addresses the ethical considerations of potential societal upheavals resulting from AI advancements and the moral frameworks necessary to navigate these changes. The core of AI ethics lies in ensuring that AI developments are aligned with human dignity, fairness, and societal well-being, through a deep understanding of their broader societal impact. Research in AI ethics would encompass, for example, privacy norms, identifying and mitigating bias in systems (Huang et al., 2022; Harvard, 2025; Khan et al., 2022).
Ethics complements technical safety approaches by providing normative guidance on what constitutes beneficial AI outcomes. Alignment focuses on ensuring AI systems pursue intended objectives, research in ethics focuses on which objectives are worth pursuing (Huang et al., 2023; LaCroix & Luccioni, 2022). AI ethics might also include discussions of digital rights and potentially even the rights of digital minds, and AIs in the future.
This chapter focuses primarily on safety frameworks as they inform technical safety and governance strategies rather than exploring ethics, meta-ethics or digital rights.
AI Control #
AI control ensures systems remain under human authority despite potential misalignment. AI control implements mechanisms to ensure AI systems remain under human direction, even when they might act against our interests. Unlike alignment approaches that focus on giving AI systems the right goals, control addresses what happens if those goals diverge from human intentions (Greenblatt et al., 2024).
Control and alignment work as complementary safety approaches. While alignment aims to prevent preference divergence by designing systems with the right objectives, control creates layers of security that function even when alignment fails. Control measures include monitoring AI actions, restricting system capabilities, human auditing processes, and mechanisms to terminate AI systems when necessary (Greenblatt et al., 2023). Some researchers argue that even if alignment is needed for superintelligence-level AIs, control through monitoring may be a working strategy for less capable systems (Greenblatt et al., 2024). Ideally, an AGI would be aligned and controllable, meaning it would have the right goals and be subject to human oversight and intervention if something goes wrong.
The control line of AI safety work is discussed in much more detail in our chapter on AI evaluations.
Footnotes
While AI alignment does not necessarily encompass all systemic risks and misuse, there is some overlap. Some alignment techniques could help mitigate specific misuse scenarios—for instance, alignment methods could ensure that models refuse to cooperate with users intending to use AI for harmful purposes, such as bioterrorism. Similarly, from a systemic risk perspective, a well-aligned AI might recognize and refuse to participate in problematic processes embedded within systems, such as financial markets. However, challenges remain, as malicious actors might attempt to circumvent these protections through targeted fine-tuning of models for harmful purposes, and in this case, even a perfectly aligned model wouldn't be able to resist
↩
References
- Dewey, D. (2011). Learning What to Value. Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings.Dewey, D. (2011). Learning What to Value. Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings, Lecture Notes in Computer Science, 6830, 309–314. https://doi.org/10.1007/978-3-642-22887-2_35Dewey, D. 2011. “Learning What to Value”. Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings (Berlin), Lecture Notes in Computer Science, vol. 6830: 309–14. https://doi.org/10.1007/978-3-642-22887-2_35.Dewey, D. “Learning What to Value”. Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings [Berlin], Lecture Notes in Computer Science, vol. 6830, 2011, pp. 309–14, https://doi.org/10.1007/978-3-642-22887-2_35.Dewey, D. Learning What to Value. in Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings vol. 6830 309–314 (Springer, Berlin, 2011).D. Dewey, “Learning What to Value”, in Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings, in Lecture Notes in Computer Science, vol. 6830. Berlin: Springer, 2011, pp. 309–314. doi: 10.1007/978-3-642-22887-2_35.
- Gil (2023). Don't Call It AI Alignment. EA Forum.Gil. (2023). Don't Call It AI Alignment. EA Forum. Internet Archive (https://web.archive.org/web/20260516150521/https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment). https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignmentGil. 2023. “Don't Call It AI Alignment”. EA Forum. Https://web.archive.org/web/20260516150521/https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment. Internet Archive. https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment.Gil. “Don't Call It AI Alignment”. EA Forum, 2023, Internet Archive, https://web.archive.org/web/20260516150521/https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment, https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment.Gil. Don't Call It AI Alignment. EA Forum https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment (2023).Gil, “Don't Call It AI Alignment”, EA Forum. Accessed: May 16, 2026. [Online]. Available: https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment
- Harvard (2025). What is AI ethics?. Harvard FAS | Mignone Center for Career Success.Harvard. (2025). What is AI ethics?. Internet Archive (https://web.archive.org/web/20260123144850/https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics/). Harvard FAS | Mignone Center for Career Success. https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethicsHarvard. 2025. “What Is AI Ethics?”. Harvard FAS | Mignone Center for Career Success. Https://web.archive.org/web/20260123144850/https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics/. Internet Archive. https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics.Harvard. “What Is AI Ethics?”. Harvard FAS | Mignone Center for Career Success, 2025, Internet Archive, https://web.archive.org/web/20260123144850/https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics/, https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics.Harvard. What is AI ethics?. Harvard FAS | Mignone Center for Career Success https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics (2025).Harvard, “What is AI ethics?”, Harvard FAS | Mignone Center for Career Success. Accessed: Jan. 23, 2026. [Online]. Available: https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics
- Huang et al. (2023). An Overview of Artificial Intelligence Ethics.Huang et al. (2023). An Overview of Artificial Intelligence Ethics. https://ieeexplore.ieee.org/abstract/document/9844014Huang et al. 2023. “An Overview of Artificial Intelligence Ethics”. https://ieeexplore.ieee.org/abstract/document/9844014.Huang et al. An Overview of Artificial Intelligence Ethics. 2023, https://ieeexplore.ieee.org/abstract/document/9844014.Huang et al. An Overview of Artificial Intelligence Ethics. https://ieeexplore.ieee.org/abstract/document/9844014 (2023).Huang et al., “An Overview of Artificial Intelligence Ethics”. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/9844014
- Jonker et al. (2024). What Is AI Alignment?. IBM.Jonker et al. (2024, October 16). What Is AI Alignment?. IBM. https://ibm.com/think/topics/ai-alignmentJonker et al. 2024. “What Is AI Alignment?”. IBM, October 16. https://ibm.com/think/topics/ai-alignment.Jonker et al. “What Is AI Alignment?”. IBM, 16 Oct. 2024, https://ibm.com/think/topics/ai-alignment.Jonker et al. What Is AI Alignment?. IBM https://ibm.com/think/topics/ai-alignment (2024).Jonker et al., “What Is AI Alignment?”, IBM. [Online]. Available: https://ibm.com/think/topics/ai-alignment
- Khan, A. A. et al. (2021). Ethics of AI: A Systematic Literature Review of Principles and Challenges. arXiv.Khan, A. A., Badshah, S., Liang, P., Khan, B., Waseem, M., Niazi, M., & Akbar, M. A. (2021). Ethics of AI: A Systematic Literature Review of Principles and Challenges. In arXiv. https://arxiv.org/abs/2109.07906Khan, A. A., S. Badshah, P. Liang, et al. 2021. “Ethics of AI: A Systematic Literature Review of Principles and Challenges”. In arXiv. Preprint, September 12. https://arxiv.org/abs/2109.07906.Khan, A. A., et al. “Ethics of AI: A Systematic Literature Review of Principles and Challenges”. arXiv, 12 Sept. 2021, https://arxiv.org/abs/2109.07906.Khan, A. A. et al. Ethics of AI: A Systematic Literature Review of Principles and Challenges. arXiv Preprint at https://arxiv.org/abs/2109.07906 (2021).A. A. Khan et al., “Ethics of AI: A Systematic Literature Review of Principles and Challenges”, Sep. 12, 2021. [Online]. Available: https://arxiv.org/abs/2109.07906
- LaCroix, T. & Luccioni, A. S. (2022). Metaethical Perspectives on 'Benchmarking' AI Ethics. arXiv.LaCroix, T., & Luccioni, A. S. (2022). Metaethical Perspectives on 'Benchmarking' AI Ethics. In arXiv. https://arxiv.org/abs/2204.05151LaCroix, T., and A. S. Luccioni. 2022. “Metaethical Perspectives on 'Benchmarking' AI Ethics”. In arXiv. Preprint, April 11. https://arxiv.org/abs/2204.05151.LaCroix, T., and A. S. Luccioni. “Metaethical Perspectives on 'Benchmarking' AI Ethics”. arXiv, 11 Apr. 2022, https://arxiv.org/abs/2204.05151.LaCroix, T. & Luccioni, A. S. Metaethical Perspectives on 'Benchmarking' AI Ethics. arXiv Preprint at https://arxiv.org/abs/2204.05151 (2022).T. LaCroix and A. S. Luccioni, “Metaethical Perspectives on 'Benchmarking' AI Ethics”, Apr. 11, 2022. [Online]. Available: https://arxiv.org/abs/2204.05151
- Lizka (2023). Beware safety-washing. EA Forum.Lizka. (2023, January 13). Beware safety-washing. EA Forum. Internet Archive (https://web.archive.org/web/20260519093736/https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing). https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washingLizka. 2023. “Beware Safety-washing”. EA Forum, January 13. Https://web.archive.org/web/20260519093736/https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing. Internet Archive. https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing.Lizka. “Beware Safety-washing”. EA Forum, 13 Jan. 2023, Internet Archive, https://web.archive.org/web/20260519093736/https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing, https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing.Lizka. Beware safety-washing. EA Forum https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing (2023).Lizka, “Beware safety-washing”, EA Forum. Accessed: May 19, 2026. [Online]. Available: https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing
- Miller (2022). AI alignment with humans... but with which humans?. EA Forum.Miller. (2022). AI alignment with humans... but with which humans?. EA Forum. Internet Archive (https://web.archive.org/web/20260422074220/https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans). https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humansMiller. 2022. “AI Alignment with Humans... But with Which Humans?”. EA Forum. Https://web.archive.org/web/20260422074220/https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans. Internet Archive. https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans.Miller. “AI Alignment with Humans... But with Which Humans?”. EA Forum, 2022, Internet Archive, https://web.archive.org/web/20260422074220/https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans, https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans.Miller. AI alignment with humans... but with which humans?. EA Forum https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans (2022).Miller, “AI alignment with humans... but with which humans?”, EA Forum. Accessed: Apr. 22, 2026. [Online]. Available: https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans
- paulfchristiano (2018). Clarifying "AI Alignment". AI Alignment Forum.paulfchristiano. (2018, November 15). Clarifying "AI Alignment". AI Alignment Forum. https://alignmentforum.org/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignmentpaulfchristiano. 2018. “Clarifying "AI Alignment"”. AI Alignment Forum, November 15. https://alignmentforum.org/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment.paulfchristiano. “Clarifying "AI Alignment"”. AI Alignment Forum, 15 Nov. 2018, https://alignmentforum.org/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment.paulfchristiano. Clarifying "AI Alignment". AI Alignment Forum https://alignmentforum.org/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment (2018).paulfchristiano, “Clarifying "AI Alignment"”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment
- Ren, R. et al. (2024). Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. arXiv.Ren, R., Basart, S., Khoja, A., Gatti, A., Phan, L., Yin, X., Mazeika, M., Pan, A., Mukobi, G., Kim, R. H., Fitz, S., & Hendrycks, D. (2024). Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. In arXiv. https://arxiv.org/abs/2407.21792Ren, R., S. Basart, A. Khoja, et al. 2024. “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”. In arXiv. Preprint, July 31. https://arxiv.org/abs/2407.21792.Ren, R., et al. “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”. arXiv, 31 July 2024, https://arxiv.org/abs/2407.21792.Ren, R. et al. Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. arXiv Preprint at https://arxiv.org/abs/2407.21792 (2024).R. Ren et al., “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”, Jul. 31, 2024. [Online]. Available: https://arxiv.org/abs/2407.21792
- Rudner et al. (2021). Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology.Rudner et al. (2021). Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology. Center for Security and Emerging Technology. https://cset.georgetown.edu/publication/key-concepts-in-ai-safety-an-overviewRudner et al. 2021. “Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology”. Center for Security and Emerging Technology. https://cset.georgetown.edu/publication/key-concepts-in-ai-safety-an-overview.Rudner et al. “Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology”. Center for Security and Emerging Technology, 2021, https://cset.georgetown.edu/publication/key-concepts-in-ai-safety-an-overview.Rudner et al. Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology. Center for Security and Emerging Technology https://cset.georgetown.edu/publication/key-concepts-in-ai-safety-an-overview (2021).Rudner et al., “Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology”, Center for Security and Emerging Technology. [Online]. Available: https://cset.georgetown.edu/publication/key-concepts-in-ai-safety-an-overview
- Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan & Fabien Roger (2023). AI Control: Improving Safety Despite Intentional Subversion. arXiv.Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, & Fabien Roger. (2023). AI Control: Improving Safety Despite Intentional Subversion. In arXiv. https://arxiv.org/abs/2312.06942Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. 2023. “AI Control: Improving Safety Despite Intentional Subversion”. In arXiv. Preprint, December 12. https://arxiv.org/abs/2312.06942.Ryan Greenblatt, et al. “AI Control: Improving Safety Despite Intentional Subversion”. arXiv, 12 Dec. 2023, https://arxiv.org/abs/2312.06942.Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan & Fabien Roger. AI Control: Improving Safety Despite Intentional Subversion. arXiv Preprint at https://arxiv.org/abs/2312.06942 (2023).Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger, “AI Control: Improving Safety Despite Intentional Subversion”, Dec. 12, 2023. [Online]. Available: https://arxiv.org/abs/2312.06942
- ryan_greenblatt & Buck (2024). The case for ensuring that powerful AIs are controlled. AI Alignment Forum.ryan_greenblatt, & Buck. (2024, January 24). The case for ensuring that powerful AIs are controlled. AI Alignment Forum. https://alignmentforum.org/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlledryan_greenblatt, and Buck. 2024. “The Case for Ensuring That Powerful AIs Are Controlled”. AI Alignment Forum, January 24. https://alignmentforum.org/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled.ryan_greenblatt, and Buck. “The Case for Ensuring That Powerful AIs Are Controlled”. AI Alignment Forum, 24 Jan. 2024, https://alignmentforum.org/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled.ryan_greenblatt & Buck. The case for ensuring that powerful AIs are controlled. AI Alignment Forum https://alignmentforum.org/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled (2024).ryan_greenblatt and Buck, “The case for ensuring that powerful AIs are controlled”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled
Was this section useful?
Thank you for your feedback
Your input helps improve the Atlas.