AI governance is not the same as traditional technology governance. Traditional technology governance relies on several key assumptions that break down when applied to AI. We typically assume we can predict how a technology will be used and its likely impacts, that we can effectively control its development pathway, and that we can regulate specific applications or end-uses. For example, pharmaceutical governance uses clinical trials and approval processes based on intended medical applications, while nuclear technology is controlled through international treaties, safeguards, and monitoring of specific facilities and materials. These approaches work when technologies follow relatively predictable development paths and have clear applications. To understand what makes AI governance uniquely challenging, we can examine AI through three different lenses that each require different governance approaches (Dafoe, 2022; Buchanan, 2020).
AI as general-purpose technology. AI transforms many sectors simultaneously, making sector-specific regulation insufficient. Like electricity or computers before it, AI can reshape healthcare, finance, transportation, and education all at once. Traditional technology governance typically focuses on specific applications - we regulate medical devices differently from automobiles. But when a single AI system can diagnose diseases, trade stocks, and drive cars, our regulatory silos break down. The impacts span across society in ways that make targeted regulation insufficient (Buchanan, 2020).
AI as information technology. AI processes and generates information in unprecedented ways. Unlike traditional information systems that store and retrieve data, AI can create entirely new content - from photorealistic images to convincing text to synthetic voices. This creates unprecedented challenges around security, privacy, and information integrity. Traditional governance frameworks weren't designed to handle technologies that can rapidly generate and manipulate information at massive scale (Brundage et al., 2018). The speed and scope of potential information impacts outstrip traditional control mechanisms.
AI as intelligence technology. AI introduces unique control challenges as systems become more capable. As AI systems approach and potentially exceed human cognitive abilities in various domains, they may develop sophisticated ways to evade controls or pursue unintended objectives. We're already seeing glimpses of this with language models that can engage in deception or manipulation when pursuing goals (Ganguli et al., 2022). There are several dangerous capabilities (refer back to chapters 1 and 2) which become even more acute when considering that AI systems might develop these capabilities without being explicitly programmed for them (Woodside, 2024). The intelligence aspect of AI creates a dynamic where the technology being governed might actively resist or circumvent governance measures, a challenge without precedent in technology regulation.
The combination of AI as a general-purpose, information, intelligence technology creates unique governance challenges. The mixed nature of AI as a general-purpose, information processing, and potentially intelligent technology gives rise to three fundamental problems that make traditional governance approaches inadequate.
Unexpected Capabilities #
AI systems develop surprising abilities that weren't part of their intended design. Through several of our chapters now, we have shown that foundation models can show "emergent" capabilities that appear suddenly as models scale up with more data, parameters and compute. GPT-3 unexpectedly demonstrated the ability to perform basic arithmetic, while later models showed emergent reasoning capabilities that surprised even their creators (Ganguli et al., 2022; Wei et al., 2022). Evaluations have found that frontier models can autonomously conduct basic scientific research, hack into computer systems, and manipulate humans through persuasion, none of which were explicitly trained for (Phuong et al., 2024; Boiko et al., 2023; Turpin et al., 2023; Fang et al., 2024).
AI evaluations are still in their early stages in 2025. Testing frameworks lack established best practices, and the field has yet to mature into a reliable science (Trusilo, 2024). While evaluations can reveal some capabilities, they cannot guarantee absence of unknown threats, forecast new emergent abilities, or assess risks from autonomous systems (Barnett & Thiergart, 2024). Predictability itself is a nascent research area, with major gaps in our ability to anticipate how present models behave, let alone future ones (Zhou et al., 2024). Even the most comprehensive test-and-evaluation frameworks struggle with complex, unpredictable AI behavior (Wojton et al., 2020).
Deployment Safety #
Once deployed, AI systems can be repurposed for harmful applications beyond their intended use. The same language model trained for helpful dialogue can generate misinformation, assist with cyberattacks, or help design biological weapons. Users regularly discover new capabilities through clever prompting that bypasses safety measures called "jailbreaks" that unlock dangerous functionalities (Solaiman et al., 2024; Marchal et al., 2024; Hendrycks et al., 2023).
AI agents amplify deployment risks. We're now seeing autonomous AI agents that can chain together model capabilities in novel ways, using tools and taking actions in the real world. These agents can pursue complex goals over extended periods, making their behavior even harder to predict and control post-deployment (Fang et al., 2024).
Proliferation #
AI capabilities spread rapidly through multiple channels, making containment nearly impossible. Models can be stolen through cyberattacks, leaked by insiders, or reproduced by competitors within months. The rapid open-source replication of ChatGPT-like capabilities led to models with safety features removed and new dangerous capabilities discovered through community experimentation (Seger et al., 2023). With API-based models, techniques like model distillation can even extract capabilities without direct access to model weights (Nevo et al., 2024).
Physical containment doesn't work for digital goods. Unlike nuclear materials or dangerous pathogens, AI models are just patterns of numbers that can be copied instantly and transmitted globally. Once capabilities exist, controlling their spread becomes a losing battle against the fundamental nature of digital information.
Governance Targets #
The unique challenges associated with AI governance mean we need to carefully choose where and how to intervene in AI development. This requires identifying both what to govern (targets) and how to govern it (mechanisms) (Anderljung et al., 2023; Reuel & Bucknall, 2024). Governance must intervene at points that address core challenges before they manifest. We can't wait for dangerous capabilities to emerge or proliferate before acting. Instead, we need to identify intervention points in the AI development pipeline that will help us shape AI development proactively.
Effective governance targets share three essential properties:
- Measurability: We must be able to track and verify what's happening. The amount of computing power used for training can be measured in precise units (floating-point operations), making it possible to set clear thresholds and monitor compliance (Sastry et al., 2024).
- Controllability: There must be concrete mechanisms to influence the target. It's not enough to identify what matters, we need practical ways to shape it. The semiconductor supply chain, for instance, has clear chokepoints where export controls can effectively limit access to advanced chips (Heim et al., 2024).
- Meaningfulness: Targets should address fundamental aspects of AI development that actually shape capabilities and risks. Regulating superficial aspects like user interfaces might be easy but won't prevent the emergence of dangerous capabilities. Core inputs like compute and data, however, directly determine what kinds of AI systems can be built (Anderljung et al., 2023)
In the AI development pipeline, several intervention points meet these criteria. Early in development, we can target the compute infrastructure required for training and the data that shapes model capabilities. During and after development, we can implement safety frameworks, monitoring systems, and deployment controls (Anderljung et. al, 2023; Heim et al., 2024; Hausenloy et al., 2024). Each target offers different opportunities and faces different challenges, which we'll explore in the following sections.
References
- Anderljung, M. et al. (2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv.Anderljung, M., Barnhart, J., Korinek, A., Leung, J., O'Keefe, C., Whittlestone, J., Avin, S., Brundage, M., Bullock, J., Cass-Beggs, D., Chang, B., Collins, T., Fist, T., Hadfield, G., Hayes, A., Ho, L., Hooker, S., Horvitz, E., Kolt, N., … Wolf, K. (2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. In arXiv. https://arxiv.org/abs/2307.03718Anderljung, M., J. Barnhart, A. Korinek, et al. 2023. “Frontier AI Regulation: Managing Emerging Risks to Public Safety”. In arXiv. Preprint, July 6. https://arxiv.org/abs/2307.03718.Anderljung, M., et al. “Frontier AI Regulation: Managing Emerging Risks to Public Safety”. arXiv, 6 July 2023, https://arxiv.org/abs/2307.03718.Anderljung, M. et al. Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv Preprint at https://arxiv.org/abs/2307.03718 (2023).M. Anderljung et al., “Frontier AI Regulation: Managing Emerging Risks to Public Safety”, Jul. 06, 2023. [Online]. Available: https://arxiv.org/abs/2307.03718
- Barnett, P. & Thiergart, L. (2024). What AI evaluations for preventing catastrophic risks can and cannot do. arXiv.Barnett, P., & Thiergart, L. (2024). What AI evaluations for preventing catastrophic risks can and cannot do. In arXiv. https://arxiv.org/abs/2412.08653Barnett, P., and L. Thiergart. 2024. “What AI Evaluations for Preventing Catastrophic Risks Can and Cannot Do”. In arXiv. Preprint, November 26. https://arxiv.org/abs/2412.08653.Barnett, P., and L. Thiergart. “What AI Evaluations for Preventing Catastrophic Risks Can and Cannot Do”. arXiv, 26 Nov. 2024, https://arxiv.org/abs/2412.08653.Barnett, P. & Thiergart, L. What AI evaluations for preventing catastrophic risks can and cannot do. arXiv Preprint at https://arxiv.org/abs/2412.08653 (2024).P. Barnett and L. Thiergart, “What AI evaluations for preventing catastrophic risks can and cannot do”, Nov. 26, 2024. [Online]. Available: https://arxiv.org/abs/2412.08653
- Boiko, D. A., MacKnight, R. & Gomes, G. (2023). Emergent autonomous scientific research capabilities of large language models. arXiv.Boiko, D. A., MacKnight, R., & Gomes, G. (2023). Emergent autonomous scientific research capabilities of large language models. In arXiv. https://arxiv.org/abs/2304.05332Boiko, D. A., R. MacKnight, and G. Gomes. 2023. “Emergent Autonomous Scientific Research Capabilities of Large Language Models”. In arXiv. Preprint, April 11. https://arxiv.org/abs/2304.05332.Boiko, D. A., et al. “Emergent Autonomous Scientific Research Capabilities of Large Language Models”. arXiv, 11 Apr. 2023, https://arxiv.org/abs/2304.05332.Boiko, D. A., MacKnight, R. & Gomes, G. Emergent autonomous scientific research capabilities of large language models. arXiv Preprint at https://arxiv.org/abs/2304.05332 (2023).D. A. Boiko, R. MacKnight, and G. Gomes, “Emergent autonomous scientific research capabilities of large language models”, Apr. 11, 2023. [Online]. Available: https://arxiv.org/abs/2304.05332
- Brundage, M. et al. (2018). The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation. arXiv.Brundage, M., Avin, S., Clark, J., Toner, H., Eckersley, P., Garfinkel, B., Dafoe, A., Scharre, P., Zeitzoff, T., Filar, B., Anderson, H., Roff, H., Allen, G. C., Steinhardt, J., Flynn, C., hÉigeartaigh, S. Ó., Beard, S., Belfield, H., Farquhar, S., … Amodei, D. (2018). The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation. In arXiv. https://arxiv.org/abs/1802.07228Brundage, M., S. Avin, J. Clark, et al. 2018. “The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation”. In arXiv. Preprint, February 20. https://arxiv.org/abs/1802.07228.Brundage, M., et al. “The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation”. arXiv, 20 Feb. 2018, https://arxiv.org/abs/1802.07228.Brundage, M. et al. The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation. arXiv Preprint at https://arxiv.org/abs/1802.07228 (2018).M. Brundage et al., “The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation”, Feb. 20, 2018. [Online]. Available: https://arxiv.org/abs/1802.07228
- Buchanan (2020). The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology.Buchanan. (2020). The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology. Center for Security and Emerging Technology. https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategyBuchanan. 2020. “The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology”. Center for Security and Emerging Technology. https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategy.Buchanan. “The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology”. Center for Security and Emerging Technology, 2020, https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategy.Buchanan. The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology. Center for Security and Emerging Technology https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategy (2020).Buchanan, “The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology”, Center for Security and Emerging Technology. [Online]. Available: https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategy
- Dafoe, A. (2024). AI Governance: Overview and Theoretical Lenses. The Oxford Handbook of AI Governance.Dafoe, A. (2024). AI Governance: Overview and Theoretical Lenses. In The Oxford Handbook of AI Governance (pp. 21–44). Oxford University Press. https://doi.org/10.1093/oxfordhb/9780197579329.013.2Dafoe, A. 2024. “AI Governance: Overview and Theoretical Lenses”. In The Oxford Handbook of AI Governance. Oxford University Press. https://doi.org/10.1093/oxfordhb/9780197579329.013.2.Dafoe, A. “AI Governance: Overview and Theoretical Lenses”. The Oxford Handbook of AI Governance, Oxford University Press, 2024, pp. 21–44, https://doi.org/10.1093/oxfordhb/9780197579329.013.2.Dafoe, A. AI Governance: Overview and Theoretical Lenses. in The Oxford Handbook of AI Governance 21–44 (Oxford University Press, 2024). doi:10.1093/oxfordhb/9780197579329.013.2.A. Dafoe, “AI Governance: Overview and Theoretical Lenses”, in The Oxford Handbook of AI Governance, Oxford University Press, 2024, pp. 21–44. doi: 10.1093/oxfordhb/9780197579329.013.2.
- Fang, R., Bindu, R., Gupta, A., Zhan, Q. & Kang, D. (2024). LLM Agents can Autonomously Hack Websites. arXiv.Fang, R., Bindu, R., Gupta, A., Zhan, Q., & Kang, D. (2024). LLM Agents can Autonomously Hack Websites. In arXiv. https://arxiv.org/abs/2402.06664Fang, R., R. Bindu, A. Gupta, Q. Zhan, and D. Kang. 2024. “LLM Agents Can Autonomously Hack Websites”. In arXiv. Preprint, February 6. https://arxiv.org/abs/2402.06664.Fang, R., et al. “LLM Agents Can Autonomously Hack Websites”. arXiv, 6 Feb. 2024, https://arxiv.org/abs/2402.06664.Fang, R., Bindu, R., Gupta, A., Zhan, Q. & Kang, D. LLM Agents can Autonomously Hack Websites. arXiv Preprint at https://arxiv.org/abs/2402.06664 (2024).R. Fang, R. Bindu, A. Gupta, Q. Zhan, and D. Kang, “LLM Agents can Autonomously Hack Websites”, Feb. 06, 2024. [Online]. Available: https://arxiv.org/abs/2402.06664
- Ganguli, D. et al. (2022). Predictability and Surprise in Large Generative Models. arXiv.Ganguli, D., Hernandez, D., Lovitt, L., DasSarma, N., Henighan, T., Jones, A., Joseph, N., Kernion, J., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Elhage, N., Showk, S. E., Fort, S., Hatfield-Dodds, Z., Johnston, S., … Clark, J. (2022). Predictability and Surprise in Large Generative Models. In arXiv. https://doi.org/10.1145/3531146.3533229Ganguli, D., D. Hernandez, L. Lovitt, et al. 2022. “Predictability and Surprise in Large Generative Models”. In arXiv. Preprint, February 15. https://doi.org/10.1145/3531146.3533229.Ganguli, D., et al. “Predictability and Surprise in Large Generative Models”. arXiv, 15 Feb. 2022, https://doi.org/10.1145/3531146.3533229.Ganguli, D. et al. Predictability and Surprise in Large Generative Models. arXiv Preprint at https://doi.org/10.1145/3531146.3533229 (2022).D. Ganguli et al., “Predictability and Surprise in Large Generative Models”, Feb. 15, 2022. doi: 10.1145/3531146.3533229.
- Hausenloy, J., McClements, D. & Thakur, M. (2024). Towards Data Governance of Frontier AI Models. arXiv.Hausenloy, J., McClements, D., & Thakur, M. (2024). Towards Data Governance of Frontier AI Models. In arXiv. https://arxiv.org/abs/2412.03824Hausenloy, J., D. McClements, and M. Thakur. 2024. “Towards Data Governance of Frontier AI Models”. In arXiv. Preprint, December 5. https://arxiv.org/abs/2412.03824.Hausenloy, J., et al. “Towards Data Governance of Frontier AI Models”. arXiv, 5 Dec. 2024, https://arxiv.org/abs/2412.03824.Hausenloy, J., McClements, D. & Thakur, M. Towards Data Governance of Frontier AI Models. arXiv Preprint at https://arxiv.org/abs/2412.03824 (2024).J. Hausenloy, D. McClements, and M. Thakur, “Towards Data Governance of Frontier AI Models”, Dec. 05, 2024. [Online]. Available: https://arxiv.org/abs/2412.03824
- Hendrycks, D., Mazeika, M. & Woodside, T. (2023). An Overview of Catastrophic AI Risks. arXiv.Hendrycks, D., Mazeika, M., & Woodside, T. (2023). An Overview of Catastrophic AI Risks. In arXiv. https://arxiv.org/abs/2306.12001Hendrycks, D., M. Mazeika, and T. Woodside. 2023. “An Overview of Catastrophic AI Risks”. In arXiv. Preprint, June 21. https://arxiv.org/abs/2306.12001.Hendrycks, D., et al. “An Overview of Catastrophic AI Risks”. arXiv, 21 June 2023, https://arxiv.org/abs/2306.12001.Hendrycks, D., Mazeika, M. & Woodside, T. An Overview of Catastrophic AI Risks. arXiv Preprint at https://arxiv.org/abs/2306.12001 (2023).D. Hendrycks, M. Mazeika, and T. Woodside, “An Overview of Catastrophic AI Risks”, Jun. 21, 2023. [Online]. Available: https://arxiv.org/abs/2306.12001
- Lennart Heim et al. (2024). Governing Through the Cloud.Lennart Heim, Tim Fist, Janet Egan, Sihao Huang, Stephen Zekany, Robert Trager, Michael A Osborne, & Noa Zilberman. (2024, March 13). Governing Through the Cloud. https://governance.ai/research-paper/governing-through-the-cloudLennart Heim, Tim Fist, Janet Egan, et al. 2024. “Governing Through the Cloud”. March 13. https://governance.ai/research-paper/governing-through-the-cloud.Lennart Heim, et al. Governing Through the Cloud. 13 Mar. 2024, https://governance.ai/research-paper/governing-through-the-cloud.Lennart Heim et al. Governing Through the Cloud. https://governance.ai/research-paper/governing-through-the-cloud (2024).Lennart Heim et al., “Governing Through the Cloud”. [Online]. Available: https://governance.ai/research-paper/governing-through-the-cloud
- Lennart Heim, * Markus Anderljung, Emma Bluemke & Robert Trager (2024). Computing Power and the Governance of AI.Lennart Heim, * Markus Anderljung, Emma Bluemke, & Robert Trager. (2024, February 14). Computing Power and the Governance of AI. https://governance.ai/analysis/computing-power-and-the-governance-of-aiLennart Heim, * Markus Anderljung, Emma Bluemke, and Robert Trager. 2024. “Computing Power and the Governance of AI”. February 14. https://governance.ai/analysis/computing-power-and-the-governance-of-ai.Lennart Heim, et al. Computing Power and the Governance of AI. 14 Feb. 2024, https://governance.ai/analysis/computing-power-and-the-governance-of-ai.Lennart Heim, * Markus Anderljung, Emma Bluemke & Robert Trager. Computing Power and the Governance of AI. https://governance.ai/analysis/computing-power-and-the-governance-of-ai (2024).Lennart Heim, * Markus Anderljung, Emma Bluemke, and Robert Trager, “Computing Power and the Governance of AI”. [Online]. Available: https://governance.ai/analysis/computing-power-and-the-governance-of-ai
- Marchal, N., Xu, R., Elasmar, R., Gabriel, I., Goldberg, B. & Isaac, W. (2024). Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data. arXiv.Marchal, N., Xu, R., Elasmar, R., Gabriel, I., Goldberg, B., & Isaac, W. (2024). Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data. In arXiv. https://arxiv.org/abs/2406.13843Marchal, N., R. Xu, R. Elasmar, I. Gabriel, B. Goldberg, and W. Isaac. 2024. “Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data”. In arXiv. Preprint, June 19. https://arxiv.org/abs/2406.13843.Marchal, N., et al. “Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data”. arXiv, 19 June 2024, https://arxiv.org/abs/2406.13843.Marchal, N. et al. Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data. arXiv Preprint at https://arxiv.org/abs/2406.13843 (2024).N. Marchal, R. Xu, R. Elasmar, I. Gabriel, B. Goldberg, and W. Isaac, “Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data”, Jun. 19, 2024. [Online]. Available: https://arxiv.org/abs/2406.13843
- Mary Phuong et al. (2024). Evaluating Frontier Models for Dangerous Capabilities. arXiv.Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, … Toby Shevlane. (2024). Evaluating Frontier Models for Dangerous Capabilities. In arXiv. https://arxiv.org/abs/2403.13793Mary Phuong, Matthew Aitchison, Elliot Catt, et al. 2024. “Evaluating Frontier Models for Dangerous Capabilities”. In arXiv. Preprint, March 20. https://arxiv.org/abs/2403.13793.Mary Phuong, et al. “Evaluating Frontier Models for Dangerous Capabilities”. arXiv, 20 Mar. 2024, https://arxiv.org/abs/2403.13793.Mary Phuong et al. Evaluating Frontier Models for Dangerous Capabilities. arXiv Preprint at https://arxiv.org/abs/2403.13793 (2024).Mary Phuong et al., “Evaluating Frontier Models for Dangerous Capabilities”, Mar. 20, 2024. [Online]. Available: https://arxiv.org/abs/2403.13793
- Reuel, A. et al. (2024). Open Problems in Technical AI Governance. arXiv.Reuel, A., Bucknall, B., Casper, S., Fist, T., Soder, L., Aarne, O., Hammond, L., Ibrahim, L., Chan, A., Wills, P., Anderljung, M., Garfinkel, B., Heim, L., Trask, A., Mukobi, G., Schaeffer, R., Baker, M., Hooker, S., Solaiman, I., … Trager, R. (2024). Open Problems in Technical AI Governance. In arXiv. https://arxiv.org/abs/2407.14981Reuel, A., B. Bucknall, S. Casper, et al. 2024. “Open Problems in Technical AI Governance”. In arXiv. Preprint, July 20. https://arxiv.org/abs/2407.14981.Reuel, A., et al. “Open Problems in Technical AI Governance”. arXiv, 20 July 2024, https://arxiv.org/abs/2407.14981.Reuel, A. et al. Open Problems in Technical AI Governance. arXiv Preprint at https://arxiv.org/abs/2407.14981 (2024).A. Reuel et al., “Open Problems in Technical AI Governance”, Jul. 20, 2024. [Online]. Available: https://arxiv.org/abs/2407.14981
- Sastry, G. et al. (2024). Computing Power and the Governance of Artificial Intelligence. arXiv.Sastry, G., Heim, L., Belfield, H., Anderljung, M., Brundage, M., Hazell, J., O'Keefe, C., Hadfield, G. K., Ngo, R., Pilz, K., Gor, G., Bluemke, E., Shoker, S., Egan, J., Trager, R. F., Avin, S., Weller, A., Bengio, Y., & Coyle, D. (2024). Computing Power and the Governance of Artificial Intelligence. In arXiv. https://arxiv.org/abs/2402.08797Sastry, G., L. Heim, H. Belfield, et al. 2024. “Computing Power and the Governance of Artificial Intelligence”. In arXiv. Preprint, February 13. https://arxiv.org/abs/2402.08797.Sastry, G., et al. “Computing Power and the Governance of Artificial Intelligence”. arXiv, 13 Feb. 2024, https://arxiv.org/abs/2402.08797.Sastry, G. et al. Computing Power and the Governance of Artificial Intelligence. arXiv Preprint at https://arxiv.org/abs/2402.08797 (2024).G. Sastry et al., “Computing Power and the Governance of Artificial Intelligence”, Feb. 13, 2024. [Online]. Available: https://arxiv.org/abs/2402.08797
- (2024). Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models.Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, Jeff Alstott. (2024). Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. https://rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdfSella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, Jeff Alstott. 2024. “Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models”. https://rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdf.Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, Jeff Alstott. Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. 2024, https://rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdf.Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, Jeff Alstott. Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. https://rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdf (2024).Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, Jeff Alstott, “Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models”. [Online]. Available: https://rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdf
- Seger, E. et al. (2023). Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives. arXiv.Seger, E., Dreksler, N., Moulange, R., Dardaman, E., Schuett, J., Wei, K., Winter, C., Arnold, M., hÉigeartaigh, S. Ó., Korinek, A., Anderljung, M., Bucknall, B., Chan, A., Stafford, E., Koessler, L., Ovadya, A., Garfinkel, B., Bluemke, E., Aird, M., … Gupta, A. (2023). Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives. In arXiv. https://arxiv.org/abs/2311.09227Seger, E., N. Dreksler, R. Moulange, et al. 2023. “Open-Sourcing Highly Capable Foundation Models: An Evaluation of Risks, Benefits, and Alternative Methods for Pursuing Open-source Objectives”. In arXiv. Preprint, September 29. https://arxiv.org/abs/2311.09227.Seger, E., et al. “Open-Sourcing Highly Capable Foundation Models: An Evaluation of Risks, Benefits, and Alternative Methods for Pursuing Open-source Objectives”. arXiv, 29 Sept. 2023, https://arxiv.org/abs/2311.09227.Seger, E. et al. Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives. arXiv Preprint at https://arxiv.org/abs/2311.09227 (2023).E. Seger et al., “Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives”, Sep. 29, 2023. [Online]. Available: https://arxiv.org/abs/2311.09227
- Solaiman, I. et al. (2023). Evaluating the Social Impact of Generative AI Systems in Systems and Society. arXiv.Solaiman, I., Talat, Z., Agnew, W., Ahmad, L., Baker, D., Blodgett, S. L., Chen, C., Daumé, H., Dodge, J., Duan, I., Evans, E., Friedrich, F., Ghosh, A., Gohar, U., Hooker, S., Jernite, Y., Kalluri, R., Lusoli, A., Leidinger, A., … Subramonian, A. (2023). Evaluating the Social Impact of Generative AI Systems in Systems and Society. In arXiv. https://doi.org/10.1093/oxfordhb/9780198940272.013.0025Solaiman, I., Z. Talat, W. Agnew, et al. 2023. “Evaluating the Social Impact of Generative AI Systems in Systems and Society”. In arXiv. Preprint, June 9. https://doi.org/10.1093/oxfordhb/9780198940272.013.0025.Solaiman, I., et al. “Evaluating the Social Impact of Generative AI Systems in Systems and Society”. arXiv, 9 June 2023, https://doi.org/10.1093/oxfordhb/9780198940272.013.0025.Solaiman, I. et al. Evaluating the Social Impact of Generative AI Systems in Systems and Society. arXiv Preprint at https://doi.org/10.1093/oxfordhb/9780198940272.013.0025 (2023).I. Solaiman et al., “Evaluating the Social Impact of Generative AI Systems in Systems and Society”, Jun. 09, 2023. doi: 10.1093/oxfordhb/9780198940272.013.0025.
- Trusilo, D. (2023). Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability. Journal of Military Ethics.Trusilo, D. (2023). Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability. Journal of Military Ethics. https://doi.org/10.1080/15027570.2023.2213985Trusilo, D. 2023. “Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability”. Journal of Military Ethics, ahead of print, January 2. https://doi.org/10.1080/15027570.2023.2213985.Trusilo, D. “Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability”. Journal of Military Ethics, Jan. 2023, https://doi.org/10.1080/15027570.2023.2213985.Trusilo, D. Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability. Journal of Military Ethics https://doi.org/10.1080/15027570.2023.2213985 (2023) doi:10.1080/15027570.2023.2213985.D. Trusilo, “Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability”, Journal of Military Ethics, Jan. 2023, doi: 10.1080/15027570.2023.2213985.
- Turpin, M., Michael, J., Perez, E. & Bowman, S. R. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. arXiv.Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. In arXiv. https://arxiv.org/abs/2305.04388Turpin, M., J. Michael, E. Perez, and S. R. Bowman. 2023. “Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting”. In arXiv. Preprint, May 7. https://arxiv.org/abs/2305.04388.Turpin, M., et al. “Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting”. arXiv, 7 May 2023, https://arxiv.org/abs/2305.04388.Turpin, M., Michael, J., Perez, E. & Bowman, S. R. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. arXiv Preprint at https://arxiv.org/abs/2305.04388 (2023).M. Turpin, J. Michael, E. Perez, and S. R. Bowman, “Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting”, May 07, 2023. [Online]. Available: https://arxiv.org/abs/2305.04388
- Wei, J. et al. (2022). Emergent Abilities of Large Language Models. arXiv.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., & Fedus, W. (2022). Emergent Abilities of Large Language Models. In arXiv. https://arxiv.org/abs/2206.07682Wei, J., Y. Tay, R. Bommasani, et al. 2022. “Emergent Abilities of Large Language Models”. In arXiv. Preprint, June 15. https://arxiv.org/abs/2206.07682.Wei, J., et al. “Emergent Abilities of Large Language Models”. arXiv, 15 June 2022, https://arxiv.org/abs/2206.07682.Wei, J. et al. Emergent Abilities of Large Language Models. arXiv Preprint at https://arxiv.org/abs/2206.07682 (2022).J. Wei et al., “Emergent Abilities of Large Language Models”, Jun. 15, 2022. [Online]. Available: https://arxiv.org/abs/2206.07682
- Wojton, H. M., Porter, D. J. & Dennis, J. W. (2020). Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review.Wojton, H. M., Porter, D. J., & Dennis, J. W. (2020). Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review (IDA Document NS-D-14331). Institute for Defense Analyses. https://testscience.org/wp-content/uploads/formidable/20/Autonomy-Lit-Review.pdfWojton, H. M., D. J. Porter, and J. W. Dennis. 2020. Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review. IDA Document NS-D-14331. Institute for Defense Analyses. https://testscience.org/wp-content/uploads/formidable/20/Autonomy-Lit-Review.pdf.Wojton, H. M., et al. Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review. IDA Document NS-D-14331, Institute for Defense Analyses, Sept. 2020, https://testscience.org/wp-content/uploads/formidable/20/Autonomy-Lit-Review.pdf.Wojton, H. M., Porter, D. J. & Dennis, J. W. Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review. https://testscience.org/wp-content/uploads/formidable/20/Autonomy-Lit-Review.pdf (2020).H. M. Wojton, D. J. Porter, and J. W. Dennis, “Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review”, Institute for Defense Analyses, IDA Document NS-D-14331, Sep. 2020. [Online]. Available: https://testscience.org/wp-content/uploads/formidable/20/Autonomy-Lit-Review.pdf
- Zhou, L. et al. (2023). Predictable Artificial Intelligence. arXiv.Zhou, L., Moreno-Casares, P. A., Martínez-Plumed, F., Burden, J., Burnell, R., Cheke, L., Ferri, C., Marcoci, A., Mehrbakhsh, B., Moros-Daval, Y., hÉigeartaigh, S. Ó., Rutar, D., Schellaert, W., Voudouris, K., & Hernández-Orallo, J. (2023). Predictable Artificial Intelligence. In arXiv. https://arxiv.org/abs/2310.06167Zhou, L., P. A. Moreno-Casares, F. Martínez-Plumed, et al. 2023. “Predictable Artificial Intelligence”. In arXiv. Preprint, October 9. https://arxiv.org/abs/2310.06167.Zhou, L., et al. “Predictable Artificial Intelligence”. arXiv, 9 Oct. 2023, https://arxiv.org/abs/2310.06167.Zhou, L. et al. Predictable Artificial Intelligence. arXiv Preprint at https://arxiv.org/abs/2310.06167 (2023).L. Zhou et al., “Predictable Artificial Intelligence”, Oct. 09, 2023. [Online]. Available: https://arxiv.org/abs/2310.06167
Was this section useful?
Thank you for your feedback
Your input helps improve the Atlas.