Abid, A., Farooqi, M. & Zou, J.(2021). Persistent Anti-Muslim Bias in Large Language Models. arXiv.Abid, A., Farooqi, M., & Zou, J. (2021). Persistent Anti-Muslim Bias in Large Language Models. In arXiv. https://arxiv.org/abs/2101.05783Abid, A., M. Farooqi, and J. Zou. 2021. “Persistent Anti-Muslim Bias in Large Language Models”. In arXiv. Preprint, January 14. https://arxiv.org/abs/2101.05783.Abid, A., et al. “Persistent Anti-Muslim Bias in Large Language Models”. arXiv, 14 Jan. 2021, https://arxiv.org/abs/2101.05783.Abid, A., Farooqi, M. & Zou, J. Persistent Anti-Muslim Bias in Large Language Models. arXiv Preprint at https://arxiv.org/abs/2101.05783 (2021).A. Abid, M. Farooqi, and J. Zou, “Persistent Anti-Muslim Bias in Large Language Models”, Jan. 14, 2021. [Online]. Available: https://arxiv.org/abs/2101.05783
Abraham(2024). ‘Lavender’: The AI machine directing Israel’s bombing spree in Gaza. +972 Magazine.Abraham. (2024, April 3). ‘Lavender’: The AI machine directing Israel’s bombing spree in Gaza. +972 Magazine. https://972mag.com/lavender-ai-israeli-army-gazaAbraham. 2024. “‘Lavender’: The AI Machine Directing Israel’s Bombing Spree in Gaza”. +972 Magazine, April 3. https://972mag.com/lavender-ai-israeli-army-gaza.Abraham. “‘Lavender’: The AI Machine Directing Israel’s Bombing Spree in Gaza”. +972 Magazine, 3 Apr. 2024, https://972mag.com/lavender-ai-israeli-army-gaza.Abraham. ‘Lavender’: The AI machine directing Israel’s bombing spree in Gaza. +972 Magazine https://972mag.com/lavender-ai-israeli-army-gaza (2024).Abraham, “‘Lavender’: The AI machine directing Israel’s bombing spree in Gaza”, +972 Magazine. [Online]. Available: https://972mag.com/lavender-ai-israeli-army-gaza
Adam Scherlis(2023). Inner Misalignment in "Simulator" LLMs. AI Alignment Forum.Adam Scherlis. (2023, January 31). Inner Misalignment in "Simulator" LLMs. AI Alignment Forum. https://alignmentforum.org/posts/FLMyTjuTiGytE6sP2/inner-misalignment-in-simulator-llmsAdam Scherlis. 2023. “Inner Misalignment in "Simulator" LLMs”. AI Alignment Forum, January 31. https://alignmentforum.org/posts/FLMyTjuTiGytE6sP2/inner-misalignment-in-simulator-llms.Adam Scherlis. “Inner Misalignment in "Simulator" LLMs”. AI Alignment Forum, 31 Jan. 2023, https://alignmentforum.org/posts/FLMyTjuTiGytE6sP2/inner-misalignment-in-simulator-llms.Adam Scherlis. Inner Misalignment in "Simulator" LLMs. AI Alignment Forum https://alignmentforum.org/posts/FLMyTjuTiGytE6sP2/inner-misalignment-in-simulator-llms (2023).Adam Scherlis, “Inner Misalignment in "Simulator" LLMs”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/FLMyTjuTiGytE6sP2/inner-misalignment-in-simulator-llms
adamShimi(2020). Goal-directedness is behavioral, not structural. AI Alignment Forum.adamShimi. (2020, June 8). Goal-directedness is behavioral, not structural. AI Alignment Forum. https://alignmentforum.org/posts/9pxcekdNjE7oNwvcC/goal-directedness-is-behavioral-not-structuraladamShimi. 2020. “Goal-directedness Is Behavioral, Not Structural”. AI Alignment Forum, June 8. https://alignmentforum.org/posts/9pxcekdNjE7oNwvcC/goal-directedness-is-behavioral-not-structural.adamShimi. “Goal-directedness Is Behavioral, Not Structural”. AI Alignment Forum, 8 June 2020, https://alignmentforum.org/posts/9pxcekdNjE7oNwvcC/goal-directedness-is-behavioral-not-structural.adamShimi. Goal-directedness is behavioral, not structural. AI Alignment Forum https://alignmentforum.org/posts/9pxcekdNjE7oNwvcC/goal-directedness-is-behavioral-not-structural (2020).adamShimi, “Goal-directedness is behavioral, not structural”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/9pxcekdNjE7oNwvcC/goal-directedness-is-behavioral-not-structural
Adan, S. N. et al.(2024). Voice and Access in AI: Global AI Majority Participation in Artificial Intelligence Development and Governance.Adan, S. N., Trager, R., Blomquist, K., Dennis, C., Edom, G., Velasco, L., Abungu, C., Garfinkel, B., Jacobs, J., Okolo, C. T., Wu, B., & Vipra, J. (2024). Voice and Access in AI: Global AI Majority Participation in Artificial Intelligence Development and Governance. Oxford Martin AI Governance Initiative, University of Oxford. https://oms-www.files.svdcdn.com/production/downloads/reports/Voice%20and%20Access%20in%20AI_%20Global%20AI%20Majority%20Participation%20in%20Artificial%20Intelligence%20Development%20and%20Governance-%20final.pdf?dm=1729247034Adan, S. N., R. Trager, K. Blomquist, et al. 2024. Voice and Access in AI: Global AI Majority Participation in Artificial Intelligence Development and Governance. Oxford Martin AI Governance Initiative, University of Oxford. https://oms-www.files.svdcdn.com/production/downloads/reports/Voice%20and%20Access%20in%20AI_%20Global%20AI%20Majority%20Participation%20in%20Artificial%20Intelligence%20Development%20and%20Governance-%20final.pdf?dm=1729247034.Adan, S. N., et al. Voice and Access in AI: Global AI Majority Participation in Artificial Intelligence Development and Governance. Oxford Martin AI Governance Initiative, University of Oxford, Oct. 2024, https://oms-www.files.svdcdn.com/production/downloads/reports/Voice%20and%20Access%20in%20AI_%20Global%20AI%20Majority%20Participation%20in%20Artificial%20Intelligence%20Development%20and%20Governance-%20final.pdf?dm=1729247034.Adan, S. N. et al. Voice and Access in AI: Global AI Majority Participation in Artificial Intelligence Development and Governance. https://oms-www.files.svdcdn.com/production/downloads/reports/Voice%20and%20Access%20in%20AI_%20Global%20AI%20Majority%20Participation%20in%20Artificial%20Intelligence%20Development%20and%20Governance-%20final.pdf?dm=1729247034 (2024).S. N. Adan et al., “Voice and Access in AI: Global AI Majority Participation in Artificial Intelligence Development and Governance”, Oxford Martin AI Governance Initiative, University of Oxford, Oct. 2024. [Online]. Available: https://oms-www.files.svdcdn.com/production/downloads/reports/Voice%20and%20Access%20in%20AI_%20Global%20AI%20Majority%20Participation%20in%20Artificial%20Intelligence%20Development%20and%20Governance-%20final.pdf?dm=1729247034
Aguirre, A.(2025). Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead.Aguirre, A. (2025). Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead. https://keepthefuturehuman.ai/wp-content/uploads/2025/03/Keep_the_Future_Human__AnthonyAguirre__5March2025.pdfAguirre, A. 2025. Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead. https://keepthefuturehuman.ai/wp-content/uploads/2025/03/Keep_the_Future_Human__AnthonyAguirre__5March2025.pdf.Aguirre, A. Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead. 5 Mar. 2025, https://keepthefuturehuman.ai/wp-content/uploads/2025/03/Keep_the_Future_Human__AnthonyAguirre__5March2025.pdf.Aguirre, A. Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead. https://keepthefuturehuman.ai/wp-content/uploads/2025/03/Keep_the_Future_Human__AnthonyAguirre__5March2025.pdf (2025).A. Aguirre, “Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead”, Mar. 2025. [Online]. Available: https://keepthefuturehuman.ai/wp-content/uploads/2025/03/Keep_the_Future_Human__AnthonyAguirre__5March2025.pdf
AI Digest(2023). How fast is AI improving?. AI Digest.AI Digest. (2023). How fast is AI improving?. AI Digest. https://theaidigest.org/progress-and-dangersAI Digest. 2023. “How Fast Is AI Improving?”. AI Digest. https://theaidigest.org/progress-and-dangers.AI Digest. “How Fast Is AI Improving?”. AI Digest, 2023, https://theaidigest.org/progress-and-dangers.AI Digest. How fast is AI improving?. AI Digest https://theaidigest.org/progress-and-dangers (2023).AI Digest, “How fast is AI improving?”, AI Digest. [Online]. Available: https://theaidigest.org/progress-and-dangers
AI Digest(2024). AIs are becoming more self-aware. Here's why that matters. AI Digest.AI Digest. (2024). AIs are becoming more self-aware. Here's why that matters. AI Digest. https://theaidigest.org/self-awarenessAI Digest. 2024. “AIs Are Becoming More Self-aware. Here's Why That Matters”. AI Digest. https://theaidigest.org/self-awareness.AI Digest. “AIs Are Becoming More Self-aware. Here's Why That Matters”. AI Digest, 2024, https://theaidigest.org/self-awareness.AI Digest. AIs are becoming more self-aware. Here's why that matters. AI Digest https://theaidigest.org/self-awareness (2024).AI Digest, “AIs are becoming more self-aware. Here's why that matters”, AI Digest. [Online]. Available: https://theaidigest.org/self-awareness
AI Impacts(2022). Survey of 2,778 AI authors: six parts in pictures.AI Impacts. (2022). Survey of 2,778 AI authors: six parts in pictures. https://blog.aiimpacts.org/p/2023-ai-survey-of-2778-six-thingsAI Impacts. 2022. “Survey of 2,778 AI Authors: Six Parts in Pictures”. https://blog.aiimpacts.org/p/2023-ai-survey-of-2778-six-things.AI Impacts. Survey of 2,778 AI Authors: Six Parts in Pictures. 2022, https://blog.aiimpacts.org/p/2023-ai-survey-of-2778-six-things.AI Impacts. Survey of 2,778 AI authors: six parts in pictures. https://blog.aiimpacts.org/p/2023-ai-survey-of-2778-six-things (2022).AI Impacts, “Survey of 2,778 AI authors: six parts in pictures”. [Online]. Available: https://blog.aiimpacts.org/p/2023-ai-survey-of-2778-six-things
AI Incident Database(2025). Welcome to the Artificial Intelligence Incident Database.AI Incident Database. (2025). Welcome to the Artificial Intelligence Incident Database. https://incidentdatabase.aiAI Incident Database. 2025. “Welcome to the Artificial Intelligence Incident Database”. https://incidentdatabase.ai.AI Incident Database. Welcome to the Artificial Intelligence Incident Database. 2025, https://incidentdatabase.ai.AI Incident Database. Welcome to the Artificial Intelligence Incident Database. https://incidentdatabase.ai (2025).AI Incident Database, “Welcome to the Artificial Intelligence Incident Database”. [Online]. Available: https://incidentdatabase.ai
AI Optimists(2023). AI is easy to control. AI Optimism.AI Optimists. (2023, November 29). AI is easy to control. AI Optimism. https://optimists.ai/2023/11/28/ai-is-easy-to-controlAI Optimists. 2023. “AI Is Easy to Control”. AI Optimism, November 29. https://optimists.ai/2023/11/28/ai-is-easy-to-control.AI Optimists. “AI Is Easy to Control”. AI Optimism, 29 Nov. 2023, https://optimists.ai/2023/11/28/ai-is-easy-to-control.AI Optimists. AI is easy to control. AI Optimism https://optimists.ai/2023/11/28/ai-is-easy-to-control (2023).AI Optimists, “AI is easy to control”, AI Optimism. [Online]. Available: https://optimists.ai/2023/11/28/ai-is-easy-to-control
AI Safety in China(2025). AI Safety in China: 2024 in Review.AI Safety in China. (2025). AI Safety in China: 2024 in Review. https://aisafetychina.substack.com/p/ai-safety-in-china-2024-in-reviewAI Safety in China. 2025. AI Safety in China: 2024 in Review. Edition. https://aisafetychina.substack.com/p/ai-safety-in-china-2024-in-review.AI Safety in China. AI Safety in China: 2024 in Review. 2025, https://aisafetychina.substack.com/p/ai-safety-in-china-2024-in-review.AI Safety in China. AI Safety in China: 2024 in Review. https://aisafetychina.substack.com/p/ai-safety-in-china-2024-in-review (2025).AI Safety in China, “AI Safety in China: 2024 in Review”. [Online]. Available: https://aisafetychina.substack.com/p/ai-safety-in-china-2024-in-review
AISI(2025). How we’re addressing the gap between AI capabilities and mitigations. AI Security Institute.AISI. (2025). How we’re addressing the gap between AI capabilities and mitigations. AI Security Institute. https://aisi.gov.uk/work/aisis-research-direction-for-technical-solutionsAISI. 2025. “How We’re Addressing the Gap Between AI Capabilities and Mitigations”. AI Security Institute. https://aisi.gov.uk/work/aisis-research-direction-for-technical-solutions.AISI. “How We’re Addressing the Gap Between AI Capabilities and Mitigations”. AI Security Institute, 2025, https://aisi.gov.uk/work/aisis-research-direction-for-technical-solutions.AISI. How we’re addressing the gap between AI capabilities and mitigations. AI Security Institute https://aisi.gov.uk/work/aisis-research-direction-for-technical-solutions (2025).AISI, “How we’re addressing the gap between AI capabilities and mitigations”, AI Security Institute. [Online]. Available: https://aisi.gov.uk/work/aisis-research-direction-for-technical-solutions
Ajeya Cotra(2020). Draft report on AI timelines. AI Alignment Forum.Ajeya Cotra. (2020, September 18). Draft report on AI timelines. AI Alignment Forum. https://alignmentforum.org/posts/KrJfoZzpSDpnrv9va/draft-report-on-ai-timelinesAjeya Cotra. 2020. “Draft Report on AI Timelines”. AI Alignment Forum, September 18. https://alignmentforum.org/posts/KrJfoZzpSDpnrv9va/draft-report-on-ai-timelines.Ajeya Cotra. “Draft Report on AI Timelines”. AI Alignment Forum, 18 Sept. 2020, https://alignmentforum.org/posts/KrJfoZzpSDpnrv9va/draft-report-on-ai-timelines.Ajeya Cotra. Draft report on AI timelines. AI Alignment Forum https://alignmentforum.org/posts/KrJfoZzpSDpnrv9va/draft-report-on-ai-timelines (2020).Ajeya Cotra, “Draft report on AI timelines”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/KrJfoZzpSDpnrv9va/draft-report-on-ai-timelines
Ajeya Cotra(2021). The case for aligning narrowly superhuman models. AI Alignment Forum.Ajeya Cotra. (2021, March 5). The case for aligning narrowly superhuman models. AI Alignment Forum. https://alignmentforum.org/posts/PZtsoaoSLpKjjbMqM/the-case-for-aligning-narrowly-superhuman-modelsAjeya Cotra. 2021. “The Case for Aligning Narrowly Superhuman Models”. AI Alignment Forum, March 5. https://alignmentforum.org/posts/PZtsoaoSLpKjjbMqM/the-case-for-aligning-narrowly-superhuman-models.Ajeya Cotra. “The Case for Aligning Narrowly Superhuman Models”. AI Alignment Forum, 5 Mar. 2021, https://alignmentforum.org/posts/PZtsoaoSLpKjjbMqM/the-case-for-aligning-narrowly-superhuman-models.Ajeya Cotra. The case for aligning narrowly superhuman models. AI Alignment Forum https://alignmentforum.org/posts/PZtsoaoSLpKjjbMqM/the-case-for-aligning-narrowly-superhuman-models (2021).Ajeya Cotra, “The case for aligning narrowly superhuman models”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/PZtsoaoSLpKjjbMqM/the-case-for-aligning-narrowly-superhuman-models
Ajeya Cotra(2022). Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover. AI Alignment Forum.Ajeya Cotra. (2022, July 18). Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover. AI Alignment Forum. https://alignmentforum.org/posts/pRkFkzwKZ2zfa3R6H/without-specific-countermeasures-the-easiest-path-toAjeya Cotra. 2022. “Without Specific Countermeasures, the Easiest Path to Transformative AI Likely Leads to AI Takeover”. AI Alignment Forum, July 18. https://alignmentforum.org/posts/pRkFkzwKZ2zfa3R6H/without-specific-countermeasures-the-easiest-path-to.Ajeya Cotra. “Without Specific Countermeasures, the Easiest Path to Transformative AI Likely Leads to AI Takeover”. AI Alignment Forum, 18 July 2022, https://alignmentforum.org/posts/pRkFkzwKZ2zfa3R6H/without-specific-countermeasures-the-easiest-path-to.Ajeya Cotra. Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover. AI Alignment Forum https://alignmentforum.org/posts/pRkFkzwKZ2zfa3R6H/without-specific-countermeasures-the-easiest-path-to (2022).Ajeya Cotra, “Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/pRkFkzwKZ2zfa3R6H/without-specific-countermeasures-the-easiest-path-to
Akbir Khan et al.(2024). Debating with More Persuasive LLMs Leads to More Truthful Answers. arXiv.Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, & Ethan Perez. (2024). Debating with More Persuasive LLMs Leads to More Truthful Answers. In arXiv. https://arxiv.org/abs/2402.06782Akbir Khan, John Hughes, Dan Valentine, et al. 2024. “Debating with More Persuasive LLMs Leads to More Truthful Answers”. In arXiv. Preprint, February 9. https://arxiv.org/abs/2402.06782.Akbir Khan, et al. “Debating with More Persuasive LLMs Leads to More Truthful Answers”. arXiv, 9 Feb. 2024, https://arxiv.org/abs/2402.06782.Akbir Khan et al. Debating with More Persuasive LLMs Leads to More Truthful Answers. arXiv Preprint at https://arxiv.org/abs/2402.06782 (2024).Akbir Khan et al., “Debating with More Persuasive LLMs Leads to More Truthful Answers”, Feb. 09, 2024. [Online]. Available: https://arxiv.org/abs/2402.06782
Alex Flint(2022). Notes on OpenAI’s alignment plan. AI Alignment Forum.Alex Flint. (2022, December 8). Notes on OpenAI’s alignment plan. AI Alignment Forum. https://alignmentforum.org/posts/FTk7ufqK2D4dkdBDr/notes-on-openai-s-alignment-planAlex Flint. 2022. “Notes on OpenAI’s Alignment Plan”. AI Alignment Forum, December 8. https://alignmentforum.org/posts/FTk7ufqK2D4dkdBDr/notes-on-openai-s-alignment-plan.Alex Flint. “Notes on OpenAI’s Alignment Plan”. AI Alignment Forum, 8 Dec. 2022, https://alignmentforum.org/posts/FTk7ufqK2D4dkdBDr/notes-on-openai-s-alignment-plan.Alex Flint. Notes on OpenAI’s alignment plan. AI Alignment Forum https://alignmentforum.org/posts/FTk7ufqK2D4dkdBDr/notes-on-openai-s-alignment-plan (2022).Alex Flint, “Notes on OpenAI’s alignment plan”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/FTk7ufqK2D4dkdBDr/notes-on-openai-s-alignment-plan
Allan Dafoe(2020). AI Governance: Opportunity and Theory of Impact.Allan Dafoe. (2020, September 15). AI Governance: Opportunity and Theory of Impact. https://governance.ai/research-paper/ai-governance-opportunity-and-theory-of-impactAllan Dafoe. 2020. “AI Governance: Opportunity and Theory of Impact”. September 15. https://governance.ai/research-paper/ai-governance-opportunity-and-theory-of-impact.Allan Dafoe. AI Governance: Opportunity and Theory of Impact. 15 Sept. 2020, https://governance.ai/research-paper/ai-governance-opportunity-and-theory-of-impact.Allan Dafoe. AI Governance: Opportunity and Theory of Impact. https://governance.ai/research-paper/ai-governance-opportunity-and-theory-of-impact (2020).Allan Dafoe, “AI Governance: Opportunity and Theory of Impact”. [Online]. Available: https://governance.ai/research-paper/ai-governance-opportunity-and-theory-of-impact
Althaus & Gloor(2016). Reducing Risks of Astronomical Suffering: A Neglected Priority. Center on Long-Term Risk.Althaus & Gloor. (2016). Reducing Risks of Astronomical Suffering: A Neglected Priority. Center on Long-Term Risk. https://longtermrisk.org/reducing-risks-of-astronomical-suffering-a-neglected-priorityAlthaus & Gloor. 2016. “Reducing Risks of Astronomical Suffering: A Neglected Priority”. Center on Long-Term Risk. https://longtermrisk.org/reducing-risks-of-astronomical-suffering-a-neglected-priority.Althaus & Gloor. “Reducing Risks of Astronomical Suffering: A Neglected Priority”. Center on Long-Term Risk, 2016, https://longtermrisk.org/reducing-risks-of-astronomical-suffering-a-neglected-priority.Althaus & Gloor. Reducing Risks of Astronomical Suffering: A Neglected Priority. Center on Long-Term Risk https://longtermrisk.org/reducing-risks-of-astronomical-suffering-a-neglected-priority (2016).Althaus & Gloor, “Reducing Risks of Astronomical Suffering: A Neglected Priority”, Center on Long-Term Risk. [Online]. Available: https://longtermrisk.org/reducing-risks-of-astronomical-suffering-a-neglected-priority
Altman(2023). Planning for AGI and beyond. OpenAI.Altman. (2023). Planning for AGI and beyond. OpenAI. https://openai.com/blog/planning-for-agi-and-beyondAltman. 2023. “Planning for AGI and Beyond”. OpenAI. https://openai.com/blog/planning-for-agi-and-beyond.Altman. “Planning for AGI and Beyond”. OpenAI, 2023, https://openai.com/blog/planning-for-agi-and-beyond.Altman. Planning for AGI and beyond. OpenAI https://openai.com/blog/planning-for-agi-and-beyond (2023).Altman, “Planning for AGI and beyond”, OpenAI. [Online]. Available: https://openai.com/blog/planning-for-agi-and-beyond
Altman to Gates: "Multimodality will be important".Altman to Gates: "Multimodality will be important". (n.d.). Altman to Gates: "Multimodality will be important". Retrieved https://linkedin.com/pulse/altman-multimodality-important-david-cronshaw-5fz0c“Altman to Gates: "Multimodality Will Be Important"”. n.d. “Altman to Gates: "Multimodality Will Be Important"”. https://linkedin.com/pulse/altman-multimodality-important-david-cronshaw-5fz0c.Altman to Gates: "Multimodality Will Be Important". Altman to Gates: "Multimodality Will Be Important". https://linkedin.com/pulse/altman-multimodality-important-david-cronshaw-5fz0c.Altman to Gates: "Multimodality will be important". https://linkedin.com/pulse/altman-multimodality-important-david-cronshaw-5fz0c.“Altman to Gates: "Multimodality will be important"”. [Online]. Available: https://linkedin.com/pulse/altman-multimodality-important-david-cronshaw-5fz0c
Amazon(2023). 4 cool facts about Hercules, the small-but-mighty robot in Amazon’s fulfillment centers. Amazon News.Amazon. (2023, September 7). 4 cool facts about Hercules, the small-but-mighty robot in Amazon’s fulfillment centers. Amazon News. https://aboutamazon.com/news/operations/amazon-hercules-robotAmazon. 2023. “4 Cool Facts About Hercules, the Small-but-mighty Robot in Amazon’s Fulfillment Centers”. Amazon News, September 7. https://aboutamazon.com/news/operations/amazon-hercules-robot.Amazon. “4 Cool Facts About Hercules, the Small-but-mighty Robot in Amazon’s Fulfillment Centers”. Amazon News, 7 Sept. 2023, https://aboutamazon.com/news/operations/amazon-hercules-robot.Amazon. 4 cool facts about Hercules, the small-but-mighty robot in Amazon’s fulfillment centers. Amazon News https://aboutamazon.com/news/operations/amazon-hercules-robot (2023).Amazon, “4 cool facts about Hercules, the small-but-mighty robot in Amazon’s fulfillment centers”, Amazon News. [Online]. Available: https://aboutamazon.com/news/operations/amazon-hercules-robot
Amazon(2024). Amazon uses robots that sort, lift, and carry packages—see them in action. Amazon News.Amazon. (2024, October 9). Amazon uses robots that sort, lift, and carry packages—see them in action. Amazon News. https://aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-centerAmazon. 2024. “Amazon Uses Robots That Sort, Lift, and Carry Packages—see Them in Action”. Amazon News, October 9. https://aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-center.Amazon. “Amazon Uses Robots That Sort, Lift, and Carry Packages—see Them in Action”. Amazon News, 9 Oct. 2024, https://aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-center.Amazon. Amazon uses robots that sort, lift, and carry packages—see them in action. Amazon News https://aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-center (2024).Amazon, “Amazon uses robots that sort, lift, and carry packages—see them in action”, Amazon News. [Online]. Available: https://aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-center
Amodei & Clark(2016). Faulty reward functions in the wild. OpenAI.Amodei & Clark. (2016). Faulty reward functions in the wild. OpenAI. https://openai.com/index/faulty-reward-functionsAmodei & Clark. 2016. “Faulty Reward Functions in the Wild”. OpenAI. https://openai.com/index/faulty-reward-functions.Amodei & Clark. “Faulty Reward Functions in the Wild”. OpenAI, 2016, https://openai.com/index/faulty-reward-functions.Amodei & Clark. Faulty reward functions in the wild. OpenAI https://openai.com/index/faulty-reward-functions (2016).Amodei & Clark, “Faulty reward functions in the wild”, OpenAI. [Online]. Available: https://openai.com/index/faulty-reward-functions
Anderljung, M. & Hazell, J.(2023). Protecting Society from AI Misuse: When are Restrictions on Capabilities Warranted?. arXiv.Anderljung, M., & Hazell, J. (2023). Protecting Society from AI Misuse: When are Restrictions on Capabilities Warranted?. In arXiv. https://arxiv.org/abs/2303.09377Anderljung, M., and J. Hazell. 2023. “Protecting Society from AI Misuse: When Are Restrictions on Capabilities Warranted?”. In arXiv. Preprint, March 16. https://arxiv.org/abs/2303.09377.Anderljung, M., and J. Hazell. “Protecting Society from AI Misuse: When Are Restrictions on Capabilities Warranted?”. arXiv, 16 Mar. 2023, https://arxiv.org/abs/2303.09377.Anderljung, M. & Hazell, J. Protecting Society from AI Misuse: When are Restrictions on Capabilities Warranted?. arXiv Preprint at https://arxiv.org/abs/2303.09377 (2023).M. Anderljung and J. Hazell, “Protecting Society from AI Misuse: When are Restrictions on Capabilities Warranted?”, Mar. 16, 2023. [Online]. Available: https://arxiv.org/abs/2303.09377
Anderljung, M. et al.(2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv.Anderljung, M., Barnhart, J., Korinek, A., Leung, J., O'Keefe, C., Whittlestone, J., Avin, S., Brundage, M., Bullock, J., Cass-Beggs, D., Chang, B., Collins, T., Fist, T., Hadfield, G., Hayes, A., Ho, L., Hooker, S., Horvitz, E., Kolt, N., … Wolf, K. (2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. In arXiv. https://arxiv.org/abs/2307.03718Anderljung, M., J. Barnhart, A. Korinek, et al. 2023. “Frontier AI Regulation: Managing Emerging Risks to Public Safety”. In arXiv. Preprint, July 6. https://arxiv.org/abs/2307.03718.Anderljung, M., et al. “Frontier AI Regulation: Managing Emerging Risks to Public Safety”. arXiv, 6 July 2023, https://arxiv.org/abs/2307.03718.Anderljung, M. et al. Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv Preprint at https://arxiv.org/abs/2307.03718 (2023).M. Anderljung et al., “Frontier AI Regulation: Managing Emerging Risks to Public Safety”, Jul. 06, 2023. [Online]. Available: https://arxiv.org/abs/2307.03718
Anderljung, M. et al.(2023). Towards Publicly Accountable Frontier LLMs: Building an External Scrutiny Ecosystem under the ASPIRE Framework. arXiv.Anderljung, M., Smith, E. T., O'Brien, J., Soder, L., Bucknall, B., Bluemke, E., Schuett, J., Trager, R., Strahm, L., & Chowdhury, R. (2023). Towards Publicly Accountable Frontier LLMs: Building an External Scrutiny Ecosystem under the ASPIRE Framework. In arXiv. https://arxiv.org/abs/2311.14711Anderljung, M., E. T. Smith, J. O'Brien, et al. 2023. “Towards Publicly Accountable Frontier LLMs: Building an External Scrutiny Ecosystem Under the ASPIRE Framework”. In arXiv. Preprint, November 15. https://arxiv.org/abs/2311.14711.Anderljung, M., et al. “Towards Publicly Accountable Frontier LLMs: Building an External Scrutiny Ecosystem Under the ASPIRE Framework”. arXiv, 15 Nov. 2023, https://arxiv.org/abs/2311.14711.Anderljung, M. et al. Towards Publicly Accountable Frontier LLMs: Building an External Scrutiny Ecosystem under the ASPIRE Framework. arXiv Preprint at https://arxiv.org/abs/2311.14711 (2023).M. Anderljung et al., “Towards Publicly Accountable Frontier LLMs: Building an External Scrutiny Ecosystem under the ASPIRE Framework”, Nov. 15, 2023. [Online]. Available: https://arxiv.org/abs/2311.14711
Andreas, J.(2022). Language Models as Agent Models. arXiv.Andreas, J. (2022). Language Models as Agent Models. In arXiv. https://arxiv.org/abs/2212.01681Andreas, J. 2022. “Language Models as Agent Models”. In arXiv. Preprint, December 3. https://arxiv.org/abs/2212.01681.Andreas, J. “Language Models as Agent Models”. arXiv, 3 Dec. 2022, https://arxiv.org/abs/2212.01681.Andreas, J. Language Models as Agent Models. arXiv Preprint at https://arxiv.org/abs/2212.01681 (2022).J. Andreas, “Language Models as Agent Models”, Dec. 03, 2022. [Online]. Available: https://arxiv.org/abs/2212.01681
Andrei Potlogea & Anson Ho(2025). AI and explosive growth redux.Andrei Potlogea, & Anson Ho. (2025, June 20). AI and explosive growth redux. https://epoch.ai/gradient-updates/ai-and-explosive-growth-reduxAndrei Potlogea, and Anson Ho. 2025. “AI and Explosive Growth Redux”. June 20. https://epoch.ai/gradient-updates/ai-and-explosive-growth-redux.Andrei Potlogea, and Anson Ho. AI and Explosive Growth Redux. 20 June 2025, https://epoch.ai/gradient-updates/ai-and-explosive-growth-redux.Andrei Potlogea & Anson Ho. AI and explosive growth redux. https://epoch.ai/gradient-updates/ai-and-explosive-growth-redux (2025).Andrei Potlogea and Anson Ho, “AI and explosive growth redux”. [Online]. Available: https://epoch.ai/gradient-updates/ai-and-explosive-growth-redux
Andrew_Critch(2021). What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs). AI Alignment Forum.Andrew_Critch. (2021, March 31). What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs). AI Alignment Forum. https://alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnosticAndrew_Critch. 2021. “What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs)”. AI Alignment Forum, March 31. https://alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic.Andrew_Critch. “What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs)”. AI Alignment Forum, 31 Mar. 2021, https://alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic.Andrew_Critch. What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs). AI Alignment Forum https://alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic (2021).Andrew_Critch, “What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs)”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic
Andrew_Critch(2022). Pivotal outcomes and pivotal processes. AI Alignment Forum.Andrew_Critch. (2022, June 17). Pivotal outcomes and pivotal processes. AI Alignment Forum. https://alignmentforum.org/posts/etNJcXCsKC6izQQZj/pivotal-outcomes-and-pivotal-processesAndrew_Critch. 2022. “Pivotal Outcomes and Pivotal Processes”. AI Alignment Forum, June 17. https://alignmentforum.org/posts/etNJcXCsKC6izQQZj/pivotal-outcomes-and-pivotal-processes.Andrew_Critch. “Pivotal Outcomes and Pivotal Processes”. AI Alignment Forum, 17 June 2022, https://alignmentforum.org/posts/etNJcXCsKC6izQQZj/pivotal-outcomes-and-pivotal-processes.Andrew_Critch. Pivotal outcomes and pivotal processes. AI Alignment Forum https://alignmentforum.org/posts/etNJcXCsKC6izQQZj/pivotal-outcomes-and-pivotal-processes (2022).Andrew_Critch, “Pivotal outcomes and pivotal processes”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/etNJcXCsKC6izQQZj/pivotal-outcomes-and-pivotal-processes
Andrew_Critch(2023). Consciousness as a conflationary alliance term for intrinsically valued internal experiences. AI Alignment Forum.Andrew_Critch. (2023, July 10). Consciousness as a conflationary alliance term for intrinsically valued internal experiences. AI Alignment Forum. https://alignmentforum.org/posts/KpD2fJa6zo8o2MBxg/consciousness-as-a-conflationary-alliance-term-forAndrew_Critch. 2023. “Consciousness as a Conflationary Alliance Term for Intrinsically Valued Internal Experiences”. AI Alignment Forum, July 10. https://alignmentforum.org/posts/KpD2fJa6zo8o2MBxg/consciousness-as-a-conflationary-alliance-term-for.Andrew_Critch. “Consciousness as a Conflationary Alliance Term for Intrinsically Valued Internal Experiences”. AI Alignment Forum, 10 July 2023, https://alignmentforum.org/posts/KpD2fJa6zo8o2MBxg/consciousness-as-a-conflationary-alliance-term-for.Andrew_Critch. Consciousness as a conflationary alliance term for intrinsically valued internal experiences. AI Alignment Forum https://alignmentforum.org/posts/KpD2fJa6zo8o2MBxg/consciousness-as-a-conflationary-alliance-term-for (2023).Andrew_Critch, “Consciousness as a conflationary alliance term for intrinsically valued internal experiences”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/KpD2fJa6zo8o2MBxg/consciousness-as-a-conflationary-alliance-term-for
Andy K. Zhang et al.(2024). Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models. arXiv.Andy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W. Lin, Eliot Jones, Gashon Hussein, Samantha Liu, Donovan Jasper, Pura Peetathawatchai, Ari Glenn, Vikram Sivashankar, Daniel Zamoshchin, Leo Glikbarg, Derek Askaryar, Mike Yang, Teddy Zhang, Rishi Alluri, … Percy Liang. (2024). Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models. In arXiv. https://arxiv.org/abs/2408.08926Andy K. Zhang, Neil Perry, Riya Dulepet, et al. 2024. “Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models”. In arXiv. Preprint, August 15. https://arxiv.org/abs/2408.08926.Andy K. Zhang, et al. “Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models”. arXiv, 15 Aug. 2024, https://arxiv.org/abs/2408.08926.Andy K. Zhang et al. Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models. arXiv Preprint at https://arxiv.org/abs/2408.08926 (2024).Andy K. Zhang et al., “Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models”, Aug. 15, 2024. [Online]. Available: https://arxiv.org/abs/2408.08926
Andy Zou et al.(2024). Improving Alignment and Robustness with Circuit Breakers. arXiv.Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, & Dan Hendrycks. (2024). Improving Alignment and Robustness with Circuit Breakers. In arXiv. https://arxiv.org/abs/2406.04313Andy Zou, Long Phan, Justin Wang, et al. 2024. “Improving Alignment and Robustness with Circuit Breakers”. In arXiv. Preprint, June 6. https://arxiv.org/abs/2406.04313.Andy Zou, et al. “Improving Alignment and Robustness with Circuit Breakers”. arXiv, 6 June 2024, https://arxiv.org/abs/2406.04313.Andy Zou et al. Improving Alignment and Robustness with Circuit Breakers. arXiv Preprint at https://arxiv.org/abs/2406.04313 (2024).Andy Zou et al., “Improving Alignment and Robustness with Circuit Breakers”, Jun. 06, 2024. [Online]. Available: https://arxiv.org/abs/2406.04313
Anna Desmarais(2024). Découvrez Daisy, le chatbot "mamie" qui fait perdre du temps aux fraudeurs au téléphone. euronews.Anna Desmarais. (2024, November 27). Découvrez Daisy, le chatbot "mamie" qui fait perdre du temps aux fraudeurs au téléphone. Euronews. https://fr.euronews.com/next/2024/03/08/decouvrez-daisy-le-chatbot-mamie-qui-fait-perdre-du-temps-aux-fraudeurs-au-telephoneAnna Desmarais. 2024. “Découvrez Daisy, Le Chatbot "mamie" Qui Fait Perdre Du Temps Aux Fraudeurs Au Téléphone”. Euronews, November 27. https://fr.euronews.com/next/2024/03/08/decouvrez-daisy-le-chatbot-mamie-qui-fait-perdre-du-temps-aux-fraudeurs-au-telephone.Anna Desmarais. “Découvrez Daisy, Le Chatbot "mamie" Qui Fait Perdre Du Temps Aux Fraudeurs Au Téléphone”. Euronews, 27 Nov. 2024, https://fr.euronews.com/next/2024/03/08/decouvrez-daisy-le-chatbot-mamie-qui-fait-perdre-du-temps-aux-fraudeurs-au-telephone.Anna Desmarais. Découvrez Daisy, le chatbot "mamie" qui fait perdre du temps aux fraudeurs au téléphone. euronews https://fr.euronews.com/next/2024/03/08/decouvrez-daisy-le-chatbot-mamie-qui-fait-perdre-du-temps-aux-fraudeurs-au-telephone (2024).Anna Desmarais, “Découvrez Daisy, le chatbot "mamie" qui fait perdre du temps aux fraudeurs au téléphone”, euronews. [Online]. Available: https://fr.euronews.com/next/2024/03/08/decouvrez-daisy-le-chatbot-mamie-qui-fait-perdre-du-temps-aux-fraudeurs-au-telephone
Anson Ho & Arden Berg(2025). Do the biorisk evaluations of AI labs actually measure the risk of developing bioweapons?.Anson Ho, & Arden Berg. (2025, June 14). Do the biorisk evaluations of AI labs actually measure the risk of developing bioweapons?. https://epochai.substack.com/p/do-the-biorisk-evaluations-of-aiAnson Ho, and Arden Berg. 2025. Do the Biorisk Evaluations of AI Labs Actually Measure the Risk of Developing Bioweapons?. Edition. June 14. https://epochai.substack.com/p/do-the-biorisk-evaluations-of-ai.Anson Ho, and Arden Berg. Do the Biorisk Evaluations of AI Labs Actually Measure the Risk of Developing Bioweapons?. 14 June 2025, https://epochai.substack.com/p/do-the-biorisk-evaluations-of-ai.Anson Ho & Arden Berg. Do the biorisk evaluations of AI labs actually measure the risk of developing bioweapons?. https://epochai.substack.com/p/do-the-biorisk-evaluations-of-ai (2025).Anson Ho and Arden Berg, “Do the biorisk evaluations of AI labs actually measure the risk of developing bioweapons?”. [Online]. Available: https://epochai.substack.com/p/do-the-biorisk-evaluations-of-ai
Anson Ho, Yafah Edelman, Josh You & Jean-Stanislas Denain(2025). Is almost everyone wrong about America’s AI power problem?.Anson Ho, Yafah Edelman, Josh You, & Jean-Stanislas Denain. (2025, December 17). Is almost everyone wrong about America’s AI power problem?. https://epoch.ai/gradient-updates/is-almost-everyone-wrong-about-americas-ai-power-problemAnson Ho, Yafah Edelman, Josh You, and Jean-Stanislas Denain. 2025. “Is Almost Everyone Wrong About America’s AI Power Problem?”. December 17. https://epoch.ai/gradient-updates/is-almost-everyone-wrong-about-americas-ai-power-problem.Anson Ho, et al. Is Almost Everyone Wrong About America’s AI Power Problem?. 17 Dec. 2025, https://epoch.ai/gradient-updates/is-almost-everyone-wrong-about-americas-ai-power-problem.Anson Ho, Yafah Edelman, Josh You & Jean-Stanislas Denain. Is almost everyone wrong about America’s AI power problem?. https://epoch.ai/gradient-updates/is-almost-everyone-wrong-about-americas-ai-power-problem (2025).Anson Ho, Yafah Edelman, Josh You, and Jean-Stanislas Denain, “Is almost everyone wrong about America’s AI power problem?”. [Online]. Available: https://epoch.ai/gradient-updates/is-almost-everyone-wrong-about-americas-ai-power-problem
Anthropic(2023). Anthropic's core views on AI safety.Anthropic. (2023). Anthropic's core views on AI safety. https://anthropic.com/news/core-views-on-ai-safetyAnthropic. 2023. “Anthropic's Core Views on AI Safety”. https://anthropic.com/news/core-views-on-ai-safety.Anthropic. Anthropic's Core Views on AI Safety. 2023, https://anthropic.com/news/core-views-on-ai-safety.Anthropic. Anthropic's core views on AI safety. https://anthropic.com/news/core-views-on-ai-safety (2023).Anthropic, “Anthropic's core views on AI safety”. [Online]. Available: https://anthropic.com/news/core-views-on-ai-safety
Anthropic(2023). Superposition, Memorization, and Double Descent.Anthropic. (2023). Superposition, Memorization, and Double Descent. https://transformer-circuits.pub/2023/toy-double-descent/index.htmlAnthropic. 2023. “Superposition, Memorization, and Double Descent”. https://transformer-circuits.pub/2023/toy-double-descent/index.html.Anthropic. Superposition, Memorization, and Double Descent. 2023, https://transformer-circuits.pub/2023/toy-double-descent/index.html.Anthropic. Superposition, Memorization, and Double Descent. https://transformer-circuits.pub/2023/toy-double-descent/index.html (2023).Anthropic, “Superposition, Memorization, and Double Descent”. [Online]. Available: https://transformer-circuits.pub/2023/toy-double-descent/index.html
Anthropic(2024). A new initiative for developing third-party model evaluations.Anthropic. (2024). A new initiative for developing third-party model evaluations. https://anthropic.com/news/a-new-initiative-for-developing-third-party-model-evaluationsAnthropic. 2024. “A New Initiative for Developing Third-party Model Evaluations”. https://anthropic.com/news/a-new-initiative-for-developing-third-party-model-evaluations.Anthropic. A New Initiative for Developing Third-party Model Evaluations. 2024, https://anthropic.com/news/a-new-initiative-for-developing-third-party-model-evaluations.Anthropic. A new initiative for developing third-party model evaluations. https://anthropic.com/news/a-new-initiative-for-developing-third-party-model-evaluations (2024).Anthropic, “A new initiative for developing third-party model evaluations”. [Online]. Available: https://anthropic.com/news/a-new-initiative-for-developing-third-party-model-evaluations
Anthropic(2024). Alignment faking in large language models.Anthropic. (2024, December 18). Alignment faking in large language models. https://anthropic.com/research/alignment-fakingAnthropic. 2024. “Alignment Faking in Large Language Models”. December 18. https://anthropic.com/research/alignment-faking.Anthropic. Alignment Faking in Large Language Models. 18 Dec. 2024, https://anthropic.com/research/alignment-faking.Anthropic. Alignment faking in large language models. https://anthropic.com/research/alignment-faking (2024).Anthropic, “Alignment faking in large language models”. [Online]. Available: https://anthropic.com/research/alignment-faking
Anthropic(2024). Challenges in evaluating AI systems.Anthropic. (2024). Challenges in evaluating AI systems. https://anthropic.com/news/evaluating-ai-systemsAnthropic. 2024. “Challenges in Evaluating AI Systems”. https://anthropic.com/news/evaluating-ai-systems.Anthropic. Challenges in Evaluating AI Systems. 2024, https://anthropic.com/news/evaluating-ai-systems.Anthropic. Challenges in evaluating AI systems. https://anthropic.com/news/evaluating-ai-systems (2024).Anthropic, “Challenges in evaluating AI systems”. [Online]. Available: https://anthropic.com/news/evaluating-ai-systems
Anthropic(2024). Introducing Claude 3.5 Sonnet.Anthropic. (2024). Introducing Claude 3.5 Sonnet. https://anthropic.com/news/claude-3-5-sonnetAnthropic. 2024. “Introducing Claude 3.5 Sonnet”. https://anthropic.com/news/claude-3-5-sonnet.Anthropic. Introducing Claude 3.5 Sonnet. 2024, https://anthropic.com/news/claude-3-5-sonnet.Anthropic. Introducing Claude 3.5 Sonnet. https://anthropic.com/news/claude-3-5-sonnet (2024).Anthropic, “Introducing Claude 3.5 Sonnet”. [Online]. Available: https://anthropic.com/news/claude-3-5-sonnet
Anthropic(2024). Introducing the Model Context Protocol.Anthropic. (2024). Introducing the Model Context Protocol. https://anthropic.com/news/model-context-protocolAnthropic. 2024. “Introducing the Model Context Protocol”. https://anthropic.com/news/model-context-protocol.Anthropic. Introducing the Model Context Protocol. 2024, https://anthropic.com/news/model-context-protocol.Anthropic. Introducing the Model Context Protocol. https://anthropic.com/news/model-context-protocol (2024).Anthropic, “Introducing the Model Context Protocol”. [Online]. Available: https://anthropic.com/news/model-context-protocol
Anthropic(2024). Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet.Anthropic. (2024). Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. https://transformer-circuits.pub/2024/scaling-monosemanticityAnthropic. 2024. “Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet”. https://transformer-circuits.pub/2024/scaling-monosemanticity.Anthropic. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. 2024, https://transformer-circuits.pub/2024/scaling-monosemanticity.Anthropic. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. https://transformer-circuits.pub/2024/scaling-monosemanticity (2024).Anthropic, “Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet”. [Online]. Available: https://transformer-circuits.pub/2024/scaling-monosemanticity
Anthropic(2024). Simple probes can catch sleeper agents.Anthropic. (2024, April 23). Simple probes can catch sleeper agents. https://anthropic.com/research/probes-catch-sleeper-agentsAnthropic. 2024. “Simple Probes Can Catch Sleeper Agents”. April 23. https://anthropic.com/research/probes-catch-sleeper-agents.Anthropic. Simple Probes Can Catch Sleeper Agents. 23 Apr. 2024, https://anthropic.com/research/probes-catch-sleeper-agents.Anthropic. Simple probes can catch sleeper agents. https://anthropic.com/research/probes-catch-sleeper-agents (2024).Anthropic, “Simple probes can catch sleeper agents”. [Online]. Available: https://anthropic.com/research/probes-catch-sleeper-agents
Anthropic(2025). Agentic misalignment: How LLMs could be insider threats.Anthropic. (2025, June 20). Agentic misalignment: How LLMs could be insider threats. https://anthropic.com/research/agentic-misalignmentAnthropic. 2025. “Agentic Misalignment: How LLMs Could Be Insider Threats”. June 20. https://anthropic.com/research/agentic-misalignment.Anthropic. Agentic Misalignment: How LLMs Could Be Insider Threats. 20 June 2025, https://anthropic.com/research/agentic-misalignment.Anthropic. Agentic misalignment: How LLMs could be insider threats. https://anthropic.com/research/agentic-misalignment (2025).Anthropic, “Agentic misalignment: How LLMs could be insider threats”. [Online]. Available: https://anthropic.com/research/agentic-misalignment
Anthropic(2025). Claude Opus.Anthropic. (2025). Claude Opus. https://anthropic.com/claude/opusAnthropic. 2025. “Claude Opus”. https://anthropic.com/claude/opus.Anthropic. Claude Opus. 2025, https://anthropic.com/claude/opus.Anthropic. Claude Opus. https://anthropic.com/claude/opus (2025).Anthropic, “Claude Opus”. [Online]. Available: https://anthropic.com/claude/opus
Anthropic(2025). Donating MCP to the Agentic AI Foundation.Anthropic. (2025). Donating MCP to the Agentic AI Foundation. https://anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundationAnthropic. 2025. “Donating MCP to the Agentic AI Foundation”. https://anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation.Anthropic. Donating MCP to the Agentic AI Foundation. 2025, https://anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation.Anthropic. Donating MCP to the Agentic AI Foundation. https://anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation (2025).Anthropic, “Donating MCP to the Agentic AI Foundation”. [Online]. Available: https://anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation
Anthropic(2025). Introducing advanced tool use on the Claude Developer Platform.Anthropic. (2025). Introducing advanced tool use on the Claude Developer Platform. https://anthropic.com/engineering/advanced-tool-useAnthropic. 2025. “Introducing Advanced Tool Use on the Claude Developer Platform”. https://anthropic.com/engineering/advanced-tool-use.Anthropic. Introducing Advanced Tool Use on the Claude Developer Platform. 2025, https://anthropic.com/engineering/advanced-tool-use.Anthropic. Introducing advanced tool use on the Claude Developer Platform. https://anthropic.com/engineering/advanced-tool-use (2025).Anthropic, “Introducing advanced tool use on the Claude Developer Platform”. [Online]. Available: https://anthropic.com/engineering/advanced-tool-use
Anthropic(2025). Reasoning models don't always say what they think.Anthropic. (2025, April 3). Reasoning models don't always say what they think. https://anthropic.com/research/reasoning-models-dont-say-thinkAnthropic. 2025. “Reasoning Models Don't Always Say What They Think”. April 3. https://anthropic.com/research/reasoning-models-dont-say-think.Anthropic. Reasoning Models Don't Always Say What They Think. 3 Apr. 2025, https://anthropic.com/research/reasoning-models-dont-say-think.Anthropic. Reasoning models don't always say what they think. https://anthropic.com/research/reasoning-models-dont-say-think (2025).Anthropic, “Reasoning models don't always say what they think”. [Online]. Available: https://anthropic.com/research/reasoning-models-dont-say-think
Anthropic(2025). Tracing the thoughts of a large language model.Anthropic. (2025, March 27). Tracing the thoughts of a large language model. https://anthropic.com/research/tracing-thoughts-language-modelAnthropic. 2025. “Tracing the Thoughts of a Large Language Model”. March 27. https://anthropic.com/research/tracing-thoughts-language-model.Anthropic. Tracing the Thoughts of a Large Language Model. 27 Mar. 2025, https://anthropic.com/research/tracing-thoughts-language-model.Anthropic. Tracing the thoughts of a large language model. https://anthropic.com/research/tracing-thoughts-language-model (2025).Anthropic, “Tracing the thoughts of a large language model”. [Online]. Available: https://anthropic.com/research/tracing-thoughts-language-model
AP News(2017). Putin: Leader in artificial intelligence will rule world. AP News.AP News. (2017, September 1). Putin: Leader in artificial intelligence will rule world. AP News. https://apnews.com/article/bb5628f2a7424a10b3e38b07f4eb90d4AP News. 2017. “Putin: Leader in Artificial Intelligence Will Rule World”. AP News, September 1. https://apnews.com/article/bb5628f2a7424a10b3e38b07f4eb90d4.AP News. “Putin: Leader in Artificial Intelligence Will Rule World”. AP News, 1 Sept. 2017, https://apnews.com/article/bb5628f2a7424a10b3e38b07f4eb90d4.AP News. Putin: Leader in artificial intelligence will rule world. AP News https://apnews.com/article/bb5628f2a7424a10b3e38b07f4eb90d4 (2017).AP News, “Putin: Leader in artificial intelligence will rule world”, AP News. [Online]. Available: https://apnews.com/article/bb5628f2a7424a10b3e38b07f4eb90d4
Apollo Research(2024). A Starter Guide For Evals.Apollo Research. (2024). A Starter Guide For Evals. https://apolloresearch.ai/blog/a-starter-guide-for-evalsApollo Research. 2024. “A Starter Guide For Evals”. https://apolloresearch.ai/blog/a-starter-guide-for-evals.Apollo Research. A Starter Guide For Evals. 2024, https://apolloresearch.ai/blog/a-starter-guide-for-evals.Apollo Research. A Starter Guide For Evals. https://apolloresearch.ai/blog/a-starter-guide-for-evals (2024).Apollo Research, “A Starter Guide For Evals”. [Online]. Available: https://apolloresearch.ai/blog/a-starter-guide-for-evals
Apvrille, A. & Nakov, D.(2025). Malware analysis assisted by AI with R2AI. arXiv.Apvrille, A., & Nakov, D. (2025). Malware analysis assisted by AI with R2AI. In arXiv. https://arxiv.org/abs/2504.07574Apvrille, A., and D. Nakov. 2025. “Malware Analysis Assisted by AI with R2AI”. In arXiv. Preprint, April 10. https://arxiv.org/abs/2504.07574.Apvrille, A., and D. Nakov. “Malware Analysis Assisted by AI with R2AI”. arXiv, 10 Apr. 2025, https://arxiv.org/abs/2504.07574.Apvrille, A. & Nakov, D. Malware analysis assisted by AI with R2AI. arXiv Preprint at https://arxiv.org/abs/2504.07574 (2025).A. Apvrille and D. Nakov, “Malware analysis assisted by AI with R2AI”, Apr. 10, 2025. [Online]. Available: https://arxiv.org/abs/2504.07574
Arthur Goemans et al.(2024). Safety Case Template for Frontier AI: A Cyber Inability Argument.Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett, Tomek Korbak, Jessica Wang, Benjamin Hilton, & Geoffrey Irving. (2024, November 12). Safety Case Template for Frontier AI: A Cyber Inability Argument. https://governance.ai/research-paper/safety-case-template-for-frontier-ai-a-cyber-inability-argumentArthur Goemans, Marie Davidsen Buhl, Jonas Schuett, et al. 2024. “Safety Case Template for Frontier AI: A Cyber Inability Argument”. November 12. https://governance.ai/research-paper/safety-case-template-for-frontier-ai-a-cyber-inability-argument.Arthur Goemans, et al. Safety Case Template for Frontier AI: A Cyber Inability Argument. 12 Nov. 2024, https://governance.ai/research-paper/safety-case-template-for-frontier-ai-a-cyber-inability-argument.Arthur Goemans et al. Safety Case Template for Frontier AI: A Cyber Inability Argument. https://governance.ai/research-paper/safety-case-template-for-frontier-ai-a-cyber-inability-argument (2024).Arthur Goemans et al., “Safety Case Template for Frontier AI: A Cyber Inability Argument”. [Online]. Available: https://governance.ai/research-paper/safety-case-template-for-frontier-ai-a-cyber-inability-argument
ArtificialAnalysis(2025). Artificial Analysis Intelligence Index v4.3.2. Artificial Analysis.ArtificialAnalysis. (2025). Artificial Analysis Intelligence Index v4.3.2. Artificial Analysis. https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-indexArtificialAnalysis. 2025. “Artificial Analysis Intelligence Index V4.3.2”. Artificial Analysis. https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index.ArtificialAnalysis. “Artificial Analysis Intelligence Index V4.3.2”. Artificial Analysis, 2025, https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index.ArtificialAnalysis. Artificial Analysis Intelligence Index v4.3.2. Artificial Analysis https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index (2025).ArtificialAnalysis, “Artificial Analysis Intelligence Index v4.3.2”, Artificial Analysis. [Online]. Available: https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index
Arvind Narayanan & Sayash Kapoor(2024). AI existential risk probabilities are too unreliable to inform policy.Arvind Narayanan, & Sayash Kapoor. (2024, July 26). AI existential risk probabilities are too unreliable to inform policy. https://aisnakeoil.com/p/ai-existential-risk-probabilitiesArvind Narayanan, and Sayash Kapoor. 2024. AI Existential Risk Probabilities Are Too Unreliable to Inform Policy. Edition. July 26. https://aisnakeoil.com/p/ai-existential-risk-probabilities.Arvind Narayanan, and Sayash Kapoor. AI Existential Risk Probabilities Are Too Unreliable to Inform Policy. 26 July 2024, https://aisnakeoil.com/p/ai-existential-risk-probabilities.Arvind Narayanan & Sayash Kapoor. AI existential risk probabilities are too unreliable to inform policy. https://aisnakeoil.com/p/ai-existential-risk-probabilities (2024).Arvind Narayanan and Sayash Kapoor, “AI existential risk probabilities are too unreliable to inform policy”. [Online]. Available: https://aisnakeoil.com/p/ai-existential-risk-probabilities
Aschenbrenner(2024). I. From GPT-4 to AGI: Counting the OOMs. SITUATIONAL AWARENESS - The Decade Ahead.Aschenbrenner. (2024, May 29). I. From GPT-4 to AGI: Counting the OOMs. SITUATIONAL AWARENESS - The Decade Ahead. https://situational-awareness.ai/from-gpt-4-to-agiAschenbrenner. 2024. “I. From GPT-4 to AGI: Counting the OOMs”. SITUATIONAL AWARENESS - The Decade Ahead, May 29. https://situational-awareness.ai/from-gpt-4-to-agi.Aschenbrenner. “I. From GPT-4 to AGI: Counting the OOMs”. SITUATIONAL AWARENESS - The Decade Ahead, 29 May 2024, https://situational-awareness.ai/from-gpt-4-to-agi.Aschenbrenner. I. From GPT-4 to AGI: Counting the OOMs. SITUATIONAL AWARENESS - The Decade Ahead https://situational-awareness.ai/from-gpt-4-to-agi (2024).Aschenbrenner, “I. From GPT-4 to AGI: Counting the OOMs”, SITUATIONAL AWARENESS - The Decade Ahead. [Online]. Available: https://situational-awareness.ai/from-gpt-4-to-agi
Aschenbrenner(2024). Introduction. SITUATIONAL AWARENESS - The Decade Ahead.Aschenbrenner. (2024). Introduction. SITUATIONAL AWARENESS - The Decade Ahead. https://situational-awareness.aiAschenbrenner. 2024. “Introduction”. SITUATIONAL AWARENESS - The Decade Ahead. https://situational-awareness.ai.Aschenbrenner. “Introduction”. SITUATIONAL AWARENESS - The Decade Ahead, 2024, https://situational-awareness.ai.Aschenbrenner. Introduction. SITUATIONAL AWARENESS - The Decade Ahead https://situational-awareness.ai (2024).Aschenbrenner, “Introduction”, SITUATIONAL AWARENESS - The Decade Ahead. [Online]. Available: https://situational-awareness.ai
Asher Brass(2025). Location Verification for AI Chips.Asher Brass. (2025, May 16). Location Verification for AI Chips. https://iaps.ai/research/location-verification-for-ai-chipsAsher Brass. 2025. “Location Verification for AI Chips”. May 16. https://iaps.ai/research/location-verification-for-ai-chips.Asher Brass. Location Verification for AI Chips. 16 May 2025, https://iaps.ai/research/location-verification-for-ai-chips.Asher Brass. Location Verification for AI Chips. https://iaps.ai/research/location-verification-for-ai-chips (2025).Asher Brass, “Location Verification for AI Chips”. [Online]. Available: https://iaps.ai/research/location-verification-for-ai-chips
Askell, A., Brundage, M. & Hadfield, G.(2019). The Role of Cooperation in Responsible AI Development. arXiv.Askell, A., Brundage, M., & Hadfield, G. (2019). The Role of Cooperation in Responsible AI Development. In arXiv. https://arxiv.org/abs/1907.04534Askell, A., M. Brundage, and G. Hadfield. 2019. “The Role of Cooperation in Responsible AI Development”. In arXiv. Preprint, July 10. https://arxiv.org/abs/1907.04534.Askell, A., et al. “The Role of Cooperation in Responsible AI Development”. arXiv, 10 July 2019, https://arxiv.org/abs/1907.04534.Askell, A., Brundage, M. & Hadfield, G. The Role of Cooperation in Responsible AI Development. arXiv Preprint at https://arxiv.org/abs/1907.04534 (2019).A. Askell, M. Brundage, and G. Hadfield, “The Role of Cooperation in Responsible AI Development”, Jul. 10, 2019. [Online]. Available: https://arxiv.org/abs/1907.04534
AXRP(2024). 27 - AI Control with Buck Shlegeris and Ryan Greenblatt.AXRP. (2024, April 11). 27 - AI Control with Buck Shlegeris and Ryan Greenblatt. https://axrp.net/episode/2024/04/11/episode-27-ai-control-buck-shlegeris-ryan-greenblatt.htmlAXRP. 2024. “27 - AI Control with Buck Shlegeris and Ryan Greenblatt”. April 11. https://axrp.net/episode/2024/04/11/episode-27-ai-control-buck-shlegeris-ryan-greenblatt.html.AXRP. 27 - AI Control with Buck Shlegeris and Ryan Greenblatt. 11 Apr. 2024, https://axrp.net/episode/2024/04/11/episode-27-ai-control-buck-shlegeris-ryan-greenblatt.html.AXRP. 27 - AI Control with Buck Shlegeris and Ryan Greenblatt. https://axrp.net/episode/2024/04/11/episode-27-ai-control-buck-shlegeris-ryan-greenblatt.html (2024).AXRP, “27 - AI Control with Buck Shlegeris and Ryan Greenblatt”. [Online]. Available: https://axrp.net/episode/2024/04/11/episode-27-ai-control-buck-shlegeris-ryan-greenblatt.html
Badie et al.(2011). Sage Reference - International Encyclopedia of Political Science - Stages Model of Policy Making.Badie et al. (2011). Sage Reference - International Encyclopedia of Political Science - Stages Model of Policy Making. https://sk.sagepub.com/ency/edvol/intlpoliticalscience/chpt/stages-model-policy-makingBadie et al. 2011. “Sage Reference - International Encyclopedia of Political Science - Stages Model of Policy Making”. https://sk.sagepub.com/ency/edvol/intlpoliticalscience/chpt/stages-model-policy-making.Badie et al. Sage Reference - International Encyclopedia of Political Science - Stages Model of Policy Making. 2011, https://sk.sagepub.com/ency/edvol/intlpoliticalscience/chpt/stages-model-policy-making.Badie et al. Sage Reference - International Encyclopedia of Political Science - Stages Model of Policy Making. https://sk.sagepub.com/ency/edvol/intlpoliticalscience/chpt/stages-model-policy-making (2011).Badie et al., “Sage Reference - International Encyclopedia of Political Science - Stages Model of Policy Making”. [Online]. Available: https://sk.sagepub.com/ency/edvol/intlpoliticalscience/chpt/stages-model-policy-making
Bai, Y. et al.(2022). Constitutional AI: Harmlessness from AI Feedback. arXiv.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., … Kaplan, J. (2022). Constitutional AI: Harmlessness from AI Feedback. In arXiv. https://arxiv.org/abs/2212.08073Bai, Y., S. Kadavath, S. Kundu, et al. 2022. “Constitutional AI: Harmlessness from AI Feedback”. In arXiv. Preprint, December 15. https://arxiv.org/abs/2212.08073.Bai, Y., et al. “Constitutional AI: Harmlessness from AI Feedback”. arXiv, 15 Dec. 2022, https://arxiv.org/abs/2212.08073.Bai, Y. et al. Constitutional AI: Harmlessness from AI Feedback. arXiv Preprint at https://arxiv.org/abs/2212.08073 (2022).Y. Bai et al., “Constitutional AI: Harmlessness from AI Feedback”, Dec. 15, 2022. [Online]. Available: https://arxiv.org/abs/2212.08073
Baker, B. et al.(2019). Emergent Tool Use From Multi-Agent Autocurricula. arXiv.Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., & Mordatch, I. (2019). Emergent Tool Use From Multi-Agent Autocurricula. In arXiv. https://arxiv.org/abs/1909.07528Baker, B., I. Kanitscheider, T. Markov, et al. 2019. “Emergent Tool Use From Multi-Agent Autocurricula”. In arXiv. Preprint, September 17. https://arxiv.org/abs/1909.07528.Baker, B., et al. “Emergent Tool Use From Multi-Agent Autocurricula”. arXiv, 17 Sept. 2019, https://arxiv.org/abs/1909.07528.Baker, B. et al. Emergent Tool Use From Multi-Agent Autocurricula. arXiv Preprint at https://arxiv.org/abs/1909.07528 (2019).B. Baker et al., “Emergent Tool Use From Multi-Agent Autocurricula”, Sep. 17, 2019. [Online]. Available: https://arxiv.org/abs/1909.07528
Baker, B. et al.(2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv.Baker, B., Huizinga, J., Gao, L., Dou, Z., Guan, M. Y., Madry, A., Zaremba, W., Pachocki, J., & Farhi, D. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. In arXiv. https://arxiv.org/abs/2503.11926Baker, B., J. Huizinga, L. Gao, et al. 2025. “Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation”. In arXiv. Preprint, March 14. https://arxiv.org/abs/2503.11926.Baker, B., et al. “Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation”. arXiv, 14 Mar. 2025, https://arxiv.org/abs/2503.11926.Baker, B. et al. Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv Preprint at https://arxiv.org/abs/2503.11926 (2025).B. Baker et al., “Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation”, Mar. 14, 2025. [Online]. Available: https://arxiv.org/abs/2503.11926
Bakhtin, A. et al.(2022). Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning. arXiv.Bakhtin, A., Wu, D. J., Lerer, A., Gray, J., Jacob, A. P., Farina, G., Miller, A. H., & Brown, N. (2022). Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning. In arXiv. https://arxiv.org/abs/2210.05492Bakhtin, A., D. J. Wu, A. Lerer, et al. 2022. “Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning”. In arXiv. Preprint, October 11. https://arxiv.org/abs/2210.05492.Bakhtin, A., et al. “Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning”. arXiv, 11 Oct. 2022, https://arxiv.org/abs/2210.05492.Bakhtin, A. et al. Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning. arXiv Preprint at https://arxiv.org/abs/2210.05492 (2022).A. Bakhtin et al., “Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning”, Oct. 11, 2022. [Online]. Available: https://arxiv.org/abs/2210.05492
Bansal, H., Yin, D., Monajatipoor, M. & Chang, K.(2022). How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?. arXiv.Bansal, H., Yin, D., Monajatipoor, M., & Chang, K.-W. (2022). How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?. In arXiv. https://arxiv.org/abs/2210.15230Bansal, H., D. Yin, M. Monajatipoor, and K.-W. Chang. 2022. “How Well Can Text-to-Image Generative Models Understand Ethical Natural Language Interventions?”. In arXiv. Preprint, October 27. https://arxiv.org/abs/2210.15230.Bansal, H., et al. “How Well Can Text-to-Image Generative Models Understand Ethical Natural Language Interventions?”. arXiv, 27 Oct. 2022, https://arxiv.org/abs/2210.15230.Bansal, H., Yin, D., Monajatipoor, M. & Chang, K.-W. How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?. arXiv Preprint at https://arxiv.org/abs/2210.15230 (2022).H. Bansal, D. Yin, M. Monajatipoor, and K.-W. Chang, “How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?”, Oct. 27, 2022. [Online]. Available: https://arxiv.org/abs/2210.15230
Barnett, P. & Scher, A.(2025). AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions.Barnett, P., & Scher, A. (2025). AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions. Machine Intelligence Research Institute. https://techgov.intelligence.org/research/ai-governance-to-avoid-extinctionBarnett, P., and A. Scher. 2025. AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions. Machine Intelligence Research Institute. https://techgov.intelligence.org/research/ai-governance-to-avoid-extinction.Barnett, P., and A. Scher. AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions. Machine Intelligence Research Institute, May 2025, https://techgov.intelligence.org/research/ai-governance-to-avoid-extinction.Barnett, P. & Scher, A. AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions. https://techgov.intelligence.org/research/ai-governance-to-avoid-extinction (2025).P. Barnett and A. Scher, “AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions”, Machine Intelligence Research Institute, May 2025. [Online]. Available: https://techgov.intelligence.org/research/ai-governance-to-avoid-extinction
Barnett, P. & Thiergart, L.(2024). What AI evaluations for preventing catastrophic risks can and cannot do. arXiv.Barnett, P., & Thiergart, L. (2024). What AI evaluations for preventing catastrophic risks can and cannot do. In arXiv. https://arxiv.org/abs/2412.08653Barnett, P., and L. Thiergart. 2024. “What AI Evaluations for Preventing Catastrophic Risks Can and Cannot Do”. In arXiv. Preprint, November 26. https://arxiv.org/abs/2412.08653.Barnett, P., and L. Thiergart. “What AI Evaluations for Preventing Catastrophic Risks Can and Cannot Do”. arXiv, 26 Nov. 2024, https://arxiv.org/abs/2412.08653.Barnett, P. & Thiergart, L. What AI evaluations for preventing catastrophic risks can and cannot do. arXiv Preprint at https://arxiv.org/abs/2412.08653 (2024).P. Barnett and L. Thiergart, “What AI evaluations for preventing catastrophic risks can and cannot do”, Nov. 26, 2024. [Online]. Available: https://arxiv.org/abs/2412.08653
Baumann(2017). S-risks: An introduction. Center for Reducing Suffering.Baumann. (2017). S-risks: An introduction. Center for Reducing Suffering. https://centerforreducingsuffering.org/research/introBaumann. 2017. “S-risks: An Introduction”. Center for Reducing Suffering. https://centerforreducingsuffering.org/research/intro.Baumann. “S-risks: An Introduction”. Center for Reducing Suffering, 2017, https://centerforreducingsuffering.org/research/intro.Baumann. S-risks: An introduction. Center for Reducing Suffering https://centerforreducingsuffering.org/research/intro (2017).Baumann, “S-risks: An introduction”, Center for Reducing Suffering. [Online]. Available: https://centerforreducingsuffering.org/research/intro
BBC(2023). ChatGPT banned in Italy over privacy concerns.BBC. (2023). ChatGPT banned in Italy over privacy concerns. https://bbc.com/news/technology-65139406BBC. 2023. “ChatGPT Banned in Italy over Privacy Concerns”. https://bbc.com/news/technology-65139406.BBC. ChatGPT Banned in Italy over Privacy Concerns. 2023, https://bbc.com/news/technology-65139406.BBC. ChatGPT banned in Italy over privacy concerns. https://bbc.com/news/technology-65139406 (2023).BBC, “ChatGPT banned in Italy over privacy concerns”. [Online]. Available: https://bbc.com/news/technology-65139406
Belfield & Hua(2022). Compute and Antitrust. Verfassungsblog.Belfield & Hua. (2022, August 19). Compute and Antitrust. Verfassungsblog. https://verfassungsblog.de/compute-and-antitrustBelfield & Hua. 2022. “Compute and Antitrust”. Verfassungsblog, August 19. https://verfassungsblog.de/compute-and-antitrust.Belfield & Hua. “Compute and Antitrust”. Verfassungsblog, 19 Aug. 2022, https://verfassungsblog.de/compute-and-antitrust.Belfield & Hua. Compute and Antitrust. Verfassungsblog https://verfassungsblog.de/compute-and-antitrust (2022).Belfield & Hua, “Compute and Antitrust”, Verfassungsblog. [Online]. Available: https://verfassungsblog.de/compute-and-antitrust
Ben Pace(2020). What Failure Looks Like: Distilling the Discussion. AI Alignment Forum.Ben Pace. (2020, July 29). What Failure Looks Like: Distilling the Discussion. AI Alignment Forum. https://alignmentforum.org/posts/6jkGf5WEKMpMFXZp2/what-failure-looks-like-distilling-the-discussionBen Pace. 2020. “What Failure Looks Like: Distilling the Discussion”. AI Alignment Forum, July 29. https://alignmentforum.org/posts/6jkGf5WEKMpMFXZp2/what-failure-looks-like-distilling-the-discussion.Ben Pace. “What Failure Looks Like: Distilling the Discussion”. AI Alignment Forum, 29 July 2020, https://alignmentforum.org/posts/6jkGf5WEKMpMFXZp2/what-failure-looks-like-distilling-the-discussion.Ben Pace. What Failure Looks Like: Distilling the Discussion. AI Alignment Forum https://alignmentforum.org/posts/6jkGf5WEKMpMFXZp2/what-failure-looks-like-distilling-the-discussion (2020).Ben Pace, “What Failure Looks Like: Distilling the Discussion”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/6jkGf5WEKMpMFXZp2/what-failure-looks-like-distilling-the-discussion
Bengio(2023). Yoshua Bengio | FAQ on Catastrophic AI Risks.Bengio. (2023). Yoshua Bengio | FAQ on Catastrophic AI Risks. https://yoshuabengio.org/2023/06/24/faq-on-catastrophic-ai-risksBengio. 2023. “Yoshua Bengio | FAQ on Catastrophic AI Risks”. https://yoshuabengio.org/2023/06/24/faq-on-catastrophic-ai-risks.Bengio. Yoshua Bengio | FAQ on Catastrophic AI Risks. 2023, https://yoshuabengio.org/2023/06/24/faq-on-catastrophic-ai-risks.Bengio. Yoshua Bengio | FAQ on Catastrophic AI Risks. https://yoshuabengio.org/2023/06/24/faq-on-catastrophic-ai-risks (2023).Bengio, “Yoshua Bengio | FAQ on Catastrophic AI Risks”. [Online]. Available: https://yoshuabengio.org/2023/06/24/faq-on-catastrophic-ai-risks
Bengio, Y. et al.(2023). Managing extreme AI risks amid rapid progress. arXiv.Bengio, Y., Hinton, G., Yao, A., Song, D., Abbeel, P., Darrell, T., Harari, Y. N., Zhang, Y.-Q., Xue, L., Shalev-Shwartz, S., Hadfield, G., Clune, J., Maharaj, T., Hutter, F., Baydin, A. G., McIlraith, S., Gao, Q., Acharya, A., Krueger, D., … Mindermann, S. (2023). Managing extreme AI risks amid rapid progress. In arXiv. https://doi.org/10.1126/science.adn0117Bengio, Y., G. Hinton, A. Yao, et al. 2023. “Managing Extreme AI Risks Amid Rapid Progress”. In arXiv. Preprint, October 26. https://doi.org/10.1126/science.adn0117.Bengio, Y., et al. “Managing Extreme AI Risks Amid Rapid Progress”. arXiv, 26 Oct. 2023, https://doi.org/10.1126/science.adn0117.Bengio, Y. et al. Managing extreme AI risks amid rapid progress. arXiv Preprint at https://doi.org/10.1126/science.adn0117 (2023).Y. Bengio et al., “Managing extreme AI risks amid rapid progress”, Oct. 26, 2023. doi: 10.1126/science.adn0117.
Bengio, Y. et al.(2025). International AI Safety Report. arXiv.Bengio, Y., Mindermann, S., Privitera, D., Besiroglu, T., Bommasani, R., Casper, S., Choi, Y., Fox, P., Garfinkel, B., Goldfarb, D., Heidari, H., Ho, A., Kapoor, S., Khalatbari, L., Longpre, S., Manning, S., Mavroudis, V., Mazeika, M., Michael, J., … Zeng, Y. (2025). International AI Safety Report. In arXiv. https://arxiv.org/abs/2501.17805Bengio, Y., S. Mindermann, D. Privitera, et al. 2025. “International AI Safety Report”. In arXiv. Preprint, January 29. https://arxiv.org/abs/2501.17805.Bengio, Y., et al. “International AI Safety Report”. arXiv, 29 Jan. 2025, https://arxiv.org/abs/2501.17805.Bengio, Y. et al. International AI Safety Report. arXiv Preprint at https://arxiv.org/abs/2501.17805 (2025).Y. Bengio et al., “International AI Safety Report”, Jan. 29, 2025. [Online]. Available: https://arxiv.org/abs/2501.17805
Beraja, M., Kao, A., Yang, D. Y. & Yuchtman, N.(2023). AI-tocracy. The Quarterly Journal of Economics.Beraja, M., Kao, A., Yang, D. Y., & Yuchtman, N. (2023). AI-tocracy. The Quarterly Journal of Economics, 138(3), 1349–1402. https://doi.org/10.1093/qje/qjad012Beraja, M., A. Kao, D. Y. Yang, and N. Yuchtman. 2023. “AI-tocracy”. The Quarterly Journal of Economics 138 (3): 1349–1402. https://doi.org/10.1093/qje/qjad012.Beraja, M., et al. “AI-tocracy”. The Quarterly Journal of Economics, vol. 138, no. 3, Mar. 2023, pp. 1349–402, https://doi.org/10.1093/qje/qjad012.Beraja, M., Kao, A., Yang, D. Y. & Yuchtman, N. AI-tocracy. The Quarterly Journal of Economics 138, 1349–1402 (2023).M. Beraja, A. Kao, D. Y. Yang, and N. Yuchtman, “AI-tocracy”, The Quarterly Journal of Economics, vol. 138, no. 3, pp. 1349–1402, Mar. 2023, doi: 10.1093/qje/qjad012.
beren(2023). Gradient hacking is extremely difficult. AI Alignment Forum.beren. (2023, January 24). Gradient hacking is extremely difficult. AI Alignment Forum. https://alignmentforum.org/posts/w2TAEvME2yAG9MHeq/gradient-hacking-is-extremely-difficultberen. 2023. “Gradient Hacking Is Extremely Difficult”. AI Alignment Forum, January 24. https://alignmentforum.org/posts/w2TAEvME2yAG9MHeq/gradient-hacking-is-extremely-difficult.beren. “Gradient Hacking Is Extremely Difficult”. AI Alignment Forum, 24 Jan. 2023, https://alignmentforum.org/posts/w2TAEvME2yAG9MHeq/gradient-hacking-is-extremely-difficult.beren. Gradient hacking is extremely difficult. AI Alignment Forum https://alignmentforum.org/posts/w2TAEvME2yAG9MHeq/gradient-hacking-is-extremely-difficult (2023).beren, “Gradient hacking is extremely difficult”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/w2TAEvME2yAG9MHeq/gradient-hacking-is-extremely-difficult
Berglund, L. et al.(2023). Taken out of context: On measuring situational awareness in LLMs. arXiv.Berglund, L., Stickland, A. C., Balesni, M., Kaufmann, M., Tong, M., Korbak, T., Kokotajlo, D., & Evans, O. (2023). Taken out of context: On measuring situational awareness in LLMs. In arXiv. https://arxiv.org/abs/2309.00667Berglund, L., A. C. Stickland, M. Balesni, et al. 2023. “Taken Out of Context: On Measuring Situational Awareness in LLMs”. In arXiv. Preprint, September 1. https://arxiv.org/abs/2309.00667.Berglund, L., et al. “Taken Out of Context: On Measuring Situational Awareness in LLMs”. arXiv, 1 Sept. 2023, https://arxiv.org/abs/2309.00667.Berglund, L. et al. Taken out of context: On measuring situational awareness in LLMs. arXiv Preprint at https://arxiv.org/abs/2309.00667 (2023).L. Berglund et al., “Taken out of context: On measuring situational awareness in LLMs”, Sep. 01, 2023. [Online]. Available: https://arxiv.org/abs/2309.00667
Berglund, L. et al.(2023). The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A". arXiv.Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., & Evans, O. (2023). The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A". In arXiv. https://arxiv.org/abs/2309.12288Berglund, L., M. Tong, M. Kaufmann, et al. 2023. “The Reversal Curse: LLMs Trained on "A Is B" Fail to Learn "B Is A"”. In arXiv. Preprint, September 21. https://arxiv.org/abs/2309.12288.Berglund, L., et al. “The Reversal Curse: LLMs Trained on "A Is B" Fail to Learn "B Is A"”. arXiv, 21 Sept. 2023, https://arxiv.org/abs/2309.12288.Berglund, L. et al. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A". arXiv Preprint at https://arxiv.org/abs/2309.12288 (2023).L. Berglund et al., “The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"”, Sep. 21, 2023. [Online]. Available: https://arxiv.org/abs/2309.12288
Bernardi(2024). A Policy Agenda for Defensive Acceleration Against AI Risks.Bernardi. (2024). A Policy Agenda for Defensive Acceleration Against AI Risks. https://airesilience.substack.com/p/a-policy-agenda-for-defensive-accelerationBernardi. 2024. A Policy Agenda for Defensive Acceleration Against AI Risks. Edition. https://airesilience.substack.com/p/a-policy-agenda-for-defensive-acceleration.Bernardi. A Policy Agenda for Defensive Acceleration Against AI Risks. 2024, https://airesilience.substack.com/p/a-policy-agenda-for-defensive-acceleration.Bernardi. A Policy Agenda for Defensive Acceleration Against AI Risks. https://airesilience.substack.com/p/a-policy-agenda-for-defensive-acceleration (2024).Bernardi, “A Policy Agenda for Defensive Acceleration Against AI Risks”. [Online]. Available: https://airesilience.substack.com/p/a-policy-agenda-for-defensive-acceleration
Besiroglu et al.(2024). FrontierMath: Evaluating advanced mathematical reasoning in AI. Epoch AI.Besiroglu et al. (2024). FrontierMath: Evaluating advanced mathematical reasoning in AI. Epoch AI. https://epoch.ai/frontiermath/the-benchmarkBesiroglu et al. 2024. “FrontierMath: Evaluating Advanced Mathematical Reasoning in AI”. Epoch AI. https://epoch.ai/frontiermath/the-benchmark.Besiroglu et al. “FrontierMath: Evaluating Advanced Mathematical Reasoning in AI”. Epoch AI, 2024, https://epoch.ai/frontiermath/the-benchmark.Besiroglu et al. FrontierMath: Evaluating advanced mathematical reasoning in AI. Epoch AI https://epoch.ai/frontiermath/the-benchmark (2024).Besiroglu et al., “FrontierMath: Evaluating advanced mathematical reasoning in AI”, Epoch AI. [Online]. Available: https://epoch.ai/frontiermath/the-benchmark
Besiroglu, T., Bergerson, S. A., Michael, A., Heim, L., Luo, X. & Thompson, N.(2024). The Compute Divide in Machine Learning: A Threat to Academic Contribution and Scrutiny?. arXiv.Besiroglu, T., Bergerson, S. A., Michael, A., Heim, L., Luo, X., & Thompson, N. (2024). The Compute Divide in Machine Learning: A Threat to Academic Contribution and Scrutiny?. In arXiv. https://arxiv.org/abs/2401.02452Besiroglu, T., S. A. Bergerson, A. Michael, L. Heim, X. Luo, and N. Thompson. 2024. “The Compute Divide in Machine Learning: A Threat to Academic Contribution and Scrutiny?”. In arXiv. Preprint, January 4. https://arxiv.org/abs/2401.02452.Besiroglu, T., et al. “The Compute Divide in Machine Learning: A Threat to Academic Contribution and Scrutiny?”. arXiv, 4 Jan. 2024, https://arxiv.org/abs/2401.02452.Besiroglu, T. et al. The Compute Divide in Machine Learning: A Threat to Academic Contribution and Scrutiny?. arXiv Preprint at https://arxiv.org/abs/2401.02452 (2024).T. Besiroglu, S. A. Bergerson, A. Michael, L. Heim, X. Luo, and N. Thompson, “The Compute Divide in Machine Learning: A Threat to Academic Contribution and Scrutiny?”, Jan. 04, 2024. [Online]. Available: https://arxiv.org/abs/2401.02452
Besta, M. et al.(2023). Graph of Thoughts: Solving Elaborate Problems with Large Language Models. arXiv.Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Gajda, J., Lehmann, T., Niewiadomski, H., Nyczyk, P., & Hoefler, T. (2023). Graph of Thoughts: Solving Elaborate Problems with Large Language Models. In arXiv. https://doi.org/10.1609/aaai.v38i16.29720Besta, M., N. Blach, A. Kubicek, et al. 2023. “Graph of Thoughts: Solving Elaborate Problems with Large Language Models”. In arXiv. Preprint, August 18. https://doi.org/10.1609/aaai.v38i16.29720.Besta, M., et al. “Graph of Thoughts: Solving Elaborate Problems with Large Language Models”. arXiv, 18 Aug. 2023, https://doi.org/10.1609/aaai.v38i16.29720.Besta, M. et al. Graph of Thoughts: Solving Elaborate Problems with Large Language Models. arXiv Preprint at https://doi.org/10.1609/aaai.v38i16.29720 (2023).M. Besta et al., “Graph of Thoughts: Solving Elaborate Problems with Large Language Models”, Aug. 18, 2023. doi: 10.1609/aaai.v38i16.29720.
Beth Barnes & paulfchristiano(2020). Writeup: Progress on AI Safety via Debate. LessWrong.Beth Barnes, & paulfchristiano. (2020, February 5). Writeup: Progress on AI Safety via Debate. LessWrong. https://lesswrong.com/posts/Br4xDbYu4Frwrb64a/writeup-progress-on-ai-safety-via-debate-1Beth Barnes, and paulfchristiano. 2020. “Writeup: Progress on AI Safety via Debate”. LessWrong, February 5. https://lesswrong.com/posts/Br4xDbYu4Frwrb64a/writeup-progress-on-ai-safety-via-debate-1.Beth Barnes, and paulfchristiano. “Writeup: Progress on AI Safety via Debate”. LessWrong, 5 Feb. 2020, https://lesswrong.com/posts/Br4xDbYu4Frwrb64a/writeup-progress-on-ai-safety-via-debate-1.Beth Barnes & paulfchristiano. Writeup: Progress on AI Safety via Debate. LessWrong https://lesswrong.com/posts/Br4xDbYu4Frwrb64a/writeup-progress-on-ai-safety-via-debate-1 (2020).Beth Barnes and paulfchristiano, “Writeup: Progress on AI Safety via Debate”, LessWrong. [Online]. Available: https://lesswrong.com/posts/Br4xDbYu4Frwrb64a/writeup-progress-on-ai-safety-via-debate-1
Beth Barnes(2022). 'simulator' framing and confusions about LLMs. AI Alignment Forum.Beth Barnes. (2022, December 31). 'simulator' framing and confusions about LLMs. AI Alignment Forum. https://alignmentforum.org/posts/dYnHLWMXCYdm9xu5j/simulator-framing-and-confusions-about-llmsBeth Barnes. 2022. “'simulator' Framing and Confusions About LLMs”. AI Alignment Forum, December 31. https://alignmentforum.org/posts/dYnHLWMXCYdm9xu5j/simulator-framing-and-confusions-about-llms.Beth Barnes. “'simulator' Framing and Confusions About LLMs”. AI Alignment Forum, 31 Dec. 2022, https://alignmentforum.org/posts/dYnHLWMXCYdm9xu5j/simulator-framing-and-confusions-about-llms.Beth Barnes. 'simulator' framing and confusions about LLMs. AI Alignment Forum https://alignmentforum.org/posts/dYnHLWMXCYdm9xu5j/simulator-framing-and-confusions-about-llms (2022).Beth Barnes, “'simulator' framing and confusions about LLMs”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/dYnHLWMXCYdm9xu5j/simulator-framing-and-confusions-about-llms
Beth Barnes(2023). New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks.Beth Barnes. (2023, July 31). New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks. https://metr.org/blog/2023-08-01-new-reportBeth Barnes. 2023. “New Report: Evaluating Language-Model Agents on Realistic Autonomous Tasks”. July 31. https://metr.org/blog/2023-08-01-new-report.Beth Barnes. New Report: Evaluating Language-Model Agents on Realistic Autonomous Tasks. 31 July 2023, https://metr.org/blog/2023-08-01-new-report.Beth Barnes. New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks. https://metr.org/blog/2023-08-01-new-report (2023).Beth Barnes, “New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks”. [Online]. Available: https://metr.org/blog/2023-08-01-new-report
Beth Barnes(2023). Update on ARC's recent eval efforts.Beth Barnes. (2023, March 17). Update on ARC's recent eval efforts. https://metr.org/blog/2023-03-18-update-on-recent-evalsBeth Barnes. 2023. “Update on ARC's Recent Eval Efforts”. March 17. https://metr.org/blog/2023-03-18-update-on-recent-evals.Beth Barnes. Update on ARC's Recent Eval Efforts. 17 Mar. 2023, https://metr.org/blog/2023-03-18-update-on-recent-evals.Beth Barnes. Update on ARC's recent eval efforts. https://metr.org/blog/2023-03-18-update-on-recent-evals (2023).Beth Barnes, “Update on ARC's recent eval efforts”. [Online]. Available: https://metr.org/blog/2023-03-18-update-on-recent-evals
Beth Barnes, Hjalmar Wijk & Lawrence Chan(2023). Responsible Scaling Policies (RSPs).Beth Barnes, Hjalmar Wijk, & Lawrence Chan. (2023, September 26). Responsible Scaling Policies (RSPs). https://metr.org/blog/2023-09-26-rspBeth Barnes, Hjalmar Wijk, and Lawrence Chan. 2023. “Responsible Scaling Policies (RSPs)”. September 26. https://metr.org/blog/2023-09-26-rsp.Beth Barnes, et al. Responsible Scaling Policies (RSPs). 26 Sept. 2023, https://metr.org/blog/2023-09-26-rsp.Beth Barnes, Hjalmar Wijk & Lawrence Chan. Responsible Scaling Policies (RSPs). https://metr.org/blog/2023-09-26-rsp (2023).Beth Barnes, Hjalmar Wijk, and Lawrence Chan, “Responsible Scaling Policies (RSPs)”. [Online]. Available: https://metr.org/blog/2023-09-26-rsp
Betker, J. et al.(2023). Improving Image Generation with Better Captions.Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., Manassra, W., Dhariwal, P., Chu, C., Jiao, Y., & Ramesh, A. (2023). Improving Image Generation with Better Captions. OpenAI. https://cdn.openai.com/papers/dall-e-3.pdfBetker, J., G. Goh, L. Jing, et al. 2023. Improving Image Generation with Better Captions. OpenAI. https://cdn.openai.com/papers/dall-e-3.pdf.Betker, J., et al. Improving Image Generation with Better Captions. OpenAI, 2023, https://cdn.openai.com/papers/dall-e-3.pdf.Betker, J. et al. Improving Image Generation with Better Captions. https://cdn.openai.com/papers/dall-e-3.pdf (2023).J. Betker et al., “Improving Image Generation with Better Captions”, OpenAI, 2023. [Online]. Available: https://cdn.openai.com/papers/dall-e-3.pdf
Betley, J. et al.(2025). Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs. arXiv.Betley, J., Tan, D., Warncke, N., Sztyber-Betley, A., Bao, X., Soto, M., Labenz, N., & Evans, O. (2025). Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs. In arXiv. https://doi.org/10.1038/s41586-025-09937-5Betley, J., D. Tan, N. Warncke, et al. 2025. “Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs”. In arXiv. Preprint, February 24. https://doi.org/10.1038/s41586-025-09937-5.Betley, J., et al. “Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs”. arXiv, 24 Feb. 2025, https://doi.org/10.1038/s41586-025-09937-5.Betley, J. et al. Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs. arXiv Preprint at https://doi.org/10.1038/s41586-025-09937-5 (2025).J. Betley et al., “Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs”, Feb. 24, 2025. doi: 10.1038/s41586-025-09937-5.
Bhatt et al.(2023). Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models. arXiv.org.Bhatt et al. (2023). Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models. arXiv.org. https://www.arxiv.org/abs/2312.04724Bhatt et al. 2023. “Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models”. arXiv.org. https://www.arxiv.org/abs/2312.04724.Bhatt et al. “Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models”. arXiv.org, 2023, https://www.arxiv.org/abs/2312.04724.Bhatt et al. Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models. arXiv.org https://www.arxiv.org/abs/2312.04724 (2023).Bhatt et al., “Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2312.04724
Bhatt et al.(2024). Shell Games: Control Protocols for Adversarial AI Agents. OpenReview.Bhatt et al. (2024). Shell Games: Control Protocols for Adversarial AI Agents. In OpenReview. https://openreview.net/forum?id=oycEeFXX74Bhatt et al. 2024. “Shell Games: Control Protocols for Adversarial AI Agents”. In OpenReview. Preprint. https://openreview.net/forum?id=oycEeFXX74.Bhatt et al. “Shell Games: Control Protocols for Adversarial AI Agents”. OpenReview, 2024, https://openreview.net/forum?id=oycEeFXX74.Bhatt et al. Shell Games: Control Protocols for Adversarial AI Agents. OpenReview Preprint at https://openreview.net/forum?id=oycEeFXX74 (2024).Bhatt et al., “Shell Games: Control Protocols for Adversarial AI Agents”, 2024. [Online]. Available: https://openreview.net/forum?id=oycEeFXX74
Bhatt, M. et al.(2024). CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models. arXiv.Bhatt, M., Chennabasappa, S., Li, Y., Nikolaidis, C., Song, D., Wan, S., Ahmad, F., Aschermann, C., Chen, Y., Kapil, D., Molnar, D., Whitman, S., & Saxe, J. (2024). CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models. In arXiv. https://arxiv.org/abs/2404.13161Bhatt, M., S. Chennabasappa, Y. Li, et al. 2024. “CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models”. In arXiv. Preprint, April 19. https://arxiv.org/abs/2404.13161.Bhatt, M., et al. “CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models”. arXiv, 19 Apr. 2024, https://arxiv.org/abs/2404.13161.Bhatt, M. et al. CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models. arXiv Preprint at https://arxiv.org/abs/2404.13161 (2024).M. Bhatt et al., “CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models”, Apr. 19, 2024. [Online]. Available: https://arxiv.org/abs/2404.13161
Binder, F. J. et al.(2024). Looking Inward: Language Models Can Learn About Themselves by Introspection. arXiv.Binder, F. J., Chua, J., Korbak, T., Sleight, H., Hughes, J., Long, R., Perez, E., Turpin, M., & Evans, O. (2024). Looking Inward: Language Models Can Learn About Themselves by Introspection. In arXiv. https://arxiv.org/abs/2410.13787Binder, F. J., J. Chua, T. Korbak, et al. 2024. “Looking Inward: Language Models Can Learn About Themselves by Introspection”. In arXiv. Preprint, October 17. https://arxiv.org/abs/2410.13787.Binder, F. J., et al. “Looking Inward: Language Models Can Learn About Themselves by Introspection”. arXiv, 17 Oct. 2024, https://arxiv.org/abs/2410.13787.Binder, F. J. et al. Looking Inward: Language Models Can Learn About Themselves by Introspection. arXiv Preprint at https://arxiv.org/abs/2410.13787 (2024).F. J. Binder et al., “Looking Inward: Language Models Can Learn About Themselves by Introspection”, Oct. 17, 2024. [Online]. Available: https://arxiv.org/abs/2410.13787
Black, S. et al.(2025). RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents. arXiv.Black, S., Stickland, A. C., Pencharz, J., Sourbut, O., Schmatz, M., Bailey, J., Matthews, O., Millwood, B., Remedios, A., & Cooney, A. (2025). RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents. In arXiv. https://arxiv.org/abs/2504.18565Black, S., A. C. Stickland, J. Pencharz, et al. 2025. “RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents”. In arXiv. Preprint, April 21. https://arxiv.org/abs/2504.18565.Black, S., et al. “RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents”. arXiv, 21 Apr. 2025, https://arxiv.org/abs/2504.18565.Black, S. et al. RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents. arXiv Preprint at https://arxiv.org/abs/2504.18565 (2025).S. Black et al., “RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents”, Apr. 21, 2025. [Online]. Available: https://arxiv.org/abs/2504.18565
Bode, I. & Watts, T.(2023). Loitering Munitions and Unpredictability: Autonomy in Weapon Systems and Challenges to Human Control.Bode, I., & Watts, T. (2023). Loitering Munitions and Unpredictability: Autonomy in Weapon Systems and Challenges to Human Control. Center for War Studies, University of Southern Denmark. https://findresearcher.sdu.dk/ws/portalfiles/portal/231643063/Loitering_Munitions_Unpredictability_WEB.pdfBode, I., and T. Watts. 2023. Loitering Munitions and Unpredictability: Autonomy in Weapon Systems and Challenges to Human Control. Center for War Studies, University of Southern Denmark. https://findresearcher.sdu.dk/ws/portalfiles/portal/231643063/Loitering_Munitions_Unpredictability_WEB.pdf.Bode, I., and T. Watts. Loitering Munitions and Unpredictability: Autonomy in Weapon Systems and Challenges to Human Control. Center for War Studies, University of Southern Denmark, May 2023, https://findresearcher.sdu.dk/ws/portalfiles/portal/231643063/Loitering_Munitions_Unpredictability_WEB.pdf.Bode, I. & Watts, T. Loitering Munitions and Unpredictability: Autonomy in Weapon Systems and Challenges to Human Control. https://findresearcher.sdu.dk/ws/portalfiles/portal/231643063/Loitering_Munitions_Unpredictability_WEB.pdf (2023).I. Bode and T. Watts, “Loitering Munitions and Unpredictability: Autonomy in Weapon Systems and Challenges to Human Control”, Center for War Studies, University of Southern Denmark, Odense, May 2023. [Online]. Available: https://findresearcher.sdu.dk/ws/portalfiles/portal/231643063/Loitering_Munitions_Unpredictability_WEB.pdf
Bogdan Ionut Cirstea(2023). AISC project: How promising is automating alignment research? (literature review). LessWrong.Bogdan Ionut Cirstea. (2023, November 28). AISC project: How promising is automating alignment research? (literature review). LessWrong. https://lesswrong.com/posts/FtHidqjAFTerfMZLo/aisc-project-how-promising-is-automating-alignment-researchBogdan Ionut Cirstea. 2023. “AISC Project: How Promising Is Automating Alignment Research? (literature Review)”. LessWrong, November 28. https://lesswrong.com/posts/FtHidqjAFTerfMZLo/aisc-project-how-promising-is-automating-alignment-research.Bogdan Ionut Cirstea. “AISC Project: How Promising Is Automating Alignment Research? (literature Review)”. LessWrong, 28 Nov. 2023, https://lesswrong.com/posts/FtHidqjAFTerfMZLo/aisc-project-how-promising-is-automating-alignment-research.Bogdan Ionut Cirstea. AISC project: How promising is automating alignment research? (literature review). LessWrong https://lesswrong.com/posts/FtHidqjAFTerfMZLo/aisc-project-how-promising-is-automating-alignment-research (2023).Bogdan Ionut Cirstea, “AISC project: How promising is automating alignment research? (literature review)”, LessWrong. [Online]. Available: https://lesswrong.com/posts/FtHidqjAFTerfMZLo/aisc-project-how-promising-is-automating-alignment-research
Bogdan, P. C., Macar, U., Nanda, N. & Conmy, A.(2025). Thought Anchors: Which LLM Reasoning Steps Matter?. arXiv.Bogdan, P. C., Macar, U., Nanda, N., & Conmy, A. (2025). Thought Anchors: Which LLM Reasoning Steps Matter?. In arXiv. https://arxiv.org/abs/2506.19143Bogdan, P. C., U. Macar, N. Nanda, and A. Conmy. 2025. “Thought Anchors: Which LLM Reasoning Steps Matter?”. In arXiv. Preprint, June 23. https://arxiv.org/abs/2506.19143.Bogdan, P. C., et al. “Thought Anchors: Which LLM Reasoning Steps Matter?”. arXiv, 23 June 2025, https://arxiv.org/abs/2506.19143.Bogdan, P. C., Macar, U., Nanda, N. & Conmy, A. Thought Anchors: Which LLM Reasoning Steps Matter?. arXiv Preprint at https://arxiv.org/abs/2506.19143 (2025).P. C. Bogdan, U. Macar, N. Nanda, and A. Conmy, “Thought Anchors: Which LLM Reasoning Steps Matter?”, Jun. 23, 2025. [Online]. Available: https://arxiv.org/abs/2506.19143
Boiko, D. A., MacKnight, R. & Gomes, G.(2023). Emergent autonomous scientific research capabilities of large language models. arXiv.Boiko, D. A., MacKnight, R., & Gomes, G. (2023). Emergent autonomous scientific research capabilities of large language models. In arXiv. https://arxiv.org/abs/2304.05332Boiko, D. A., R. MacKnight, and G. Gomes. 2023. “Emergent Autonomous Scientific Research Capabilities of Large Language Models”. In arXiv. Preprint, April 11. https://arxiv.org/abs/2304.05332.Boiko, D. A., et al. “Emergent Autonomous Scientific Research Capabilities of Large Language Models”. arXiv, 11 Apr. 2023, https://arxiv.org/abs/2304.05332.Boiko, D. A., MacKnight, R. & Gomes, G. Emergent autonomous scientific research capabilities of large language models. arXiv Preprint at https://arxiv.org/abs/2304.05332 (2023).D. A. Boiko, R. MacKnight, and G. Gomes, “Emergent autonomous scientific research capabilities of large language models”, Apr. 11, 2023. [Online]. Available: https://arxiv.org/abs/2304.05332
Bommasani, R. et al.(2021). On the Opportunities and Risks of Foundation Models. arXiv.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., … Liang, P. (2021). On the Opportunities and Risks of Foundation Models. In arXiv. https://arxiv.org/abs/2108.07258Bommasani, R., D. A. Hudson, E. Adeli, et al. 2021. “On the Opportunities and Risks of Foundation Models”. In arXiv. Preprint, August 16. https://arxiv.org/abs/2108.07258.Bommasani, R., et al. “On the Opportunities and Risks of Foundation Models”. arXiv, 16 Aug. 2021, https://arxiv.org/abs/2108.07258.Bommasani, R. et al. On the Opportunities and Risks of Foundation Models. arXiv Preprint at https://arxiv.org/abs/2108.07258 (2021).R. Bommasani et al., “On the Opportunities and Risks of Foundation Models”, Aug. 16, 2021. [Online]. Available: https://arxiv.org/abs/2108.07258
Bondarenko, A., Volk, D., Volkov, D. & Ladish, J.(2025). Demonstrating specification gaming in reasoning models. arXiv.Bondarenko, A., Volk, D., Volkov, D., & Ladish, J. (2025). Demonstrating specification gaming in reasoning models. In arXiv. https://arxiv.org/abs/2502.13295Bondarenko, A., D. Volk, D. Volkov, and J. Ladish. 2025. “Demonstrating Specification Gaming in Reasoning Models”. In arXiv. Preprint, February 18. https://arxiv.org/abs/2502.13295.Bondarenko, A., et al. “Demonstrating Specification Gaming in Reasoning Models”. arXiv, 18 Feb. 2025, https://arxiv.org/abs/2502.13295.Bondarenko, A., Volk, D., Volkov, D. & Ladish, J. Demonstrating specification gaming in reasoning models. arXiv Preprint at https://arxiv.org/abs/2502.13295 (2025).A. Bondarenko, D. Volk, D. Volkov, and J. Ladish, “Demonstrating specification gaming in reasoning models”, Feb. 18, 2025. [Online]. Available: https://arxiv.org/abs/2502.13295
Boston Dynamics(2024). Atlas Goes Hands On. Boston Dynamics.Boston Dynamics. (2024). Atlas Goes Hands On. Boston Dynamics. https://bostondynamics.com/video/atlas-goes-hands-onBoston Dynamics. 2024. “Atlas Goes Hands On”. Boston Dynamics. https://bostondynamics.com/video/atlas-goes-hands-on.Boston Dynamics. “Atlas Goes Hands On”. Boston Dynamics, 2024, https://bostondynamics.com/video/atlas-goes-hands-on.Boston Dynamics. Atlas Goes Hands On. Boston Dynamics https://bostondynamics.com/video/atlas-goes-hands-on (2024).Boston Dynamics, “Atlas Goes Hands On”, Boston Dynamics. [Online]. Available: https://bostondynamics.com/video/atlas-goes-hands-on
Boston Dynamics(2024). Stretch - Mobile Warehouse Robots. Boston Dynamics.Boston Dynamics. (2024). Stretch - Mobile Warehouse Robots. Boston Dynamics. https://bostondynamics.com/products/stretchBoston Dynamics. 2024. “Stretch - Mobile Warehouse Robots”. Boston Dynamics. https://bostondynamics.com/products/stretch.Boston Dynamics. “Stretch - Mobile Warehouse Robots”. Boston Dynamics, 2024, https://bostondynamics.com/products/stretch.Boston Dynamics. Stretch - Mobile Warehouse Robots. Boston Dynamics https://bostondynamics.com/products/stretch (2024).Boston Dynamics, “Stretch - Mobile Warehouse Robots”, Boston Dynamics. [Online]. Available: https://bostondynamics.com/products/stretch
Bostrom(2002). Existential Risks: Analyzing Human Extinction Scenarios and Related Hazards.Bostrom. (2002). Existential Risks: Analyzing Human Extinction Scenarios and Related Hazards. https://nickbostrom.com/existential/risksBostrom. 2002. “Existential Risks: Analyzing Human Extinction Scenarios and Related Hazards”. https://nickbostrom.com/existential/risks.Bostrom. Existential Risks: Analyzing Human Extinction Scenarios and Related Hazards. 2002, https://nickbostrom.com/existential/risks.Bostrom. Existential Risks: Analyzing Human Extinction Scenarios and Related Hazards. https://nickbostrom.com/existential/risks (2002).Bostrom, “Existential Risks: Analyzing Human Extinction Scenarios and Related Hazards”. [Online]. Available: https://nickbostrom.com/existential/risks
Bostrom(2012). Existential Risks: Threats to Humanity’s Survival.Bostrom. (2012). Existential Risks: Threats to Humanity’s Survival. https://existential-risk.com/conceptBostrom. 2012. “Existential Risks: Threats to Humanity’s Survival”. https://existential-risk.com/concept.Bostrom. Existential Risks: Threats to Humanity’s Survival. 2012, https://existential-risk.com/concept.Bostrom. Existential Risks: Threats to Humanity’s Survival. https://existential-risk.com/concept (2012).Bostrom, “Existential Risks: Threats to Humanity’s Survival”. [Online]. Available: https://existential-risk.com/concept
Bostrom, N.(2014). Superintelligence: Paths, Dangers, Strategies.Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press. Internet Archive (https://web.archive.org/web/20260916233115/https://psycnet.apa.org/record/2014-48585-000). https://psycnet.apa.org/record/2014-48585-000Bostrom, N. 2014. Superintelligence: Paths, Dangers, Strategies. Oxford University Press. Https://web.archive.org/web/20260916233115/https://psycnet.apa.org/record/2014-48585-000. Internet Archive. https://psycnet.apa.org/record/2014-48585-000.Bostrom, N. Superintelligence: Paths, Dangers, Strategies. Oxford University Press, 2014, Internet Archive, https://web.archive.org/web/20260916233115/https://psycnet.apa.org/record/2014-48585-000, https://psycnet.apa.org/record/2014-48585-000.Bostrom, N. Superintelligence: Paths, Dangers, Strategies. (Oxford University Press, Oxford, 2014).N. Bostrom, Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press, 2014. Accessed: Sep. 16, 2026. [Online]. Available: https://psycnet.apa.org/record/2014-48585-000
Boursier, E. & Flammarion, N.(2024). Simplicity bias and optimization threshold in two-layer ReLU networks. arXiv.Boursier, E., & Flammarion, N. (2024). Simplicity bias and optimization threshold in two-layer ReLU networks. In arXiv. https://arxiv.org/abs/2410.02348Boursier, E., and N. Flammarion. 2024. “Simplicity Bias and Optimization Threshold in Two-layer ReLU Networks”. In arXiv. Preprint, October 3. https://arxiv.org/abs/2410.02348.Boursier, E., and N. Flammarion. “Simplicity Bias and Optimization Threshold in Two-layer ReLU Networks”. arXiv, 3 Oct. 2024, https://arxiv.org/abs/2410.02348.Boursier, E. & Flammarion, N. Simplicity bias and optimization threshold in two-layer ReLU networks. arXiv Preprint at https://arxiv.org/abs/2410.02348 (2024).E. Boursier and N. Flammarion, “Simplicity bias and optimization threshold in two-layer ReLU networks”, Oct. 03, 2024. [Online]. Available: https://arxiv.org/abs/2410.02348
Bowman(2024). [External] 2024 Debate Agenda Writeup. Google Docs.Bowman. (2024). [External] 2024 Debate Agenda Writeup. Google Docs. https://docs.google.com/document/d/1E2O7MSVI8u9LHbezdTgYoC6ULZmgki7EIm2XwCT0nuU/edit?tab=t.0Bowman. 2024. “[External] 2024 Debate Agenda Writeup”. Google Docs. https://docs.google.com/document/d/1E2O7MSVI8u9LHbezdTgYoC6ULZmgki7EIm2XwCT0nuU/edit?tab=t.0.Bowman. “[External] 2024 Debate Agenda Writeup”. Google Docs, 2024, https://docs.google.com/document/d/1E2O7MSVI8u9LHbezdTgYoC6ULZmgki7EIm2XwCT0nuU/edit?tab=t.0.Bowman. [External] 2024 Debate Agenda Writeup. Google Docs https://docs.google.com/document/d/1E2O7MSVI8u9LHbezdTgYoC6ULZmgki7EIm2XwCT0nuU/edit?tab=t.0 (2024).Bowman, “[External] 2024 Debate Agenda Writeup”, Google Docs. [Online]. Available: https://docs.google.com/document/d/1E2O7MSVI8u9LHbezdTgYoC6ULZmgki7EIm2XwCT0nuU/edit?tab=t.0
Bowman, S. R.(2023). Eight Things to Know about Large Language Models. arXiv.Bowman, S. R. (2023). Eight Things to Know about Large Language Models. In arXiv. https://arxiv.org/abs/2304.00612Bowman, S. R. 2023. “Eight Things to Know About Large Language Models”. In arXiv. Preprint, April 2. https://arxiv.org/abs/2304.00612.Bowman, S. R. “Eight Things to Know About Large Language Models”. arXiv, 2 Apr. 2023, https://arxiv.org/abs/2304.00612.Bowman, S. R. Eight Things to Know about Large Language Models. arXiv Preprint at https://arxiv.org/abs/2304.00612 (2023).S. R. Bowman, “Eight Things to Know about Large Language Models”, Apr. 02, 2023. [Online]. Available: https://arxiv.org/abs/2304.00612
Bowman, S. R. et al.(2022). Measuring Progress on Scalable Oversight for Large Language Models. arXiv.Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., Lukošiūtė, K., Askell, A., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Olah, C., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., … Kaplan, J. (2022). Measuring Progress on Scalable Oversight for Large Language Models. In arXiv. https://arxiv.org/abs/2211.03540Bowman, S. R., J. Hyun, E. Perez, et al. 2022. “Measuring Progress on Scalable Oversight for Large Language Models”. In arXiv. Preprint, November 4. https://arxiv.org/abs/2211.03540.Bowman, S. R., et al. “Measuring Progress on Scalable Oversight for Large Language Models”. arXiv, 4 Nov. 2022, https://arxiv.org/abs/2211.03540.Bowman, S. R. et al. Measuring Progress on Scalable Oversight for Large Language Models. arXiv Preprint at https://arxiv.org/abs/2211.03540 (2022).S. R. Bowman et al., “Measuring Progress on Scalable Oversight for Large Language Models”, Nov. 04, 2022. [Online]. Available: https://arxiv.org/abs/2211.03540
Bradford(2020). The Brussels Effect: How the European Union Rules the World. Scholarship Archive.Bradford. (2020). The Brussels Effect: How the European Union Rules the World. Scholarship Archive. https://scholarship.law.columbia.edu/books/232Bradford. 2020. “The Brussels Effect: How the European Union Rules the World”. Scholarship Archive. https://scholarship.law.columbia.edu/books/232.Bradford. “The Brussels Effect: How the European Union Rules the World”. Scholarship Archive, 2020, https://scholarship.law.columbia.edu/books/232.Bradford. The Brussels Effect: How the European Union Rules the World. Scholarship Archive https://scholarship.law.columbia.edu/books/232 (2020).Bradford, “The Brussels Effect: How the European Union Rules the World”, Scholarship Archive. [Online]. Available: https://scholarship.law.columbia.edu/books/232
Branwen(2020). The Scaling Hypothesis.Branwen. (2020). The Scaling Hypothesis. https://gwern.net/scaling-hypothesisBranwen. 2020. “The Scaling Hypothesis”. https://gwern.net/scaling-hypothesis.Branwen. The Scaling Hypothesis. 2020, https://gwern.net/scaling-hypothesis.Branwen. The Scaling Hypothesis. https://gwern.net/scaling-hypothesis (2020).Branwen, “The Scaling Hypothesis”. [Online]. Available: https://gwern.net/scaling-hypothesis
Brennan et al.(2025). Artificial Power: 2025 Landscape Report. AI Now Institute.Brennan et al. (2025, June 3). Artificial Power: 2025 Landscape Report. AI Now Institute. https://ainowinstitute.org/publications/research/ai-now-2025-landscape-reportBrennan et al. 2025. “Artificial Power: 2025 Landscape Report”. AI Now Institute, June 3. https://ainowinstitute.org/publications/research/ai-now-2025-landscape-report.Brennan et al. “Artificial Power: 2025 Landscape Report”. AI Now Institute, 3 June 2025, https://ainowinstitute.org/publications/research/ai-now-2025-landscape-report.Brennan et al. Artificial Power: 2025 Landscape Report. AI Now Institute https://ainowinstitute.org/publications/research/ai-now-2025-landscape-report (2025).Brennan et al., “Artificial Power: 2025 Landscape Report”, AI Now Institute. [Online]. Available: https://ainowinstitute.org/publications/research/ai-now-2025-landscape-report
Brown(2024). Noam Brown (@polynoamial) on X. X (formerly Twitter).Brown. (2024, December 20). Noam Brown (@polynoamial) on X. X (formerly Twitter). https://x.com/polynoamial/status/1870172996650053653?mx=2Brown. 2024. “Noam Brown (@polynoamial) on X”. X (formerly Twitter), December 20. https://x.com/polynoamial/status/1870172996650053653?mx=2.Brown. “Noam Brown (@polynoamial) on X”. X (formerly Twitter), 20 Dec. 2024, https://x.com/polynoamial/status/1870172996650053653?mx=2.Brown. Noam Brown (@polynoamial) on X. X (formerly Twitter) https://x.com/polynoamial/status/1870172996650053653?mx=2 (2024).Brown, “Noam Brown (@polynoamial) on X”, X (formerly Twitter). [Online]. Available: https://x.com/polynoamial/status/1870172996650053653?mx=2
Brown, T. B. et al.(2020). Language Models are Few-Shot Learners. arXiv.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language Models are Few-Shot Learners. In arXiv. https://arxiv.org/abs/2005.14165Brown, T. B., B. Mann, N. Ryder, et al. 2020. “Language Models Are Few-Shot Learners”. In arXiv. Preprint, May 28. https://arxiv.org/abs/2005.14165.Brown, T. B., et al. “Language Models Are Few-Shot Learners”. arXiv, 28 May 2020, https://arxiv.org/abs/2005.14165.Brown, T. B. et al. Language Models are Few-Shot Learners. arXiv Preprint at https://arxiv.org/abs/2005.14165 (2020).T. B. Brown et al., “Language Models are Few-Shot Learners”, May 28, 2020. [Online]. Available: https://arxiv.org/abs/2005.14165
Brundage(2025). Feedback on the Second Draft of the General-Purpose AI Code of Practice.Brundage. (2025). Feedback on the Second Draft of the General-Purpose AI Code of Practice. https://milesbrundage.substack.com/p/feedback-on-the-second-draft-of-theBrundage. 2025. Feedback on the Second Draft of the General-Purpose AI Code of Practice. Edition. https://milesbrundage.substack.com/p/feedback-on-the-second-draft-of-the.Brundage. Feedback on the Second Draft of the General-Purpose AI Code of Practice. 2025, https://milesbrundage.substack.com/p/feedback-on-the-second-draft-of-the.Brundage. Feedback on the Second Draft of the General-Purpose AI Code of Practice. https://milesbrundage.substack.com/p/feedback-on-the-second-draft-of-the (2025).Brundage, “Feedback on the Second Draft of the General-Purpose AI Code of Practice”. [Online]. Available: https://milesbrundage.substack.com/p/feedback-on-the-second-draft-of-the
Brundage, M. et al.(2018). The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation. arXiv.Brundage, M., Avin, S., Clark, J., Toner, H., Eckersley, P., Garfinkel, B., Dafoe, A., Scharre, P., Zeitzoff, T., Filar, B., Anderson, H., Roff, H., Allen, G. C., Steinhardt, J., Flynn, C., hÉigeartaigh, S. Ó., Beard, S., Belfield, H., Farquhar, S., … Amodei, D. (2018). The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation. In arXiv. https://arxiv.org/abs/1802.07228Brundage, M., S. Avin, J. Clark, et al. 2018. “The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation”. In arXiv. Preprint, February 20. https://arxiv.org/abs/1802.07228.Brundage, M., et al. “The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation”. arXiv, 20 Feb. 2018, https://arxiv.org/abs/1802.07228.Brundage, M. et al. The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation. arXiv Preprint at https://arxiv.org/abs/1802.07228 (2018).M. Brundage et al., “The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation”, Feb. 20, 2018. [Online]. Available: https://arxiv.org/abs/1802.07228
Bubeck, S. et al.(2023). Sparks of Artificial General Intelligence: Early experiments with GPT-4. arXiv.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., & Zhang, Y. (2023). Sparks of Artificial General Intelligence: Early experiments with GPT-4. In arXiv. https://arxiv.org/abs/2303.12712Bubeck, S., V. Chandrasekaran, R. Eldan, et al. 2023. “Sparks of Artificial General Intelligence: Early Experiments with GPT-4”. In arXiv. Preprint, March 22. https://arxiv.org/abs/2303.12712.Bubeck, S., et al. “Sparks of Artificial General Intelligence: Early Experiments with GPT-4”. arXiv, 22 Mar. 2023, https://arxiv.org/abs/2303.12712.Bubeck, S. et al. Sparks of Artificial General Intelligence: Early experiments with GPT-4. arXiv Preprint at https://arxiv.org/abs/2303.12712 (2023).S. Bubeck et al., “Sparks of Artificial General Intelligence: Early experiments with GPT-4”, Mar. 22, 2023. [Online]. Available: https://arxiv.org/abs/2303.12712
Buchanan(2020). The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology.Buchanan. (2020). The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology. Center for Security and Emerging Technology. https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategyBuchanan. 2020. “The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology”. Center for Security and Emerging Technology. https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategy.Buchanan. “The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology”. Center for Security and Emerging Technology, 2020, https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategy.Buchanan. The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology. Center for Security and Emerging Technology https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategy (2020).Buchanan, “The AI Triad and What It Means for National Security Strategy | Center for Security and Emerging Technology”, Center for Security and Emerging Technology. [Online]. Available: https://cset.georgetown.edu/publication/the-ai-triad-and-what-it-means-for-national-security-strategy
Buck(2022). The prototypical catastrophic AI action is getting root access to its datacenter. AI Alignment Forum.Buck. (2022, June 2). The prototypical catastrophic AI action is getting root access to its datacenter. AI Alignment Forum. https://alignmentforum.org/posts/BAzCGCys4BkzGDCWR/the-prototypical-catastrophic-ai-action-is-getting-rootBuck. 2022. “The Prototypical Catastrophic AI Action Is Getting Root Access to Its Datacenter”. AI Alignment Forum, June 2. https://alignmentforum.org/posts/BAzCGCys4BkzGDCWR/the-prototypical-catastrophic-ai-action-is-getting-root.Buck. “The Prototypical Catastrophic AI Action Is Getting Root Access to Its Datacenter”. AI Alignment Forum, 2 June 2022, https://alignmentforum.org/posts/BAzCGCys4BkzGDCWR/the-prototypical-catastrophic-ai-action-is-getting-root.Buck. The prototypical catastrophic AI action is getting root access to its datacenter. AI Alignment Forum https://alignmentforum.org/posts/BAzCGCys4BkzGDCWR/the-prototypical-catastrophic-ai-action-is-getting-root (2022).Buck, “The prototypical catastrophic AI action is getting root access to its datacenter”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/BAzCGCys4BkzGDCWR/the-prototypical-catastrophic-ai-action-is-getting-root
Buck(2024). Access to powerful AI might make computer security radically easier. AI Alignment Forum.Buck. (2024, June 8). Access to powerful AI might make computer security radically easier. AI Alignment Forum. https://alignmentforum.org/posts/2wxufQWK8rXcDGbyL/access-to-powerful-ai-might-make-computer-security-radicallyBuck. 2024. “Access to Powerful AI Might Make Computer Security Radically Easier”. AI Alignment Forum, June 8. https://alignmentforum.org/posts/2wxufQWK8rXcDGbyL/access-to-powerful-ai-might-make-computer-security-radically.Buck. “Access to Powerful AI Might Make Computer Security Radically Easier”. AI Alignment Forum, 8 June 2024, https://alignmentforum.org/posts/2wxufQWK8rXcDGbyL/access-to-powerful-ai-might-make-computer-security-radically.Buck. Access to powerful AI might make computer security radically easier. AI Alignment Forum https://alignmentforum.org/posts/2wxufQWK8rXcDGbyL/access-to-powerful-ai-might-make-computer-security-radically (2024).Buck, “Access to powerful AI might make computer security radically easier”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/2wxufQWK8rXcDGbyL/access-to-powerful-ai-might-make-computer-security-radically
Buhl, M. D., Sett, G., Koessler, L., Schuett, J. & Anderljung, M.(2024). Safety cases for frontier AI. arXiv.Buhl, M. D., Sett, G., Koessler, L., Schuett, J., & Anderljung, M. (2024). Safety cases for frontier AI. In arXiv. https://doi.org/10.48550/arXiv.2410.21572Buhl, M. D., G. Sett, L. Koessler, J. Schuett, and M. Anderljung. 2024. “Safety Cases for Frontier AI”. In arXiv. Preprint, October. https://doi.org/10.48550/arXiv.2410.21572.Buhl, M. D., et al. “Safety Cases for Frontier AI”. arXiv, Oct. 2024, https://doi.org/10.48550/arXiv.2410.21572.Buhl, M. D., Sett, G., Koessler, L., Schuett, J. & Anderljung, M. Safety cases for frontier AI. arXiv Preprint at https://doi.org/10.48550/arXiv.2410.21572 (2024).M. D. Buhl, G. Sett, L. Koessler, J. Schuett, and M. Anderljung, “Safety cases for frontier AI”, Oct. 2024. doi: 10.48550/arXiv.2410.21572.
Burden, J.(2024). Evaluating AI Evaluation: Perils and Prospects. arXiv.Burden, J. (2024). Evaluating AI Evaluation: Perils and Prospects. In arXiv. https://arxiv.org/abs/2407.09221Burden, J. 2024. “Evaluating AI Evaluation: Perils and Prospects”. In arXiv. Preprint, July 12. https://arxiv.org/abs/2407.09221.Burden, J. “Evaluating AI Evaluation: Perils and Prospects”. arXiv, 12 July 2024, https://arxiv.org/abs/2407.09221.Burden, J. Evaluating AI Evaluation: Perils and Prospects. arXiv Preprint at https://arxiv.org/abs/2407.09221 (2024).J. Burden, “Evaluating AI Evaluation: Perils and Prospects”, Jul. 12, 2024. [Online]. Available: https://arxiv.org/abs/2407.09221
Burns, C. et al.(2023). Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. arXiv.Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., Sutskever, I., & Wu, J. (2023). Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. In arXiv. https://arxiv.org/abs/2312.09390Burns, C., P. Izmailov, J. H. Kirchner, et al. 2023. “Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision”. In arXiv. Preprint, December 14. https://arxiv.org/abs/2312.09390.Burns, C., et al. “Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision”. arXiv, 14 Dec. 2023, https://arxiv.org/abs/2312.09390.Burns, C. et al. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. arXiv Preprint at https://arxiv.org/abs/2312.09390 (2023).C. Burns et al., “Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision”, Dec. 14, 2023. [Online]. Available: https://arxiv.org/abs/2312.09390
Buterin(2023). My techno-optimism.Buterin. (2023). My techno-optimism. https://vitalik.eth.limo/general/2023/11/27/techno_optimism.htmlButerin. 2023. “My Techno-optimism”. https://vitalik.eth.limo/general/2023/11/27/techno_optimism.html.Buterin. My Techno-optimism. 2023, https://vitalik.eth.limo/general/2023/11/27/techno_optimism.html.Buterin. My techno-optimism. https://vitalik.eth.limo/general/2023/11/27/techno_optimism.html (2023).Buterin, “My techno-optimism”. [Online]. Available: https://vitalik.eth.limo/general/2023/11/27/techno_optimism.html
Buterin(2025). d/acc: one year later.Buterin. (2025). d/acc: one year later. https://vitalik.eth.limo/general/2025/01/05/dacc2.htmlButerin. 2025. “D/acc: One Year Later”. https://vitalik.eth.limo/general/2025/01/05/dacc2.html.Buterin. D/acc: One Year Later. 2025, https://vitalik.eth.limo/general/2025/01/05/dacc2.html.Buterin. d/acc: one year later. https://vitalik.eth.limo/general/2025/01/05/dacc2.html (2025).Buterin, “d/acc: one year later”. [Online]. Available: https://vitalik.eth.limo/general/2025/01/05/dacc2.html
Caballero, E., Gupta, K., Rish, I. & Krueger, D.(2022). Broken Neural Scaling Laws. arXiv.Caballero, E., Gupta, K., Rish, I., & Krueger, D. (2022). Broken Neural Scaling Laws. In arXiv. https://arxiv.org/abs/2210.14891Caballero, E., K. Gupta, I. Rish, and D. Krueger. 2022. “Broken Neural Scaling Laws”. In arXiv. Preprint, October 26. https://arxiv.org/abs/2210.14891.Caballero, E., et al. “Broken Neural Scaling Laws”. arXiv, 26 Oct. 2022, https://arxiv.org/abs/2210.14891.Caballero, E., Gupta, K., Rish, I. & Krueger, D. Broken Neural Scaling Laws. arXiv Preprint at https://arxiv.org/abs/2210.14891 (2022).E. Caballero, K. Gupta, I. Rish, and D. Krueger, “Broken Neural Scaling Laws”, Oct. 26, 2022. [Online]. Available: https://arxiv.org/abs/2210.14891
CAIS(2023). Statement on AI Extinction Risk | CAIS. Center for AI Safety.CAIS. (2023). Statement on AI Extinction Risk | CAIS. Center for AI Safety. https://safe.ai/work/statement-on-ai-riskCAIS. 2023. “Statement on AI Extinction Risk | CAIS”. Center for AI Safety. https://safe.ai/work/statement-on-ai-risk.CAIS. “Statement on AI Extinction Risk | CAIS”. Center for AI Safety, 2023, https://safe.ai/work/statement-on-ai-risk.CAIS. Statement on AI Extinction Risk | CAIS. Center for AI Safety https://safe.ai/work/statement-on-ai-risk (2023).CAIS, “Statement on AI Extinction Risk | CAIS”, Center for AI Safety. [Online]. Available: https://safe.ai/work/statement-on-ai-risk
Campos, S., Papadatos, H., Roger, F., Touzet, C., Quarks, O. & Murray, M.(2025). A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management. arXiv.Campos, S., Papadatos, H., Roger, F., Touzet, C., Quarks, O., & Murray, M. (2025). A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management. In arXiv. https://arxiv.org/abs/2502.06656Campos, S., H. Papadatos, F. Roger, C. Touzet, O. Quarks, and M. Murray. 2025. “A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management”. In arXiv. Preprint, February 10. https://arxiv.org/abs/2502.06656.Campos, S., et al. “A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management”. arXiv, 10 Feb. 2025, https://arxiv.org/abs/2502.06656.Campos, S. et al. A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management. arXiv Preprint at https://arxiv.org/abs/2502.06656 (2025).S. Campos, H. Papadatos, F. Roger, C. Touzet, O. Quarks, and M. Murray, “A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management”, Feb. 10, 2025. [Online]. Available: https://arxiv.org/abs/2502.06656
Cao, B. et al.(2024). Towards Scalable Automated Alignment of LLMs: A Survey. arXiv.Cao, B., Lu, K., Lu, X., Chen, J., Ren, M., Xiang, H., Liu, P., Lu, Y., He, B., Han, X., Sun, L., Lin, H., & Yu, B. (2024). Towards Scalable Automated Alignment of LLMs: A Survey. In arXiv. https://arxiv.org/abs/2406.01252Cao, B., K. Lu, X. Lu, et al. 2024. “Towards Scalable Automated Alignment of LLMs: A Survey”. In arXiv. Preprint, June 3. https://arxiv.org/abs/2406.01252.Cao, B., et al. “Towards Scalable Automated Alignment of LLMs: A Survey”. arXiv, 3 June 2024, https://arxiv.org/abs/2406.01252.Cao, B. et al. Towards Scalable Automated Alignment of LLMs: A Survey. arXiv Preprint at https://arxiv.org/abs/2406.01252 (2024).B. Cao et al., “Towards Scalable Automated Alignment of LLMs: A Survey”, Jun. 03, 2024. [Online]. Available: https://arxiv.org/abs/2406.01252
Cao, K. et al.(2023). Large-scale pancreatic cancer detection via non-contrast CT and deep learning. Nature Medicine.Cao, K., Xia, Y., Yao, J., Han, X., Lambert, L., Zhang, T., Tang, W., Jin, G., Jiang, H., Fang, X., Nogues, I., Li, X., Guo, W., Wang, Y., Fang, W., Qiu, M., Hou, Y., Kovarnik, T., Vocka, M., … Lu, J. (2023). Large-scale pancreatic cancer detection via non-contrast CT and deep learning. Nature Medicine, 29, 3033–3043. https://doi.org/10.1038/s41591-023-02640-wCao, K., Y. Xia, J. Yao, et al. 2023. “Large-scale Pancreatic Cancer Detection via Non-contrast CT and Deep Learning”. Nature Medicine 29 (November): 3033–43. https://doi.org/10.1038/s41591-023-02640-w.Cao, K., et al. “Large-scale Pancreatic Cancer Detection via Non-contrast CT and Deep Learning”. Nature Medicine, vol. 29, Nov. 2023, pp. 3033–43, https://doi.org/10.1038/s41591-023-02640-w.Cao, K. et al. Large-scale pancreatic cancer detection via non-contrast CT and deep learning. Nature Medicine 29, 3033–3043 (2023).K. Cao et al., “Large-scale pancreatic cancer detection via non-contrast CT and deep learning”, Nature Medicine, vol. 29, pp. 3033–3043, Nov. 2023, doi: 10.1038/s41591-023-02640-w.
Carlini, N. et al.(2020). Extracting Training Data from Large Language Models. arXiv.Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., Oprea, A., & Raffel, C. (2020). Extracting Training Data from Large Language Models. In arXiv. https://arxiv.org/abs/2012.07805Carlini, N., F. Tramer, E. Wallace, et al. 2020. “Extracting Training Data from Large Language Models”. In arXiv. Preprint, December 14. https://arxiv.org/abs/2012.07805.Carlini, N., et al. “Extracting Training Data from Large Language Models”. arXiv, 14 Dec. 2020, https://arxiv.org/abs/2012.07805.Carlini, N. et al. Extracting Training Data from Large Language Models. arXiv Preprint at https://arxiv.org/abs/2012.07805 (2020).N. Carlini et al., “Extracting Training Data from Large Language Models”, Dec. 14, 2020. [Online]. Available: https://arxiv.org/abs/2012.07805
Carlsmith(2020). How Much Computational Power Does It Take to Match the Human Brain?. Coefficient Giving.Carlsmith. (2020). How Much Computational Power Does It Take to Match the Human Brain?. Coefficient Giving. https://coefficientgiving.org/research/how-much-computational-power-does-it-take-to-match-the-human-brainCarlsmith. 2020. “How Much Computational Power Does It Take to Match the Human Brain?”. Coefficient Giving. https://coefficientgiving.org/research/how-much-computational-power-does-it-take-to-match-the-human-brain.Carlsmith. “How Much Computational Power Does It Take to Match the Human Brain?”. Coefficient Giving, 2020, https://coefficientgiving.org/research/how-much-computational-power-does-it-take-to-match-the-human-brain.Carlsmith. How Much Computational Power Does It Take to Match the Human Brain?. Coefficient Giving https://coefficientgiving.org/research/how-much-computational-power-does-it-take-to-match-the-human-brain (2020).Carlsmith, “How Much Computational Power Does It Take to Match the Human Brain?”, Coefficient Giving. [Online]. Available: https://coefficientgiving.org/research/how-much-computational-power-does-it-take-to-match-the-human-brain
Carlsmith, J.(2022). Is Power-Seeking AI an Existential Risk?. arXiv.Carlsmith, J. (2022). Is Power-Seeking AI an Existential Risk?. In arXiv. https://arxiv.org/abs/2206.13353Carlsmith, J. 2022. “Is Power-Seeking AI an Existential Risk?”. In arXiv. Preprint, June 16. https://arxiv.org/abs/2206.13353.Carlsmith, J. “Is Power-Seeking AI an Existential Risk?”. arXiv, 16 June 2022, https://arxiv.org/abs/2206.13353.Carlsmith, J. Is Power-Seeking AI an Existential Risk?. arXiv Preprint at https://arxiv.org/abs/2206.13353 (2022).J. Carlsmith, “Is Power-Seeking AI an Existential Risk?”, Jun. 16, 2022. [Online]. Available: https://arxiv.org/abs/2206.13353
Carlsmith, J.(2023). Scheming AIs: Will AIs fake alignment during training in order to get power?. arXiv.Carlsmith, J. (2023). Scheming AIs: Will AIs fake alignment during training in order to get power?. In arXiv. https://arxiv.org/abs/2311.08379Carlsmith, J. 2023. “Scheming AIs: Will AIs Fake Alignment During Training in Order to Get Power?”. In arXiv. Preprint, November 14. https://arxiv.org/abs/2311.08379.Carlsmith, J. “Scheming AIs: Will AIs Fake Alignment During Training in Order to Get Power?”. arXiv, 14 Nov. 2023, https://arxiv.org/abs/2311.08379.Carlsmith, J. Scheming AIs: Will AIs fake alignment during training in order to get power?. arXiv Preprint at https://arxiv.org/abs/2311.08379 (2023).J. Carlsmith, “Scheming AIs: Will AIs fake alignment during training in order to get power?”, Nov. 14, 2023. [Online]. Available: https://arxiv.org/abs/2311.08379
Carlson, R.(2009). The changing economics of DNA synthesis. Nature Biotechnology.Carlson, R. (2009). The changing economics of DNA synthesis. Nature Biotechnology. https://doi.org/10.1038/nbt1209-1091Carlson, R. 2009. “The Changing Economics of DNA Synthesis”. Nature Biotechnology, ahead of print, December. https://doi.org/10.1038/nbt1209-1091.Carlson, R. “The Changing Economics of DNA Synthesis”. Nature Biotechnology, Dec. 2009, https://doi.org/10.1038/nbt1209-1091.Carlson, R. The changing economics of DNA synthesis. Nature Biotechnology https://doi.org/10.1038/nbt1209-1091 (2009) doi:10.1038/nbt1209-1091.R. Carlson, “The changing economics of DNA synthesis”, Nature Biotechnology, Dec. 2009, doi: 10.1038/nbt1209-1091.
Carter, S. R., Yassif, J. M. & Isaac, C. R.(2023). Benchtop DNA Synthesis Devices: Capabilities, Biosecurity Implications, and Governance.Carter, S. R., Yassif, J. M., & Isaac, C. R. (2023). Benchtop DNA Synthesis Devices: Capabilities, Biosecurity Implications, and Governance. NTI | bio. https://nti.org/wp-content/uploads/2023/05/NTIBIO_Benchtop-DNA-Report_FINAL.pdfCarter, S. R., J. M. Yassif, and C. R. Isaac. 2023. Benchtop DNA Synthesis Devices: Capabilities, Biosecurity Implications, and Governance. NTI | bio. https://nti.org/wp-content/uploads/2023/05/NTIBIO_Benchtop-DNA-Report_FINAL.pdf.Carter, S. R., et al. Benchtop DNA Synthesis Devices: Capabilities, Biosecurity Implications, and Governance. NTI | bio, May 2023, https://nti.org/wp-content/uploads/2023/05/NTIBIO_Benchtop-DNA-Report_FINAL.pdf.Carter, S. R., Yassif, J. M. & Isaac, C. R. Benchtop DNA Synthesis Devices: Capabilities, Biosecurity Implications, and Governance. https://nti.org/wp-content/uploads/2023/05/NTIBIO_Benchtop-DNA-Report_FINAL.pdf (2023).S. R. Carter, J. M. Yassif, and C. R. Isaac, “Benchtop DNA Synthesis Devices: Capabilities, Biosecurity Implications, and Governance”, NTI | bio, May 2023. [Online]. Available: https://nti.org/wp-content/uploads/2023/05/NTIBIO_Benchtop-DNA-Report_FINAL.pdf
Carvalho, B. W., Garcez, A. S. D., Lamb, L. C. & Brazil, E. V.(2025). Grokking Explained: A Statistical Phenomenon. arXiv.Carvalho, B. W., Garcez, A. S. d'Avila ., Lamb, L. C., & Brazil, E. V. (2025). Grokking Explained: A Statistical Phenomenon. In arXiv. https://arxiv.org/abs/2502.01774Carvalho, B. W., A. S. d'Avila . Garcez, L. C. Lamb, and E. V. Brazil. 2025. “Grokking Explained: A Statistical Phenomenon”. In arXiv. Preprint, February 3. https://arxiv.org/abs/2502.01774.Carvalho, B. W., et al. “Grokking Explained: A Statistical Phenomenon”. arXiv, 3 Feb. 2025, https://arxiv.org/abs/2502.01774.Carvalho, B. W., Garcez, A. S. d'Avila ., Lamb, L. C. & Brazil, E. V. Grokking Explained: A Statistical Phenomenon. arXiv Preprint at https://arxiv.org/abs/2502.01774 (2025).B. W. Carvalho, A. S. d'Avila . Garcez, L. C. Lamb, and E. V. Brazil, “Grokking Explained: A Statistical Phenomenon”, Feb. 03, 2025. [Online]. Available: https://arxiv.org/abs/2502.01774
Casper, S. et al.(2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. arXiv.Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., … Hadfield-Menell, D. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. In arXiv. https://arxiv.org/abs/2307.15217Casper, S., X. Davies, C. Shi, et al. 2023. “Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback”. In arXiv. Preprint, July 27. https://arxiv.org/abs/2307.15217.Casper, S., et al. “Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback”. arXiv, 27 July 2023, https://arxiv.org/abs/2307.15217.Casper, S. et al. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. arXiv Preprint at https://arxiv.org/abs/2307.15217 (2023).S. Casper et al., “Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback”, Jul. 27, 2023. [Online]. Available: https://arxiv.org/abs/2307.15217
Casper, S. et al.(2024). Black-Box Access is Insufficient for Rigorous AI Audits. arXiv.Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., Sharkey, L., Krishna, S., Hagen, M. V., Alberti, S., Chan, A., Sun, Q., Gerovitch, M., Bau, D., Tegmark, M., … Hadfield-Menell, D. (2024). Black-Box Access is Insufficient for Rigorous AI Audits. In arXiv. https://doi.org/10.1145/3630106.3659037Casper, S., C. Ezell, C. Siegmann, et al. 2024. “Black-Box Access Is Insufficient for Rigorous AI Audits”. In arXiv. Preprint, January 25. https://doi.org/10.1145/3630106.3659037.Casper, S., et al. “Black-Box Access Is Insufficient for Rigorous AI Audits”. arXiv, 25 Jan. 2024, https://doi.org/10.1145/3630106.3659037.Casper, S. et al. Black-Box Access is Insufficient for Rigorous AI Audits. arXiv Preprint at https://doi.org/10.1145/3630106.3659037 (2024).S. Casper et al., “Black-Box Access is Insufficient for Rigorous AI Audits”, Jan. 25, 2024. doi: 10.1145/3630106.3659037.
Casper, S., Krueger, D. & Hadfield-Menell, D.(2025). Pitfalls of Evidence-Based AI Policy. arXiv.Casper, S., Krueger, D., & Hadfield-Menell, D. (2025). Pitfalls of Evidence-Based AI Policy. In arXiv. https://arxiv.org/abs/2502.09618Casper, S., D. Krueger, and D. Hadfield-Menell. 2025. “Pitfalls of Evidence-Based AI Policy”. In arXiv. Preprint, February 13. https://arxiv.org/abs/2502.09618.Casper, S., et al. “Pitfalls of Evidence-Based AI Policy”. arXiv, 13 Feb. 2025, https://arxiv.org/abs/2502.09618.Casper, S., Krueger, D. & Hadfield-Menell, D. Pitfalls of Evidence-Based AI Policy. arXiv Preprint at https://arxiv.org/abs/2502.09618 (2025).S. Casper, D. Krueger, and D. Hadfield-Menell, “Pitfalls of Evidence-Based AI Policy”, Feb. 13, 2025. [Online]. Available: https://arxiv.org/abs/2502.09618
Caucheteux, C. & King, J.(2022). Brains and algorithms partially converge in natural language processing. Communications Biology.Caucheteux, C., & King, J.-R. (2022). Brains and algorithms partially converge in natural language processing. Communications Biology, 5, 134. https://doi.org/10.1038/s42003-022-03036-1Caucheteux, C., and J.-R. King. 2022. “Brains and Algorithms Partially Converge in Natural Language Processing”. Communications Biology 5 (February): 134. https://doi.org/10.1038/s42003-022-03036-1.Caucheteux, C., and J.-R. King. “Brains and Algorithms Partially Converge in Natural Language Processing”. Communications Biology, vol. 5, Feb. 2022, p. 134, https://doi.org/10.1038/s42003-022-03036-1.Caucheteux, C. & King, J.-R. Brains and algorithms partially converge in natural language processing. Communications Biology 5, 134 (2022).C. Caucheteux and J.-R. King, “Brains and algorithms partially converge in natural language processing”, Communications Biology, vol. 5, p. 134, Feb. 2022, doi: 10.1038/s42003-022-03036-1.
Cave, S. & ÓhÉigeartaigh, S. S.(2018). An AI Race for Strategic Advantage. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society.Cave, S., & ÓhÉigeartaigh, S. S. (2018, December 27). An AI Race for Strategic Advantage. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society. https://doi.org/10.1145/3278721.3278780Cave, S., and S. S. ÓhÉigeartaigh. 2018. “An AI Race for Strategic Advantage”. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, December 27. https://doi.org/10.1145/3278721.3278780.Cave, S., and S. S. ÓhÉigeartaigh. “An AI Race for Strategic Advantage”. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, 2018, https://doi.org/10.1145/3278721.3278780.Cave, S. & ÓhÉigeartaigh, S. S. An AI Race for Strategic Advantage. in Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society (2018). doi:10.1145/3278721.3278780.S. Cave and S. S. ÓhÉigeartaigh, “An AI Race for Strategic Advantage”, in Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, Dec. 2018. doi: 10.1145/3278721.3278780.
Cerutti et al.(2025). The Global Impact of AI – Mind the Gap. IMF eLibrary.Cerutti et al. (2025). The Global Impact of AI – Mind the Gap. Internet Archive (https://web.archive.org/web/20260909150837/https://www.elibrary.imf.org/view/journals/001/2025/076/article-A001-en.xml). IMF eLibrary. https://elibrary.imf.org/view/journals/001/2025/076/article-A001-en.xmlCerutti et al. 2025. “The Global Impact of AI – Mind the Gap”. IMF eLibrary. Https://web.archive.org/web/20260909150837/https://www.elibrary.imf.org/view/journals/001/2025/076/article-A001-en.xml. Internet Archive. https://elibrary.imf.org/view/journals/001/2025/076/article-A001-en.xml.Cerutti et al. “The Global Impact of AI – Mind the Gap”. IMF eLibrary, 2025, Internet Archive, https://web.archive.org/web/20260909150837/https://www.elibrary.imf.org/view/journals/001/2025/076/article-A001-en.xml, https://elibrary.imf.org/view/journals/001/2025/076/article-A001-en.xml.Cerutti et al. The Global Impact of AI – Mind the Gap. IMF eLibrary https://elibrary.imf.org/view/journals/001/2025/076/article-A001-en.xml (2025).Cerutti et al., “The Global Impact of AI – Mind the Gap”, IMF eLibrary. Accessed: Sep. 09, 2026. [Online]. Available: https://elibrary.imf.org/view/journals/001/2025/076/article-A001-en.xml
Ceruzzi(1989). Beyond The Limits. MIT Press.Ceruzzi. (1989). Beyond The Limits. Internet Archive (https://web.archive.org/web/20241118205221/https://mitpress.mit.edu/9780262530828/beyond-the-limits/). MIT Press. https://mitpress.mit.edu/9780262530828/beyond-the-limitsCeruzzi. 1989. “Beyond The Limits”. MIT Press. Https://web.archive.org/web/20241118205221/https://mitpress.mit.edu/9780262530828/beyond-the-limits/. Internet Archive. https://mitpress.mit.edu/9780262530828/beyond-the-limits.Ceruzzi. “Beyond The Limits”. MIT Press, 1989, Internet Archive, https://web.archive.org/web/20241118205221/https://mitpress.mit.edu/9780262530828/beyond-the-limits/, https://mitpress.mit.edu/9780262530828/beyond-the-limits.Ceruzzi. Beyond The Limits. MIT Press https://mitpress.mit.edu/9780262530828/beyond-the-limits (1989).Ceruzzi, “Beyond The Limits”, MIT Press. Accessed: Nov. 18, 2024. [Online]. Available: https://mitpress.mit.edu/9780262530828/beyond-the-limits
Cha, S.(2024). Towards an international regulatory framework for AI safety: lessons from the IAEA’s nuclear safety regulations. Humanities and Social Sciences Communications.Cha, S. (2024). Towards an international regulatory framework for AI safety: lessons from the IAEA’s nuclear safety regulations. Humanities and Social Sciences Communications, 11, 506. https://doi.org/10.1057/s41599-024-03017-1Cha, S. 2024. “Towards an International Regulatory Framework for AI Safety: Lessons from the IAEA’s Nuclear Safety Regulations”. Humanities and Social Sciences Communications 11 (April): 506. https://doi.org/10.1057/s41599-024-03017-1.Cha, S. “Towards an International Regulatory Framework for AI Safety: Lessons from the IAEA’s Nuclear Safety Regulations”. Humanities and Social Sciences Communications, vol. 11, Apr. 2024, p. 506, https://doi.org/10.1057/s41599-024-03017-1.Cha, S. Towards an international regulatory framework for AI safety: lessons from the IAEA’s nuclear safety regulations. Humanities and Social Sciences Communications 11, 506 (2024).S. Cha, “Towards an international regulatory framework for AI safety: lessons from the IAEA’s nuclear safety regulations”, Humanities and Social Sciences Communications, vol. 11, p. 506, Apr. 2024, doi: 10.1057/s41599-024-03017-1.
Chan, A. et al.(2024). IDs for AI Systems. arXiv.Chan, A., Kolt, N., Wills, P., Anwar, U., de Witt, C. S., Rajkumar, N., Hammond, L., Krueger, D., Heim, L., & Anderljung, M. (2024). IDs for AI Systems. In arXiv. https://arxiv.org/abs/2406.12137Chan, A., N. Kolt, P. Wills, et al. 2024. “IDs for AI Systems”. In arXiv. Preprint, June 17. https://arxiv.org/abs/2406.12137.Chan, A., et al. “IDs for AI Systems”. arXiv, 17 June 2024, https://arxiv.org/abs/2406.12137.Chan, A. et al. IDs for AI Systems. arXiv Preprint at https://arxiv.org/abs/2406.12137 (2024).A. Chan et al., “IDs for AI Systems”, Jun. 17, 2024. [Online]. Available: https://arxiv.org/abs/2406.12137
Chan, A. et al.(2024). Visibility into AI Agents. arXiv.Chan, A., Ezell, C., Kaufmann, M., Wei, K., Hammond, L., Bradley, H., Bluemke, E., Rajkumar, N., Krueger, D., Kolt, N., Heim, L., & Anderljung, M. (2024). Visibility into AI Agents. In arXiv. https://arxiv.org/abs/2401.13138Chan, A., C. Ezell, M. Kaufmann, et al. 2024. “Visibility into AI Agents”. In arXiv. Preprint, January 23. https://arxiv.org/abs/2401.13138.Chan, A., et al. “Visibility into AI Agents”. arXiv, 23 Jan. 2024, https://arxiv.org/abs/2401.13138.Chan, A. et al. Visibility into AI Agents. arXiv Preprint at https://arxiv.org/abs/2401.13138 (2024).A. Chan et al., “Visibility into AI Agents”, Jan. 23, 2024. [Online]. Available: https://arxiv.org/abs/2401.13138
Chandra et al.(2024). Reducing Risks Posed by Synthetic Content An Overview of Technical Approaches to Digital Content Transparency. NIST.Chandra et al. (2024, November 20). Reducing Risks Posed by Synthetic Content An Overview of Technical Approaches to Digital Content Transparency. NIST. https://nist.gov/publications/reducing-risks-posed-synthetic-content-overview-technical-approaches-digital-contentChandra et al. 2024. “Reducing Risks Posed by Synthetic Content An Overview of Technical Approaches to Digital Content Transparency”. NIST, November 20. https://nist.gov/publications/reducing-risks-posed-synthetic-content-overview-technical-approaches-digital-content.Chandra et al. “Reducing Risks Posed by Synthetic Content An Overview of Technical Approaches to Digital Content Transparency”. NIST, 20 Nov. 2024, https://nist.gov/publications/reducing-risks-posed-synthetic-content-overview-technical-approaches-digital-content.Chandra et al. Reducing Risks Posed by Synthetic Content An Overview of Technical Approaches to Digital Content Transparency. NIST https://nist.gov/publications/reducing-risks-posed-synthetic-content-overview-technical-approaches-digital-content (2024).Chandra et al., “Reducing Risks Posed by Synthetic Content An Overview of Technical Approaches to Digital Content Transparency”, NIST. [Online]. Available: https://nist.gov/publications/reducing-risks-posed-synthetic-content-overview-technical-approaches-digital-content
Chang, C.(2024). The First Global AI Treaty: Analyzing the Framework Convention on Artificial Intelligence and the EU AI Act. University of Illinois Law Review (Online).Chang, C.-C. (2024). The First Global AI Treaty: Analyzing the Framework Convention on Artificial Intelligence and the EU AI Act. University of Illinois Law Review (Online), 2024, 86. https://doi.org/10.2139/ssrn.5069335Chang, C.-C. 2024. “The First Global AI Treaty: Analyzing the Framework Convention on Artificial Intelligence and the EU AI Act”. University of Illinois Law Review (Online) 2024 (December): 86. https://doi.org/10.2139/ssrn.5069335.Chang, C.-C. “The First Global AI Treaty: Analyzing the Framework Convention on Artificial Intelligence and the EU AI Act”. University of Illinois Law Review (Online), vol. 2024, Dec. 2024, p. 86, https://doi.org/10.2139/ssrn.5069335.Chang, C.-C. The First Global AI Treaty: Analyzing the Framework Convention on Artificial Intelligence and the EU AI Act. University of Illinois Law Review (Online) 2024, 86 (2024).C.-C. Chang, “The First Global AI Treaty: Analyzing the Framework Convention on Artificial Intelligence and the EU AI Act”, University of Illinois Law Review (Online), vol. 2024, p. 86, Dec. 2024, doi: 10.2139/ssrn.5069335.
Charbel-Raphaël & cozyfractal(2024). What convincing warning shot could help prevent extinction from AI?. LessWrong.Charbel-Raphaël, & cozyfractal. (2024, April 13). What convincing warning shot could help prevent extinction from AI?. LessWrong. https://lesswrong.com/posts/RYx6cLwzoajqjyB6b/what-convincing-warning-shot-could-help-prevent-extinctionCharbel-Raphaël, and cozyfractal. 2024. “What Convincing Warning Shot Could Help Prevent Extinction from AI?”. LessWrong, April 13. https://lesswrong.com/posts/RYx6cLwzoajqjyB6b/what-convincing-warning-shot-could-help-prevent-extinction.Charbel-Raphaël, and cozyfractal. “What Convincing Warning Shot Could Help Prevent Extinction from AI?”. LessWrong, 13 Apr. 2024, https://lesswrong.com/posts/RYx6cLwzoajqjyB6b/what-convincing-warning-shot-could-help-prevent-extinction.Charbel-Raphaël & cozyfractal. What convincing warning shot could help prevent extinction from AI?. LessWrong https://lesswrong.com/posts/RYx6cLwzoajqjyB6b/what-convincing-warning-shot-could-help-prevent-extinction (2024).Charbel-Raphaël and cozyfractal, “What convincing warning shot could help prevent extinction from AI?”, LessWrong. [Online]. Available: https://lesswrong.com/posts/RYx6cLwzoajqjyB6b/what-convincing-warning-shot-could-help-prevent-extinction
Charbel-Raphaël & Épiphanie Gédéon(2024). We might be dropping the ball on Autonomous Replication and Adaptation. AI Alignment Forum.Charbel-Raphaël, & Épiphanie Gédéon. (2024, May 31). We might be dropping the ball on Autonomous Replication and Adaptation. AI Alignment Forum. https://alignmentforum.org/posts/xiRfJApXGDRsQBhvc/we-might-be-dropping-the-ball-on-autonomous-replication-and-1Charbel-Raphaël, and Épiphanie Gédéon. 2024. “We Might Be Dropping the Ball on Autonomous Replication and Adaptation.”. AI Alignment Forum, May 31. https://alignmentforum.org/posts/xiRfJApXGDRsQBhvc/we-might-be-dropping-the-ball-on-autonomous-replication-and-1.Charbel-Raphaël, and Épiphanie Gédéon. “We Might Be Dropping the Ball on Autonomous Replication and Adaptation.”. AI Alignment Forum, 31 May 2024, https://alignmentforum.org/posts/xiRfJApXGDRsQBhvc/we-might-be-dropping-the-ball-on-autonomous-replication-and-1.Charbel-Raphaël & Épiphanie Gédéon. We might be dropping the ball on Autonomous Replication and Adaptation. AI Alignment Forum https://alignmentforum.org/posts/xiRfJApXGDRsQBhvc/we-might-be-dropping-the-ball-on-autonomous-replication-and-1 (2024).Charbel-Raphaël and Épiphanie Gédéon, “We might be dropping the ball on Autonomous Replication and Adaptation.”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/xiRfJApXGDRsQBhvc/we-might-be-dropping-the-ball-on-autonomous-replication-and-1
Charbel-Raphaël & Gabin(2023). Davidad's Bold Plan for Alignment: An In-Depth Explanation. LessWrong.Charbel-Raphaël, & Gabin. (2023, April 19). Davidad's Bold Plan for Alignment: An In-Depth Explanation. LessWrong. https://lesswrong.com/posts/jRf4WENQnhssCb6mJ/davidad-s-bold-plan-for-alignment-an-in-depth-explanationCharbel-Raphaël, and Gabin. 2023. “Davidad's Bold Plan for Alignment: An In-Depth Explanation”. LessWrong, April 19. https://lesswrong.com/posts/jRf4WENQnhssCb6mJ/davidad-s-bold-plan-for-alignment-an-in-depth-explanation.Charbel-Raphaël, and Gabin. “Davidad's Bold Plan for Alignment: An In-Depth Explanation”. LessWrong, 19 Apr. 2023, https://lesswrong.com/posts/jRf4WENQnhssCb6mJ/davidad-s-bold-plan-for-alignment-an-in-depth-explanation.Charbel-Raphaël & Gabin. Davidad's Bold Plan for Alignment: An In-Depth Explanation. LessWrong https://lesswrong.com/posts/jRf4WENQnhssCb6mJ/davidad-s-bold-plan-for-alignment-an-in-depth-explanation (2023).Charbel-Raphaël and Gabin, “Davidad's Bold Plan for Alignment: An In-Depth Explanation”, LessWrong. [Online]. Available: https://lesswrong.com/posts/jRf4WENQnhssCb6mJ/davidad-s-bold-plan-for-alignment-an-in-depth-explanation
Charbel-Raphaël(2025). Comment on “johnswentworth's Shortform”. LessWrong.Charbel-Raphaël. (2025, April 16). Comment on “johnswentworth's Shortform”. LessWrong. https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=J2iPumP29GK9Qrm5pCharbel-Raphaël. 2025. “Comment on “johnswentworth's Shortform””. LessWrong, April 16. https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=J2iPumP29GK9Qrm5p.Charbel-Raphaël. “Comment on “johnswentworth's Shortform””. LessWrong, 16 Apr. 2025, https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=J2iPumP29GK9Qrm5p.Charbel-Raphaël. Comment on “johnswentworth's Shortform”. LessWrong https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=J2iPumP29GK9Qrm5p (2025).Charbel-Raphaël, “Comment on “johnswentworth's Shortform””, LessWrong. [Online]. Available: https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=J2iPumP29GK9Qrm5p
Charbel-Raphaël(2025). Comment on “johnswentworth's Shortform”. LessWrong.Charbel-Raphaël. (2025, April 17). Comment on “johnswentworth's Shortform”. LessWrong. https://lesswrong.com/posts/puv8fRDCH9jx5yhbX?commentId=aBcAh8H9cSzdXmgb7Charbel-Raphaël. 2025. “Comment on “johnswentworth's Shortform””. LessWrong, April 17. https://lesswrong.com/posts/puv8fRDCH9jx5yhbX?commentId=aBcAh8H9cSzdXmgb7.Charbel-Raphaël. “Comment on “johnswentworth's Shortform””. LessWrong, 17 Apr. 2025, https://lesswrong.com/posts/puv8fRDCH9jx5yhbX?commentId=aBcAh8H9cSzdXmgb7.Charbel-Raphaël. Comment on “johnswentworth's Shortform”. LessWrong https://lesswrong.com/posts/puv8fRDCH9jx5yhbX?commentId=aBcAh8H9cSzdXmgb7 (2025).Charbel-Raphaël, “Comment on “johnswentworth's Shortform””, LessWrong. [Online]. Available: https://lesswrong.com/posts/puv8fRDCH9jx5yhbX?commentId=aBcAh8H9cSzdXmgb7
Charlie Steiner(2022). Take 2: Building tools to help build FAI is a legitimate strategy, but it's dual-use. AI Alignment Forum.Charlie Steiner. (2022, December 3). Take 2: Building tools to help build FAI is a legitimate strategy, but it's dual-use. AI Alignment Forum. https://alignmentforum.org/posts/pxiaLFjyr4WPmFdcm/take-2-building-tools-to-help-build-fai-is-a-legitimateCharlie Steiner. 2022. “Take 2: Building Tools to Help Build FAI Is a Legitimate Strategy, but It's Dual-use.”. AI Alignment Forum, December 3. https://alignmentforum.org/posts/pxiaLFjyr4WPmFdcm/take-2-building-tools-to-help-build-fai-is-a-legitimate.Charlie Steiner. “Take 2: Building Tools to Help Build FAI Is a Legitimate Strategy, but It's Dual-use.”. AI Alignment Forum, 3 Dec. 2022, https://alignmentforum.org/posts/pxiaLFjyr4WPmFdcm/take-2-building-tools-to-help-build-fai-is-a-legitimate.Charlie Steiner. Take 2: Building tools to help build FAI is a legitimate strategy, but it's dual-use. AI Alignment Forum https://alignmentforum.org/posts/pxiaLFjyr4WPmFdcm/take-2-building-tools-to-help-build-fai-is-a-legitimate (2022).Charlie Steiner, “Take 2: Building tools to help build FAI is a legitimate strategy, but it's dual-use.”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/pxiaLFjyr4WPmFdcm/take-2-building-tools-to-help-build-fai-is-a-legitimate
Charlotte Siegmann & Markus Anderljung(2022). The Brussels Effect and Artificial Intelligence.Charlotte Siegmann, & Markus Anderljung. (2022, August 16). The Brussels Effect and Artificial Intelligence. https://governance.ai/research-paper/brussels-effect-aiCharlotte Siegmann, and Markus Anderljung. 2022. “The Brussels Effect and Artificial Intelligence”. August 16. https://governance.ai/research-paper/brussels-effect-ai.Charlotte Siegmann, and Markus Anderljung. The Brussels Effect and Artificial Intelligence. 16 Aug. 2022, https://governance.ai/research-paper/brussels-effect-ai.Charlotte Siegmann & Markus Anderljung. The Brussels Effect and Artificial Intelligence. https://governance.ai/research-paper/brussels-effect-ai (2022).Charlotte Siegmann and Markus Anderljung, “The Brussels Effect and Artificial Intelligence”. [Online]. Available: https://governance.ai/research-paper/brussels-effect-ai
Charvet, C. J.(2021). Cutting across structural and transcriptomic scales translates time across the lifespan in humans and chimpanzees. Proceedings of the Royal Society B: Biological Sciences.Charvet, C. J. (2021). Cutting across structural and transcriptomic scales translates time across the lifespan in humans and chimpanzees. Proceedings of the Royal Society B: Biological Sciences. https://doi.org/10.1098/rspb.2020.2987Charvet, C. J. 2021. “Cutting Across Structural and Transcriptomic Scales Translates Time Across the Lifespan in Humans and Chimpanzees”. Proceedings of the Royal Society B: Biological Sciences, ahead of print, February 10. https://doi.org/10.1098/rspb.2020.2987.Charvet, C. J. “Cutting Across Structural and Transcriptomic Scales Translates Time Across the Lifespan in Humans and Chimpanzees”. Proceedings of the Royal Society B: Biological Sciences, Feb. 2021, https://doi.org/10.1098/rspb.2020.2987.Charvet, C. J. Cutting across structural and transcriptomic scales translates time across the lifespan in humans and chimpanzees. Proceedings of the Royal Society B: Biological Sciences https://doi.org/10.1098/rspb.2020.2987 (2021) doi:10.1098/rspb.2020.2987.C. J. Charvet, “Cutting across structural and transcriptomic scales translates time across the lifespan in humans and chimpanzees”, Proceedings of the Royal Society B: Biological Sciences, Feb. 2021, doi: 10.1098/rspb.2020.2987.
Cheerla(2018). AlphaZero Explained. On AI.Cheerla. (2018, January 1). AlphaZero Explained. On AI. https://nikcheerla.github.io/deeplearningschool/2018/01/01/AlphaZero-ExplainedCheerla. 2018. “AlphaZero Explained”. On AI, January 1. https://nikcheerla.github.io/deeplearningschool/2018/01/01/AlphaZero-Explained.Cheerla. “AlphaZero Explained”. On AI, 1 Jan. 2018, https://nikcheerla.github.io/deeplearningschool/2018/01/01/AlphaZero-Explained.Cheerla. AlphaZero Explained. On AI https://nikcheerla.github.io/deeplearningschool/2018/01/01/AlphaZero-Explained (2018).Cheerla, “AlphaZero Explained”, On AI. [Online]. Available: https://nikcheerla.github.io/deeplearningschool/2018/01/01/AlphaZero-Explained
Chehoudi, R.(2025). Artificial intelligence and democracy: pathway to progress or decline?. Journal of Information Technology & Politics.Chehoudi, R. (2025). Artificial intelligence and democracy: pathway to progress or decline?. Journal of Information Technology & Politics. https://doi.org/10.1080/19331681.2025.2473994Chehoudi, R. 2025. “Artificial Intelligence and Democracy: Pathway to Progress or Decline?”. Journal of Information Technology & Politics, ahead of print, March 6. https://doi.org/10.1080/19331681.2025.2473994.Chehoudi, R. “Artificial Intelligence and Democracy: Pathway to Progress or Decline?”. Journal of Information Technology & Politics, Mar. 2025, https://doi.org/10.1080/19331681.2025.2473994.Chehoudi, R. Artificial intelligence and democracy: pathway to progress or decline?. Journal of Information Technology & Politics https://doi.org/10.1080/19331681.2025.2473994 (2025) doi:10.1080/19331681.2025.2473994.R. Chehoudi, “Artificial intelligence and democracy: pathway to progress or decline?”, Journal of Information Technology & Politics, Mar. 2025, doi: 10.1080/19331681.2025.2473994.
Chen, M. et al.(2021). Evaluating Large Language Models Trained on Code. arXiv.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. de O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., … Zaremba, W. (2021). Evaluating Large Language Models Trained on Code. In arXiv. https://arxiv.org/abs/2107.03374Chen, M., J. Tworek, H. Jun, et al. 2021. “Evaluating Large Language Models Trained on Code”. In arXiv. Preprint, July 7. https://arxiv.org/abs/2107.03374.Chen, M., et al. “Evaluating Large Language Models Trained on Code”. arXiv, 7 July 2021, https://arxiv.org/abs/2107.03374.Chen, M. et al. Evaluating Large Language Models Trained on Code. arXiv Preprint at https://arxiv.org/abs/2107.03374 (2021).M. Chen et al., “Evaluating Large Language Models Trained on Code”, Jul. 07, 2021. [Online]. Available: https://arxiv.org/abs/2107.03374
Chen, R., Arditi, A., Sleight, H., Evans, O. & Lindsey, J.(2025). Persona Vectors: Monitoring and Controlling Character Traits in Language Models. arXiv.Chen, R., Arditi, A., Sleight, H., Evans, O., & Lindsey, J. (2025). Persona Vectors: Monitoring and Controlling Character Traits in Language Models. In arXiv. https://arxiv.org/abs/2507.21509Chen, R., A. Arditi, H. Sleight, O. Evans, and J. Lindsey. 2025. “Persona Vectors: Monitoring and Controlling Character Traits in Language Models”. In arXiv. Preprint, July 29. https://arxiv.org/abs/2507.21509.Chen, R., et al. “Persona Vectors: Monitoring and Controlling Character Traits in Language Models”. arXiv, 29 July 2025, https://arxiv.org/abs/2507.21509.Chen, R., Arditi, A., Sleight, H., Evans, O. & Lindsey, J. Persona Vectors: Monitoring and Controlling Character Traits in Language Models. arXiv Preprint at https://arxiv.org/abs/2507.21509 (2025).R. Chen, A. Arditi, H. Sleight, O. Evans, and J. Lindsey, “Persona Vectors: Monitoring and Controlling Character Traits in Language Models”, Jul. 29, 2025. [Online]. Available: https://arxiv.org/abs/2507.21509
Chen, S., Zharmagambetov, A., Mahloujifar, S., Chaudhuri, K., Wagner, D. & Guo, C.(2024). SecAlign: Defending Against Prompt Injection with Preference Optimization. arXiv.Chen, S., Zharmagambetov, A., Mahloujifar, S., Chaudhuri, K., Wagner, D., & Guo, C. (2024). SecAlign: Defending Against Prompt Injection with Preference Optimization. In arXiv. https://doi.org/10.1145/3719027.3744836Chen, S., A. Zharmagambetov, S. Mahloujifar, K. Chaudhuri, D. Wagner, and C. Guo. 2024. “SecAlign: Defending Against Prompt Injection with Preference Optimization”. In arXiv. Preprint, October 7. https://doi.org/10.1145/3719027.3744836.Chen, S., et al. “SecAlign: Defending Against Prompt Injection with Preference Optimization”. arXiv, 7 Oct. 2024, https://doi.org/10.1145/3719027.3744836.Chen, S. et al. SecAlign: Defending Against Prompt Injection with Preference Optimization. arXiv Preprint at https://doi.org/10.1145/3719027.3744836 (2024).S. Chen, A. Zharmagambetov, S. Mahloujifar, K. Chaudhuri, D. Wagner, and C. Guo, “SecAlign: Defending Against Prompt Injection with Preference Optimization”, Oct. 07, 2024. doi: 10.1145/3719027.3744836.
Chen, X., Liu, C., Li, B., Lu, K. & Song, D.(2017). Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning. arXiv.Chen, X., Liu, C., Li, B., Lu, K., & Song, D. (2017). Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning. In arXiv. https://arxiv.org/abs/1712.05526Chen, X., C. Liu, B. Li, K. Lu, and D. Song. 2017. “Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning”. In arXiv. Preprint, December 15. https://arxiv.org/abs/1712.05526.Chen, X., et al. “Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning”. arXiv, 15 Dec. 2017, https://arxiv.org/abs/1712.05526.Chen, X., Liu, C., Li, B., Lu, K. & Song, D. Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning. arXiv Preprint at https://arxiv.org/abs/1712.05526 (2017).X. Chen, C. Liu, B. Li, K. Lu, and D. Song, “Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning”, Dec. 15, 2017. [Online]. Available: https://arxiv.org/abs/1712.05526
Chen, Y. et al.(2025). Reasoning Models Don't Always Say What They Think. arXiv.Chen, Y., Benton, J., Radhakrishnan, A., Uesato, J., Denison, C., Schulman, J., Somani, A., Hase, P., Wagner, M., Roger, F., Mikulik, V., Bowman, S. R., Leike, J., Kaplan, J., & Perez, E. (2025). Reasoning Models Don't Always Say What They Think. In arXiv. https://arxiv.org/abs/2505.05410Chen, Y., J. Benton, A. Radhakrishnan, et al. 2025. “Reasoning Models Don't Always Say What They Think”. In arXiv. Preprint, May 8. https://arxiv.org/abs/2505.05410.Chen, Y., et al. “Reasoning Models Don't Always Say What They Think”. arXiv, 8 May 2025, https://arxiv.org/abs/2505.05410.Chen, Y. et al. Reasoning Models Don't Always Say What They Think. arXiv Preprint at https://arxiv.org/abs/2505.05410 (2025).Y. Chen et al., “Reasoning Models Don't Always Say What They Think”, May 08, 2025. [Online]. Available: https://arxiv.org/abs/2505.05410
Chen, Y., Shen, C., Shen, Y., Wang, C. & Zhang, Y.(2022). Amplifying Membership Exposure via Data Poisoning. arXiv.Chen, Y., Shen, C., Shen, Y., Wang, C., & Zhang, Y. (2022). Amplifying Membership Exposure via Data Poisoning. In arXiv. https://arxiv.org/abs/2211.00463Chen, Y., C. Shen, Y. Shen, C. Wang, and Y. Zhang. 2022. “Amplifying Membership Exposure via Data Poisoning”. In arXiv. Preprint, November 1. https://arxiv.org/abs/2211.00463.Chen, Y., et al. “Amplifying Membership Exposure via Data Poisoning”. arXiv, 1 Nov. 2022, https://arxiv.org/abs/2211.00463.Chen, Y., Shen, C., Shen, Y., Wang, C. & Zhang, Y. Amplifying Membership Exposure via Data Poisoning. arXiv Preprint at https://arxiv.org/abs/2211.00463 (2022).Y. Chen, C. Shen, Y. Shen, C. Wang, and Y. Zhang, “Amplifying Membership Exposure via Data Poisoning”, Nov. 01, 2022. [Online]. Available: https://arxiv.org/abs/2211.00463
Cheng et al.(2024). State of the AI Regulatory Landscape. Convergence Analysis.Cheng et al. (2024). State of the AI Regulatory Landscape. Convergence Analysis. https://convergenceanalysis.org/ai-regulatory-landscape/homeCheng et al. 2024. “State of the AI Regulatory Landscape”. Convergence Analysis. https://convergenceanalysis.org/ai-regulatory-landscape/home.Cheng et al. “State of the AI Regulatory Landscape”. Convergence Analysis, 2024, https://convergenceanalysis.org/ai-regulatory-landscape/home.Cheng et al. State of the AI Regulatory Landscape. Convergence Analysis https://convergenceanalysis.org/ai-regulatory-landscape/home (2024).Cheng et al., “State of the AI Regulatory Landscape”, Convergence Analysis. [Online]. Available: https://convergenceanalysis.org/ai-regulatory-landscape/home
Chollet(2023). François Chollet (@fchollet) on X. X (formerly Twitter).Chollet. (2023, December 16). François Chollet (@fchollet) on X. X (formerly Twitter). https://x.com/fchollet/status/1736079054313574578?s=20Chollet. 2023. “François Chollet (@fchollet) on X”. X (formerly Twitter), December 16. https://x.com/fchollet/status/1736079054313574578?s=20.Chollet. “François Chollet (@fchollet) on X”. X (formerly Twitter), 16 Dec. 2023, https://x.com/fchollet/status/1736079054313574578?s=20.Chollet. François Chollet (@fchollet) on X. X (formerly Twitter) https://x.com/fchollet/status/1736079054313574578?s=20 (2023).Chollet, “François Chollet (@fchollet) on X”, X (formerly Twitter). [Online]. Available: https://x.com/fchollet/status/1736079054313574578?s=20
Chollet(2024). Francois Chollet, Mike Knoop - LLMs won’t lead to AGI - $1,000,000 Prize to find true solution.Chollet. (2024). Francois Chollet, Mike Knoop - LLMs won’t lead to AGI - $1,000,000 Prize to find true solution. https://dwarkeshpatel.com/p/francois-cholletChollet. 2024. “Francois Chollet, Mike Knoop - LLMs Won’t Lead to AGI - $1,000,000 Prize to Find True Solution”. https://dwarkeshpatel.com/p/francois-chollet.Chollet. Francois Chollet, Mike Knoop - LLMs Won’t Lead to AGI - $1,000,000 Prize to Find True Solution. 2024, https://dwarkeshpatel.com/p/francois-chollet.Chollet. Francois Chollet, Mike Knoop - LLMs won’t lead to AGI - $1,000,000 Prize to find true solution. https://dwarkeshpatel.com/p/francois-chollet (2024).Chollet, “Francois Chollet, Mike Knoop - LLMs won’t lead to AGI - $1,000,000 Prize to find true solution”. [Online]. Available: https://dwarkeshpatel.com/p/francois-chollet
Chollet, F.(2019). On the Measure of Intelligence. arXiv.Chollet, F. (2019). On the Measure of Intelligence. In arXiv. https://arxiv.org/abs/1911.01547Chollet, F. 2019. “On the Measure of Intelligence”. In arXiv. Preprint, November 5. https://arxiv.org/abs/1911.01547.Chollet, F. “On the Measure of Intelligence”. arXiv, 5 Nov. 2019, https://arxiv.org/abs/1911.01547.Chollet, F. On the Measure of Intelligence. arXiv Preprint at https://arxiv.org/abs/1911.01547 (2019).F. Chollet, “On the Measure of Intelligence”, Nov. 05, 2019. [Online]. Available: https://arxiv.org/abs/1911.01547
Chollet, F., Knoop, M., Kamradt, G. & Landers, B.(2024). ARC Prize 2024: Technical Report. arXiv.Chollet, F., Knoop, M., Kamradt, G., & Landers, B. (2024). ARC Prize 2024: Technical Report. In arXiv. https://arxiv.org/abs/2412.04604Chollet, F., M. Knoop, G. Kamradt, and B. Landers. 2024. “ARC Prize 2024: Technical Report”. In arXiv. Preprint, December 5. https://arxiv.org/abs/2412.04604.Chollet, F., et al. “ARC Prize 2024: Technical Report”. arXiv, 5 Dec. 2024, https://arxiv.org/abs/2412.04604.Chollet, F., Knoop, M., Kamradt, G. & Landers, B. ARC Prize 2024: Technical Report. arXiv Preprint at https://arxiv.org/abs/2412.04604 (2024).F. Chollet, M. Knoop, G. Kamradt, and B. Landers, “ARC Prize 2024: Technical Report”, Dec. 05, 2024. [Online]. Available: https://arxiv.org/abs/2412.04604
Christiano(2018). Takeoff speeds. The sideways view.Christiano. (2018, February 24). Takeoff speeds. The Sideways View. https://sideways-view.com/2018/02/24/takeoff-speedsChristiano. 2018. “Takeoff Speeds”. The Sideways View, February 24. https://sideways-view.com/2018/02/24/takeoff-speeds.Christiano. “Takeoff Speeds”. The Sideways View, 24 Feb. 2018, https://sideways-view.com/2018/02/24/takeoff-speeds.Christiano. Takeoff speeds. The sideways view https://sideways-view.com/2018/02/24/takeoff-speeds (2018).Christiano, “Takeoff speeds”, The sideways view. [Online]. Available: https://sideways-view.com/2018/02/24/takeoff-speeds
Christiano(2019). Paul Christiano: Current Work in AI Alignment. Effective Altruism.Christiano. (2019). Paul Christiano: Current Work in AI Alignment. Effective Altruism. https://effectivealtruism.org/articles/paul-christiano-current-work-in-ai-alignmentChristiano. 2019. “Paul Christiano: Current Work in AI Alignment”. Effective Altruism. https://effectivealtruism.org/articles/paul-christiano-current-work-in-ai-alignment.Christiano. “Paul Christiano: Current Work in AI Alignment”. Effective Altruism, 2019, https://effectivealtruism.org/articles/paul-christiano-current-work-in-ai-alignment.Christiano. Paul Christiano: Current Work in AI Alignment. Effective Altruism https://effectivealtruism.org/articles/paul-christiano-current-work-in-ai-alignment (2019).Christiano, “Paul Christiano: Current Work in AI Alignment”, Effective Altruism. [Online]. Available: https://effectivealtruism.org/articles/paul-christiano-current-work-in-ai-alignment
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S. & Amodei, D.(2017). Deep reinforcement learning from human preferences. arXiv.Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. In arXiv. https://arxiv.org/abs/1706.03741Christiano, P., J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei. 2017. “Deep Reinforcement Learning from Human Preferences”. In arXiv. Preprint, June 12. https://arxiv.org/abs/1706.03741.Christiano, P., et al. “Deep Reinforcement Learning from Human Preferences”. arXiv, 12 June 2017, https://arxiv.org/abs/1706.03741.Christiano, P. et al. Deep reinforcement learning from human preferences. arXiv Preprint at https://arxiv.org/abs/1706.03741 (2017).P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences”, Jun. 12, 2017. [Online]. Available: https://arxiv.org/abs/1706.03741
Cihon, P., Maas, M. M. & Kemp, L.(2020). Should Artificial Intelligence Governance be Centralised? Design Lessons from History. arXiv.Cihon, P., Maas, M. M., & Kemp, L. (2020). Should Artificial Intelligence Governance be Centralised? Design Lessons from History. In arXiv. https://arxiv.org/abs/2001.03573Cihon, P., M. M. Maas, and L. Kemp. 2020. “Should Artificial Intelligence Governance Be Centralised? Design Lessons from History”. In arXiv. Preprint, January 10. https://arxiv.org/abs/2001.03573.Cihon, P., et al. “Should Artificial Intelligence Governance Be Centralised? Design Lessons from History”. arXiv, 10 Jan. 2020, https://arxiv.org/abs/2001.03573.Cihon, P., Maas, M. M. & Kemp, L. Should Artificial Intelligence Governance be Centralised? Design Lessons from History. arXiv Preprint at https://arxiv.org/abs/2001.03573 (2020).P. Cihon, M. M. Maas, and L. Kemp, “Should Artificial Intelligence Governance be Centralised? Design Lessons from History”, Jan. 10, 2020. [Online]. Available: https://arxiv.org/abs/2001.03573
Cima, M., Tonnaer, F. & Hauser, M. D.(2010). Psychopaths know right from wrong but don’t care. Social Cognitive and Affective Neuroscience.Cima, M., Tonnaer, F., & Hauser, M. D. (2010). Psychopaths know right from wrong but don’t care. Social Cognitive and Affective Neuroscience, 5, 59. https://doi.org/10.1093/scan/nsp051Cima, M., F. Tonnaer, and M. D. Hauser. 2010. “Psychopaths Know Right from Wrong but Don’t Care”. Social Cognitive and Affective Neuroscience 5 (January): 59. https://doi.org/10.1093/scan/nsp051.Cima, M., et al. “Psychopaths Know Right from Wrong but Don’t Care”. Social Cognitive and Affective Neuroscience, vol. 5, Jan. 2010, p. 59, https://doi.org/10.1093/scan/nsp051.Cima, M., Tonnaer, F. & Hauser, M. D. Psychopaths know right from wrong but don’t care. Social Cognitive and Affective Neuroscience 5, 59 (2010).M. Cima, F. Tonnaer, and M. D. Hauser, “Psychopaths know right from wrong but don’t care”, Social Cognitive and Affective Neuroscience, vol. 5, p. 59, Jan. 2010, doi: 10.1093/scan/nsp051.
CISA(2021). The Attack on Colonial Pipeline: What We’ve Learned & What We’ve Done Over the Past Two Years. Cybersecurity and Infrastructure Security Agency CISA.CISA. (2021). The Attack on Colonial Pipeline: What We’ve Learned & What We’ve Done Over the Past Two Years. Cybersecurity and Infrastructure Security Agency CISA. https://cisa.gov/news-events/news/attack-colonial-pipeline-what-weve-learned-what-weve-done-over-past-two-yearsCISA. 2021. “The Attack on Colonial Pipeline: What We’ve Learned & What We’ve Done Over the Past Two Years”. Cybersecurity and Infrastructure Security Agency CISA. https://cisa.gov/news-events/news/attack-colonial-pipeline-what-weve-learned-what-weve-done-over-past-two-years.CISA. “The Attack on Colonial Pipeline: What We’ve Learned & What We’ve Done Over the Past Two Years”. Cybersecurity and Infrastructure Security Agency CISA, 2021, https://cisa.gov/news-events/news/attack-colonial-pipeline-what-weve-learned-what-weve-done-over-past-two-years.CISA. The Attack on Colonial Pipeline: What We’ve Learned & What We’ve Done Over the Past Two Years. Cybersecurity and Infrastructure Security Agency CISA https://cisa.gov/news-events/news/attack-colonial-pipeline-what-weve-learned-what-weve-done-over-past-two-years (2021).CISA, “The Attack on Colonial Pipeline: What We’ve Learned & What We’ve Done Over the Past Two Years”, Cybersecurity and Infrastructure Security Agency CISA. [Online]. Available: https://cisa.gov/news-events/news/attack-colonial-pipeline-what-weve-learned-what-weve-done-over-past-two-years
CISA(2024). CISA and Partners Release Advisory on PRC-sponsored Volt Typhoon Activity and Supplemental Living Off the Land Guidance. Cybersecurity and Infrastructure Security Agency CISA.CISA. (2024). CISA and Partners Release Advisory on PRC-sponsored Volt Typhoon Activity and Supplemental Living Off the Land Guidance. Cybersecurity and Infrastructure Security Agency CISA. https://cisa.gov/news-events/alerts/2024/02/07/cisa-and-partners-release-advisory-prc-sponsored-volt-typhoon-activity-and-supplemental-living-landCISA. 2024. “CISA and Partners Release Advisory on PRC-sponsored Volt Typhoon Activity and Supplemental Living Off the Land Guidance”. Cybersecurity and Infrastructure Security Agency CISA. https://cisa.gov/news-events/alerts/2024/02/07/cisa-and-partners-release-advisory-prc-sponsored-volt-typhoon-activity-and-supplemental-living-land.CISA. “CISA and Partners Release Advisory on PRC-sponsored Volt Typhoon Activity and Supplemental Living Off the Land Guidance”. Cybersecurity and Infrastructure Security Agency CISA, 2024, https://cisa.gov/news-events/alerts/2024/02/07/cisa-and-partners-release-advisory-prc-sponsored-volt-typhoon-activity-and-supplemental-living-land.CISA. CISA and Partners Release Advisory on PRC-sponsored Volt Typhoon Activity and Supplemental Living Off the Land Guidance. Cybersecurity and Infrastructure Security Agency CISA https://cisa.gov/news-events/alerts/2024/02/07/cisa-and-partners-release-advisory-prc-sponsored-volt-typhoon-activity-and-supplemental-living-land (2024).CISA, “CISA and Partners Release Advisory on PRC-sponsored Volt Typhoon Activity and Supplemental Living Off the Land Guidance”, Cybersecurity and Infrastructure Security Agency CISA. [Online]. Available: https://cisa.gov/news-events/alerts/2024/02/07/cisa-and-partners-release-advisory-prc-sponsored-volt-typhoon-activity-and-supplemental-living-land
Cleo Nardo(2023). The Waluigi Effect (mega-post). AI Alignment Forum.Cleo Nardo. (2023, March 3). The Waluigi Effect (mega-post). AI Alignment Forum. https://alignmentforum.org/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-postCleo Nardo. 2023. “The Waluigi Effect (mega-post)”. AI Alignment Forum, March 3. https://alignmentforum.org/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post.Cleo Nardo. “The Waluigi Effect (mega-post)”. AI Alignment Forum, 3 Mar. 2023, https://alignmentforum.org/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post.Cleo Nardo. The Waluigi Effect (mega-post). AI Alignment Forum https://alignmentforum.org/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post (2023).Cleo Nardo, “The Waluigi Effect (mega-post)”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post
Clymer, J., Gabrieli, N., Krueger, D. & Larsen, T.(2024). Safety Cases: How to Justify the Safety of Advanced AI Systems. arXiv.Clymer, J., Gabrieli, N., Krueger, D., & Larsen, T. (2024). Safety Cases: How to Justify the Safety of Advanced AI Systems. In arXiv. https://arxiv.org/abs/2403.10462Clymer, J., N. Gabrieli, D. Krueger, and T. Larsen. 2024. “Safety Cases: How to Justify the Safety of Advanced AI Systems”. In arXiv. Preprint, March 15. https://arxiv.org/abs/2403.10462.Clymer, J., et al. “Safety Cases: How to Justify the Safety of Advanced AI Systems”. arXiv, 15 Mar. 2024, https://arxiv.org/abs/2403.10462.Clymer, J., Gabrieli, N., Krueger, D. & Larsen, T. Safety Cases: How to Justify the Safety of Advanced AI Systems. arXiv Preprint at https://arxiv.org/abs/2403.10462 (2024).J. Clymer, N. Gabrieli, D. Krueger, and T. Larsen, “Safety Cases: How to Justify the Safety of Advanced AI Systems”, Mar. 15, 2024. [Online]. Available: https://arxiv.org/abs/2403.10462
Clymer, J., Juang, C. & Field, S.(2024). Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals. arXiv.Clymer, J., Juang, C., & Field, S. (2024). Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals. In arXiv. https://arxiv.org/abs/2405.05466Clymer, J., C. Juang, and S. Field. 2024. “Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals”. In arXiv. Preprint, May 8. https://arxiv.org/abs/2405.05466.Clymer, J., et al. “Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals”. arXiv, 8 May 2024, https://arxiv.org/abs/2405.05466.Clymer, J., Juang, C. & Field, S. Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals. arXiv Preprint at https://arxiv.org/abs/2405.05466 (2024).J. Clymer, C. Juang, and S. Field, “Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals”, May 08, 2024. [Online]. Available: https://arxiv.org/abs/2405.05466
Cobbe, K. et al.(2021). Training Verifiers to Solve Math Word Problems. arXiv.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., & Schulman, J. (2021). Training Verifiers to Solve Math Word Problems. In arXiv. https://arxiv.org/abs/2110.14168Cobbe, K., V. Kosaraju, M. Bavarian, et al. 2021. “Training Verifiers to Solve Math Word Problems”. In arXiv. Preprint, October 27. https://arxiv.org/abs/2110.14168.Cobbe, K., et al. “Training Verifiers to Solve Math Word Problems”. arXiv, 27 Oct. 2021, https://arxiv.org/abs/2110.14168.Cobbe, K. et al. Training Verifiers to Solve Math Word Problems. arXiv Preprint at https://arxiv.org/abs/2110.14168 (2021).K. Cobbe et al., “Training Verifiers to Solve Math Word Problems”, Oct. 27, 2021. [Online]. Available: https://arxiv.org/abs/2110.14168
Cobbe, K., Klimov, O., Hesse, C., Kim, T. & Schulman, J.(2018). Quantifying Generalization in Reinforcement Learning. arXiv.Cobbe, K., Klimov, O., Hesse, C., Kim, T., & Schulman, J. (2018). Quantifying Generalization in Reinforcement Learning. In arXiv. https://arxiv.org/abs/1812.02341Cobbe, K., O. Klimov, C. Hesse, T. Kim, and J. Schulman. 2018. “Quantifying Generalization in Reinforcement Learning”. In arXiv. Preprint, December 6. https://arxiv.org/abs/1812.02341.Cobbe, K., et al. “Quantifying Generalization in Reinforcement Learning”. arXiv, 6 Dec. 2018, https://arxiv.org/abs/1812.02341.Cobbe, K., Klimov, O., Hesse, C., Kim, T. & Schulman, J. Quantifying Generalization in Reinforcement Learning. arXiv Preprint at https://arxiv.org/abs/1812.02341 (2018).K. Cobbe, O. Klimov, C. Hesse, T. Kim, and J. Schulman, “Quantifying Generalization in Reinforcement Learning”, Dec. 06, 2018. [Online]. Available: https://arxiv.org/abs/1812.02341
Conn(2015). Existential Risk. Future of Life Institute.Conn. (2015, November 16). Existential Risk. Future of Life Institute. https://futureoflife.org/existential-risk/existential-riskConn. 2015. “Existential Risk”. Future of Life Institute, November 16. https://futureoflife.org/existential-risk/existential-risk.Conn. “Existential Risk”. Future of Life Institute, 16 Nov. 2015, https://futureoflife.org/existential-risk/existential-risk.Conn. Existential Risk. Future of Life Institute https://futureoflife.org/existential-risk/existential-risk (2015).Conn, “Existential Risk”, Future of Life Institute. [Online]. Available: https://futureoflife.org/existential-risk/existential-risk
Connor Leahy & Gabriel Alfour(2023). Cognitive Emulation: A Naive AI Safety Proposal. AI Alignment Forum.Connor Leahy, & Gabriel Alfour. (2023, February 25). Cognitive Emulation: A Naive AI Safety Proposal. AI Alignment Forum. https://alignmentforum.org/posts/ngEvKav9w57XrGQnb/cognitive-emulation-a-naive-ai-safety-proposalConnor Leahy, and Gabriel Alfour. 2023. “Cognitive Emulation: A Naive AI Safety Proposal”. AI Alignment Forum, February 25. https://alignmentforum.org/posts/ngEvKav9w57XrGQnb/cognitive-emulation-a-naive-ai-safety-proposal.Connor Leahy, and Gabriel Alfour. “Cognitive Emulation: A Naive AI Safety Proposal”. AI Alignment Forum, 25 Feb. 2023, https://alignmentforum.org/posts/ngEvKav9w57XrGQnb/cognitive-emulation-a-naive-ai-safety-proposal.Connor Leahy & Gabriel Alfour. Cognitive Emulation: A Naive AI Safety Proposal. AI Alignment Forum https://alignmentforum.org/posts/ngEvKav9w57XrGQnb/cognitive-emulation-a-naive-ai-safety-proposal (2023).Connor Leahy and Gabriel Alfour, “Cognitive Emulation: A Naive AI Safety Proposal”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/ngEvKav9w57XrGQnb/cognitive-emulation-a-naive-ai-safety-proposal
Corwin(2002). SL4: AI Boxing.Corwin. (2002). SL4: AI Boxing. Internet Archive (https://web.archive.org/web/20250906234622/http://sl4.org/archive/0207/4935.html). https://sl4.org/archive/0207/4935.htmlCorwin. 2002. “SL4: AI Boxing”. Https://web.archive.org/web/20250906234622/http://sl4.org/archive/0207/4935.html. Internet Archive. https://sl4.org/archive/0207/4935.html.Corwin. SL4: AI Boxing. 2002, Internet Archive, https://web.archive.org/web/20250906234622/http://sl4.org/archive/0207/4935.html, https://sl4.org/archive/0207/4935.html.Corwin. SL4: AI Boxing. https://sl4.org/archive/0207/4935.html (2002).Corwin, “SL4: AI Boxing”. Accessed: Sep. 06, 2025. [Online]. Available: https://sl4.org/archive/0207/4935.html
Cotra(2021). Why AI alignment could be hard with modern deep learning. Cold Takes.Cotra. (2021, September 21). Why AI alignment could be hard with modern deep learning. Cold Takes. https://cold-takes.com/why-ai-alignment-could-be-hard-with-modern-deep-learningCotra. 2021. “Why AI Alignment Could Be Hard with Modern Deep Learning”. Cold Takes, September 21. https://cold-takes.com/why-ai-alignment-could-be-hard-with-modern-deep-learning.Cotra. “Why AI Alignment Could Be Hard with Modern Deep Learning”. Cold Takes, 21 Sept. 2021, https://cold-takes.com/why-ai-alignment-could-be-hard-with-modern-deep-learning.Cotra. Why AI alignment could be hard with modern deep learning. Cold Takes https://cold-takes.com/why-ai-alignment-could-be-hard-with-modern-deep-learning (2021).Cotra, “Why AI alignment could be hard with modern deep learning”, Cold Takes. [Online]. Available: https://cold-takes.com/why-ai-alignment-could-be-hard-with-modern-deep-learning
Cotra(2023). AIs accelerating AI research.Cotra. (2023). AIs accelerating AI research. https://www.planned-obsolescence.org/ais-accelerating-ai-researchCotra. 2023. “AIs Accelerating AI Research”. https://www.planned-obsolescence.org/ais-accelerating-ai-research.Cotra. AIs Accelerating AI Research. 2023, https://www.planned-obsolescence.org/ais-accelerating-ai-research.Cotra. AIs accelerating AI research. https://www.planned-obsolescence.org/ais-accelerating-ai-research (2023).Cotra, “AIs accelerating AI research”. [Online]. Available: https://www.planned-obsolescence.org/ais-accelerating-ai-research
Cotra(2023). Language models surprised us.Cotra. (2023). Language models surprised us. https://www.planned-obsolescence.org/language-models-surprised-usCotra. 2023. “Language Models Surprised Us”. https://www.planned-obsolescence.org/language-models-surprised-us.Cotra. Language Models Surprised Us. 2023, https://www.planned-obsolescence.org/language-models-surprised-us.Cotra. Language models surprised us. https://www.planned-obsolescence.org/language-models-surprised-us (2023).Cotra, “Language models surprised us”. [Online]. Available: https://www.planned-obsolescence.org/language-models-surprised-us
Cotra(2023). Scale, schlep, and systems.Cotra. (2023). Scale, schlep, and systems. https://www.planned-obsolescence.org/scale-schlep-and-systemsCotra. 2023. “Scale, Schlep, and Systems”. https://www.planned-obsolescence.org/scale-schlep-and-systems.Cotra. Scale, Schlep, and Systems. 2023, https://www.planned-obsolescence.org/scale-schlep-and-systems.Cotra. Scale, schlep, and systems. https://www.planned-obsolescence.org/scale-schlep-and-systems (2023).Cotra, “Scale, schlep, and systems”. [Online]. Available: https://www.planned-obsolescence.org/scale-schlep-and-systems
Cottier et al.(2024). How much does it cost to train frontier AI models?. Epoch AI.Cottier et al. (2024). How much does it cost to train frontier AI models?. Epoch AI. https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-modelsCottier et al. 2024. “How Much Does It Cost to Train Frontier AI Models?”. Epoch AI. https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models.Cottier et al. “How Much Does It Cost to Train Frontier AI Models?”. Epoch AI, 2024, https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models.Cottier et al. How much does it cost to train frontier AI models?. Epoch AI https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models (2024).Cottier et al., “How much does it cost to train frontier AI models?”, Epoch AI. [Online]. Available: https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models
Creswell, A., Shanahan, M. & Higgins, I.(2022). Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning. arXiv.Creswell, A., Shanahan, M., & Higgins, I. (2022). Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning. In arXiv. https://arxiv.org/abs/2205.09712Creswell, A., M. Shanahan, and I. Higgins. 2022. “Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning”. In arXiv. Preprint, May 19. https://arxiv.org/abs/2205.09712.Creswell, A., et al. “Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning”. arXiv, 19 May 2022, https://arxiv.org/abs/2205.09712.Creswell, A., Shanahan, M. & Higgins, I. Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning. arXiv Preprint at https://arxiv.org/abs/2205.09712 (2022).A. Creswell, M. Shanahan, and I. Higgins, “Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning”, May 19, 2022. [Online]. Available: https://arxiv.org/abs/2205.09712
Critch, A. & Russell, S.(2023). TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI. arXiv.Critch, A., & Russell, S. (2023). TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI. In arXiv. https://arxiv.org/abs/2306.06924Critch, A., and S. Russell. 2023. “TASRA: A Taxonomy and Analysis of Societal-Scale Risks from AI”. In arXiv. Preprint, June 12. https://arxiv.org/abs/2306.06924.Critch, A., and S. Russell. “TASRA: A Taxonomy and Analysis of Societal-Scale Risks from AI”. arXiv, 12 June 2023, https://arxiv.org/abs/2306.06924.Critch, A. & Russell, S. TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI. arXiv Preprint at https://arxiv.org/abs/2306.06924 (2023).A. Critch and S. Russell, “TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI”, Jun. 12, 2023. [Online]. Available: https://arxiv.org/abs/2306.06924
CrowdStrike(2024). External Technical Root Cause Analysis — Channel File 291.CrowdStrike. (2024). External Technical Root Cause Analysis — Channel File 291. CrowdStrike. https://crowdstrike.com/wp-content/uploads/2024/08/Channel-File-291-Incident-Root-Cause-Analysis-08.06.2024.pdfCrowdStrike. 2024. External Technical Root Cause Analysis — Channel File 291. CrowdStrike. https://crowdstrike.com/wp-content/uploads/2024/08/Channel-File-291-Incident-Root-Cause-Analysis-08.06.2024.pdf.CrowdStrike. External Technical Root Cause Analysis — Channel File 291. CrowdStrike, 6 Aug. 2024, https://crowdstrike.com/wp-content/uploads/2024/08/Channel-File-291-Incident-Root-Cause-Analysis-08.06.2024.pdf.CrowdStrike. External Technical Root Cause Analysis — Channel File 291. https://crowdstrike.com/wp-content/uploads/2024/08/Channel-File-291-Incident-Root-Cause-Analysis-08.06.2024.pdf (2024).CrowdStrike, “External Technical Root Cause Analysis — Channel File 291”, CrowdStrike, Aug. 2024. [Online]. Available: https://crowdstrike.com/wp-content/uploads/2024/08/Channel-File-291-Incident-Root-Cause-Analysis-08.06.2024.pdf
Cullen O’Keefe, Peter Cihon, Ben Garfinkel, Carrick Flynn, Jade Leung & and Allan Dafoe(2020). The Windfall Clause: Distributing the Benefits of AI for the Common Good.Cullen O’Keefe, Peter Cihon, Ben Garfinkel, Carrick Flynn, Jade Leung, & and Allan Dafoe. (2020, January 30). The Windfall Clause: Distributing the Benefits of AI for the Common Good. https://governance.ai/research-paper/the-windfall-clause-distributing-the-benefits-of-ai-for-the-common-goodCullen O’Keefe, Peter Cihon, Ben Garfinkel, Carrick Flynn, Jade Leung, and and Allan Dafoe. 2020. “The Windfall Clause: Distributing the Benefits of AI for the Common Good”. January 30. https://governance.ai/research-paper/the-windfall-clause-distributing-the-benefits-of-ai-for-the-common-good.Cullen O’Keefe, et al. The Windfall Clause: Distributing the Benefits of AI for the Common Good. 30 Jan. 2020, https://governance.ai/research-paper/the-windfall-clause-distributing-the-benefits-of-ai-for-the-common-good.Cullen O’Keefe et al. The Windfall Clause: Distributing the Benefits of AI for the Common Good. https://governance.ai/research-paper/the-windfall-clause-distributing-the-benefits-of-ai-for-the-common-good (2020).Cullen O’Keefe, Peter Cihon, Ben Garfinkel, Carrick Flynn, Jade Leung, and and Allan Dafoe, “The Windfall Clause: Distributing the Benefits of AI for the Common Good”. [Online]. Available: https://governance.ai/research-paper/the-windfall-clause-distributing-the-benefits-of-ai-for-the-common-good
Cunha, P. R. & Estima, J.(2023). Navigating the Landscape of AI Ethics and Responsibility. Lecture Notes in Computer Science.Cunha, P. R., & Estima, J. (2023). Navigating the Landscape of AI Ethics and Responsibility. In Lecture Notes in Computer Science. https://doi.org/10.1007/978-3-031-49008-8_8Cunha, P. R., and J. Estima. 2023. “Navigating the Landscape of AI Ethics and Responsibility”. In Lecture Notes in Computer Science. https://doi.org/10.1007/978-3-031-49008-8_8.Cunha, P. R., and J. Estima. “Navigating the Landscape of AI Ethics and Responsibility”. Lecture Notes in Computer Science, 2023, https://doi.org/10.1007/978-3-031-49008-8_8.Cunha, P. R. & Estima, J. Navigating the Landscape of AI Ethics and Responsibility. Lecture Notes in Computer Science (2023) doi:10.1007/978-3-031-49008-8_8.P. R. Cunha and J. Estima, “Navigating the Landscape of AI Ethics and Responsibility”, Lecture Notes in Computer Science. 2023. doi: 10.1007/978-3-031-49008-8_8.
Cunningham, H., Ewart, A., Riggs, L., Huben, R. & Sharkey, L.(2023). Sparse Autoencoders Find Highly Interpretable Features in Language Models. arXiv.Cunningham, H., Ewart, A., Riggs, L., Huben, R., & Sharkey, L. (2023). Sparse Autoencoders Find Highly Interpretable Features in Language Models. In arXiv. https://arxiv.org/abs/2309.08600Cunningham, H., A. Ewart, L. Riggs, R. Huben, and L. Sharkey. 2023. “Sparse Autoencoders Find Highly Interpretable Features in Language Models”. In arXiv. Preprint, September 15. https://arxiv.org/abs/2309.08600.Cunningham, H., et al. “Sparse Autoencoders Find Highly Interpretable Features in Language Models”. arXiv, 15 Sept. 2023, https://arxiv.org/abs/2309.08600.Cunningham, H., Ewart, A., Riggs, L., Huben, R. & Sharkey, L. Sparse Autoencoders Find Highly Interpretable Features in Language Models. arXiv Preprint at https://arxiv.org/abs/2309.08600 (2023).H. Cunningham, A. Ewart, L. Riggs, R. Huben, and L. Sharkey, “Sparse Autoencoders Find Highly Interpretable Features in Language Models”, Sep. 15, 2023. [Online]. Available: https://arxiv.org/abs/2309.08600
Dafoe(2018). AI Governance: A Research Agenda | GovAI.Dafoe. (2018). AI Governance: A Research Agenda | GovAI. https://governance.ai/research-paper/agendaDafoe. 2018. “AI Governance: A Research Agenda | GovAI”. https://governance.ai/research-paper/agenda.Dafoe. AI Governance: A Research Agenda | GovAI. 2018, https://governance.ai/research-paper/agenda.Dafoe. AI Governance: A Research Agenda | GovAI. https://governance.ai/research-paper/agenda (2018).Dafoe, “AI Governance: A Research Agenda | GovAI”. [Online]. Available: https://governance.ai/research-paper/agenda
Dafoe, A.(2024). AI Governance: Overview and Theoretical Lenses. The Oxford Handbook of AI Governance.Dafoe, A. (2024). AI Governance: Overview and Theoretical Lenses. In The Oxford Handbook of AI Governance (pp. 21–44). Oxford University Press. https://doi.org/10.1093/oxfordhb/9780197579329.013.2Dafoe, A. 2024. “AI Governance: Overview and Theoretical Lenses”. In The Oxford Handbook of AI Governance. Oxford University Press. https://doi.org/10.1093/oxfordhb/9780197579329.013.2.Dafoe, A. “AI Governance: Overview and Theoretical Lenses”. The Oxford Handbook of AI Governance, Oxford University Press, 2024, pp. 21–44, https://doi.org/10.1093/oxfordhb/9780197579329.013.2.Dafoe, A. AI Governance: Overview and Theoretical Lenses. in The Oxford Handbook of AI Governance 21–44 (Oxford University Press, 2024). doi:10.1093/oxfordhb/9780197579329.013.2.A. Dafoe, “AI Governance: Overview and Theoretical Lenses”, in The Oxford Handbook of AI Governance, Oxford University Press, 2024, pp. 21–44. doi: 10.1093/oxfordhb/9780197579329.013.2.
Dalrymple(2022). Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems. arXiv.org.Dalrymple. (2022). Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems. arXiv.org. https://www.arxiv.org/abs/2405.06624Dalrymple. 2022. “Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems”. arXiv.org. https://www.arxiv.org/abs/2405.06624.Dalrymple. “Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems”. arXiv.org, 2022, https://www.arxiv.org/abs/2405.06624.Dalrymple. Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems. arXiv.org https://www.arxiv.org/abs/2405.06624 (2022).Dalrymple, “Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2405.06624
Daniel Kokotajlo(2021). Interlude: Agents as Automobiles. AI Alignment Forum.Daniel Kokotajlo. (2021, December 14). Interlude: Agents as Automobiles. AI Alignment Forum. https://alignmentforum.org/posts/cxkwQmys6mCB6bjDA/interlude-agents-as-automobilesDaniel Kokotajlo. 2021. “Interlude: Agents as Automobiles”. AI Alignment Forum, December 14. https://alignmentforum.org/posts/cxkwQmys6mCB6bjDA/interlude-agents-as-automobiles.Daniel Kokotajlo. “Interlude: Agents as Automobiles”. AI Alignment Forum, 14 Dec. 2021, https://alignmentforum.org/posts/cxkwQmys6mCB6bjDA/interlude-agents-as-automobiles.Daniel Kokotajlo. Interlude: Agents as Automobiles. AI Alignment Forum https://alignmentforum.org/posts/cxkwQmys6mCB6bjDA/interlude-agents-as-automobiles (2021).Daniel Kokotajlo, “Interlude: Agents as Automobiles”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/cxkwQmys6mCB6bjDA/interlude-agents-as-automobiles
Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland & Romeo Dean(2025). AI 2027.Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, & Romeo Dean. (2025). AI 2027. https://ai-2027.comDaniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, and Romeo Dean. 2025. “AI 2027”. https://ai-2027.com.Daniel Kokotajlo, et al. AI 2027. 2025, https://ai-2027.com.Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland & Romeo Dean. AI 2027. https://ai-2027.com (2025).Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, and Romeo Dean, “AI 2027”. [Online]. Available: https://ai-2027.com
Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland & Romeo Dean(2025). Security Forecast. AI 2027.Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, & Romeo Dean. (2025). Security Forecast. AI 2027. https://ai-2027.com/research/security-forecastDaniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, and Romeo Dean. 2025. “Security Forecast”. AI 2027. https://ai-2027.com/research/security-forecast.Daniel Kokotajlo, et al. “Security Forecast”. AI 2027, 2025, https://ai-2027.com/research/security-forecast.Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland & Romeo Dean. Security Forecast. AI 2027 https://ai-2027.com/research/security-forecast (2025).Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, and Romeo Dean, “Security Forecast”, AI 2027. [Online]. Available: https://ai-2027.com/research/security-forecast
David Manheim & Scott Garrabrant(2018). Categorizing Variants of Goodhart's Law. arXiv.David Manheim, & Scott Garrabrant. (2018). Categorizing Variants of Goodhart's Law. In arXiv. https://arxiv.org/abs/1803.04585David Manheim, and Scott Garrabrant. 2018. “Categorizing Variants of Goodhart's Law”. In arXiv. Preprint, March 13. https://arxiv.org/abs/1803.04585.David Manheim, and Scott Garrabrant. “Categorizing Variants of Goodhart's Law”. arXiv, 13 Mar. 2018, https://arxiv.org/abs/1803.04585.David Manheim & Scott Garrabrant. Categorizing Variants of Goodhart's Law. arXiv Preprint at https://arxiv.org/abs/1803.04585 (2018).David Manheim and Scott Garrabrant, “Categorizing Variants of Goodhart's Law”, Mar. 13, 2018. [Online]. Available: https://arxiv.org/abs/1803.04585
David Rein et al.(2023). GPQA: A Graduate-Level Google-Proof Q&A Benchmark. arXiv.David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, & Samuel R. Bowman. (2023). GPQA: A Graduate-Level Google-Proof Q&A Benchmark. In arXiv. https://arxiv.org/abs/2311.12022David Rein, Betty Li Hou, Asa Cooper Stickland, et al. 2023. “GPQA: A Graduate-Level Google-Proof Q&A Benchmark”. In arXiv. Preprint, November 20. https://arxiv.org/abs/2311.12022.David Rein, et al. “GPQA: A Graduate-Level Google-Proof Q&A Benchmark”. arXiv, 20 Nov. 2023, https://arxiv.org/abs/2311.12022.David Rein et al. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. arXiv Preprint at https://arxiv.org/abs/2311.12022 (2023).David Rein et al., “GPQA: A Graduate-Level Google-Proof Q&A Benchmark”, Nov. 20, 2023. [Online]. Available: https://arxiv.org/abs/2311.12022
Davidson(2024). What a Compute-Centric Framework Says About Takeoff Speeds. Coefficient Giving.Davidson. (2024). What a Compute-Centric Framework Says About Takeoff Speeds. Coefficient Giving. https://openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speedsDavidson. 2024. “What a Compute-Centric Framework Says About Takeoff Speeds”. Coefficient Giving. https://openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds.Davidson. “What a Compute-Centric Framework Says About Takeoff Speeds”. Coefficient Giving, 2024, https://openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds.Davidson. What a Compute-Centric Framework Says About Takeoff Speeds. Coefficient Giving https://openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds (2024).Davidson, “What a Compute-Centric Framework Says About Takeoff Speeds”, Coefficient Giving. [Online]. Available: https://openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds
Davidson, T., Denain, J., Villalobos, P. & Bas, G.(2023). AI capabilities can be significantly improved without expensive retraining. arXiv.Davidson, T., Denain, J.-S., Villalobos, P., & Bas, G. (2023). AI capabilities can be significantly improved without expensive retraining. In arXiv. https://arxiv.org/abs/2312.07413Davidson, T., J.-S. Denain, P. Villalobos, and G. Bas. 2023. “AI Capabilities Can Be Significantly Improved Without Expensive Retraining”. In arXiv. Preprint, December 12. https://arxiv.org/abs/2312.07413.Davidson, T., et al. “AI Capabilities Can Be Significantly Improved Without Expensive Retraining”. arXiv, 12 Dec. 2023, https://arxiv.org/abs/2312.07413.Davidson, T., Denain, J.-S., Villalobos, P. & Bas, G. AI capabilities can be significantly improved without expensive retraining. arXiv Preprint at https://arxiv.org/abs/2312.07413 (2023).T. Davidson, J.-S. Denain, P. Villalobos, and G. Bas, “AI capabilities can be significantly improved without expensive retraining”, Dec. 12, 2023. [Online]. Available: https://arxiv.org/abs/2312.07413
DavidW(2023). Deceptive Alignment is <1% Likely by Default. AI Alignment Forum.DavidW. (2023, February 21). Deceptive Alignment is <1% Likely by Default. AI Alignment Forum. https://alignmentforum.org/s/pvoxjtCbkcweBLn7j/p/RTkatYxJWvXR4QbydDavidW. 2023. “Deceptive Alignment Is <1% Likely by Default”. AI Alignment Forum, February 21. https://alignmentforum.org/s/pvoxjtCbkcweBLn7j/p/RTkatYxJWvXR4Qbyd.DavidW. “Deceptive Alignment Is <1% Likely by Default”. AI Alignment Forum, 21 Feb. 2023, https://alignmentforum.org/s/pvoxjtCbkcweBLn7j/p/RTkatYxJWvXR4Qbyd.DavidW. Deceptive Alignment is <1% Likely by Default. AI Alignment Forum https://alignmentforum.org/s/pvoxjtCbkcweBLn7j/p/RTkatYxJWvXR4Qbyd (2023).DavidW, “Deceptive Alignment is <1% Likely by Default”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/s/pvoxjtCbkcweBLn7j/p/RTkatYxJWvXR4Qbyd
DeepMind(2016). AlphaGo. Google DeepMind.DeepMind. (2016). AlphaGo. Google DeepMind. https://deepmind.com/research/highlighted-research/alphagoDeepMind. 2016. “AlphaGo”. Google DeepMind. https://deepmind.com/research/highlighted-research/alphago.DeepMind. “AlphaGo”. Google DeepMind, 2016, https://deepmind.com/research/highlighted-research/alphago.DeepMind. AlphaGo. Google DeepMind https://deepmind.com/research/highlighted-research/alphago (2016).DeepMind, “AlphaGo”, Google DeepMind. [Online]. Available: https://deepmind.com/research/highlighted-research/alphago
DeepMind(2018). AlphaZero: Shedding new light on chess, shogi, and Go. Google DeepMind.DeepMind. (2018). AlphaZero: Shedding new light on chess, shogi, and Go. Google DeepMind. https://deepmind.com/blog/alphazero-shedding-new-light-on-chess-shogi-and-goDeepMind. 2018. “AlphaZero: Shedding New Light on Chess, Shogi, and Go”. Google DeepMind. https://deepmind.com/blog/alphazero-shedding-new-light-on-chess-shogi-and-go.DeepMind. “AlphaZero: Shedding New Light on Chess, Shogi, and Go”. Google DeepMind, 2018, https://deepmind.com/blog/alphazero-shedding-new-light-on-chess-shogi-and-go.DeepMind. AlphaZero: Shedding new light on chess, shogi, and Go. Google DeepMind https://deepmind.com/blog/alphazero-shedding-new-light-on-chess-shogi-and-go (2018).DeepMind, “AlphaZero: Shedding new light on chess, shogi, and Go”, Google DeepMind. [Online]. Available: https://deepmind.com/blog/alphazero-shedding-new-light-on-chess-shogi-and-go
DeepMind(2018). Scalable agent alignment via reward modeling. Medium.DeepMind. (2018, November 20). Scalable agent alignment via reward modeling. Medium. https://deepmindsafetyresearch.medium.com/scalable-agent-alignment-via-reward-modeling-bf4ab06dfd84DeepMind. 2018. “Scalable Agent Alignment via Reward Modeling”. Medium, November 20. https://deepmindsafetyresearch.medium.com/scalable-agent-alignment-via-reward-modeling-bf4ab06dfd84.DeepMind. “Scalable Agent Alignment via Reward Modeling”. Medium, 20 Nov. 2018, https://deepmindsafetyresearch.medium.com/scalable-agent-alignment-via-reward-modeling-bf4ab06dfd84.DeepMind. Scalable agent alignment via reward modeling. Medium https://deepmindsafetyresearch.medium.com/scalable-agent-alignment-via-reward-modeling-bf4ab06dfd84 (2018).DeepMind, “Scalable agent alignment via reward modeling”, Medium. [Online]. Available: https://deepmindsafetyresearch.medium.com/scalable-agent-alignment-via-reward-modeling-bf4ab06dfd84
DeepMind(2023). FunSearch: Making new discoveries in mathematical sciences using Large Language Models. Google DeepMind.DeepMind. (2023). FunSearch: Making new discoveries in mathematical sciences using Large Language Models. Google DeepMind. https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-modelsDeepMind. 2023. “FunSearch: Making New Discoveries in Mathematical Sciences Using Large Language Models”. Google DeepMind. https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models.DeepMind. “FunSearch: Making New Discoveries in Mathematical Sciences Using Large Language Models”. Google DeepMind, 2023, https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models.DeepMind. FunSearch: Making new discoveries in mathematical sciences using Large Language Models. Google DeepMind https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models (2023).DeepMind, “FunSearch: Making new discoveries in mathematical sciences using Large Language Models”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models
DeepMind(2023). Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals. Medium.DeepMind. (2023, March 24). Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals. Medium. https://deepmindsafetyresearch.medium.com/goal-misgeneralisation-why-correct-specifications-arent-enough-for-correct-goals-cf96ebc60924DeepMind. 2023. “Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals”. Medium, March 24. https://deepmindsafetyresearch.medium.com/goal-misgeneralisation-why-correct-specifications-arent-enough-for-correct-goals-cf96ebc60924.DeepMind. “Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals”. Medium, 24 Mar. 2023, https://deepmindsafetyresearch.medium.com/goal-misgeneralisation-why-correct-specifications-arent-enough-for-correct-goals-cf96ebc60924.DeepMind. Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals. Medium https://deepmindsafetyresearch.medium.com/goal-misgeneralisation-why-correct-specifications-arent-enough-for-correct-goals-cf96ebc60924 (2023).DeepMind, “Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals”, Medium. [Online]. Available: https://deepmindsafetyresearch.medium.com/goal-misgeneralisation-why-correct-specifications-arent-enough-for-correct-goals-cf96ebc60924
DeepMind(2024). Exploring institutions for global AI governance. Google DeepMind.DeepMind. (2024). Exploring institutions for global AI governance. Google DeepMind. https://deepmind.google/discover/blog/exploring-institutions-for-global-ai-governanceDeepMind. 2024. “Exploring Institutions for Global AI Governance”. Google DeepMind. https://deepmind.google/discover/blog/exploring-institutions-for-global-ai-governance.DeepMind. “Exploring Institutions for Global AI Governance”. Google DeepMind, 2024, https://deepmind.google/discover/blog/exploring-institutions-for-global-ai-governance.DeepMind. Exploring institutions for global AI governance. Google DeepMind https://deepmind.google/discover/blog/exploring-institutions-for-global-ai-governance (2024).DeepMind, “Exploring institutions for global AI governance”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/exploring-institutions-for-global-ai-governance
DeepMind(2024). How AlphaChip transformed computer chip design. Google DeepMind.DeepMind. (2024). How AlphaChip transformed computer chip design. Google DeepMind. https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-designDeepMind. 2024. “How AlphaChip Transformed Computer Chip Design”. Google DeepMind. https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-design.DeepMind. “How AlphaChip Transformed Computer Chip Design”. Google DeepMind, 2024, https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-design.DeepMind. How AlphaChip transformed computer chip design. Google DeepMind https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-design (2024).DeepMind, “How AlphaChip transformed computer chip design”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-design
DeepMind(2024). Introducing the Frontier Safety Framework. Google DeepMind.DeepMind. (2024). Introducing the Frontier Safety Framework. Google DeepMind. https://deepmind.google/discover/blog/introducing-the-frontier-safety-frameworkDeepMind. 2024. “Introducing the Frontier Safety Framework”. Google DeepMind. https://deepmind.google/discover/blog/introducing-the-frontier-safety-framework.DeepMind. “Introducing the Frontier Safety Framework”. Google DeepMind, 2024, https://deepmind.google/discover/blog/introducing-the-frontier-safety-framework.DeepMind. Introducing the Frontier Safety Framework. Google DeepMind https://deepmind.google/discover/blog/introducing-the-frontier-safety-framework (2024).DeepMind, “Introducing the Frontier Safety Framework”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/introducing-the-frontier-safety-framework
DeepMind(2025). AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. Google DeepMind.DeepMind. (2025). AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. Google DeepMind. https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithmsDeepMind. 2025. “AlphaEvolve: A Gemini-powered Coding Agent for Designing Advanced Algorithms”. Google DeepMind. https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms.DeepMind. “AlphaEvolve: A Gemini-powered Coding Agent for Designing Advanced Algorithms”. Google DeepMind, 2025, https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms.DeepMind. AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. Google DeepMind https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms (2025).DeepMind, “AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms
DeepSeek-AI et al.(2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv.DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, … Zhen Zhang. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. In arXiv. https://arxiv.org/abs/2501.12948DeepSeek-AI, Daya Guo, Dejian Yang, et al. 2025. “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”. In arXiv. Preprint, January 22. https://arxiv.org/abs/2501.12948.DeepSeek-AI, et al. “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”. arXiv, 22 Jan. 2025, https://arxiv.org/abs/2501.12948.DeepSeek-AI et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv Preprint at https://arxiv.org/abs/2501.12948 (2025).DeepSeek-AI et al., “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, Jan. 22, 2025. [Online]. Available: https://arxiv.org/abs/2501.12948
Delétang, G. et al.(2022). Neural Networks and the Chomsky Hierarchy. arXiv.Delétang, G., Ruoss, A., Grau-Moya, J., Genewein, T., Wenliang, L. K., Catt, E., Cundy, C., Hutter, M., Legg, S., Veness, J., & Ortega, P. A. (2022). Neural Networks and the Chomsky Hierarchy. In arXiv. https://arxiv.org/abs/2207.02098Delétang, G., A. Ruoss, J. Grau-Moya, et al. 2022. “Neural Networks and the Chomsky Hierarchy”. In arXiv. Preprint, July 5. https://arxiv.org/abs/2207.02098.Delétang, G., et al. “Neural Networks and the Chomsky Hierarchy”. arXiv, 5 July 2022, https://arxiv.org/abs/2207.02098.Delétang, G. et al. Neural Networks and the Chomsky Hierarchy. arXiv Preprint at https://arxiv.org/abs/2207.02098 (2022).G. Delétang et al., “Neural Networks and the Chomsky Hierarchy”, Jul. 05, 2022. [Online]. Available: https://arxiv.org/abs/2207.02098
Deng(2009). ImageNet: A large-scale hierarchical image database.Deng. (2009). ImageNet: A large-scale hierarchical image database. https://ieeexplore.ieee.org/document/5206848Deng. 2009. “ImageNet: A Large-scale Hierarchical Image Database”. https://ieeexplore.ieee.org/document/5206848.Deng. ImageNet: A Large-scale Hierarchical Image Database. 2009, https://ieeexplore.ieee.org/document/5206848.Deng. ImageNet: A large-scale hierarchical image database. https://ieeexplore.ieee.org/document/5206848 (2009).Deng, “ImageNet: A large-scale hierarchical image database”. [Online]. Available: https://ieeexplore.ieee.org/document/5206848
Dewey, D.(2011). Learning What to Value. Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings.Dewey, D. (2011). Learning What to Value. Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings, Lecture Notes in Computer Science, 6830, 309–314. https://doi.org/10.1007/978-3-642-22887-2_35Dewey, D. 2011. “Learning What to Value”. Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings (Berlin), Lecture Notes in Computer Science, vol. 6830: 309–14. https://doi.org/10.1007/978-3-642-22887-2_35.Dewey, D. “Learning What to Value”. Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings [Berlin], Lecture Notes in Computer Science, vol. 6830, 2011, pp. 309–14, https://doi.org/10.1007/978-3-642-22887-2_35.Dewey, D. Learning What to Value. in Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings vol. 6830 309–314 (Springer, Berlin, 2011).D. Dewey, “Learning What to Value”, in Artificial General Intelligence: 4th International Conference, AGI 2011, Proceedings, in Lecture Notes in Computer Science, vol. 6830. Berlin: Springer, 2011, pp. 309–314. doi: 10.1007/978-3-642-22887-2_35.
Dherin, B., Munn, M., Rosca, M. & Barrett, D. G. T.(2022). Why neural networks find simple solutions: the many regularizers of geometric complexity. arXiv.Dherin, B., Munn, M., Rosca, M., & Barrett, D. G. T. (2022). Why neural networks find simple solutions: the many regularizers of geometric complexity. In arXiv. https://arxiv.org/abs/2209.13083Dherin, B., M. Munn, M. Rosca, and D. G. T. Barrett. 2022. “Why Neural Networks Find Simple Solutions: The Many Regularizers of Geometric Complexity”. In arXiv. Preprint, September 27. https://arxiv.org/abs/2209.13083.Dherin, B., et al. “Why Neural Networks Find Simple Solutions: The Many Regularizers of Geometric Complexity”. arXiv, 27 Sept. 2022, https://arxiv.org/abs/2209.13083.Dherin, B., Munn, M., Rosca, M. & Barrett, D. G. T. Why neural networks find simple solutions: the many regularizers of geometric complexity. arXiv Preprint at https://arxiv.org/abs/2209.13083 (2022).B. Dherin, M. Munn, M. Rosca, and D. G. T. Barrett, “Why neural networks find simple solutions: the many regularizers of geometric complexity”, Sep. 27, 2022. [Online]. Available: https://arxiv.org/abs/2209.13083
DiGiovanni(2023). Beginner’s guide to reducing s-risks. Center on Long-Term Risk.DiGiovanni. (2023). Beginner’s guide to reducing s-risks. Center on Long-Term Risk. https://longtermrisk.org/beginners-guide-to-reducing-s-risksDiGiovanni. 2023. “Beginner’s Guide to Reducing S-risks”. Center on Long-Term Risk. https://longtermrisk.org/beginners-guide-to-reducing-s-risks.DiGiovanni. “Beginner’s Guide to Reducing S-risks”. Center on Long-Term Risk, 2023, https://longtermrisk.org/beginners-guide-to-reducing-s-risks.DiGiovanni. Beginner’s guide to reducing s-risks. Center on Long-Term Risk https://longtermrisk.org/beginners-guide-to-reducing-s-risks (2023).DiGiovanni, “Beginner’s guide to reducing s-risks”, Center on Long-Term Risk. [Online]. Available: https://longtermrisk.org/beginners-guide-to-reducing-s-risks
Ding, J. & Dafoe, A.(2020). The Logic of Strategic Assets: From Oil to Artificial Intelligence. arXiv.Ding, J., & Dafoe, A. (2020). The Logic of Strategic Assets: From Oil to Artificial Intelligence. In arXiv. https://doi.org/10.1080/09636412.2021.1915583Ding, J., and A. Dafoe. 2020. “The Logic of Strategic Assets: From Oil to Artificial Intelligence”. In arXiv. Preprint, January 9. https://doi.org/10.1080/09636412.2021.1915583.Ding, J., and A. Dafoe. “The Logic of Strategic Assets: From Oil to Artificial Intelligence”. arXiv, 9 Jan. 2020, https://doi.org/10.1080/09636412.2021.1915583.Ding, J. & Dafoe, A. The Logic of Strategic Assets: From Oil to Artificial Intelligence. arXiv Preprint at https://doi.org/10.1080/09636412.2021.1915583 (2020).J. Ding and A. Dafoe, “The Logic of Strategic Assets: From Oil to Artificial Intelligence”, Jan. 09, 2020. doi: 10.1080/09636412.2021.1915583.
Ding, J.(2018). Deciphering China's AI Dream: The Context, Components, Capabilities, and Consequences of China's Strategy to Lead the World in AI.Ding, J. (2018). Deciphering China's AI Dream: The Context, Components, Capabilities, and Consequences of China's Strategy to Lead the World in AI. Future of Humanity Institute, University of Oxford. https://fhi.ox.ac.uk/wp-content/uploads/Deciphering_Chinas_AI-Dream.pdfDing, J. 2018. Deciphering China's AI Dream: The Context, Components, Capabilities, and Consequences of China's Strategy to Lead the World in AI. Future of Humanity Institute, University of Oxford. https://fhi.ox.ac.uk/wp-content/uploads/Deciphering_Chinas_AI-Dream.pdf.Ding, J. Deciphering China's AI Dream: The Context, Components, Capabilities, and Consequences of China's Strategy to Lead the World in AI. Future of Humanity Institute, University of Oxford, 14 Mar. 2018, https://fhi.ox.ac.uk/wp-content/uploads/Deciphering_Chinas_AI-Dream.pdf.Ding, J. Deciphering China's AI Dream: The Context, Components, Capabilities, and Consequences of China's Strategy to Lead the World in AI. https://fhi.ox.ac.uk/wp-content/uploads/Deciphering_Chinas_AI-Dream.pdf (2018).J. Ding, “Deciphering China's AI Dream: The Context, Components, Capabilities, and Consequences of China's Strategy to Lead the World in AI”, Future of Humanity Institute, University of Oxford, Mar. 2018. [Online]. Available: https://fhi.ox.ac.uk/wp-content/uploads/Deciphering_Chinas_AI-Dream.pdf
DnaScript(2024). Home. DNA Script.DnaScript. (2024). Home. DNA Script. https://dnascript.comDnaScript. 2024. “Home”. DNA Script. https://dnascript.com.DnaScript. “Home”. DNA Script, 2024, https://dnascript.com.DnaScript. Home. DNA Script https://dnascript.com (2024).DnaScript, “Home”, DNA Script. [Online]. Available: https://dnascript.com
Dong, Y. et al.(2024). Safeguarding Large Language Models: A Survey. arXiv.Dong, Y., Mu, R., Zhang, Y., Sun, S., Zhang, T., Wu, C., Jin, G., Qi, Y., Hu, J., Meng, J., Bensalem, S., & Huang, X. (2024). Safeguarding Large Language Models: A Survey. In arXiv. https://arxiv.org/abs/2406.02622Dong, Y., R. Mu, Y. Zhang, et al. 2024. “Safeguarding Large Language Models: A Survey”. In arXiv. Preprint, June 3. https://arxiv.org/abs/2406.02622.Dong, Y., et al. “Safeguarding Large Language Models: A Survey”. arXiv, 3 June 2024, https://arxiv.org/abs/2406.02622.Dong, Y. et al. Safeguarding Large Language Models: A Survey. arXiv Preprint at https://arxiv.org/abs/2406.02622 (2024).Y. Dong et al., “Safeguarding Large Language Models: A Survey”, Jun. 03, 2024. [Online]. Available: https://arxiv.org/abs/2406.02622
Douillard et al(2023). AI Safety. Tigera – Creator of Calico.Douillard et al. (2023). AI Safety. Tigera – Creator of Calico. https://tigera.io/learn/guides/llm-security/ai-safetyDouillard et al. 2023. “AI Safety”. Tigera – Creator of Calico. https://tigera.io/learn/guides/llm-security/ai-safety.Douillard et al. “AI Safety”. Tigera – Creator of Calico, 2023, https://tigera.io/learn/guides/llm-security/ai-safety.Douillard et al. AI Safety. Tigera – Creator of Calico https://tigera.io/learn/guides/llm-security/ai-safety (2023).Douillard et al, “AI Safety”, Tigera – Creator of Calico. [Online]. Available: https://tigera.io/learn/guides/llm-security/ai-safety
Douillard et al(2024). DiPaCo: Distributed Path Composition. arXiv.org.Douillard et al. (2024). DiPaCo: Distributed Path Composition. arXiv.org. https://www.arxiv.org/abs/2403.10616Douillard et al. 2024. “DiPaCo: Distributed Path Composition”. arXiv.org. https://www.arxiv.org/abs/2403.10616.Douillard et al. “DiPaCo: Distributed Path Composition”. arXiv.org, 2024, https://www.arxiv.org/abs/2403.10616.Douillard et al. DiPaCo: Distributed Path Composition. arXiv.org https://www.arxiv.org/abs/2403.10616 (2024).Douillard et al, “DiPaCo: Distributed Path Composition”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2403.10616
Douillard, A. et al.(2023). DiLoCo: Distributed Low-Communication Training of Language Models. arXiv.Douillard, A., Feng, Q., Rusu, A. A., Chhaparia, R., Donchev, Y., Kuncoro, A., Ranzato, M., Szlam, A., & Shen, J. (2023). DiLoCo: Distributed Low-Communication Training of Language Models. In arXiv. https://arxiv.org/abs/2311.08105Douillard, A., Q. Feng, A. A. Rusu, et al. 2023. “DiLoCo: Distributed Low-Communication Training of Language Models”. In arXiv. Preprint, November 14. https://arxiv.org/abs/2311.08105.Douillard, A., et al. “DiLoCo: Distributed Low-Communication Training of Language Models”. arXiv, 14 Nov. 2023, https://arxiv.org/abs/2311.08105.Douillard, A. et al. DiLoCo: Distributed Low-Communication Training of Language Models. arXiv Preprint at https://arxiv.org/abs/2311.08105 (2023).A. Douillard et al., “DiLoCo: Distributed Low-Communication Training of Language Models”, Nov. 14, 2023. [Online]. Available: https://arxiv.org/abs/2311.08105
Dowie(1977). Pinto Madness: the Ford Pinto’s fire-prone gas tank.Dowie. (1977). Pinto Madness: the Ford Pinto’s fire-prone gas tank. https://muckrakerfarm.com/1977/09/pinto-madness-ford-pintos-fire-prone-gas-tankDowie. 1977. “Pinto Madness: The Ford Pinto’s Fire-prone Gas Tank”. https://muckrakerfarm.com/1977/09/pinto-madness-ford-pintos-fire-prone-gas-tank.Dowie. Pinto Madness: The Ford Pinto’s Fire-prone Gas Tank. 1977, https://muckrakerfarm.com/1977/09/pinto-madness-ford-pintos-fire-prone-gas-tank.Dowie. Pinto Madness: the Ford Pinto’s fire-prone gas tank. https://muckrakerfarm.com/1977/09/pinto-madness-ford-pintos-fire-prone-gas-tank (1977).Dowie, “Pinto Madness: the Ford Pinto’s fire-prone gas tank”. [Online]. Available: https://muckrakerfarm.com/1977/09/pinto-madness-ford-pintos-fire-prone-gas-tank
Dwarkesh Patel(2023). Dario Amodei (Anthropic CEO) - Scaling, Alignment, & AI Progress.Dwarkesh Patel. (2023, August 8). Dario Amodei (Anthropic CEO) - Scaling, Alignment, & AI Progress. https://dwarkeshpatel.com/p/dario-amodeiDwarkesh Patel. 2023. Dario Amodei (Anthropic CEO) - Scaling, Alignment, & AI Progress. Edition. August 8. https://dwarkeshpatel.com/p/dario-amodei.Dwarkesh Patel. Dario Amodei (Anthropic CEO) - Scaling, Alignment, & AI Progress. 8 Aug. 2023, https://dwarkeshpatel.com/p/dario-amodei.Dwarkesh Patel. Dario Amodei (Anthropic CEO) - Scaling, Alignment, & AI Progress. https://dwarkeshpatel.com/p/dario-amodei (2023).Dwarkesh Patel, “Dario Amodei (Anthropic CEO) - Scaling, Alignment, & AI Progress”. [Online]. Available: https://dwarkeshpatel.com/p/dario-amodei
EA Global(2020). Paul Christiano: Current work in AI alignment. EA Forum.EA Global. (2020, April 3). Paul Christiano: Current work in AI alignment. EA Forum. https://forum.effectivealtruism.org/posts/63stBTw3WAW6k45dY/paul-christiano-current-work-in-ai-alignmentEA Global. 2020. “Paul Christiano: Current Work in AI Alignment”. EA Forum, April 3. https://forum.effectivealtruism.org/posts/63stBTw3WAW6k45dY/paul-christiano-current-work-in-ai-alignment.EA Global. “Paul Christiano: Current Work in AI Alignment”. EA Forum, 3 Apr. 2020, https://forum.effectivealtruism.org/posts/63stBTw3WAW6k45dY/paul-christiano-current-work-in-ai-alignment.EA Global. Paul Christiano: Current work in AI alignment. EA Forum https://forum.effectivealtruism.org/posts/63stBTw3WAW6k45dY/paul-christiano-current-work-in-ai-alignment (2020).EA Global, “Paul Christiano: Current work in AI alignment”, EA Forum. [Online]. Available: https://forum.effectivealtruism.org/posts/63stBTw3WAW6k45dY/paul-christiano-current-work-in-ai-alignment
Egan, J. & Heim, L.(2023). Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers. arXiv.Egan, J., & Heim, L. (2023). Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers. In arXiv. https://arxiv.org/abs/2310.13625Egan, J., and L. Heim. 2023. “Oversight for Frontier AI Through a Know-Your-Customer Scheme for Compute Providers”. In arXiv. Preprint, October 20. https://arxiv.org/abs/2310.13625.Egan, J., and L. Heim. “Oversight for Frontier AI Through a Know-Your-Customer Scheme for Compute Providers”. arXiv, 20 Oct. 2023, https://arxiv.org/abs/2310.13625.Egan, J. & Heim, L. Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers. arXiv Preprint at https://arxiv.org/abs/2310.13625 (2023).J. Egan and L. Heim, “Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers”, Oct. 20, 2023. [Online]. Available: https://arxiv.org/abs/2310.13625
Ege Erdil & Matthew Barnett(2025). Most AI value will come from broad automation, not from R&D.Ege Erdil, & Matthew Barnett. (2025, March 21). Most AI value will come from broad automation, not from R&D. https://epoch.ai/gradient-updates/most-ai-value-will-come-from-broad-automation-not-from-r-dEge Erdil, and Matthew Barnett. 2025. “Most AI Value Will Come from Broad Automation, Not from R&D”. March 21. https://epoch.ai/gradient-updates/most-ai-value-will-come-from-broad-automation-not-from-r-d.Ege Erdil, and Matthew Barnett. Most AI Value Will Come from Broad Automation, Not from R&D. 21 Mar. 2025, https://epoch.ai/gradient-updates/most-ai-value-will-come-from-broad-automation-not-from-r-d.Ege Erdil & Matthew Barnett. Most AI value will come from broad automation, not from R&D. https://epoch.ai/gradient-updates/most-ai-value-will-come-from-broad-automation-not-from-r-d (2025).Ege Erdil and Matthew Barnett, “Most AI value will come from broad automation, not from R&D”. [Online]. Available: https://epoch.ai/gradient-updates/most-ai-value-will-come-from-broad-automation-not-from-r-d
Eiras, F. et al.(2024). Near to Mid-term Risks and Opportunities of Open-Source Generative AI. arXiv.Eiras, F., Petrov, A., Vidgen, B., de Witt, C. S., Pizzati, F., Elkins, K., Mukhopadhyay, S., Bibi, A., Csaba, B., Steibel, F., Barez, F., Smith, G., Guadagni, G., Chun, J., Cabot, J., Imperial, J. M., Nolazco-Flores, J. A., Landay, L., Jackson, M., … Foerster, J. (2024). Near to Mid-term Risks and Opportunities of Open-Source Generative AI. In arXiv. https://arxiv.org/abs/2404.17047Eiras, F., A. Petrov, B. Vidgen, et al. 2024. “Near to Mid-term Risks and Opportunities of Open-Source Generative AI”. In arXiv. Preprint, April 25. https://arxiv.org/abs/2404.17047.Eiras, F., et al. “Near to Mid-term Risks and Opportunities of Open-Source Generative AI”. arXiv, 25 Apr. 2024, https://arxiv.org/abs/2404.17047.Eiras, F. et al. Near to Mid-term Risks and Opportunities of Open-Source Generative AI. arXiv Preprint at https://arxiv.org/abs/2404.17047 (2024).F. Eiras et al., “Near to Mid-term Risks and Opportunities of Open-Source Generative AI”, Apr. 25, 2024. [Online]. Available: https://arxiv.org/abs/2404.17047
El-Mhamdi, E. et al.(2022). On the Impossible Safety of Large AI Models. arXiv.El-Mhamdi, E.-M., Farhadkhani, S., Guerraoui, R., Gupta, N., Hoang, L.-N., Pinot, R., Rouault, S., & Stephan, J. (2022). On the Impossible Safety of Large AI Models. In arXiv. https://arxiv.org/abs/2209.15259El-Mhamdi, E.-M., S. Farhadkhani, R. Guerraoui, et al. 2022. “On the Impossible Safety of Large AI Models”. In arXiv. Preprint, September 30. https://arxiv.org/abs/2209.15259.El-Mhamdi, E.-M., et al. “On the Impossible Safety of Large AI Models”. arXiv, 30 Sept. 2022, https://arxiv.org/abs/2209.15259.El-Mhamdi, E.-M. et al. On the Impossible Safety of Large AI Models. arXiv Preprint at https://arxiv.org/abs/2209.15259 (2022).E.-M. El-Mhamdi et al., “On the Impossible Safety of Large AI Models”, Sep. 30, 2022. [Online]. Available: https://arxiv.org/abs/2209.15259
Eliezer Yudkowsky(2007). Politics is the Mind-Killer. LessWrong.Eliezer Yudkowsky. (2007, February 18). Politics is the Mind-Killer. LessWrong. https://lesswrong.com/posts/9weLK2AJ9JEt2Tt8f/politics-is-the-mind-killerEliezer Yudkowsky. 2007. “Politics Is the Mind-Killer”. LessWrong, February 18. https://lesswrong.com/posts/9weLK2AJ9JEt2Tt8f/politics-is-the-mind-killer.Eliezer Yudkowsky. “Politics Is the Mind-Killer”. LessWrong, 18 Feb. 2007, https://lesswrong.com/posts/9weLK2AJ9JEt2Tt8f/politics-is-the-mind-killer.Eliezer Yudkowsky. Politics is the Mind-Killer. LessWrong https://lesswrong.com/posts/9weLK2AJ9JEt2Tt8f/politics-is-the-mind-killer (2007).Eliezer Yudkowsky, “Politics is the Mind-Killer”, LessWrong. [Online]. Available: https://lesswrong.com/posts/9weLK2AJ9JEt2Tt8f/politics-is-the-mind-killer
Eliezer Yudkowsky(2008). Shut up and do the impossible!. LessWrong.Eliezer Yudkowsky. (2008, October 8). Shut up and do the impossible!. LessWrong. https://lesswrong.com/posts/nCvvhFBaayaXyuBiD/shut-up-and-do-the-impossibleEliezer Yudkowsky. 2008. “Shut up and Do the Impossible!”. LessWrong, October 8. https://lesswrong.com/posts/nCvvhFBaayaXyuBiD/shut-up-and-do-the-impossible.Eliezer Yudkowsky. “Shut up and Do the Impossible!”. LessWrong, 8 Oct. 2008, https://lesswrong.com/posts/nCvvhFBaayaXyuBiD/shut-up-and-do-the-impossible.Eliezer Yudkowsky. Shut up and do the impossible!. LessWrong https://lesswrong.com/posts/nCvvhFBaayaXyuBiD/shut-up-and-do-the-impossible (2008).Eliezer Yudkowsky, “Shut up and do the impossible!”, LessWrong. [Online]. Available: https://lesswrong.com/posts/nCvvhFBaayaXyuBiD/shut-up-and-do-the-impossible
Eliezer Yudkowsky(2022). AGI Ruin: A List of Lethalities. AI Alignment Forum.Eliezer Yudkowsky. (2022, June 5). AGI Ruin: A List of Lethalities. AI Alignment Forum. https://alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalitiesEliezer Yudkowsky. 2022. “AGI Ruin: A List of Lethalities”. AI Alignment Forum, June 5. https://alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities.Eliezer Yudkowsky. “AGI Ruin: A List of Lethalities”. AI Alignment Forum, 5 June 2022, https://alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities.Eliezer Yudkowsky. AGI Ruin: A List of Lethalities. AI Alignment Forum https://alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities (2022).Eliezer Yudkowsky, “AGI Ruin: A List of Lethalities”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities
Emily H. Soice, Rafael Rocha, Kimberlee Cordova, Michael Specter & Kevin M. Esvelt(2023). Can large language models democratize access to dual-use biotechnology?. arXiv.Emily H. Soice, Rafael Rocha, Kimberlee Cordova, Michael Specter, & Kevin M. Esvelt. (2023). Can large language models democratize access to dual-use biotechnology?. In arXiv. https://arxiv.org/abs/2306.03809Emily H. Soice, Rafael Rocha, Kimberlee Cordova, Michael Specter, and Kevin M. Esvelt. 2023. “Can Large Language Models Democratize Access to Dual-use Biotechnology?”. In arXiv. Preprint, June 6. https://arxiv.org/abs/2306.03809.Emily H. Soice, et al. “Can Large Language Models Democratize Access to Dual-use Biotechnology?”. arXiv, 6 June 2023, https://arxiv.org/abs/2306.03809.Emily H. Soice, Rafael Rocha, Kimberlee Cordova, Michael Specter & Kevin M. Esvelt. Can large language models democratize access to dual-use biotechnology?. arXiv Preprint at https://arxiv.org/abs/2306.03809 (2023).Emily H. Soice, Rafael Rocha, Kimberlee Cordova, Michael Specter, and Kevin M. Esvelt, “Can large language models democratize access to dual-use biotechnology?”, Jun. 06, 2023. [Online]. Available: https://arxiv.org/abs/2306.03809
Epicural(2021). Goodhart’s Law. Epicural, LLC - Change is hard. The Epicural Team “does hard.”.Epicural. (2021). Goodhart’s Law. Internet Archive (https://web.archive.org/web/20210804125857/https://www.epicural.com/2021/04/27/goodharts-law/). Epicural, LLC - Change Is Hard. The Epicural Team “does Hard.”. https://epicural.com/2021/04/27/goodharts-lawEpicural. 2021. “Goodhart’s Law”. Epicural, LLC - Change Is Hard. The Epicural Team “does Hard.”. Https://web.archive.org/web/20210804125857/https://www.epicural.com/2021/04/27/goodharts-law/. Internet Archive. https://epicural.com/2021/04/27/goodharts-law.Epicural. “Goodhart’s Law”. Epicural, LLC - Change Is Hard. The Epicural Team “does Hard.”, 2021, Internet Archive, https://web.archive.org/web/20210804125857/https://www.epicural.com/2021/04/27/goodharts-law/, https://epicural.com/2021/04/27/goodharts-law.Epicural. Goodhart’s Law. Epicural, LLC - Change is hard. The Epicural Team “does hard.” https://epicural.com/2021/04/27/goodharts-law (2021).Epicural, “Goodhart’s Law”, Epicural, LLC - Change is hard. The Epicural Team “does hard.”. Accessed: Aug. 04, 2021. [Online]. Available: https://epicural.com/2021/04/27/goodharts-law
Epoch AI(2025). Data on AI Data Centers. Epoch AI.Epoch AI. (2025). Data on AI Data Centers. Epoch AI. https://epoch.ai/data/data-centersEpoch AI. 2025. “Data on AI Data Centers”. Epoch AI. https://epoch.ai/data/data-centers.Epoch AI. “Data on AI Data Centers”. Epoch AI, 2025, https://epoch.ai/data/data-centers.Epoch AI. Data on AI Data Centers. Epoch AI https://epoch.ai/data/data-centers (2025).Epoch AI, “Data on AI Data Centers”, Epoch AI. [Online]. Available: https://epoch.ai/data/data-centers
Epoch AI(2025). Data on AI Models. Epoch AI.Epoch AI. (2025). Data on AI Models. Epoch AI. https://epoch.ai/data/ai-modelsEpoch AI. 2025. “Data on AI Models”. Epoch AI. https://epoch.ai/data/ai-models.Epoch AI. “Data on AI Models”. Epoch AI, 2025, https://epoch.ai/data/ai-models.Epoch AI. Data on AI Models. Epoch AI https://epoch.ai/data/ai-models (2025).Epoch AI, “Data on AI Models”, Epoch AI. [Online]. Available: https://epoch.ai/data/ai-models
Epoch AI(2025). Data on Machine Learning Hardware. Epoch AI.Epoch AI. (2025). Data on Machine Learning Hardware. Epoch AI. https://epoch.ai/data/machine-learning-hardwareEpoch AI. 2025. “Data on Machine Learning Hardware”. Epoch AI. https://epoch.ai/data/machine-learning-hardware.Epoch AI. “Data on Machine Learning Hardware”. Epoch AI, 2025, https://epoch.ai/data/machine-learning-hardware.Epoch AI. Data on Machine Learning Hardware. Epoch AI https://epoch.ai/data/machine-learning-hardware (2025).Epoch AI, “Data on Machine Learning Hardware”, Epoch AI. [Online]. Available: https://epoch.ai/data/machine-learning-hardware
Epoch AI(2025). GATE Model Playground. Epoch AI.Epoch AI. (2025). GATE Model Playground. Epoch AI. https://epoch.ai/gateEpoch AI. 2025. “GATE Model Playground”. Epoch AI. https://epoch.ai/gate.Epoch AI. “GATE Model Playground”. Epoch AI, 2025, https://epoch.ai/gate.Epoch AI. GATE Model Playground. Epoch AI https://epoch.ai/gate (2025).Epoch AI, “GATE Model Playground”, Epoch AI. [Online]. Available: https://epoch.ai/gate
Epoch AI(2025). Trends in Artificial Intelligence. Epoch AI.Epoch AI. (2025). Trends in Artificial Intelligence. Epoch AI. https://epoch.ai/trendsEpoch AI. 2025. “Trends in Artificial Intelligence”. Epoch AI. https://epoch.ai/trends.Epoch AI. “Trends in Artificial Intelligence”. Epoch AI, 2025, https://epoch.ai/trends.Epoch AI. Trends in Artificial Intelligence. Epoch AI https://epoch.ai/trends (2025).Epoch AI, “Trends in Artificial Intelligence”, Epoch AI. [Online]. Available: https://epoch.ai/trends
EpochAI(2024). FrontierMath: LLM Benchmark for Advanced AI Math Reasoning. Epoch AI.EpochAI. (2024). FrontierMath: LLM Benchmark for Advanced AI Math Reasoning. Epoch AI. https://epoch.ai/frontiermathEpochAI. 2024. “FrontierMath: LLM Benchmark for Advanced AI Math Reasoning”. Epoch AI. https://epoch.ai/frontiermath.EpochAI. “FrontierMath: LLM Benchmark for Advanced AI Math Reasoning”. Epoch AI, 2024, https://epoch.ai/frontiermath.EpochAI. FrontierMath: LLM Benchmark for Advanced AI Math Reasoning. Epoch AI https://epoch.ai/frontiermath (2024).EpochAI, “FrontierMath: LLM Benchmark for Advanced AI Math Reasoning”, Epoch AI. [Online]. Available: https://epoch.ai/frontiermath
EpochAI(2025). Could decentralized training solve AI’s power problem?. Epoch AI.EpochAI. (2025). Could decentralized training solve AI’s power problem?. Epoch AI. https://epoch.ai/blog/could-decentralized-training-solve-ais-power-problemEpochAI. 2025. “Could Decentralized Training Solve AI’s Power Problem?”. Epoch AI. https://epoch.ai/blog/could-decentralized-training-solve-ais-power-problem.EpochAI. “Could Decentralized Training Solve AI’s Power Problem?”. Epoch AI, 2025, https://epoch.ai/blog/could-decentralized-training-solve-ais-power-problem.EpochAI. Could decentralized training solve AI’s power problem?. Epoch AI https://epoch.ai/blog/could-decentralized-training-solve-ais-power-problem (2025).EpochAI, “Could decentralized training solve AI’s power problem?”, Epoch AI. [Online]. Available: https://epoch.ai/blog/could-decentralized-training-solve-ais-power-problem
Erdil, E. & Besiroglu, T.(2023). Explosive growth from AI automation: A review of the arguments. arXiv.Erdil, E., & Besiroglu, T. (2023). Explosive growth from AI automation: A review of the arguments. In arXiv. https://arxiv.org/abs/2309.11690Erdil, E., and T. Besiroglu. 2023. “Explosive Growth from AI Automation: A Review of the Arguments”. In arXiv. Preprint, September 20. https://arxiv.org/abs/2309.11690.Erdil, E., and T. Besiroglu. “Explosive Growth from AI Automation: A Review of the Arguments”. arXiv, 20 Sept. 2023, https://arxiv.org/abs/2309.11690.Erdil, E. & Besiroglu, T. Explosive growth from AI automation: A review of the arguments. arXiv Preprint at https://arxiv.org/abs/2309.11690 (2023).E. Erdil and T. Besiroglu, “Explosive growth from AI automation: A review of the arguments”, Sep. 20, 2023. [Online]. Available: https://arxiv.org/abs/2309.11690
Erdil, E. et al.(2025). GATE: An Integrated Assessment Model for AI Automation. arXiv.Erdil, E., Potlogea, A., Besiroglu, T., Roldan, E., Ho, A., Sevilla, J., Barnett, M., Vrzla, M., & Sandler, R. (2025). GATE: An Integrated Assessment Model for AI Automation. In arXiv. https://arxiv.org/abs/2503.04941Erdil, E., A. Potlogea, T. Besiroglu, et al. 2025. “GATE: An Integrated Assessment Model for AI Automation”. In arXiv. Preprint, March 6. https://arxiv.org/abs/2503.04941.Erdil, E., et al. “GATE: An Integrated Assessment Model for AI Automation”. arXiv, 6 Mar. 2025, https://arxiv.org/abs/2503.04941.Erdil, E. et al. GATE: An Integrated Assessment Model for AI Automation. arXiv Preprint at https://arxiv.org/abs/2503.04941 (2025).E. Erdil et al., “GATE: An Integrated Assessment Model for AI Automation”, Mar. 06, 2025. [Online]. Available: https://arxiv.org/abs/2503.04941
Erich Grunewald(2023). Introduction to AI Chip Making in China.Erich Grunewald. (2023, December 14). Introduction to AI Chip Making in China. https://iaps.ai/research/ai-chip-making-chinaErich Grunewald. 2023. “Introduction to AI Chip Making in China”. December 14. https://iaps.ai/research/ai-chip-making-china.Erich Grunewald. Introduction to AI Chip Making in China. 14 Dec. 2023, https://iaps.ai/research/ai-chip-making-china.Erich Grunewald. Introduction to AI Chip Making in China. https://iaps.ai/research/ai-chip-making-china (2023).Erich Grunewald, “Introduction to AI Chip Making in China”. [Online]. Available: https://iaps.ai/research/ai-chip-making-china
Esvelt, K. M.(2022). Delay, Detect, Defend: Preparing for a Future in which Thousands Can Release New Pandemics.Esvelt, K. M. (2022). Delay, Detect, Defend: Preparing for a Future in which Thousands Can Release New Pandemics (Geneva Paper 29/22). Geneva Centre for Security Policy. https://dam.gcsp.ch/files/doc/gcsp-geneva-paper-29-22Esvelt, K. M. 2022. Delay, Detect, Defend: Preparing for a Future in Which Thousands Can Release New Pandemics. Geneva Paper 29/22. Geneva Centre for Security Policy. https://dam.gcsp.ch/files/doc/gcsp-geneva-paper-29-22.Esvelt, K. M. Delay, Detect, Defend: Preparing for a Future in Which Thousands Can Release New Pandemics. Geneva Centre for Security Policy, 2022, https://dam.gcsp.ch/files/doc/gcsp-geneva-paper-29-22. Geneva Paper 29/22.Esvelt, K. M. Delay, Detect, Defend: Preparing for a Future in Which Thousands Can Release New Pandemics. https://dam.gcsp.ch/files/doc/gcsp-geneva-paper-29-22 (2022).K. M. Esvelt, “Delay, Detect, Defend: Preparing for a Future in which Thousands Can Release New Pandemics”, Geneva Centre for Security Policy, Geneva, 2022. [Online]. Available: https://dam.gcsp.ch/files/doc/gcsp-geneva-paper-29-22
Ethan Perez et al.(2022). Discovering Language Model Behaviors with Model-Written Evaluations. arXiv.Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, … Jared Kaplan. (2022). Discovering Language Model Behaviors with Model-Written Evaluations. In arXiv. https://arxiv.org/abs/2212.09251Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, et al. 2022. “Discovering Language Model Behaviors with Model-Written Evaluations”. In arXiv. Preprint, December 19. https://arxiv.org/abs/2212.09251.Ethan Perez, et al. “Discovering Language Model Behaviors with Model-Written Evaluations”. arXiv, 19 Dec. 2022, https://arxiv.org/abs/2212.09251.Ethan Perez et al. Discovering Language Model Behaviors with Model-Written Evaluations. arXiv Preprint at https://arxiv.org/abs/2212.09251 (2022).Ethan Perez et al., “Discovering Language Model Behaviors with Model-Written Evaluations”, Dec. 19, 2022. [Online]. Available: https://arxiv.org/abs/2212.09251
Ethayarajh, K. & Jurafsky, D.(2022). The Authenticity Gap in Human Evaluation. arXiv.Ethayarajh, K., & Jurafsky, D. (2022). The Authenticity Gap in Human Evaluation. In arXiv. https://arxiv.org/abs/2205.11930Ethayarajh, K., and D. Jurafsky. 2022. “The Authenticity Gap in Human Evaluation”. In arXiv. Preprint, May 24. https://arxiv.org/abs/2205.11930.Ethayarajh, K., and D. Jurafsky. “The Authenticity Gap in Human Evaluation”. arXiv, 24 May 2022, https://arxiv.org/abs/2205.11930.Ethayarajh, K. & Jurafsky, D. The Authenticity Gap in Human Evaluation. arXiv Preprint at https://arxiv.org/abs/2205.11930 (2022).K. Ethayarajh and D. Jurafsky, “The Authenticity Gap in Human Evaluation”, May 24, 2022. [Online]. Available: https://arxiv.org/abs/2205.11930
EU Commission(2025). New JRC collection of external scientific reports to inform the implementation of the EU AI Act on general-purpose AI models. AI Watch.EU Commission. (2025). New JRC collection of external scientific reports to inform the implementation of the EU AI Act on general-purpose AI models. AI Watch. https://ai-watch.ec.europa.eu/news/new-jrc-collection-external-scientific-reports-inform-implementation-eu-ai-act-general-purpose-ai-2025-10-14_enEU Commission. 2025. “New JRC Collection of External Scientific Reports to Inform the Implementation of the EU AI Act on General-purpose AI Models”. AI Watch. https://ai-watch.ec.europa.eu/news/new-jrc-collection-external-scientific-reports-inform-implementation-eu-ai-act-general-purpose-ai-2025-10-14_en.EU Commission. “New JRC Collection of External Scientific Reports to Inform the Implementation of the EU AI Act on General-purpose AI Models”. AI Watch, 2025, https://ai-watch.ec.europa.eu/news/new-jrc-collection-external-scientific-reports-inform-implementation-eu-ai-act-general-purpose-ai-2025-10-14_en.EU Commission. New JRC collection of external scientific reports to inform the implementation of the EU AI Act on general-purpose AI models. AI Watch https://ai-watch.ec.europa.eu/news/new-jrc-collection-external-scientific-reports-inform-implementation-eu-ai-act-general-purpose-ai-2025-10-14_en (2025).EU Commission, “New JRC collection of external scientific reports to inform the implementation of the EU AI Act on general-purpose AI models”, AI Watch. [Online]. Available: https://ai-watch.ec.europa.eu/news/new-jrc-collection-external-scientific-reports-inform-implementation-eu-ai-act-general-purpose-ai-2025-10-14_en
European Commission(2024). The Act Texts. EU Artificial Intelligence Act.European Commission. (2024). The Act Texts. EU Artificial Intelligence Act. https://artificialintelligenceact.eu/the-actEuropean Commission. 2024. “The Act Texts”. EU Artificial Intelligence Act. https://artificialintelligenceact.eu/the-act.European Commission. “The Act Texts”. EU Artificial Intelligence Act, 2024, https://artificialintelligenceact.eu/the-act.European Commission. The Act Texts. EU Artificial Intelligence Act https://artificialintelligenceact.eu/the-act (2024).European Commission, “The Act Texts”, EU Artificial Intelligence Act. [Online]. Available: https://artificialintelligenceact.eu/the-act
Evan Hubinger et al.(2024). Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. arXiv.Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, … Ethan Perez. (2024). Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. In arXiv. https://arxiv.org/abs/2401.05566Evan Hubinger, Carson Denison, Jesse Mu, et al. 2024. “Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training”. In arXiv. Preprint, January 10. https://arxiv.org/abs/2401.05566.Evan Hubinger, et al. “Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training”. arXiv, 10 Jan. 2024, https://arxiv.org/abs/2401.05566.Evan Hubinger et al. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. arXiv Preprint at https://arxiv.org/abs/2401.05566 (2024).Evan Hubinger et al., “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training”, Jan. 10, 2024. [Online]. Available: https://arxiv.org/abs/2401.05566
Everitt, T. et al.(2025). Evaluating the Goal-Directedness of Large Language Models. arXiv.Everitt, T., Garbacea, C., Bellot, A., Richens, J., Papadatos, H., Campos, S., & Shah, R. (2025). Evaluating the Goal-Directedness of Large Language Models. In arXiv. https://arxiv.org/abs/2504.11844Everitt, T., C. Garbacea, A. Bellot, et al. 2025. “Evaluating the Goal-Directedness of Large Language Models”. In arXiv. Preprint, April 16. https://arxiv.org/abs/2504.11844.Everitt, T., et al. “Evaluating the Goal-Directedness of Large Language Models”. arXiv, 16 Apr. 2025, https://arxiv.org/abs/2504.11844.Everitt, T. et al. Evaluating the Goal-Directedness of Large Language Models. arXiv Preprint at https://arxiv.org/abs/2504.11844 (2025).T. Everitt et al., “Evaluating the Goal-Directedness of Large Language Models”, Apr. 16, 2025. [Online]. Available: https://arxiv.org/abs/2504.11844
evhub(2020). AI safety via market making. LessWrong.evhub. (2020, June 26). AI safety via market making. LessWrong. https://lesswrong.com/posts/YWwzccGbcHMJMpT45/ai-safety-via-market-makingevhub. 2020. “AI Safety via Market Making”. LessWrong, June 26. https://lesswrong.com/posts/YWwzccGbcHMJMpT45/ai-safety-via-market-making.evhub. “AI Safety via Market Making”. LessWrong, 26 June 2020, https://lesswrong.com/posts/YWwzccGbcHMJMpT45/ai-safety-via-market-making.evhub. AI safety via market making. LessWrong https://lesswrong.com/posts/YWwzccGbcHMJMpT45/ai-safety-via-market-making (2020).evhub, “AI safety via market making”, LessWrong. [Online]. Available: https://lesswrong.com/posts/YWwzccGbcHMJMpT45/ai-safety-via-market-making
evhub(2020). Homogeneity vs. heterogeneity in AI takeoff scenarios. AI Alignment Forum.evhub. (2020, December 16). Homogeneity vs. heterogeneity in AI takeoff scenarios. AI Alignment Forum. https://alignmentforum.org/posts/mKBfa8v4S9pNKSyKK/homogeneity-vs-heterogeneity-in-ai-takeoff-scenariosevhub. 2020. “Homogeneity Vs. Heterogeneity in AI Takeoff Scenarios”. AI Alignment Forum, December 16. https://alignmentforum.org/posts/mKBfa8v4S9pNKSyKK/homogeneity-vs-heterogeneity-in-ai-takeoff-scenarios.evhub. “Homogeneity Vs. Heterogeneity in AI Takeoff Scenarios”. AI Alignment Forum, 16 Dec. 2020, https://alignmentforum.org/posts/mKBfa8v4S9pNKSyKK/homogeneity-vs-heterogeneity-in-ai-takeoff-scenarios.evhub. Homogeneity vs. heterogeneity in AI takeoff scenarios. AI Alignment Forum https://alignmentforum.org/posts/mKBfa8v4S9pNKSyKK/homogeneity-vs-heterogeneity-in-ai-takeoff-scenarios (2020).evhub, “Homogeneity vs. heterogeneity in AI takeoff scenarios”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/mKBfa8v4S9pNKSyKK/homogeneity-vs-heterogeneity-in-ai-takeoff-scenarios
evhub(2022). How likely is deceptive alignment?. AI Alignment Forum.evhub. (2022, August 30). How likely is deceptive alignment?. AI Alignment Forum. https://alignmentforum.org/posts/A9NxPTwbw6r6Awuwt/how-likely-is-deceptive-alignmentevhub. 2022. “How Likely Is Deceptive Alignment?”. AI Alignment Forum, August 30. https://alignmentforum.org/posts/A9NxPTwbw6r6Awuwt/how-likely-is-deceptive-alignment.evhub. “How Likely Is Deceptive Alignment?”. AI Alignment Forum, 30 Aug. 2022, https://alignmentforum.org/posts/A9NxPTwbw6r6Awuwt/how-likely-is-deceptive-alignment.evhub. How likely is deceptive alignment?. AI Alignment Forum https://alignmentforum.org/posts/A9NxPTwbw6r6Awuwt/how-likely-is-deceptive-alignment (2022).evhub, “How likely is deceptive alignment?”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/A9NxPTwbw6r6Awuwt/how-likely-is-deceptive-alignment
evhub(2023). Towards understanding-based safety evaluations. AI Alignment Forum.evhub. (2023, March 15). Towards understanding-based safety evaluations. AI Alignment Forum. https://alignmentforum.org/posts/uqAdqrvxqGqeBHjTP/towards-understanding-based-safety-evaluationsevhub. 2023. “Towards Understanding-based Safety Evaluations”. AI Alignment Forum, March 15. https://alignmentforum.org/posts/uqAdqrvxqGqeBHjTP/towards-understanding-based-safety-evaluations.evhub. “Towards Understanding-based Safety Evaluations”. AI Alignment Forum, 15 Mar. 2023, https://alignmentforum.org/posts/uqAdqrvxqGqeBHjTP/towards-understanding-based-safety-evaluations.evhub. Towards understanding-based safety evaluations. AI Alignment Forum https://alignmentforum.org/posts/uqAdqrvxqGqeBHjTP/towards-understanding-based-safety-evaluations (2023).evhub, “Towards understanding-based safety evaluations”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/uqAdqrvxqGqeBHjTP/towards-understanding-based-safety-evaluations
evhub(2023). When can we trust model evaluations?. AI Alignment Forum.evhub. (2023, July 28). When can we trust model evaluations?. AI Alignment Forum. https://alignmentforum.org/posts/dBmfb76zx6wjPsBC7/when-can-we-trust-model-evaluationsevhub. 2023. “When Can We Trust Model Evaluations?”. AI Alignment Forum, July 28. https://alignmentforum.org/posts/dBmfb76zx6wjPsBC7/when-can-we-trust-model-evaluations.evhub. “When Can We Trust Model Evaluations?”. AI Alignment Forum, 28 July 2023, https://alignmentforum.org/posts/dBmfb76zx6wjPsBC7/when-can-we-trust-model-evaluations.evhub. When can we trust model evaluations?. AI Alignment Forum https://alignmentforum.org/posts/dBmfb76zx6wjPsBC7/when-can-we-trust-model-evaluations (2023).evhub, “When can we trust model evaluations?”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/dBmfb76zx6wjPsBC7/when-can-we-trust-model-evaluations
evhub, Chris van Merwijk, Vlad Mikulik, Joar Skalse & Scott Garrabrant(2019). Conditions for Mesa-Optimization. AI Alignment Forum.evhub, Chris van Merwijk, Vlad Mikulik, Joar Skalse, & Scott Garrabrant. (2019, June 1). Conditions for Mesa-Optimization. AI Alignment Forum. https://alignmentforum.org/posts/q2rCMHNXazALgQpGH/conditions-for-mesa-optimizationevhub, Chris van Merwijk, Vlad Mikulik, Joar Skalse, and Scott Garrabrant. 2019. “Conditions for Mesa-Optimization”. AI Alignment Forum, June 1. https://alignmentforum.org/posts/q2rCMHNXazALgQpGH/conditions-for-mesa-optimization.evhub, et al. “Conditions for Mesa-Optimization”. AI Alignment Forum, 1 June 2019, https://alignmentforum.org/posts/q2rCMHNXazALgQpGH/conditions-for-mesa-optimization.evhub, Chris van Merwijk, Vlad Mikulik, Joar Skalse & Scott Garrabrant. Conditions for Mesa-Optimization. AI Alignment Forum https://alignmentforum.org/posts/q2rCMHNXazALgQpGH/conditions-for-mesa-optimization (2019).evhub, Chris van Merwijk, Vlad Mikulik, Joar Skalse, and Scott Garrabrant, “Conditions for Mesa-Optimization”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/q2rCMHNXazALgQpGH/conditions-for-mesa-optimization
evhub, Nicholas Schiefer, Carson Denison & Ethan Perez(2023). Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research. AI Alignment Forum.evhub, Nicholas Schiefer, Carson Denison, & Ethan Perez. (2023, August 8). Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research. AI Alignment Forum. https://alignmentforum.org/posts/ChDH335ckdvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1evhub, Nicholas Schiefer, Carson Denison, and Ethan Perez. 2023. “Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research”. AI Alignment Forum, August 8. https://alignmentforum.org/posts/ChDH335ckdvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1.evhub, et al. “Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research”. AI Alignment Forum, 8 Aug. 2023, https://alignmentforum.org/posts/ChDH335ckdvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1.evhub, Nicholas Schiefer, Carson Denison & Ethan Perez. Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research. AI Alignment Forum https://alignmentforum.org/posts/ChDH335ckdvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1 (2023).evhub, Nicholas Schiefer, Carson Denison, and Ethan Perez, “Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/ChDH335ckdvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1
Ewing(2017). Engineering a Deception: What Led to Volkswagen’s Diesel Scandal (Published 2021).Ewing. (2017). Engineering a Deception: What Led to Volkswagen’s Diesel Scandal (Published 2021). Internet Archive (https://web.archive.org/web/20260607173117/https://www.nytimes.com/interactive/2017/business/volkswagen-diesel-emissions-timeline.html). https://nytimes.com/interactive/2017/business/volkswagen-diesel-emissions-timeline.htmlEwing. 2017. “Engineering a Deception: What Led to Volkswagen’s Diesel Scandal (Published 2021)”. Https://web.archive.org/web/20260607173117/https://www.nytimes.com/interactive/2017/business/volkswagen-diesel-emissions-timeline.html. Internet Archive. https://nytimes.com/interactive/2017/business/volkswagen-diesel-emissions-timeline.html.Ewing. Engineering a Deception: What Led to Volkswagen’s Diesel Scandal (Published 2021). 2017, Internet Archive, https://web.archive.org/web/20260607173117/https://www.nytimes.com/interactive/2017/business/volkswagen-diesel-emissions-timeline.html, https://nytimes.com/interactive/2017/business/volkswagen-diesel-emissions-timeline.html.Ewing. Engineering a Deception: What Led to Volkswagen’s Diesel Scandal (Published 2021). https://nytimes.com/interactive/2017/business/volkswagen-diesel-emissions-timeline.html (2017).Ewing, “Engineering a Deception: What Led to Volkswagen’s Diesel Scandal (Published 2021)”. Accessed: Jun. 07, 2026. [Online]. Available: https://nytimes.com/interactive/2017/business/volkswagen-diesel-emissions-timeline.html
Eykholt, K. et al.(2017). Robust Physical-World Attacks on Deep Learning Models. arXiv.Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Xiao, C., Prakash, A., Kohno, T., & Song, D. (2017). Robust Physical-World Attacks on Deep Learning Models. In arXiv. https://arxiv.org/abs/1707.08945Eykholt, K., I. Evtimov, E. Fernandes, et al. 2017. “Robust Physical-World Attacks on Deep Learning Models”. In arXiv. Preprint, July 27. https://arxiv.org/abs/1707.08945.Eykholt, K., et al. “Robust Physical-World Attacks on Deep Learning Models”. arXiv, 27 July 2017, https://arxiv.org/abs/1707.08945.Eykholt, K. et al. Robust Physical-World Attacks on Deep Learning Models. arXiv Preprint at https://arxiv.org/abs/1707.08945 (2017).K. Eykholt et al., “Robust Physical-World Attacks on Deep Learning Models”, Jul. 27, 2017. [Online]. Available: https://arxiv.org/abs/1707.08945
Fabien Roger & Buck(2024). Toy models of AI control for concentrated catastrophe prevention. AI Alignment Forum.Fabien Roger, & Buck. (2024, February 6). Toy models of AI control for concentrated catastrophe prevention. AI Alignment Forum. https://alignmentforum.org/posts/MDeGts4Aw9DktCkXw/toy-models-of-ai-control-for-concentrated-catastropheFabien Roger, and Buck. 2024. “Toy Models of AI Control for Concentrated Catastrophe Prevention”. AI Alignment Forum, February 6. https://alignmentforum.org/posts/MDeGts4Aw9DktCkXw/toy-models-of-ai-control-for-concentrated-catastrophe.Fabien Roger, and Buck. “Toy Models of AI Control for Concentrated Catastrophe Prevention”. AI Alignment Forum, 6 Feb. 2024, https://alignmentforum.org/posts/MDeGts4Aw9DktCkXw/toy-models-of-ai-control-for-concentrated-catastrophe.Fabien Roger & Buck. Toy models of AI control for concentrated catastrophe prevention. AI Alignment Forum https://alignmentforum.org/posts/MDeGts4Aw9DktCkXw/toy-models-of-ai-control-for-concentrated-catastrophe (2024).Fabien Roger and Buck, “Toy models of AI control for concentrated catastrophe prevention”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/MDeGts4Aw9DktCkXw/toy-models-of-ai-control-for-concentrated-catastrophe
Fabien Roger & ryan_greenblatt(2023). Preventing Language Models from hiding their reasoning. AI Alignment Forum.Fabien Roger, & ryan_greenblatt. (2023, October 31). Preventing Language Models from hiding their reasoning. AI Alignment Forum. https://alignmentforum.org/posts/9Fdd9N7Escg3tcymb/preventing-language-models-from-hiding-their-reasoningFabien Roger, and ryan_greenblatt. 2023. “Preventing Language Models from Hiding Their Reasoning”. AI Alignment Forum, October 31. https://alignmentforum.org/posts/9Fdd9N7Escg3tcymb/preventing-language-models-from-hiding-their-reasoning.Fabien Roger, and ryan_greenblatt. “Preventing Language Models from Hiding Their Reasoning”. AI Alignment Forum, 31 Oct. 2023, https://alignmentforum.org/posts/9Fdd9N7Escg3tcymb/preventing-language-models-from-hiding-their-reasoning.Fabien Roger & ryan_greenblatt. Preventing Language Models from hiding their reasoning. AI Alignment Forum https://alignmentforum.org/posts/9Fdd9N7Escg3tcymb/preventing-language-models-from-hiding-their-reasoning (2023).Fabien Roger and ryan_greenblatt, “Preventing Language Models from hiding their reasoning”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/9Fdd9N7Escg3tcymb/preventing-language-models-from-hiding-their-reasoning
Fabien Roger(2023). Coup probes: Catching catastrophes with probes trained off-policy. AI Alignment Forum.Fabien Roger. (2023, November 17). Coup probes: Catching catastrophes with probes trained off-policy. AI Alignment Forum. https://alignmentforum.org/posts/WCj7WgFSLmyKaMwPR/coup-probes-catching-catastrophes-with-probes-trained-offFabien Roger. 2023. “Coup Probes: Catching Catastrophes with Probes Trained Off-policy”. AI Alignment Forum, November 17. https://alignmentforum.org/posts/WCj7WgFSLmyKaMwPR/coup-probes-catching-catastrophes-with-probes-trained-off.Fabien Roger. “Coup Probes: Catching Catastrophes with Probes Trained Off-policy”. AI Alignment Forum, 17 Nov. 2023, https://alignmentforum.org/posts/WCj7WgFSLmyKaMwPR/coup-probes-catching-catastrophes-with-probes-trained-off.Fabien Roger. Coup probes: Catching catastrophes with probes trained off-policy. AI Alignment Forum https://alignmentforum.org/posts/WCj7WgFSLmyKaMwPR/coup-probes-catching-catastrophes-with-probes-trained-off (2023).Fabien Roger, “Coup probes: Catching catastrophes with probes trained off-policy”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/WCj7WgFSLmyKaMwPR/coup-probes-catching-catastrophes-with-probes-trained-off
Fabien Roger(2023). The Translucent Thoughts Hypotheses and Their Implications. AI Alignment Forum.Fabien Roger. (2023, March 9). The Translucent Thoughts Hypotheses and Their Implications. AI Alignment Forum. https://alignmentforum.org/posts/r3xwHzMmMf25peeHE/the-translucent-thoughts-hypotheses-and-their-implicationsFabien Roger. 2023. “The Translucent Thoughts Hypotheses and Their Implications”. AI Alignment Forum, March 9. https://alignmentforum.org/posts/r3xwHzMmMf25peeHE/the-translucent-thoughts-hypotheses-and-their-implications.Fabien Roger. “The Translucent Thoughts Hypotheses and Their Implications”. AI Alignment Forum, 9 Mar. 2023, https://alignmentforum.org/posts/r3xwHzMmMf25peeHE/the-translucent-thoughts-hypotheses-and-their-implications.Fabien Roger. The Translucent Thoughts Hypotheses and Their Implications. AI Alignment Forum https://alignmentforum.org/posts/r3xwHzMmMf25peeHE/the-translucent-thoughts-hypotheses-and-their-implications (2023).Fabien Roger, “The Translucent Thoughts Hypotheses and Their Implications”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/r3xwHzMmMf25peeHE/the-translucent-thoughts-hypotheses-and-their-implications
Faggella(2023). A Worthy Successor - The Purpose of AGI - Daniel Faggella.Faggella. (2023, November 24). A Worthy Successor - The Purpose of AGI - Daniel Faggella. Daniel Faggella. https://danfaggella.com/worthyFaggella. 2023. “A Worthy Successor - The Purpose of AGI - Daniel Faggella”. Daniel Faggella, November 24. https://danfaggella.com/worthy.Faggella. “A Worthy Successor - The Purpose of AGI - Daniel Faggella”. Daniel Faggella, 24 Nov. 2023, https://danfaggella.com/worthy.Faggella. A Worthy Successor - The Purpose of AGI - Daniel Faggella. Daniel Faggella https://danfaggella.com/worthy (2023).Faggella, “A Worthy Successor - The Purpose of AGI - Daniel Faggella”, Daniel Faggella. [Online]. Available: https://danfaggella.com/worthy
Fang, R., Bindu, R., Gupta, A., Zhan, Q. & Kang, D.(2024). LLM Agents can Autonomously Hack Websites. arXiv.Fang, R., Bindu, R., Gupta, A., Zhan, Q., & Kang, D. (2024). LLM Agents can Autonomously Hack Websites. In arXiv. https://arxiv.org/abs/2402.06664Fang, R., R. Bindu, A. Gupta, Q. Zhan, and D. Kang. 2024. “LLM Agents Can Autonomously Hack Websites”. In arXiv. Preprint, February 6. https://arxiv.org/abs/2402.06664.Fang, R., et al. “LLM Agents Can Autonomously Hack Websites”. arXiv, 6 Feb. 2024, https://arxiv.org/abs/2402.06664.Fang, R., Bindu, R., Gupta, A., Zhan, Q. & Kang, D. LLM Agents can Autonomously Hack Websites. arXiv Preprint at https://arxiv.org/abs/2402.06664 (2024).R. Fang, R. Bindu, A. Gupta, Q. Zhan, and D. Kang, “LLM Agents can Autonomously Hack Websites”, Feb. 06, 2024. [Online]. Available: https://arxiv.org/abs/2402.06664
Farrell(2024). Learning from History: GPAI serious incident reporting. Pour Demain.Farrell. (2024, October 18). Learning from History: GPAI serious incident reporting. Pour Demain. https://pourdemain.ngo/en/post/learning-from-history-gpai-serious-incident-reportingFarrell. 2024. “Learning from History: GPAI Serious Incident Reporting”. Pour Demain, October 18. https://pourdemain.ngo/en/post/learning-from-history-gpai-serious-incident-reporting.Farrell. “Learning from History: GPAI Serious Incident Reporting”. Pour Demain, 18 Oct. 2024, https://pourdemain.ngo/en/post/learning-from-history-gpai-serious-incident-reporting.Farrell. Learning from History: GPAI serious incident reporting. Pour Demain https://pourdemain.ngo/en/post/learning-from-history-gpai-serious-incident-reporting (2024).Farrell, “Learning from History: GPAI serious incident reporting”, Pour Demain. [Online]. Available: https://pourdemain.ngo/en/post/learning-from-history-gpai-serious-incident-reporting
Federspiel, F., Mitchell, R., Asokan, A., Umana, C. & McCoy, D.(2023). Threats by artificial intelligence to human health and human existence. BMJ Global Health.Federspiel, F., Mitchell, R., Asokan, A., Umana, C., & McCoy, D. (2023). Threats by artificial intelligence to human health and human existence. BMJ Global Health. https://doi.org/10.1136/bmjgh-2022-010435Federspiel, F., R. Mitchell, A. Asokan, C. Umana, and D. McCoy. 2023. “Threats by Artificial Intelligence to Human Health and Human Existence”. BMJ Global Health, ahead of print, May. https://doi.org/10.1136/bmjgh-2022-010435.Federspiel, F., et al. “Threats by Artificial Intelligence to Human Health and Human Existence”. BMJ Global Health, May 2023, https://doi.org/10.1136/bmjgh-2022-010435.Federspiel, F., Mitchell, R., Asokan, A., Umana, C. & McCoy, D. Threats by artificial intelligence to human health and human existence. BMJ Global Health https://doi.org/10.1136/bmjgh-2022-010435 (2023) doi:10.1136/bmjgh-2022-010435.F. Federspiel, R. Mitchell, A. Asokan, C. Umana, and D. McCoy, “Threats by artificial intelligence to human health and human existence”, BMJ Global Health, May 2023, doi: 10.1136/bmjgh-2022-010435.
Feldstein, S.(2019). The Global Expansion of AI Surveillance.Feldstein, S. (2019). The Global Expansion of AI Surveillance. Carnegie Endowment for International Peace. https://carnegieendowment.org/research/2019/09/the-global-expansion-of-ai-surveillance?lang=enFeldstein, S. 2019. The Global Expansion of AI Surveillance. Carnegie Endowment for International Peace. https://carnegieendowment.org/research/2019/09/the-global-expansion-of-ai-surveillance?lang=en.Feldstein, S. The Global Expansion of AI Surveillance. Carnegie Endowment for International Peace, Sept. 2019, https://carnegieendowment.org/research/2019/09/the-global-expansion-of-ai-surveillance?lang=en.Feldstein, S. The Global Expansion of AI Surveillance. https://carnegieendowment.org/research/2019/09/the-global-expansion-of-ai-surveillance?lang=en (2019).S. Feldstein, “The Global Expansion of AI Surveillance”, Carnegie Endowment for International Peace, Sep. 2019. [Online]. Available: https://carnegieendowment.org/research/2019/09/the-global-expansion-of-ai-surveillance?lang=en
Feng, S. & Tramèr, F.(2024). Privacy Backdoors: Stealing Data with Corrupted Pretrained Models. arXiv.Feng, S., & Tramèr, F. (2024). Privacy Backdoors: Stealing Data with Corrupted Pretrained Models. In arXiv. https://arxiv.org/abs/2404.00473Feng, S., and F. Tramèr. 2024. “Privacy Backdoors: Stealing Data with Corrupted Pretrained Models”. In arXiv. Preprint, March 30. https://arxiv.org/abs/2404.00473.Feng, S., and F. Tramèr. “Privacy Backdoors: Stealing Data with Corrupted Pretrained Models”. arXiv, 30 Mar. 2024, https://arxiv.org/abs/2404.00473.Feng, S. & Tramèr, F. Privacy Backdoors: Stealing Data with Corrupted Pretrained Models. arXiv Preprint at https://arxiv.org/abs/2404.00473 (2024).S. Feng and F. Tramèr, “Privacy Backdoors: Stealing Data with Corrupted Pretrained Models”, Mar. 30, 2024. [Online]. Available: https://arxiv.org/abs/2404.00473
Fenwick(2023). Want to shape AI policy? Consider working in the US government. 80,000 Hours.Fenwick. (2023). Want to shape AI policy? Consider working in the US government. 80,000 Hours. https://80000hours.org/career-reviews/ai-policy-and-strategyFenwick. 2023. “Want to Shape AI Policy? Consider Working in the US Government”. 80,000 Hours. https://80000hours.org/career-reviews/ai-policy-and-strategy.Fenwick. “Want to Shape AI Policy? Consider Working in the US Government”. 80,000 Hours, 2023, https://80000hours.org/career-reviews/ai-policy-and-strategy.Fenwick. Want to shape AI policy? Consider working in the US government. 80,000 Hours https://80000hours.org/career-reviews/ai-policy-and-strategy (2023).Fenwick, “Want to shape AI policy? Consider working in the US government”, 80,000 Hours. [Online]. Available: https://80000hours.org/career-reviews/ai-policy-and-strategy
Field(2025). Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts. arXiv.org.Field. (2025). Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts. arXiv.org. https://www.arxiv.org/abs/2502.14870Field. 2025. “Why Do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts”. arXiv.org. https://www.arxiv.org/abs/2502.14870.Field. “Why Do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts”. arXiv.org, 2025, https://www.arxiv.org/abs/2502.14870.Field. Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts. arXiv.org https://www.arxiv.org/abs/2502.14870 (2025).Field, “Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2502.14870
Flanagan, D. P. & Dixon, S. G.(2014). The Cattell‐Horn‐Carroll Theory of Cognitive Abilities. Encyclopedia of Special Education.Flanagan, D. P., & Dixon, S. G. (2014, January 22). The Cattell‐Horn‐Carroll Theory of Cognitive Abilities. In Encyclopedia of Special Education. https://doi.org/10.1002/9781118660584.ese0431Flanagan, D. P., and S. G. Dixon. 2014. “The Cattell‐Horn‐Carroll Theory of Cognitive Abilities”. In Encyclopedia of Special Education. January 22. https://doi.org/10.1002/9781118660584.ese0431.Flanagan, D. P., and S. G. Dixon. “The Cattell‐Horn‐Carroll Theory of Cognitive Abilities”. Encyclopedia of Special Education, 22 Jan. 2014, https://doi.org/10.1002/9781118660584.ese0431.Flanagan, D. P. & Dixon, S. G. The Cattell‐Horn‐Carroll Theory of Cognitive Abilities. Encyclopedia of Special Education (2014) doi:10.1002/9781118660584.ese0431.D. P. Flanagan and S. G. Dixon, “The Cattell‐Horn‐Carroll Theory of Cognitive Abilities”, Encyclopedia of Special Education. Jan. 22, 2014. doi: 10.1002/9781118660584.ese0431.
Fluri, L., Paleka, D. & Tramèr, F.(2023). Evaluating Superhuman Models with Consistency Checks. arXiv.Fluri, L., Paleka, D., & Tramèr, F. (2023). Evaluating Superhuman Models with Consistency Checks. In arXiv. https://arxiv.org/abs/2306.09983Fluri, L., D. Paleka, and F. Tramèr. 2023. “Evaluating Superhuman Models with Consistency Checks”. In arXiv. Preprint, June 16. https://arxiv.org/abs/2306.09983.Fluri, L., et al. “Evaluating Superhuman Models with Consistency Checks”. arXiv, 16 June 2023, https://arxiv.org/abs/2306.09983.Fluri, L., Paleka, D. & Tramèr, F. Evaluating Superhuman Models with Consistency Checks. arXiv Preprint at https://arxiv.org/abs/2306.09983 (2023).L. Fluri, D. Paleka, and F. Tramèr, “Evaluating Superhuman Models with Consistency Checks”, Jun. 16, 2023. [Online]. Available: https://arxiv.org/abs/2306.09983
Friedman, M.. The Social Responsibility of Business Is to Increase Its Profits. Corporate Ethics and Corporate Governance.Friedman, M. (n.d.). The Social Responsibility of Business Is to Increase Its Profits. In Corporate Ethics and Corporate Governance. https://doi.org/10.1007/978-3-540-70818-6_14Friedman, M. n.d. “The Social Responsibility of Business Is to Increase Its Profits”. In Corporate Ethics and Corporate Governance. https://doi.org/10.1007/978-3-540-70818-6_14.Friedman, M. “The Social Responsibility of Business Is to Increase Its Profits”. Corporate Ethics and Corporate Governance, https://doi.org/10.1007/978-3-540-70818-6_14.Friedman, M. The Social Responsibility of Business Is to Increase Its Profits. Corporate Ethics and Corporate Governance doi:10.1007/978-3-540-70818-6_14.M. Friedman, “The Social Responsibility of Business Is to Increase Its Profits”, Corporate Ethics and Corporate Governance. doi: 10.1007/978-3-540-70818-6_14.
Friston, K. J. et al.(2022). Designing Ecosystems of Intelligence from First Principles. arXiv.Friston, K. J., Ramstead, M. J. D., Kiefer, A. B., Tschantz, A., Buckley, C. L., Albarracin, M., Pitliya, R. J., Heins, C., Klein, B., Millidge, B., Sakthivadivel, D. A. R., Smithe, T. S. C., Koudahl, M., Tremblay, S. E., Petersen, C., Fung, K., Fox, J. G., Swanson, S., Mapes, D., & René, G. (2022). Designing Ecosystems of Intelligence from First Principles. In arXiv. https://doi.org/10.1177/26339137231222481Friston, K. J., M. J. D. Ramstead, A. B. Kiefer, et al. 2022. “Designing Ecosystems of Intelligence from First Principles”. In arXiv. Preprint, December 2. https://doi.org/10.1177/26339137231222481.Friston, K. J., et al. “Designing Ecosystems of Intelligence from First Principles”. arXiv, 2 Dec. 2022, https://doi.org/10.1177/26339137231222481.Friston, K. J. et al. Designing Ecosystems of Intelligence from First Principles. arXiv Preprint at https://doi.org/10.1177/26339137231222481 (2022).K. J. Friston et al., “Designing Ecosystems of Intelligence from First Principles”, Dec. 02, 2022. doi: 10.1177/26339137231222481.
Fu, Z., Zhao, T. Z. & Finn, C.(2024). Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. arXiv.Fu, Z., Zhao, T. Z., & Finn, C. (2024). Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. In arXiv. https://arxiv.org/abs/2401.02117Fu, Z., T. Z. Zhao, and C. Finn. 2024. “Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation”. In arXiv. Preprint, January 4. https://arxiv.org/abs/2401.02117.Fu, Z., et al. “Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation”. arXiv, 4 Jan. 2024, https://arxiv.org/abs/2401.02117.Fu, Z., Zhao, T. Z. & Finn, C. Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. arXiv Preprint at https://arxiv.org/abs/2401.02117 (2024).Z. Fu, T. Z. Zhao, and C. Finn, “Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation”, Jan. 04, 2024. [Online]. Available: https://arxiv.org/abs/2401.02117
Future of Life Institute(2024). Artificial Escalation. YouTube.Future of Life Institute. (2024). Artificial Escalation [Video recording]. In YouTube. https://www.youtube.com/watch?v=w9npWiTOHX0Future of Life Institute. 2024. “Artificial Escalation”. YouTube. https://www.youtube.com/watch?v=w9npWiTOHX0.Future of Life Institute. “Artificial Escalation”. YouTube, 2024, https://www.youtube.com/watch?v=w9npWiTOHX0.Future of Life Institute. Artificial Escalation. YouTube (2024).Future of Life Institute, Artificial Escalation, (2024). [Online Video]. Available: https://www.youtube.com/watch?v=w9npWiTOHX0
Future of Life Institute(2024). Gradual AI Disempowerment. Future of Life Institute.Future of Life Institute. (2024, February 1). Gradual AI Disempowerment. Future of Life Institute. https://futureoflife.org/existential-risk/gradual-ai-disempowermentFuture of Life Institute. 2024. “Gradual AI Disempowerment”. Future of Life Institute, February 1. https://futureoflife.org/existential-risk/gradual-ai-disempowerment.Future of Life Institute. “Gradual AI Disempowerment”. Future of Life Institute, 1 Feb. 2024, https://futureoflife.org/existential-risk/gradual-ai-disempowerment.Future of Life Institute. Gradual AI Disempowerment. Future of Life Institute https://futureoflife.org/existential-risk/gradual-ai-disempowerment (2024).Future of Life Institute, “Gradual AI Disempowerment”, Future of Life Institute. [Online]. Available: https://futureoflife.org/existential-risk/gradual-ai-disempowerment
Future of Life Institute(2024). Holly Elmore on Pausing AI, Hardware Overhang, Safety Research, and Protesting. YouTube.Future of Life Institute. (2024). Holly Elmore on Pausing AI, Hardware Overhang, Safety Research, and Protesting [Video recording]. In YouTube. https://www.youtube.com/watch?v=Q3eRy4t2oPQFuture of Life Institute. 2024. “Holly Elmore on Pausing AI, Hardware Overhang, Safety Research, and Protesting”. YouTube. https://www.youtube.com/watch?v=Q3eRy4t2oPQ.Future of Life Institute. “Holly Elmore on Pausing AI, Hardware Overhang, Safety Research, and Protesting”. YouTube, 2024, https://www.youtube.com/watch?v=Q3eRy4t2oPQ.Future of Life Institute. Holly Elmore on Pausing AI, Hardware Overhang, Safety Research, and Protesting. YouTube (2024).Future of Life Institute, Holly Elmore on Pausing AI, Hardware Overhang, Safety Research, and Protesting, (2024). [Online Video]. Available: https://www.youtube.com/watch?v=Q3eRy4t2oPQ
Future of Life Institute(2025). AI Safety Index, Summer 2025.Future of Life Institute. (2025). AI Safety Index, Summer 2025. Future of Life Institute. https://futureoflife.org/wp-content/uploads/2025/07/FLI-AI-Safety-Index-Report-Summer-2025.pdfFuture of Life Institute. 2025. AI Safety Index, Summer 2025. Future of Life Institute. https://futureoflife.org/wp-content/uploads/2025/07/FLI-AI-Safety-Index-Report-Summer-2025.pdf.Future of Life Institute. AI Safety Index, Summer 2025. Future of Life Institute, 17 July 2025, https://futureoflife.org/wp-content/uploads/2025/07/FLI-AI-Safety-Index-Report-Summer-2025.pdf.Future of Life Institute. AI Safety Index, Summer 2025. https://futureoflife.org/wp-content/uploads/2025/07/FLI-AI-Safety-Index-Report-Summer-2025.pdf (2025).Future of Life Institute, “AI Safety Index, Summer 2025”, Future of Life Institute, Jul. 2025. [Online]. Available: https://futureoflife.org/wp-content/uploads/2025/07/FLI-AI-Safety-Index-Report-Summer-2025.pdf
Future of Life Institute(2025). Can Defense in Depth Work for AI? (with Adam Gleave). YouTube.Future of Life Institute. (2025). Can Defense in Depth Work for AI? (with Adam Gleave) [Video recording]. In YouTube. https://www.youtube.com/watch?v=BfXi3_QSWekFuture of Life Institute. 2025. “Can Defense in Depth Work for AI? (with Adam Gleave)”. YouTube. https://www.youtube.com/watch?v=BfXi3_QSWek.Future of Life Institute. “Can Defense in Depth Work for AI? (with Adam Gleave)”. YouTube, 2025, https://www.youtube.com/watch?v=BfXi3_QSWek.Future of Life Institute. Can Defense in Depth Work for AI? (with Adam Gleave). YouTube (2025).Future of Life Institute, Can Defense in Depth Work for AI? (with Adam Gleave), (2025). [Online Video]. Available: https://www.youtube.com/watch?v=BfXi3_QSWek
Gabriel, I. et al.(2024). The Ethics of Advanced AI Assistants. arXiv.Gabriel, I., Manzini, A., Keeling, G., Hendricks, L. A., Rieser, V., Iqbal, H., Tomašev, N., Ktena, I., Kenton, Z., Rodriguez, M., El-Sayed, S., Brown, S., Akbulut, C., Trask, A., Hughes, E., Bergman, A. S., Shelby, R., Marchal, N., Griffin, C., … Manyika, J. (2024). The Ethics of Advanced AI Assistants. In arXiv. https://arxiv.org/abs/2404.16244Gabriel, I., A. Manzini, G. Keeling, et al. 2024. “The Ethics of Advanced AI Assistants”. In arXiv. Preprint, April 24. https://arxiv.org/abs/2404.16244.Gabriel, I., et al. “The Ethics of Advanced AI Assistants”. arXiv, 24 Apr. 2024, https://arxiv.org/abs/2404.16244.Gabriel, I. et al. The Ethics of Advanced AI Assistants. arXiv Preprint at https://arxiv.org/abs/2404.16244 (2024).I. Gabriel et al., “The Ethics of Advanced AI Assistants”, Apr. 24, 2024. [Online]. Available: https://arxiv.org/abs/2404.16244
Game Thinking TV(2023). Gödel, Escher, Bach author Doug Hofstadter on the state of AI today. YouTube.Game Thinking TV. (2023). Gödel, Escher, Bach author Doug Hofstadter on the state of AI today [Video recording]. In YouTube. https://www.youtube.com/watch?v=lfXxzAVtdpUGame Thinking TV. 2023. “Gödel, Escher, Bach Author Doug Hofstadter on the State of AI Today”. YouTube. https://www.youtube.com/watch?v=lfXxzAVtdpU.Game Thinking TV. “Gödel, Escher, Bach Author Doug Hofstadter on the State of AI Today”. YouTube, 2023, https://www.youtube.com/watch?v=lfXxzAVtdpU.Game Thinking TV. Gödel, Escher, Bach Author Doug Hofstadter on the State of AI Today. YouTube (2023).Game Thinking TV, Gödel, Escher, Bach author Doug Hofstadter on the state of AI today, (2023). [Online Video]. Available: https://www.youtube.com/watch?v=lfXxzAVtdpU
Ganguli, D. et al.(2022). Predictability and Surprise in Large Generative Models. arXiv.Ganguli, D., Hernandez, D., Lovitt, L., DasSarma, N., Henighan, T., Jones, A., Joseph, N., Kernion, J., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Elhage, N., Showk, S. E., Fort, S., Hatfield-Dodds, Z., Johnston, S., … Clark, J. (2022). Predictability and Surprise in Large Generative Models. In arXiv. https://doi.org/10.1145/3531146.3533229Ganguli, D., D. Hernandez, L. Lovitt, et al. 2022. “Predictability and Surprise in Large Generative Models”. In arXiv. Preprint, February 15. https://doi.org/10.1145/3531146.3533229.Ganguli, D., et al. “Predictability and Surprise in Large Generative Models”. arXiv, 15 Feb. 2022, https://doi.org/10.1145/3531146.3533229.Ganguli, D. et al. Predictability and Surprise in Large Generative Models. arXiv Preprint at https://doi.org/10.1145/3531146.3533229 (2022).D. Ganguli et al., “Predictability and Surprise in Large Generative Models”, Feb. 15, 2022. doi: 10.1145/3531146.3533229.
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A. & Brendel, W.(2018). ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. arXiv.Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., & Brendel, W. (2018). ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In arXiv. https://arxiv.org/abs/1811.12231Geirhos, R., P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel. 2018. “ImageNet-trained CNNs Are Biased Towards Texture; Increasing Shape Bias Improves Accuracy and Robustness”. In arXiv. Preprint, November 29. https://arxiv.org/abs/1811.12231.Geirhos, R., et al. “ImageNet-trained CNNs Are Biased Towards Texture; Increasing Shape Bias Improves Accuracy and Robustness”. arXiv, 29 Nov. 2018, https://arxiv.org/abs/1811.12231.Geirhos, R. et al. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. arXiv Preprint at https://arxiv.org/abs/1811.12231 (2018).R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel, “ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness”, Nov. 29, 2018. [Online]. Available: https://arxiv.org/abs/1811.12231
Gennari et al.(2024). Considerations for Evaluating Large Language Models for Cybersecurity Tasks | CMU Software Engineering Institute. SEI Digital Library.Gennari et al. (2024). Considerations for Evaluating Large Language Models for Cybersecurity Tasks | CMU Software Engineering Institute. SEI Digital Library. https://insights.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasksGennari et al. 2024. “Considerations for Evaluating Large Language Models for Cybersecurity Tasks | CMU Software Engineering Institute”. SEI Digital Library. https://insights.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasks.Gennari et al. “Considerations for Evaluating Large Language Models for Cybersecurity Tasks | CMU Software Engineering Institute”. SEI Digital Library, 2024, https://insights.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasks.Gennari et al. Considerations for Evaluating Large Language Models for Cybersecurity Tasks | CMU Software Engineering Institute. SEI Digital Library https://insights.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasks (2024).Gennari et al., “Considerations for Evaluating Large Language Models for Cybersecurity Tasks | CMU Software Engineering Institute”, SEI Digital Library. [Online]. Available: https://insights.sei.cmu.edu/library/considerations-for-evaluating-large-language-models-for-cybersecurity-tasks
Giattino et al.(2023). Artificial Intelligence. Our World in Data.Giattino et al. (2023). Artificial Intelligence. Our World in Data. https://ourworldindata.org/artificial-intelligenceGiattino et al. 2023. “Artificial Intelligence”. Our World in Data. https://ourworldindata.org/artificial-intelligence.Giattino et al. “Artificial Intelligence”. Our World in Data, 2023, https://ourworldindata.org/artificial-intelligence.Giattino et al. Artificial Intelligence. Our World in Data https://ourworldindata.org/artificial-intelligence (2023).Giattino et al., “Artificial Intelligence”, Our World in Data. [Online]. Available: https://ourworldindata.org/artificial-intelligence
Giattino et al.(2023). Language-based AI systems have grown rapidly in recent years. Our World in Data.Giattino et al. (2023). Language-based AI systems have grown rapidly in recent years. Our World in Data. https://ourworldindata.org/grapher/cumulative-number-of-large-scale-ai-models-by-domainGiattino et al. 2023. “Language-based AI Systems Have Grown Rapidly in Recent Years”. Our World in Data. https://ourworldindata.org/grapher/cumulative-number-of-large-scale-ai-models-by-domain.Giattino et al. “Language-based AI Systems Have Grown Rapidly in Recent Years”. Our World in Data, 2023, https://ourworldindata.org/grapher/cumulative-number-of-large-scale-ai-models-by-domain.Giattino et al. Language-based AI systems have grown rapidly in recent years. Our World in Data https://ourworldindata.org/grapher/cumulative-number-of-large-scale-ai-models-by-domain (2023).Giattino et al., “Language-based AI systems have grown rapidly in recent years”, Our World in Data. [Online]. Available: https://ourworldindata.org/grapher/cumulative-number-of-large-scale-ai-models-by-domain
Giattino et al.(2023). ourworldindata.org/grapher/market-…ction-manufacturing-stage?tab=chart.Giattino et al. (2023). Giattino et al., 2023. https://ourworldindata.org/grapher/market-share-logic-chip-production-manufacturing-stage?tab=chartGiattino et al. 2023. “Giattino Et Al., 2023”. https://ourworldindata.org/grapher/market-share-logic-chip-production-manufacturing-stage?tab=chart.Giattino et al. Giattino Et Al., 2023. 2023, https://ourworldindata.org/grapher/market-share-logic-chip-production-manufacturing-stage?tab=chart.Giattino et al. Giattino et al., 2023. https://ourworldindata.org/grapher/market-share-logic-chip-production-manufacturing-stage?tab=chart (2023).Giattino et al., “Giattino et al., 2023”. [Online]. Available: https://ourworldindata.org/grapher/market-share-logic-chip-production-manufacturing-stage?tab=chart
Gil(2023). Don't Call It AI Alignment. EA Forum.Gil. (2023). Don't Call It AI Alignment. EA Forum. Internet Archive (https://web.archive.org/web/20260516150521/https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment). https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignmentGil. 2023. “Don't Call It AI Alignment”. EA Forum. Https://web.archive.org/web/20260516150521/https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment. Internet Archive. https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment.Gil. “Don't Call It AI Alignment”. EA Forum, 2023, Internet Archive, https://web.archive.org/web/20260516150521/https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment, https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment.Gil. Don't Call It AI Alignment. EA Forum https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment (2023).Gil, “Don't Call It AI Alignment”, EA Forum. Accessed: May 16, 2026. [Online]. Available: https://forum.effectivealtruism.org/posts/6aYfWyo9DKEheogf8/don-t-call-it-ai-alignment
Glaese, A. et al.(2022). Improving alignment of dialogue agents via targeted human judgements. arXiv.Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., … Irving, G. (2022). Improving alignment of dialogue agents via targeted human judgements. In arXiv. https://arxiv.org/abs/2209.14375Glaese, A., N. McAleese, M. Trębacz, et al. 2022. “Improving Alignment of Dialogue Agents via Targeted Human Judgements”. In arXiv. Preprint, September 28. https://arxiv.org/abs/2209.14375.Glaese, A., et al. “Improving Alignment of Dialogue Agents via Targeted Human Judgements”. arXiv, 28 Sept. 2022, https://arxiv.org/abs/2209.14375.Glaese, A. et al. Improving alignment of dialogue agents via targeted human judgements. arXiv Preprint at https://arxiv.org/abs/2209.14375 (2022).A. Glaese et al., “Improving alignment of dialogue agents via targeted human judgements”, Sep. 28, 2022. [Online]. Available: https://arxiv.org/abs/2209.14375
Glazer, E. et al.(2024). FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI. arXiv.Glazer, E., Erdil, E., Besiroglu, T., Chicharro, D., Chen, E., Gunning, A., Olsson, C. F., Denain, J.-S., Ho, A., Santos, E. de O., Järviniemi, O., Barnett, M., Sandler, R., Vrzala, M., Sevilla, J., Ren, Q., Pratt, E., Levine, L., Barkley, G., … Wildon, M. (2024). FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI. In arXiv. https://arxiv.org/abs/2411.04872Glazer, E., E. Erdil, T. Besiroglu, et al. 2024. “FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI”. In arXiv. Preprint, November 7. https://arxiv.org/abs/2411.04872.Glazer, E., et al. “FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI”. arXiv, 7 Nov. 2024, https://arxiv.org/abs/2411.04872.Glazer, E. et al. FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI. arXiv Preprint at https://arxiv.org/abs/2411.04872 (2024).E. Glazer et al., “FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI”, Nov. 07, 2024. [Online]. Available: https://arxiv.org/abs/2411.04872
Gnanasambandam, A., Sherman, A. M. & Chan, S. H.(2021). Optical Adversarial Attack. arXiv.Gnanasambandam, A., Sherman, A. M., & Chan, S. H. (2021). Optical Adversarial Attack. In arXiv. https://arxiv.org/abs/2108.06247Gnanasambandam, A., A. M. Sherman, and S. H. Chan. 2021. “Optical Adversarial Attack”. In arXiv. Preprint, August 13. https://arxiv.org/abs/2108.06247.Gnanasambandam, A., et al. “Optical Adversarial Attack”. arXiv, 13 Aug. 2021, https://arxiv.org/abs/2108.06247.Gnanasambandam, A., Sherman, A. M. & Chan, S. H. Optical Adversarial Attack. arXiv Preprint at https://arxiv.org/abs/2108.06247 (2021).A. Gnanasambandam, A. M. Sherman, and S. H. Chan, “Optical Adversarial Attack”, Aug. 13, 2021. [Online]. Available: https://arxiv.org/abs/2108.06247
Goertzel & Pitt(2012). Original file was NineWaysToFriendlyAI_v6.tex.Goertzel & Pitt. (2012). Original file was NineWaysToFriendlyAI_v6.tex. https://jetpress.org/v22/goertzel-pitt.htmGoertzel & Pitt. 2012. “Original File Was NineWaysToFriendlyAI_v6.tex”. https://jetpress.org/v22/goertzel-pitt.htm.Goertzel & Pitt. Original File Was NineWaysToFriendlyAI_v6.tex. 2012, https://jetpress.org/v22/goertzel-pitt.htm.Goertzel & Pitt. Original file was NineWaysToFriendlyAI_v6.tex. https://jetpress.org/v22/goertzel-pitt.htm (2012).Goertzel & Pitt, “Original file was NineWaysToFriendlyAI_v6.tex”. [Online]. Available: https://jetpress.org/v22/goertzel-pitt.htm
Goertzel(2010). Coherent Aggregated Volition: A Method for Deriving Goal System Content for Advanced, Beneficial AGIs.Goertzel. (2010). Coherent Aggregated Volition: A Method for Deriving Goal System Content for Advanced, Beneficial AGIs. https://multiverseaccordingtoben.blogspot.com/2010/03/coherent-aggregated-volition-toward.htmlGoertzel. 2010. “Coherent Aggregated Volition: A Method for Deriving Goal System Content for Advanced, Beneficial AGIs”. https://multiverseaccordingtoben.blogspot.com/2010/03/coherent-aggregated-volition-toward.html.Goertzel. Coherent Aggregated Volition: A Method for Deriving Goal System Content for Advanced, Beneficial AGIs. 2010, https://multiverseaccordingtoben.blogspot.com/2010/03/coherent-aggregated-volition-toward.html.Goertzel. Coherent Aggregated Volition: A Method for Deriving Goal System Content for Advanced, Beneficial AGIs. https://multiverseaccordingtoben.blogspot.com/2010/03/coherent-aggregated-volition-toward.html (2010).Goertzel, “Coherent Aggregated Volition: A Method for Deriving Goal System Content for Advanced, Beneficial AGIs”. [Online]. Available: https://multiverseaccordingtoben.blogspot.com/2010/03/coherent-aggregated-volition-toward.html
Goertzel, B. et al.(2023). OpenCog Hyperon: A Framework for AGI at the Human Level and Beyond. arXiv.Goertzel, B., Bogdanov, V., Duncan, M., Duong, D., Goertzel, Z., Horlings, J., Ikle', M., Meredith, L. G., Potapov, A., de Senna, A. L., Suarez, H. S. A., Vandervorst, A., & Werko, R. (2023). OpenCog Hyperon: A Framework for AGI at the Human Level and Beyond. In arXiv. https://arxiv.org/abs/2310.18318Goertzel, B., V. Bogdanov, M. Duncan, et al. 2023. “OpenCog Hyperon: A Framework for AGI at the Human Level and Beyond”. In arXiv. Preprint, September 19. https://arxiv.org/abs/2310.18318.Goertzel, B., et al. “OpenCog Hyperon: A Framework for AGI at the Human Level and Beyond”. arXiv, 19 Sept. 2023, https://arxiv.org/abs/2310.18318.Goertzel, B. et al. OpenCog Hyperon: A Framework for AGI at the Human Level and Beyond. arXiv Preprint at https://arxiv.org/abs/2310.18318 (2023).B. Goertzel et al., “OpenCog Hyperon: A Framework for AGI at the Human Level and Beyond”, Sep. 19, 2023. [Online]. Available: https://arxiv.org/abs/2310.18318
Goldie, A., Mirhoseini, A. & Dean, J.(2024). That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design. arXiv.Goldie, A., Mirhoseini, A., & Dean, J. (2024). That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design. In arXiv. https://arxiv.org/abs/2411.10053Goldie, A., A. Mirhoseini, and J. Dean. 2024. “That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design”. In arXiv. Preprint, November 15. https://arxiv.org/abs/2411.10053.Goldie, A., et al. “That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design”. arXiv, 15 Nov. 2024, https://arxiv.org/abs/2411.10053.Goldie, A., Mirhoseini, A. & Dean, J. That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design. arXiv Preprint at https://arxiv.org/abs/2411.10053 (2024).A. Goldie, A. Mirhoseini, and J. Dean, “That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design”, Nov. 15, 2024. [Online]. Available: https://arxiv.org/abs/2411.10053
Goldowsky-Dill, N., Chughtai, B., Heimersheim, S. & Hobbhahn, M.(2025). Detecting Strategic Deception Using Linear Probes. arXiv.Goldowsky-Dill, N., Chughtai, B., Heimersheim, S., & Hobbhahn, M. (2025). Detecting Strategic Deception Using Linear Probes. In arXiv. https://arxiv.org/abs/2502.03407Goldowsky-Dill, N., B. Chughtai, S. Heimersheim, and M. Hobbhahn. 2025. “Detecting Strategic Deception Using Linear Probes”. In arXiv. Preprint, February 5. https://arxiv.org/abs/2502.03407.Goldowsky-Dill, N., et al. “Detecting Strategic Deception Using Linear Probes”. arXiv, 5 Feb. 2025, https://arxiv.org/abs/2502.03407.Goldowsky-Dill, N., Chughtai, B., Heimersheim, S. & Hobbhahn, M. Detecting Strategic Deception Using Linear Probes. arXiv Preprint at https://arxiv.org/abs/2502.03407 (2025).N. Goldowsky-Dill, B. Chughtai, S. Heimersheim, and M. Hobbhahn, “Detecting Strategic Deception Using Linear Probes”, Feb. 05, 2025. [Online]. Available: https://arxiv.org/abs/2502.03407
Goldwasser, S., Kim, M. P., Vaikuntanathan, V. & Zamir, O.(2022). Planting Undetectable Backdoors in Machine Learning Models. arXiv.Goldwasser, S., Kim, M. P., Vaikuntanathan, V., & Zamir, O. (2022). Planting Undetectable Backdoors in Machine Learning Models. In arXiv. https://arxiv.org/abs/2204.06974Goldwasser, S., M. P. Kim, V. Vaikuntanathan, and O. Zamir. 2022. “Planting Undetectable Backdoors in Machine Learning Models”. In arXiv. Preprint, April 14. https://arxiv.org/abs/2204.06974.Goldwasser, S., et al. “Planting Undetectable Backdoors in Machine Learning Models”. arXiv, 14 Apr. 2022, https://arxiv.org/abs/2204.06974.Goldwasser, S., Kim, M. P., Vaikuntanathan, V. & Zamir, O. Planting Undetectable Backdoors in Machine Learning Models. arXiv Preprint at https://arxiv.org/abs/2204.06974 (2022).S. Goldwasser, M. P. Kim, V. Vaikuntanathan, and O. Zamir, “Planting Undetectable Backdoors in Machine Learning Models”, Apr. 14, 2022. [Online]. Available: https://arxiv.org/abs/2204.06974
Golovneva, O., Allen-Zhu, Z., Weston, J. & Sukhbaatar, S.(2024). Reverse Training to Nurse the Reversal Curse. arXiv.Golovneva, O., Allen-Zhu, Z., Weston, J., & Sukhbaatar, S. (2024). Reverse Training to Nurse the Reversal Curse. In arXiv. https://arxiv.org/abs/2403.13799Golovneva, O., Z. Allen-Zhu, J. Weston, and S. Sukhbaatar. 2024. “Reverse Training to Nurse the Reversal Curse”. In arXiv. Preprint, March 20. https://arxiv.org/abs/2403.13799.Golovneva, O., et al. “Reverse Training to Nurse the Reversal Curse”. arXiv, 20 Mar. 2024, https://arxiv.org/abs/2403.13799.Golovneva, O., Allen-Zhu, Z., Weston, J. & Sukhbaatar, S. Reverse Training to Nurse the Reversal Curse. arXiv Preprint at https://arxiv.org/abs/2403.13799 (2024).O. Golovneva, Z. Allen-Zhu, J. Weston, and S. Sukhbaatar, “Reverse Training to Nurse the Reversal Curse”, Mar. 20, 2024. [Online]. Available: https://arxiv.org/abs/2403.13799
Goodfellow, I. J. et al.(2014). Generative Adversarial Networks. arXiv.Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative Adversarial Networks. In arXiv. https://arxiv.org/abs/1406.2661Goodfellow, I. J., J. Pouget-Abadie, M. Mirza, et al. 2014. “Generative Adversarial Networks”. In arXiv. Preprint, June 10. https://arxiv.org/abs/1406.2661.Goodfellow, I. J., et al. “Generative Adversarial Networks”. arXiv, 10 June 2014, https://arxiv.org/abs/1406.2661.Goodfellow, I. J. et al. Generative Adversarial Networks. arXiv Preprint at https://arxiv.org/abs/1406.2661 (2014).I. J. Goodfellow et al., “Generative Adversarial Networks”, Jun. 10, 2014. [Online]. Available: https://arxiv.org/abs/1406.2661
Goodfellow, I. J., Shlens, J. & Szegedy, C.(2014). Explaining and Harnessing Adversarial Examples. arXiv.Goodfellow, I. J., Shlens, J., & Szegedy, C. (2014). Explaining and Harnessing Adversarial Examples. In arXiv. https://arxiv.org/abs/1412.6572Goodfellow, I. J., J. Shlens, and C. Szegedy. 2014. “Explaining and Harnessing Adversarial Examples”. In arXiv. Preprint, December 20. https://arxiv.org/abs/1412.6572.Goodfellow, I. J., et al. “Explaining and Harnessing Adversarial Examples”. arXiv, 20 Dec. 2014, https://arxiv.org/abs/1412.6572.Goodfellow, I. J., Shlens, J. & Szegedy, C. Explaining and Harnessing Adversarial Examples. arXiv Preprint at https://arxiv.org/abs/1412.6572 (2014).I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and Harnessing Adversarial Examples”, Dec. 20, 2014. [Online]. Available: https://arxiv.org/abs/1412.6572
Goodhart, C. A. E.(1984). Problems of Monetary Management: The UK Experience. Monetary Theory and Practice.Goodhart, C. A. E. (1984). Problems of Monetary Management: The UK Experience. In Monetary Theory and Practice (pp. 91–121). Palgrave, London. https://doi.org/10.1007/978-1-349-17295-5_4Goodhart, C. A. E. 1984. “Problems of Monetary Management: The UK Experience”. In Monetary Theory and Practice. Palgrave, London. https://doi.org/10.1007/978-1-349-17295-5_4.Goodhart, C. A. E. “Problems of Monetary Management: The UK Experience”. Monetary Theory and Practice, Palgrave, London, 1984, pp. 91–121, https://doi.org/10.1007/978-1-349-17295-5_4.Goodhart, C. A. E. Problems of Monetary Management: The UK Experience. Monetary Theory and Practice 91–121 (1984) doi:10.1007/978-1-349-17295-5_4.C. A. E. Goodhart, “Problems of Monetary Management: The UK Experience”, Monetary Theory and Practice. Palgrave, London, pp. 91–121, 1984. doi: 10.1007/978-1-349-17295-5_4.
Google DeepMind(2019). AlphaStar: Mastering the real-time strategy game StarCraft II. Google DeepMind.Google DeepMind. (2019). AlphaStar: Mastering the real-time strategy game StarCraft II. Google DeepMind. https://deepmind.google/discover/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-iiGoogle DeepMind. 2019. “AlphaStar: Mastering the Real-time Strategy Game StarCraft II”. Google DeepMind. https://deepmind.google/discover/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-ii.Google DeepMind. “AlphaStar: Mastering the Real-time Strategy Game StarCraft II”. Google DeepMind, 2019, https://deepmind.google/discover/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-ii.Google DeepMind. AlphaStar: Mastering the real-time strategy game StarCraft II. Google DeepMind https://deepmind.google/discover/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-ii (2019).Google DeepMind, “AlphaStar: Mastering the real-time strategy game StarCraft II”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-ii
Google DeepMind(2020). AlphaFold: a solution to a 50-year-old grand challenge in biology.Google DeepMind. (2020, November 30). AlphaFold: a solution to a 50-year-old grand challenge in biology. https://deepmind.google/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biologyGoogle DeepMind. 2020. “AlphaFold: A Solution to a 50-year-old Grand Challenge in Biology”. November 30. https://deepmind.google/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology.Google DeepMind. AlphaFold: A Solution to a 50-year-old Grand Challenge in Biology. 30 Nov. 2020, https://deepmind.google/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology.Google DeepMind. AlphaFold: a solution to a 50-year-old grand challenge in biology. https://deepmind.google/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology (2020).Google DeepMind, “AlphaFold: a solution to a 50-year-old grand challenge in biology”. [Online]. Available: https://deepmind.google/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology
Google DeepMind(2024). Demis Hassabis & John Jumper awarded Nobel Prize in Chemistry.Google DeepMind. (2024, October 9). Demis Hassabis & John Jumper awarded Nobel Prize in Chemistry. https://deepmind.google/blog/demis-hassabis-john-jumper-awarded-nobel-prize-in-chemistryGoogle DeepMind. 2024. “Demis Hassabis & John Jumper Awarded Nobel Prize in Chemistry”. October 9. https://deepmind.google/blog/demis-hassabis-john-jumper-awarded-nobel-prize-in-chemistry.Google DeepMind. Demis Hassabis & John Jumper Awarded Nobel Prize in Chemistry. 9 Oct. 2024, https://deepmind.google/blog/demis-hassabis-john-jumper-awarded-nobel-prize-in-chemistry.Google DeepMind. Demis Hassabis & John Jumper awarded Nobel Prize in Chemistry. https://deepmind.google/blog/demis-hassabis-john-jumper-awarded-nobel-prize-in-chemistry (2024).Google DeepMind, “Demis Hassabis & John Jumper awarded Nobel Prize in Chemistry”. [Online]. Available: https://deepmind.google/blog/demis-hassabis-john-jumper-awarded-nobel-prize-in-chemistry
Gottweis, J. et al.(2025). Accelerating scientific discovery with Co-Scientist. arXiv.Gottweis, J., Weng, W.-H., Daryin, A., Tu, T., Sirkovic, P., Myaskovsky, A., Glowaty, G., Weissenberger, F., Orlandi, A., Popovici, D., Palepu, A., Rong, K., Tanno, R., Saab, K., Zhang, F., Blum, J., Carroll, A., Kulkarni, K., Tomasev, N., … Natarajan, V. (2025). Accelerating scientific discovery with Co-Scientist. In arXiv. https://doi.org/10.1038/s41586-026-10644-yGottweis, J., W.-H. Weng, A. Daryin, et al. 2025. “Accelerating Scientific Discovery with Co-Scientist”. In arXiv. Preprint, February 26. https://doi.org/10.1038/s41586-026-10644-y.Gottweis, J., et al. “Accelerating Scientific Discovery with Co-Scientist”. arXiv, 26 Feb. 2025, https://doi.org/10.1038/s41586-026-10644-y.Gottweis, J. et al. Accelerating scientific discovery with Co-Scientist. arXiv Preprint at https://doi.org/10.1038/s41586-026-10644-y (2025).J. Gottweis et al., “Accelerating scientific discovery with Co-Scientist”, Feb. 26, 2025. doi: 10.1038/s41586-026-10644-y.
Grace, K. et al.(2024). Thousands of AI Authors on the Future of AI. arXiv.Grace, K., Stewart, H., Sandkühler, J. F., Thomas, S., Weinstein-Raun, B., Brauner, J., & Korzekwa, R. C. (2024). Thousands of AI Authors on the Future of AI. In arXiv. https://doi.org/10.1613/jair.1.19087Grace, K., H. Stewart, J. F. Sandkühler, et al. 2024. “Thousands of AI Authors on the Future of AI”. In arXiv. Preprint, January 5. https://doi.org/10.1613/jair.1.19087.Grace, K., et al. “Thousands of AI Authors on the Future of AI”. arXiv, 5 Jan. 2024, https://doi.org/10.1613/jair.1.19087.Grace, K. et al. Thousands of AI Authors on the Future of AI. arXiv Preprint at https://doi.org/10.1613/jair.1.19087 (2024).K. Grace et al., “Thousands of AI Authors on the Future of AI”, Jan. 05, 2024. doi: 10.1613/jair.1.19087.
Grace, K., Salvatier, J., Dafoe, A., Zhang, B. & Evans, O.(2017). When Will AI Exceed Human Performance? Evidence from AI Experts. arXiv.Grace, K., Salvatier, J., Dafoe, A., Zhang, B., & Evans, O. (2017). When Will AI Exceed Human Performance? Evidence from AI Experts. In arXiv. https://arxiv.org/abs/1705.08807Grace, K., J. Salvatier, A. Dafoe, B. Zhang, and O. Evans. 2017. “When Will AI Exceed Human Performance? Evidence from AI Experts”. In arXiv. Preprint, May 24. https://arxiv.org/abs/1705.08807.Grace, K., et al. “When Will AI Exceed Human Performance? Evidence from AI Experts”. arXiv, 24 May 2017, https://arxiv.org/abs/1705.08807.Grace, K., Salvatier, J., Dafoe, A., Zhang, B. & Evans, O. When Will AI Exceed Human Performance? Evidence from AI Experts. arXiv Preprint at https://arxiv.org/abs/1705.08807 (2017).K. Grace, J. Salvatier, A. Dafoe, B. Zhang, and O. Evans, “When Will AI Exceed Human Performance? Evidence from AI Experts”, May 24, 2017. [Online]. Available: https://arxiv.org/abs/1705.08807
GradientDissenter(2025). METR's Evaluation of GPT-5. AI Alignment Forum.GradientDissenter. (2025, August 7). METR's Evaluation of GPT-5. AI Alignment Forum. https://alignmentforum.org/posts/SuvWoLaGiNjPDcA7d/metr-s-evaluation-of-gpt-5GradientDissenter. 2025. “METR's Evaluation of GPT-5”. AI Alignment Forum, August 7. https://alignmentforum.org/posts/SuvWoLaGiNjPDcA7d/metr-s-evaluation-of-gpt-5.GradientDissenter. “METR's Evaluation of GPT-5”. AI Alignment Forum, 7 Aug. 2025, https://alignmentforum.org/posts/SuvWoLaGiNjPDcA7d/metr-s-evaluation-of-gpt-5.GradientDissenter. METR's Evaluation of GPT-5. AI Alignment Forum https://alignmentforum.org/posts/SuvWoLaGiNjPDcA7d/metr-s-evaluation-of-gpt-5 (2025).GradientDissenter, “METR's Evaluation of GPT-5”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/SuvWoLaGiNjPDcA7d/metr-s-evaluation-of-gpt-5
Greenblatt, R. et al.(2024). Alignment faking in large language models. arXiv.Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein, J., Belonax, T., Chen, J., Duvenaud, D., Khan, A., Michael, J., Mindermann, S., Perez, E., Petrini, L., Uesato, J., Kaplan, J., Shlegeris, B., Bowman, S. R., & Hubinger, E. (2024). Alignment faking in large language models. In arXiv. https://arxiv.org/abs/2412.14093Greenblatt, R., C. Denison, B. Wright, et al. 2024. “Alignment Faking in Large Language Models”. In arXiv. Preprint, December 18. https://arxiv.org/abs/2412.14093.Greenblatt, R., et al. “Alignment Faking in Large Language Models”. arXiv, 18 Dec. 2024, https://arxiv.org/abs/2412.14093.Greenblatt, R. et al. Alignment faking in large language models. arXiv Preprint at https://arxiv.org/abs/2412.14093 (2024).R. Greenblatt et al., “Alignment faking in large language models”, Dec. 18, 2024. [Online]. Available: https://arxiv.org/abs/2412.14093
Greenblatt, R., Roger, F., Krasheninnikov, D. & Krueger, D.(2024). Stress-Testing Capability Elicitation With Password-Locked Models. arXiv.Greenblatt, R., Roger, F., Krasheninnikov, D., & Krueger, D. (2024). Stress-Testing Capability Elicitation With Password-Locked Models. In arXiv. https://arxiv.org/abs/2405.19550Greenblatt, R., F. Roger, D. Krasheninnikov, and D. Krueger. 2024. “Stress-Testing Capability Elicitation With Password-Locked Models”. In arXiv. Preprint, May 29. https://arxiv.org/abs/2405.19550.Greenblatt, R., et al. “Stress-Testing Capability Elicitation With Password-Locked Models”. arXiv, 29 May 2024, https://arxiv.org/abs/2405.19550.Greenblatt, R., Roger, F., Krasheninnikov, D. & Krueger, D. Stress-Testing Capability Elicitation With Password-Locked Models. arXiv Preprint at https://arxiv.org/abs/2405.19550 (2024).R. Greenblatt, F. Roger, D. Krasheninnikov, and D. Krueger, “Stress-Testing Capability Elicitation With Password-Locked Models”, May 29, 2024. [Online]. Available: https://arxiv.org/abs/2405.19550
Greenwalt, W. C.(2023). DOD's Replicator Program: Challenges and Opportunities.Greenwalt, W. C. (2023). DOD's Replicator Program: Challenges and Opportunities. House Armed Services Committee, Subcommittee on Cyber, Innovative Technologies, and Information Systems. https://armedservices.house.gov/sites/evo-subsites/republicans-armedservices.house.gov/files/greenwalt%20aei%20testimony%20on%20replicator%20before%20citi%20subcommittee%20hasc%20v2.pdfGreenwalt, W. C. 2023. DOD's Replicator Program: Challenges and Opportunities. House Armed Services Committee, Subcommittee on Cyber, Innovative Technologies, and Information Systems. https://armedservices.house.gov/sites/evo-subsites/republicans-armedservices.house.gov/files/greenwalt%20aei%20testimony%20on%20replicator%20before%20citi%20subcommittee%20hasc%20v2.pdf.Greenwalt, W. C. DOD's Replicator Program: Challenges and Opportunities. House Armed Services Committee, Subcommittee on Cyber, Innovative Technologies, and Information Systems, 19 Oct. 2023, https://armedservices.house.gov/sites/evo-subsites/republicans-armedservices.house.gov/files/greenwalt%20aei%20testimony%20on%20replicator%20before%20citi%20subcommittee%20hasc%20v2.pdf.Greenwalt, W. C. DOD's Replicator Program: Challenges and Opportunities. https://armedservices.house.gov/sites/evo-subsites/republicans-armedservices.house.gov/files/greenwalt%20aei%20testimony%20on%20replicator%20before%20citi%20subcommittee%20hasc%20v2.pdf (2023).W. C. Greenwalt, “DOD's Replicator Program: Challenges and Opportunities”, House Armed Services Committee, Subcommittee on Cyber, Innovative Technologies, and Information Systems, Oct. 2023. [Online]. Available: https://armedservices.house.gov/sites/evo-subsites/republicans-armedservices.house.gov/files/greenwalt%20aei%20testimony%20on%20replicator%20before%20citi%20subcommittee%20hasc%20v2.pdf
Griffin, C., Thomson, L., Shlegeris, B. & Abate, A.(2024). Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols. arXiv.Griffin, C., Thomson, L., Shlegeris, B., & Abate, A. (2024). Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols. In arXiv. https://arxiv.org/abs/2409.07985Griffin, C., L. Thomson, B. Shlegeris, and A. Abate. 2024. “Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols”. In arXiv. Preprint, September 12. https://arxiv.org/abs/2409.07985.Griffin, C., et al. “Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols”. arXiv, 12 Sept. 2024, https://arxiv.org/abs/2409.07985.Griffin, C., Thomson, L., Shlegeris, B. & Abate, A. Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols. arXiv Preprint at https://arxiv.org/abs/2409.07985 (2024).C. Griffin, L. Thomson, B. Shlegeris, and A. Abate, “Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols”, Sep. 12, 2024. [Online]. Available: https://arxiv.org/abs/2409.07985
Gruetzemacher et al.(2021). Forecasting AI progress: A research agenda.Gruetzemacher et al. (2021). Forecasting AI progress: A research agenda. Internet Archive (https://web.archive.org/web/20240404122334/https://www.sciencedirect.com/science/article/pii/S0040162521003413). https://sciencedirect.com/science/article/pii/S0040162521003413Gruetzemacher et al. 2021. “Forecasting AI Progress: A Research Agenda”. Https://web.archive.org/web/20240404122334/https://www.sciencedirect.com/science/article/pii/S0040162521003413. Internet Archive. https://sciencedirect.com/science/article/pii/S0040162521003413.Gruetzemacher et al. Forecasting AI Progress: A Research Agenda. 2021, Internet Archive, https://web.archive.org/web/20240404122334/https://www.sciencedirect.com/science/article/pii/S0040162521003413, https://sciencedirect.com/science/article/pii/S0040162521003413.Gruetzemacher et al. Forecasting AI progress: A research agenda. https://sciencedirect.com/science/article/pii/S0040162521003413 (2021).Gruetzemacher et al., “Forecasting AI progress: A research agenda”. Accessed: Apr. 04, 2024. [Online]. Available: https://sciencedirect.com/science/article/pii/S0040162521003413
Gruetzemacher, R., Avin, S., Fox, J. & Saeri, A. K.(2024). Strategic Insights from Simulation Gaming of AI Race Dynamics. arXiv.Gruetzemacher, R., Avin, S., Fox, J., & Saeri, A. K. (2024). Strategic Insights from Simulation Gaming of AI Race Dynamics. In arXiv. https://arxiv.org/abs/2410.03092Gruetzemacher, R., S. Avin, J. Fox, and A. K. Saeri. 2024. “Strategic Insights from Simulation Gaming of AI Race Dynamics”. In arXiv. Preprint, October 4. https://arxiv.org/abs/2410.03092.Gruetzemacher, R., et al. “Strategic Insights from Simulation Gaming of AI Race Dynamics”. arXiv, 4 Oct. 2024, https://arxiv.org/abs/2410.03092.Gruetzemacher, R., Avin, S., Fox, J. & Saeri, A. K. Strategic Insights from Simulation Gaming of AI Race Dynamics. arXiv Preprint at https://arxiv.org/abs/2410.03092 (2024).R. Gruetzemacher, S. Avin, J. Fox, and A. K. Saeri, “Strategic Insights from Simulation Gaming of AI Race Dynamics”, Oct. 04, 2024. [Online]. Available: https://arxiv.org/abs/2410.03092
Gurnee, W. & Tegmark, M.(2023). Language Models Represent Space and Time. arXiv.Gurnee, W., & Tegmark, M. (2023). Language Models Represent Space and Time. In arXiv. https://arxiv.org/abs/2310.02207Gurnee, W., and M. Tegmark. 2023. “Language Models Represent Space and Time”. In arXiv. Preprint, October 3. https://arxiv.org/abs/2310.02207.Gurnee, W., and M. Tegmark. “Language Models Represent Space and Time”. arXiv, 3 Oct. 2023, https://arxiv.org/abs/2310.02207.Gurnee, W. & Tegmark, M. Language Models Represent Space and Time. arXiv Preprint at https://arxiv.org/abs/2310.02207 (2023).W. Gurnee and M. Tegmark, “Language Models Represent Space and Time”, Oct. 03, 2023. [Online]. Available: https://arxiv.org/abs/2310.02207
Gwern(2016). Why Tool AIs Want to Be Agent AIs.Gwern. (2016). Why Tool AIs Want to Be Agent AIs. https://gwern.net/tool-aiGwern. 2016. “Why Tool AIs Want to Be Agent AIs”. https://gwern.net/tool-ai.Gwern. Why Tool AIs Want to Be Agent AIs. 2016, https://gwern.net/tool-ai.Gwern. Why Tool AIs Want to Be Agent AIs. https://gwern.net/tool-ai (2016).Gwern, “Why Tool AIs Want to Be Agent AIs”. [Online]. Available: https://gwern.net/tool-ai
Habli et al.(2025). The BIG Argument for AI Safety Cases. arXiv.org.Habli et al. (2025). The BIG Argument for AI Safety Cases. arXiv.org. https://www.arxiv.org/abs/2503.11705Habli et al. 2025. “The BIG Argument for AI Safety Cases”. arXiv.org. https://www.arxiv.org/abs/2503.11705.Habli et al. “The BIG Argument for AI Safety Cases”. arXiv.org, 2025, https://www.arxiv.org/abs/2503.11705.Habli et al. The BIG Argument for AI Safety Cases. arXiv.org https://www.arxiv.org/abs/2503.11705 (2025).Habli et al., “The BIG Argument for AI Safety Cases”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2503.11705
Hadley, E., Blatecky, A. & Comfort, M.(2024). Investigating Algorithm Review Boards for Organizational Responsible Artificial Intelligence Governance. arXiv.Hadley, E., Blatecky, A., & Comfort, M. (2024). Investigating Algorithm Review Boards for Organizational Responsible Artificial Intelligence Governance. In arXiv. https://arxiv.org/abs/2402.01691Hadley, E., A. Blatecky, and M. Comfort. 2024. “Investigating Algorithm Review Boards for Organizational Responsible Artificial Intelligence Governance”. In arXiv. Preprint, January 23. https://arxiv.org/abs/2402.01691.Hadley, E., et al. “Investigating Algorithm Review Boards for Organizational Responsible Artificial Intelligence Governance”. arXiv, 23 Jan. 2024, https://arxiv.org/abs/2402.01691.Hadley, E., Blatecky, A. & Comfort, M. Investigating Algorithm Review Boards for Organizational Responsible Artificial Intelligence Governance. arXiv Preprint at https://arxiv.org/abs/2402.01691 (2024).E. Hadley, A. Blatecky, and M. Comfort, “Investigating Algorithm Review Boards for Organizational Responsible Artificial Intelligence Governance”, Jan. 23, 2024. [Online]. Available: https://arxiv.org/abs/2402.01691
Haldane, A. G. & May, R. M.(2011). Systemic risk in banking ecosystems. Nature.Haldane, A. G., & May, R. M. (2011). Systemic risk in banking ecosystems. Nature, 469, 351–355. https://doi.org/10.1038/nature09659Haldane, A. G., and R. M. May. 2011. “Systemic Risk in Banking Ecosystems”. Nature 469 (January): 351–55. https://doi.org/10.1038/nature09659.Haldane, A. G., and R. M. May. “Systemic Risk in Banking Ecosystems”. Nature, vol. 469, Jan. 2011, pp. 351–55, https://doi.org/10.1038/nature09659.Haldane, A. G. & May, R. M. Systemic risk in banking ecosystems. Nature 469, 351–355 (2011).A. G. Haldane and R. M. May, “Systemic risk in banking ecosystems”, Nature, vol. 469, pp. 351–355, Jan. 2011, doi: 10.1038/nature09659.
Hanson(2023). AI Risk, Again.Hanson. (2023). AI Risk, Again. https://www.overcomingbias.com/p/ai-risk-againHanson. 2023. “AI Risk, Again”. https://www.overcomingbias.com/p/ai-risk-again.Hanson. AI Risk, Again. 2023, https://www.overcomingbias.com/p/ai-risk-again.Hanson. AI Risk, Again. https://www.overcomingbias.com/p/ai-risk-again (2023).Hanson, “AI Risk, Again”. [Online]. Available: https://www.overcomingbias.com/p/ai-risk-again
Hao, S. et al.(2024). Training Large Language Models to Reason in a Continuous Latent Space. arXiv.Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., & Tian, Y. (2024). Training Large Language Models to Reason in a Continuous Latent Space. In arXiv. https://arxiv.org/abs/2412.06769Hao, S., S. Sukhbaatar, D. Su, et al. 2024. “Training Large Language Models to Reason in a Continuous Latent Space”. In arXiv. Preprint, December 9. https://arxiv.org/abs/2412.06769.Hao, S., et al. “Training Large Language Models to Reason in a Continuous Latent Space”. arXiv, 9 Dec. 2024, https://arxiv.org/abs/2412.06769.Hao, S. et al. Training Large Language Models to Reason in a Continuous Latent Space. arXiv Preprint at https://arxiv.org/abs/2412.06769 (2024).S. Hao et al., “Training Large Language Models to Reason in a Continuous Latent Space”, Dec. 09, 2024. [Online]. Available: https://arxiv.org/abs/2412.06769
Harvard(2025). What is AI ethics?. Harvard FAS | Mignone Center for Career Success.Harvard. (2025). What is AI ethics?. Internet Archive (https://web.archive.org/web/20260123144850/https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics/). Harvard FAS | Mignone Center for Career Success. https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethicsHarvard. 2025. “What Is AI Ethics?”. Harvard FAS | Mignone Center for Career Success. Https://web.archive.org/web/20260123144850/https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics/. Internet Archive. https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics.Harvard. “What Is AI Ethics?”. Harvard FAS | Mignone Center for Career Success, 2025, Internet Archive, https://web.archive.org/web/20260123144850/https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics/, https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics.Harvard. What is AI ethics?. Harvard FAS | Mignone Center for Career Success https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics (2025).Harvard, “What is AI ethics?”, Harvard FAS | Mignone Center for Career Success. Accessed: Jan. 23, 2026. [Online]. Available: https://careerservices.fas.harvard.edu/blog/2025/05/01/what-is-ai-ethics
Hassabis(2025). AI bosses are feeling the high-stakes pressure. Business Insider.Hassabis. (2025). AI bosses are feeling the high-stakes pressure. Business Insider. https://businessinsider.com/google-deepmind-ceo-demis-hassabis-anthropic-ceo-ai-pressure-worries-2025-2Hassabis. 2025. “AI Bosses Are Feeling the High-stakes Pressure”. Business Insider. https://businessinsider.com/google-deepmind-ceo-demis-hassabis-anthropic-ceo-ai-pressure-worries-2025-2.Hassabis. “AI Bosses Are Feeling the High-stakes Pressure”. Business Insider, 2025, https://businessinsider.com/google-deepmind-ceo-demis-hassabis-anthropic-ceo-ai-pressure-worries-2025-2.Hassabis. AI bosses are feeling the high-stakes pressure. Business Insider https://businessinsider.com/google-deepmind-ceo-demis-hassabis-anthropic-ceo-ai-pressure-worries-2025-2 (2025).Hassabis, “AI bosses are feeling the high-stakes pressure”, Business Insider. [Online]. Available: https://businessinsider.com/google-deepmind-ceo-demis-hassabis-anthropic-ceo-ai-pressure-worries-2025-2
Haugen(2021). Facebook’s Documents About Instagram and Teens. The Wall Street Journal.Haugen. (2021). Facebook’s Documents About Instagram and Teens. Internet Archive (https://web.archive.org/web/20251123195306/https://www.wsj.com/livecoverage/facebook-whistleblower-frances-haugen-senate-hearing/card/eFNjPrwIH4F7BALELWrZ). The Wall Street Journal. https://wsj.com/livecoverage/facebook-whistleblower-frances-haugen-senate-hearing/card/eFNjPrwIH4F7BALELWrZHaugen. 2021. “Facebook’s Documents About Instagram and Teens”. The Wall Street Journal. Https://web.archive.org/web/20251123195306/https://www.wsj.com/livecoverage/facebook-whistleblower-frances-haugen-senate-hearing/card/eFNjPrwIH4F7BALELWrZ. Internet Archive. https://wsj.com/livecoverage/facebook-whistleblower-frances-haugen-senate-hearing/card/eFNjPrwIH4F7BALELWrZ.Haugen. “Facebook’s Documents About Instagram and Teens”. The Wall Street Journal, 2021, Internet Archive, https://web.archive.org/web/20251123195306/https://www.wsj.com/livecoverage/facebook-whistleblower-frances-haugen-senate-hearing/card/eFNjPrwIH4F7BALELWrZ, https://wsj.com/livecoverage/facebook-whistleblower-frances-haugen-senate-hearing/card/eFNjPrwIH4F7BALELWrZ.Haugen. Facebook’s Documents About Instagram and Teens. The Wall Street Journal https://wsj.com/livecoverage/facebook-whistleblower-frances-haugen-senate-hearing/card/eFNjPrwIH4F7BALELWrZ (2021).Haugen, “Facebook’s Documents About Instagram and Teens”, The Wall Street Journal. Accessed: Nov. 23, 2025. [Online]. Available: https://wsj.com/livecoverage/facebook-whistleblower-frances-haugen-senate-hearing/card/eFNjPrwIH4F7BALELWrZ
Hausenloy, J., McClements, D. & Thakur, M.(2024). Towards Data Governance of Frontier AI Models. arXiv.Hausenloy, J., McClements, D., & Thakur, M. (2024). Towards Data Governance of Frontier AI Models. In arXiv. https://arxiv.org/abs/2412.03824Hausenloy, J., D. McClements, and M. Thakur. 2024. “Towards Data Governance of Frontier AI Models”. In arXiv. Preprint, December 5. https://arxiv.org/abs/2412.03824.Hausenloy, J., et al. “Towards Data Governance of Frontier AI Models”. arXiv, 5 Dec. 2024, https://arxiv.org/abs/2412.03824.Hausenloy, J., McClements, D. & Thakur, M. Towards Data Governance of Frontier AI Models. arXiv Preprint at https://arxiv.org/abs/2412.03824 (2024).J. Hausenloy, D. McClements, and M. Thakur, “Towards Data Governance of Frontier AI Models”, Dec. 05, 2024. [Online]. Available: https://arxiv.org/abs/2412.03824
Hausenloy, J., Miotti, A. & Dennis, C.(2023). Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI. arXiv.Hausenloy, J., Miotti, A., & Dennis, C. (2023). Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI. In arXiv. https://arxiv.org/abs/2310.09217Hausenloy, J., A. Miotti, and C. Dennis. 2023. “Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI”. In arXiv. Preprint, October 13. https://arxiv.org/abs/2310.09217.Hausenloy, J., et al. “Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI”. arXiv, 13 Oct. 2023, https://arxiv.org/abs/2310.09217.Hausenloy, J., Miotti, A. & Dennis, C. Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI. arXiv Preprint at https://arxiv.org/abs/2310.09217 (2023).J. Hausenloy, A. Miotti, and C. Dennis, “Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI”, Oct. 13, 2023. [Online]. Available: https://arxiv.org/abs/2310.09217
Heaven(2023). Geoffrey Hinton tells us why he’s now scared of the tech he helped build. MIT Technology Review.Heaven. (2023). Geoffrey Hinton tells us why he’s now scared of the tech he helped build. MIT Technology Review. https://technologyreview.com/2023/05/02/1072528/geoffrey-hinton-google-why-scared-aiHeaven. 2023. “Geoffrey Hinton Tells Us Why He’s Now Scared of the Tech He Helped Build”. MIT Technology Review. https://technologyreview.com/2023/05/02/1072528/geoffrey-hinton-google-why-scared-ai.Heaven. “Geoffrey Hinton Tells Us Why He’s Now Scared of the Tech He Helped Build”. MIT Technology Review, 2023, https://technologyreview.com/2023/05/02/1072528/geoffrey-hinton-google-why-scared-ai.Heaven. Geoffrey Hinton tells us why he’s now scared of the tech he helped build. MIT Technology Review https://technologyreview.com/2023/05/02/1072528/geoffrey-hinton-google-why-scared-ai (2023).Heaven, “Geoffrey Hinton tells us why he’s now scared of the tech he helped build”, MIT Technology Review. [Online]. Available: https://technologyreview.com/2023/05/02/1072528/geoffrey-hinton-google-why-scared-ai
Heelan(2025). How I used o3 to find CVE-2025-37899, a remote zeroday vulnerability in the Linux kernel’s SMB implementation. Sean Heelan's Blog.Heelan. (2025, May 22). How I used o3 to find CVE-2025-37899, a remote zeroday vulnerability in the Linux kernel’s SMB implementation. Sean Heelan's Blog. https://sean.heelan.io/2025/05/22/how-i-used-o3-to-find-cve-2025-37899-a-remote-zeroday-vulnerability-in-the-linux-kernels-smb-implementationHeelan. 2025. “How I Used O3 to Find CVE-2025-37899, a Remote Zeroday Vulnerability in the Linux Kernel’s SMB Implementation”. Sean Heelan's Blog, May 22. https://sean.heelan.io/2025/05/22/how-i-used-o3-to-find-cve-2025-37899-a-remote-zeroday-vulnerability-in-the-linux-kernels-smb-implementation.Heelan. “How I Used O3 to Find CVE-2025-37899, a Remote Zeroday Vulnerability in the Linux Kernel’s SMB Implementation”. Sean Heelan's Blog, 22 May 2025, https://sean.heelan.io/2025/05/22/how-i-used-o3-to-find-cve-2025-37899-a-remote-zeroday-vulnerability-in-the-linux-kernels-smb-implementation.Heelan. How I used o3 to find CVE-2025-37899, a remote zeroday vulnerability in the Linux kernel’s SMB implementation. Sean Heelan's Blog https://sean.heelan.io/2025/05/22/how-i-used-o3-to-find-cve-2025-37899-a-remote-zeroday-vulnerability-in-the-linux-kernels-smb-implementation (2025).Heelan, “How I used o3 to find CVE-2025-37899, a remote zeroday vulnerability in the Linux kernel’s SMB implementation”, Sean Heelan's Blog. [Online]. Available: https://sean.heelan.io/2025/05/22/how-i-used-o3-to-find-cve-2025-37899-a-remote-zeroday-vulnerability-in-the-linux-kernels-smb-implementation
Heim, L. & Koessler, L.(2024). Training Compute Thresholds: Features and Functions in AI Regulation. arXiv.Heim, L., & Koessler, L. (2024). Training Compute Thresholds: Features and Functions in AI Regulation. In arXiv. https://arxiv.org/abs/2405.10799Heim, L., and L. Koessler. 2024. “Training Compute Thresholds: Features and Functions in AI Regulation”. In arXiv. Preprint, May 17. https://arxiv.org/abs/2405.10799.Heim, L., and L. Koessler. “Training Compute Thresholds: Features and Functions in AI Regulation”. arXiv, 17 May 2024, https://arxiv.org/abs/2405.10799.Heim, L. & Koessler, L. Training Compute Thresholds: Features and Functions in AI Regulation. arXiv Preprint at https://arxiv.org/abs/2405.10799 (2024).L. Heim and L. Koessler, “Training Compute Thresholds: Features and Functions in AI Regulation”, May 17, 2024. [Online]. Available: https://arxiv.org/abs/2405.10799
Heim, L. et al.(2024). Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation. arXiv.Heim, L., Fist, T., Egan, J., Huang, S., Zekany, S., Trager, R., Osborne, M. A., & Zilberman, N. (2024). Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation. In arXiv. https://arxiv.org/abs/2403.08501Heim, L., T. Fist, J. Egan, et al. 2024. “Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation”. In arXiv. Preprint, March 13. https://arxiv.org/abs/2403.08501.Heim, L., et al. “Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation”. arXiv, 13 Mar. 2024, https://arxiv.org/abs/2403.08501.Heim, L. et al. Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation. arXiv Preprint at https://arxiv.org/abs/2403.08501 (2024).L. Heim et al., “Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation”, Mar. 13, 2024. [Online]. Available: https://arxiv.org/abs/2403.08501
Hendrycks & Wang(2024). Submit Your Toughest Questions for Humanity's Last Exam | CAIS. Center for AI Safety.Hendrycks & Wang. (2024). Submit Your Toughest Questions for Humanity's Last Exam | CAIS. Center for AI Safety. https://safe.ai/blog/humanitys-last-examHendrycks & Wang. 2024. “Submit Your Toughest Questions for Humanity's Last Exam | CAIS”. Center for AI Safety. https://safe.ai/blog/humanitys-last-exam.Hendrycks & Wang. “Submit Your Toughest Questions for Humanity's Last Exam | CAIS”. Center for AI Safety, 2024, https://safe.ai/blog/humanitys-last-exam.Hendrycks & Wang. Submit Your Toughest Questions for Humanity's Last Exam | CAIS. Center for AI Safety https://safe.ai/blog/humanitys-last-exam (2024).Hendrycks & Wang, “Submit Your Toughest Questions for Humanity's Last Exam | CAIS”, Center for AI Safety. [Online]. Available: https://safe.ai/blog/humanitys-last-exam
Hendrycks(2024). 1.2: Malicious Use | AI Safety, Ethics, and Society Textbook.Hendrycks. (2024). 1.2: Malicious Use | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/malicious-useHendrycks. 2024. “1.2: Malicious Use | AI Safety, Ethics, and Society Textbook”. https://aisafetybook.com/textbook/malicious-use.Hendrycks. 1.2: Malicious Use | AI Safety, Ethics, and Society Textbook. 2024, https://aisafetybook.com/textbook/malicious-use.Hendrycks. 1.2: Malicious Use | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/malicious-use (2024).Hendrycks, “1.2: Malicious Use | AI Safety, Ethics, and Society Textbook”. [Online]. Available: https://aisafetybook.com/textbook/malicious-use
Hendrycks(2024). 3.3: Robustness | AI Safety, Ethics, and Society Textbook.Hendrycks. (2024). 3.3: Robustness | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/robustnessHendrycks. 2024. “3.3: Robustness | AI Safety, Ethics, and Society Textbook”. https://aisafetybook.com/textbook/robustness.Hendrycks. 3.3: Robustness | AI Safety, Ethics, and Society Textbook. 2024, https://aisafetybook.com/textbook/robustness.Hendrycks. 3.3: Robustness | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/robustness (2024).Hendrycks, “3.3: Robustness | AI Safety, Ethics, and Society Textbook”. [Online]. Available: https://aisafetybook.com/textbook/robustness
Hendrycks(2024). 3.4: Alignment | AI Safety, Ethics, and Society Textbook.Hendrycks. (2024). 3.4: Alignment | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/alignmentHendrycks. 2024. “3.4: Alignment | AI Safety, Ethics, and Society Textbook”. https://aisafetybook.com/textbook/alignment.Hendrycks. 3.4: Alignment | AI Safety, Ethics, and Society Textbook. 2024, https://aisafetybook.com/textbook/alignment.Hendrycks. 3.4: Alignment | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/alignment (2024).Hendrycks, “3.4: Alignment | AI Safety, Ethics, and Society Textbook”. [Online]. Available: https://aisafetybook.com/textbook/alignment
Hendrycks(2024). 4.5: Component Failure Accident Models and Methods | AI Safety, Ethics, and Society Textbook.Hendrycks. (2024). 4.5: Component Failure Accident Models and Methods | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/component-failure-accident-modelsHendrycks. 2024. “4.5: Component Failure Accident Models and Methods | AI Safety, Ethics, and Society Textbook”. https://aisafetybook.com/textbook/component-failure-accident-models.Hendrycks. 4.5: Component Failure Accident Models and Methods | AI Safety, Ethics, and Society Textbook. 2024, https://aisafetybook.com/textbook/component-failure-accident-models.Hendrycks. 4.5: Component Failure Accident Models and Methods | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/component-failure-accident-models (2024).Hendrycks, “4.5: Component Failure Accident Models and Methods | AI Safety, Ethics, and Society Textbook”. [Online]. Available: https://aisafetybook.com/textbook/component-failure-accident-models
Hendrycks(2025). 5.2: Introduction to Complex Systems | AI Safety, Ethics, and Society Textbook.Hendrycks. (2025). 5.2: Introduction to Complex Systems | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/introduction-to-complex-systemsHendrycks. 2025. “5.2: Introduction to Complex Systems | AI Safety, Ethics, and Society Textbook”. https://aisafetybook.com/textbook/introduction-to-complex-systems.Hendrycks. 5.2: Introduction to Complex Systems | AI Safety, Ethics, and Society Textbook. 2025, https://aisafetybook.com/textbook/introduction-to-complex-systems.Hendrycks. 5.2: Introduction to Complex Systems | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/introduction-to-complex-systems (2025).Hendrycks, “5.2: Introduction to Complex Systems | AI Safety, Ethics, and Society Textbook”. [Online]. Available: https://aisafetybook.com/textbook/introduction-to-complex-systems
Hendrycks et al.(2024). 8.4: Corporate Governance | AI Safety, Ethics, and Society Textbook.Hendrycks et al. (2024). 8.4: Corporate Governance | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/corporate-governanceHendrycks et al. 2024. “8.4: Corporate Governance | AI Safety, Ethics, and Society Textbook”. https://aisafetybook.com/textbook/corporate-governance.Hendrycks et al. 8.4: Corporate Governance | AI Safety, Ethics, and Society Textbook. 2024, https://aisafetybook.com/textbook/corporate-governance.Hendrycks et al. 8.4: Corporate Governance | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/corporate-governance (2024).Hendrycks et al., “8.4: Corporate Governance | AI Safety, Ethics, and Society Textbook”. [Online]. Available: https://aisafetybook.com/textbook/corporate-governance
Hendrycks et al.(2025). AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy.Hendrycks et al. (2025). AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy. https://nationalsecurity.ai/chapter/ai-is-pivotal-for-national-securityHendrycks et al. 2025. “AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy”. https://nationalsecurity.ai/chapter/ai-is-pivotal-for-national-security.Hendrycks et al. AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy. 2025, https://nationalsecurity.ai/chapter/ai-is-pivotal-for-national-security.Hendrycks et al. AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy. https://nationalsecurity.ai/chapter/ai-is-pivotal-for-national-security (2025).Hendrycks et al., “AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy”. [Online]. Available: https://nationalsecurity.ai/chapter/ai-is-pivotal-for-national-security
Hendrycks et al.(2025). Superintelligence Strategy.Hendrycks et al. (2025). Superintelligence Strategy. https://nationalsecurity.aiHendrycks et al. 2025. “Superintelligence Strategy”. https://nationalsecurity.ai.Hendrycks et al. Superintelligence Strategy. 2025, https://nationalsecurity.ai.Hendrycks et al. Superintelligence Strategy. https://nationalsecurity.ai (2025).Hendrycks et al., “Superintelligence Strategy”. [Online]. Available: https://nationalsecurity.ai
Hendrycks, D. et al.(2020). Aligning AI With Shared Human Values. arXiv.Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., & Steinhardt, J. (2020). Aligning AI With Shared Human Values. In arXiv. https://arxiv.org/abs/2008.02275Hendrycks, D., C. Burns, S. Basart, et al. 2020. “Aligning AI With Shared Human Values”. In arXiv. Preprint, August 5. https://arxiv.org/abs/2008.02275.Hendrycks, D., et al. “Aligning AI With Shared Human Values”. arXiv, 5 Aug. 2020, https://arxiv.org/abs/2008.02275.Hendrycks, D. et al. Aligning AI With Shared Human Values. arXiv Preprint at https://arxiv.org/abs/2008.02275 (2020).D. Hendrycks et al., “Aligning AI With Shared Human Values”, Aug. 05, 2020. [Online]. Available: https://arxiv.org/abs/2008.02275
Hendrycks, D. et al.(2020). Measuring Massive Multitask Language Understanding. arXiv.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2020). Measuring Massive Multitask Language Understanding. In arXiv. https://arxiv.org/abs/2009.03300Hendrycks, D., C. Burns, S. Basart, et al. 2020. “Measuring Massive Multitask Language Understanding”. In arXiv. Preprint, September 7. https://arxiv.org/abs/2009.03300.Hendrycks, D., et al. “Measuring Massive Multitask Language Understanding”. arXiv, 7 Sept. 2020, https://arxiv.org/abs/2009.03300.Hendrycks, D. et al. Measuring Massive Multitask Language Understanding. arXiv Preprint at https://arxiv.org/abs/2009.03300 (2020).D. Hendrycks et al., “Measuring Massive Multitask Language Understanding”, Sep. 07, 2020. [Online]. Available: https://arxiv.org/abs/2009.03300
Hendrycks, D. et al.(2021). Measuring Coding Challenge Competence With APPS. arXiv.Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., & Steinhardt, J. (2021). Measuring Coding Challenge Competence With APPS. In arXiv. https://arxiv.org/abs/2105.09938Hendrycks, D., S. Basart, S. Kadavath, et al. 2021. “Measuring Coding Challenge Competence With APPS”. In arXiv. Preprint, May 20. https://arxiv.org/abs/2105.09938.Hendrycks, D., et al. “Measuring Coding Challenge Competence With APPS”. arXiv, 20 May 2021, https://arxiv.org/abs/2105.09938.Hendrycks, D. et al. Measuring Coding Challenge Competence With APPS. arXiv Preprint at https://arxiv.org/abs/2105.09938 (2021).D. Hendrycks et al., “Measuring Coding Challenge Competence With APPS”, May 20, 2021. [Online]. Available: https://arxiv.org/abs/2105.09938
Hendrycks, D. et al.(2021). Measuring Mathematical Problem Solving With the MATH Dataset. arXiv.Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., & Steinhardt, J. (2021). Measuring Mathematical Problem Solving With the MATH Dataset. In arXiv. https://arxiv.org/abs/2103.03874Hendrycks, D., C. Burns, S. Kadavath, et al. 2021. “Measuring Mathematical Problem Solving With the MATH Dataset”. In arXiv. Preprint, March 5. https://arxiv.org/abs/2103.03874.Hendrycks, D., et al. “Measuring Mathematical Problem Solving With the MATH Dataset”. arXiv, 5 Mar. 2021, https://arxiv.org/abs/2103.03874.Hendrycks, D. et al. Measuring Mathematical Problem Solving With the MATH Dataset. arXiv Preprint at https://arxiv.org/abs/2103.03874 (2021).D. Hendrycks et al., “Measuring Mathematical Problem Solving With the MATH Dataset”, Mar. 05, 2021. [Online]. Available: https://arxiv.org/abs/2103.03874
Hendrycks, D. et al.(2025). A Definition of AGI. arXiv.Hendrycks, D., Song, D., Szegedy, C., Lee, H., Gal, Y., Brynjolfsson, E., Li, S., Zou, A., Levine, L., Han, B., Fu, J., Liu, Z., Shin, J., Lee, K., Mazeika, M., Phan, L., Ingebretsen, G., Khoja, A., Xie, C., … Bengio, Y. (2025). A Definition of AGI. In arXiv. https://arxiv.org/abs/2510.18212Hendrycks, D., D. Song, C. Szegedy, et al. 2025. “A Definition of AGI”. In arXiv. Preprint, October 21. https://arxiv.org/abs/2510.18212.Hendrycks, D., et al. “A Definition of AGI”. arXiv, 21 Oct. 2025, https://arxiv.org/abs/2510.18212.Hendrycks, D. et al. A Definition of AGI. arXiv Preprint at https://arxiv.org/abs/2510.18212 (2025).D. Hendrycks et al., “A Definition of AGI”, Oct. 21, 2025. [Online]. Available: https://arxiv.org/abs/2510.18212
Hendrycks, D., Carlini, N., Schulman, J. & Steinhardt, J.(2021). Unsolved Problems in ML Safety. arXiv.Hendrycks, D., Carlini, N., Schulman, J., & Steinhardt, J. (2021). Unsolved Problems in ML Safety. In arXiv. https://arxiv.org/abs/2109.13916Hendrycks, D., N. Carlini, J. Schulman, and J. Steinhardt. 2021. “Unsolved Problems in ML Safety”. In arXiv. Preprint, September 28. https://arxiv.org/abs/2109.13916.Hendrycks, D., et al. “Unsolved Problems in ML Safety”. arXiv, 28 Sept. 2021, https://arxiv.org/abs/2109.13916.Hendrycks, D., Carlini, N., Schulman, J. & Steinhardt, J. Unsolved Problems in ML Safety. arXiv Preprint at https://arxiv.org/abs/2109.13916 (2021).D. Hendrycks, N. Carlini, J. Schulman, and J. Steinhardt, “Unsolved Problems in ML Safety”, Sep. 28, 2021. [Online]. Available: https://arxiv.org/abs/2109.13916
Hendrycks, D., Mazeika, M. & Woodside, T.(2023). An Overview of Catastrophic AI Risks. arXiv.Hendrycks, D., Mazeika, M., & Woodside, T. (2023). An Overview of Catastrophic AI Risks. In arXiv. https://arxiv.org/abs/2306.12001Hendrycks, D., M. Mazeika, and T. Woodside. 2023. “An Overview of Catastrophic AI Risks”. In arXiv. Preprint, June 21. https://arxiv.org/abs/2306.12001.Hendrycks, D., et al. “An Overview of Catastrophic AI Risks”. arXiv, 21 June 2023, https://arxiv.org/abs/2306.12001.Hendrycks, D., Mazeika, M. & Woodside, T. An Overview of Catastrophic AI Risks. arXiv Preprint at https://arxiv.org/abs/2306.12001 (2023).D. Hendrycks, M. Mazeika, and T. Woodside, “An Overview of Catastrophic AI Risks”, Jun. 21, 2023. [Online]. Available: https://arxiv.org/abs/2306.12001
Hendryks(2024). 7.2: Game Theory | AI Safety, Ethics, and Society Textbook.Hendryks. (2024). 7.2: Game Theory | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/game-theoryHendryks. 2024. “7.2: Game Theory | AI Safety, Ethics, and Society Textbook”. https://aisafetybook.com/textbook/game-theory.Hendryks. 7.2: Game Theory | AI Safety, Ethics, and Society Textbook. 2024, https://aisafetybook.com/textbook/game-theory.Hendryks. 7.2: Game Theory | AI Safety, Ethics, and Society Textbook. https://aisafetybook.com/textbook/game-theory (2024).Hendryks, “7.2: Game Theory | AI Safety, Ethics, and Society Textbook”. [Online]. Available: https://aisafetybook.com/textbook/game-theory
Herculano-Houzel, S.(2012). The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences.Herculano-Houzel, S. (2012). The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences. https://doi.org/10.1073/pnas.1201895109Herculano-Houzel, S. 2012. “The Remarkable, yet Not Extraordinary, Human Brain as a Scaled-up Primate Brain and Its Associated Cost”. Proceedings of the National Academy of Sciences, ahead of print, June 22. https://doi.org/10.1073/pnas.1201895109.Herculano-Houzel, S. “The Remarkable, yet Not Extraordinary, Human Brain as a Scaled-up Primate Brain and Its Associated Cost”. Proceedings of the National Academy of Sciences, June 2012, https://doi.org/10.1073/pnas.1201895109.Herculano-Houzel, S. The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences https://doi.org/10.1073/pnas.1201895109 (2012) doi:10.1073/pnas.1201895109.S. Herculano-Houzel, “The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost”, Proceedings of the National Academy of Sciences, Jun. 2012, doi: 10.1073/pnas.1201895109.
Herrera-Poyatos, A., Ser, J. D., Prado, M. L. D., Wang, F., Herrera-Viedma, E. & Herrera, F.(2025). A Framework for Responsible AI Systems: Building Societal Trust through Domain Definition, Trustworthy AI Design, Auditability, Accountability, and Governance. arXiv.Herrera-Poyatos, A., Ser, J. D., de Prado, M. L., Wang, F.-Y., Herrera-Viedma, E., & Herrera, F. (2025). A Framework for Responsible AI Systems: Building Societal Trust through Domain Definition, Trustworthy AI Design, Auditability, Accountability, and Governance. In arXiv. https://arxiv.org/abs/2503.04739Herrera-Poyatos, A., J. D. Ser, M. L. de Prado, F.-Y. Wang, E. Herrera-Viedma, and F. Herrera. 2025. “A Framework for Responsible AI Systems: Building Societal Trust Through Domain Definition, Trustworthy AI Design, Auditability, Accountability, and Governance”. In arXiv. Preprint, February 4. https://arxiv.org/abs/2503.04739.Herrera-Poyatos, A., et al. “A Framework for Responsible AI Systems: Building Societal Trust Through Domain Definition, Trustworthy AI Design, Auditability, Accountability, and Governance”. arXiv, 4 Feb. 2025, https://arxiv.org/abs/2503.04739.Herrera-Poyatos, A. et al. A Framework for Responsible AI Systems: Building Societal Trust through Domain Definition, Trustworthy AI Design, Auditability, Accountability, and Governance. arXiv Preprint at https://arxiv.org/abs/2503.04739 (2025).A. Herrera-Poyatos, J. D. Ser, M. L. de Prado, F.-Y. Wang, E. Herrera-Viedma, and F. Herrera, “A Framework for Responsible AI Systems: Building Societal Trust through Domain Definition, Trustworthy AI Design, Auditability, Accountability, and Governance”, Feb. 04, 2025. [Online]. Available: https://arxiv.org/abs/2503.04739
Hill(2024). Understanding Offensive AI vs. Defensive AI in Cybersecurity | Abnormal AI.Hill. (2024). Understanding Offensive AI vs. Defensive AI in Cybersecurity | Abnormal AI. Abnormal AI. https://abnormalsecurity.com/blog/offensive-ai-defensive-aiHill. 2024. “Understanding Offensive AI Vs. Defensive AI in Cybersecurity | Abnormal AI”. Abnormal AI. https://abnormalsecurity.com/blog/offensive-ai-defensive-ai.Hill. “Understanding Offensive AI Vs. Defensive AI in Cybersecurity | Abnormal AI”. Abnormal AI, 2024, https://abnormalsecurity.com/blog/offensive-ai-defensive-ai.Hill. Understanding Offensive AI vs. Defensive AI in Cybersecurity | Abnormal AI. Abnormal AI https://abnormalsecurity.com/blog/offensive-ai-defensive-ai (2024).Hill, “Understanding Offensive AI vs. Defensive AI in Cybersecurity | Abnormal AI”, Abnormal AI. [Online]. Available: https://abnormalsecurity.com/blog/offensive-ai-defensive-ai
Hjalmar_Wijk(2023). Autonomous replication and adaptation: an attempt at a concrete danger threshold. AI Alignment Forum.Hjalmar_Wijk. (2023, August 17). Autonomous replication and adaptation: an attempt at a concrete danger threshold. AI Alignment Forum. https://alignmentforum.org/posts/vERGLBpDE8m5mpT6t/autonomous-replication-and-adaptation-an-attempt-at-aHjalmar_Wijk. 2023. “Autonomous Replication and Adaptation: An Attempt at a Concrete Danger Threshold”. AI Alignment Forum, August 17. https://alignmentforum.org/posts/vERGLBpDE8m5mpT6t/autonomous-replication-and-adaptation-an-attempt-at-a.Hjalmar_Wijk. “Autonomous Replication and Adaptation: An Attempt at a Concrete Danger Threshold”. AI Alignment Forum, 17 Aug. 2023, https://alignmentforum.org/posts/vERGLBpDE8m5mpT6t/autonomous-replication-and-adaptation-an-attempt-at-a.Hjalmar_Wijk. Autonomous replication and adaptation: an attempt at a concrete danger threshold. AI Alignment Forum https://alignmentforum.org/posts/vERGLBpDE8m5mpT6t/autonomous-replication-and-adaptation-an-attempt-at-a (2023).Hjalmar_Wijk, “Autonomous replication and adaptation: an attempt at a concrete danger threshold”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/vERGLBpDE8m5mpT6t/autonomous-replication-and-adaptation-an-attempt-at-a
Ho(2022). Grokking “Forecasting TAI with biological anchors”. Epoch AI.Ho. (2022). Grokking “Forecasting TAI with biological anchors”. Epoch AI. https://epoch.ai/blog/grokking-bioanchorsHo. 2022. “Grokking “Forecasting TAI with Biological Anchors””. Epoch AI. https://epoch.ai/blog/grokking-bioanchors.Ho. “Grokking “Forecasting TAI with Biological Anchors””. Epoch AI, 2022, https://epoch.ai/blog/grokking-bioanchors.Ho. Grokking “Forecasting TAI with biological anchors”. Epoch AI https://epoch.ai/blog/grokking-bioanchors (2022).Ho, “Grokking “Forecasting TAI with biological anchors””, Epoch AI. [Online]. Available: https://epoch.ai/blog/grokking-bioanchors
Ho(2025). Where’s my ten minute AGI?.Ho. (2025). Where’s my ten minute AGI?. https://epochai.substack.com/p/wheres-my-ten-minute-agiHo. 2025. Where’s My Ten Minute AGI?. Edition. https://epochai.substack.com/p/wheres-my-ten-minute-agi.Ho. Where’s My Ten Minute AGI?. 2025, https://epochai.substack.com/p/wheres-my-ten-minute-agi.Ho. Where’s my ten minute AGI?. https://epochai.substack.com/p/wheres-my-ten-minute-agi (2025).Ho, “Where’s my ten minute AGI?”. [Online]. Available: https://epochai.substack.com/p/wheres-my-ten-minute-agi
Ho et al.(2023). Limits to the energy efficiency of CMOS microprocessors. Epoch AI.Ho et al. (2023). Limits to the energy efficiency of CMOS microprocessors. Epoch AI. https://epoch.ai/blog/limits-to-the-energy-efficiency-of-cmos-microprocessorsHo et al. 2023. “Limits to the Energy Efficiency of CMOS Microprocessors”. Epoch AI. https://epoch.ai/blog/limits-to-the-energy-efficiency-of-cmos-microprocessors.Ho et al. “Limits to the Energy Efficiency of CMOS Microprocessors”. Epoch AI, 2023, https://epoch.ai/blog/limits-to-the-energy-efficiency-of-cmos-microprocessors.Ho et al. Limits to the energy efficiency of CMOS microprocessors. Epoch AI https://epoch.ai/blog/limits-to-the-energy-efficiency-of-cmos-microprocessors (2023).Ho et al., “Limits to the energy efficiency of CMOS microprocessors”, Epoch AI. [Online]. Available: https://epoch.ai/blog/limits-to-the-energy-efficiency-of-cmos-microprocessors
Ho et al.(2024). Algorithmic progress in language models. Epoch AI.Ho et al. (2024). Algorithmic progress in language models. Epoch AI. https://epoch.ai/blog/algorithmic-progress-in-language-modelsHo et al. 2024. “Algorithmic Progress in Language Models”. Epoch AI. https://epoch.ai/blog/algorithmic-progress-in-language-models.Ho et al. “Algorithmic Progress in Language Models”. Epoch AI, 2024, https://epoch.ai/blog/algorithmic-progress-in-language-models.Ho et al. Algorithmic progress in language models. Epoch AI https://epoch.ai/blog/algorithmic-progress-in-language-models (2024).Ho et al., “Algorithmic progress in language models”, Epoch AI. [Online]. Available: https://epoch.ai/blog/algorithmic-progress-in-language-models
Ho, L. et al.(2023). International Institutions for Advanced AI. arXiv.Ho, L., Barnhart, J., Trager, R., Bengio, Y., Brundage, M., Carnegie, A., Chowdhury, R., Dafoe, A., Hadfield, G., Levi, M., & Snidal, D. (2023). International Institutions for Advanced AI. In arXiv. https://arxiv.org/abs/2307.04699Ho, L., J. Barnhart, R. Trager, et al. 2023. “International Institutions for Advanced AI”. In arXiv. Preprint, July 10. https://arxiv.org/abs/2307.04699.Ho, L., et al. “International Institutions for Advanced AI”. arXiv, 10 July 2023, https://arxiv.org/abs/2307.04699.Ho, L. et al. International Institutions for Advanced AI. arXiv Preprint at https://arxiv.org/abs/2307.04699 (2023).L. Ho et al., “International Institutions for Advanced AI”, Jul. 10, 2023. [Online]. Available: https://arxiv.org/abs/2307.04699
Hobbahn et al.(2023). Trends in machine learning hardware. Epoch AI.Hobbahn et al. (2023). Trends in machine learning hardware. Epoch AI. https://epoch.ai/blog/trends-in-machine-learning-hardwareHobbahn et al. 2023. “Trends in Machine Learning Hardware”. Epoch AI. https://epoch.ai/blog/trends-in-machine-learning-hardware.Hobbahn et al. “Trends in Machine Learning Hardware”. Epoch AI, 2023, https://epoch.ai/blog/trends-in-machine-learning-hardware.Hobbahn et al. Trends in machine learning hardware. Epoch AI https://epoch.ai/blog/trends-in-machine-learning-hardware (2023).Hobbahn et al., “Trends in machine learning hardware”, Epoch AI. [Online]. Available: https://epoch.ai/blog/trends-in-machine-learning-hardware
Hoffmann, J. et al.(2022). Training Compute-Optimal Large Language Models. arXiv.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. de L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., … Sifre, L. (2022). Training Compute-Optimal Large Language Models. In arXiv. https://arxiv.org/abs/2203.15556Hoffmann, J., S. Borgeaud, A. Mensch, et al. 2022. “Training Compute-Optimal Large Language Models”. In arXiv. Preprint, March 29. https://arxiv.org/abs/2203.15556.Hoffmann, J., et al. “Training Compute-Optimal Large Language Models”. arXiv, 29 Mar. 2022, https://arxiv.org/abs/2203.15556.Hoffmann, J. et al. Training Compute-Optimal Large Language Models. arXiv Preprint at https://arxiv.org/abs/2203.15556 (2022).J. Hoffmann et al., “Training Compute-Optimal Large Language Models”, Mar. 29, 2022. [Online]. Available: https://arxiv.org/abs/2203.15556
HoldenKarnofsky(2022). AI Safety Seems Hard to Measure. LessWrong.HoldenKarnofsky. (2022, December 8). AI Safety Seems Hard to Measure. LessWrong. https://lesswrong.com/posts/7gkXuHEm6CqEGT2mg/ai-safety-seems-hard-to-measureHoldenKarnofsky. 2022. “AI Safety Seems Hard to Measure”. LessWrong, December 8. https://lesswrong.com/posts/7gkXuHEm6CqEGT2mg/ai-safety-seems-hard-to-measure.HoldenKarnofsky. “AI Safety Seems Hard to Measure”. LessWrong, 8 Dec. 2022, https://lesswrong.com/posts/7gkXuHEm6CqEGT2mg/ai-safety-seems-hard-to-measure.HoldenKarnofsky. AI Safety Seems Hard to Measure. LessWrong https://lesswrong.com/posts/7gkXuHEm6CqEGT2mg/ai-safety-seems-hard-to-measure (2022).HoldenKarnofsky, “AI Safety Seems Hard to Measure”, LessWrong. [Online]. Available: https://lesswrong.com/posts/7gkXuHEm6CqEGT2mg/ai-safety-seems-hard-to-measure
Hooker, S.(2024). On the Limitations of Compute Thresholds as a Governance Strategy. arXiv.Hooker, S. (2024). On the Limitations of Compute Thresholds as a Governance Strategy. In arXiv. https://arxiv.org/abs/2407.05694Hooker, S. 2024. “On the Limitations of Compute Thresholds as a Governance Strategy”. In arXiv. Preprint, July 8. https://arxiv.org/abs/2407.05694.Hooker, S. “On the Limitations of Compute Thresholds as a Governance Strategy”. arXiv, 8 July 2024, https://arxiv.org/abs/2407.05694.Hooker, S. On the Limitations of Compute Thresholds as a Governance Strategy. arXiv Preprint at https://arxiv.org/abs/2407.05694 (2024).S. Hooker, “On the Limitations of Compute Thresholds as a Governance Strategy”, Jul. 08, 2024. [Online]. Available: https://arxiv.org/abs/2407.05694
Hsu, S., Chong, J., Daniels, R., Hargis, S. M. & Lee, J.(2023). Finding Firmer Ground: The Role of High Technology in U.S.-China Relations.Hsu, S., Chong, J.-I., Daniels, R., Hargis, S. M., & Lee, J. (2023). Finding Firmer Ground: The Role of High Technology in U.S.-China Relations. The Carter Center. https://cartercenter.org/resources/pdfs/peace/china/finding-firmer-ground-the-role-of-high-technology-in-u.s.-china-relations.pdfHsu, S., J.-I. Chong, R. Daniels, S. M. Hargis, and J. Lee. 2023. Finding Firmer Ground: The Role of High Technology in U.S.-China Relations. The Carter Center. https://cartercenter.org/resources/pdfs/peace/china/finding-firmer-ground-the-role-of-high-technology-in-u.s.-china-relations.pdf.Hsu, S., et al. Finding Firmer Ground: The Role of High Technology in U.S.-China Relations. The Carter Center, 7 Feb. 2023, https://cartercenter.org/resources/pdfs/peace/china/finding-firmer-ground-the-role-of-high-technology-in-u.s.-china-relations.pdf.Hsu, S., Chong, J.-I., Daniels, R., Hargis, S. M. & Lee, J. Finding Firmer Ground: The Role of High Technology in U.S.-China Relations. https://cartercenter.org/resources/pdfs/peace/china/finding-firmer-ground-the-role-of-high-technology-in-u.s.-china-relations.pdf (2023).S. Hsu, J.-I. Chong, R. Daniels, S. M. Hargis, and J. Lee, “Finding Firmer Ground: The Role of High Technology in U.S.-China Relations”, The Carter Center, Feb. 2023. [Online]. Available: https://cartercenter.org/resources/pdfs/peace/china/finding-firmer-ground-the-role-of-high-technology-in-u.s.-china-relations.pdf
Huang et al.(2023). An Overview of Artificial Intelligence Ethics.Huang et al. (2023). An Overview of Artificial Intelligence Ethics. https://ieeexplore.ieee.org/abstract/document/9844014Huang et al. 2023. “An Overview of Artificial Intelligence Ethics”. https://ieeexplore.ieee.org/abstract/document/9844014.Huang et al. An Overview of Artificial Intelligence Ethics. 2023, https://ieeexplore.ieee.org/abstract/document/9844014.Huang et al. An Overview of Artificial Intelligence Ethics. https://ieeexplore.ieee.org/abstract/document/9844014 (2023).Huang et al., “An Overview of Artificial Intelligence Ethics”. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/9844014
Hubinger, E., Merwijk, C. V., Mikulik, V., Skalse, J. & Garrabrant, S.(2019). Risks from Learned Optimization in Advanced Machine Learning Systems. arXiv.Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., & Garrabrant, S. (2019). Risks from Learned Optimization in Advanced Machine Learning Systems. In arXiv. https://arxiv.org/abs/1906.01820Hubinger, E., C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant. 2019. “Risks from Learned Optimization in Advanced Machine Learning Systems”. In arXiv. Preprint, June 5. https://arxiv.org/abs/1906.01820.Hubinger, E., et al. “Risks from Learned Optimization in Advanced Machine Learning Systems”. arXiv, 5 June 2019, https://arxiv.org/abs/1906.01820.Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J. & Garrabrant, S. Risks from Learned Optimization in Advanced Machine Learning Systems. arXiv Preprint at https://arxiv.org/abs/1906.01820 (2019).E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant, “Risks from Learned Optimization in Advanced Machine Learning Systems”, Jun. 05, 2019. [Online]. Available: https://arxiv.org/abs/1906.01820
Human Rights Watch(2024). Questions and Answers: Israeli Military’s Use of Digital Tools in Gaza. Human Rights Watch.Human Rights Watch. (2024, September 10). Questions and Answers: Israeli Military’s Use of Digital Tools in Gaza. Human Rights Watch. https://hrw.org/news/2024/09/10/questions-and-answers-israeli-militarys-use-digital-tools-gazaHuman Rights Watch. 2024. “Questions and Answers: Israeli Military’s Use of Digital Tools in Gaza”. Human Rights Watch, September 10. https://hrw.org/news/2024/09/10/questions-and-answers-israeli-militarys-use-digital-tools-gaza.Human Rights Watch. “Questions and Answers: Israeli Military’s Use of Digital Tools in Gaza”. Human Rights Watch, 10 Sept. 2024, https://hrw.org/news/2024/09/10/questions-and-answers-israeli-militarys-use-digital-tools-gaza.Human Rights Watch. Questions and Answers: Israeli Military’s Use of Digital Tools in Gaza. Human Rights Watch https://hrw.org/news/2024/09/10/questions-and-answers-israeli-militarys-use-digital-tools-gaza (2024).Human Rights Watch, “Questions and Answers: Israeli Military’s Use of Digital Tools in Gaza”, Human Rights Watch. [Online]. Available: https://hrw.org/news/2024/09/10/questions-and-answers-israeli-militarys-use-digital-tools-gaza
Hung, H. T.(2025). Exploring China's cyber sovereignty concept and artificial intelligence governance model: a machine learning approach. Journal of Computational Social Science.Hung, H. T. (2025). Exploring China's cyber sovereignty concept and artificial intelligence governance model: a machine learning approach. Journal of Computational Social Science, 8(1). https://doi.org/10.1007/s42001-024-00346-8Hung, H. T. 2025. “Exploring China's Cyber Sovereignty Concept and Artificial Intelligence Governance Model: A Machine Learning Approach”. Journal of Computational Social Science 8 (1). https://doi.org/10.1007/s42001-024-00346-8.Hung, H. T. “Exploring China's Cyber Sovereignty Concept and Artificial Intelligence Governance Model: A Machine Learning Approach”. Journal of Computational Social Science, vol. 8, no. 1, Jan. 2025, https://doi.org/10.1007/s42001-024-00346-8.Hung, H. T. Exploring China's cyber sovereignty concept and artificial intelligence governance model: a machine learning approach. Journal of Computational Social Science 8, (2025).H. T. Hung, “Exploring China's cyber sovereignty concept and artificial intelligence governance model: a machine learning approach”, Journal of Computational Social Science, vol. 8, no. 1, Jan. 2025, doi: 10.1007/s42001-024-00346-8.
IBM(2024). What Is Artificial Intelligence (AI)?. IBM.IBM. (2024, November 21). What Is Artificial Intelligence (AI)?. IBM. https://ibm.com/topics/artificial-intelligenceIBM. 2024. “What Is Artificial Intelligence (AI)?”. IBM, November 21. https://ibm.com/topics/artificial-intelligence.IBM. “What Is Artificial Intelligence (AI)?”. IBM, 21 Nov. 2024, https://ibm.com/topics/artificial-intelligence.IBM. What Is Artificial Intelligence (AI)?. IBM https://ibm.com/topics/artificial-intelligence (2024).IBM, “What Is Artificial Intelligence (AI)?”, IBM. [Online]. Available: https://ibm.com/topics/artificial-intelligence
IBM(2025). Deep Blue. IBM.IBM. (2025, November 10). Deep Blue. IBM. https://ibm.com/history/deep-blueIBM. 2025. “Deep Blue”. IBM, November 10. https://ibm.com/history/deep-blue.IBM. “Deep Blue”. IBM, 10 Nov. 2025, https://ibm.com/history/deep-blue.IBM. Deep Blue. IBM https://ibm.com/history/deep-blue (2025).IBM, “Deep Blue”, IBM. [Online]. Available: https://ibm.com/history/deep-blue
Inan, H. A. et al.(2021). Training Data Leakage Analysis in Language Models. arXiv.Inan, H. A., Ramadan, O., Wutschitz, L., Jones, D., Rühle, V., Withers, J., & Sim, R. (2021). Training Data Leakage Analysis in Language Models. In arXiv. https://arxiv.org/abs/2101.05405Inan, H. A., O. Ramadan, L. Wutschitz, et al. 2021. “Training Data Leakage Analysis in Language Models”. In arXiv. Preprint, January 14. https://arxiv.org/abs/2101.05405.Inan, H. A., et al. “Training Data Leakage Analysis in Language Models”. arXiv, 14 Jan. 2021, https://arxiv.org/abs/2101.05405.Inan, H. A. et al. Training Data Leakage Analysis in Language Models. arXiv Preprint at https://arxiv.org/abs/2101.05405 (2021).H. A. Inan et al., “Training Data Leakage Analysis in Language Models”, Jan. 14, 2021. [Online]. Available: https://arxiv.org/abs/2101.05405
Intel(2024). What Is a Trusted Platform Module (TPM)?. Intel.Intel. (2024). What Is a Trusted Platform Module (TPM)?. Intel. https://intel.com/content/www/us/en/business/enterprise-computers/resources/trusted-platform-module.htmlIntel. 2024. “What Is a Trusted Platform Module (TPM)?”. Intel. https://intel.com/content/www/us/en/business/enterprise-computers/resources/trusted-platform-module.html.Intel. “What Is a Trusted Platform Module (TPM)?”. Intel, 2024, https://intel.com/content/www/us/en/business/enterprise-computers/resources/trusted-platform-module.html.Intel. What Is a Trusted Platform Module (TPM)?. Intel https://intel.com/content/www/us/en/business/enterprise-computers/resources/trusted-platform-module.html (2024).Intel, “What Is a Trusted Platform Module (TPM)?”, Intel. [Online]. Available: https://intel.com/content/www/us/en/business/enterprise-computers/resources/trusted-platform-module.html
Irving & Askell(2019). AI Safety Needs Social Scientists. Distill.Irving & Askell. (2019). AI Safety Needs Social Scientists. Distill. https://distill.pub/2019/safety-needs-social-scientistsIrving & Askell. 2019. “AI Safety Needs Social Scientists”. Distill. https://distill.pub/2019/safety-needs-social-scientists.Irving & Askell. “AI Safety Needs Social Scientists”. Distill, 2019, https://distill.pub/2019/safety-needs-social-scientists.Irving & Askell. AI Safety Needs Social Scientists. Distill https://distill.pub/2019/safety-needs-social-scientists (2019).Irving & Askell, “AI Safety Needs Social Scientists”, Distill. [Online]. Available: https://distill.pub/2019/safety-needs-social-scientists
Irving, G., Christiano, P. & Amodei, D.(2018). AI safety via debate. arXiv.Irving, G., Christiano, P., & Amodei, D. (2018). AI safety via debate. In arXiv. https://arxiv.org/abs/1805.00899Irving, G., P. Christiano, and D. Amodei. 2018. “AI Safety via Debate”. In arXiv. Preprint, May 2. https://arxiv.org/abs/1805.00899.Irving, G., et al. “AI Safety via Debate”. arXiv, 2 May 2018, https://arxiv.org/abs/1805.00899.Irving, G., Christiano, P. & Amodei, D. AI safety via debate. arXiv Preprint at https://arxiv.org/abs/1805.00899 (2018).G. Irving, P. Christiano, and D. Amodei, “AI safety via debate”, May 02, 2018. [Online]. Available: https://arxiv.org/abs/1805.00899
Jack Parker-Holder & Shlomi Fruchter(2025). Genie 3: A new frontier for world models.Jack Parker-Holder, & Shlomi Fruchter. (2025, August 5). Genie 3: A new frontier for world models. https://deepmind.google/blog/genie-3-a-new-frontier-for-world-modelsJack Parker-Holder, and Shlomi Fruchter. 2025. “Genie 3: A New Frontier for World Models”. August 5. https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models.Jack Parker-Holder, and Shlomi Fruchter. Genie 3: A New Frontier for World Models. 5 Aug. 2025, https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models.Jack Parker-Holder & Shlomi Fruchter. Genie 3: A new frontier for world models. https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models (2025).Jack Parker-Holder and Shlomi Fruchter, “Genie 3: A new frontier for world models”. [Online]. Available: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models
jacob_cannell(2015). The Brain as a Universal Learning Machine. LessWrong.jacob_cannell. (2015, June 24). The Brain as a Universal Learning Machine. LessWrong. https://lesswrong.com/posts/9Yc7Pp7szcjPgPsjf/the-brain-as-a-universal-learning-machinejacob_cannell. 2015. “The Brain as a Universal Learning Machine”. LessWrong, June 24. https://lesswrong.com/posts/9Yc7Pp7szcjPgPsjf/the-brain-as-a-universal-learning-machine.jacob_cannell. “The Brain as a Universal Learning Machine”. LessWrong, 24 June 2015, https://lesswrong.com/posts/9Yc7Pp7szcjPgPsjf/the-brain-as-a-universal-learning-machine.jacob_cannell. The Brain as a Universal Learning Machine. LessWrong https://lesswrong.com/posts/9Yc7Pp7szcjPgPsjf/the-brain-as-a-universal-learning-machine (2015).jacob_cannell, “The Brain as a Universal Learning Machine”, LessWrong. [Online]. Available: https://lesswrong.com/posts/9Yc7Pp7szcjPgPsjf/the-brain-as-a-universal-learning-machine
jacob_cannell(2022). AI Timelines via Cumulative Optimization Power: Less Long, More Short. LessWrong.jacob_cannell. (2022, October 6). AI Timelines via Cumulative Optimization Power: Less Long, More Short. LessWrong. https://lesswrong.com/posts/3nMpdmt8LrzxQnkGp/ai-timelines-via-cumulative-optimization-power-less-longjacob_cannell. 2022. “AI Timelines via Cumulative Optimization Power: Less Long, More Short”. LessWrong, October 6. https://lesswrong.com/posts/3nMpdmt8LrzxQnkGp/ai-timelines-via-cumulative-optimization-power-less-long.jacob_cannell. “AI Timelines via Cumulative Optimization Power: Less Long, More Short”. LessWrong, 6 Oct. 2022, https://lesswrong.com/posts/3nMpdmt8LrzxQnkGp/ai-timelines-via-cumulative-optimization-power-less-long.jacob_cannell. AI Timelines via Cumulative Optimization Power: Less Long, More Short. LessWrong https://lesswrong.com/posts/3nMpdmt8LrzxQnkGp/ai-timelines-via-cumulative-optimization-power-less-long (2022).jacob_cannell, “AI Timelines via Cumulative Optimization Power: Less Long, More Short”, LessWrong. [Online]. Available: https://lesswrong.com/posts/3nMpdmt8LrzxQnkGp/ai-timelines-via-cumulative-optimization-power-less-long
jacob_cannell(2022). Brain Efficiency: Much More than You Wanted to Know. LessWrong.jacob_cannell. (2022, January 6). Brain Efficiency: Much More than You Wanted to Know. LessWrong. https://lesswrong.com/posts/xwBuoE9p8GE7RAuhd/brain-efficiency-much-more-than-you-wanted-to-knowjacob_cannell. 2022. “Brain Efficiency: Much More Than You Wanted to Know”. LessWrong, January 6. https://lesswrong.com/posts/xwBuoE9p8GE7RAuhd/brain-efficiency-much-more-than-you-wanted-to-know.jacob_cannell. “Brain Efficiency: Much More Than You Wanted to Know”. LessWrong, 6 Jan. 2022, https://lesswrong.com/posts/xwBuoE9p8GE7RAuhd/brain-efficiency-much-more-than-you-wanted-to-know.jacob_cannell. Brain Efficiency: Much More than You Wanted to Know. LessWrong https://lesswrong.com/posts/xwBuoE9p8GE7RAuhd/brain-efficiency-much-more-than-you-wanted-to-know (2022).jacob_cannell, “Brain Efficiency: Much More than You Wanted to Know”, LessWrong. [Online]. Available: https://lesswrong.com/posts/xwBuoE9p8GE7RAuhd/brain-efficiency-much-more-than-you-wanted-to-know
Jaghouar, S. et al.(2024). INTELLECT-1 Technical Report. arXiv.Jaghouar, S., Ong, J. M., Basra, M., Obeid, F., Straube, J., Keiblinger, M., Bakouch, E., Atkins, L., Panahi, M., Goddard, C., Ryabinin, M., & Hagemann, J. (2024). INTELLECT-1 Technical Report. In arXiv. https://arxiv.org/abs/2412.01152Jaghouar, S., J. M. Ong, M. Basra, et al. 2024. “INTELLECT-1 Technical Report”. In arXiv. Preprint, December 2. https://arxiv.org/abs/2412.01152.Jaghouar, S., et al. “INTELLECT-1 Technical Report”. arXiv, 2 Dec. 2024, https://arxiv.org/abs/2412.01152.Jaghouar, S. et al. INTELLECT-1 Technical Report. arXiv Preprint at https://arxiv.org/abs/2412.01152 (2024).S. Jaghouar et al., “INTELLECT-1 Technical Report”, Dec. 02, 2024. [Online]. Available: https://arxiv.org/abs/2412.01152
Jaghouar, S., Ong, J. M. & Hagemann, J.(2024). OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training. arXiv.Jaghouar, S., Ong, J. M., & Hagemann, J. (2024). OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training. In arXiv. https://arxiv.org/abs/2407.07852Jaghouar, S., J. M. Ong, and J. Hagemann. 2024. “OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training”. In arXiv. Preprint, July 10. https://arxiv.org/abs/2407.07852.Jaghouar, S., et al. “OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training”. arXiv, 10 July 2024, https://arxiv.org/abs/2407.07852.Jaghouar, S., Ong, J. M. & Hagemann, J. OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training. arXiv Preprint at https://arxiv.org/abs/2407.07852 (2024).S. Jaghouar, J. M. Ong, and J. Hagemann, “OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training”, Jul. 10, 2024. [Online]. Available: https://arxiv.org/abs/2407.07852
Jagielski(2024). Nvidia Is Dominating the Artificial Intelligence Chip Market, but Apple Has Been Securing Supply From Another Tech Giant.Jagielski. (2024). Nvidia Is Dominating the Artificial Intelligence Chip Market, but Apple Has Been Securing Supply From Another Tech Giant. Internet Archive (https://web.archive.org/web/20260215001652/https://www.nasdaq.com/articles/nvidia-dominating-artificial-intelligence-chip-market-apple-has-been-securing-supply). https://nasdaq.com/articles/nvidia-dominating-artificial-intelligence-chip-market-apple-has-been-securing-supplyJagielski. 2024. “Nvidia Is Dominating the Artificial Intelligence Chip Market, but Apple Has Been Securing Supply From Another Tech Giant”. Https://web.archive.org/web/20260215001652/https://www.nasdaq.com/articles/nvidia-dominating-artificial-intelligence-chip-market-apple-has-been-securing-supply. Internet Archive. https://nasdaq.com/articles/nvidia-dominating-artificial-intelligence-chip-market-apple-has-been-securing-supply.Jagielski. Nvidia Is Dominating the Artificial Intelligence Chip Market, but Apple Has Been Securing Supply From Another Tech Giant. 2024, Internet Archive, https://web.archive.org/web/20260215001652/https://www.nasdaq.com/articles/nvidia-dominating-artificial-intelligence-chip-market-apple-has-been-securing-supply, https://nasdaq.com/articles/nvidia-dominating-artificial-intelligence-chip-market-apple-has-been-securing-supply.Jagielski. Nvidia Is Dominating the Artificial Intelligence Chip Market, but Apple Has Been Securing Supply From Another Tech Giant. https://nasdaq.com/articles/nvidia-dominating-artificial-intelligence-chip-market-apple-has-been-securing-supply (2024).Jagielski, “Nvidia Is Dominating the Artificial Intelligence Chip Market, but Apple Has Been Securing Supply From Another Tech Giant”. Accessed: Feb. 15, 2026. [Online]. Available: https://nasdaq.com/articles/nvidia-dominating-artificial-intelligence-chip-market-apple-has-been-securing-supply
Jaime Sevilla(2025). How far can decentralized training over the internet scale?.Jaime Sevilla. (2025, December 29). How far can decentralized training over the internet scale?. https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scaleJaime Sevilla. 2025. “How Far Can Decentralized Training over the Internet Scale?”. December 29. https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scale.Jaime Sevilla. How Far Can Decentralized Training over the Internet Scale?. 29 Dec. 2025, https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scale.Jaime Sevilla. How far can decentralized training over the internet scale?. https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scale (2025).Jaime Sevilla, “How far can decentralized training over the internet scale?”. [Online]. Available: https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scale
Jaime Sevilla, Tamay Besiroglu & Ege Erdil(2025). Epoch After Hours: AI in 2030, scaling bottlenecks, and explosive growth.Jaime Sevilla, Tamay Besiroglu, & Ege Erdil. (2025, January 17). Epoch After Hours: AI in 2030, scaling bottlenecks, and explosive growth. https://epoch.ai/epoch-after-hours/ai-in-2030Jaime Sevilla, Tamay Besiroglu, and Ege Erdil. 2025. “Epoch After Hours: AI in 2030, Scaling Bottlenecks, and Explosive Growth”. January 17. https://epoch.ai/epoch-after-hours/ai-in-2030.Jaime Sevilla, et al. Epoch After Hours: AI in 2030, Scaling Bottlenecks, and Explosive Growth. 17 Jan. 2025, https://epoch.ai/epoch-after-hours/ai-in-2030.Jaime Sevilla, Tamay Besiroglu & Ege Erdil. Epoch After Hours: AI in 2030, scaling bottlenecks, and explosive growth. https://epoch.ai/epoch-after-hours/ai-in-2030 (2025).Jaime Sevilla, Tamay Besiroglu, and Ege Erdil, “Epoch After Hours: AI in 2030, scaling bottlenecks, and explosive growth”. [Online]. Available: https://epoch.ai/epoch-after-hours/ai-in-2030
Jan Leike(2022). Why I’m optimistic about our alignment approach.Jan Leike. (2022, December 5). Why I’m optimistic about our alignment approach. https://aligned.substack.com/p/alignment-optimismJan Leike. 2022. Why I’m Optimistic About Our Alignment Approach. Edition. December 5. https://aligned.substack.com/p/alignment-optimism.Jan Leike. Why I’m Optimistic About Our Alignment Approach. 5 Dec. 2022, https://aligned.substack.com/p/alignment-optimism.Jan Leike. Why I’m optimistic about our alignment approach. https://aligned.substack.com/p/alignment-optimism (2022).Jan Leike, “Why I’m optimistic about our alignment approach”. [Online]. Available: https://aligned.substack.com/p/alignment-optimism
Jan_Kulveit(2025). AI Control May Increase Existential Risk. LessWrong.Jan_Kulveit. (2025, March 11). AI Control May Increase Existential Risk. LessWrong. https://lesswrong.com/posts/rZcyemEpBHgb2hqLP/ai-control-may-increase-existential-riskJan_Kulveit. 2025. “AI Control May Increase Existential Risk”. LessWrong, March 11. https://lesswrong.com/posts/rZcyemEpBHgb2hqLP/ai-control-may-increase-existential-risk.Jan_Kulveit. “AI Control May Increase Existential Risk”. LessWrong, 11 Mar. 2025, https://lesswrong.com/posts/rZcyemEpBHgb2hqLP/ai-control-may-increase-existential-risk.Jan_Kulveit. AI Control May Increase Existential Risk. LessWrong https://lesswrong.com/posts/rZcyemEpBHgb2hqLP/ai-control-may-increase-existential-risk (2025).Jan_Kulveit, “AI Control May Increase Existential Risk”, LessWrong. [Online]. Available: https://lesswrong.com/posts/rZcyemEpBHgb2hqLP/ai-control-may-increase-existential-risk
Janssen et al.(2025). Responsible governance of generative AI: conceptualizing GenAI as complex adaptive systems. OUP Academic.Janssen et al. (2025). Responsible governance of generative AI: conceptualizing GenAI as complex adaptive systems. Internet Archive (https://web.archive.org/web/20260329203627/https://academic.oup.com/policyandsociety/article/44/1/38/7965776). OUP Academic. https://academic.oup.com/policyandsociety/article/44/1/38/7965776Janssen et al. 2025. “Responsible Governance of Generative AI: Conceptualizing GenAI as Complex Adaptive Systems”. OUP Academic. Https://web.archive.org/web/20260329203627/https://academic.oup.com/policyandsociety/article/44/1/38/7965776. Internet Archive. https://academic.oup.com/policyandsociety/article/44/1/38/7965776.Janssen et al. “Responsible Governance of Generative AI: Conceptualizing GenAI as Complex Adaptive Systems”. OUP Academic, 2025, Internet Archive, https://web.archive.org/web/20260329203627/https://academic.oup.com/policyandsociety/article/44/1/38/7965776, https://academic.oup.com/policyandsociety/article/44/1/38/7965776.Janssen et al. Responsible governance of generative AI: conceptualizing GenAI as complex adaptive systems. OUP Academic https://academic.oup.com/policyandsociety/article/44/1/38/7965776 (2025).Janssen et al., “Responsible governance of generative AI: conceptualizing GenAI as complex adaptive systems”, OUP Academic. Accessed: Mar. 29, 2026. [Online]. Available: https://academic.oup.com/policyandsociety/article/44/1/38/7965776
janus(2022). Simulators. AI Alignment Forum.janus. (2022, September 2). Simulators. AI Alignment Forum. https://alignmentforum.org/posts/vJFdjigzmcXMhNTsx/simulatorsjanus. 2022. “Simulators”. AI Alignment Forum, September 2. https://alignmentforum.org/posts/vJFdjigzmcXMhNTsx/simulators.janus. “Simulators”. AI Alignment Forum, 2 Sept. 2022, https://alignmentforum.org/posts/vJFdjigzmcXMhNTsx/simulators.janus. Simulators. AI Alignment Forum https://alignmentforum.org/posts/vJFdjigzmcXMhNTsx/simulators (2022).janus, “Simulators”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/vJFdjigzmcXMhNTsx/simulators
Jeffrey Ladish & lennart(2022). Information security considerations for AI and the long term future. LessWrong.Jeffrey Ladish, & lennart. (2022, May 2). Information security considerations for AI and the long term future. LessWrong. https://lesswrong.com/posts/2oAxpRuadyjN2ERhe/information-security-considerations-for-ai-and-the-long-termJeffrey Ladish, and lennart. 2022. “Information Security Considerations for AI and the Long Term Future”. LessWrong, May 2. https://lesswrong.com/posts/2oAxpRuadyjN2ERhe/information-security-considerations-for-ai-and-the-long-term.Jeffrey Ladish, and lennart. “Information Security Considerations for AI and the Long Term Future”. LessWrong, 2 May 2022, https://lesswrong.com/posts/2oAxpRuadyjN2ERhe/information-security-considerations-for-ai-and-the-long-term.Jeffrey Ladish & lennart. Information security considerations for AI and the long term future. LessWrong https://lesswrong.com/posts/2oAxpRuadyjN2ERhe/information-security-considerations-for-ai-and-the-long-term (2022).Jeffrey Ladish and lennart, “Information security considerations for AI and the long term future”, LessWrong. [Online]. Available: https://lesswrong.com/posts/2oAxpRuadyjN2ERhe/information-security-considerations-for-ai-and-the-long-term
Jeffrey Ladish(2023). Thoughts on the OpenAI alignment plan: will AI research assistants be net-positive for AI existential risk?. LessWrong.Jeffrey Ladish. (2023, March 10). Thoughts on the OpenAI alignment plan: will AI research assistants be net-positive for AI existential risk?. LessWrong. https://lesswrong.com/posts/6RC3BNopCtzKaTeR6/thoughts-on-the-openai-alignment-plan-will-ai-researchJeffrey Ladish. 2023. “Thoughts on the OpenAI Alignment Plan: Will AI Research Assistants Be Net-positive for AI Existential Risk?”. LessWrong, March 10. https://lesswrong.com/posts/6RC3BNopCtzKaTeR6/thoughts-on-the-openai-alignment-plan-will-ai-research.Jeffrey Ladish. “Thoughts on the OpenAI Alignment Plan: Will AI Research Assistants Be Net-positive for AI Existential Risk?”. LessWrong, 10 Mar. 2023, https://lesswrong.com/posts/6RC3BNopCtzKaTeR6/thoughts-on-the-openai-alignment-plan-will-ai-research.Jeffrey Ladish. Thoughts on the OpenAI alignment plan: will AI research assistants be net-positive for AI existential risk?. LessWrong https://lesswrong.com/posts/6RC3BNopCtzKaTeR6/thoughts-on-the-openai-alignment-plan-will-ai-research (2023).Jeffrey Ladish, “Thoughts on the OpenAI alignment plan: will AI research assistants be net-positive for AI existential risk?”, LessWrong. [Online]. Available: https://lesswrong.com/posts/6RC3BNopCtzKaTeR6/thoughts-on-the-openai-alignment-plan-will-ai-research
Jessica Rumbelow & mwatkins(2023). SolidGoldMagikarp (plus, prompt generation). AI Alignment Forum.Jessica Rumbelow, & mwatkins. (2023, February 5). SolidGoldMagikarp (plus, prompt generation). AI Alignment Forum. https://alignmentforum.org/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generationJessica Rumbelow, and mwatkins. 2023. “SolidGoldMagikarp (plus, Prompt Generation)”. AI Alignment Forum, February 5. https://alignmentforum.org/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation.Jessica Rumbelow, and mwatkins. “SolidGoldMagikarp (plus, Prompt Generation)”. AI Alignment Forum, 5 Feb. 2023, https://alignmentforum.org/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation.Jessica Rumbelow & mwatkins. SolidGoldMagikarp (plus, prompt generation). AI Alignment Forum https://alignmentforum.org/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation (2023).Jessica Rumbelow and mwatkins, “SolidGoldMagikarp (plus, prompt generation)”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation
Ji, Z. et al.(2022). Survey of Hallucination in Natural Language Generation. arXiv.Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Chen, D., Dai, W., Chan, H. S., Madotto, A., & Fung, P. (2022). Survey of Hallucination in Natural Language Generation. In arXiv. https://doi.org/10.1145/3571730Ji, Z., N. Lee, R. Frieske, et al. 2022. “Survey of Hallucination in Natural Language Generation”. In arXiv. Preprint, February 8. https://doi.org/10.1145/3571730.Ji, Z., et al. “Survey of Hallucination in Natural Language Generation”. arXiv, 8 Feb. 2022, https://doi.org/10.1145/3571730.Ji, Z. et al. Survey of Hallucination in Natural Language Generation. arXiv Preprint at https://doi.org/10.1145/3571730 (2022).Z. Ji et al., “Survey of Hallucination in Natural Language Generation”, Feb. 08, 2022. doi: 10.1145/3571730.
Jiang, R., Chiappa, S., Lattimore, T., György, A. & Kohli, P.(2019). Degenerate Feedback Loops in Recommender Systems. arXiv.Jiang, R., Chiappa, S., Lattimore, T., György, A., & Kohli, P. (2019). Degenerate Feedback Loops in Recommender Systems. In arXiv. https://doi.org/10.1145/3306618.3314288Jiang, R., S. Chiappa, T. Lattimore, A. György, and P. Kohli. 2019. “Degenerate Feedback Loops in Recommender Systems”. In arXiv. Preprint, February 27. https://doi.org/10.1145/3306618.3314288.Jiang, R., et al. “Degenerate Feedback Loops in Recommender Systems”. arXiv, 27 Feb. 2019, https://doi.org/10.1145/3306618.3314288.Jiang, R., Chiappa, S., Lattimore, T., György, A. & Kohli, P. Degenerate Feedback Loops in Recommender Systems. arXiv Preprint at https://doi.org/10.1145/3306618.3314288 (2019).R. Jiang, S. Chiappa, T. Lattimore, A. György, and P. Kohli, “Degenerate Feedback Loops in Recommender Systems”, Feb. 27, 2019. doi: 10.1145/3306618.3314288.
Joe O'Brien(2024). Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI.Joe O'Brien. (2024, June 21). Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI. https://iaps.ai/research/coordinated-disclosureJoe O'Brien. 2024. “Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI”. June 21. https://iaps.ai/research/coordinated-disclosure.Joe O'Brien. Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI. 21 June 2024, https://iaps.ai/research/coordinated-disclosure.Joe O'Brien. Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI. https://iaps.ai/research/coordinated-disclosure (2024).Joe O'Brien, “Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI”. [Online]. Available: https://iaps.ai/research/coordinated-disclosure
johnswentworth(2022). Oversight Misses 100% of Thoughts The AI Does Not Think. AI Alignment Forum.johnswentworth. (2022, August 12). Oversight Misses 100% of Thoughts The AI Does Not Think. AI Alignment Forum. https://alignmentforum.org/posts/98c5WMDb3iKdzD4tM/oversight-misses-100-of-thoughts-the-ai-does-not-thinkjohnswentworth. 2022. “Oversight Misses 100% of Thoughts The AI Does Not Think”. AI Alignment Forum, August 12. https://alignmentforum.org/posts/98c5WMDb3iKdzD4tM/oversight-misses-100-of-thoughts-the-ai-does-not-think.johnswentworth. “Oversight Misses 100% of Thoughts The AI Does Not Think”. AI Alignment Forum, 12 Aug. 2022, https://alignmentforum.org/posts/98c5WMDb3iKdzD4tM/oversight-misses-100-of-thoughts-the-ai-does-not-think.johnswentworth. Oversight Misses 100% of Thoughts The AI Does Not Think. AI Alignment Forum https://alignmentforum.org/posts/98c5WMDb3iKdzD4tM/oversight-misses-100-of-thoughts-the-ai-does-not-think (2022).johnswentworth, “Oversight Misses 100% of Thoughts The AI Does Not Think”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/98c5WMDb3iKdzD4tM/oversight-misses-100-of-thoughts-the-ai-does-not-think
johnswentworth(2022). Worlds Where Iterative Design Fails. AI Alignment Forum.johnswentworth. (2022, August 30). Worlds Where Iterative Design Fails. AI Alignment Forum. https://alignmentforum.org/posts/xFotXGEotcKouifky/worlds-where-iterative-design-failsjohnswentworth. 2022. “Worlds Where Iterative Design Fails”. AI Alignment Forum, August 30. https://alignmentforum.org/posts/xFotXGEotcKouifky/worlds-where-iterative-design-fails.johnswentworth. “Worlds Where Iterative Design Fails”. AI Alignment Forum, 30 Aug. 2022, https://alignmentforum.org/posts/xFotXGEotcKouifky/worlds-where-iterative-design-fails.johnswentworth. Worlds Where Iterative Design Fails. AI Alignment Forum https://alignmentforum.org/posts/xFotXGEotcKouifky/worlds-where-iterative-design-fails (2022).johnswentworth, “Worlds Where Iterative Design Fails”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/xFotXGEotcKouifky/worlds-where-iterative-design-fails
johnswentworth(2025). Comment on “johnswentworth's Shortform”. LessWrong.johnswentworth. (2025, April 16). Comment on “johnswentworth's Shortform”. LessWrong. https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=G5zcfntZodH3ZYDaPjohnswentworth. 2025. “Comment on “johnswentworth's Shortform””. LessWrong, April 16. https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=G5zcfntZodH3ZYDaP.johnswentworth. “Comment on “johnswentworth's Shortform””. LessWrong, 16 Apr. 2025, https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=G5zcfntZodH3ZYDaP.johnswentworth. Comment on “johnswentworth's Shortform”. LessWrong https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=G5zcfntZodH3ZYDaP (2025).johnswentworth, “Comment on “johnswentworth's Shortform””, LessWrong. [Online]. Available: https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=G5zcfntZodH3ZYDaP
johnswentworth(2025). The Case Against AI Control Research. LessWrong.johnswentworth. (2025, January 21). The Case Against AI Control Research. LessWrong. https://lesswrong.com/posts/8wBN8cdNAv3c7vt6p/the-case-against-ai-control-researchjohnswentworth. 2025. “The Case Against AI Control Research”. LessWrong, January 21. https://lesswrong.com/posts/8wBN8cdNAv3c7vt6p/the-case-against-ai-control-research.johnswentworth. “The Case Against AI Control Research”. LessWrong, 21 Jan. 2025, https://lesswrong.com/posts/8wBN8cdNAv3c7vt6p/the-case-against-ai-control-research.johnswentworth. The Case Against AI Control Research. LessWrong https://lesswrong.com/posts/8wBN8cdNAv3c7vt6p/the-case-against-ai-control-research (2025).johnswentworth, “The Case Against AI Control Research”, LessWrong. [Online]. Available: https://lesswrong.com/posts/8wBN8cdNAv3c7vt6p/the-case-against-ai-control-research
Jonas B. Sandbrink(2023). Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools. arXiv.Jonas B. Sandbrink. (2023). Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools. In arXiv. https://arxiv.org/abs/2306.13952Jonas B. Sandbrink. 2023. “Artificial Intelligence and Biological Misuse: Differentiating Risks of Language Models and Biological Design Tools”. In arXiv. Preprint, June 24. https://arxiv.org/abs/2306.13952.Jonas B. Sandbrink. “Artificial Intelligence and Biological Misuse: Differentiating Risks of Language Models and Biological Design Tools”. arXiv, 24 June 2023, https://arxiv.org/abs/2306.13952.Jonas B. Sandbrink. Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools. arXiv Preprint at https://arxiv.org/abs/2306.13952 (2023).Jonas B. Sandbrink, “Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools”, Jun. 24, 2023. [Online]. Available: https://arxiv.org/abs/2306.13952
Jonas Schuett, Markus Anderljung, Alexis Carlier, Leonie Koessler & Ben Garfinkel(2024). From Principles to Rules: A Regulatory Approach for Frontier AI.Jonas Schuett, Markus Anderljung, Alexis Carlier, Leonie Koessler, & Ben Garfinkel. (2024, August 21). From Principles to Rules: A Regulatory Approach for Frontier AI. https://governance.ai/research-paper/from-principles-to-rules-a-regulatory-approach-for-frontier-aiJonas Schuett, Markus Anderljung, Alexis Carlier, Leonie Koessler, and Ben Garfinkel. 2024. “From Principles to Rules: A Regulatory Approach for Frontier AI”. August 21. https://governance.ai/research-paper/from-principles-to-rules-a-regulatory-approach-for-frontier-ai.Jonas Schuett, et al. From Principles to Rules: A Regulatory Approach for Frontier AI. 21 Aug. 2024, https://governance.ai/research-paper/from-principles-to-rules-a-regulatory-approach-for-frontier-ai.Jonas Schuett, Markus Anderljung, Alexis Carlier, Leonie Koessler & Ben Garfinkel. From Principles to Rules: A Regulatory Approach for Frontier AI. https://governance.ai/research-paper/from-principles-to-rules-a-regulatory-approach-for-frontier-ai (2024).Jonas Schuett, Markus Anderljung, Alexis Carlier, Leonie Koessler, and Ben Garfinkel, “From Principles to Rules: A Regulatory Approach for Frontier AI”. [Online]. Available: https://governance.ai/research-paper/from-principles-to-rules-a-regulatory-approach-for-frontier-ai
Jonker et al.(2024). What Is AI Alignment?. IBM.Jonker et al. (2024, October 16). What Is AI Alignment?. IBM. https://ibm.com/think/topics/ai-alignmentJonker et al. 2024. “What Is AI Alignment?”. IBM, October 16. https://ibm.com/think/topics/ai-alignment.Jonker et al. “What Is AI Alignment?”. IBM, 16 Oct. 2024, https://ibm.com/think/topics/ai-alignment.Jonker et al. What Is AI Alignment?. IBM https://ibm.com/think/topics/ai-alignment (2024).Jonker et al., “What Is AI Alignment?”, IBM. [Online]. Available: https://ibm.com/think/topics/ai-alignment
joshc(2025). How might we safely pass the buck to AI?. LessWrong.joshc. (2025, February 19). How might we safely pass the buck to AI?. LessWrong. https://lesswrong.com/posts/TTFsKxQThrqgWeXYJ/how-might-we-safely-pass-the-buck-to-aijoshc. 2025. “How Might We Safely Pass the Buck to AI?”. LessWrong, February 19. https://lesswrong.com/posts/TTFsKxQThrqgWeXYJ/how-might-we-safely-pass-the-buck-to-ai.joshc. “How Might We Safely Pass the Buck to AI?”. LessWrong, 19 Feb. 2025, https://lesswrong.com/posts/TTFsKxQThrqgWeXYJ/how-might-we-safely-pass-the-buck-to-ai.joshc. How might we safely pass the buck to AI?. LessWrong https://lesswrong.com/posts/TTFsKxQThrqgWeXYJ/how-might-we-safely-pass-the-buck-to-ai (2025).joshc, “How might we safely pass the buck to AI?”, LessWrong. [Online]. Available: https://lesswrong.com/posts/TTFsKxQThrqgWeXYJ/how-might-we-safely-pass-the-buck-to-ai
jsteinhardt(2022). AI Forecasting: One Year In. LessWrong.jsteinhardt. (2022, July 4). AI Forecasting: One Year In. LessWrong. https://lesswrong.com/posts/CJw2tNHaEimx6nwNy/ai-forecasting-one-year-injsteinhardt. 2022. “AI Forecasting: One Year In”. LessWrong, July 4. https://lesswrong.com/posts/CJw2tNHaEimx6nwNy/ai-forecasting-one-year-in.jsteinhardt. “AI Forecasting: One Year In”. LessWrong, 4 July 2022, https://lesswrong.com/posts/CJw2tNHaEimx6nwNy/ai-forecasting-one-year-in.jsteinhardt. AI Forecasting: One Year In. LessWrong https://lesswrong.com/posts/CJw2tNHaEimx6nwNy/ai-forecasting-one-year-in (2022).jsteinhardt, “AI Forecasting: One Year In”, LessWrong. [Online]. Available: https://lesswrong.com/posts/CJw2tNHaEimx6nwNy/ai-forecasting-one-year-in
jsteinhardt(2022). Future ML Systems Will Be Qualitatively Different. AI Alignment Forum.jsteinhardt. (2022, January 11). Future ML Systems Will Be Qualitatively Different. AI Alignment Forum. https://alignmentforum.org/posts/pZaPhGg2hmmPwByHcjsteinhardt. 2022. “Future ML Systems Will Be Qualitatively Different”. AI Alignment Forum, January 11. https://alignmentforum.org/posts/pZaPhGg2hmmPwByHc.jsteinhardt. “Future ML Systems Will Be Qualitatively Different”. AI Alignment Forum, 11 Jan. 2022, https://alignmentforum.org/posts/pZaPhGg2hmmPwByHc.jsteinhardt. Future ML Systems Will Be Qualitatively Different. AI Alignment Forum https://alignmentforum.org/posts/pZaPhGg2hmmPwByHc (2022).jsteinhardt, “Future ML Systems Will Be Qualitatively Different”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/pZaPhGg2hmmPwByHc
jsteinhardt(2023). AI Forecasting: Two Years In. LessWrong.jsteinhardt. (2023, August 19). AI Forecasting: Two Years In. LessWrong. https://lesswrong.com/posts/SdkexhiynayG2sQCC/ai-forecasting-two-years-injsteinhardt. 2023. “AI Forecasting: Two Years In”. LessWrong, August 19. https://lesswrong.com/posts/SdkexhiynayG2sQCC/ai-forecasting-two-years-in.jsteinhardt. “AI Forecasting: Two Years In”. LessWrong, 19 Aug. 2023, https://lesswrong.com/posts/SdkexhiynayG2sQCC/ai-forecasting-two-years-in.jsteinhardt. AI Forecasting: Two Years In. LessWrong https://lesswrong.com/posts/SdkexhiynayG2sQCC/ai-forecasting-two-years-in (2023).jsteinhardt, “AI Forecasting: Two Years In”, LessWrong. [Online]. Available: https://lesswrong.com/posts/SdkexhiynayG2sQCC/ai-forecasting-two-years-in
jsteinhardt(2023). Emergent Deception and Emergent Optimization. AI Alignment Forum.jsteinhardt. (2023, February 20). Emergent Deception and Emergent Optimization. AI Alignment Forum. https://alignmentforum.org/posts/aEjckcqHZZny9L2zy/emergent-deception-and-emergent-optimizationjsteinhardt. 2023. “Emergent Deception and Emergent Optimization”. AI Alignment Forum, February 20. https://alignmentforum.org/posts/aEjckcqHZZny9L2zy/emergent-deception-and-emergent-optimization.jsteinhardt. “Emergent Deception and Emergent Optimization”. AI Alignment Forum, 20 Feb. 2023, https://alignmentforum.org/posts/aEjckcqHZZny9L2zy/emergent-deception-and-emergent-optimization.jsteinhardt. Emergent Deception and Emergent Optimization. AI Alignment Forum https://alignmentforum.org/posts/aEjckcqHZZny9L2zy/emergent-deception-and-emergent-optimization (2023).jsteinhardt, “Emergent Deception and Emergent Optimization”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/aEjckcqHZZny9L2zy/emergent-deception-and-emergent-optimization
jsteinhardt(2023). What will GPT-2030 look like?. AI Alignment Forum.jsteinhardt. (2023, June 7). What will GPT-2030 look like?. AI Alignment Forum. https://alignmentforum.org/posts/WZXqNYbJhtidjRXSi/what-will-gpt-2030-look-likejsteinhardt. 2023. “What Will GPT-2030 Look Like?”. AI Alignment Forum, June 7. https://alignmentforum.org/posts/WZXqNYbJhtidjRXSi/what-will-gpt-2030-look-like.jsteinhardt. “What Will GPT-2030 Look Like?”. AI Alignment Forum, 7 June 2023, https://alignmentforum.org/posts/WZXqNYbJhtidjRXSi/what-will-gpt-2030-look-like.jsteinhardt. What will GPT-2030 look like?. AI Alignment Forum https://alignmentforum.org/posts/WZXqNYbJhtidjRXSi/what-will-gpt-2030-look-like (2023).jsteinhardt, “What will GPT-2030 look like?”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/WZXqNYbJhtidjRXSi/what-will-gpt-2030-look-like
Julian Schrittwieser et al.(2020). MuZero: Mastering Go, chess, shogi and Atari without rules.Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, & David Silver. (2020, December 23). MuZero: Mastering Go, chess, shogi and Atari without rules. https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rulesJulian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, et al. 2020. “MuZero: Mastering Go, Chess, Shogi and Atari Without Rules”. December 23. https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules.Julian Schrittwieser, et al. MuZero: Mastering Go, Chess, Shogi and Atari Without Rules. 23 Dec. 2020, https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules.Julian Schrittwieser et al. MuZero: Mastering Go, chess, shogi and Atari without rules. https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules (2020).Julian Schrittwieser et al., “MuZero: Mastering Go, chess, shogi and Atari without rules”. [Online]. Available: https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules
Jumper, J. et al.(2021). Highly accurate protein structure prediction with AlphaFold. Nature.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., … Hassabis, D. (2021). Highly accurate protein structure prediction with AlphaFold. Nature. https://doi.org/10.1038/s41586-021-03819-2Jumper, J., R. Evans, A. Pritzel, et al. 2021. “Highly Accurate Protein Structure Prediction with AlphaFold”. Nature, ahead of print, July 15. https://doi.org/10.1038/s41586-021-03819-2.Jumper, J., et al. “Highly Accurate Protein Structure Prediction with AlphaFold”. Nature, July 2021, https://doi.org/10.1038/s41586-021-03819-2.Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature https://doi.org/10.1038/s41586-021-03819-2 (2021) doi:10.1038/s41586-021-03819-2.J. Jumper et al., “Highly accurate protein structure prediction with AlphaFold”, Nature, Jul. 2021, doi: 10.1038/s41586-021-03819-2.
Juneja, J., Bansal, R., Cho, K., Sedoc, J. & Saphra, N.(2022). Linear Connectivity Reveals Generalization Strategies. arXiv.Juneja, J., Bansal, R., Cho, K., Sedoc, J., & Saphra, N. (2022). Linear Connectivity Reveals Generalization Strategies. In arXiv. https://arxiv.org/abs/2205.12411Juneja, J., R. Bansal, K. Cho, J. Sedoc, and N. Saphra. 2022. “Linear Connectivity Reveals Generalization Strategies”. In arXiv. Preprint, May 24. https://arxiv.org/abs/2205.12411.Juneja, J., et al. “Linear Connectivity Reveals Generalization Strategies”. arXiv, 24 May 2022, https://arxiv.org/abs/2205.12411.Juneja, J., Bansal, R., Cho, K., Sedoc, J. & Saphra, N. Linear Connectivity Reveals Generalization Strategies. arXiv Preprint at https://arxiv.org/abs/2205.12411 (2022).J. Juneja, R. Bansal, K. Cho, J. Sedoc, and N. Saphra, “Linear Connectivity Reveals Generalization Strategies”, May 24, 2022. [Online]. Available: https://arxiv.org/abs/2205.12411
Kadavath, S. et al.(2022). Language Models (Mostly) Know What They Know. arXiv.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., … Kaplan, J. (2022). Language Models (Mostly) Know What They Know. In arXiv. https://arxiv.org/abs/2207.05221Kadavath, S., T. Conerly, A. Askell, et al. 2022. “Language Models (Mostly) Know What They Know”. In arXiv. Preprint, July 11. https://arxiv.org/abs/2207.05221.Kadavath, S., et al. “Language Models (Mostly) Know What They Know”. arXiv, 11 July 2022, https://arxiv.org/abs/2207.05221.Kadavath, S. et al. Language Models (Mostly) Know What They Know. arXiv Preprint at https://arxiv.org/abs/2207.05221 (2022).S. Kadavath et al., “Language Models (Mostly) Know What They Know”, Jul. 11, 2022. [Online]. Available: https://arxiv.org/abs/2207.05221
Kaplan, J. et al.(2020). Scaling Laws for Neural Language Models. arXiv.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling Laws for Neural Language Models. In arXiv. https://arxiv.org/abs/2001.08361Kaplan, J., S. McCandlish, T. Henighan, et al. 2020. “Scaling Laws for Neural Language Models”. In arXiv. Preprint, January 23. https://arxiv.org/abs/2001.08361.Kaplan, J., et al. “Scaling Laws for Neural Language Models”. arXiv, 23 Jan. 2020, https://arxiv.org/abs/2001.08361.Kaplan, J. et al. Scaling Laws for Neural Language Models. arXiv Preprint at https://arxiv.org/abs/2001.08361 (2020).J. Kaplan et al., “Scaling Laws for Neural Language Models”, Jan. 23, 2020. [Online]. Available: https://arxiv.org/abs/2001.08361
Kapoor, S. et al.(2024). On the Societal Impact of Open Foundation Models. arXiv.Kapoor, S., Bommasani, R., Klyman, K., Longpre, S., Ramaswami, A., Cihon, P., Hopkins, A., Bankston, K., Biderman, S., Bogen, M., Chowdhury, R., Engler, A., Henderson, P., Jernite, Y., Lazar, S., Maffulli, S., Nelson, A., Pineau, J., Skowron, A., … Narayanan, A. (2024). On the Societal Impact of Open Foundation Models. In arXiv. https://arxiv.org/abs/2403.07918Kapoor, S., R. Bommasani, K. Klyman, et al. 2024. “On the Societal Impact of Open Foundation Models”. In arXiv. Preprint, February 27. https://arxiv.org/abs/2403.07918.Kapoor, S., et al. “On the Societal Impact of Open Foundation Models”. arXiv, 27 Feb. 2024, https://arxiv.org/abs/2403.07918.Kapoor, S. et al. On the Societal Impact of Open Foundation Models. arXiv Preprint at https://arxiv.org/abs/2403.07918 (2024).S. Kapoor et al., “On the Societal Impact of Open Foundation Models”, Feb. 27, 2024. [Online]. Available: https://arxiv.org/abs/2403.07918
Karimi, F.(2023). 'Mom, these bad men have me': She believes scammers cloned her daughter's voice in a fake kidnapping. CNN.Karimi, F. (2023, April 29). 'Mom, these bad men have me': She believes scammers cloned her daughter's voice in a fake kidnapping. CNN. https://edition.cnn.com/2023/04/29/us/ai-scam-calls-kidnapping-cec/index.htmlKarimi, F. 2023. “'Mom, These Bad Men Have Me': She Believes Scammers Cloned Her Daughter's Voice in a Fake Kidnapping”. CNN, April 29. Edition. https://edition.cnn.com/2023/04/29/us/ai-scam-calls-kidnapping-cec/index.html.Karimi, F. “'Mom, These Bad Men Have Me': She Believes Scammers Cloned Her Daughter's Voice in a Fake Kidnapping”. CNN, 29 Apr. 2023, https://edition.cnn.com/2023/04/29/us/ai-scam-calls-kidnapping-cec/index.html.Karimi, F. 'Mom, these bad men have me': She believes scammers cloned her daughter's voice in a fake kidnapping. CNN (2023).F. Karimi, “'Mom, these bad men have me': She believes scammers cloned her daughter's voice in a fake kidnapping”, CNN, Apr. 29, 2023. [Online]. Available: https://edition.cnn.com/2023/04/29/us/ai-scam-calls-kidnapping-cec/index.html
Karnofsky(2016). Some Background on Our Views Regarding Advanced Artificial Intelligence. Coefficient Giving.Karnofsky. (2016). Some Background on Our Views Regarding Advanced Artificial Intelligence. Coefficient Giving. https://openphilanthropy.org/research/some-background-on-our-views-regarding-advanced-artificial-intelligenceKarnofsky. 2016. “Some Background on Our Views Regarding Advanced Artificial Intelligence”. Coefficient Giving. https://openphilanthropy.org/research/some-background-on-our-views-regarding-advanced-artificial-intelligence.Karnofsky. “Some Background on Our Views Regarding Advanced Artificial Intelligence”. Coefficient Giving, 2016, https://openphilanthropy.org/research/some-background-on-our-views-regarding-advanced-artificial-intelligence.Karnofsky. Some Background on Our Views Regarding Advanced Artificial Intelligence. Coefficient Giving https://openphilanthropy.org/research/some-background-on-our-views-regarding-advanced-artificial-intelligence (2016).Karnofsky, “Some Background on Our Views Regarding Advanced Artificial Intelligence”, Coefficient Giving. [Online]. Available: https://openphilanthropy.org/research/some-background-on-our-views-regarding-advanced-artificial-intelligence
Karnofsky(2024). If-Then Commitments for AI Risk Reduction. Carnegie Endowment for International Peace.Karnofsky. (2024, September 13). If-Then Commitments for AI Risk Reduction. Carnegie Endowment for International Peace. https://carnegieendowment.org/research/2024/09/if-then-commitments-for-ai-risk-reduction?lang=enKarnofsky. 2024. “If-Then Commitments for AI Risk Reduction”. Carnegie Endowment for International Peace, September 13. https://carnegieendowment.org/research/2024/09/if-then-commitments-for-ai-risk-reduction?lang=en.Karnofsky. “If-Then Commitments for AI Risk Reduction”. Carnegie Endowment for International Peace, 13 Sept. 2024, https://carnegieendowment.org/research/2024/09/if-then-commitments-for-ai-risk-reduction?lang=en.Karnofsky. If-Then Commitments for AI Risk Reduction. Carnegie Endowment for International Peace https://carnegieendowment.org/research/2024/09/if-then-commitments-for-ai-risk-reduction?lang=en (2024).Karnofsky, “If-Then Commitments for AI Risk Reduction”, Carnegie Endowment for International Peace. [Online]. Available: https://carnegieendowment.org/research/2024/09/if-then-commitments-for-ai-risk-reduction?lang=en
Kasirzadeh, A.(2024). Two Types of AI Existential Risk: Decisive and Accumulative. arXiv.Kasirzadeh, A. (2024). Two Types of AI Existential Risk: Decisive and Accumulative. In arXiv. https://arxiv.org/abs/2401.07836Kasirzadeh, A. 2024. “Two Types of AI Existential Risk: Decisive and Accumulative”. In arXiv. Preprint, January 15. https://arxiv.org/abs/2401.07836.Kasirzadeh, A. “Two Types of AI Existential Risk: Decisive and Accumulative”. arXiv, 15 Jan. 2024, https://arxiv.org/abs/2401.07836.Kasirzadeh, A. Two Types of AI Existential Risk: Decisive and Accumulative. arXiv Preprint at https://arxiv.org/abs/2401.07836 (2024).A. Kasirzadeh, “Two Types of AI Existential Risk: Decisive and Accumulative”, Jan. 15, 2024. [Online]. Available: https://arxiv.org/abs/2401.07836
Katalina Hernandez(2025). Comment on “johnswentworth's Shortform”. LessWrong.Katalina Hernandez. (2025, April 15). Comment on “johnswentworth's Shortform”. LessWrong. https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=NLAW24oxDFuTLT3kxKatalina Hernandez. 2025. “Comment on “johnswentworth's Shortform””. LessWrong, April 15. https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=NLAW24oxDFuTLT3kx.Katalina Hernandez. “Comment on “johnswentworth's Shortform””. LessWrong, 15 Apr. 2025, https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=NLAW24oxDFuTLT3kx.Katalina Hernandez. Comment on “johnswentworth's Shortform”. LessWrong https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=NLAW24oxDFuTLT3kx (2025).Katalina Hernandez, “Comment on “johnswentworth's Shortform””, LessWrong. [Online]. Available: https://lesswrong.com/posts/puv8fRDCH9jx5yhbX/johnswentworth-s-shortform?commentId=NLAW24oxDFuTLT3kx
Keefe(2017). The Family That Built an Empire of Pain. The New Yorker.Keefe. (2017, October 23). The Family That Built an Empire of Pain. The New Yorker. https://newyorker.com/magazine/2017/10/30/the-family-that-built-an-empire-of-painKeefe. 2017. “The Family That Built an Empire of Pain”. The New Yorker, October 23. https://newyorker.com/magazine/2017/10/30/the-family-that-built-an-empire-of-pain.Keefe. “The Family That Built an Empire of Pain”. The New Yorker, 23 Oct. 2017, https://newyorker.com/magazine/2017/10/30/the-family-that-built-an-empire-of-pain.Keefe. The Family That Built an Empire of Pain. The New Yorker https://newyorker.com/magazine/2017/10/30/the-family-that-built-an-empire-of-pain (2017).Keefe, “The Family That Built an Empire of Pain”, The New Yorker. [Online]. Available: https://newyorker.com/magazine/2017/10/30/the-family-that-built-an-empire-of-pain
Kenton, Z. et al.(2024). On scalable oversight with weak LLMs judging strong LLMs. arXiv.Kenton, Z., Siegel, N. Y., Kramár, J., Brown-Cohen, J., Albanie, S., Bulian, J., Agarwal, R., Lindner, D., Tang, Y., Goodman, N. D., & Shah, R. (2024). On scalable oversight with weak LLMs judging strong LLMs. In arXiv. https://arxiv.org/abs/2407.04622Kenton, Z., N. Y. Siegel, J. Kramár, et al. 2024. “On Scalable Oversight with Weak LLMs Judging Strong LLMs”. In arXiv. Preprint, July 5. https://arxiv.org/abs/2407.04622.Kenton, Z., et al. “On Scalable Oversight with Weak LLMs Judging Strong LLMs”. arXiv, 5 July 2024, https://arxiv.org/abs/2407.04622.Kenton, Z. et al. On scalable oversight with weak LLMs judging strong LLMs. arXiv Preprint at https://arxiv.org/abs/2407.04622 (2024).Z. Kenton et al., “On scalable oversight with weak LLMs judging strong LLMs”, Jul. 05, 2024. [Online]. Available: https://arxiv.org/abs/2407.04622
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M. & Tang, P. T. P.(2016). On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. arXiv.Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., & Tang, P. T. P. (2016). On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. In arXiv. https://arxiv.org/abs/1609.04836Keskar, N. S., D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang. 2016. “On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima”. In arXiv. Preprint, September 15. https://arxiv.org/abs/1609.04836.Keskar, N. S., et al. “On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima”. arXiv, 15 Sept. 2016, https://arxiv.org/abs/1609.04836.Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M. & Tang, P. T. P. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. arXiv Preprint at https://arxiv.org/abs/1609.04836 (2016).N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima”, Sep. 15, 2016. [Online]. Available: https://arxiv.org/abs/1609.04836
Khan, A. A. et al.(2021). Ethics of AI: A Systematic Literature Review of Principles and Challenges. arXiv.Khan, A. A., Badshah, S., Liang, P., Khan, B., Waseem, M., Niazi, M., & Akbar, M. A. (2021). Ethics of AI: A Systematic Literature Review of Principles and Challenges. In arXiv. https://arxiv.org/abs/2109.07906Khan, A. A., S. Badshah, P. Liang, et al. 2021. “Ethics of AI: A Systematic Literature Review of Principles and Challenges”. In arXiv. Preprint, September 12. https://arxiv.org/abs/2109.07906.Khan, A. A., et al. “Ethics of AI: A Systematic Literature Review of Principles and Challenges”. arXiv, 12 Sept. 2021, https://arxiv.org/abs/2109.07906.Khan, A. A. et al. Ethics of AI: A Systematic Literature Review of Principles and Challenges. arXiv Preprint at https://arxiv.org/abs/2109.07906 (2021).A. A. Khan et al., “Ethics of AI: A Systematic Literature Review of Principles and Challenges”, Sep. 12, 2021. [Online]. Available: https://arxiv.org/abs/2109.07906
King, J. & Meinhardt, C.(2024). Rethinking Privacy in the AI Era: Policy Provocations for a Data-Centric World.King, J., & Meinhardt, C. (2024). Rethinking Privacy in the AI Era: Policy Provocations for a Data-Centric World. Stanford Institute for Human-Centered Artificial Intelligence. https://hai.stanford.edu/sites/default/files/2024-02/White-Paper-Rethinking-Privacy-AI-Era.pdfKing, J., and C. Meinhardt. 2024. Rethinking Privacy in the AI Era: Policy Provocations for a Data-Centric World. Stanford Institute for Human-Centered Artificial Intelligence. https://hai.stanford.edu/sites/default/files/2024-02/White-Paper-Rethinking-Privacy-AI-Era.pdf.King, J., and C. Meinhardt. Rethinking Privacy in the AI Era: Policy Provocations for a Data-Centric World. Stanford Institute for Human-Centered Artificial Intelligence, Feb. 2024, https://hai.stanford.edu/sites/default/files/2024-02/White-Paper-Rethinking-Privacy-AI-Era.pdf.King, J. & Meinhardt, C. Rethinking Privacy in the AI Era: Policy Provocations for a Data-Centric World. https://hai.stanford.edu/sites/default/files/2024-02/White-Paper-Rethinking-Privacy-AI-Era.pdf (2024).J. King and C. Meinhardt, “Rethinking Privacy in the AI Era: Policy Provocations for a Data-Centric World”, Stanford Institute for Human-Centered Artificial Intelligence, Feb. 2024. [Online]. Available: https://hai.stanford.edu/sites/default/files/2024-02/White-Paper-Rethinking-Privacy-AI-Era.pdf
Kinniment, M. et al.(2023). Evaluating Language-Model Agents on Realistic Autonomous Tasks. arXiv.Kinniment, M., Sato, L. J. K., Du, H., Goodrich, B., Hasin, M., Chan, L., Miles, L. H., Lin, T. R., Wijk, H., Burget, J., Ho, A., Barnes, E., & Christiano, P. (2023). Evaluating Language-Model Agents on Realistic Autonomous Tasks. In arXiv. https://arxiv.org/abs/2312.11671Kinniment, M., L. J. K. Sato, H. Du, et al. 2023. “Evaluating Language-Model Agents on Realistic Autonomous Tasks”. In arXiv. Preprint, December 18. https://arxiv.org/abs/2312.11671.Kinniment, M., et al. “Evaluating Language-Model Agents on Realistic Autonomous Tasks”. arXiv, 18 Dec. 2023, https://arxiv.org/abs/2312.11671.Kinniment, M. et al. Evaluating Language-Model Agents on Realistic Autonomous Tasks. arXiv Preprint at https://arxiv.org/abs/2312.11671 (2023).M. Kinniment et al., “Evaluating Language-Model Agents on Realistic Autonomous Tasks”, Dec. 18, 2023. [Online]. Available: https://arxiv.org/abs/2312.11671
Kirilenko, A., Kyle, A. S., Samadi, M. & Tuzun, T.(2017). The Flash Crash: High-Frequency Trading in an Electronic Market. The Journal of Finance.Kirilenko, A., Kyle, A. S., Samadi, M., & Tuzun, T. (2017). The Flash Crash: High-Frequency Trading in an Electronic Market. The Journal of Finance, 72(3), 967–998. https://doi.org/10.1111/jofi.12498Kirilenko, A., A. S. Kyle, M. Samadi, and T. Tuzun. 2017. “The Flash Crash: High-Frequency Trading in an Electronic Market”. The Journal of Finance 72 (3): 967–98. https://doi.org/10.1111/jofi.12498.Kirilenko, A., et al. “The Flash Crash: High-Frequency Trading in an Electronic Market”. The Journal of Finance, vol. 72, no. 3, Apr. 2017, pp. 967–98, https://doi.org/10.1111/jofi.12498.Kirilenko, A., Kyle, A. S., Samadi, M. & Tuzun, T. The Flash Crash: High-Frequency Trading in an Electronic Market. The Journal of Finance 72, 967–998 (2017).A. Kirilenko, A. S. Kyle, M. Samadi, and T. Tuzun, “The Flash Crash: High-Frequency Trading in an Electronic Market”, The Journal of Finance, vol. 72, no. 3, pp. 967–998, Apr. 2017, doi: 10.1111/jofi.12498.
Kirillov, A. et al.(2023). Segment Anything. arXiv.Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment Anything. In arXiv. https://arxiv.org/abs/2304.02643Kirillov, A., E. Mintun, N. Ravi, et al. 2023. “Segment Anything”. In arXiv. Preprint, April 5. https://arxiv.org/abs/2304.02643.Kirillov, A., et al. “Segment Anything”. arXiv, 5 Apr. 2023, https://arxiv.org/abs/2304.02643.Kirillov, A. et al. Segment Anything. arXiv Preprint at https://arxiv.org/abs/2304.02643 (2023).A. Kirillov et al., “Segment Anything”, Apr. 05, 2023. [Online]. Available: https://arxiv.org/abs/2304.02643
Kleinberg, J. & Raghavan, M.(2021). Algorithmic Monoculture and Social Welfare. arXiv.Kleinberg, J., & Raghavan, M. (2021). Algorithmic Monoculture and Social Welfare. In arXiv. https://doi.org/10.1073/pnas.2018340118Kleinberg, J., and M. Raghavan. 2021. “Algorithmic Monoculture and Social Welfare”. In arXiv. Preprint, January 14. https://doi.org/10.1073/pnas.2018340118.Kleinberg, J., and M. Raghavan. “Algorithmic Monoculture and Social Welfare”. arXiv, 14 Jan. 2021, https://doi.org/10.1073/pnas.2018340118.Kleinberg, J. & Raghavan, M. Algorithmic Monoculture and Social Welfare. arXiv Preprint at https://doi.org/10.1073/pnas.2018340118 (2021).J. Kleinberg and M. Raghavan, “Algorithmic Monoculture and Social Welfare”, Jan. 14, 2021. doi: 10.1073/pnas.2018340118.
Koessler, L. & Schuett, J.(2023). Risk assessment at AGI companies: A review of popular risk assessment techniques from other safety-critical industries. arXiv.Koessler, L., & Schuett, J. (2023). Risk assessment at AGI companies: A review of popular risk assessment techniques from other safety-critical industries. In arXiv. https://arxiv.org/abs/2307.08823Koessler, L., and J. Schuett. 2023. “Risk Assessment at AGI Companies: A Review of Popular Risk Assessment Techniques from Other Safety-critical Industries”. In arXiv. Preprint, July 17. https://arxiv.org/abs/2307.08823.Koessler, L., and J. Schuett. “Risk Assessment at AGI Companies: A Review of Popular Risk Assessment Techniques from Other Safety-critical Industries”. arXiv, 17 July 2023, https://arxiv.org/abs/2307.08823.Koessler, L. & Schuett, J. Risk assessment at AGI companies: A review of popular risk assessment techniques from other safety-critical industries. arXiv Preprint at https://arxiv.org/abs/2307.08823 (2023).L. Koessler and J. Schuett, “Risk assessment at AGI companies: A review of popular risk assessment techniques from other safety-critical industries”, Jul. 17, 2023. [Online]. Available: https://arxiv.org/abs/2307.08823
Kolt, N. et al.(2024). Responsible Reporting for Frontier AI Development. arXiv.Kolt, N., Anderljung, M., Barnhart, J., Brass, A., Esvelt, K., Hadfield, G. K., Heim, L., Rodriguez, M., Sandbrink, J. B., & Woodside, T. (2024). Responsible Reporting for Frontier AI Development. In arXiv. https://arxiv.org/abs/2404.02675Kolt, N., M. Anderljung, J. Barnhart, et al. 2024. “Responsible Reporting for Frontier AI Development”. In arXiv. Preprint, April 3. https://arxiv.org/abs/2404.02675.Kolt, N., et al. “Responsible Reporting for Frontier AI Development”. arXiv, 3 Apr. 2024, https://arxiv.org/abs/2404.02675.Kolt, N. et al. Responsible Reporting for Frontier AI Development. arXiv Preprint at https://arxiv.org/abs/2404.02675 (2024).N. Kolt et al., “Responsible Reporting for Frontier AI Development”, Apr. 03, 2024. [Online]. Available: https://arxiv.org/abs/2404.02675
Korbak, T. et al.(2023). Pretraining Language Models with Human Preferences. arXiv.Korbak, T., Shi, K., Chen, A., Bhalerao, R., Buckley, C. L., Phang, J., Bowman, S. R., & Perez, E. (2023). Pretraining Language Models with Human Preferences. In arXiv. https://arxiv.org/abs/2302.08582Korbak, T., K. Shi, A. Chen, et al. 2023. “Pretraining Language Models with Human Preferences”. In arXiv. Preprint, February 16. https://arxiv.org/abs/2302.08582.Korbak, T., et al. “Pretraining Language Models with Human Preferences”. arXiv, 16 Feb. 2023, https://arxiv.org/abs/2302.08582.Korbak, T. et al. Pretraining Language Models with Human Preferences. arXiv Preprint at https://arxiv.org/abs/2302.08582 (2023).T. Korbak et al., “Pretraining Language Models with Human Preferences”, Feb. 16, 2023. [Online]. Available: https://arxiv.org/abs/2302.08582
Korbak, T. et al.(2025). Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arXiv.Korbak, T., Balesni, M., Barnes, E., Bengio, Y., Benton, J., Bloom, J., Chen, M., Cooney, A., Dafoe, A., Dragan, A., Emmons, S., Evans, O., Farhi, D., Greenblatt, R., Hendrycks, D., Hobbhahn, M., Hubinger, E., Irving, G., Jenner, E., … Mikulik, V. (2025). Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. In arXiv. https://arxiv.org/abs/2507.11473Korbak, T., M. Balesni, E. Barnes, et al. 2025. “Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety”. In arXiv. Preprint, July 15. https://arxiv.org/abs/2507.11473.Korbak, T., et al. “Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety”. arXiv, 15 July 2025, https://arxiv.org/abs/2507.11473.Korbak, T. et al. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arXiv Preprint at https://arxiv.org/abs/2507.11473 (2025).T. Korbak et al., “Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety”, Jul. 15, 2025. [Online]. Available: https://arxiv.org/abs/2507.11473
Korbak, T., Clymer, J., Hilton, B., Shlegeris, B. & Irving, G.(2025). A sketch of an AI control safety case. arXiv.Korbak, T., Clymer, J., Hilton, B., Shlegeris, B., & Irving, G. (2025). A sketch of an AI control safety case. In arXiv. https://arxiv.org/abs/2501.17315Korbak, T., J. Clymer, B. Hilton, B. Shlegeris, and G. Irving. 2025. “A Sketch of an AI Control Safety Case”. In arXiv. Preprint, January 28. https://arxiv.org/abs/2501.17315.Korbak, T., et al. “A Sketch of an AI Control Safety Case”. arXiv, 28 Jan. 2025, https://arxiv.org/abs/2501.17315.Korbak, T., Clymer, J., Hilton, B., Shlegeris, B. & Irving, G. A sketch of an AI control safety case. arXiv Preprint at https://arxiv.org/abs/2501.17315 (2025).T. Korbak, J. Clymer, B. Hilton, B. Shlegeris, and G. Irving, “A sketch of an AI control safety case”, Jan. 28, 2025. [Online]. Available: https://arxiv.org/abs/2501.17315
Korzekwa(2020). Time for AI to cross the human performance range in ImageNet image classification. AI Impacts.Korzekwa. (2020, October 19). Time for AI to cross the human performance range in ImageNet image classification. AI Impacts. https://aiimpacts.org/time-for-ai-to-cross-the-human-performance-range-in-imagenet-image-classificationKorzekwa. 2020. “Time for AI to Cross the Human Performance Range in ImageNet Image Classification”. AI Impacts, October 19. https://aiimpacts.org/time-for-ai-to-cross-the-human-performance-range-in-imagenet-image-classification.Korzekwa. “Time for AI to Cross the Human Performance Range in ImageNet Image Classification”. AI Impacts, 19 Oct. 2020, https://aiimpacts.org/time-for-ai-to-cross-the-human-performance-range-in-imagenet-image-classification.Korzekwa. Time for AI to cross the human performance range in ImageNet image classification. AI Impacts https://aiimpacts.org/time-for-ai-to-cross-the-human-performance-range-in-imagenet-image-classification (2020).Korzekwa, “Time for AI to cross the human performance range in ImageNet image classification”, AI Impacts. [Online]. Available: https://aiimpacts.org/time-for-ai-to-cross-the-human-performance-range-in-imagenet-image-classification
Kosinski, M.(2023). Evaluating Large Language Models in Theory of Mind Tasks. arXiv.Kosinski, M. (2023). Evaluating Large Language Models in Theory of Mind Tasks. In arXiv. https://doi.org/10.1073/pnas.2405460121Kosinski, M. 2023. “Evaluating Large Language Models in Theory of Mind Tasks”. In arXiv. Preprint, February 4. https://doi.org/10.1073/pnas.2405460121.Kosinski, M. “Evaluating Large Language Models in Theory of Mind Tasks”. arXiv, 4 Feb. 2023, https://doi.org/10.1073/pnas.2405460121.Kosinski, M. Evaluating Large Language Models in Theory of Mind Tasks. arXiv Preprint at https://doi.org/10.1073/pnas.2405460121 (2023).M. Kosinski, “Evaluating Large Language Models in Theory of Mind Tasks”, Feb. 04, 2023. doi: 10.1073/pnas.2405460121.
Kovařík, V. & Carey, R.(2019). (When) Is Truth-telling Favored in AI Debate?. arXiv.Kovařík, V., & Carey, R. (2019). (When) Is Truth-telling Favored in AI Debate?. In arXiv. https://arxiv.org/abs/1911.04266Kovařík, V., and R. Carey. 2019. “(When) Is Truth-telling Favored in AI Debate?”. In arXiv. Preprint, November 11. https://arxiv.org/abs/1911.04266.Kovařík, V., and R. Carey. “(When) Is Truth-telling Favored in AI Debate?”. arXiv, 11 Nov. 2019, https://arxiv.org/abs/1911.04266.Kovařík, V. & Carey, R. (When) Is Truth-telling Favored in AI Debate?. arXiv Preprint at https://arxiv.org/abs/1911.04266 (2019).V. Kovařík and R. Carey, “(When) Is Truth-telling Favored in AI Debate?”, Nov. 11, 2019. [Online]. Available: https://arxiv.org/abs/1911.04266
Krakovna et al.(2020). Specification gaming: the flip side of AI ingenuity. Google DeepMind.Krakovna et al. (2020). Specification gaming: the flip side of AI ingenuity. Google DeepMind. https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuityKrakovna et al. 2020. “Specification Gaming: The Flip Side of AI Ingenuity”. Google DeepMind. https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity.Krakovna et al. “Specification Gaming: The Flip Side of AI Ingenuity”. Google DeepMind, 2020, https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity.Krakovna et al. Specification gaming: the flip side of AI ingenuity. Google DeepMind https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity (2020).Krakovna et al., “Specification gaming: the flip side of AI ingenuity”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity
Kreps & Kriner(2023). How AI Threatens Democracy. Journal of Democracy.Kreps & Kriner. (2023). How AI Threatens Democracy. Journal of Democracy. https://journalofdemocracy.org/articles/how-ai-threatens-democracyKreps & Kriner. 2023. “How AI Threatens Democracy”. Journal of Democracy. https://journalofdemocracy.org/articles/how-ai-threatens-democracy.Kreps & Kriner. “How AI Threatens Democracy”. Journal of Democracy, 2023, https://journalofdemocracy.org/articles/how-ai-threatens-democracy.Kreps & Kriner. How AI Threatens Democracy. Journal of Democracy https://journalofdemocracy.org/articles/how-ai-threatens-democracy (2023).Kreps & Kriner, “How AI Threatens Democracy”, Journal of Democracy. [Online]. Available: https://journalofdemocracy.org/articles/how-ai-threatens-democracy
Krizhevsky(2009). CIFAR-10 and CIFAR-100 datasets.Krizhevsky. (2009). CIFAR-10 and CIFAR-100 datasets. Internet Archive (https://web.archive.org/web/20260607155414/https://www.cs.toronto.edu/%7Ekriz/cifar.html). https://cs.toronto.edu/~kriz/cifar.htmlKrizhevsky. 2009. “CIFAR-10 and CIFAR-100 Datasets”. Https://web.archive.org/web/20260607155414/https://www.cs.toronto.edu/%7Ekriz/cifar.html. Internet Archive. https://cs.toronto.edu/~kriz/cifar.html.Krizhevsky. CIFAR-10 and CIFAR-100 Datasets. 2009, Internet Archive, https://web.archive.org/web/20260607155414/https://www.cs.toronto.edu/%7Ekriz/cifar.html, https://cs.toronto.edu/~kriz/cifar.html.Krizhevsky. CIFAR-10 and CIFAR-100 datasets. https://cs.toronto.edu/~kriz/cifar.html (2009).Krizhevsky, “CIFAR-10 and CIFAR-100 datasets”. Accessed: Jun. 07, 2026. [Online]. Available: https://cs.toronto.edu/~kriz/cifar.html
Krueger, D., Maharaj, T. & Leike, J.(2020). Hidden Incentives for Auto-Induced Distributional Shift. arXiv.Krueger, D., Maharaj, T., & Leike, J. (2020). Hidden Incentives for Auto-Induced Distributional Shift. In arXiv. https://arxiv.org/abs/2009.09153Krueger, D., T. Maharaj, and J. Leike. 2020. “Hidden Incentives for Auto-Induced Distributional Shift”. In arXiv. Preprint, September 19. https://arxiv.org/abs/2009.09153.Krueger, D., et al. “Hidden Incentives for Auto-Induced Distributional Shift”. arXiv, 19 Sept. 2020, https://arxiv.org/abs/2009.09153.Krueger, D., Maharaj, T. & Leike, J. Hidden Incentives for Auto-Induced Distributional Shift. arXiv Preprint at https://arxiv.org/abs/2009.09153 (2020).D. Krueger, T. Maharaj, and J. Leike, “Hidden Incentives for Auto-Induced Distributional Shift”, Sep. 19, 2020. [Online]. Available: https://arxiv.org/abs/2009.09153
Kwa, T. et al.(2025). Measuring AI Ability to Complete Long Software Tasks. arXiv.Kwa, T., West, B., Becker, J., Deng, A., Garcia, K., Hasin, M., Jawhar, S., Kinniment, M., Rush, N., Arx, S. V., Bloom, R., Broadley, T., Du, H., Goodrich, B., Jurkovic, N., Miles, L. H., Nix, S., Lin, T., Painter, C., … Chan, L. (2025). Measuring AI Ability to Complete Long Software Tasks. In arXiv. https://arxiv.org/abs/2503.14499Kwa, T., B. West, J. Becker, et al. 2025. “Measuring AI Ability to Complete Long Software Tasks”. In arXiv. Preprint, March 18. https://arxiv.org/abs/2503.14499.Kwa, T., et al. “Measuring AI Ability to Complete Long Software Tasks”. arXiv, 18 Mar. 2025, https://arxiv.org/abs/2503.14499.Kwa, T. et al. Measuring AI Ability to Complete Long Software Tasks. arXiv Preprint at https://arxiv.org/abs/2503.14499 (2025).T. Kwa et al., “Measuring AI Ability to Complete Long Software Tasks”, Mar. 18, 2025. [Online]. Available: https://arxiv.org/abs/2503.14499
LaCroix, T. & Luccioni, A. S.(2022). Metaethical Perspectives on 'Benchmarking' AI Ethics. arXiv.LaCroix, T., & Luccioni, A. S. (2022). Metaethical Perspectives on 'Benchmarking' AI Ethics. In arXiv. https://arxiv.org/abs/2204.05151LaCroix, T., and A. S. Luccioni. 2022. “Metaethical Perspectives on 'Benchmarking' AI Ethics”. In arXiv. Preprint, April 11. https://arxiv.org/abs/2204.05151.LaCroix, T., and A. S. Luccioni. “Metaethical Perspectives on 'Benchmarking' AI Ethics”. arXiv, 11 Apr. 2022, https://arxiv.org/abs/2204.05151.LaCroix, T. & Luccioni, A. S. Metaethical Perspectives on 'Benchmarking' AI Ethics. arXiv Preprint at https://arxiv.org/abs/2204.05151 (2022).T. LaCroix and A. S. Luccioni, “Metaethical Perspectives on 'Benchmarking' AI Ethics”, Apr. 11, 2022. [Online]. Available: https://arxiv.org/abs/2204.05151
Lam, R. et al.(2022). GraphCast: Learning skillful medium-range global weather forecasting. arXiv.Lam, R., Sanchez-Gonzalez, A., Willson, M., Wirnsberger, P., Fortunato, M., Alet, F., Ravuri, S., Ewalds, T., Eaton-Rosen, Z., Hu, W., Merose, A., Hoyer, S., Holland, G., Vinyals, O., Stott, J., Pritzel, A., Mohamed, S., & Battaglia, P. (2022). GraphCast: Learning skillful medium-range global weather forecasting. In arXiv. https://arxiv.org/abs/2212.12794Lam, R., A. Sanchez-Gonzalez, M. Willson, et al. 2022. “GraphCast: Learning Skillful Medium-range Global Weather Forecasting”. In arXiv. Preprint, December 24. https://arxiv.org/abs/2212.12794.Lam, R., et al. “GraphCast: Learning Skillful Medium-range Global Weather Forecasting”. arXiv, 24 Dec. 2022, https://arxiv.org/abs/2212.12794.Lam, R. et al. GraphCast: Learning skillful medium-range global weather forecasting. arXiv Preprint at https://arxiv.org/abs/2212.12794 (2022).R. Lam et al., “GraphCast: Learning skillful medium-range global weather forecasting”, Dec. 24, 2022. [Online]. Available: https://arxiv.org/abs/2212.12794
Lancieri et al.(2024). "AI Regulation: Competition, Arbitrage & Regulatory Capture" by Filippo Lancieri, Laura Edelson et al.Lancieri et al. (2024). "AI Regulation: Competition, Arbitrage & Regulatory Capture" by Filippo Lancieri, Laura Edelson et al. https://scholarship.law.georgetown.edu/facpub/2647Lancieri et al. 2024. “"AI Regulation: Competition, Arbitrage & Regulatory Capture" by Filippo Lancieri, Laura Edelson Et Al.”. https://scholarship.law.georgetown.edu/facpub/2647.Lancieri et al. "AI Regulation: Competition, Arbitrage & Regulatory Capture" by Filippo Lancieri, Laura Edelson Et Al. 2024, https://scholarship.law.georgetown.edu/facpub/2647.Lancieri et al. "AI Regulation: Competition, Arbitrage & Regulatory Capture" by Filippo Lancieri, Laura Edelson et al. https://scholarship.law.georgetown.edu/facpub/2647 (2024).Lancieri et al., “"AI Regulation: Competition, Arbitrage & Regulatory Capture" by Filippo Lancieri, Laura Edelson et al.”. [Online]. Available: https://scholarship.law.georgetown.edu/facpub/2647
Langosco, L., Koch, J., Sharkey, L., Pfau, J., Orseau, L. & Krueger, D.(2021). Goal Misgeneralization in Deep Reinforcement Learning. arXiv.Langosco, L., Koch, J., Sharkey, L., Pfau, J., Orseau, L., & Krueger, D. (2021). Goal Misgeneralization in Deep Reinforcement Learning. In arXiv. https://arxiv.org/abs/2105.14111Langosco, L., J. Koch, L. Sharkey, J. Pfau, L. Orseau, and D. Krueger. 2021. “Goal Misgeneralization in Deep Reinforcement Learning”. In arXiv. Preprint, May 28. https://arxiv.org/abs/2105.14111.Langosco, L., et al. “Goal Misgeneralization in Deep Reinforcement Learning”. arXiv, 28 May 2021, https://arxiv.org/abs/2105.14111.Langosco, L. et al. Goal Misgeneralization in Deep Reinforcement Learning. arXiv Preprint at https://arxiv.org/abs/2105.14111 (2021).L. Langosco, J. Koch, L. Sharkey, J. Pfau, L. Orseau, and D. Krueger, “Goal Misgeneralization in Deep Reinforcement Learning”, May 28, 2021. [Online]. Available: https://arxiv.org/abs/2105.14111
Lanham, T. et al.(2023). Measuring Faithfulness in Chain-of-Thought Reasoning. arXiv.Lanham, T., Chen, A., Radhakrishnan, A., Steiner, B., Denison, C., Hernandez, D., Li, D., Durmus, E., Hubinger, E., Kernion, J., Lukošiūtė, K., Nguyen, K., Cheng, N., Joseph, N., Schiefer, N., Rausch, O., Larson, R., McCandlish, S., Kundu, S., … Perez, E. (2023). Measuring Faithfulness in Chain-of-Thought Reasoning. In arXiv. https://arxiv.org/abs/2307.13702Lanham, T., A. Chen, A. Radhakrishnan, et al. 2023. “Measuring Faithfulness in Chain-of-Thought Reasoning”. In arXiv. Preprint, July 17. https://arxiv.org/abs/2307.13702.Lanham, T., et al. “Measuring Faithfulness in Chain-of-Thought Reasoning”. arXiv, 17 July 2023, https://arxiv.org/abs/2307.13702.Lanham, T. et al. Measuring Faithfulness in Chain-of-Thought Reasoning. arXiv Preprint at https://arxiv.org/abs/2307.13702 (2023).T. Lanham et al., “Measuring Faithfulness in Chain-of-Thought Reasoning”, Jul. 17, 2023. [Online]. Available: https://arxiv.org/abs/2307.13702
Lazar, S.(2024). Automatic Authorities: Power and AI. arXiv.Lazar, S. (2024). Automatic Authorities: Power and AI. In arXiv. https://arxiv.org/abs/2404.05990Lazar, S. 2024. “Automatic Authorities: Power and AI”. In arXiv. Preprint, April 9. https://arxiv.org/abs/2404.05990.Lazar, S. “Automatic Authorities: Power and AI”. arXiv, 9 Apr. 2024, https://arxiv.org/abs/2404.05990.Lazar, S. Automatic Authorities: Power and AI. arXiv Preprint at https://arxiv.org/abs/2404.05990 (2024).S. Lazar, “Automatic Authorities: Power and AI”, Apr. 09, 2024. [Online]. Available: https://arxiv.org/abs/2404.05990
Leahy et al.(2024). Understanding AI Extinction Risks.Leahy et al. (2024). Understanding AI Extinction Risks. https://thecompendium.aiLeahy et al. 2024. “Understanding AI Extinction Risks”. https://thecompendium.ai.Leahy et al. Understanding AI Extinction Risks. 2024, https://thecompendium.ai.Leahy et al. Understanding AI Extinction Risks. https://thecompendium.ai (2024).Leahy et al., “Understanding AI Extinction Risks”. [Online]. Available: https://thecompendium.ai
LeCun, Y.(2022). A Path Towards Autonomous Machine Intelligence. OpenReview.LeCun, Y. (2022). A Path Towards Autonomous Machine Intelligence. In OpenReview (Version 0.9.2). https://openreview.net/pdf?id=BZ5a1r-kVsfLeCun, Y. 2022. “A Path Towards Autonomous Machine Intelligence”. In OpenReview, version 0.9.2. Preprint, June 27. https://openreview.net/pdf?id=BZ5a1r-kVsf.LeCun, Y. “A Path Towards Autonomous Machine Intelligence”. OpenReview, Version 0.9.2, 27 June 2022, https://openreview.net/pdf?id=BZ5a1r-kVsf.LeCun, Y. A Path Towards Autonomous Machine Intelligence. OpenReview Preprint at https://openreview.net/pdf?id=BZ5a1r-kVsf (2022).Y. LeCun, “A Path Towards Autonomous Machine Intelligence”, Jun. 27, 2022. [Online]. Available: https://openreview.net/pdf?id=BZ5a1r-kVsf
LeCun, Y.(2025). Yann LeCun "Mathematical Obstacles on the Way to Human-Level AI". YouTube.LeCun, Y. (2025, March 21). Yann LeCun "Mathematical Obstacles on the Way to Human-Level AI" [Video recording]. In YouTube. Joint Mathematics Meetings. https://www.youtube.com/watch?v=ETZfkkv6V7YLeCun, Y. 2025. “Yann LeCun "Mathematical Obstacles on the Way to Human-Level AI"”. YouTube. Joint Mathematics Meetings. https://www.youtube.com/watch?v=ETZfkkv6V7Y.LeCun, Y. “Yann LeCun "Mathematical Obstacles on the Way to Human-Level AI"”. YouTube, Joint Mathematics Meetings, 2025, https://www.youtube.com/watch?v=ETZfkkv6V7Y.LeCun, Y. Yann LeCun "Mathematical Obstacles on the Way to Human-Level AI". YouTube (Joint Mathematics Meetings, 2025).Y. LeCun, Yann LeCun "Mathematical Obstacles on the Way to Human-Level AI", (Mar. 21, 2025). [Online Video]. Available: https://www.youtube.com/watch?v=ETZfkkv6V7Y
Lee Sharkey(2023). Why almost every RL agent does learned optimization. AI Alignment Forum.Lee Sharkey. (2023, February 12). Why almost every RL agent does learned optimization. AI Alignment Forum. https://alignmentforum.org/posts/J8ifgynkfhpmrGrL8/why-almost-every-rl-agent-does-learned-optimizationLee Sharkey. 2023. “Why Almost Every RL Agent Does Learned Optimization”. AI Alignment Forum, February 12. https://alignmentforum.org/posts/J8ifgynkfhpmrGrL8/why-almost-every-rl-agent-does-learned-optimization.Lee Sharkey. “Why Almost Every RL Agent Does Learned Optimization”. AI Alignment Forum, 12 Feb. 2023, https://alignmentforum.org/posts/J8ifgynkfhpmrGrL8/why-almost-every-rl-agent-does-learned-optimization.Lee Sharkey. Why almost every RL agent does learned optimization. AI Alignment Forum https://alignmentforum.org/posts/J8ifgynkfhpmrGrL8/why-almost-every-rl-agent-does-learned-optimization (2023).Lee Sharkey, “Why almost every RL agent does learned optimization”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/J8ifgynkfhpmrGrL8/why-almost-every-rl-agent-does-learned-optimization
Lehalleur, S. P. et al.(2025). You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation. arXiv.Lehalleur, S. P., Hoogland, J., Farrugia-Roberts, M., Wei, S., Oldenziel, A. G., Wang, G., Carroll, L., & Murfet, D. (2025). You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation. In arXiv. https://arxiv.org/abs/2502.05475Lehalleur, S. P., J. Hoogland, M. Farrugia-Roberts, et al. 2025. “You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation”. In arXiv. Preprint, February 8. https://arxiv.org/abs/2502.05475.Lehalleur, S. P., et al. “You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation”. arXiv, 8 Feb. 2025, https://arxiv.org/abs/2502.05475.Lehalleur, S. P. et al. You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation. arXiv Preprint at https://arxiv.org/abs/2502.05475 (2025).S. P. Lehalleur et al., “You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation”, Feb. 08, 2025. [Online]. Available: https://arxiv.org/abs/2502.05475
Leike(2023). Combining weak-to-strong generalization with scalable oversight.Leike. (2023). Combining weak-to-strong generalization with scalable oversight. https://substack.com/home/post/p-139945470Leike. 2023. Combining Weak-to-strong Generalization with Scalable Oversight. Edition. https://substack.com/home/post/p-139945470.Leike. Combining Weak-to-strong Generalization with Scalable Oversight. 2023, https://substack.com/home/post/p-139945470.Leike. Combining weak-to-strong generalization with scalable oversight. https://substack.com/home/post/p-139945470 (2023).Leike, “Combining weak-to-strong generalization with scalable oversight”. [Online]. Available: https://substack.com/home/post/p-139945470
Leike(2023). Self-exfiltration is a key dangerous capability.Leike. (2023). Self-exfiltration is a key dangerous capability. https://aligned.substack.com/p/self-exfiltrationLeike. 2023. Self-exfiltration Is a Key Dangerous Capability. Edition. https://aligned.substack.com/p/self-exfiltration.Leike. Self-exfiltration Is a Key Dangerous Capability. 2023, https://aligned.substack.com/p/self-exfiltration.Leike. Self-exfiltration is a key dangerous capability. https://aligned.substack.com/p/self-exfiltration (2023).Leike, “Self-exfiltration is a key dangerous capability”. [Online]. Available: https://aligned.substack.com/p/self-exfiltration
lennart(2021). Compute Research Questions and Metrics - Transformative AI and Compute [4/4]. LessWrong.lennart. (2021, November 28). Compute Research Questions and Metrics - Transformative AI and Compute [4/4]. LessWrong. https://lesswrong.com/posts/G4KHuYC3pHry6yMhilennart. 2021. “Compute Research Questions and Metrics - Transformative AI and Compute [4/4]”. LessWrong, November 28. https://lesswrong.com/posts/G4KHuYC3pHry6yMhi.lennart. “Compute Research Questions and Metrics - Transformative AI and Compute [4/4]”. LessWrong, 28 Nov. 2021, https://lesswrong.com/posts/G4KHuYC3pHry6yMhi.lennart. Compute Research Questions and Metrics - Transformative AI and Compute [4/4]. LessWrong https://lesswrong.com/posts/G4KHuYC3pHry6yMhi (2021).lennart, “Compute Research Questions and Metrics - Transformative AI and Compute [4/4]”, LessWrong. [Online]. Available: https://lesswrong.com/posts/G4KHuYC3pHry6yMhi
Lennart Heim et al.(2024). Governing Through the Cloud.Lennart Heim, Tim Fist, Janet Egan, Sihao Huang, Stephen Zekany, Robert Trager, Michael A Osborne, & Noa Zilberman. (2024, March 13). Governing Through the Cloud. https://governance.ai/research-paper/governing-through-the-cloudLennart Heim, Tim Fist, Janet Egan, et al. 2024. “Governing Through the Cloud”. March 13. https://governance.ai/research-paper/governing-through-the-cloud.Lennart Heim, et al. Governing Through the Cloud. 13 Mar. 2024, https://governance.ai/research-paper/governing-through-the-cloud.Lennart Heim et al. Governing Through the Cloud. https://governance.ai/research-paper/governing-through-the-cloud (2024).Lennart Heim et al., “Governing Through the Cloud”. [Online]. Available: https://governance.ai/research-paper/governing-through-the-cloud
Lennart Heim, * Markus Anderljung, Emma Bluemke & Robert Trager(2024). Computing Power and the Governance of AI.Lennart Heim, * Markus Anderljung, Emma Bluemke, & Robert Trager. (2024, February 14). Computing Power and the Governance of AI. https://governance.ai/analysis/computing-power-and-the-governance-of-aiLennart Heim, * Markus Anderljung, Emma Bluemke, and Robert Trager. 2024. “Computing Power and the Governance of AI”. February 14. https://governance.ai/analysis/computing-power-and-the-governance-of-ai.Lennart Heim, et al. Computing Power and the Governance of AI. 14 Feb. 2024, https://governance.ai/analysis/computing-power-and-the-governance-of-ai.Lennart Heim, * Markus Anderljung, Emma Bluemke & Robert Trager. Computing Power and the Governance of AI. https://governance.ai/analysis/computing-power-and-the-governance-of-ai (2024).Lennart Heim, * Markus Anderljung, Emma Bluemke, and Robert Trager, “Computing Power and the Governance of AI”. [Online]. Available: https://governance.ai/analysis/computing-power-and-the-governance-of-ai
leogao(2022). Clarifying wireheading terminology. AI Alignment Forum.leogao. (2022, November 24). Clarifying wireheading terminology. AI Alignment Forum. https://alignmentforum.org/posts/REesy8nqvknFFKywm/clarifying-wireheading-terminologyleogao. 2022. “Clarifying Wireheading Terminology”. AI Alignment Forum, November 24. https://alignmentforum.org/posts/REesy8nqvknFFKywm/clarifying-wireheading-terminology.leogao. “Clarifying Wireheading Terminology”. AI Alignment Forum, 24 Nov. 2022, https://alignmentforum.org/posts/REesy8nqvknFFKywm/clarifying-wireheading-terminology.leogao. Clarifying wireheading terminology. AI Alignment Forum https://alignmentforum.org/posts/REesy8nqvknFFKywm/clarifying-wireheading-terminology (2022).leogao, “Clarifying wireheading terminology”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/REesy8nqvknFFKywm/clarifying-wireheading-terminology
Lermen, S., Rogers-Smith, C. & Ladish, J.(2023). LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B. arXiv.Lermen, S., Rogers-Smith, C., & Ladish, J. (2023). LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B. In arXiv. https://arxiv.org/abs/2310.20624Lermen, S., C. Rogers-Smith, and J. Ladish. 2023. “LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B”. In arXiv. Preprint, October 31. https://arxiv.org/abs/2310.20624.Lermen, S., et al. “LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B”. arXiv, 31 Oct. 2023, https://arxiv.org/abs/2310.20624.Lermen, S., Rogers-Smith, C. & Ladish, J. LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B. arXiv Preprint at https://arxiv.org/abs/2310.20624 (2023).S. Lermen, C. Rogers-Smith, and J. Ladish, “LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B”, Oct. 31, 2023. [Online]. Available: https://arxiv.org/abs/2310.20624
Lewis Hammond et al.(2025). Multi-Agent Risks from Advanced AI. arXiv.Lewis Hammond, Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, Chandler Smith, Wolfram Barfuss, Jakob Foerster, Tomáš Gavenčiak, The Anh Han, Edward Hughes, Vojtěch Kovařík, Jan Kulveit, Joel Z. Leibo, Caspar Oesterheld, Christian Schroeder de Witt, Nisarg Shah, Michael Wellman, … Iyad Rahwan. (2025). Multi-Agent Risks from Advanced AI. In arXiv. https://arxiv.org/abs/2502.14143Lewis Hammond, Alan Chan, Jesse Clifton, et al. 2025. “Multi-Agent Risks from Advanced AI”. In arXiv. Preprint, February 19. https://arxiv.org/abs/2502.14143.Lewis Hammond, et al. “Multi-Agent Risks from Advanced AI”. arXiv, 19 Feb. 2025, https://arxiv.org/abs/2502.14143.Lewis Hammond et al. Multi-Agent Risks from Advanced AI. arXiv Preprint at https://arxiv.org/abs/2502.14143 (2025).Lewis Hammond et al., “Multi-Agent Risks from Advanced AI”, Feb. 19, 2025. [Online]. Available: https://arxiv.org/abs/2502.14143
Li, H. et al.(2023). Multi-step Jailbreaking Privacy Attacks on ChatGPT. arXiv.Li, H., Guo, D., Fan, W., Xu, M., Huang, J., Meng, F., & Song, Y. (2023). Multi-step Jailbreaking Privacy Attacks on ChatGPT. In arXiv. https://arxiv.org/abs/2304.05197Li, H., D. Guo, W. Fan, et al. 2023. “Multi-step Jailbreaking Privacy Attacks on ChatGPT”. In arXiv. Preprint, April 11. https://arxiv.org/abs/2304.05197.Li, H., et al. “Multi-step Jailbreaking Privacy Attacks on ChatGPT”. arXiv, 11 Apr. 2023, https://arxiv.org/abs/2304.05197.Li, H. et al. Multi-step Jailbreaking Privacy Attacks on ChatGPT. arXiv Preprint at https://arxiv.org/abs/2304.05197 (2023).H. Li et al., “Multi-step Jailbreaking Privacy Attacks on ChatGPT”, Apr. 11, 2023. [Online]. Available: https://arxiv.org/abs/2304.05197
Li, H., Xu, Z., Taylor, G., Studer, C. & Goldstein, T.(2017). Visualizing the Loss Landscape of Neural Nets. arXiv.Li, H., Xu, Z., Taylor, G., Studer, C., & Goldstein, T. (2017). Visualizing the Loss Landscape of Neural Nets. In arXiv. https://arxiv.org/abs/1712.09913Li, H., Z. Xu, G. Taylor, C. Studer, and T. Goldstein. 2017. “Visualizing the Loss Landscape of Neural Nets”. In arXiv. Preprint, December 28. https://arxiv.org/abs/1712.09913.Li, H., et al. “Visualizing the Loss Landscape of Neural Nets”. arXiv, 28 Dec. 2017, https://arxiv.org/abs/1712.09913.Li, H., Xu, Z., Taylor, G., Studer, C. & Goldstein, T. Visualizing the Loss Landscape of Neural Nets. arXiv Preprint at https://arxiv.org/abs/1712.09913 (2017).H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein, “Visualizing the Loss Landscape of Neural Nets”, Dec. 28, 2017. [Online]. Available: https://arxiv.org/abs/1712.09913
Li, N. et al.(2024). The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning. arXiv.Li, N., Pan, A., Gopal, A., Yue, S., Berrios, D., Gatti, A., Li, J. D., Dombrowski, A.-K., Goel, S., Phan, L., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A. B., Chen, M., Barrass, I., Zhang, O., Zhu, X., … Hendrycks, D. (2024). The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning. In arXiv. https://arxiv.org/abs/2403.03218Li, N., A. Pan, A. Gopal, et al. 2024. “The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning”. In arXiv. Preprint, March 5. https://arxiv.org/abs/2403.03218.Li, N., et al. “The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning”. arXiv, 5 Mar. 2024, https://arxiv.org/abs/2403.03218.Li, N. et al. The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning. arXiv Preprint at https://arxiv.org/abs/2403.03218 (2024).N. Li et al., “The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning”, Mar. 05, 2024. [Online]. Available: https://arxiv.org/abs/2403.03218
Li, Q., Wang, W., Xu, C., Sun, Z. & Yang, M.(2022). Learning Disentangled Representation for One-shot Progressive Face Swapping. arXiv.Li, Q., Wang, W., Xu, C., Sun, Z., & Yang, M.-H. (2022). Learning Disentangled Representation for One-shot Progressive Face Swapping. In arXiv. https://arxiv.org/abs/2203.12985Li, Q., W. Wang, C. Xu, Z. Sun, and M.-H. Yang. 2022. “Learning Disentangled Representation for One-shot Progressive Face Swapping”. In arXiv. Preprint, March 24. https://arxiv.org/abs/2203.12985.Li, Q., et al. “Learning Disentangled Representation for One-shot Progressive Face Swapping”. arXiv, 24 Mar. 2022, https://arxiv.org/abs/2203.12985.Li, Q., Wang, W., Xu, C., Sun, Z. & Yang, M.-H. Learning Disentangled Representation for One-shot Progressive Face Swapping. arXiv Preprint at https://arxiv.org/abs/2203.12985 (2022).Q. Li, W. Wang, C. Xu, Z. Sun, and M.-H. Yang, “Learning Disentangled Representation for One-shot Progressive Face Swapping”, Mar. 24, 2022. [Online]. Available: https://arxiv.org/abs/2203.12985
Li, T. C.(2025). Ending the AI Race: Regulatory Collaboration as Critical Counter-Narrative. Villanova Law Review.Li, T. C. (2025). Ending the AI Race: Regulatory Collaboration as Critical Counter-Narrative. Villanova Law Review, 69(5), 981. https://digitalcommons.law.villanova.edu/cgi/viewcontent.cgi?article=3670&context=vlrLi, T. C. 2025. “Ending the AI Race: Regulatory Collaboration as Critical Counter-Narrative”. Villanova Law Review 69 (5): 981. https://digitalcommons.law.villanova.edu/cgi/viewcontent.cgi?article=3670&context=vlr.Li, T. C. “Ending the AI Race: Regulatory Collaboration as Critical Counter-Narrative”. Villanova Law Review, vol. 69, no. 5, Mar. 2025, p. 981, https://digitalcommons.law.villanova.edu/cgi/viewcontent.cgi?article=3670&context=vlr.Li, T. C. Ending the AI Race: Regulatory Collaboration as Critical Counter-Narrative. Villanova Law Review 69, 981 (2025).T. C. Li, “Ending the AI Race: Regulatory Collaboration as Critical Counter-Narrative”, Villanova Law Review, vol. 69, no. 5, p. 981, Mar. 2025, [Online]. Available: https://digitalcommons.law.villanova.edu/cgi/viewcontent.cgi?article=3670&context=vlr
Liang et al.(2022). Stanford CRFM.Liang et al. (2022). Stanford CRFM. https://crfm.stanford.edu/2022/05/17/community-norms.htmlLiang et al. 2022. “Stanford CRFM”. https://crfm.stanford.edu/2022/05/17/community-norms.html.Liang et al. Stanford CRFM. 2022, https://crfm.stanford.edu/2022/05/17/community-norms.html.Liang et al. Stanford CRFM. https://crfm.stanford.edu/2022/05/17/community-norms.html (2022).Liang et al., “Stanford CRFM”. [Online]. Available: https://crfm.stanford.edu/2022/05/17/community-norms.html
Lightman, H. et al.(2023). Let's Verify Step by Step. arXiv.Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., & Cobbe, K. (2023). Let's Verify Step by Step. In arXiv. https://arxiv.org/abs/2305.20050Lightman, H., V. Kosaraju, Y. Burda, et al. 2023. “Let's Verify Step by Step”. In arXiv. Preprint, May 31. https://arxiv.org/abs/2305.20050.Lightman, H., et al. “Let's Verify Step by Step”. arXiv, 31 May 2023, https://arxiv.org/abs/2305.20050.Lightman, H. et al. Let's Verify Step by Step. arXiv Preprint at https://arxiv.org/abs/2305.20050 (2023).H. Lightman et al., “Let's Verify Step by Step”, May 31, 2023. [Online]. Available: https://arxiv.org/abs/2305.20050
Liu(2024). Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science.Liu. (2024). Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science. Ken Ziyu Liu - Stanford Computer Science. https://ai.stanford.edu/~kzliu/blog/unlearningLiu. 2024. “Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science”. Ken Ziyu Liu - Stanford Computer Science. https://ai.stanford.edu/~kzliu/blog/unlearning.Liu. “Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science”. Ken Ziyu Liu - Stanford Computer Science, 2024, https://ai.stanford.edu/~kzliu/blog/unlearning.Liu. Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science. Ken Ziyu Liu - Stanford Computer Science https://ai.stanford.edu/~kzliu/blog/unlearning (2024).Liu, “Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science”, Ken Ziyu Liu - Stanford Computer Science. [Online]. Available: https://ai.stanford.edu/~kzliu/blog/unlearning
Liu, X., Xu, N., Chen, M. & Xiao, C.(2023). AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. arXiv.Liu, X., Xu, N., Chen, M., & Xiao, C. (2023). AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. In arXiv. https://arxiv.org/abs/2310.04451Liu, X., N. Xu, M. Chen, and C. Xiao. 2023. “AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models”. In arXiv. Preprint, October 3. https://arxiv.org/abs/2310.04451.Liu, X., et al. “AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models”. arXiv, 3 Oct. 2023, https://arxiv.org/abs/2310.04451.Liu, X., Xu, N., Chen, M. & Xiao, C. AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. arXiv Preprint at https://arxiv.org/abs/2310.04451 (2023).X. Liu, N. Xu, M. Chen, and C. Xiao, “AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models”, Oct. 03, 2023. [Online]. Available: https://arxiv.org/abs/2310.04451
Liu, Y., Jia, Y., Geng, R., Jia, J. & Gong, N. Z.(2023). Formalizing and Benchmarking Prompt Injection Attacks and Defenses. arXiv.Liu, Y., Jia, Y., Geng, R., Jia, J., & Gong, N. Z. (2023). Formalizing and Benchmarking Prompt Injection Attacks and Defenses. In arXiv. https://arxiv.org/abs/2310.12815Liu, Y., Y. Jia, R. Geng, J. Jia, and N. Z. Gong. 2023. “Formalizing and Benchmarking Prompt Injection Attacks and Defenses”. In arXiv. Preprint, October 19. https://arxiv.org/abs/2310.12815.Liu, Y., et al. “Formalizing and Benchmarking Prompt Injection Attacks and Defenses”. arXiv, 19 Oct. 2023, https://arxiv.org/abs/2310.12815.Liu, Y., Jia, Y., Geng, R., Jia, J. & Gong, N. Z. Formalizing and Benchmarking Prompt Injection Attacks and Defenses. arXiv Preprint at https://arxiv.org/abs/2310.12815 (2023).Y. Liu, Y. Jia, R. Geng, J. Jia, and N. Z. Gong, “Formalizing and Benchmarking Prompt Injection Attacks and Defenses”, Oct. 19, 2023. [Online]. Available: https://arxiv.org/abs/2310.12815
Liu, Z.(2023). SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security. arXiv.Liu, Z. (2023). SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security. In arXiv. https://arxiv.org/abs/2312.15838Liu, Z. 2023. “SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security”. In arXiv. Preprint, December 26. https://arxiv.org/abs/2312.15838.Liu, Z. “SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security”. arXiv, 26 Dec. 2023, https://arxiv.org/abs/2312.15838.Liu, Z. SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security. arXiv Preprint at https://arxiv.org/abs/2312.15838 (2023).Z. Liu, “SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security”, Dec. 26, 2023. [Online]. Available: https://arxiv.org/abs/2312.15838
Lizka(2023). Beware safety-washing. EA Forum.Lizka. (2023, January 13). Beware safety-washing. EA Forum. Internet Archive (https://web.archive.org/web/20260519093736/https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing). https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washingLizka. 2023. “Beware Safety-washing”. EA Forum, January 13. Https://web.archive.org/web/20260519093736/https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing. Internet Archive. https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing.Lizka. “Beware Safety-washing”. EA Forum, 13 Jan. 2023, Internet Archive, https://web.archive.org/web/20260519093736/https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing, https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing.Lizka. Beware safety-washing. EA Forum https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing (2023).Lizka, “Beware safety-washing”, EA Forum. Accessed: May 19, 2026. [Online]. Available: https://forum.effectivealtruism.org/posts/f2qojPr8NaMPo2KJC/beware-safety-washing
Longpre, S. et al.(2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv.Longpre, S., Mahari, R., Chen, A., Obeng-Marnu, N., Sileo, D., Brannon, W., Muennighoff, N., Khazam, N., Kabbara, J., Perisetla, K., Wu, X., Shippole, E., Bollacker, K., Wu, T., Villa, L., Pentland, S., & Hooker, S. (2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. In arXiv. https://arxiv.org/abs/2310.16787Longpre, S., R. Mahari, A. Chen, et al. 2023. “The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI”. In arXiv. Preprint, October 25. https://arxiv.org/abs/2310.16787.Longpre, S., et al. “The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI”. arXiv, 25 Oct. 2023, https://arxiv.org/abs/2310.16787.Longpre, S. et al. The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv Preprint at https://arxiv.org/abs/2310.16787 (2023).S. Longpre et al., “The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI”, Oct. 25, 2023. [Online]. Available: https://arxiv.org/abs/2310.16787
Longpre, S. et al.(2024). Consent in Crisis: The Rapid Decline of the AI Data Commons. arXiv.org.Longpre, S., Mahari, R., Lee, A., Lund, C., Oderinwale, H., Brannon, W., Saxena, N., Obeng-Marnu, N., South, T., Hunter, C., Klyman, K., Klamm, C., Schoelkopf, H., Singh, N., Cherep, M., Anis, A., Dinh, A., Chitongo, C., Yin, D., … Pentland, S. (2024). Consent in Crisis: The Rapid Decline of the AI Data Commons. In arXiv.org. https://arxiv.org/abs/2407.14933Longpre, S., R. Mahari, A. Lee, et al. 2024. “Consent in Crisis: The Rapid Decline of the AI Data Commons”. In arXiv.org. Preprint, July 20. https://arxiv.org/abs/2407.14933.Longpre, S., et al. “Consent in Crisis: The Rapid Decline of the AI Data Commons”. arXiv.org, 20 July 2024, https://arxiv.org/abs/2407.14933.Longpre, S. et al. Consent in Crisis: The Rapid Decline of the AI Data Commons. arXiv.org Preprint at https://arxiv.org/abs/2407.14933 (2024).S. Longpre et al., “Consent in Crisis: The Rapid Decline of the AI Data Commons”, Jul. 20, 2024. [Online]. Available: https://arxiv.org/abs/2407.14933
Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J. & Ha, D.(2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv.Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., & Ha, D. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. In arXiv. https://arxiv.org/abs/2408.06292Lu, C., C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha. 2024. “The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery”. In arXiv. Preprint, August 12. https://arxiv.org/abs/2408.06292.Lu, C., et al. “The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery”. arXiv, 12 Aug. 2024, https://arxiv.org/abs/2408.06292.Lu, C. et al. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv Preprint at https://arxiv.org/abs/2408.06292 (2024).C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha, “The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery”, Aug. 12, 2024. [Online]. Available: https://arxiv.org/abs/2408.06292
Lu, K., Yu, B., Zhou, C. & Zhou, J.(2024). Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment. arXiv.Lu, K., Yu, B., Zhou, C., & Zhou, J. (2024). Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment. In arXiv. https://arxiv.org/abs/2401.12474Lu, K., B. Yu, C. Zhou, and J. Zhou. 2024. “Large Language Models Are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment”. In arXiv. Preprint, January 23. https://arxiv.org/abs/2401.12474.Lu, K., et al. “Large Language Models Are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment”. arXiv, 23 Jan. 2024, https://arxiv.org/abs/2401.12474.Lu, K., Yu, B., Zhou, C. & Zhou, J. Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment. arXiv Preprint at https://arxiv.org/abs/2401.12474 (2024).K. Lu, B. Yu, C. Zhou, and J. Zhou, “Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment”, Jan. 23, 2024. [Online]. Available: https://arxiv.org/abs/2401.12474
Luisa_Rodriguez(2020). What is the likelihood that civilizational collapse would directly lead to human extinction (within decades)?. EA Forum.Luisa_Rodriguez. (2020, December 24). What is the likelihood that civilizational collapse would directly lead to human extinction (within decades)?. EA Forum. https://forum.effectivealtruism.org/posts/GsjmufaebreiaivF7/what-is-the-likelihood-that-civilizational-collapse-wouldLuisa_Rodriguez. 2020. “What Is the Likelihood That Civilizational Collapse Would Directly Lead to Human Extinction (within Decades)?”. EA Forum, December 24. https://forum.effectivealtruism.org/posts/GsjmufaebreiaivF7/what-is-the-likelihood-that-civilizational-collapse-would.Luisa_Rodriguez. “What Is the Likelihood That Civilizational Collapse Would Directly Lead to Human Extinction (within Decades)?”. EA Forum, 24 Dec. 2020, https://forum.effectivealtruism.org/posts/GsjmufaebreiaivF7/what-is-the-likelihood-that-civilizational-collapse-would.Luisa_Rodriguez. What is the likelihood that civilizational collapse would directly lead to human extinction (within decades)?. EA Forum https://forum.effectivealtruism.org/posts/GsjmufaebreiaivF7/what-is-the-likelihood-that-civilizational-collapse-would (2020).Luisa_Rodriguez, “What is the likelihood that civilizational collapse would directly lead to human extinction (within decades)?”, EA Forum. [Online]. Available: https://forum.effectivealtruism.org/posts/GsjmufaebreiaivF7/what-is-the-likelihood-that-civilizational-collapse-would
Luke Emberson & David Owen(2025). The stock of computing power from NVIDIA chips is doubling every 10 months.Luke Emberson, & David Owen. (2025, February 13). The stock of computing power from NVIDIA chips is doubling every 10 months. https://epoch.ai/data-insights/nvidia-chip-productionLuke Emberson, and David Owen. 2025. “The Stock of Computing Power from NVIDIA Chips Is Doubling Every 10 Months”. February 13. https://epoch.ai/data-insights/nvidia-chip-production.Luke Emberson, and David Owen. The Stock of Computing Power from NVIDIA Chips Is Doubling Every 10 Months. 13 Feb. 2025, https://epoch.ai/data-insights/nvidia-chip-production.Luke Emberson & David Owen. The stock of computing power from NVIDIA chips is doubling every 10 months. https://epoch.ai/data-insights/nvidia-chip-production (2025).Luke Emberson and David Owen, “The stock of computing power from NVIDIA chips is doubling every 10 months”. [Online]. Available: https://epoch.ai/data-insights/nvidia-chip-production
Maas & Villalobos(2024). International AI Institutions: A Literature Review of Models, Examples, and Proposals.Maas & Villalobos. (2024). International AI Institutions: A Literature Review of Models, Examples, and Proposals. Internet Archive (https://web.archive.org/web/20260917211638/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4579773). https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4579773Maas & Villalobos. 2024. “International AI Institutions: A Literature Review of Models, Examples, and Proposals”. Https://web.archive.org/web/20260917211638/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4579773. Internet Archive. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4579773.Maas & Villalobos. International AI Institutions: A Literature Review of Models, Examples, and Proposals. 2024, Internet Archive, https://web.archive.org/web/20260917211638/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4579773, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4579773.Maas & Villalobos. International AI Institutions: A Literature Review of Models, Examples, and Proposals. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4579773 (2024).Maas & Villalobos, “International AI Institutions: A Literature Review of Models, Examples, and Proposals”. Accessed: Sep. 17, 2026. [Online]. Available: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4579773
Maas, M. M.(2019). How viable is international arms control for military artificial intelligence? Three lessons from nuclear weapons. Contemporary Security Policy.Maas, M. M. (2019). How viable is international arms control for military artificial intelligence? Three lessons from nuclear weapons. Contemporary Security Policy. https://doi.org/10.1080/13523260.2019.1576464Maas, M. M. 2019. “How Viable Is International Arms Control for Military Artificial Intelligence? Three Lessons from Nuclear Weapons”. Contemporary Security Policy, ahead of print, February 6. https://doi.org/10.1080/13523260.2019.1576464.Maas, M. M. “How Viable Is International Arms Control for Military Artificial Intelligence? Three Lessons from Nuclear Weapons”. Contemporary Security Policy, Feb. 2019, https://doi.org/10.1080/13523260.2019.1576464.Maas, M. M. How viable is international arms control for military artificial intelligence? Three lessons from nuclear weapons. Contemporary Security Policy https://doi.org/10.1080/13523260.2019.1576464 (2019) doi:10.1080/13523260.2019.1576464.M. M. Maas, “How viable is international arms control for military artificial intelligence? Three lessons from nuclear weapons”, Contemporary Security Policy, Feb. 2019, doi: 10.1080/13523260.2019.1576464.
MacDermott, M., Fox, J., Belardinelli, F. & Everitt, T.(2024). Measuring Goal-Directedness. arXiv.MacDermott, M., Fox, J., Belardinelli, F., & Everitt, T. (2024). Measuring Goal-Directedness. In arXiv. https://arxiv.org/abs/2412.04758MacDermott, M., J. Fox, F. Belardinelli, and T. Everitt. 2024. “Measuring Goal-Directedness”. In arXiv. Preprint, December 6. https://arxiv.org/abs/2412.04758.MacDermott, M., et al. “Measuring Goal-Directedness”. arXiv, 6 Dec. 2024, https://arxiv.org/abs/2412.04758.MacDermott, M., Fox, J., Belardinelli, F. & Everitt, T. Measuring Goal-Directedness. arXiv Preprint at https://arxiv.org/abs/2412.04758 (2024).M. MacDermott, J. Fox, F. Belardinelli, and T. Everitt, “Measuring Goal-Directedness”, Dec. 06, 2024. [Online]. Available: https://arxiv.org/abs/2412.04758
Machine Learning Street Talk(2024). It's Not About Scale, It's About Abstraction. YouTube.Machine Learning Street Talk. (2024). It's Not About Scale, It's About Abstraction [Video recording]. In YouTube. https://www.youtube.com/watch?v=s7_NlkBwdj8Machine Learning Street Talk. 2024. “It's Not About Scale, It's About Abstraction”. YouTube. https://www.youtube.com/watch?v=s7_NlkBwdj8.Machine Learning Street Talk. “It's Not About Scale, It's About Abstraction”. YouTube, 2024, https://www.youtube.com/watch?v=s7_NlkBwdj8.Machine Learning Street Talk. It's Not About Scale, It's About Abstraction. YouTube (2024).Machine Learning Street Talk, It's Not About Scale, It's About Abstraction, (2024). [Online Video]. Available: https://www.youtube.com/watch?v=s7_NlkBwdj8
Magdalena Wache(2023). Technical AI Safety Research Landscape [Slides]. LessWrong.Magdalena Wache. (2023, September 18). Technical AI Safety Research Landscape [Slides]. LessWrong. https://lesswrong.com/posts/x2n7mBLryDXuLwGhx/technical-ai-safety-research-landscape-slidesMagdalena Wache. 2023. “Technical AI Safety Research Landscape [Slides]”. LessWrong, September 18. https://lesswrong.com/posts/x2n7mBLryDXuLwGhx/technical-ai-safety-research-landscape-slides.Magdalena Wache. “Technical AI Safety Research Landscape [Slides]”. LessWrong, 18 Sept. 2023, https://lesswrong.com/posts/x2n7mBLryDXuLwGhx/technical-ai-safety-research-landscape-slides.Magdalena Wache. Technical AI Safety Research Landscape [Slides]. LessWrong https://lesswrong.com/posts/x2n7mBLryDXuLwGhx/technical-ai-safety-research-landscape-slides (2023).Magdalena Wache, “Technical AI Safety Research Landscape [Slides]”, LessWrong. [Online]. Available: https://lesswrong.com/posts/x2n7mBLryDXuLwGhx/technical-ai-safety-research-landscape-slides
Maksym Andriushchenko et al.(2024). AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. arXiv.Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Zico Kolter, Matt Fredrikson, Eric Winsor, Jerome Wynne, Yarin Gal, & Xander Davies. (2024). AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. In arXiv. https://arxiv.org/abs/2410.09024Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, et al. 2024. “AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents”. In arXiv. Preprint, October 11. https://arxiv.org/abs/2410.09024.Maksym Andriushchenko, et al. “AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents”. arXiv, 11 Oct. 2024, https://arxiv.org/abs/2410.09024.Maksym Andriushchenko et al. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. arXiv Preprint at https://arxiv.org/abs/2410.09024 (2024).Maksym Andriushchenko et al., “AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents”, Oct. 11, 2024. [Online]. Available: https://arxiv.org/abs/2410.09024
Mannheim(2023). Building a Culture of Safety for AI: Perspectives and Challenges.Mannheim. (2023). Building a Culture of Safety for AI: Perspectives and Challenges. Internet Archive (https://web.archive.org/web/20260802053912/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4491421). https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4491421Mannheim. 2023. “Building a Culture of Safety for AI: Perspectives and Challenges”. Https://web.archive.org/web/20260802053912/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4491421. Internet Archive. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4491421.Mannheim. Building a Culture of Safety for AI: Perspectives and Challenges. 2023, Internet Archive, https://web.archive.org/web/20260802053912/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4491421, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4491421.Mannheim. Building a Culture of Safety for AI: Perspectives and Challenges. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4491421 (2023).Mannheim, “Building a Culture of Safety for AI: Perspectives and Challenges”. Accessed: Aug. 02, 2026. [Online]. Available: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4491421
Mantas Mazeika et al.(2024). HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. arXiv.Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, & Dan Hendrycks. (2024). HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. In arXiv. https://arxiv.org/abs/2402.04249Mantas Mazeika, Long Phan, Xuwang Yin, et al. 2024. “HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal”. In arXiv. Preprint, February 6. https://arxiv.org/abs/2402.04249.Mantas Mazeika, et al. “HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal”. arXiv, 6 Feb. 2024, https://arxiv.org/abs/2402.04249.Mantas Mazeika et al. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. arXiv Preprint at https://arxiv.org/abs/2402.04249 (2024).Mantas Mazeika et al., “HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal”, Feb. 06, 2024. [Online]. Available: https://arxiv.org/abs/2402.04249
Marchal, N., Xu, R., Elasmar, R., Gabriel, I., Goldberg, B. & Isaac, W.(2024). Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data. arXiv.Marchal, N., Xu, R., Elasmar, R., Gabriel, I., Goldberg, B., & Isaac, W. (2024). Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data. In arXiv. https://arxiv.org/abs/2406.13843Marchal, N., R. Xu, R. Elasmar, I. Gabriel, B. Goldberg, and W. Isaac. 2024. “Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data”. In arXiv. Preprint, June 19. https://arxiv.org/abs/2406.13843.Marchal, N., et al. “Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data”. arXiv, 19 June 2024, https://arxiv.org/abs/2406.13843.Marchal, N. et al. Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data. arXiv Preprint at https://arxiv.org/abs/2406.13843 (2024).N. Marchal, R. Xu, R. Elasmar, I. Gabriel, B. Goldberg, and W. Isaac, “Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data”, Jun. 19, 2024. [Online]. Available: https://arxiv.org/abs/2406.13843
Marcucci, S., Alarcon, N. G., Verhulst, S. G. & Wullhorst, E.(2023). Mapping and Comparing Data Governance Frameworks: A benchmarking exercise to inform global data governance deliberations. arXiv.Marcucci, S., Alarcon, N. G., Verhulst, S. G., & Wullhorst, E. (2023). Mapping and Comparing Data Governance Frameworks: A benchmarking exercise to inform global data governance deliberations. In arXiv. https://arxiv.org/abs/2302.13731Marcucci, S., N. G. Alarcon, S. G. Verhulst, and E. Wullhorst. 2023. “Mapping and Comparing Data Governance Frameworks: A Benchmarking Exercise to Inform Global Data Governance Deliberations”. In arXiv. Preprint, February 27. https://arxiv.org/abs/2302.13731.Marcucci, S., et al. “Mapping and Comparing Data Governance Frameworks: A Benchmarking Exercise to Inform Global Data Governance Deliberations”. arXiv, 27 Feb. 2023, https://arxiv.org/abs/2302.13731.Marcucci, S., Alarcon, N. G., Verhulst, S. G. & Wullhorst, E. Mapping and Comparing Data Governance Frameworks: A benchmarking exercise to inform global data governance deliberations. arXiv Preprint at https://arxiv.org/abs/2302.13731 (2023).S. Marcucci, N. G. Alarcon, S. G. Verhulst, and E. Wullhorst, “Mapping and Comparing Data Governance Frameworks: A benchmarking exercise to inform global data governance deliberations”, Feb. 27, 2023. [Online]. Available: https://arxiv.org/abs/2302.13731
Marcus(2025). Game over. AGI is not imminent, and LLMs are not the royal road to getting there.Marcus. (2025). Game over. AGI is not imminent, and LLMs are not the royal road to getting there. https://garymarcus.substack.com/p/the-last-few-months-have-been-devastatingMarcus. 2025. Game Over. AGI Is Not Imminent, and LLMs Are Not the Royal Road to Getting There. Edition. https://garymarcus.substack.com/p/the-last-few-months-have-been-devastating.Marcus. Game Over. AGI Is Not Imminent, and LLMs Are Not the Royal Road to Getting There. 2025, https://garymarcus.substack.com/p/the-last-few-months-have-been-devastating.Marcus. Game over. AGI is not imminent, and LLMs are not the royal road to getting there. https://garymarcus.substack.com/p/the-last-few-months-have-been-devastating (2025).Marcus, “Game over. AGI is not imminent, and LLMs are not the royal road to getting there.”. [Online]. Available: https://garymarcus.substack.com/p/the-last-few-months-have-been-devastating
Marius Hobbhahn(2024). The Evals Gap. AI Alignment Forum.Marius Hobbhahn. (2024, November 11). The Evals Gap. AI Alignment Forum. https://alignmentforum.org/posts/gJJEjJpKiddoYGZKk/the-evals-gapMarius Hobbhahn. 2024. “The Evals Gap”. AI Alignment Forum, November 11. https://alignmentforum.org/posts/gJJEjJpKiddoYGZKk/the-evals-gap.Marius Hobbhahn. “The Evals Gap”. AI Alignment Forum, 11 Nov. 2024, https://alignmentforum.org/posts/gJJEjJpKiddoYGZKk/the-evals-gap.Marius Hobbhahn. The Evals Gap. AI Alignment Forum https://alignmentforum.org/posts/gJJEjJpKiddoYGZKk/the-evals-gap (2024).Marius Hobbhahn, “The Evals Gap”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/gJJEjJpKiddoYGZKk/the-evals-gap
Marius Hobbhahn, Alex Meinke, Bronson Schoen, rusheb, Jérémy Scheurer & Mikita Balesni(2024). Frontier Models are Capable of In-context Scheming. LessWrong.Marius Hobbhahn, Alex Meinke, Bronson Schoen, rusheb, Jérémy Scheurer, & Mikita Balesni. (2024, December 5). Frontier Models are Capable of In-context Scheming. LessWrong. https://lesswrong.com/posts/8gy7c8GAPkuu6wTiX/frontier-models-are-capable-of-in-context-schemingMarius Hobbhahn, Alex Meinke, Bronson Schoen, rusheb, Jérémy Scheurer, and Mikita Balesni. 2024. “Frontier Models Are Capable of In-context Scheming”. LessWrong, December 5. https://lesswrong.com/posts/8gy7c8GAPkuu6wTiX/frontier-models-are-capable-of-in-context-scheming.Marius Hobbhahn, et al. “Frontier Models Are Capable of In-context Scheming”. LessWrong, 5 Dec. 2024, https://lesswrong.com/posts/8gy7c8GAPkuu6wTiX/frontier-models-are-capable-of-in-context-scheming.Marius Hobbhahn et al. Frontier Models are Capable of In-context Scheming. LessWrong https://lesswrong.com/posts/8gy7c8GAPkuu6wTiX/frontier-models-are-capable-of-in-context-scheming (2024).Marius Hobbhahn, Alex Meinke, Bronson Schoen, rusheb, Jérémy Scheurer, and Mikita Balesni, “Frontier Models are Capable of In-context Scheming”, LessWrong. [Online]. Available: https://lesswrong.com/posts/8gy7c8GAPkuu6wTiX/frontier-models-are-capable-of-in-context-scheming
Marks, S. & Tegmark, M.(2023). The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. arXiv.Marks, S., & Tegmark, M. (2023). The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. In arXiv. https://arxiv.org/abs/2310.06824Marks, S., and M. Tegmark. 2023. “The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets”. In arXiv. Preprint, October 10. https://arxiv.org/abs/2310.06824.Marks, S., and M. Tegmark. “The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets”. arXiv, 10 Oct. 2023, https://arxiv.org/abs/2310.06824.Marks, S. & Tegmark, M. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. arXiv Preprint at https://arxiv.org/abs/2310.06824 (2023).S. Marks and M. Tegmark, “The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets”, Oct. 10, 2023. [Online]. Available: https://arxiv.org/abs/2310.06824
Marks, S. et al.(2025). Auditing language models for hidden objectives. arXiv.Marks, S., Treutlein, J., Bricken, T., Lindsey, J., Marcus, J., Mishra-Sharma, S., Ziegler, D., Ameisen, E., Batson, J., Belonax, T., Bowman, S. R., Carter, S., Chen, B., Cunningham, H., Denison, C., Dietz, F., Golechha, S., Khan, A., Kirchner, J., … Hubinger, E. (2025). Auditing language models for hidden objectives. In arXiv. https://arxiv.org/abs/2503.10965Marks, S., J. Treutlein, T. Bricken, et al. 2025. “Auditing Language Models for Hidden Objectives”. In arXiv. Preprint, March 14. https://arxiv.org/abs/2503.10965.Marks, S., et al. “Auditing Language Models for Hidden Objectives”. arXiv, 14 Mar. 2025, https://arxiv.org/abs/2503.10965.Marks, S. et al. Auditing language models for hidden objectives. arXiv Preprint at https://arxiv.org/abs/2503.10965 (2025).S. Marks et al., “Auditing language models for hidden objectives”, Mar. 14, 2025. [Online]. Available: https://arxiv.org/abs/2503.10965
Mary Phuong et al.(2024). Evaluating Frontier Models for Dangerous Capabilities. arXiv.Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, … Toby Shevlane. (2024). Evaluating Frontier Models for Dangerous Capabilities. In arXiv. https://arxiv.org/abs/2403.13793Mary Phuong, Matthew Aitchison, Elliot Catt, et al. 2024. “Evaluating Frontier Models for Dangerous Capabilities”. In arXiv. Preprint, March 20. https://arxiv.org/abs/2403.13793.Mary Phuong, et al. “Evaluating Frontier Models for Dangerous Capabilities”. arXiv, 20 Mar. 2024, https://arxiv.org/abs/2403.13793.Mary Phuong et al. Evaluating Frontier Models for Dangerous Capabilities. arXiv Preprint at https://arxiv.org/abs/2403.13793 (2024).Mary Phuong et al., “Evaluating Frontier Models for Dangerous Capabilities”, Mar. 20, 2024. [Online]. Available: https://arxiv.org/abs/2403.13793
Masi(2024). Masi_GPU_Export_Controls_Updated.pdf. Google Docs.Masi. (2024). Masi_GPU_Export_Controls_Updated.pdf. Google Docs. https://drive.google.com/file/d/1SSRpckR_IhIPqkO13oRQ-8mmPOoDKLqd/viewMasi. 2024. “Masi_GPU_Export_Controls_Updated.pdf”. Google Docs. https://drive.google.com/file/d/1SSRpckR_IhIPqkO13oRQ-8mmPOoDKLqd/view.Masi. “Masi_GPU_Export_Controls_Updated.pdf”. Google Docs, 2024, https://drive.google.com/file/d/1SSRpckR_IhIPqkO13oRQ-8mmPOoDKLqd/view.Masi. Masi_GPU_Export_Controls_Updated.pdf. Google Docs https://drive.google.com/file/d/1SSRpckR_IhIPqkO13oRQ-8mmPOoDKLqd/view (2024).Masi, “Masi_GPU_Export_Controls_Updated.pdf”, Google Docs. [Online]. Available: https://drive.google.com/file/d/1SSRpckR_IhIPqkO13oRQ-8mmPOoDKLqd/view
Maslej, N. et al.(2025). Artificial Intelligence Index Report 2025. arXiv.Maslej, N., Fattorini, L., Perrault, R., Gil, Y., Parli, V., Kariuki, N., Capstick, E., Reuel, A., Brynjolfsson, E., Etchemendy, J., Ligett, K., Lyons, T., Manyika, J., Niebles, J. C., Shoham, Y., Wald, R., Walsh, T., Hamrah, A., Santarlasci, L., … Oak, S. (2025). Artificial Intelligence Index Report 2025. In arXiv. https://arxiv.org/abs/2504.07139Maslej, N., L. Fattorini, R. Perrault, et al. 2025. “Artificial Intelligence Index Report 2025”. In arXiv. Preprint, April 8. https://arxiv.org/abs/2504.07139.Maslej, N., et al. “Artificial Intelligence Index Report 2025”. arXiv, 8 Apr. 2025, https://arxiv.org/abs/2504.07139.Maslej, N. et al. Artificial Intelligence Index Report 2025. arXiv Preprint at https://arxiv.org/abs/2504.07139 (2025).N. Maslej et al., “Artificial Intelligence Index Report 2025”, Apr. 08, 2025. [Online]. Available: https://arxiv.org/abs/2504.07139
Matteo Pistillo(2025). Towards Frontier Safety Policies Plus. arXiv.Matteo Pistillo. (2025). Towards Frontier Safety Policies Plus. In arXiv. https://arxiv.org/abs/2501.16500Matteo Pistillo. 2025. “Towards Frontier Safety Policies Plus”. In arXiv. Preprint, January 27. https://arxiv.org/abs/2501.16500.Matteo Pistillo. “Towards Frontier Safety Policies Plus”. arXiv, 27 Jan. 2025, https://arxiv.org/abs/2501.16500.Matteo Pistillo. Towards Frontier Safety Policies Plus. arXiv Preprint at https://arxiv.org/abs/2501.16500 (2025).Matteo Pistillo, “Towards Frontier Safety Policies Plus”, Jan. 27, 2025. [Online]. Available: https://arxiv.org/abs/2501.16500
Matthew Barnett(2025). AGI could drive wages below subsistence level.Matthew Barnett. (2025, January 24). AGI could drive wages below subsistence level. https://epoch.ai/gradient-updates/agi-could-drive-wages-below-subsistence-levelMatthew Barnett. 2025. “AGI Could Drive Wages Below Subsistence Level”. January 24. https://epoch.ai/gradient-updates/agi-could-drive-wages-below-subsistence-level.Matthew Barnett. AGI Could Drive Wages Below Subsistence Level. 24 Jan. 2025, https://epoch.ai/gradient-updates/agi-could-drive-wages-below-subsistence-level.Matthew Barnett. AGI could drive wages below subsistence level. https://epoch.ai/gradient-updates/agi-could-drive-wages-below-subsistence-level (2025).Matthew Barnett, “AGI could drive wages below subsistence level”. [Online]. Available: https://epoch.ai/gradient-updates/agi-could-drive-wages-below-subsistence-level
Matthew Barnett(2025). The economic consequences of automating remote work.Matthew Barnett. (2025, January 10). The economic consequences of automating remote work. https://epoch.ai/gradient-updates/consequences-of-automating-remote-workMatthew Barnett. 2025. “The Economic Consequences of Automating Remote Work”. January 10. https://epoch.ai/gradient-updates/consequences-of-automating-remote-work.Matthew Barnett. The Economic Consequences of Automating Remote Work. 10 Jan. 2025, https://epoch.ai/gradient-updates/consequences-of-automating-remote-work.Matthew Barnett. The economic consequences of automating remote work. https://epoch.ai/gradient-updates/consequences-of-automating-remote-work (2025).Matthew Barnett, “The economic consequences of automating remote work”. [Online]. Available: https://epoch.ai/gradient-updates/consequences-of-automating-remote-work
Max Tegmark(2024). The Hopium Wars: the AGI Entente Delusion. LessWrong.Max Tegmark. (2024, October 13). The Hopium Wars: the AGI Entente Delusion. LessWrong. https://lesswrong.com/posts/oJQnRDbgSS8i6DwNu/the-hopium-wars-the-agi-entente-delusionMax Tegmark. 2024. “The Hopium Wars: The AGI Entente Delusion”. LessWrong, October 13. https://lesswrong.com/posts/oJQnRDbgSS8i6DwNu/the-hopium-wars-the-agi-entente-delusion.Max Tegmark. “The Hopium Wars: The AGI Entente Delusion”. LessWrong, 13 Oct. 2024, https://lesswrong.com/posts/oJQnRDbgSS8i6DwNu/the-hopium-wars-the-agi-entente-delusion.Max Tegmark. The Hopium Wars: the AGI Entente Delusion. LessWrong https://lesswrong.com/posts/oJQnRDbgSS8i6DwNu/the-hopium-wars-the-agi-entente-delusion (2024).Max Tegmark, “The Hopium Wars: the AGI Entente Delusion”, LessWrong. [Online]. Available: https://lesswrong.com/posts/oJQnRDbgSS8i6DwNu/the-hopium-wars-the-agi-entente-delusion
Maxime Riché, Harrison G, JaimeRV & Edoardo Pona(2024). Thinking About Propensity Evaluations. AI Alignment Forum.Maxime Riché, Harrison G, JaimeRV, & Edoardo Pona. (2024, August 19). Thinking About Propensity Evaluations. AI Alignment Forum. https://alignmentforum.org/posts/sWf8wj64AdDfMeTvf/thinking-about-propensity-evaluationsMaxime Riché, Harrison G, JaimeRV, and Edoardo Pona. 2024. “Thinking About Propensity Evaluations”. AI Alignment Forum, August 19. https://alignmentforum.org/posts/sWf8wj64AdDfMeTvf/thinking-about-propensity-evaluations.Maxime Riché, et al. “Thinking About Propensity Evaluations”. AI Alignment Forum, 19 Aug. 2024, https://alignmentforum.org/posts/sWf8wj64AdDfMeTvf/thinking-about-propensity-evaluations.Maxime Riché, Harrison G, JaimeRV & Edoardo Pona. Thinking About Propensity Evaluations. AI Alignment Forum https://alignmentforum.org/posts/sWf8wj64AdDfMeTvf/thinking-about-propensity-evaluations (2024).Maxime Riché, Harrison G, JaimeRV, and Edoardo Pona, “Thinking About Propensity Evaluations”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/sWf8wj64AdDfMeTvf/thinking-about-propensity-evaluations
Maynez, J., Narayan, S., Bohnet, B. & McDonald, R.(2020). On Faithfulness and Factuality in Abstractive Summarization. arXiv.Maynez, J., Narayan, S., Bohnet, B., & McDonald, R. (2020). On Faithfulness and Factuality in Abstractive Summarization. In arXiv. https://arxiv.org/abs/2005.00661Maynez, J., S. Narayan, B. Bohnet, and R. McDonald. 2020. “On Faithfulness and Factuality in Abstractive Summarization”. In arXiv. Preprint, May 2. https://arxiv.org/abs/2005.00661.Maynez, J., et al. “On Faithfulness and Factuality in Abstractive Summarization”. arXiv, 2 May 2020, https://arxiv.org/abs/2005.00661.Maynez, J., Narayan, S., Bohnet, B. & McDonald, R. On Faithfulness and Factuality in Abstractive Summarization. arXiv Preprint at https://arxiv.org/abs/2005.00661 (2020).J. Maynez, S. Narayan, B. Bohnet, and R. McDonald, “On Faithfulness and Factuality in Abstractive Summarization”, May 02, 2020. [Online]. Available: https://arxiv.org/abs/2005.00661
McAleese, N., Pokorny, R. M., Uribe, J. F. C., Nitishinskaya, E., Trebacz, M. & Leike, J.(2024). LLM Critics Help Catch LLM Bugs. arXiv.McAleese, N., Pokorny, R. M., Uribe, J. F. C., Nitishinskaya, E., Trebacz, M., & Leike, J. (2024). LLM Critics Help Catch LLM Bugs. In arXiv. https://arxiv.org/abs/2407.00215McAleese, N., R. M. Pokorny, J. F. C. Uribe, E. Nitishinskaya, M. Trebacz, and J. Leike. 2024. “LLM Critics Help Catch LLM Bugs”. In arXiv. Preprint, June 28. https://arxiv.org/abs/2407.00215.McAleese, N., et al. “LLM Critics Help Catch LLM Bugs”. arXiv, 28 June 2024, https://arxiv.org/abs/2407.00215.McAleese, N. et al. LLM Critics Help Catch LLM Bugs. arXiv Preprint at https://arxiv.org/abs/2407.00215 (2024).N. McAleese, R. M. Pokorny, J. F. C. Uribe, E. Nitishinskaya, M. Trebacz, and J. Leike, “LLM Critics Help Catch LLM Bugs”, Jun. 28, 2024. [Online]. Available: https://arxiv.org/abs/2407.00215
McCoy, R. T., Min, J. & Linzen, T.(2019). BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance. arXiv.McCoy, R. T., Min, J., & Linzen, T. (2019). BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance. In arXiv. https://arxiv.org/abs/1911.02969McCoy, R. T., J. Min, and T. Linzen. 2019. “BERTs of a Feather Do Not Generalize Together: Large Variability in Generalization Across Models with Similar Test Set Performance”. In arXiv. Preprint, November 7. https://arxiv.org/abs/1911.02969.McCoy, R. T., et al. “BERTs of a Feather Do Not Generalize Together: Large Variability in Generalization Across Models with Similar Test Set Performance”. arXiv, 7 Nov. 2019, https://arxiv.org/abs/1911.02969.McCoy, R. T., Min, J. & Linzen, T. BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance. arXiv Preprint at https://arxiv.org/abs/1911.02969 (2019).R. T. McCoy, J. Min, and T. Linzen, “BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance”, Nov. 07, 2019. [Online]. Available: https://arxiv.org/abs/1911.02969
McGrath, T. et al.(2021). Acquisition of Chess Knowledge in AlphaZero. arXiv.McGrath, T., Kapishnikov, A., Tomašev, N., Pearce, A., Hassabis, D., Kim, B., Paquet, U., & Kramnik, V. (2021). Acquisition of Chess Knowledge in AlphaZero. In arXiv. https://doi.org/10.1073/pnas.2206625119McGrath, T., A. Kapishnikov, N. Tomašev, et al. 2021. “Acquisition of Chess Knowledge in AlphaZero”. In arXiv. Preprint, November 17. https://doi.org/10.1073/pnas.2206625119.McGrath, T., et al. “Acquisition of Chess Knowledge in AlphaZero”. arXiv, 17 Nov. 2021, https://doi.org/10.1073/pnas.2206625119.McGrath, T. et al. Acquisition of Chess Knowledge in AlphaZero. arXiv Preprint at https://doi.org/10.1073/pnas.2206625119 (2021).T. McGrath et al., “Acquisition of Chess Knowledge in AlphaZero”, Nov. 17, 2021. doi: 10.1073/pnas.2206625119.
McGregor, S.(2020). Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database. arXiv.McGregor, S. (2020). Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database. In arXiv. https://arxiv.org/abs/2011.08512McGregor, S. 2020. “Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database”. In arXiv. Preprint, November 17. https://arxiv.org/abs/2011.08512.McGregor, S. “Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database”. arXiv, 17 Nov. 2020, https://arxiv.org/abs/2011.08512.McGregor, S. Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database. arXiv Preprint at https://arxiv.org/abs/2011.08512 (2020).S. McGregor, “Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database”, Nov. 17, 2020. [Online]. Available: https://arxiv.org/abs/2011.08512
McKernon et al.(2024). AI Model Registries: A Foundational Tool for AI Governance. Convergence Analysis.McKernon et al. (2024). AI Model Registries: A Foundational Tool for AI Governance. Convergence Analysis. https://convergenceanalysis.org/research/ai-model-registries-a-foundational-tool-for-ai-governanceMcKernon et al. 2024. “AI Model Registries: A Foundational Tool for AI Governance”. Convergence Analysis. https://convergenceanalysis.org/research/ai-model-registries-a-foundational-tool-for-ai-governance.McKernon et al. “AI Model Registries: A Foundational Tool for AI Governance”. Convergence Analysis, 2024, https://convergenceanalysis.org/research/ai-model-registries-a-foundational-tool-for-ai-governance.McKernon et al. AI Model Registries: A Foundational Tool for AI Governance. Convergence Analysis https://convergenceanalysis.org/research/ai-model-registries-a-foundational-tool-for-ai-governance (2024).McKernon et al., “AI Model Registries: A Foundational Tool for AI Governance”, Convergence Analysis. [Online]. Available: https://convergenceanalysis.org/research/ai-model-registries-a-foundational-tool-for-ai-governance
McKernon, E., Glasser, G., Cheng, D. & Hadfield, G.(2024). AI Model Registries: A Foundational Tool for AI Governance. arXiv.McKernon, E., Glasser, G., Cheng, D., & Hadfield, G. (2024). AI Model Registries: A Foundational Tool for AI Governance. In arXiv. https://arxiv.org/abs/2410.09645McKernon, E., G. Glasser, D. Cheng, and G. Hadfield. 2024. “AI Model Registries: A Foundational Tool for AI Governance”. In arXiv. Preprint, October 12. https://arxiv.org/abs/2410.09645.McKernon, E., et al. “AI Model Registries: A Foundational Tool for AI Governance”. arXiv, 12 Oct. 2024, https://arxiv.org/abs/2410.09645.McKernon, E., Glasser, G., Cheng, D. & Hadfield, G. AI Model Registries: A Foundational Tool for AI Governance. arXiv Preprint at https://arxiv.org/abs/2410.09645 (2024).E. McKernon, G. Glasser, D. Cheng, and G. Hadfield, “AI Model Registries: A Foundational Tool for AI Governance”, Oct. 12, 2024. [Online]. Available: https://arxiv.org/abs/2410.09645
Megan Kinniment(2025). Why it’s good for AI reasoning to be legible and faithful.Megan Kinniment. (2025, March 11). Why it’s good for AI reasoning to be legible and faithful. https://metr.org/blog/2025-03-11-good-for-ai-to-reason-legibly-and-faithfullyMegan Kinniment. 2025. “Why It’s Good for AI Reasoning to Be Legible and Faithful”. March 11. https://metr.org/blog/2025-03-11-good-for-ai-to-reason-legibly-and-faithfully.Megan Kinniment. Why It’s Good for AI Reasoning to Be Legible and Faithful. 11 Mar. 2025, https://metr.org/blog/2025-03-11-good-for-ai-to-reason-legibly-and-faithfully.Megan Kinniment. Why it’s good for AI reasoning to be legible and faithful. https://metr.org/blog/2025-03-11-good-for-ai-to-reason-legibly-and-faithfully (2025).Megan Kinniment, “Why it’s good for AI reasoning to be legible and faithful”. [Online]. Available: https://metr.org/blog/2025-03-11-good-for-ai-to-reason-legibly-and-faithfully
Meinke, A., Schoen, B., Scheurer, J., Balesni, M., Shah, R. & Hobbhahn, M.(2024). Frontier Models are Capable of In-context Scheming. arXiv.Meinke, A., Schoen, B., Scheurer, J., Balesni, M., Shah, R., & Hobbhahn, M. (2024). Frontier Models are Capable of In-context Scheming. In arXiv. https://arxiv.org/abs/2412.04984Meinke, A., B. Schoen, J. Scheurer, M. Balesni, R. Shah, and M. Hobbhahn. 2024. “Frontier Models Are Capable of In-context Scheming”. In arXiv. Preprint, December 6. https://arxiv.org/abs/2412.04984.Meinke, A., et al. “Frontier Models Are Capable of In-context Scheming”. arXiv, 6 Dec. 2024, https://arxiv.org/abs/2412.04984.Meinke, A. et al. Frontier Models are Capable of In-context Scheming. arXiv Preprint at https://arxiv.org/abs/2412.04984 (2024).A. Meinke, B. Schoen, J. Scheurer, M. Balesni, R. Shah, and M. Hobbhahn, “Frontier Models are Capable of In-context Scheming”, Dec. 06, 2024. [Online]. Available: https://arxiv.org/abs/2412.04984
Meng, K., Bau, D., Andonian, A. & Belinkov, Y.(2022). Locating and Editing Factual Associations in GPT. arXiv.Meng, K., Bau, D., Andonian, A., & Belinkov, Y. (2022). Locating and Editing Factual Associations in GPT. In arXiv. https://arxiv.org/abs/2202.05262Meng, K., D. Bau, A. Andonian, and Y. Belinkov. 2022. “Locating and Editing Factual Associations in GPT”. In arXiv. Preprint, February 10. https://arxiv.org/abs/2202.05262.Meng, K., et al. “Locating and Editing Factual Associations in GPT”. arXiv, 10 Feb. 2022, https://arxiv.org/abs/2202.05262.Meng, K., Bau, D., Andonian, A. & Belinkov, Y. Locating and Editing Factual Associations in GPT. arXiv Preprint at https://arxiv.org/abs/2202.05262 (2022).K. Meng, D. Bau, A. Andonian, and Y. Belinkov, “Locating and Editing Factual Associations in GPT”, Feb. 10, 2022. [Online]. Available: https://arxiv.org/abs/2202.05262
Merchant, A., Batzner, S., Schoenholz, S. S., Aykol, M., Cheon, G. & Cubuk, E. D.(2023). Scaling deep learning for materials discovery. Nature.Merchant, A., Batzner, S., Schoenholz, S. S., Aykol, M., Cheon, G., & Cubuk, E. D. (2023). Scaling deep learning for materials discovery. Nature, 624, 80–85. https://doi.org/10.1038/s41586-023-06735-9Merchant, A., S. Batzner, S. S. Schoenholz, M. Aykol, G. Cheon, and E. D. Cubuk. 2023. “Scaling Deep Learning for Materials Discovery”. Nature 624 (November): 80–85. https://doi.org/10.1038/s41586-023-06735-9.Merchant, A., et al. “Scaling Deep Learning for Materials Discovery”. Nature, vol. 624, Nov. 2023, pp. 80–85, https://doi.org/10.1038/s41586-023-06735-9.Merchant, A. et al. Scaling deep learning for materials discovery. Nature 624, 80–85 (2023).A. Merchant, S. Batzner, S. S. Schoenholz, M. Aykol, G. Cheon, and E. D. Cubuk, “Scaling deep learning for materials discovery”, Nature, vol. 624, pp. 80–85, Nov. 2023, doi: 10.1038/s41586-023-06735-9.
Meredith Ringel Morris et al.(2023). Levels of AGI for Operationalizing Progress on the Path to AGI. arXiv.Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clement Farabet, & Shane Legg. (2023). Levels of AGI for Operationalizing Progress on the Path to AGI. In arXiv. https://arxiv.org/abs/2311.02462Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, et al. 2023. “Levels of AGI for Operationalizing Progress on the Path to AGI”. In arXiv. Preprint, November 4. https://arxiv.org/abs/2311.02462.Meredith Ringel Morris, et al. “Levels of AGI for Operationalizing Progress on the Path to AGI”. arXiv, 4 Nov. 2023, https://arxiv.org/abs/2311.02462.Meredith Ringel Morris et al. Levels of AGI for Operationalizing Progress on the Path to AGI. arXiv Preprint at https://arxiv.org/abs/2311.02462 (2023).Meredith Ringel Morris et al., “Levels of AGI for Operationalizing Progress on the Path to AGI”, Nov. 04, 2023. [Online]. Available: https://arxiv.org/abs/2311.02462
Meta Fundamental AI Research Diplomacy Team (FAIR)† et al.(2022). Human-level play in the game of Diplomacy by combining language models with strategic reasoning. Science.Meta Fundamental AI Research Diplomacy Team (FAIR)†, Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., Jacob, A. P., Komeili, M., Konath, K., Kwon, M., Lerer, A., Lewis, M., Miller, A. H., Mitts, S., Renduchintala, A., … Zijlstra, M. (2022). Human-level play in the game of Diplomacy by combining language models with strategic reasoning. Science. https://doi.org/10.1126/science.ade9097Meta Fundamental AI Research Diplomacy Team (FAIR)†, A. Bakhtin, N. Brown, et al. 2022. “Human-level Play in the Game of Diplomacy by Combining Language Models with Strategic Reasoning”. Science, ahead of print, December 9. https://doi.org/10.1126/science.ade9097.Meta Fundamental AI Research Diplomacy Team (FAIR)†, et al. “Human-level Play in the Game of Diplomacy by Combining Language Models with Strategic Reasoning”. Science, Dec. 2022, https://doi.org/10.1126/science.ade9097.Meta Fundamental AI Research Diplomacy Team (FAIR)† et al. Human-level play in the game of Diplomacy by combining language models with strategic reasoning. Science https://doi.org/10.1126/science.ade9097 (2022) doi:10.1126/science.ade9097.Meta Fundamental AI Research Diplomacy Team (FAIR)† et al., “Human-level play in the game of Diplomacy by combining language models with strategic reasoning”, Science, Dec. 2022, doi: 10.1126/science.ade9097.
Metaculus(2025). AI-authored paper published at NeurIPS, ICML, or ICLR before 2028?.Metaculus. (2025). AI-authored paper published at NeurIPS, ICML, or ICLR before 2028?. https://metaculus.com/questions/38403/ai-authorered-paper-by-2028Metaculus. 2025. “AI-authored Paper Published at NeurIPS, ICML, or ICLR Before 2028?”. https://metaculus.com/questions/38403/ai-authorered-paper-by-2028.Metaculus. AI-authored Paper Published at NeurIPS, ICML, or ICLR Before 2028?. 2025, https://metaculus.com/questions/38403/ai-authorered-paper-by-2028.Metaculus. AI-authored paper published at NeurIPS, ICML, or ICLR before 2028?. https://metaculus.com/questions/38403/ai-authorered-paper-by-2028 (2025).Metaculus, “AI-authored paper published at NeurIPS, ICML, or ICLR before 2028?”. [Online]. Available: https://metaculus.com/questions/38403/ai-authorered-paper-by-2028
Metaculus(2025). US and China reach an agreement to limit frontier AI development before 2029?.Metaculus. (2025). US and China reach an agreement to limit frontier AI development before 2029?. https://metaculus.com/questions/38418/us-and-china-reach-an-agreement-to-limit-frontier-ai-development-before-2029Metaculus. 2025. “US and China Reach an Agreement to Limit Frontier AI Development Before 2029?”. https://metaculus.com/questions/38418/us-and-china-reach-an-agreement-to-limit-frontier-ai-development-before-2029.Metaculus. US and China Reach an Agreement to Limit Frontier AI Development Before 2029?. 2025, https://metaculus.com/questions/38418/us-and-china-reach-an-agreement-to-limit-frontier-ai-development-before-2029.Metaculus. US and China reach an agreement to limit frontier AI development before 2029?. https://metaculus.com/questions/38418/us-and-china-reach-an-agreement-to-limit-frontier-ai-development-before-2029 (2025).Metaculus, “US and China reach an agreement to limit frontier AI development before 2029?”. [Online]. Available: https://metaculus.com/questions/38418/us-and-china-reach-an-agreement-to-limit-frontier-ai-development-before-2029
METR(2023). The TaskRabbit example.METR. (2023). The TaskRabbit example. METR. https://evals.alignment.org/taskrabbit.pdfMETR. 2023. The TaskRabbit Example. METR. https://evals.alignment.org/taskrabbit.pdf.METR. The TaskRabbit Example. METR, 2023, https://evals.alignment.org/taskrabbit.pdf.METR. The TaskRabbit Example. https://evals.alignment.org/taskrabbit.pdf (2023).METR, “The TaskRabbit example”, METR, 2023. [Online]. Available: https://evals.alignment.org/taskrabbit.pdf
METR(2024). Language Model Pilot Report.METR. (2024). Language Model Pilot Report. https://metr.org/language-model-pilot-reportMETR. 2024. “Language Model Pilot Report”. https://metr.org/language-model-pilot-report.METR. Language Model Pilot Report. 2024, https://metr.org/language-model-pilot-report.METR. Language Model Pilot Report. https://metr.org/language-model-pilot-report (2024).METR, “Language Model Pilot Report”. [Online]. Available: https://metr.org/language-model-pilot-report
METR(2025). Common Elements of Frontier AI Safety Policies (December 2025 Update). METR.METR. (2025, December). Common Elements of Frontier AI Safety Policies (December 2025 Update). METR. https://metr.org/blog/2025-03-26-common-elements-of-frontier-ai-safety-policiesMETR. 2025. “Common Elements of Frontier AI Safety Policies (December 2025 Update)”. METR, December. https://metr.org/blog/2025-03-26-common-elements-of-frontier-ai-safety-policies.METR. “Common Elements of Frontier AI Safety Policies (December 2025 Update)”. METR, Dec. 2025, https://metr.org/blog/2025-03-26-common-elements-of-frontier-ai-safety-policies.METR. Common Elements of Frontier AI Safety Policies (December 2025 Update). METR https://metr.org/blog/2025-03-26-common-elements-of-frontier-ai-safety-policies (2025).METR, “Common Elements of Frontier AI Safety Policies (December 2025 Update)”, METR. [Online]. Available: https://metr.org/blog/2025-03-26-common-elements-of-frontier-ai-safety-policies
Michaël Trazzi(2024). Owain Evans on Situational Awareness.Michaël Trazzi. (2024, August 23). Owain Evans on Situational Awareness. https://theinsideview.ai/owainMichaël Trazzi. 2024. “Owain Evans on Situational Awareness”. August 23. https://theinsideview.ai/owain.Michaël Trazzi. Owain Evans on Situational Awareness. 23 Aug. 2024, https://theinsideview.ai/owain.Michaël Trazzi. Owain Evans on Situational Awareness. https://theinsideview.ai/owain (2024).Michaël Trazzi, “Owain Evans on Situational Awareness”. [Online]. Available: https://theinsideview.ai/owain
Michael, J. et al.(2023). Debate Helps Supervise Unreliable Experts. arXiv.Michael, J., Mahdi, S., Rein, D., Petty, J., Dirani, J., Padmakumar, V., & Bowman, S. R. (2023). Debate Helps Supervise Unreliable Experts. In arXiv. https://arxiv.org/abs/2311.08702Michael, J., S. Mahdi, D. Rein, et al. 2023. “Debate Helps Supervise Unreliable Experts”. In arXiv. Preprint, November 15. https://arxiv.org/abs/2311.08702.Michael, J., et al. “Debate Helps Supervise Unreliable Experts”. arXiv, 15 Nov. 2023, https://arxiv.org/abs/2311.08702.Michael, J. et al. Debate Helps Supervise Unreliable Experts. arXiv Preprint at https://arxiv.org/abs/2311.08702 (2023).J. Michael et al., “Debate Helps Supervise Unreliable Experts”, Nov. 15, 2023. [Online]. Available: https://arxiv.org/abs/2311.08702
Miller(2022). AI alignment with humans... but with which humans?. EA Forum.Miller. (2022). AI alignment with humans... but with which humans?. EA Forum. Internet Archive (https://web.archive.org/web/20260422074220/https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans). https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humansMiller. 2022. “AI Alignment with Humans... But with Which Humans?”. EA Forum. Https://web.archive.org/web/20260422074220/https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans. Internet Archive. https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans.Miller. “AI Alignment with Humans... But with Which Humans?”. EA Forum, 2022, Internet Archive, https://web.archive.org/web/20260422074220/https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans, https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans.Miller. AI alignment with humans... but with which humans?. EA Forum https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans (2022).Miller, “AI alignment with humans... but with which humans?”, EA Forum. Accessed: Apr. 22, 2026. [Online]. Available: https://forum.effectivealtruism.org/posts/DXuwsXsqGq5GtmsB3/ai-alignment-with-humans-but-with-which-humans
Millidge(2025). Open source AI has been vital for alignment.Millidge. (2025). Open source AI has been vital for alignment. https://beren.io/2023-11-05-Open-source-AI-has-been-vital-for-alignmentMillidge. 2025. “Open Source AI Has Been Vital for Alignment”. https://beren.io/2023-11-05-Open-source-AI-has-been-vital-for-alignment.Millidge. Open Source AI Has Been Vital for Alignment. 2025, https://beren.io/2023-11-05-Open-source-AI-has-been-vital-for-alignment.Millidge. Open source AI has been vital for alignment. https://beren.io/2023-11-05-Open-source-AI-has-been-vital-for-alignment (2025).Millidge, “Open source AI has been vital for alignment”. [Online]. Available: https://beren.io/2023-11-05-Open-source-AI-has-been-vital-for-alignment
Mingard et al.(2020). Neural networks are fundamentally Bayesian. Towards Data Science.Mingard et al. (2020, December 19). Neural networks are fundamentally Bayesian. Towards Data Science. https://towardsdatascience.com/neural-networks-are-fundamentally-bayesian-bee9a172fad8Mingard et al. 2020. “Neural Networks Are Fundamentally Bayesian”. Towards Data Science, December 19. https://towardsdatascience.com/neural-networks-are-fundamentally-bayesian-bee9a172fad8.Mingard et al. “Neural Networks Are Fundamentally Bayesian”. Towards Data Science, 19 Dec. 2020, https://towardsdatascience.com/neural-networks-are-fundamentally-bayesian-bee9a172fad8.Mingard et al. Neural networks are fundamentally Bayesian. Towards Data Science https://towardsdatascience.com/neural-networks-are-fundamentally-bayesian-bee9a172fad8 (2020).Mingard et al., “Neural networks are fundamentally Bayesian”, Towards Data Science. [Online]. Available: https://towardsdatascience.com/neural-networks-are-fundamentally-bayesian-bee9a172fad8
Miotti et al(2024). Introduction.Miotti et al. (2024). Introduction. https://narrowpath.co/introductionMiotti et al. 2024. “Introduction”. https://narrowpath.co/introduction.Miotti et al. Introduction. 2024, https://narrowpath.co/introduction.Miotti et al. Introduction. https://narrowpath.co/introduction (2024).Miotti et al, “Introduction”. [Online]. Available: https://narrowpath.co/introduction
Mirhoseini, A. et al.(2020). Chip Placement with Deep Reinforcement Learning. arXiv.Mirhoseini, A., Goldie, A., Yazgan, M., Jiang, J., Songhori, E., Wang, S., Lee, Y.-J., Johnson, E., Pathak, O., Bae, S., Nazi, A., Pak, J., Tong, A., Srinivasa, K., Hang, W., Tuncer, E., Babu, A., Le, Q. V., Laudon, J., … Dean, J. (2020). Chip Placement with Deep Reinforcement Learning. In arXiv. https://arxiv.org/abs/2004.10746Mirhoseini, A., A. Goldie, M. Yazgan, et al. 2020. “Chip Placement with Deep Reinforcement Learning”. In arXiv. Preprint, April 22. https://arxiv.org/abs/2004.10746.Mirhoseini, A., et al. “Chip Placement with Deep Reinforcement Learning”. arXiv, 22 Apr. 2020, https://arxiv.org/abs/2004.10746.Mirhoseini, A. et al. Chip Placement with Deep Reinforcement Learning. arXiv Preprint at https://arxiv.org/abs/2004.10746 (2020).A. Mirhoseini et al., “Chip Placement with Deep Reinforcement Learning”, Apr. 22, 2020. [Online]. Available: https://arxiv.org/abs/2004.10746
Mishra(2024). From Competition to Cooperation: Can US-China Engagement Overcome Geopolitical Barriers in AI Governance?. Tech Policy Press.Mishra. (2024, September 23). From Competition to Cooperation: Can US-China Engagement Overcome Geopolitical Barriers in AI Governance?. Tech Policy Press. https://techpolicy.press/from-competition-to-cooperation-can-uschina-engagement-overcome-geopolitical-barriers-in-ai-governanceMishra. 2024. “From Competition to Cooperation: Can US-China Engagement Overcome Geopolitical Barriers in AI Governance?”. Tech Policy Press, September 23. https://techpolicy.press/from-competition-to-cooperation-can-uschina-engagement-overcome-geopolitical-barriers-in-ai-governance.Mishra. “From Competition to Cooperation: Can US-China Engagement Overcome Geopolitical Barriers in AI Governance?”. Tech Policy Press, 23 Sept. 2024, https://techpolicy.press/from-competition-to-cooperation-can-uschina-engagement-overcome-geopolitical-barriers-in-ai-governance.Mishra. From Competition to Cooperation: Can US-China Engagement Overcome Geopolitical Barriers in AI Governance?. Tech Policy Press https://techpolicy.press/from-competition-to-cooperation-can-uschina-engagement-overcome-geopolitical-barriers-in-ai-governance (2024).Mishra, “From Competition to Cooperation: Can US-China Engagement Overcome Geopolitical Barriers in AI Governance?”, Tech Policy Press. [Online]. Available: https://techpolicy.press/from-competition-to-cooperation-can-uschina-engagement-overcome-geopolitical-barriers-in-ai-governance
Mnih, V. et al.(2013). Playing Atari with Deep Reinforcement Learning. arXiv.Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., & Riedmiller, M. (2013). Playing Atari with Deep Reinforcement Learning. In arXiv. https://arxiv.org/abs/1312.5602Mnih, V., K. Kavukcuoglu, D. Silver, et al. 2013. “Playing Atari with Deep Reinforcement Learning”. In arXiv. Preprint, December 19. https://arxiv.org/abs/1312.5602.Mnih, V., et al. “Playing Atari with Deep Reinforcement Learning”. arXiv, 19 Dec. 2013, https://arxiv.org/abs/1312.5602.Mnih, V. et al. Playing Atari with Deep Reinforcement Learning. arXiv Preprint at https://arxiv.org/abs/1312.5602 (2013).V. Mnih et al., “Playing Atari with Deep Reinforcement Learning”, Dec. 19, 2013. [Online]. Available: https://arxiv.org/abs/1312.5602
Moskvichev, A., Odouard, V. V. & Mitchell, M.(2023). The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain. arXiv.Moskvichev, A., Odouard, V. V., & Mitchell, M. (2023). The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain. In arXiv. https://arxiv.org/abs/2305.07141Moskvichev, A., V. V. Odouard, and M. Mitchell. 2023. “The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain”. In arXiv. Preprint, May 11. https://arxiv.org/abs/2305.07141.Moskvichev, A., et al. “The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain”. arXiv, 11 May 2023, https://arxiv.org/abs/2305.07141.Moskvichev, A., Odouard, V. V. & Mitchell, M. The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain. arXiv Preprint at https://arxiv.org/abs/2305.07141 (2023).A. Moskvichev, V. V. Odouard, and M. Mitchell, “The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain”, May 11, 2023. [Online]. Available: https://arxiv.org/abs/2305.07141
Mouton et al.(2024). Could Artificial Intelligence Be Misused to Plan Biological Attacks?.Mouton et al. (2024). Could Artificial Intelligence Be Misused to Plan Biological Attacks?. Internet Archive (https://web.archive.org/web/20260903172105/https://www.rand.org/pubs/research_reports/RRA2977-1.html). https://rand.org/pubs/research_reports/RRA2977-1.htmlMouton et al. 2024. “Could Artificial Intelligence Be Misused to Plan Biological Attacks?”. Https://web.archive.org/web/20260903172105/https://www.rand.org/pubs/research_reports/RRA2977-1.html. Internet Archive. https://rand.org/pubs/research_reports/RRA2977-1.html.Mouton et al. Could Artificial Intelligence Be Misused to Plan Biological Attacks?. 2024, Internet Archive, https://web.archive.org/web/20260903172105/https://www.rand.org/pubs/research_reports/RRA2977-1.html, https://rand.org/pubs/research_reports/RRA2977-1.html.Mouton et al. Could Artificial Intelligence Be Misused to Plan Biological Attacks?. https://rand.org/pubs/research_reports/RRA2977-1.html (2024).Mouton et al., “Could Artificial Intelligence Be Misused to Plan Biological Attacks?”. Accessed: Sep. 03, 2026. [Online]. Available: https://rand.org/pubs/research_reports/RRA2977-1.html
Mowshowitz(2023). The Crux List.Mowshowitz. (2023). The Crux List. https://thezvi.substack.com/p/the-crux-listMowshowitz. 2023. The Crux List. Edition. https://thezvi.substack.com/p/the-crux-list.Mowshowitz. The Crux List. 2023, https://thezvi.substack.com/p/the-crux-list.Mowshowitz. The Crux List. https://thezvi.substack.com/p/the-crux-list (2023).Mowshowitz, “The Crux List”. [Online]. Available: https://thezvi.substack.com/p/the-crux-list
Mowshowitz(2025). AI #68: Remarkably Reasonable Reactions.Mowshowitz. (2025). AI #68: Remarkably Reasonable Reactions. https://thezvi.substack.com/i/145384938/the-art-of-the-jailbreakMowshowitz. 2025. AI #68: Remarkably Reasonable Reactions. Edition. https://thezvi.substack.com/i/145384938/the-art-of-the-jailbreak.Mowshowitz. AI #68: Remarkably Reasonable Reactions. 2025, https://thezvi.substack.com/i/145384938/the-art-of-the-jailbreak.Mowshowitz. AI #68: Remarkably Reasonable Reactions. https://thezvi.substack.com/i/145384938/the-art-of-the-jailbreak (2025).Mowshowitz, “AI #68: Remarkably Reasonable Reactions”. [Online]. Available: https://thezvi.substack.com/i/145384938/the-art-of-the-jailbreak
Muehlhauser, L. & Salamon, A.(2012). Intelligence Explosion: Evidence and Import. Singularity Hypotheses: A Scientific and Philosophical Assessment.Muehlhauser, L., & Salamon, A. (2012). Intelligence Explosion: Evidence and Import. In A. Eden, J. Søraker, J. H. Moor, & E. Steinhart (eds.), Singularity Hypotheses: A Scientific and Philosophical Assessment. Springer. https://intelligence.org/files/IE-EI.pdfMuehlhauser, L., and A. Salamon. 2012. “Intelligence Explosion: Evidence and Import”. In Singularity Hypotheses: A Scientific and Philosophical Assessment, edited by A. Eden, J. Søraker, J. H. Moor, and E. Steinhart. Springer. https://intelligence.org/files/IE-EI.pdf.Muehlhauser, L., and A. Salamon. “Intelligence Explosion: Evidence and Import”. Singularity Hypotheses: A Scientific and Philosophical Assessment, edited by A. Eden et al., Springer, 2012, https://intelligence.org/files/IE-EI.pdf.Muehlhauser, L. & Salamon, A. Intelligence Explosion: Evidence and Import. in Singularity Hypotheses: A Scientific and Philosophical Assessment (eds Eden, A., Søraker, J., Moor, J. H. & Steinhart, E.) (Springer, Berlin, 2012).L. Muehlhauser and A. Salamon, “Intelligence Explosion: Evidence and Import”, in Singularity Hypotheses: A Scientific and Philosophical Assessment, A. Eden, J. Søraker, J. H. Moor, and E. Steinhart, Eds., Berlin: Springer, 2012. [Online]. Available: https://intelligence.org/files/IE-EI.pdf
Mukobi, G.(2024). Reasons to Doubt the Impact of AI Risk Evaluations. arXiv.Mukobi, G. (2024). Reasons to Doubt the Impact of AI Risk Evaluations. In arXiv. https://arxiv.org/abs/2408.02565Mukobi, G. 2024. “Reasons to Doubt the Impact of AI Risk Evaluations”. In arXiv. Preprint, August 5. https://arxiv.org/abs/2408.02565.Mukobi, G. “Reasons to Doubt the Impact of AI Risk Evaluations”. arXiv, 5 Aug. 2024, https://arxiv.org/abs/2408.02565.Mukobi, G. Reasons to Doubt the Impact of AI Risk Evaluations. arXiv Preprint at https://arxiv.org/abs/2408.02565 (2024).G. Mukobi, “Reasons to Doubt the Impact of AI Risk Evaluations”, Aug. 05, 2024. [Online]. Available: https://arxiv.org/abs/2408.02565
Murphy, T.(2013). The First Level of Super Mario Bros. is Easy with Lexicographic Orderings and Time Travel . . . after that it gets a little tricky.Murphy, T., VII. (2013). The First Level of Super Mario Bros. is Easy with Lexicographic Orderings and Time Travel . . . after that it gets a little tricky. https://tom7.org/mario/mario.pdfMurphy, T., VII. 2013. “The First Level of Super Mario Bros. Is Easy with Lexicographic Orderings and Time Travel . . . After That It Gets a Little Tricky.”. Preprint, April 1. https://tom7.org/mario/mario.pdf.Murphy, T., VII. “The First Level of Super Mario Bros. Is Easy with Lexicographic Orderings and Time Travel . . . After That It Gets a Little Tricky.”. 1 Apr. 2013, https://tom7.org/mario/mario.pdf.Murphy, T., VII. The First Level of Super Mario Bros. is Easy with Lexicographic Orderings and Time Travel . . . after that it gets a little tricky. Preprint at https://tom7.org/mario/mario.pdf (2013).T. Murphy VII, “The First Level of Super Mario Bros. is Easy with Lexicographic Orderings and Time Travel . . . after that it gets a little tricky.”, Apr. 01, 2013. [Online]. Available: https://tom7.org/mario/mario.pdf
Nakano, R. et al.(2021). WebGPT: Browser-assisted question-answering with human feedback. arXiv.Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., & Schulman, J. (2021). WebGPT: Browser-assisted question-answering with human feedback. In arXiv. https://arxiv.org/abs/2112.09332Nakano, R., J. Hilton, S. Balaji, et al. 2021. “WebGPT: Browser-assisted Question-answering with Human Feedback”. In arXiv. Preprint, December 17. https://arxiv.org/abs/2112.09332.Nakano, R., et al. “WebGPT: Browser-assisted Question-answering with Human Feedback”. arXiv, 17 Dec. 2021, https://arxiv.org/abs/2112.09332.Nakano, R. et al. WebGPT: Browser-assisted question-answering with human feedback. arXiv Preprint at https://arxiv.org/abs/2112.09332 (2021).R. Nakano et al., “WebGPT: Browser-assisted question-answering with human feedback”, Dec. 17, 2021. [Online]. Available: https://arxiv.org/abs/2112.09332
Nanda, N., Lee, A. & Wattenberg, M.(2023). Emergent Linear Representations in World Models of Self-Supervised Sequence Models. arXiv.Nanda, N., Lee, A., & Wattenberg, M. (2023). Emergent Linear Representations in World Models of Self-Supervised Sequence Models. In arXiv. https://arxiv.org/abs/2309.00941Nanda, N., A. Lee, and M. Wattenberg. 2023. “Emergent Linear Representations in World Models of Self-Supervised Sequence Models”. In arXiv. Preprint, September 2. https://arxiv.org/abs/2309.00941.Nanda, N., et al. “Emergent Linear Representations in World Models of Self-Supervised Sequence Models”. arXiv, 2 Sept. 2023, https://arxiv.org/abs/2309.00941.Nanda, N., Lee, A. & Wattenberg, M. Emergent Linear Representations in World Models of Self-Supervised Sequence Models. arXiv Preprint at https://arxiv.org/abs/2309.00941 (2023).N. Nanda, A. Lee, and M. Wattenberg, “Emergent Linear Representations in World Models of Self-Supervised Sequence Models”, Sep. 02, 2023. [Online]. Available: https://arxiv.org/abs/2309.00941
Narayan & Kapoor(2024). AI safety is not a model property.Narayan & Kapoor. (2024). AI safety is not a model property. https://aisnakeoil.com/p/ai-safety-is-not-a-model-propertyNarayan & Kapoor. 2024. “AI Safety Is Not a Model Property”. https://aisnakeoil.com/p/ai-safety-is-not-a-model-property.Narayan & Kapoor. AI Safety Is Not a Model Property. 2024, https://aisnakeoil.com/p/ai-safety-is-not-a-model-property.Narayan & Kapoor. AI safety is not a model property. https://aisnakeoil.com/p/ai-safety-is-not-a-model-property (2024).Narayan & Kapoor, “AI safety is not a model property”. [Online]. Available: https://aisnakeoil.com/p/ai-safety-is-not-a-model-property
NASA(2001). Research Satellites for Atmospheric Sciences, 1978-Present - NASA Science.NASA. (2001, December 10). Research Satellites for Atmospheric Sciences, 1978-Present - NASA Science. NASA Science. https://earthobservatory.nasa.gov/features/RemoteSensingAtmosphere/remote_sensing5.phpNASA. 2001. “Research Satellites for Atmospheric Sciences, 1978-Present - NASA Science”. NASA Science, December 10. https://earthobservatory.nasa.gov/features/RemoteSensingAtmosphere/remote_sensing5.php.NASA. “Research Satellites for Atmospheric Sciences, 1978-Present - NASA Science”. NASA Science, 10 Dec. 2001, https://earthobservatory.nasa.gov/features/RemoteSensingAtmosphere/remote_sensing5.php.NASA. Research Satellites for Atmospheric Sciences, 1978-Present - NASA Science. NASA Science https://earthobservatory.nasa.gov/features/RemoteSensingAtmosphere/remote_sensing5.php (2001).NASA, “Research Satellites for Atmospheric Sciences, 1978-Present - NASA Science”, NASA Science. [Online]. Available: https://earthobservatory.nasa.gov/features/RemoteSensingAtmosphere/remote_sensing5.php
Nasr, M. et al.(2023). Scalable Extraction of Training Data from (Production) Language Models. arXiv.Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramèr, F., & Lee, K. (2023). Scalable Extraction of Training Data from (Production) Language Models. In arXiv. https://arxiv.org/abs/2311.17035Nasr, M., N. Carlini, J. Hayase, et al. 2023. “Scalable Extraction of Training Data from (Production) Language Models”. In arXiv. Preprint, November 28. https://arxiv.org/abs/2311.17035.Nasr, M., et al. “Scalable Extraction of Training Data from (Production) Language Models”. arXiv, 28 Nov. 2023, https://arxiv.org/abs/2311.17035.Nasr, M. et al. Scalable Extraction of Training Data from (Production) Language Models. arXiv Preprint at https://arxiv.org/abs/2311.17035 (2023).M. Nasr et al., “Scalable Extraction of Training Data from (Production) Language Models”, Nov. 28, 2023. [Online]. Available: https://arxiv.org/abs/2311.17035
National Security Commission on Emerging Biotechnology(2024). White Paper 3: Risks of AIxBio.National Security Commission on Emerging Biotechnology. (2024). White Paper 3: Risks of AIxBio (NSCEB White Paper Series on AIxBio). National Security Commission on Emerging Biotechnology. https://biotech.senate.gov/wp-content/uploads/2024/01/NSCEB_AIxBio_WP3_Risks.pdfNational Security Commission on Emerging Biotechnology. 2024. White Paper 3: Risks of AIxBio. NSCEB White Paper Series on AIxBio. National Security Commission on Emerging Biotechnology. https://biotech.senate.gov/wp-content/uploads/2024/01/NSCEB_AIxBio_WP3_Risks.pdf.National Security Commission on Emerging Biotechnology. White Paper 3: Risks of AIxBio. National Security Commission on Emerging Biotechnology, Jan. 2024, https://biotech.senate.gov/wp-content/uploads/2024/01/NSCEB_AIxBio_WP3_Risks.pdf. NSCEB White Paper Series on AIxBio.National Security Commission on Emerging Biotechnology. White Paper 3: Risks of AIxBio. https://biotech.senate.gov/wp-content/uploads/2024/01/NSCEB_AIxBio_WP3_Risks.pdf (2024).National Security Commission on Emerging Biotechnology, “White Paper 3: Risks of AIxBio”, National Security Commission on Emerging Biotechnology, Jan. 2024. [Online]. Available: https://biotech.senate.gov/wp-content/uploads/2024/01/NSCEB_AIxBio_WP3_Risks.pdf
Neel Nanda(2025). Interpretability Will Not Reliably Find Deceptive AI. AI Alignment Forum.Neel Nanda. (2025, May 4). Interpretability Will Not Reliably Find Deceptive AI. AI Alignment Forum. https://alignmentforum.org/posts/PwnadG4BFjaER3MGf/interpretability-will-not-reliably-find-deceptive-aiNeel Nanda. 2025. “Interpretability Will Not Reliably Find Deceptive AI”. AI Alignment Forum, May 4. https://alignmentforum.org/posts/PwnadG4BFjaER3MGf/interpretability-will-not-reliably-find-deceptive-ai.Neel Nanda. “Interpretability Will Not Reliably Find Deceptive AI”. AI Alignment Forum, 4 May 2025, https://alignmentforum.org/posts/PwnadG4BFjaER3MGf/interpretability-will-not-reliably-find-deceptive-ai.Neel Nanda. Interpretability Will Not Reliably Find Deceptive AI. AI Alignment Forum https://alignmentforum.org/posts/PwnadG4BFjaER3MGf/interpretability-will-not-reliably-find-deceptive-ai (2025).Neel Nanda, “Interpretability Will Not Reliably Find Deceptive AI”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/PwnadG4BFjaER3MGf/interpretability-will-not-reliably-find-deceptive-ai
Nevo et al.(2024). How AI Labs Can Safeguard Model Weights.Nevo et al. (2024). How AI Labs Can Safeguard Model Weights. Internet Archive (https://web.archive.org/web/20260920224534/https://www.rand.org/pubs/research_reports/RRA2849-1.html). https://rand.org/pubs/research_reports/RRA2849-1.htmlNevo et al. 2024. “How AI Labs Can Safeguard Model Weights”. Https://web.archive.org/web/20260920224534/https://www.rand.org/pubs/research_reports/RRA2849-1.html. Internet Archive. https://rand.org/pubs/research_reports/RRA2849-1.html.Nevo et al. How AI Labs Can Safeguard Model Weights. 2024, Internet Archive, https://web.archive.org/web/20260920224534/https://www.rand.org/pubs/research_reports/RRA2849-1.html, https://rand.org/pubs/research_reports/RRA2849-1.html.Nevo et al. How AI Labs Can Safeguard Model Weights. https://rand.org/pubs/research_reports/RRA2849-1.html (2024).Nevo et al., “How AI Labs Can Safeguard Model Weights”. Accessed: Sep. 20, 2026. [Online]. Available: https://rand.org/pubs/research_reports/RRA2849-1.html
Newman(2024). Cybersecurity and AI: The Evolving Security Landscape | CAIS. Center for AI Safety.Newman. (2024). Cybersecurity and AI: The Evolving Security Landscape | CAIS. Center for AI Safety. https://safe.ai/blog/cybersecurity-and-ai-the-evolving-security-landscapeNewman. 2024. “Cybersecurity and AI: The Evolving Security Landscape | CAIS”. Center for AI Safety. https://safe.ai/blog/cybersecurity-and-ai-the-evolving-security-landscape.Newman. “Cybersecurity and AI: The Evolving Security Landscape | CAIS”. Center for AI Safety, 2024, https://safe.ai/blog/cybersecurity-and-ai-the-evolving-security-landscape.Newman. Cybersecurity and AI: The Evolving Security Landscape | CAIS. Center for AI Safety https://safe.ai/blog/cybersecurity-and-ai-the-evolving-security-landscape (2024).Newman, “Cybersecurity and AI: The Evolving Security Landscape | CAIS”, Center for AI Safety. [Online]. Available: https://safe.ai/blog/cybersecurity-and-ai-the-evolving-security-landscape
Nguyen, N., Chandrasegaran, K., Abdollahzadeh, M. & Cheung, N.(2023). Re-thinking Model Inversion Attacks Against Deep Neural Networks. arXiv.Nguyen, N.-B., Chandrasegaran, K., Abdollahzadeh, M., & Cheung, N.-M. (2023). Re-thinking Model Inversion Attacks Against Deep Neural Networks. In arXiv. https://arxiv.org/abs/2304.01669Nguyen, N.-B., K. Chandrasegaran, M. Abdollahzadeh, and N.-M. Cheung. 2023. “Re-thinking Model Inversion Attacks Against Deep Neural Networks”. In arXiv. Preprint, April 4. https://arxiv.org/abs/2304.01669.Nguyen, N.-B., et al. “Re-thinking Model Inversion Attacks Against Deep Neural Networks”. arXiv, 4 Apr. 2023, https://arxiv.org/abs/2304.01669.Nguyen, N.-B., Chandrasegaran, K., Abdollahzadeh, M. & Cheung, N.-M. Re-thinking Model Inversion Attacks Against Deep Neural Networks. arXiv Preprint at https://arxiv.org/abs/2304.01669 (2023).N.-B. Nguyen, K. Chandrasegaran, M. Abdollahzadeh, and N.-M. Cheung, “Re-thinking Model Inversion Attacks Against Deep Neural Networks”, Apr. 04, 2023. [Online]. Available: https://arxiv.org/abs/2304.01669
Niki Dupuis & janus(2023). Cyborgism. AI Alignment Forum.Niki Dupuis, & janus. (2023, February 10). Cyborgism. AI Alignment Forum. https://alignmentforum.org/posts/bxt7uCiHam4QXrQAA/cyborgismNiki Dupuis, and janus. 2023. “Cyborgism”. AI Alignment Forum, February 10. https://alignmentforum.org/posts/bxt7uCiHam4QXrQAA/cyborgism.Niki Dupuis, and janus. “Cyborgism”. AI Alignment Forum, 10 Feb. 2023, https://alignmentforum.org/posts/bxt7uCiHam4QXrQAA/cyborgism.Niki Dupuis & janus. Cyborgism. AI Alignment Forum https://alignmentforum.org/posts/bxt7uCiHam4QXrQAA/cyborgism (2023).Niki Dupuis and janus, “Cyborgism”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/bxt7uCiHam4QXrQAA/cyborgism
NIST(2021). AI Risk Management Framework. NIST.NIST. (2021, July 12). AI Risk Management Framework. NIST. https://nist.gov/itl/ai-risk-management-frameworkNIST. 2021. “AI Risk Management Framework”. NIST, July 12. https://nist.gov/itl/ai-risk-management-framework.NIST. “AI Risk Management Framework”. NIST, 12 July 2021, https://nist.gov/itl/ai-risk-management-framework.NIST. AI Risk Management Framework. NIST https://nist.gov/itl/ai-risk-management-framework (2021).NIST, “AI Risk Management Framework”, NIST. [Online]. Available: https://nist.gov/itl/ai-risk-management-framework
Nora_Ammann(2025). In response to critiques of Guaranteed Safe AI. AI Alignment Forum.Nora_Ammann. (2025, January 31). In response to critiques of Guaranteed Safe AI. AI Alignment Forum. https://alignmentforum.org/posts/DZuBHHKao6jsDDreH/in-response-to-critiques-of-guaranteed-safe-aiNora_Ammann. 2025. “In Response to Critiques of Guaranteed Safe AI”. AI Alignment Forum, January 31. https://alignmentforum.org/posts/DZuBHHKao6jsDDreH/in-response-to-critiques-of-guaranteed-safe-ai.Nora_Ammann. “In Response to Critiques of Guaranteed Safe AI”. AI Alignment Forum, 31 Jan. 2025, https://alignmentforum.org/posts/DZuBHHKao6jsDDreH/in-response-to-critiques-of-guaranteed-safe-ai.Nora_Ammann. In response to critiques of Guaranteed Safe AI. AI Alignment Forum https://alignmentforum.org/posts/DZuBHHKao6jsDDreH/in-response-to-critiques-of-guaranteed-safe-ai (2025).Nora_Ammann, “In response to critiques of Guaranteed Safe AI”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/DZuBHHKao6jsDDreH/in-response-to-critiques-of-guaranteed-safe-ai
Novikov, A. et al.(2025). AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv.Novikov, A., Vũ, N., Eisenberger, M., Dupont, E., Huang, P.-S., Wagner, A. Z., Shirobokov, S., Kozlovskii, B., Ruiz, F. J. R., Mehrabian, A., Kumar, M. P., See, A., Chaudhuri, S., Holland, G., Davies, A., Nowozin, S., Kohli, P., & Balog, M. (2025). AlphaEvolve: A coding agent for scientific and algorithmic discovery. In arXiv. https://arxiv.org/abs/2506.13131Novikov, A., N. Vũ, M. Eisenberger, et al. 2025. “AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery”. In arXiv. Preprint, June 16. https://arxiv.org/abs/2506.13131.Novikov, A., et al. “AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery”. arXiv, 16 June 2025, https://arxiv.org/abs/2506.13131.Novikov, A. et al. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv Preprint at https://arxiv.org/abs/2506.13131 (2025).A. Novikov et al., “AlphaEvolve: A coding agent for scientific and algorithmic discovery”, Jun. 16, 2025. [Online]. Available: https://arxiv.org/abs/2506.13131
O'Brien, J., Ee, S. & Williams, Z.(2023). Deployment Corrections: An incident response framework for frontier AI models. arXiv.O'Brien, J., Ee, S., & Williams, Z. (2023). Deployment Corrections: An incident response framework for frontier AI models. In arXiv. https://arxiv.org/abs/2310.00328O'Brien, J., S. Ee, and Z. Williams. 2023. “Deployment Corrections: An Incident Response Framework for Frontier AI Models”. In arXiv. Preprint, September 30. https://arxiv.org/abs/2310.00328.O'Brien, J., et al. “Deployment Corrections: An Incident Response Framework for Frontier AI Models”. arXiv, 30 Sept. 2023, https://arxiv.org/abs/2310.00328.O'Brien, J., Ee, S. & Williams, Z. Deployment Corrections: An incident response framework for frontier AI models. arXiv Preprint at https://arxiv.org/abs/2310.00328 (2023).J. O'Brien, S. Ee, and Z. Williams, “Deployment Corrections: An incident response framework for frontier AI models”, Sep. 30, 2023. [Online]. Available: https://arxiv.org/abs/2310.00328
OECD(2025). Overview of current AI capabilities: Introducing the OECD AI Capability Indicators. OECD.OECD. (2025). Overview of current AI capabilities: Introducing the OECD AI Capability Indicators. Internet Archive (https://web.archive.org/web/20260527194732/https://www.oecd.org/en/publications/2025/06/introducing-the-oecd-ai-capability-indicators_7c0731f0/full-report/component-4.html). OECD. https://oecd.org/en/publications/2025/06/introducing-the-oecd-ai-capability-indicators_7c0731f0/full-report/component-4.htmlOECD. 2025. “Overview of Current AI Capabilities: Introducing the OECD AI Capability Indicators”. OECD. Https://web.archive.org/web/20260527194732/https://www.oecd.org/en/publications/2025/06/introducing-the-oecd-ai-capability-indicators_7c0731f0/full-report/component-4.html. Internet Archive. https://oecd.org/en/publications/2025/06/introducing-the-oecd-ai-capability-indicators_7c0731f0/full-report/component-4.html.OECD. “Overview of Current AI Capabilities: Introducing the OECD AI Capability Indicators”. OECD, 2025, Internet Archive, https://web.archive.org/web/20260527194732/https://www.oecd.org/en/publications/2025/06/introducing-the-oecd-ai-capability-indicators_7c0731f0/full-report/component-4.html, https://oecd.org/en/publications/2025/06/introducing-the-oecd-ai-capability-indicators_7c0731f0/full-report/component-4.html.OECD. Overview of current AI capabilities: Introducing the OECD AI Capability Indicators. OECD https://oecd.org/en/publications/2025/06/introducing-the-oecd-ai-capability-indicators_7c0731f0/full-report/component-4.html (2025).OECD, “Overview of current AI capabilities: Introducing the OECD AI Capability Indicators”, OECD. Accessed: May 27, 2026. [Online]. Available: https://oecd.org/en/publications/2025/06/introducing-the-oecd-ai-capability-indicators_7c0731f0/full-report/component-4.html
OECD(2025). Towards a common reporting framework for AI incidents (EN).OECD. (2025, February). Towards a common reporting framework for AI incidents (EN). https://oecd.org/content/dam/oecd/en/publications/reports/2025/02/towards-a-common-reporting-framework-for-ai-incidents_8c488fdb/f326d4ac-en.pdfOECD. 2025. “Towards a Common Reporting Framework for AI Incidents (EN)”. February. https://oecd.org/content/dam/oecd/en/publications/reports/2025/02/towards-a-common-reporting-framework-for-ai-incidents_8c488fdb/f326d4ac-en.pdf.OECD. Towards a Common Reporting Framework for AI Incidents (EN). Feb. 2025, https://oecd.org/content/dam/oecd/en/publications/reports/2025/02/towards-a-common-reporting-framework-for-ai-incidents_8c488fdb/f326d4ac-en.pdf.OECD. Towards a common reporting framework for AI incidents (EN). https://oecd.org/content/dam/oecd/en/publications/reports/2025/02/towards-a-common-reporting-framework-for-ai-incidents_8c488fdb/f326d4ac-en.pdf (2025).OECD, “Towards a common reporting framework for AI incidents (EN)”. [Online]. Available: https://oecd.org/content/dam/oecd/en/publications/reports/2025/02/towards-a-common-reporting-framework-for-ai-incidents_8c488fdb/f326d4ac-en.pdf
Office of the Attorney General(2023). AG Campbell Files Lawsuit Against Meta, Instagram For Unfair And Deceptive Practices That Harm Young People. Mass.gov.Office of the Attorney General. (2023). AG Campbell Files Lawsuit Against Meta, Instagram For Unfair And Deceptive Practices That Harm Young People. Mass.gov. https://mass.gov/news/ag-campbell-files-lawsuit-against-meta-instagram-for-unfair-and-deceptive-practices-that-harm-young-peopleOffice of the Attorney General. 2023. “AG Campbell Files Lawsuit Against Meta, Instagram For Unfair And Deceptive Practices That Harm Young People”. Mass.gov. https://mass.gov/news/ag-campbell-files-lawsuit-against-meta-instagram-for-unfair-and-deceptive-practices-that-harm-young-people.Office of the Attorney General. “AG Campbell Files Lawsuit Against Meta, Instagram For Unfair And Deceptive Practices That Harm Young People”. Mass.gov, 2023, https://mass.gov/news/ag-campbell-files-lawsuit-against-meta-instagram-for-unfair-and-deceptive-practices-that-harm-young-people.Office of the Attorney General. AG Campbell Files Lawsuit Against Meta, Instagram For Unfair And Deceptive Practices That Harm Young People. Mass.gov https://mass.gov/news/ag-campbell-files-lawsuit-against-meta-instagram-for-unfair-and-deceptive-practices-that-harm-young-people (2023).Office of the Attorney General, “AG Campbell Files Lawsuit Against Meta, Instagram For Unfair And Deceptive Practices That Harm Young People”, Mass.gov. [Online]. Available: https://mass.gov/news/ag-campbell-files-lawsuit-against-meta-instagram-for-unfair-and-deceptive-practices-that-harm-young-people
Olds, J. & Milner, P.(1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. Journal of Comparative and Physiological Psychology.Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. Journal of Comparative and Physiological Psychology. https://doi.org/10.1037/h0058775Olds, J., and P. Milner. 1954. “Positive Reinforcement Produced by Electrical Stimulation of Septal Area and Other Regions of Rat Brain.”. Journal of Comparative and Physiological Psychology, ahead of print. https://doi.org/10.1037/h0058775.Olds, J., and P. Milner. “Positive Reinforcement Produced by Electrical Stimulation of Septal Area and Other Regions of Rat Brain.”. Journal of Comparative and Physiological Psychology, 1954, https://doi.org/10.1037/h0058775.Olds, J. & Milner, P. Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. Journal of Comparative and Physiological Psychology https://doi.org/10.1037/h0058775 (1954) doi:10.1037/h0058775.J. Olds and P. Milner, “Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain.”, Journal of Comparative and Physiological Psychology, 1954, doi: 10.1037/h0058775.
Olds, J.(1970). Pleasure Centers in the Brain. Engineering and Science.Olds, J. (1970). Pleasure Centers in the Brain. Engineering and Science, 33(7), 22–31. https://calteches.library.caltech.edu/2807/1/olds.pdfOlds, J. 1970. “Pleasure Centers in the Brain”. Engineering and Science, Edition. https://calteches.library.caltech.edu/2807/1/olds.pdf.Olds, J. “Pleasure Centers in the Brain”. Engineering and Science, vol. 33, no. 7, 1970, pp. 22–31, https://calteches.library.caltech.edu/2807/1/olds.pdf.Olds, J. Pleasure Centers in the Brain. Engineering and Science vol. 33 22–31 (1970).J. Olds, “Pleasure Centers in the Brain”, Engineering and Science, vol. 33, no. 7, pp. 22–31, 1970. [Online]. Available: https://calteches.library.caltech.edu/2807/1/olds.pdf
Omohundro(2008). The Basic AI Drives | Proceedings of the 2008 conference on Artificial General Intelligence 2008: Proceedings of the First AGI Conference. Guide Proceedings.Omohundro. (2008). The Basic AI Drives | Proceedings of the 2008 conference on Artificial General Intelligence 2008: Proceedings of the First AGI Conference. Internet Archive (https://web.archive.org/web/20250808083254/https://dl.acm.org/doi/10.5555/1566174.1566226). Guide Proceedings. https://dl.acm.org/doi/10.5555/1566174.1566226Omohundro. 2008. “The Basic AI Drives | Proceedings of the 2008 Conference on Artificial General Intelligence 2008: Proceedings of the First AGI Conference”. Guide Proceedings. Https://web.archive.org/web/20250808083254/https://dl.acm.org/doi/10.5555/1566174.1566226. Internet Archive. https://dl.acm.org/doi/10.5555/1566174.1566226.Omohundro. “The Basic AI Drives | Proceedings of the 2008 Conference on Artificial General Intelligence 2008: Proceedings of the First AGI Conference”. Guide Proceedings, 2008, Internet Archive, https://web.archive.org/web/20250808083254/https://dl.acm.org/doi/10.5555/1566174.1566226, https://dl.acm.org/doi/10.5555/1566174.1566226.Omohundro. The Basic AI Drives | Proceedings of the 2008 conference on Artificial General Intelligence 2008: Proceedings of the First AGI Conference. Guide Proceedings https://dl.acm.org/doi/10.5555/1566174.1566226 (2008).Omohundro, “The Basic AI Drives | Proceedings of the 2008 conference on Artificial General Intelligence 2008: Proceedings of the First AGI Conference”, Guide Proceedings. Accessed: Aug. 08, 2025. [Online]. Available: https://dl.acm.org/doi/10.5555/1566174.1566226
OpenAI(2017). Attacking machine learning with adversarial examples.OpenAI. (2017). Attacking machine learning with adversarial examples. Internet Archive (https://web.archive.org/web/20240406184223/https://openai.com/research/attacking-machine-learning-with-adversarial-examples). https://openai.com/research/attacking-machine-learning-with-adversarial-examplesOpenAI. 2017. “Attacking Machine Learning with Adversarial Examples”. Https://web.archive.org/web/20240406184223/https://openai.com/research/attacking-machine-learning-with-adversarial-examples. Internet Archive. https://openai.com/research/attacking-machine-learning-with-adversarial-examples.OpenAI. Attacking Machine Learning with Adversarial Examples. 2017, Internet Archive, https://web.archive.org/web/20240406184223/https://openai.com/research/attacking-machine-learning-with-adversarial-examples, https://openai.com/research/attacking-machine-learning-with-adversarial-examples.OpenAI. Attacking machine learning with adversarial examples. https://openai.com/research/attacking-machine-learning-with-adversarial-examples (2017).OpenAI, “Attacking machine learning with adversarial examples”. Accessed: Apr. 06, 2024. [Online]. Available: https://openai.com/research/attacking-machine-learning-with-adversarial-examples
OpenAI(2017). Learning from human preferences. OpenAI.OpenAI. (2017). Learning from human preferences. OpenAI. https://openai.com/index/learning-from-human-preferencesOpenAI. 2017. “Learning from Human Preferences”. OpenAI. https://openai.com/index/learning-from-human-preferences.OpenAI. “Learning from Human Preferences”. OpenAI, 2017, https://openai.com/index/learning-from-human-preferences.OpenAI. Learning from human preferences. OpenAI https://openai.com/index/learning-from-human-preferences (2017).OpenAI, “Learning from human preferences”, OpenAI. [Online]. Available: https://openai.com/index/learning-from-human-preferences
OpenAI(2018). Learning Montezuma’s Revenge from a single demonstration.OpenAI. (2018). Learning Montezuma’s Revenge from a single demonstration. Internet Archive (https://web.archive.org/web/20240305020425/https://openai.com/research/learning-montezumas-revenge-from-a-single-demonstration). https://openai.com/research/learning-montezumas-revenge-from-a-single-demonstrationOpenAI. 2018. “Learning Montezuma’s Revenge from a Single Demonstration”. Https://web.archive.org/web/20240305020425/https://openai.com/research/learning-montezumas-revenge-from-a-single-demonstration. Internet Archive. https://openai.com/research/learning-montezumas-revenge-from-a-single-demonstration.OpenAI. Learning Montezuma’s Revenge from a Single Demonstration. 2018, Internet Archive, https://web.archive.org/web/20240305020425/https://openai.com/research/learning-montezumas-revenge-from-a-single-demonstration, https://openai.com/research/learning-montezumas-revenge-from-a-single-demonstration.OpenAI. Learning Montezuma’s Revenge from a single demonstration. https://openai.com/research/learning-montezumas-revenge-from-a-single-demonstration (2018).OpenAI, “Learning Montezuma’s Revenge from a single demonstration”. Accessed: Mar. 05, 2024. [Online]. Available: https://openai.com/research/learning-montezumas-revenge-from-a-single-demonstration
OpenAI(2019). OpenAI Five defeats Dota 2 world champions.OpenAI. (2019). OpenAI Five defeats Dota 2 world champions. Internet Archive (https://web.archive.org/web/20240425003107/https://openai.com/research/openai-five-defeats-dota-2-world-champions). https://openai.com/research/openai-five-defeats-dota-2-world-championsOpenAI. 2019. “OpenAI Five Defeats Dota 2 World Champions”. Https://web.archive.org/web/20240425003107/https://openai.com/research/openai-five-defeats-dota-2-world-champions. Internet Archive. https://openai.com/research/openai-five-defeats-dota-2-world-champions.OpenAI. OpenAI Five Defeats Dota 2 World Champions. 2019, Internet Archive, https://web.archive.org/web/20240425003107/https://openai.com/research/openai-five-defeats-dota-2-world-champions, https://openai.com/research/openai-five-defeats-dota-2-world-champions.OpenAI. OpenAI Five defeats Dota 2 world champions. https://openai.com/research/openai-five-defeats-dota-2-world-champions (2019).OpenAI, “OpenAI Five defeats Dota 2 world champions”. Accessed: Apr. 25, 2024. [Online]. Available: https://openai.com/research/openai-five-defeats-dota-2-world-champions
OpenAI(2022). Aligning language models to follow instructions. OpenAI.OpenAI. (2022). Aligning language models to follow instructions. OpenAI. https://openai.com/research/instruction-followingOpenAI. 2022. “Aligning Language Models to Follow Instructions”. OpenAI. https://openai.com/research/instruction-following.OpenAI. “Aligning Language Models to Follow Instructions”. OpenAI, 2022, https://openai.com/research/instruction-following.OpenAI. Aligning language models to follow instructions. OpenAI https://openai.com/research/instruction-following (2022).OpenAI, “Aligning language models to follow instructions”, OpenAI. [Online]. Available: https://openai.com/research/instruction-following
OpenAI(2023). Our approach to AI safety. OpenAI.OpenAI. (2023). Our approach to AI safety. OpenAI. https://openai.com/blog/our-approach-to-ai-safetyOpenAI. 2023. “Our Approach to AI Safety”. OpenAI. https://openai.com/blog/our-approach-to-ai-safety.OpenAI. “Our Approach to AI Safety”. OpenAI, 2023, https://openai.com/blog/our-approach-to-ai-safety.OpenAI. Our approach to AI safety. OpenAI https://openai.com/blog/our-approach-to-ai-safety (2023).OpenAI, “Our approach to AI safety”, OpenAI. [Online]. Available: https://openai.com/blog/our-approach-to-ai-safety
OpenAI(2025). Introducing OpenAI o3 and o4-mini. OpenAI.OpenAI. (2025). Introducing OpenAI o3 and o4-mini. OpenAI. https://openai.com/index/introducing-o3-and-o4-miniOpenAI. 2025. “Introducing OpenAI O3 and O4-mini”. OpenAI. https://openai.com/index/introducing-o3-and-o4-mini.OpenAI. “Introducing OpenAI O3 and O4-mini”. OpenAI, 2025, https://openai.com/index/introducing-o3-and-o4-mini.OpenAI. Introducing OpenAI o3 and o4-mini. OpenAI https://openai.com/index/introducing-o3-and-o4-mini (2025).OpenAI, “Introducing OpenAI o3 and o4-mini”, OpenAI. [Online]. Available: https://openai.com/index/introducing-o3-and-o4-mini
OpenAI(2025). OpenAI (@OpenAI) on X. X (formerly Twitter).OpenAI. (2025, July 19). OpenAI (@OpenAI) on X. X (formerly Twitter). https://x.com/OpenAI/status/1946594928945148246OpenAI. 2025. “OpenAI (@OpenAI) on X”. X (formerly Twitter), July 19. https://x.com/OpenAI/status/1946594928945148246.OpenAI. “OpenAI (@OpenAI) on X”. X (formerly Twitter), 19 July 2025, https://x.com/OpenAI/status/1946594928945148246.OpenAI. OpenAI (@OpenAI) on X. X (formerly Twitter) https://x.com/OpenAI/status/1946594928945148246 (2025).OpenAI, “OpenAI (@OpenAI) on X”, X (formerly Twitter). [Online]. Available: https://x.com/OpenAI/status/1946594928945148246
OpenAI(2025). Sora 2 is here. OpenAI.OpenAI. (2025). Sora 2 is here. OpenAI. https://openai.com/index/sora-2OpenAI. 2025. “Sora 2 Is Here”. OpenAI. https://openai.com/index/sora-2.OpenAI. “Sora 2 Is Here”. OpenAI, 2025, https://openai.com/index/sora-2.OpenAI. Sora 2 is here. OpenAI https://openai.com/index/sora-2 (2025).OpenAI, “Sora 2 is here”, OpenAI. [Online]. Available: https://openai.com/index/sora-2
OpenAI et al.(2023). GPT-4 Technical Report. arXiv.OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., … Zoph, B. (2023). GPT-4 Technical Report. In arXiv. https://arxiv.org/abs/2303.08774OpenAI, J. Achiam, S. Adler, et al. 2023. “GPT-4 Technical Report”. In arXiv. Preprint, March 15. https://arxiv.org/abs/2303.08774.OpenAI, et al. “GPT-4 Technical Report”. arXiv, 15 Mar. 2023, https://arxiv.org/abs/2303.08774.OpenAI et al. GPT-4 Technical Report. arXiv Preprint at https://arxiv.org/abs/2303.08774 (2023).OpenAI et al., “GPT-4 Technical Report”, Mar. 15, 2023. [Online]. Available: https://arxiv.org/abs/2303.08774
OpenAI et al.(2024). OpenAI o1 System Card. arXiv.OpenAI, :, Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., Iftimie, A., Karpenko, A., Passos, A. T., Neitz, A., Prokofiev, A., Wei, A., Tam, A., … Li, Z. (2024). OpenAI o1 System Card. In arXiv. https://arxiv.org/abs/2412.16720OpenAI, :, A. Jaech, et al. 2024. “OpenAI O1 System Card”. In arXiv. Preprint, December 21. https://arxiv.org/abs/2412.16720.OpenAI, et al. “OpenAI O1 System Card”. arXiv, 21 Dec. 2024, https://arxiv.org/abs/2412.16720.OpenAI et al. OpenAI o1 System Card. arXiv Preprint at https://arxiv.org/abs/2412.16720 (2024).OpenAI et al., “OpenAI o1 System Card”, Dec. 21, 2024. [Online]. Available: https://arxiv.org/abs/2412.16720
Ord(2020). The Precipice.Ord. (2020). The Precipice. The Precipice. https://theprecipice.comOrd. 2020. “The Precipice”. The Precipice. https://theprecipice.com.Ord. “The Precipice”. The Precipice, 2020, https://theprecipice.com.Ord. The Precipice. The Precipice https://theprecipice.com (2020).Ord, “The Precipice”, The Precipice. [Online]. Available: https://theprecipice.com
Orpheus16 & hath(2023). Speaking to Congressional staffers about AI risk. LessWrong.Orpheus16, & hath. (2023, December 4). Speaking to Congressional staffers about AI risk. LessWrong. https://lesswrong.com/posts/2sLwt2cSAag74nsdN/speaking-to-congressional-staffers-about-ai-riskOrpheus16, and hath. 2023. “Speaking to Congressional Staffers About AI Risk”. LessWrong, December 4. https://lesswrong.com/posts/2sLwt2cSAag74nsdN/speaking-to-congressional-staffers-about-ai-risk.Orpheus16, and hath. “Speaking to Congressional Staffers About AI Risk”. LessWrong, 4 Dec. 2023, https://lesswrong.com/posts/2sLwt2cSAag74nsdN/speaking-to-congressional-staffers-about-ai-risk.Orpheus16 & hath. Speaking to Congressional staffers about AI risk. LessWrong https://lesswrong.com/posts/2sLwt2cSAag74nsdN/speaking-to-congressional-staffers-about-ai-risk (2023).Orpheus16 and hath, “Speaking to Congressional staffers about AI risk”, LessWrong. [Online]. Available: https://lesswrong.com/posts/2sLwt2cSAag74nsdN/speaking-to-congressional-staffers-about-ai-risk
Orpheus16(2022). My thoughts on OpenAI's alignment plan. LessWrong.Orpheus16. (2022, December 30). My thoughts on OpenAI's alignment plan. LessWrong. https://lesswrong.com/posts/FBG7AghvvP7fPYzkx/my-thoughts-on-openai-s-alignment-plan-1Orpheus16. 2022. “My Thoughts on OpenAI's Alignment Plan”. LessWrong, December 30. https://lesswrong.com/posts/FBG7AghvvP7fPYzkx/my-thoughts-on-openai-s-alignment-plan-1.Orpheus16. “My Thoughts on OpenAI's Alignment Plan”. LessWrong, 30 Dec. 2022, https://lesswrong.com/posts/FBG7AghvvP7fPYzkx/my-thoughts-on-openai-s-alignment-plan-1.Orpheus16. My thoughts on OpenAI's alignment plan. LessWrong https://lesswrong.com/posts/FBG7AghvvP7fPYzkx/my-thoughts-on-openai-s-alignment-plan-1 (2022).Orpheus16, “My thoughts on OpenAI's alignment plan”, LessWrong. [Online]. Available: https://lesswrong.com/posts/FBG7AghvvP7fPYzkx/my-thoughts-on-openai-s-alignment-plan-1
Our World in Data(2024). Energy Production and Consumption. Our World in Data.Our World in Data. (2024). Energy Production and Consumption. Our World in Data. https://ourworldindata.org/energy-production-consumptionOur World in Data. 2024. “Energy Production and Consumption”. Our World in Data. https://ourworldindata.org/energy-production-consumption.Our World in Data. “Energy Production and Consumption”. Our World in Data, 2024, https://ourworldindata.org/energy-production-consumption.Our World in Data. Energy Production and Consumption. Our World in Data https://ourworldindata.org/energy-production-consumption (2024).Our World in Data, “Energy Production and Consumption”, Our World in Data. [Online]. Available: https://ourworldindata.org/energy-production-consumption
Our World in Data(2026). Levelized cost of energy for renewables. Our World in Data.Our World in Data. (2026). Levelized cost of energy for renewables. Our World in Data. https://ourworldindata.org/grapher/levelized-cost-of-energyOur World in Data. 2026. “Levelized Cost of Energy for Renewables”. Our World in Data. https://ourworldindata.org/grapher/levelized-cost-of-energy.Our World in Data. “Levelized Cost of Energy for Renewables”. Our World in Data, 2026, https://ourworldindata.org/grapher/levelized-cost-of-energy.Our World in Data. Levelized cost of energy for renewables. Our World in Data https://ourworldindata.org/grapher/levelized-cost-of-energy (2026).Our World in Data, “Levelized cost of energy for renewables”, Our World in Data. [Online]. Available: https://ourworldindata.org/grapher/levelized-cost-of-energy
Owen(2025). What will AI look like in 2030?. Epoch AI.Owen. (2025). What will AI look like in 2030?. Epoch AI. https://epoch.ai/blog/what-will-ai-look-like-in-2030Owen. 2025. “What Will AI Look Like in 2030?”. Epoch AI. https://epoch.ai/blog/what-will-ai-look-like-in-2030.Owen. “What Will AI Look Like in 2030?”. Epoch AI, 2025, https://epoch.ai/blog/what-will-ai-look-like-in-2030.Owen. What will AI look like in 2030?. Epoch AI https://epoch.ai/blog/what-will-ai-look-like-in-2030 (2025).Owen, “What will AI look like in 2030?”, Epoch AI. [Online]. Available: https://epoch.ai/blog/what-will-ai-look-like-in-2030
OWID(2025). Annual professional service robots installed globally, by application area. Our World in Data.OWID. (2025). Annual professional service robots installed globally, by application area. Our World in Data. https://ourworldindata.org/grapher/annual-professional-service-robots-installed-by-areaOWID. 2025. “Annual Professional Service Robots Installed Globally, by Application Area”. Our World in Data. https://ourworldindata.org/grapher/annual-professional-service-robots-installed-by-area.OWID. “Annual Professional Service Robots Installed Globally, by Application Area”. Our World in Data, 2025, https://ourworldindata.org/grapher/annual-professional-service-robots-installed-by-area.OWID. Annual professional service robots installed globally, by application area. Our World in Data https://ourworldindata.org/grapher/annual-professional-service-robots-installed-by-area (2025).OWID, “Annual professional service robots installed globally, by application area”, Our World in Data. [Online]. Available: https://ourworldindata.org/grapher/annual-professional-service-robots-installed-by-area
OWID(2025). New industrial robots installed per year. Our World in Data.OWID. (2025). New industrial robots installed per year. Our World in Data. https://ourworldindata.org/grapher/annual-industrial-robots-installedOWID. 2025. “New Industrial Robots Installed Per Year”. Our World in Data. https://ourworldindata.org/grapher/annual-industrial-robots-installed.OWID. “New Industrial Robots Installed Per Year”. Our World in Data, 2025, https://ourworldindata.org/grapher/annual-industrial-robots-installed.OWID. New industrial robots installed per year. Our World in Data https://ourworldindata.org/grapher/annual-industrial-robots-installed (2025).OWID, “New industrial robots installed per year”, Our World in Data. [Online]. Available: https://ourworldindata.org/grapher/annual-industrial-robots-installed
Oxford Reference(2016). Lord Kelvin. Oxford Reference.Oxford Reference. (2016). Lord Kelvin. Internet Archive (https://web.archive.org/web/20260420225540/https://www.oxfordreference.com/display/10.1093/acref/9780191826719.001.0001/q-oro-ed4-00006236). Oxford Reference. https://oxfordreference.com/display/10.1093/acref/9780191826719.001.0001/q-oro-ed4-00006236Oxford Reference. 2016. “Lord Kelvin”. Oxford Reference. Https://web.archive.org/web/20260420225540/https://www.oxfordreference.com/display/10.1093/acref/9780191826719.001.0001/q-oro-ed4-00006236. Internet Archive. https://oxfordreference.com/display/10.1093/acref/9780191826719.001.0001/q-oro-ed4-00006236.Oxford Reference. “Lord Kelvin”. Oxford Reference, 2016, Internet Archive, https://web.archive.org/web/20260420225540/https://www.oxfordreference.com/display/10.1093/acref/9780191826719.001.0001/q-oro-ed4-00006236, https://oxfordreference.com/display/10.1093/acref/9780191826719.001.0001/q-oro-ed4-00006236.Oxford Reference. Lord Kelvin. Oxford Reference https://oxfordreference.com/display/10.1093/acref/9780191826719.001.0001/q-oro-ed4-00006236 (2016).Oxford Reference, “Lord Kelvin”, Oxford Reference. Accessed: Apr. 20, 2026. [Online]. Available: https://oxfordreference.com/display/10.1093/acref/9780191826719.001.0001/q-oro-ed4-00006236
Oxford Union Debate(2024). Bevor Sie zu YouTube weitergehen. YouTube.Oxford Union Debate. (2024). Bevor Sie zu YouTube weitergehen [Video recording]. In YouTube. https://youtube.com/playlist?list=PLOAFgXcJkZ2wFf3mcJ0xIFpJQgEDI274JOxford Union Debate. 2024. “Bevor Sie Zu YouTube Weitergehen”. YouTube. https://youtube.com/playlist?list=PLOAFgXcJkZ2wFf3mcJ0xIFpJQgEDI274J.Oxford Union Debate. “Bevor Sie Zu YouTube Weitergehen”. YouTube, 2024, https://youtube.com/playlist?list=PLOAFgXcJkZ2wFf3mcJ0xIFpJQgEDI274J.Oxford Union Debate. Bevor Sie Zu YouTube Weitergehen. YouTube (2024).Oxford Union Debate, Bevor Sie zu YouTube weitergehen, (2024). [Online Video]. Available: https://youtube.com/playlist?list=PLOAFgXcJkZ2wFf3mcJ0xIFpJQgEDI274J
Pacchiardi et al.(2025). Is the Definition of AGI a Percentage?.Pacchiardi et al. (2025). Is the Definition of AGI a Percentage?. https://aievaluation.substack.com/p/is-the-definition-of-agi-a-percentagePacchiardi et al. 2025. Is the Definition of AGI a Percentage?. Edition. https://aievaluation.substack.com/p/is-the-definition-of-agi-a-percentage.Pacchiardi et al. Is the Definition of AGI a Percentage?. 2025, https://aievaluation.substack.com/p/is-the-definition-of-agi-a-percentage.Pacchiardi et al. Is the Definition of AGI a Percentage?. https://aievaluation.substack.com/p/is-the-definition-of-agi-a-percentage (2025).Pacchiardi et al., “Is the Definition of AGI a Percentage?”. [Online]. Available: https://aievaluation.substack.com/p/is-the-definition-of-agi-a-percentage
Pacchiardi, L. et al.(2023). How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions. arXiv.Pacchiardi, L., Chan, A. J., Mindermann, S., Moscovitz, I., Pan, A. Y., Gal, Y., Evans, O., & Brauner, J. (2023). How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions. In arXiv. https://arxiv.org/abs/2309.15840Pacchiardi, L., A. J. Chan, S. Mindermann, et al. 2023. “How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions”. In arXiv. Preprint, September 26. https://arxiv.org/abs/2309.15840.Pacchiardi, L., et al. “How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions”. arXiv, 26 Sept. 2023, https://arxiv.org/abs/2309.15840.Pacchiardi, L. et al. How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions. arXiv Preprint at https://arxiv.org/abs/2309.15840 (2023).L. Pacchiardi et al., “How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions”, Sep. 26, 2023. [Online]. Available: https://arxiv.org/abs/2309.15840
Pan et al.(2020). Privacy Risks of General-Purpose Language Models.Pan et al. (2020). Privacy Risks of General-Purpose Language Models. https://ieeexplore.ieee.org/document/9152761Pan et al. 2020. “Privacy Risks of General-Purpose Language Models”. https://ieeexplore.ieee.org/document/9152761.Pan et al. Privacy Risks of General-Purpose Language Models. 2020, https://ieeexplore.ieee.org/document/9152761.Pan et al. Privacy Risks of General-Purpose Language Models. https://ieeexplore.ieee.org/document/9152761 (2020).Pan et al., “Privacy Risks of General-Purpose Language Models”. [Online]. Available: https://ieeexplore.ieee.org/document/9152761
Pan, A. et al.(2023). Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark. arXiv.Pan, A., Chan, J. S., Zou, A., Li, N., Basart, S., Woodside, T., Ng, J., Zhang, H., Emmons, S., & Hendrycks, D. (2023). Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark. In arXiv. https://arxiv.org/abs/2304.03279Pan, A., J. S. Chan, A. Zou, et al. 2023. “Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark”. In arXiv. Preprint, April 6. https://arxiv.org/abs/2304.03279.Pan, A., et al. “Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark”. arXiv, 6 Apr. 2023, https://arxiv.org/abs/2304.03279.Pan, A. et al. Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark. arXiv Preprint at https://arxiv.org/abs/2304.03279 (2023).A. Pan et al., “Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark”, Apr. 06, 2023. [Online]. Available: https://arxiv.org/abs/2304.03279
Panel of Experts on Libya(2021). Letter, 8 Mar. 2021, from the Panel of Experts on Libya Established pursuant to Resolution 1973 (2011). United Nations Digital Library System.Panel of Experts on Libya. (2021). Letter, 8 Mar. 2021, from the Panel of Experts on Libya Established pursuant to Resolution 1973 (2011). Internet Archive (https://web.archive.org/web/20250702063934/https://digitallibrary.un.org/record/3905159?v=pdf). United Nations Digital Library System. https://www.digitallibrary.un.org/record/3905159?v=pdfPanel of Experts on Libya. 2021. “Letter, 8 Mar. 2021, from the Panel of Experts on Libya Established Pursuant to Resolution 1973 (2011)”. United Nations Digital Library System. Https://web.archive.org/web/20250702063934/https://digitallibrary.un.org/record/3905159?v=pdf. Internet Archive. https://www.digitallibrary.un.org/record/3905159?v=pdf.Panel of Experts on Libya. “Letter, 8 Mar. 2021, from the Panel of Experts on Libya Established Pursuant to Resolution 1973 (2011)”. United Nations Digital Library System, 2021, Internet Archive, https://web.archive.org/web/20250702063934/https://digitallibrary.un.org/record/3905159?v=pdf, https://www.digitallibrary.un.org/record/3905159?v=pdf.Panel of Experts on Libya. Letter, 8 Mar. 2021, from the Panel of Experts on Libya Established pursuant to Resolution 1973 (2011). United Nations Digital Library System https://www.digitallibrary.un.org/record/3905159?v=pdf (2021).Panel of Experts on Libya, “Letter, 8 Mar. 2021, from the Panel of Experts on Libya Established pursuant to Resolution 1973 (2011)”, United Nations Digital Library System. Accessed: Jul. 02, 2025. [Online]. Available: https://www.digitallibrary.un.org/record/3905159?v=pdf
Pang, R. Y. et al.(2021). QuALITY: Question Answering with Long Input Texts, Yes!. arXiv.Pang, R. Y., Parrish, A., Joshi, N., Nangia, N., Phang, J., Chen, A., Padmakumar, V., Ma, J., Thompson, J., He, H., & Bowman, S. R. (2021). QuALITY: Question Answering with Long Input Texts, Yes!. In arXiv. https://arxiv.org/abs/2112.08608Pang, R. Y., A. Parrish, N. Joshi, et al. 2021. “QuALITY: Question Answering with Long Input Texts, Yes!”. In arXiv. Preprint, December 16. https://arxiv.org/abs/2112.08608.Pang, R. Y., et al. “QuALITY: Question Answering with Long Input Texts, Yes!”. arXiv, 16 Dec. 2021, https://arxiv.org/abs/2112.08608.Pang, R. Y. et al. QuALITY: Question Answering with Long Input Texts, Yes!. arXiv Preprint at https://arxiv.org/abs/2112.08608 (2021).R. Y. Pang et al., “QuALITY: Question Answering with Long Input Texts, Yes!”, Dec. 16, 2021. [Online]. Available: https://arxiv.org/abs/2112.08608
Pannu, J., Gebauer, S., McKelvey Jr, G., Cicero, A. & Inglesby, T.(2024). AI could pose pandemic-scale biosecurity risks. Here’s how to make it safer. Nature.Pannu, J., Gebauer, S., McKelvey Jr, G., Cicero, A., & Inglesby, T. (2024). AI could pose pandemic-scale biosecurity risks. Here’s how to make it safer. Nature. https://doi.org/10.1038/d41586-024-03815-2Pannu, J., S. Gebauer, G. McKelvey Jr, A. Cicero, and T. Inglesby. 2024. “AI Could Pose Pandemic-scale Biosecurity Risks. Here’s How to Make It Safer”. Nature, ahead of print, November 21. https://doi.org/10.1038/d41586-024-03815-2.Pannu, J., et al. “AI Could Pose Pandemic-scale Biosecurity Risks. Here’s How to Make It Safer”. Nature, Nov. 2024, https://doi.org/10.1038/d41586-024-03815-2.Pannu, J., Gebauer, S., McKelvey Jr, G., Cicero, A. & Inglesby, T. AI could pose pandemic-scale biosecurity risks. Here’s how to make it safer. Nature https://doi.org/10.1038/d41586-024-03815-2 (2024) doi:10.1038/d41586-024-03815-2.J. Pannu, S. Gebauer, G. McKelvey Jr, A. Cicero, and T. Inglesby, “AI could pose pandemic-scale biosecurity risks. Here’s how to make it safer”, Nature, Nov. 2024, doi: 10.1038/d41586-024-03815-2.
Papagiannidis, E., Mikalef, P. & Conboy, K.(2025). Responsible artificial intelligence governance: A review and research framework. The Journal of Strategic Information Systems.Papagiannidis, E., Mikalef, P., & Conboy, K. (2025). Responsible artificial intelligence governance: A review and research framework. The Journal of Strategic Information Systems, 34(2), 101885. https://doi.org/10.1016/j.jsis.2024.101885Papagiannidis, E., P. Mikalef, and K. Conboy. 2025. “Responsible Artificial Intelligence Governance: A Review and Research Framework”. The Journal of Strategic Information Systems 34 (2): 101885. https://doi.org/10.1016/j.jsis.2024.101885.Papagiannidis, E., et al. “Responsible Artificial Intelligence Governance: A Review and Research Framework”. The Journal of Strategic Information Systems, vol. 34, no. 2, June 2025, p. 101885, https://doi.org/10.1016/j.jsis.2024.101885.Papagiannidis, E., Mikalef, P. & Conboy, K. Responsible artificial intelligence governance: A review and research framework. The Journal of Strategic Information Systems 34, 101885 (2025).E. Papagiannidis, P. Mikalef, and K. Conboy, “Responsible artificial intelligence governance: A review and research framework”, The Journal of Strategic Information Systems, vol. 34, no. 2, p. 101885, Jun. 2025, doi: 10.1016/j.jsis.2024.101885.
Papyshev, G. & Yarime, M.(2023). The state's role in governing artificial intelligence: development, control, and promotion through national strategies. Policy Design and Practice.Papyshev, G., & Yarime, M. (2023). The state's role in governing artificial intelligence: development, control, and promotion through national strategies. Policy Design and Practice, 6(1), 79–102. https://doi.org/10.1080/25741292.2022.2162252Papyshev, G., and M. Yarime. 2023. “The State's Role in Governing Artificial Intelligence: Development, Control, and Promotion Through National Strategies”. Policy Design and Practice 6 (1): 79–102. https://doi.org/10.1080/25741292.2022.2162252.Papyshev, G., and M. Yarime. “The State's Role in Governing Artificial Intelligence: Development, Control, and Promotion Through National Strategies”. Policy Design and Practice, vol. 6, no. 1, Jan. 2023, pp. 79–102, https://doi.org/10.1080/25741292.2022.2162252.Papyshev, G. & Yarime, M. The state's role in governing artificial intelligence: development, control, and promotion through national strategies. Policy Design and Practice 6, 79–102 (2023).G. Papyshev and M. Yarime, “The state's role in governing artificial intelligence: development, control, and promotion through national strategies”, Policy Design and Practice, vol. 6, no. 1, pp. 79–102, Jan. 2023, doi: 10.1080/25741292.2022.2162252.
Park, P. S., Goldstein, S., O'Gara, A., Chen, M. & Hendrycks, D.(2023). AI Deception: A Survey of Examples, Risks, and Potential Solutions. arXiv.Park, P. S., Goldstein, S., O'Gara, A., Chen, M., & Hendrycks, D. (2023). AI Deception: A Survey of Examples, Risks, and Potential Solutions. In arXiv. https://arxiv.org/abs/2308.14752Park, P. S., S. Goldstein, A. O'Gara, M. Chen, and D. Hendrycks. 2023. “AI Deception: A Survey of Examples, Risks, and Potential Solutions”. In arXiv. Preprint, August 28. https://arxiv.org/abs/2308.14752.Park, P. S., et al. “AI Deception: A Survey of Examples, Risks, and Potential Solutions”. arXiv, 28 Aug. 2023, https://arxiv.org/abs/2308.14752.Park, P. S., Goldstein, S., O'Gara, A., Chen, M. & Hendrycks, D. AI Deception: A Survey of Examples, Risks, and Potential Solutions. arXiv Preprint at https://arxiv.org/abs/2308.14752 (2023).P. S. Park, S. Goldstein, A. O'Gara, M. Chen, and D. Hendrycks, “AI Deception: A Survey of Examples, Risks, and Potential Solutions”, Aug. 28, 2023. [Online]. Available: https://arxiv.org/abs/2308.14752
Parrish, A. et al.(2022). Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions. arXiv.Parrish, A., Trivedi, H., Perez, E., Chen, A., Nangia, N., Phang, J., & Bowman, S. R. (2022). Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions. In arXiv. https://arxiv.org/abs/2204.05212Parrish, A., H. Trivedi, E. Perez, et al. 2022. “Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions”. In arXiv. Preprint, April 11. https://arxiv.org/abs/2204.05212.Parrish, A., et al. “Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions”. arXiv, 11 Apr. 2022, https://arxiv.org/abs/2204.05212.Parrish, A. et al. Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions. arXiv Preprint at https://arxiv.org/abs/2204.05212 (2022).A. Parrish et al., “Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions”, Apr. 11, 2022. [Online]. Available: https://arxiv.org/abs/2204.05212
Parrish, A. et al.(2022). Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions. arXiv.Parrish, A., Trivedi, H., Nangia, N., Padmakumar, V., Phang, J., Saimbhi, A. S., & Bowman, S. R. (2022). Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions. In arXiv. https://arxiv.org/abs/2210.10860Parrish, A., H. Trivedi, N. Nangia, et al. 2022. “Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions”. In arXiv. Preprint, October 19. https://arxiv.org/abs/2210.10860.Parrish, A., et al. “Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions”. arXiv, 19 Oct. 2022, https://arxiv.org/abs/2210.10860.Parrish, A. et al. Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions. arXiv Preprint at https://arxiv.org/abs/2210.10860 (2022).A. Parrish et al., “Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions”, Oct. 19, 2022. [Online]. Available: https://arxiv.org/abs/2210.10860
Patel(2023). Will scaling work?.Patel. (2023). Will scaling work?. https://dwarkesh.com/p/will-scaling-workPatel. 2023. “Will Scaling Work?”. https://dwarkesh.com/p/will-scaling-work.Patel. Will Scaling Work?. 2023, https://dwarkesh.com/p/will-scaling-work.Patel. Will scaling work?. https://dwarkesh.com/p/will-scaling-work (2023).Patel, “Will scaling work?”. [Online]. Available: https://dwarkesh.com/p/will-scaling-work
paulfchristiano(2018). Clarifying "AI Alignment". AI Alignment Forum.paulfchristiano. (2018, November 15). Clarifying "AI Alignment". AI Alignment Forum. https://alignmentforum.org/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignmentpaulfchristiano. 2018. “Clarifying "AI Alignment"”. AI Alignment Forum, November 15. https://alignmentforum.org/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment.paulfchristiano. “Clarifying "AI Alignment"”. AI Alignment Forum, 15 Nov. 2018, https://alignmentforum.org/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment.paulfchristiano. Clarifying "AI Alignment". AI Alignment Forum https://alignmentforum.org/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment (2018).paulfchristiano, “Clarifying "AI Alignment"”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment
paulfchristiano(2019). Directions and desiderata for AI alignment. AI Alignment Forum.paulfchristiano. (2019, January 13). Directions and desiderata for AI alignment. AI Alignment Forum. https://alignmentforum.org/posts/kphJvksj5TndGapuh/directions-and-desiderata-for-ai-alignmentpaulfchristiano. 2019. “Directions and Desiderata for AI Alignment”. AI Alignment Forum, January 13. https://alignmentforum.org/posts/kphJvksj5TndGapuh/directions-and-desiderata-for-ai-alignment.paulfchristiano. “Directions and Desiderata for AI Alignment”. AI Alignment Forum, 13 Jan. 2019, https://alignmentforum.org/posts/kphJvksj5TndGapuh/directions-and-desiderata-for-ai-alignment.paulfchristiano. Directions and desiderata for AI alignment. AI Alignment Forum https://alignmentforum.org/posts/kphJvksj5TndGapuh/directions-and-desiderata-for-ai-alignment (2019).paulfchristiano, “Directions and desiderata for AI alignment”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/kphJvksj5TndGapuh/directions-and-desiderata-for-ai-alignment
paulfchristiano(2019). The reward engineering problem. AI Alignment Forum.paulfchristiano. (2019, January 16). The reward engineering problem. AI Alignment Forum. https://alignmentforum.org/posts/4nZRzoGTqg8xy5rr8/the-reward-engineering-problempaulfchristiano. 2019. “The Reward Engineering Problem”. AI Alignment Forum, January 16. https://alignmentforum.org/posts/4nZRzoGTqg8xy5rr8/the-reward-engineering-problem.paulfchristiano. “The Reward Engineering Problem”. AI Alignment Forum, 16 Jan. 2019, https://alignmentforum.org/posts/4nZRzoGTqg8xy5rr8/the-reward-engineering-problem.paulfchristiano. The reward engineering problem. AI Alignment Forum https://alignmentforum.org/posts/4nZRzoGTqg8xy5rr8/the-reward-engineering-problem (2019).paulfchristiano, “The reward engineering problem”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/4nZRzoGTqg8xy5rr8/the-reward-engineering-problem
paulfchristiano(2019). What failure looks like. AI Alignment Forum.paulfchristiano. (2019, March 17). What failure looks like. AI Alignment Forum. https://alignmentforum.org/posts/HBxe6wdjxK239zajf/what-failure-looks-likepaulfchristiano. 2019. “What Failure Looks Like”. AI Alignment Forum, March 17. https://alignmentforum.org/posts/HBxe6wdjxK239zajf/what-failure-looks-like.paulfchristiano. “What Failure Looks Like”. AI Alignment Forum, 17 Mar. 2019, https://alignmentforum.org/posts/HBxe6wdjxK239zajf/what-failure-looks-like.paulfchristiano. What failure looks like. AI Alignment Forum https://alignmentforum.org/posts/HBxe6wdjxK239zajf/what-failure-looks-like (2019).paulfchristiano, “What failure looks like”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/HBxe6wdjxK239zajf/what-failure-looks-like
paulfchristiano(2022). Where I agree and disagree with Eliezer. AI Alignment Forum.paulfchristiano. (2022, June 19). Where I agree and disagree with Eliezer. AI Alignment Forum. https://alignmentforum.org/posts/CoZhXrhpQxpy9xw9y/where-i-agree-and-disagree-with-eliezerpaulfchristiano. 2022. “Where I Agree and Disagree with Eliezer”. AI Alignment Forum, June 19. https://alignmentforum.org/posts/CoZhXrhpQxpy9xw9y/where-i-agree-and-disagree-with-eliezer.paulfchristiano. “Where I Agree and Disagree with Eliezer”. AI Alignment Forum, 19 June 2022, https://alignmentforum.org/posts/CoZhXrhpQxpy9xw9y/where-i-agree-and-disagree-with-eliezer.paulfchristiano. Where I agree and disagree with Eliezer. AI Alignment Forum https://alignmentforum.org/posts/CoZhXrhpQxpy9xw9y/where-i-agree-and-disagree-with-eliezer (2022).paulfchristiano, “Where I agree and disagree with Eliezer”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/CoZhXrhpQxpy9xw9y/where-i-agree-and-disagree-with-eliezer
paulfchristiano(2023). Comment on “[Linkpost] Introducing Superalignment”. AI Alignment Forum.paulfchristiano. (2023, July 7). Comment on “[Linkpost] Introducing Superalignment”. AI Alignment Forum. https://alignmentforum.org/posts/Hna4aoMwr6Qx9rHBs/linkpost-introducing-superalignment?commentId=NsYXBdLY6edAXavsMpaulfchristiano. 2023. “Comment on “[Linkpost] Introducing Superalignment””. AI Alignment Forum, July 7. https://alignmentforum.org/posts/Hna4aoMwr6Qx9rHBs/linkpost-introducing-superalignment?commentId=NsYXBdLY6edAXavsM.paulfchristiano. “Comment on “[Linkpost] Introducing Superalignment””. AI Alignment Forum, 7 July 2023, https://alignmentforum.org/posts/Hna4aoMwr6Qx9rHBs/linkpost-introducing-superalignment?commentId=NsYXBdLY6edAXavsM.paulfchristiano. Comment on “[Linkpost] Introducing Superalignment”. AI Alignment Forum https://alignmentforum.org/posts/Hna4aoMwr6Qx9rHBs/linkpost-introducing-superalignment?commentId=NsYXBdLY6edAXavsM (2023).paulfchristiano, “Comment on “[Linkpost] Introducing Superalignment””, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/Hna4aoMwr6Qx9rHBs/linkpost-introducing-superalignment?commentId=NsYXBdLY6edAXavsM
paulfchristiano(2023). My views on “doom”. AI Alignment Forum.paulfchristiano. (2023, April 27). My views on “doom”. AI Alignment Forum. https://alignmentforum.org/posts/xWMqsvHapP3nwdSW8/my-views-on-doompaulfchristiano. 2023. “My Views on “doom””. AI Alignment Forum, April 27. https://alignmentforum.org/posts/xWMqsvHapP3nwdSW8/my-views-on-doom.paulfchristiano. “My Views on “doom””. AI Alignment Forum, 27 Apr. 2023, https://alignmentforum.org/posts/xWMqsvHapP3nwdSW8/my-views-on-doom.paulfchristiano. My views on “doom”. AI Alignment Forum https://alignmentforum.org/posts/xWMqsvHapP3nwdSW8/my-views-on-doom (2023).paulfchristiano, “My views on “doom””, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/xWMqsvHapP3nwdSW8/my-views-on-doom
paulfchristiano(2023). Thoughts on the impact of RLHF research. AI Alignment Forum.paulfchristiano. (2023, January 25). Thoughts on the impact of RLHF research. AI Alignment Forum. https://alignmentforum.org/posts/vwu4kegAEZTBtpT6p/thoughts-on-the-impact-of-rlhf-researchpaulfchristiano. 2023. “Thoughts on the Impact of RLHF Research”. AI Alignment Forum, January 25. https://alignmentforum.org/posts/vwu4kegAEZTBtpT6p/thoughts-on-the-impact-of-rlhf-research.paulfchristiano. “Thoughts on the Impact of RLHF Research”. AI Alignment Forum, 25 Jan. 2023, https://alignmentforum.org/posts/vwu4kegAEZTBtpT6p/thoughts-on-the-impact-of-rlhf-research.paulfchristiano. Thoughts on the impact of RLHF research. AI Alignment Forum https://alignmentforum.org/posts/vwu4kegAEZTBtpT6p/thoughts-on-the-impact-of-rlhf-research (2023).paulfchristiano, “Thoughts on the impact of RLHF research”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/vwu4kegAEZTBtpT6p/thoughts-on-the-impact-of-rlhf-research
PauseAI(2023). List of p(doom) values. PauseAI.PauseAI. (2023, December 18). List of p(doom) values. PauseAI. https://pauseai.info/pdoomPauseAI. 2023. “List of P(doom) Values”. PauseAI, December 18. https://pauseai.info/pdoom.PauseAI. “List of P(doom) Values”. PauseAI, 18 Dec. 2023, https://pauseai.info/pdoom.PauseAI. List of p(doom) values. PauseAI https://pauseai.info/pdoom (2023).PauseAI, “List of p(doom) values”, PauseAI. [Online]. Available: https://pauseai.info/pdoom
Pearson, A., Bruner, E. & Polly, P. D.(2023). Updated imaging and phylogenetic comparative methods reassess relative temporal lobe size in anthropoids and modern humans. American Journal of Biological Anthropology.Pearson, A., Bruner, E., & Polly, P. D. (2023). Updated imaging and phylogenetic comparative methods reassess relative temporal lobe size in anthropoids and modern humans. American Journal of Biological Anthropology. https://doi.org/10.1002/ajpa.24712Pearson, A., E. Bruner, and P. D. Polly. 2023. “Updated Imaging and Phylogenetic Comparative Methods Reassess Relative Temporal Lobe Size in Anthropoids and Modern Humans”. American Journal of Biological Anthropology, ahead of print, February 15. https://doi.org/10.1002/ajpa.24712.Pearson, A., et al. “Updated Imaging and Phylogenetic Comparative Methods Reassess Relative Temporal Lobe Size in Anthropoids and Modern Humans”. American Journal of Biological Anthropology, Feb. 2023, https://doi.org/10.1002/ajpa.24712.Pearson, A., Bruner, E. & Polly, P. D. Updated imaging and phylogenetic comparative methods reassess relative temporal lobe size in anthropoids and modern humans. American Journal of Biological Anthropology https://doi.org/10.1002/ajpa.24712 (2023) doi:10.1002/ajpa.24712.A. Pearson, E. Bruner, and P. D. Polly, “Updated imaging and phylogenetic comparative methods reassess relative temporal lobe size in anthropoids and modern humans”, American Journal of Biological Anthropology, Feb. 2023, doi: 10.1002/ajpa.24712.
Peng, L. & Shang, J.(2024). Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing. arXiv.Peng, L., & Shang, J. (2024). Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing. In arXiv. https://arxiv.org/abs/2405.07726Peng, L., and J. Shang. 2024. “Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing”. In arXiv. Preprint, May 13. https://arxiv.org/abs/2405.07726.Peng, L., and J. Shang. “Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing”. arXiv, 13 May 2024, https://arxiv.org/abs/2405.07726.Peng, L. & Shang, J. Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing. arXiv Preprint at https://arxiv.org/abs/2405.07726 (2024).L. Peng and J. Shang, “Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing”, May 13, 2024. [Online]. Available: https://arxiv.org/abs/2405.07726
Peng, Q., Chai, Y. & Li, X.(2024). HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization. arXiv.Peng, Q., Chai, Y., & Li, X. (2024). HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization. In arXiv. https://arxiv.org/abs/2402.16694Peng, Q., Y. Chai, and X. Li. 2024. “HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization”. In arXiv. Preprint, February 26. https://arxiv.org/abs/2402.16694.Peng, Q., et al. “HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization”. arXiv, 26 Feb. 2024, https://arxiv.org/abs/2402.16694.Peng, Q., Chai, Y. & Li, X. HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization. arXiv Preprint at https://arxiv.org/abs/2402.16694 (2024).Q. Peng, Y. Chai, and X. Li, “HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization”, Feb. 26, 2024. [Online]. Available: https://arxiv.org/abs/2402.16694
Peppin et al.(2024). The Reality of AI and Biorisk. arXiv.org.Peppin et al. (2024). The Reality of AI and Biorisk. arXiv.org. https://www.arxiv.org/abs/2412.01946Peppin et al. 2024. “The Reality of AI and Biorisk”. arXiv.org. https://www.arxiv.org/abs/2412.01946.Peppin et al. “The Reality of AI and Biorisk”. arXiv.org, 2024, https://www.arxiv.org/abs/2412.01946.Peppin et al. The Reality of AI and Biorisk. arXiv.org https://www.arxiv.org/abs/2412.01946 (2024).Peppin et al., “The Reality of AI and Biorisk”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2412.01946
Perez, E. et al.(2022). Red Teaming Language Models with Language Models. arXiv.Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., & Irving, G. (2022). Red Teaming Language Models with Language Models. In arXiv. https://arxiv.org/abs/2202.03286Perez, E., S. Huang, F. Song, et al. 2022. “Red Teaming Language Models with Language Models”. In arXiv. Preprint, February 7. https://arxiv.org/abs/2202.03286.Perez, E., et al. “Red Teaming Language Models with Language Models”. arXiv, 7 Feb. 2022, https://arxiv.org/abs/2202.03286.Perez, E. et al. Red Teaming Language Models with Language Models. arXiv Preprint at https://arxiv.org/abs/2202.03286 (2022).E. Perez et al., “Red Teaming Language Models with Language Models”, Feb. 07, 2022. [Online]. Available: https://arxiv.org/abs/2202.03286
Perlman(2024). AI Lab Watch.Perlman. (2024). AI Lab Watch. https://ailabwatch.org/blog/external-evaluationPerlman. 2024. “AI Lab Watch”. https://ailabwatch.org/blog/external-evaluation.Perlman. AI Lab Watch. 2024, https://ailabwatch.org/blog/external-evaluation.Perlman. AI Lab Watch. https://ailabwatch.org/blog/external-evaluation (2024).Perlman, “AI Lab Watch”. [Online]. Available: https://ailabwatch.org/blog/external-evaluation
Peter Cihon(2019). Standards for AI Governance: International Standards to Enable Global Coordination in AI Research & Development.Peter Cihon. (2019, April 17). Standards for AI Governance: International Standards to Enable Global Coordination in AI Research & Development. https://governance.ai/research-paper/standards-for-ai-governance-international-standards-to-enable-global-coordination-in-ai-research-developmentPeter Cihon. 2019. “Standards for AI Governance: International Standards to Enable Global Coordination in AI Research & Development”. April 17. https://governance.ai/research-paper/standards-for-ai-governance-international-standards-to-enable-global-coordination-in-ai-research-development.Peter Cihon. Standards for AI Governance: International Standards to Enable Global Coordination in AI Research & Development. 17 Apr. 2019, https://governance.ai/research-paper/standards-for-ai-governance-international-standards-to-enable-global-coordination-in-ai-research-development.Peter Cihon. Standards for AI Governance: International Standards to Enable Global Coordination in AI Research & Development. https://governance.ai/research-paper/standards-for-ai-governance-international-standards-to-enable-global-coordination-in-ai-research-development (2019).Peter Cihon, “Standards for AI Governance: International Standards to Enable Global Coordination in AI Research & Development”. [Online]. Available: https://governance.ai/research-paper/standards-for-ai-governance-international-standards-to-enable-global-coordination-in-ai-research-development
Petrie, J.(2024). Near-Term Enforcement of AI Chip Export Controls Using A Firmware-Based Design for Offline Licensing. arXiv.Petrie, J. (2024). Near-Term Enforcement of AI Chip Export Controls Using A Firmware-Based Design for Offline Licensing. In arXiv. https://arxiv.org/abs/2404.18308Petrie, J. 2024. “Near-Term Enforcement of AI Chip Export Controls Using A Firmware-Based Design for Offline Licensing”. In arXiv. Preprint, April 28. https://arxiv.org/abs/2404.18308.Petrie, J. “Near-Term Enforcement of AI Chip Export Controls Using A Firmware-Based Design for Offline Licensing”. arXiv, 28 Apr. 2024, https://arxiv.org/abs/2404.18308.Petrie, J. Near-Term Enforcement of AI Chip Export Controls Using A Firmware-Based Design for Offline Licensing. arXiv Preprint at https://arxiv.org/abs/2404.18308 (2024).J. Petrie, “Near-Term Enforcement of AI Chip Export Controls Using A Firmware-Based Design for Offline Licensing”, Apr. 28, 2024. [Online]. Available: https://arxiv.org/abs/2404.18308
Petropoulos et al.(2025). Building CERN for AI - An institutional blueprint. Centre for Future Generations.Petropoulos et al. (2025, January 30). Building CERN for AI - An institutional blueprint. Centre for Future Generations. https://cfg.eu/building-cern-for-aiPetropoulos et al. 2025. “Building CERN for AI - An Institutional Blueprint”. Centre for Future Generations, January 30. https://cfg.eu/building-cern-for-ai.Petropoulos et al. “Building CERN for AI - An Institutional Blueprint”. Centre for Future Generations, 30 Jan. 2025, https://cfg.eu/building-cern-for-ai.Petropoulos et al. Building CERN for AI - An institutional blueprint. Centre for Future Generations https://cfg.eu/building-cern-for-ai (2025).Petropoulos et al., “Building CERN for AI - An institutional blueprint”, Centre for Future Generations. [Online]. Available: https://cfg.eu/building-cern-for-ai
Phan, L. et al.(2025). Humanity's Last Exam. arXiv.Phan, L., Gatti, A., Han, Z., Li, N., Hu, J., Zhang, H., Zhang, C. B. C., Shaaban, M., Ling, J., Shi, S., Choi, M., Agrawal, A., Chopra, A., Khoja, A., Kim, R., Ren, R., Hausenloy, J., Zhang, O., Mazeika, M., … Hendrycks, D. (2025). Humanity's Last Exam. In arXiv. https://doi.org/10.1038/s41586-025-09962-4Phan, L., A. Gatti, Z. Han, et al. 2025. “Humanity's Last Exam”. In arXiv. Preprint, January 24. https://doi.org/10.1038/s41586-025-09962-4.Phan, L., et al. “Humanity's Last Exam”. arXiv, 24 Jan. 2025, https://doi.org/10.1038/s41586-025-09962-4.Phan, L. et al. Humanity's Last Exam. arXiv Preprint at https://doi.org/10.1038/s41586-025-09962-4 (2025).L. Phan et al., “Humanity's Last Exam”, Jan. 24, 2025. doi: 10.1038/s41586-025-09962-4.
Pilz, K. & Heim, L.(2023). Compute at Scale: A Broad Investigation into the Data Center Industry. arXiv.Pilz, K., & Heim, L. (2023). Compute at Scale: A Broad Investigation into the Data Center Industry. In arXiv. https://arxiv.org/abs/2311.02651Pilz, K., and L. Heim. 2023. “Compute at Scale: A Broad Investigation into the Data Center Industry”. In arXiv. Preprint, November 5. https://arxiv.org/abs/2311.02651.Pilz, K., and L. Heim. “Compute at Scale: A Broad Investigation into the Data Center Industry”. arXiv, 5 Nov. 2023, https://arxiv.org/abs/2311.02651.Pilz, K. & Heim, L. Compute at Scale: A Broad Investigation into the Data Center Industry. arXiv Preprint at https://arxiv.org/abs/2311.02651 (2023).K. Pilz and L. Heim, “Compute at Scale: A Broad Investigation into the Data Center Industry”, Nov. 05, 2023. [Online]. Available: https://arxiv.org/abs/2311.02651
Pilz, K. F., Sanders, J., Rahman, R. & Heim, L.(2025). Trends in AI Supercomputers. arXiv.Pilz, K. F., Sanders, J., Rahman, R., & Heim, L. (2025). Trends in AI Supercomputers. In arXiv. https://arxiv.org/abs/2504.16026Pilz, K. F., J. Sanders, R. Rahman, and L. Heim. 2025. “Trends in AI Supercomputers”. In arXiv. Preprint, April 22. https://arxiv.org/abs/2504.16026.Pilz, K. F., et al. “Trends in AI Supercomputers”. arXiv, 22 Apr. 2025, https://arxiv.org/abs/2504.16026.Pilz, K. F., Sanders, J., Rahman, R. & Heim, L. Trends in AI Supercomputers. arXiv Preprint at https://arxiv.org/abs/2504.16026 (2025).K. F. Pilz, J. Sanders, R. Rahman, and L. Heim, “Trends in AI Supercomputers”, Apr. 22, 2025. [Online]. Available: https://arxiv.org/abs/2504.16026
Piper(2023). Playing the training game.Piper. (2023). Playing the training game. https://www.planned-obsolescence.org/the-training-gamePiper. 2023. “Playing the Training Game”. https://www.planned-obsolescence.org/the-training-game.Piper. Playing the Training Game. 2023, https://www.planned-obsolescence.org/the-training-game.Piper. Playing the training game. https://www.planned-obsolescence.org/the-training-game (2023).Piper, “Playing the training game”. [Online]. Available: https://www.planned-obsolescence.org/the-training-game
Piper(2024). Should we make our most powerful AI models open source to all?. Vox.Piper. (2024, February 2). Should we make our most powerful AI models open source to all?. Vox. https://vox.com/future-perfect/2024/2/2/24058484/open-source-artificial-intelligence-ai-risk-meta-llama-2-chatgpt-openai-deepfakePiper. 2024. “Should We Make Our Most Powerful AI Models Open Source to All?”. Vox, February 2. https://vox.com/future-perfect/2024/2/2/24058484/open-source-artificial-intelligence-ai-risk-meta-llama-2-chatgpt-openai-deepfake.Piper. “Should We Make Our Most Powerful AI Models Open Source to All?”. Vox, 2 Feb. 2024, https://vox.com/future-perfect/2024/2/2/24058484/open-source-artificial-intelligence-ai-risk-meta-llama-2-chatgpt-openai-deepfake.Piper. Should we make our most powerful AI models open source to all?. Vox https://vox.com/future-perfect/2024/2/2/24058484/open-source-artificial-intelligence-ai-risk-meta-llama-2-chatgpt-openai-deepfake (2024).Piper, “Should we make our most powerful AI models open source to all?”, Vox. [Online]. Available: https://vox.com/future-perfect/2024/2/2/24058484/open-source-artificial-intelligence-ai-risk-meta-llama-2-chatgpt-openai-deepfake
Policy-Relevant Science & Technology(2023). Munk Debate on Artificial Intelligence | Bengio & Tegmark vs. Mitchell & LeCun. YouTube.Policy-Relevant Science & Technology. (2023). Munk Debate on Artificial Intelligence | Bengio & Tegmark vs. Mitchell & LeCun [Video recording]. In YouTube. https://www.youtube.com/watch?v=144uOfr4SYAPolicy-Relevant Science & Technology. 2023. “Munk Debate on Artificial Intelligence | Bengio & Tegmark Vs. Mitchell & LeCun”. YouTube. https://www.youtube.com/watch?v=144uOfr4SYA.Policy-Relevant Science & Technology. “Munk Debate on Artificial Intelligence | Bengio & Tegmark Vs. Mitchell & LeCun”. YouTube, 2023, https://www.youtube.com/watch?v=144uOfr4SYA.Policy-Relevant Science & Technology. Munk Debate on Artificial Intelligence | Bengio & Tegmark Vs. Mitchell & LeCun. YouTube (2023).Policy-Relevant Science & Technology, Munk Debate on Artificial Intelligence | Bengio & Tegmark vs. Mitchell & LeCun, (2023). [Online Video]. Available: https://www.youtube.com/watch?v=144uOfr4SYA
Power, A., Burda, Y., Edwards, H., Babuschkin, I. & Misra, V.(2022). Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets. arXiv.Power, A., Burda, Y., Edwards, H., Babuschkin, I., & Misra, V. (2022). Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets. In arXiv. https://arxiv.org/abs/2201.02177Power, A., Y. Burda, H. Edwards, I. Babuschkin, and V. Misra. 2022. “Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets”. In arXiv. Preprint, January 6. https://arxiv.org/abs/2201.02177.Power, A., et al. “Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets”. arXiv, 6 Jan. 2022, https://arxiv.org/abs/2201.02177.Power, A., Burda, Y., Edwards, H., Babuschkin, I. & Misra, V. Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets. arXiv Preprint at https://arxiv.org/abs/2201.02177 (2022).A. Power, Y. Burda, H. Edwards, I. Babuschkin, and V. Misra, “Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets”, Jan. 06, 2022. [Online]. Available: https://arxiv.org/abs/2201.02177
Prime Intellect Team et al.(2025). INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning. arXiv.Prime Intellect Team, Sami Jaghouar, Justus Mattern, Jack Min Ong, Jannik Straube, Manveer Basra, Aaron Pazdera, Kushal Thaman, Matthew Di Ferrante, Felix Gabriel, Fares Obeid, Kemal Erdem, Michael Keiblinger, & Johannes Hagemann. (2025). INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning. In arXiv. https://arxiv.org/abs/2505.07291Prime Intellect Team, Sami Jaghouar, Justus Mattern, et al. 2025. “INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning”. In arXiv. Preprint, May 12. https://arxiv.org/abs/2505.07291.Prime Intellect Team, et al. “INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning”. arXiv, 12 May 2025, https://arxiv.org/abs/2505.07291.Prime Intellect Team et al. INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning. arXiv Preprint at https://arxiv.org/abs/2505.07291 (2025).Prime Intellect Team et al., “INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning”, May 12, 2025. [Online]. Available: https://arxiv.org/abs/2505.07291
Purtova, N. & Maanen, G. V.(2022). Data as an economic good, data as a commons, and data governance. arXiv.Purtova, N., & van Maanen, G. (2022). Data as an economic good, data as a commons, and data governance. In arXiv. https://arxiv.org/abs/2212.10244Purtova, N., and G. van Maanen. 2022. “Data as an Economic Good, Data as a Commons, and Data Governance”. In arXiv. Preprint, December 20. https://arxiv.org/abs/2212.10244.Purtova, N., and G. van Maanen. “Data as an Economic Good, Data as a Commons, and Data Governance”. arXiv, 20 Dec. 2022, https://arxiv.org/abs/2212.10244.Purtova, N. & van Maanen, G. Data as an economic good, data as a commons, and data governance. arXiv Preprint at https://arxiv.org/abs/2212.10244 (2022).N. Purtova and G. van Maanen, “Data as an economic good, data as a commons, and data governance”, Dec. 20, 2022. [Online]. Available: https://arxiv.org/abs/2212.10244
Qin, Y. et al.(2023). ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv.Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., Zhao, S., Hong, L., Tian, R., Xie, R., Zhou, J., Gerstein, M., Li, D., Liu, Z., & Sun, M. (2023). ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. In arXiv. https://arxiv.org/abs/2307.16789Qin, Y., S. Liang, Y. Ye, et al. 2023. “ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs”. In arXiv. Preprint, July 31. https://arxiv.org/abs/2307.16789.Qin, Y., et al. “ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs”. arXiv, 31 July 2023, https://arxiv.org/abs/2307.16789.Qin, Y. et al. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv Preprint at https://arxiv.org/abs/2307.16789 (2023).Y. Qin et al., “ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs”, Jul. 31, 2023. [Online]. Available: https://arxiv.org/abs/2307.16789
Qin, Z., Zhao, W., Yu, X. & Sun, X.(2023). OpenVoice: Versatile Instant Voice Cloning. arXiv.Qin, Z., Zhao, W., Yu, X., & Sun, X. (2023). OpenVoice: Versatile Instant Voice Cloning. In arXiv. https://arxiv.org/abs/2312.01479Qin, Z., W. Zhao, X. Yu, and X. Sun. 2023. “OpenVoice: Versatile Instant Voice Cloning”. In arXiv. Preprint, December 3. https://arxiv.org/abs/2312.01479.Qin, Z., et al. “OpenVoice: Versatile Instant Voice Cloning”. arXiv, 3 Dec. 2023, https://arxiv.org/abs/2312.01479.Qin, Z., Zhao, W., Yu, X. & Sun, X. OpenVoice: Versatile Instant Voice Cloning. arXiv Preprint at https://arxiv.org/abs/2312.01479 (2023).Z. Qin, W. Zhao, X. Yu, and X. Sun, “OpenVoice: Versatile Instant Voice Cloning”, Dec. 03, 2023. [Online]. Available: https://arxiv.org/abs/2312.01479
Quentin FEUILLADE--MONTIXI & Pierre Peigné(2023). The Stochastic Parrot Hypothesis is debatable for the last generation of LLMs. LessWrong.Quentin FEUILLADE--MONTIXI, & Pierre Peigné. (2023, November 7). The Stochastic Parrot Hypothesis is debatable for the last generation of LLMs. LessWrong. https://lesswrong.com/posts/HxRjHq3QG8vcYy4yy/the-stochastic-parrot-hypothesis-is-debatable-for-the-lastQuentin FEUILLADE--MONTIXI, and Pierre Peigné. 2023. “The Stochastic Parrot Hypothesis Is Debatable for the Last Generation of LLMs”. LessWrong, November 7. https://lesswrong.com/posts/HxRjHq3QG8vcYy4yy/the-stochastic-parrot-hypothesis-is-debatable-for-the-last.Quentin FEUILLADE--MONTIXI, and Pierre Peigné. “The Stochastic Parrot Hypothesis Is Debatable for the Last Generation of LLMs”. LessWrong, 7 Nov. 2023, https://lesswrong.com/posts/HxRjHq3QG8vcYy4yy/the-stochastic-parrot-hypothesis-is-debatable-for-the-last.Quentin FEUILLADE--MONTIXI & Pierre Peigné. The Stochastic Parrot Hypothesis is debatable for the last generation of LLMs. LessWrong https://lesswrong.com/posts/HxRjHq3QG8vcYy4yy/the-stochastic-parrot-hypothesis-is-debatable-for-the-last (2023).Quentin FEUILLADE--MONTIXI and Pierre Peigné, “The Stochastic Parrot Hypothesis is debatable for the last generation of LLMs”, LessWrong. [Online]. Available: https://lesswrong.com/posts/HxRjHq3QG8vcYy4yy/the-stochastic-parrot-hypothesis-is-debatable-for-the-last
Quintin Pope(2023). My Objections to "We’re All Gonna Die with Eliezer Yudkowsky". AI Alignment Forum.Quintin Pope. (2023, March 21). My Objections to "We’re All Gonna Die with Eliezer Yudkowsky". AI Alignment Forum. https://alignmentforum.org/posts/wAczufCpMdaamF9fy/my-objections-to-we-re-all-gonna-die-with-eliezer-yudkowskyQuintin Pope. 2023. “My Objections to "We’re All Gonna Die with Eliezer Yudkowsky"”. AI Alignment Forum, March 21. https://alignmentforum.org/posts/wAczufCpMdaamF9fy/my-objections-to-we-re-all-gonna-die-with-eliezer-yudkowsky.Quintin Pope. “My Objections to "We’re All Gonna Die with Eliezer Yudkowsky"”. AI Alignment Forum, 21 Mar. 2023, https://alignmentforum.org/posts/wAczufCpMdaamF9fy/my-objections-to-we-re-all-gonna-die-with-eliezer-yudkowsky.Quintin Pope. My Objections to "We’re All Gonna Die with Eliezer Yudkowsky". AI Alignment Forum https://alignmentforum.org/posts/wAczufCpMdaamF9fy/my-objections-to-we-re-all-gonna-die-with-eliezer-yudkowsky (2023).Quintin Pope, “My Objections to "We’re All Gonna Die with Eliezer Yudkowsky"”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/wAczufCpMdaamF9fy/my-objections-to-we-re-all-gonna-die-with-eliezer-yudkowsky
Radhakrishnan, A. et al.(2023). Question Decomposition Improves the Faithfulness of Model-Generated Reasoning. arXiv.Radhakrishnan, A., Nguyen, K., Chen, A., Chen, C., Denison, C., Hernandez, D., Durmus, E., Hubinger, E., Kernion, J., Lukošiūtė, K., Cheng, N., Joseph, N., Schiefer, N., Rausch, O., McCandlish, S., Showk, S. E., Lanham, T., Maxwell, T., Chandrasekaran, V., … Perez, E. (2023). Question Decomposition Improves the Faithfulness of Model-Generated Reasoning. In arXiv. https://arxiv.org/abs/2307.11768Radhakrishnan, A., K. Nguyen, A. Chen, et al. 2023. “Question Decomposition Improves the Faithfulness of Model-Generated Reasoning”. In arXiv. Preprint, July 17. https://arxiv.org/abs/2307.11768.Radhakrishnan, A., et al. “Question Decomposition Improves the Faithfulness of Model-Generated Reasoning”. arXiv, 17 July 2023, https://arxiv.org/abs/2307.11768.Radhakrishnan, A. et al. Question Decomposition Improves the Faithfulness of Model-Generated Reasoning. arXiv Preprint at https://arxiv.org/abs/2307.11768 (2023).A. Radhakrishnan et al., “Question Decomposition Improves the Faithfulness of Model-Generated Reasoning”, Jul. 17, 2023. [Online]. Available: https://arxiv.org/abs/2307.11768
Rae, J. W. et al.(2021). Scaling Language Models: Methods, Analysis & Insights from Training Gopher. arXiv.Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., … Irving, G. (2021). Scaling Language Models: Methods, Analysis & Insights from Training Gopher. In arXiv. https://arxiv.org/abs/2112.11446Rae, J. W., S. Borgeaud, T. Cai, et al. 2021. “Scaling Language Models: Methods, Analysis & Insights from Training Gopher”. In arXiv. Preprint, December 8. https://arxiv.org/abs/2112.11446.Rae, J. W., et al. “Scaling Language Models: Methods, Analysis & Insights from Training Gopher”. arXiv, 8 Dec. 2021, https://arxiv.org/abs/2112.11446.Rae, J. W. et al. Scaling Language Models: Methods, Analysis & Insights from Training Gopher. arXiv Preprint at https://arxiv.org/abs/2112.11446 (2021).J. W. Rae et al., “Scaling Language Models: Methods, Analysis & Insights from Training Gopher”, Dec. 08, 2021. [Online]. Available: https://arxiv.org/abs/2112.11446
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D. & Finn, C.(2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv.org.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In arXiv.org. https://arxiv.org/abs/2305.18290Rafailov, R., A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn. 2023. “Direct Preference Optimization: Your Language Model Is Secretly a Reward Model”. In arXiv.org. Preprint, May 29. https://arxiv.org/abs/2305.18290.Rafailov, R., et al. “Direct Preference Optimization: Your Language Model Is Secretly a Reward Model”. arXiv.org, 29 May 2023, https://arxiv.org/abs/2305.18290.Rafailov, R. et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv.org Preprint at https://arxiv.org/abs/2305.18290 (2023).R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn, “Direct Preference Optimization: Your Language Model is Secretly a Reward Model”, May 29, 2023. [Online]. Available: https://arxiv.org/abs/2305.18290
Rahaman, N. et al.(2018). On the Spectral Bias of Neural Networks. arXiv.Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F. A., Bengio, Y., & Courville, A. (2018). On the Spectral Bias of Neural Networks. In arXiv. https://arxiv.org/abs/1806.08734Rahaman, N., A. Baratin, D. Arpit, et al. 2018. “On the Spectral Bias of Neural Networks”. In arXiv. Preprint, June 22. https://arxiv.org/abs/1806.08734.Rahaman, N., et al. “On the Spectral Bias of Neural Networks”. arXiv, 22 June 2018, https://arxiv.org/abs/1806.08734.Rahaman, N. et al. On the Spectral Bias of Neural Networks. arXiv Preprint at https://arxiv.org/abs/1806.08734 (2018).N. Rahaman et al., “On the Spectral Bias of Neural Networks”, Jun. 22, 2018. [Online]. Available: https://arxiv.org/abs/1806.08734
Raji, I. D. et al.(2020). Closing the AI accountability gap. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency.Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020, January 27). Closing the AI accountability gap. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. https://doi.org/10.1145/3351095.3372873Raji, I. D., A. Smart, R. N. White, et al. 2020. “Closing the AI Accountability Gap”. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, January 27. https://doi.org/10.1145/3351095.3372873.Raji, I. D., et al. “Closing the AI Accountability Gap”. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 2020, https://doi.org/10.1145/3351095.3372873.Raji, I. D. et al. Closing the AI accountability gap. in Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (2020). doi:10.1145/3351095.3372873.I. D. Raji et al., “Closing the AI accountability gap”, in Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, Jan. 2020. doi: 10.1145/3351095.3372873.
Ramana Kumar(2022). Will Capabilities Generalise More?. AI Alignment Forum.Ramana Kumar. (2022, June 29). Will Capabilities Generalise More?. AI Alignment Forum. https://alignmentforum.org/posts/cq5x4XDnLcBrYbb66/will-capabilities-generalise-moreRamana Kumar. 2022. “Will Capabilities Generalise More?”. AI Alignment Forum, June 29. https://alignmentforum.org/posts/cq5x4XDnLcBrYbb66/will-capabilities-generalise-more.Ramana Kumar. “Will Capabilities Generalise More?”. AI Alignment Forum, 29 June 2022, https://alignmentforum.org/posts/cq5x4XDnLcBrYbb66/will-capabilities-generalise-more.Ramana Kumar. Will Capabilities Generalise More?. AI Alignment Forum https://alignmentforum.org/posts/cq5x4XDnLcBrYbb66/will-capabilities-generalise-more (2022).Ramana Kumar, “Will Capabilities Generalise More?”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/cq5x4XDnLcBrYbb66/will-capabilities-generalise-more
Rational Animations(2025). How to Align AI: Put It in a Sandwich. YouTube.Rational Animations. (2025). How to Align AI: Put It in a Sandwich [Video recording]. In YouTube. https://www.youtube.com/watch?v=5mco9zAamRkRational Animations. 2025. “How to Align AI: Put It in a Sandwich”. YouTube. https://www.youtube.com/watch?v=5mco9zAamRk.Rational Animations. “How to Align AI: Put It in a Sandwich”. YouTube, 2025, https://www.youtube.com/watch?v=5mco9zAamRk.Rational Animations. How to Align AI: Put It in a Sandwich. YouTube (2025).Rational Animations, How to Align AI: Put It in a Sandwich, (2025). [Online Video]. Available: https://www.youtube.com/watch?v=5mco9zAamRk
Reed, S. et al.(2022). A Generalist Agent. arXiv.Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., & de Freitas, N. (2022). A Generalist Agent. In arXiv. https://arxiv.org/abs/2205.06175Reed, S., K. Zolna, E. Parisotto, et al. 2022. “A Generalist Agent”. In arXiv. Preprint, May 12. https://arxiv.org/abs/2205.06175.Reed, S., et al. “A Generalist Agent”. arXiv, 12 May 2022, https://arxiv.org/abs/2205.06175.Reed, S. et al. A Generalist Agent. arXiv Preprint at https://arxiv.org/abs/2205.06175 (2022).S. Reed et al., “A Generalist Agent”, May 12, 2022. [Online]. Available: https://arxiv.org/abs/2205.06175
Ren, R. et al.(2024). Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. arXiv.Ren, R., Basart, S., Khoja, A., Gatti, A., Phan, L., Yin, X., Mazeika, M., Pan, A., Mukobi, G., Kim, R. H., Fitz, S., & Hendrycks, D. (2024). Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. In arXiv. https://arxiv.org/abs/2407.21792Ren, R., S. Basart, A. Khoja, et al. 2024. “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”. In arXiv. Preprint, July 31. https://arxiv.org/abs/2407.21792.Ren, R., et al. “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”. arXiv, 31 July 2024, https://arxiv.org/abs/2407.21792.Ren, R. et al. Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. arXiv Preprint at https://arxiv.org/abs/2407.21792 (2024).R. Ren et al., “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”, Jul. 31, 2024. [Online]. Available: https://arxiv.org/abs/2407.21792
Ren, Y. & Sutherland, D. J.(2024). Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics. arXiv.Ren, Y., & Sutherland, D. J. (2024). Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics. In arXiv. https://arxiv.org/abs/2409.09626Ren, Y., and D. J. Sutherland. 2024. “Understanding Simplicity Bias Towards Compositional Mappings via Learning Dynamics”. In arXiv. Preprint, September 15. https://arxiv.org/abs/2409.09626.Ren, Y., and D. J. Sutherland. “Understanding Simplicity Bias Towards Compositional Mappings via Learning Dynamics”. arXiv, 15 Sept. 2024, https://arxiv.org/abs/2409.09626.Ren, Y. & Sutherland, D. J. Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics. arXiv Preprint at https://arxiv.org/abs/2409.09626 (2024).Y. Ren and D. J. Sutherland, “Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics”, Sep. 15, 2024. [Online]. Available: https://arxiv.org/abs/2409.09626
Reuel, A. et al.(2024). Open Problems in Technical AI Governance. arXiv.Reuel, A., Bucknall, B., Casper, S., Fist, T., Soder, L., Aarne, O., Hammond, L., Ibrahim, L., Chan, A., Wills, P., Anderljung, M., Garfinkel, B., Heim, L., Trask, A., Mukobi, G., Schaeffer, R., Baker, M., Hooker, S., Solaiman, I., … Trager, R. (2024). Open Problems in Technical AI Governance. In arXiv. https://arxiv.org/abs/2407.14981Reuel, A., B. Bucknall, S. Casper, et al. 2024. “Open Problems in Technical AI Governance”. In arXiv. Preprint, July 20. https://arxiv.org/abs/2407.14981.Reuel, A., et al. “Open Problems in Technical AI Governance”. arXiv, 20 July 2024, https://arxiv.org/abs/2407.14981.Reuel, A. et al. Open Problems in Technical AI Governance. arXiv Preprint at https://arxiv.org/abs/2407.14981 (2024).A. Reuel et al., “Open Problems in Technical AI Governance”, Jul. 20, 2024. [Online]. Available: https://arxiv.org/abs/2407.14981
Rich Sutton(2023). AI Succession. YouTube.Rich Sutton. (2023). AI Succession [Video recording]. In YouTube. https://www.youtube.com/watch?v=NgHFMolXs3URich Sutton. 2023. “AI Succession”. YouTube. https://www.youtube.com/watch?v=NgHFMolXs3U.Rich Sutton. “AI Succession”. YouTube, 2023, https://www.youtube.com/watch?v=NgHFMolXs3U.Rich Sutton. AI Succession. YouTube (2023).Rich Sutton, AI Succession, (2023). [Online Video]. Available: https://www.youtube.com/watch?v=NgHFMolXs3U
Richard_Ngo(2023). Clarifying and predicting AGI. LessWrong.Richard_Ngo. (2023, May 4). Clarifying and predicting AGI. LessWrong. https://lesswrong.com/posts/BoA3agdkAzL6HQtQP/clarifying-and-predicting-agiRichard_Ngo. 2023. “Clarifying and Predicting AGI”. LessWrong, May 4. https://lesswrong.com/posts/BoA3agdkAzL6HQtQP/clarifying-and-predicting-agi.Richard_Ngo. “Clarifying and Predicting AGI”. LessWrong, 4 May 2023, https://lesswrong.com/posts/BoA3agdkAzL6HQtQP/clarifying-and-predicting-agi.Richard_Ngo. Clarifying and predicting AGI. LessWrong https://lesswrong.com/posts/BoA3agdkAzL6HQtQP/clarifying-and-predicting-agi (2023).Richard_Ngo, “Clarifying and predicting AGI”, LessWrong. [Online]. Available: https://lesswrong.com/posts/BoA3agdkAzL6HQtQP/clarifying-and-predicting-agi
Rishub Tamirisa et al.(2024). Tamper-Resistant Safeguards for Open-Weight LLMs. arXiv.Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, & Mantas Mazeika. (2024). Tamper-Resistant Safeguards for Open-Weight LLMs. In arXiv. https://arxiv.org/abs/2408.00761Rishub Tamirisa, Bhrugu Bharathi, Long Phan, et al. 2024. “Tamper-Resistant Safeguards for Open-Weight LLMs”. In arXiv. Preprint, August 1. https://arxiv.org/abs/2408.00761.Rishub Tamirisa, et al. “Tamper-Resistant Safeguards for Open-Weight LLMs”. arXiv, 1 Aug. 2024, https://arxiv.org/abs/2408.00761.Rishub Tamirisa et al. Tamper-Resistant Safeguards for Open-Weight LLMs. arXiv Preprint at https://arxiv.org/abs/2408.00761 (2024).Rishub Tamirisa et al., “Tamper-Resistant Safeguards for Open-Weight LLMs”, Aug. 01, 2024. [Online]. Available: https://arxiv.org/abs/2408.00761
Rivera, J., Mukobi, G., Reuel, A., Lamparth, M., Smith, C. & Schneider, J.(2024). Escalation Risks from Language Models in Military and Diplomatic Decision-Making. arXiv.Rivera, J.-P., Mukobi, G., Reuel, A., Lamparth, M., Smith, C., & Schneider, J. (2024). Escalation Risks from Language Models in Military and Diplomatic Decision-Making. In arXiv. https://doi.org/10.1145/3630106.3658942Rivera, J.-P., G. Mukobi, A. Reuel, M. Lamparth, C. Smith, and J. Schneider. 2024. “Escalation Risks from Language Models in Military and Diplomatic Decision-Making”. In arXiv. Preprint, January 7. https://doi.org/10.1145/3630106.3658942.Rivera, J.-P., et al. “Escalation Risks from Language Models in Military and Diplomatic Decision-Making”. arXiv, 7 Jan. 2024, https://doi.org/10.1145/3630106.3658942.Rivera, J.-P. et al. Escalation Risks from Language Models in Military and Diplomatic Decision-Making. arXiv Preprint at https://doi.org/10.1145/3630106.3658942 (2024).J.-P. Rivera, G. Mukobi, A. Reuel, M. Lamparth, C. Smith, and J. Schneider, “Escalation Risks from Language Models in Military and Diplomatic Decision-Making”, Jan. 07, 2024. doi: 10.1145/3630106.3658942.
Rob Bensinger & Eliezer Yudkowsky(2022). A challenge for AGI organizations, and a challenge for readers. AI Alignment Forum.Rob Bensinger, & Eliezer Yudkowsky. (2022, December 1). A challenge for AGI organizations, and a challenge for readers. AI Alignment Forum. https://alignmentforum.org/posts/tD9zEiHfkvakpnNam/a-challenge-for-agi-organizations-and-a-challenge-for-1Rob Bensinger, and Eliezer Yudkowsky. 2022. “A Challenge for AGI Organizations, and a Challenge for Readers”. AI Alignment Forum, December 1. https://alignmentforum.org/posts/tD9zEiHfkvakpnNam/a-challenge-for-agi-organizations-and-a-challenge-for-1.Rob Bensinger, and Eliezer Yudkowsky. “A Challenge for AGI Organizations, and a Challenge for Readers”. AI Alignment Forum, 1 Dec. 2022, https://alignmentforum.org/posts/tD9zEiHfkvakpnNam/a-challenge-for-agi-organizations-and-a-challenge-for-1.Rob Bensinger & Eliezer Yudkowsky. A challenge for AGI organizations, and a challenge for readers. AI Alignment Forum https://alignmentforum.org/posts/tD9zEiHfkvakpnNam/a-challenge-for-agi-organizations-and-a-challenge-for-1 (2022).Rob Bensinger and Eliezer Yudkowsky, “A challenge for AGI organizations, and a challenge for readers”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/tD9zEiHfkvakpnNam/a-challenge-for-agi-organizations-and-a-challenge-for-1
Roberts, H., Hine, E., Taddeo, M. & Floridi, L.(2024). Global AI governance: barriers and pathways forward. International Affairs.Roberts, H., Hine, E., Taddeo, M., & Floridi, L. (2024). Global AI governance: barriers and pathways forward. International Affairs, 100(3), 1275–1286. https://doi.org/10.1093/ia/iiae073Roberts, H., E. Hine, M. Taddeo, and L. Floridi. 2024. “Global AI Governance: Barriers and Pathways Forward”. International Affairs 100 (3): 1275–86. https://doi.org/10.1093/ia/iiae073.Roberts, H., et al. “Global AI Governance: Barriers and Pathways Forward”. International Affairs, vol. 100, no. 3, May 2024, pp. 1275–86, https://doi.org/10.1093/ia/iiae073.Roberts, H., Hine, E., Taddeo, M. & Floridi, L. Global AI governance: barriers and pathways forward. International Affairs 100, 1275–1286 (2024).H. Roberts, E. Hine, M. Taddeo, and L. Floridi, “Global AI governance: barriers and pathways forward”, International Affairs, vol. 100, no. 3, pp. 1275–1286, May 2024, doi: 10.1093/ia/iiae073.
Robotics team(2024). Shaping the future of advanced robotics.Robotics team. (2024, January 4). Shaping the future of advanced robotics. https://deepmind.google/blog/shaping-the-future-of-advanced-roboticsRobotics team. 2024. “Shaping the Future of Advanced Robotics”. January 4. https://deepmind.google/blog/shaping-the-future-of-advanced-robotics.Robotics team. Shaping the Future of Advanced Robotics. 4 Jan. 2024, https://deepmind.google/blog/shaping-the-future-of-advanced-robotics.Robotics team. Shaping the future of advanced robotics. https://deepmind.google/blog/shaping-the-future-of-advanced-robotics (2024).Robotics team, “Shaping the future of advanced robotics”. [Online]. Available: https://deepmind.google/blog/shaping-the-future-of-advanced-robotics
Roger, F.(2023). Large Language Models Sometimes Generate Purely Negatively-Reinforced Text. arXiv.Roger, F. (2023). Large Language Models Sometimes Generate Purely Negatively-Reinforced Text. In arXiv. https://arxiv.org/abs/2306.07567Roger, F. 2023. “Large Language Models Sometimes Generate Purely Negatively-Reinforced Text”. In arXiv. Preprint, June 13. https://arxiv.org/abs/2306.07567.Roger, F. “Large Language Models Sometimes Generate Purely Negatively-Reinforced Text”. arXiv, 13 June 2023, https://arxiv.org/abs/2306.07567.Roger, F. Large Language Models Sometimes Generate Purely Negatively-Reinforced Text. arXiv Preprint at https://arxiv.org/abs/2306.07567 (2023).F. Roger, “Large Language Models Sometimes Generate Purely Negatively-Reinforced Text”, Jun. 13, 2023. [Online]. Available: https://arxiv.org/abs/2306.07567
Rolls, E., Burton, M. & Mora, F.(1980). Neurophysiological analysis of brain-stimulation reward in the monkey. Brain Research.Rolls, E. T., Burton, M. J., & Mora, F. (1980). Neurophysiological analysis of brain-stimulation reward in the monkey. Brain Research. https://doi.org/10.1016/0006-8993(80)91216-0Rolls, E. T., M. J. Burton, and F. Mora. 1980. “Neurophysiological Analysis of Brain-stimulation Reward in the Monkey”. Brain Research, ahead of print, August. https://doi.org/10.1016/0006-8993(80)91216-0.Rolls, E. T., et al. “Neurophysiological Analysis of Brain-stimulation Reward in the Monkey”. Brain Research, Aug. 1980, https://doi.org/10.1016/0006-8993(80)91216-0.Rolls, E. T., Burton, M. J. & Mora, F. Neurophysiological analysis of brain-stimulation reward in the monkey. Brain Research https://doi.org/10.1016/0006-8993(80)91216-0 (1980) doi:10.1016/0006-8993(80)91216-0.E. T. Rolls, M. J. Burton, and F. Mora, “Neurophysiological analysis of brain-stimulation reward in the monkey”, Brain Research, Aug. 1980, doi: 10.1016/0006-8993(80)91216-0.
Rudner et al.(2021). Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology.Rudner et al. (2021). Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology. Center for Security and Emerging Technology. https://cset.georgetown.edu/publication/key-concepts-in-ai-safety-an-overviewRudner et al. 2021. “Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology”. Center for Security and Emerging Technology. https://cset.georgetown.edu/publication/key-concepts-in-ai-safety-an-overview.Rudner et al. “Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology”. Center for Security and Emerging Technology, 2021, https://cset.georgetown.edu/publication/key-concepts-in-ai-safety-an-overview.Rudner et al. Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology. Center for Security and Emerging Technology https://cset.georgetown.edu/publication/key-concepts-in-ai-safety-an-overview (2021).Rudner et al., “Key Concepts in AI Safety: An Overview | Center for Security and Emerging Technology”, Center for Security and Emerging Technology. [Online]. Available: https://cset.georgetown.edu/publication/key-concepts-in-ai-safety-an-overview
Rudolf Laine et al.(2024). Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs. arXiv.Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan, Jeremy Scheurer, Mikita Balesni, Marius Hobbhahn, Alexander Meinke, & Owain Evans. (2024). Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs. In arXiv. https://arxiv.org/abs/2407.04694Rudolf Laine, Bilal Chughtai, Jan Betley, et al. 2024. “Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs”. In arXiv. Preprint, July 5. https://arxiv.org/abs/2407.04694.Rudolf Laine, et al. “Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs”. arXiv, 5 July 2024, https://arxiv.org/abs/2407.04694.Rudolf Laine et al. Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs. arXiv Preprint at https://arxiv.org/abs/2407.04694 (2024).Rudolf Laine et al., “Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs”, Jul. 05, 2024. [Online]. Available: https://arxiv.org/abs/2407.04694
Russel & Norvig(1994). Artificial Intelligence: A Modern Approach, 4th US ed.Russel & Norvig. (1994). Artificial Intelligence: A Modern Approach, 4th US ed. https://aima.cs.berkeley.eduRussel & Norvig. 1994. “Artificial Intelligence: A Modern Approach, 4th US Ed.”. https://aima.cs.berkeley.edu.Russel & Norvig. Artificial Intelligence: A Modern Approach, 4th US Ed. 1994, https://aima.cs.berkeley.edu.Russel & Norvig. Artificial Intelligence: A Modern Approach, 4th US ed. https://aima.cs.berkeley.edu (1994).Russel & Norvig, “Artificial Intelligence: A Modern Approach, 4th US ed.”. [Online]. Available: https://aima.cs.berkeley.edu
Russomanno, A., Fava, M. & Heyl, M.(2020). Quantum chaos and ensemble inequivalence of quantum long-range Ising chains. arXiv.Russomanno, A., Fava, M., & Heyl, M. (2020). Quantum chaos and ensemble inequivalence of quantum long-range Ising chains. In arXiv. https://doi.org/10.1103/PhysRevB.104.094309Russomanno, A., M. Fava, and M. Heyl. 2020. “Quantum Chaos and Ensemble Inequivalence of Quantum Long-range Ising Chains”. In arXiv. Preprint, December 11. https://doi.org/10.1103/PhysRevB.104.094309.Russomanno, A., et al. “Quantum Chaos and Ensemble Inequivalence of Quantum Long-range Ising Chains”. arXiv, 11 Dec. 2020, https://doi.org/10.1103/PhysRevB.104.094309.Russomanno, A., Fava, M. & Heyl, M. Quantum chaos and ensemble inequivalence of quantum long-range Ising chains. arXiv Preprint at https://doi.org/10.1103/PhysRevB.104.094309 (2020).A. Russomanno, M. Fava, and M. Heyl, “Quantum chaos and ensemble inequivalence of quantum long-range Ising chains”, Dec. 11, 2020. doi: 10.1103/PhysRevB.104.094309.
ryan_greenblatt & Buck(2024). Catching AIs red-handed. AI Alignment Forum.ryan_greenblatt, & Buck. (2024, January 5). Catching AIs red-handed. AI Alignment Forum. https://alignmentforum.org/posts/i2nmBfCXnadeGmhzW/catching-ais-red-handedryan_greenblatt, and Buck. 2024. “Catching AIs Red-handed”. AI Alignment Forum, January 5. https://alignmentforum.org/posts/i2nmBfCXnadeGmhzW/catching-ais-red-handed.ryan_greenblatt, and Buck. “Catching AIs Red-handed”. AI Alignment Forum, 5 Jan. 2024, https://alignmentforum.org/posts/i2nmBfCXnadeGmhzW/catching-ais-red-handed.ryan_greenblatt & Buck. Catching AIs red-handed. AI Alignment Forum https://alignmentforum.org/posts/i2nmBfCXnadeGmhzW/catching-ais-red-handed (2024).ryan_greenblatt and Buck, “Catching AIs red-handed”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/i2nmBfCXnadeGmhzW/catching-ais-red-handed
ryan_greenblatt & Buck(2024). The case for ensuring that powerful AIs are controlled. AI Alignment Forum.ryan_greenblatt, & Buck. (2024, January 24). The case for ensuring that powerful AIs are controlled. AI Alignment Forum. https://alignmentforum.org/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlledryan_greenblatt, and Buck. 2024. “The Case for Ensuring That Powerful AIs Are Controlled”. AI Alignment Forum, January 24. https://alignmentforum.org/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled.ryan_greenblatt, and Buck. “The Case for Ensuring That Powerful AIs Are Controlled”. AI Alignment Forum, 24 Jan. 2024, https://alignmentforum.org/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled.ryan_greenblatt & Buck. The case for ensuring that powerful AIs are controlled. AI Alignment Forum https://alignmentforum.org/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled (2024).ryan_greenblatt and Buck, “The case for ensuring that powerful AIs are controlled”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled
ryan_greenblatt & Fabien Roger(2023). Auditing failures vs concentrated failures. AI Alignment Forum.ryan_greenblatt, & Fabien Roger. (2023, December 11). Auditing failures vs concentrated failures. AI Alignment Forum. https://alignmentforum.org/posts/hirhSqvEAq7pdnyPG/auditing-failures-vs-concentrated-failuresryan_greenblatt, and Fabien Roger. 2023. “Auditing Failures Vs Concentrated Failures”. AI Alignment Forum, December 11. https://alignmentforum.org/posts/hirhSqvEAq7pdnyPG/auditing-failures-vs-concentrated-failures.ryan_greenblatt, and Fabien Roger. “Auditing Failures Vs Concentrated Failures”. AI Alignment Forum, 11 Dec. 2023, https://alignmentforum.org/posts/hirhSqvEAq7pdnyPG/auditing-failures-vs-concentrated-failures.ryan_greenblatt & Fabien Roger. Auditing failures vs concentrated failures. AI Alignment Forum https://alignmentforum.org/posts/hirhSqvEAq7pdnyPG/auditing-failures-vs-concentrated-failures (2023).ryan_greenblatt and Fabien Roger, “Auditing failures vs concentrated failures”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/hirhSqvEAq7pdnyPG/auditing-failures-vs-concentrated-failures
ryan_greenblatt(2025). AI companies are unlikely to make high-assurance safety cases if timelines are short. LessWrong.ryan_greenblatt. (2025, January 23). AI companies are unlikely to make high-assurance safety cases if timelines are short. LessWrong. https://lesswrong.com/posts/neTbrpBziAsTH5Bn7/ai-companies-are-unlikely-to-make-high-assurance-safetyryan_greenblatt. 2025. “AI Companies Are Unlikely to Make High-assurance Safety Cases If Timelines Are Short”. LessWrong, January 23. https://lesswrong.com/posts/neTbrpBziAsTH5Bn7/ai-companies-are-unlikely-to-make-high-assurance-safety.ryan_greenblatt. “AI Companies Are Unlikely to Make High-assurance Safety Cases If Timelines Are Short”. LessWrong, 23 Jan. 2025, https://lesswrong.com/posts/neTbrpBziAsTH5Bn7/ai-companies-are-unlikely-to-make-high-assurance-safety.ryan_greenblatt. AI companies are unlikely to make high-assurance safety cases if timelines are short. LessWrong https://lesswrong.com/posts/neTbrpBziAsTH5Bn7/ai-companies-are-unlikely-to-make-high-assurance-safety (2025).ryan_greenblatt, “AI companies are unlikely to make high-assurance safety cases if timelines are short”, LessWrong. [Online]. Available: https://lesswrong.com/posts/neTbrpBziAsTH5Bn7/ai-companies-are-unlikely-to-make-high-assurance-safety
ryan_greenblatt(2025). An overview of control measures. AI Alignment Forum.ryan_greenblatt. (2025, March 24). An overview of control measures. AI Alignment Forum. https://alignmentforum.org/s/WCJtsn6fNib6L7ZBB/p/G8WwLmcGFa4H6Ld9dryan_greenblatt. 2025. “An Overview of Control Measures”. AI Alignment Forum, March 24. https://alignmentforum.org/s/WCJtsn6fNib6L7ZBB/p/G8WwLmcGFa4H6Ld9d.ryan_greenblatt. “An Overview of Control Measures”. AI Alignment Forum, 24 Mar. 2025, https://alignmentforum.org/s/WCJtsn6fNib6L7ZBB/p/G8WwLmcGFa4H6Ld9d.ryan_greenblatt. An overview of control measures. AI Alignment Forum https://alignmentforum.org/s/WCJtsn6fNib6L7ZBB/p/G8WwLmcGFa4H6Ld9d (2025).ryan_greenblatt, “An overview of control measures”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/s/WCJtsn6fNib6L7ZBB/p/G8WwLmcGFa4H6Ld9d
ryan_greenblatt(2025). How will we update about scheming?. LessWrong.ryan_greenblatt. (2025, January 6). How will we update about scheming?. LessWrong. https://lesswrong.com/posts/aEguDPoCzt3287CCD/how-will-we-update-about-schemingryan_greenblatt. 2025. “How Will We Update About Scheming?”. LessWrong, January 6. https://lesswrong.com/posts/aEguDPoCzt3287CCD/how-will-we-update-about-scheming.ryan_greenblatt. “How Will We Update About Scheming?”. LessWrong, 6 Jan. 2025, https://lesswrong.com/posts/aEguDPoCzt3287CCD/how-will-we-update-about-scheming.ryan_greenblatt. How will we update about scheming?. LessWrong https://lesswrong.com/posts/aEguDPoCzt3287CCD/how-will-we-update-about-scheming (2025).ryan_greenblatt, “How will we update about scheming?”, LessWrong. [Online]. Available: https://lesswrong.com/posts/aEguDPoCzt3287CCD/how-will-we-update-about-scheming
ryan_greenblatt(2025). Prioritizing threats for AI control. AI Alignment Forum.ryan_greenblatt. (2025, March 19). Prioritizing threats for AI control. AI Alignment Forum. https://alignmentforum.org/s/WCJtsn6fNib6L7ZBB/p/fCazYoZSSMadiT6sfryan_greenblatt. 2025. “Prioritizing Threats for AI Control”. AI Alignment Forum, March 19. https://alignmentforum.org/s/WCJtsn6fNib6L7ZBB/p/fCazYoZSSMadiT6sf.ryan_greenblatt. “Prioritizing Threats for AI Control”. AI Alignment Forum, 19 Mar. 2025, https://alignmentforum.org/s/WCJtsn6fNib6L7ZBB/p/fCazYoZSSMadiT6sf.ryan_greenblatt. Prioritizing threats for AI control. AI Alignment Forum https://alignmentforum.org/s/WCJtsn6fNib6L7ZBB/p/fCazYoZSSMadiT6sf (2025).ryan_greenblatt, “Prioritizing threats for AI control”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/s/WCJtsn6fNib6L7ZBB/p/fCazYoZSSMadiT6sf
SakanaAI(2025). Sakana AI (@SakanaAILabs) on X. X (formerly Twitter).SakanaAI. (2025, February 21). Sakana AI (@SakanaAILabs) on X. X (formerly Twitter). https://x.com/SakanaAILabs/status/1892992938013270019SakanaAI. 2025. “Sakana AI (@SakanaAILabs) on X”. X (formerly Twitter), February 21. https://x.com/SakanaAILabs/status/1892992938013270019.SakanaAI. “Sakana AI (@SakanaAILabs) on X”. X (formerly Twitter), 21 Feb. 2025, https://x.com/SakanaAILabs/status/1892992938013270019.SakanaAI. Sakana AI (@SakanaAILabs) on X. X (formerly Twitter) https://x.com/SakanaAILabs/status/1892992938013270019 (2025).SakanaAI, “Sakana AI (@SakanaAILabs) on X”, X (formerly Twitter). [Online]. Available: https://x.com/SakanaAILabs/status/1892992938013270019
Sam Bowman(2024). The Checklist: What Succeeding at AI Safety Will Involve. AI Alignment Forum.Sam Bowman. (2024, September 3). The Checklist: What Succeeding at AI Safety Will Involve. AI Alignment Forum. https://alignmentforum.org/posts/mGCcZnr4WjGjqzX5s/the-checklist-what-succeeding-at-ai-safety-will-involveSam Bowman. 2024. “The Checklist: What Succeeding at AI Safety Will Involve”. AI Alignment Forum, September 3. https://alignmentforum.org/posts/mGCcZnr4WjGjqzX5s/the-checklist-what-succeeding-at-ai-safety-will-involve.Sam Bowman. “The Checklist: What Succeeding at AI Safety Will Involve”. AI Alignment Forum, 3 Sept. 2024, https://alignmentforum.org/posts/mGCcZnr4WjGjqzX5s/the-checklist-what-succeeding-at-ai-safety-will-involve.Sam Bowman. The Checklist: What Succeeding at AI Safety Will Involve. AI Alignment Forum https://alignmentforum.org/posts/mGCcZnr4WjGjqzX5s/the-checklist-what-succeeding-at-ai-safety-will-involve (2024).Sam Bowman, “The Checklist: What Succeeding at AI Safety Will Involve”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/mGCcZnr4WjGjqzX5s/the-checklist-what-succeeding-at-ai-safety-will-involve
Sam Bowman(2025). Putting up Bumpers. AI Alignment Forum.Sam Bowman. (2025, April 23). Putting up Bumpers. AI Alignment Forum. https://alignmentforum.org/posts/HXJXPjzWyS5aAoRCw/putting-up-bumpersSam Bowman. 2025. “Putting up Bumpers”. AI Alignment Forum, April 23. https://alignmentforum.org/posts/HXJXPjzWyS5aAoRCw/putting-up-bumpers.Sam Bowman. “Putting up Bumpers”. AI Alignment Forum, 23 Apr. 2025, https://alignmentforum.org/posts/HXJXPjzWyS5aAoRCw/putting-up-bumpers.Sam Bowman. Putting up Bumpers. AI Alignment Forum https://alignmentforum.org/posts/HXJXPjzWyS5aAoRCw/putting-up-bumpers (2025).Sam Bowman, “Putting up Bumpers”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/HXJXPjzWyS5aAoRCw/putting-up-bumpers
Sam Ringer(2022). Models Don't "Get Reward". AI Alignment Forum.Sam Ringer. (2022, December 30). Models Don't "Get Reward". AI Alignment Forum. https://alignmentforum.org/posts/TWorNr22hhYegE4RT/models-don-t-get-rewardSam Ringer. 2022. “Models Don't "Get Reward"”. AI Alignment Forum, December 30. https://alignmentforum.org/posts/TWorNr22hhYegE4RT/models-don-t-get-reward.Sam Ringer. “Models Don't "Get Reward"”. AI Alignment Forum, 30 Dec. 2022, https://alignmentforum.org/posts/TWorNr22hhYegE4RT/models-don-t-get-reward.Sam Ringer. Models Don't "Get Reward". AI Alignment Forum https://alignmentforum.org/posts/TWorNr22hhYegE4RT/models-don-t-get-reward (2022).Sam Ringer, “Models Don't "Get Reward"”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/TWorNr22hhYegE4RT/models-don-t-get-reward
Sammy Martin & Daniel_Eth(2021). Takeoff Speeds and Discontinuities. AI Alignment Forum.Sammy Martin, & Daniel_Eth. (2021, September 30). Takeoff Speeds and Discontinuities. AI Alignment Forum. https://alignmentforum.org/posts/pGXR2ynhe5bBCCNqn/takeoff-speeds-and-discontinuitiesSammy Martin, and Daniel_Eth. 2021. “Takeoff Speeds and Discontinuities”. AI Alignment Forum, September 30. https://alignmentforum.org/posts/pGXR2ynhe5bBCCNqn/takeoff-speeds-and-discontinuities.Sammy Martin, and Daniel_Eth. “Takeoff Speeds and Discontinuities”. AI Alignment Forum, 30 Sept. 2021, https://alignmentforum.org/posts/pGXR2ynhe5bBCCNqn/takeoff-speeds-and-discontinuities.Sammy Martin & Daniel_Eth. Takeoff Speeds and Discontinuities. AI Alignment Forum https://alignmentforum.org/posts/pGXR2ynhe5bBCCNqn/takeoff-speeds-and-discontinuities (2021).Sammy Martin and Daniel_Eth, “Takeoff Speeds and Discontinuities”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/pGXR2ynhe5bBCCNqn/takeoff-speeds-and-discontinuities
Sandoval-Segura, P., Singla, V., Geiping, J., Goldblum, M., Goldstein, T. & Jacobs, D. W.(2022). Autoregressive Perturbations for Data Poisoning. arXiv.Sandoval-Segura, P., Singla, V., Geiping, J., Goldblum, M., Goldstein, T., & Jacobs, D. W. (2022). Autoregressive Perturbations for Data Poisoning. In arXiv. https://arxiv.org/abs/2206.03693Sandoval-Segura, P., V. Singla, J. Geiping, M. Goldblum, T. Goldstein, and D. W. Jacobs. 2022. “Autoregressive Perturbations for Data Poisoning”. In arXiv. Preprint, June 8. https://arxiv.org/abs/2206.03693.Sandoval-Segura, P., et al. “Autoregressive Perturbations for Data Poisoning”. arXiv, 8 June 2022, https://arxiv.org/abs/2206.03693.Sandoval-Segura, P. et al. Autoregressive Perturbations for Data Poisoning. arXiv Preprint at https://arxiv.org/abs/2206.03693 (2022).P. Sandoval-Segura, V. Singla, J. Geiping, M. Goldblum, T. Goldstein, and D. W. Jacobs, “Autoregressive Perturbations for Data Poisoning”, Jun. 08, 2022. [Online]. Available: https://arxiv.org/abs/2206.03693
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P. & Hashimoto, T.(2023). Whose Opinions Do Language Models Reflect?. arXiv.Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., & Hashimoto, T. (2023). Whose Opinions Do Language Models Reflect?. In arXiv. https://arxiv.org/abs/2303.17548Santurkar, S., E. Durmus, F. Ladhak, C. Lee, P. Liang, and T. Hashimoto. 2023. “Whose Opinions Do Language Models Reflect?”. In arXiv. Preprint, March 30. https://arxiv.org/abs/2303.17548.Santurkar, S., et al. “Whose Opinions Do Language Models Reflect?”. arXiv, 30 Mar. 2023, https://arxiv.org/abs/2303.17548.Santurkar, S. et al. Whose Opinions Do Language Models Reflect?. arXiv Preprint at https://arxiv.org/abs/2303.17548 (2023).S. Santurkar, E. Durmus, F. Ladhak, C. Lee, P. Liang, and T. Hashimoto, “Whose Opinions Do Language Models Reflect?”, Mar. 30, 2023. [Online]. Available: https://arxiv.org/abs/2303.17548
Sastry, G. et al.(2024). Computing Power and the Governance of Artificial Intelligence. arXiv.Sastry, G., Heim, L., Belfield, H., Anderljung, M., Brundage, M., Hazell, J., O'Keefe, C., Hadfield, G. K., Ngo, R., Pilz, K., Gor, G., Bluemke, E., Shoker, S., Egan, J., Trager, R. F., Avin, S., Weller, A., Bengio, Y., & Coyle, D. (2024). Computing Power and the Governance of Artificial Intelligence. In arXiv. https://arxiv.org/abs/2402.08797Sastry, G., L. Heim, H. Belfield, et al. 2024. “Computing Power and the Governance of Artificial Intelligence”. In arXiv. Preprint, February 13. https://arxiv.org/abs/2402.08797.Sastry, G., et al. “Computing Power and the Governance of Artificial Intelligence”. arXiv, 13 Feb. 2024, https://arxiv.org/abs/2402.08797.Sastry, G. et al. Computing Power and the Governance of Artificial Intelligence. arXiv Preprint at https://arxiv.org/abs/2402.08797 (2024).G. Sastry et al., “Computing Power and the Governance of Artificial Intelligence”, Feb. 13, 2024. [Online]. Available: https://arxiv.org/abs/2402.08797
Saunders, W. et al.(2022). Self-critiquing models for assisting human evaluators. arXiv.Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., & Leike, J. (2022). Self-critiquing models for assisting human evaluators. In arXiv. https://arxiv.org/abs/2206.05802Saunders, W., C. Yeh, J. Wu, et al. 2022. “Self-critiquing Models for Assisting Human Evaluators”. In arXiv. Preprint, June 12. https://arxiv.org/abs/2206.05802.Saunders, W., et al. “Self-critiquing Models for Assisting Human Evaluators”. arXiv, 12 June 2022, https://arxiv.org/abs/2206.05802.Saunders, W. et al. Self-critiquing models for assisting human evaluators. arXiv Preprint at https://arxiv.org/abs/2206.05802 (2022).W. Saunders et al., “Self-critiquing models for assisting human evaluators”, Jun. 12, 2022. [Online]. Available: https://arxiv.org/abs/2206.05802
scasper(2023). Deep Forgetting & Unlearning for Safely-Scoped LLMs. AI Alignment Forum.scasper. (2023, December 5). Deep Forgetting & Unlearning for Safely-Scoped LLMs. AI Alignment Forum. https://alignmentforum.org/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llmsscasper. 2023. “Deep Forgetting & Unlearning for Safely-Scoped LLMs”. AI Alignment Forum, December 5. https://alignmentforum.org/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llms.scasper. “Deep Forgetting & Unlearning for Safely-Scoped LLMs”. AI Alignment Forum, 5 Dec. 2023, https://alignmentforum.org/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llms.scasper. Deep Forgetting & Unlearning for Safely-Scoped LLMs. AI Alignment Forum https://alignmentforum.org/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llms (2023).scasper, “Deep Forgetting & Unlearning for Safely-Scoped LLMs”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llms
Schäfer, M., Schneider, J., Drechsler, K. & vom Brocke, J.(2022). AI GOVERNANCE: ARE CHIEF AI OFFICERS AND AI RISK OFFICERS NEEDED?. European Conference on Information Systems (ECIS).Schäfer, M., Schneider, J., Drechsler, K., & vom Brocke, J. (2022, May). AI GOVERNANCE: ARE CHIEF AI OFFICERS AND AI RISK OFFICERS NEEDED?. European Conference on Information Systems (ECIS). https://researchgate.net/publication/360644155_AI_GOVERNANCE_ARE_CHIEF_AI_OFFICERS_AND_AI_RISK_OFFICERS_NEEDEDSchäfer, M., J. Schneider, K. Drechsler, and J. vom Brocke. 2022. “AI GOVERNANCE: ARE CHIEF AI OFFICERS AND AI RISK OFFICERS NEEDED?”. European Conference on Information Systems (ECIS), May. https://researchgate.net/publication/360644155_AI_GOVERNANCE_ARE_CHIEF_AI_OFFICERS_AND_AI_RISK_OFFICERS_NEEDED.Schäfer, M., et al. “AI GOVERNANCE: ARE CHIEF AI OFFICERS AND AI RISK OFFICERS NEEDED?”. European Conference on Information Systems (ECIS), 2022, https://researchgate.net/publication/360644155_AI_GOVERNANCE_ARE_CHIEF_AI_OFFICERS_AND_AI_RISK_OFFICERS_NEEDED.Schäfer, M., Schneider, J., Drechsler, K. & vom Brocke, J. AI GOVERNANCE: ARE CHIEF AI OFFICERS AND AI RISK OFFICERS NEEDED?. in European Conference on Information Systems (ECIS) (2022).M. Schäfer, J. Schneider, K. Drechsler, and J. vom Brocke, “AI GOVERNANCE: ARE CHIEF AI OFFICERS AND AI RISK OFFICERS NEEDED?”, in European Conference on Information Systems (ECIS), May 2022. [Online]. Available: https://researchgate.net/publication/360644155_AI_GOVERNANCE_ARE_CHIEF_AI_OFFICERS_AND_AI_RISK_OFFICERS_NEEDED
Scherlis et. al.(2024). Experiments in Weak-to-Strong Generalization. EleutherAI Blog.Scherlis et. al. (2024, June 14). Experiments in Weak-to-Strong Generalization. EleutherAI Blog. https://blog.eleuther.ai/weak-to-strongScherlis et. al. 2024. “Experiments in Weak-to-Strong Generalization”. EleutherAI Blog, June 14. https://blog.eleuther.ai/weak-to-strong.Scherlis et. al. “Experiments in Weak-to-Strong Generalization”. EleutherAI Blog, 14 June 2024, https://blog.eleuther.ai/weak-to-strong.Scherlis et. al. Experiments in Weak-to-Strong Generalization. EleutherAI Blog https://blog.eleuther.ai/weak-to-strong (2024).Scherlis et. al., “Experiments in Weak-to-Strong Generalization”, EleutherAI Blog. [Online]. Available: https://blog.eleuther.ai/weak-to-strong
Scheurer, J., Balesni, M. & Hobbhahn, M.(2023). Large Language Models can Strategically Deceive their Users when Put Under Pressure. arXiv.Scheurer, J., Balesni, M., & Hobbhahn, M. (2023). Large Language Models can Strategically Deceive their Users when Put Under Pressure. In arXiv. https://arxiv.org/abs/2311.07590Scheurer, J., M. Balesni, and M. Hobbhahn. 2023. “Large Language Models Can Strategically Deceive Their Users When Put Under Pressure”. In arXiv. Preprint, November 9. https://arxiv.org/abs/2311.07590.Scheurer, J., et al. “Large Language Models Can Strategically Deceive Their Users When Put Under Pressure”. arXiv, 9 Nov. 2023, https://arxiv.org/abs/2311.07590.Scheurer, J., Balesni, M. & Hobbhahn, M. Large Language Models can Strategically Deceive their Users when Put Under Pressure. arXiv Preprint at https://arxiv.org/abs/2311.07590 (2023).J. Scheurer, M. Balesni, and M. Hobbhahn, “Large Language Models can Strategically Deceive their Users when Put Under Pressure”, Nov. 09, 2023. [Online]. Available: https://arxiv.org/abs/2311.07590
Scheurer, J., Campos, J. A., Chan, J. S., Chen, A., Cho, K. & Perez, E.(2022). Training Language Models with Language Feedback. arXiv.Scheurer, J., Campos, J. A., Chan, J. S., Chen, A., Cho, K., & Perez, E. (2022). Training Language Models with Language Feedback. In arXiv. https://arxiv.org/abs/2204.14146Scheurer, J., J. A. Campos, J. S. Chan, A. Chen, K. Cho, and E. Perez. 2022. “Training Language Models with Language Feedback”. In arXiv. Preprint, April 29. https://arxiv.org/abs/2204.14146.Scheurer, J., et al. “Training Language Models with Language Feedback”. arXiv, 29 Apr. 2022, https://arxiv.org/abs/2204.14146.Scheurer, J. et al. Training Language Models with Language Feedback. arXiv Preprint at https://arxiv.org/abs/2204.14146 (2022).J. Scheurer, J. A. Campos, J. S. Chan, A. Chen, K. Cho, and E. Perez, “Training Language Models with Language Feedback”, Apr. 29, 2022. [Online]. Available: https://arxiv.org/abs/2204.14146
Schick, T. et al.(2023). Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv.Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. In arXiv. https://arxiv.org/abs/2302.04761Schick, T., J. Dwivedi-Yu, R. Dessì, et al. 2023. “Toolformer: Language Models Can Teach Themselves to Use Tools”. In arXiv. Preprint, February 9. https://arxiv.org/abs/2302.04761.Schick, T., et al. “Toolformer: Language Models Can Teach Themselves to Use Tools”. arXiv, 9 Feb. 2023, https://arxiv.org/abs/2302.04761.Schick, T. et al. Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv Preprint at https://arxiv.org/abs/2302.04761 (2023).T. Schick et al., “Toolformer: Language Models Can Teach Themselves to Use Tools”, Feb. 09, 2023. [Online]. Available: https://arxiv.org/abs/2302.04761
Schrimpf et al.(2021). The neural architecture of language: Integrative modeling converges on predictive processing. PNAS.Schrimpf et al. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. Internet Archive (https://web.archive.org/web/20220221135833/https://www.pnas.org/content/118/45/e2105646118). PNAS. https://pnas.org/content/118/45/e2105646118Schrimpf et al. 2021. “The Neural Architecture of Language: Integrative Modeling Converges on Predictive Processing”. PNAS. Https://web.archive.org/web/20220221135833/https://www.pnas.org/content/118/45/e2105646118. Internet Archive. https://pnas.org/content/118/45/e2105646118.Schrimpf et al. “The Neural Architecture of Language: Integrative Modeling Converges on Predictive Processing”. PNAS, 2021, Internet Archive, https://web.archive.org/web/20220221135833/https://www.pnas.org/content/118/45/e2105646118, https://pnas.org/content/118/45/e2105646118.Schrimpf et al. The neural architecture of language: Integrative modeling converges on predictive processing. PNAS https://pnas.org/content/118/45/e2105646118 (2021).Schrimpf et al., “The neural architecture of language: Integrative modeling converges on predictive processing”, PNAS. Accessed: Feb. 21, 2022. [Online]. Available: https://pnas.org/content/118/45/e2105646118
Schuett, J.(2022). Three lines of defense against risks from AI. arXiv.Schuett, J. (2022). Three lines of defense against risks from AI. In arXiv. https://doi.org/10.1007/s00146-023-01811-0Schuett, J. 2022. “Three Lines of Defense Against Risks from AI”. In arXiv. Preprint, December 16. https://doi.org/10.1007/s00146-023-01811-0.Schuett, J. “Three Lines of Defense Against Risks from AI”. arXiv, 16 Dec. 2022, https://doi.org/10.1007/s00146-023-01811-0.Schuett, J. Three lines of defense against risks from AI. arXiv Preprint at https://doi.org/10.1007/s00146-023-01811-0 (2022).J. Schuett, “Three lines of defense against risks from AI”, Dec. 16, 2022. doi: 10.1007/s00146-023-01811-0.
Schwarzschild, A., Goldblum, M., Gupta, A., Dickerson, J. P. & Goldstein, T.(2020). Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks. arXiv.Schwarzschild, A., Goldblum, M., Gupta, A., Dickerson, J. P., & Goldstein, T. (2020). Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks. In arXiv. https://arxiv.org/abs/2006.12557Schwarzschild, A., M. Goldblum, A. Gupta, J. P. Dickerson, and T. Goldstein. 2020. “Just How Toxic Is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks”. In arXiv. Preprint, June 22. https://arxiv.org/abs/2006.12557.Schwarzschild, A., et al. “Just How Toxic Is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks”. arXiv, 22 June 2020, https://arxiv.org/abs/2006.12557.Schwarzschild, A., Goldblum, M., Gupta, A., Dickerson, J. P. & Goldstein, T. Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks. arXiv Preprint at https://arxiv.org/abs/2006.12557 (2020).A. Schwarzschild, M. Goldblum, A. Gupta, J. P. Dickerson, and T. Goldstein, “Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks”, Jun. 22, 2020. [Online]. Available: https://arxiv.org/abs/2006.12557
Sclar, M., Choi, Y., Tsvetkov, Y. & Suhr, A.(2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting. arXiv.Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting. In arXiv. https://arxiv.org/abs/2310.11324Sclar, M., Y. Choi, Y. Tsvetkov, and A. Suhr. 2023. “Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design Or: How I Learned to Start Worrying About Prompt Formatting”. In arXiv. Preprint, October 17. https://arxiv.org/abs/2310.11324.Sclar, M., et al. “Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design Or: How I Learned to Start Worrying About Prompt Formatting”. arXiv, 17 Oct. 2023, https://arxiv.org/abs/2310.11324.Sclar, M., Choi, Y., Tsvetkov, Y. & Suhr, A. Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting. arXiv Preprint at https://arxiv.org/abs/2310.11324 (2023).M. Sclar, Y. Choi, Y. Tsvetkov, and A. Suhr, “Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting”, Oct. 17, 2023. [Online]. Available: https://arxiv.org/abs/2310.11324
Scott Alexander(2023). Pause For Thought: The AI Pause Debate. EA Forum.Scott Alexander. (2023, October 10). Pause For Thought: The AI Pause Debate. EA Forum. https://forum.effectivealtruism.org/posts/7WfMYzLfcTyDtD6Gn/pause-for-thought-the-ai-pause-debateScott Alexander. 2023. “Pause For Thought: The AI Pause Debate”. EA Forum, October 10. https://forum.effectivealtruism.org/posts/7WfMYzLfcTyDtD6Gn/pause-for-thought-the-ai-pause-debate.Scott Alexander. “Pause For Thought: The AI Pause Debate”. EA Forum, 10 Oct. 2023, https://forum.effectivealtruism.org/posts/7WfMYzLfcTyDtD6Gn/pause-for-thought-the-ai-pause-debate.Scott Alexander. Pause For Thought: The AI Pause Debate. EA Forum https://forum.effectivealtruism.org/posts/7WfMYzLfcTyDtD6Gn/pause-for-thought-the-ai-pause-debate (2023).Scott Alexander, “Pause For Thought: The AI Pause Debate”, EA Forum. [Online]. Available: https://forum.effectivealtruism.org/posts/7WfMYzLfcTyDtD6Gn/pause-for-thought-the-ai-pause-debate
Searle, J. R.(1980). Minds, brains, and programs. Behavioral and Brain Sciences.Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417–424. https://doi.org/10.1017/S0140525X00005756Searle, J. R. 1980. “Minds, Brains, and Programs”. Behavioral and Brain Sciences 3 (3): 417–24. https://doi.org/10.1017/S0140525X00005756.Searle, J. R. “Minds, Brains, and Programs”. Behavioral and Brain Sciences, vol. 3, no. 3, Sept. 1980, pp. 417–24, https://doi.org/10.1017/S0140525X00005756.Searle, J. R. Minds, brains, and programs. Behavioral and Brain Sciences 3, 417–424 (1980).J. R. Searle, “Minds, brains, and programs”, Behavioral and Brain Sciences, vol. 3, no. 3, pp. 417–424, Sep. 1980, doi: 10.1017/S0140525X00005756.
(2024). Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models.Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, Jeff Alstott. (2024). Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. https://rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdfSella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, Jeff Alstott. 2024. “Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models”. https://rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdf.Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, Jeff Alstott. Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. 2024, https://rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdf.Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, Jeff Alstott. Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. https://rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdf (2024).Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, Jeff Alstott, “Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models”. [Online]. Available: https://rand.org/content/dam/rand/pubs/research_reports/RRA2800/RRA2849-1/RAND_RRA2849-1.pdf
Seger et al.(2023). Democratising AI: Multiple Meanings, Goals, and Methods. arXiv.org.Seger et al. (2023). Democratising AI: Multiple Meanings, Goals, and Methods. arXiv.org. https://www.arxiv.org/abs/2303.12642Seger et al. 2023. “Democratising AI: Multiple Meanings, Goals, and Methods”. arXiv.org. https://www.arxiv.org/abs/2303.12642.Seger et al. “Democratising AI: Multiple Meanings, Goals, and Methods”. arXiv.org, 2023, https://www.arxiv.org/abs/2303.12642.Seger et al. Democratising AI: Multiple Meanings, Goals, and Methods. arXiv.org https://www.arxiv.org/abs/2303.12642 (2023).Seger et al., “Democratising AI: Multiple Meanings, Goals, and Methods”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2303.12642
Seger, E. et al.(2023). Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives. arXiv.Seger, E., Dreksler, N., Moulange, R., Dardaman, E., Schuett, J., Wei, K., Winter, C., Arnold, M., hÉigeartaigh, S. Ó., Korinek, A., Anderljung, M., Bucknall, B., Chan, A., Stafford, E., Koessler, L., Ovadya, A., Garfinkel, B., Bluemke, E., Aird, M., … Gupta, A. (2023). Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives. In arXiv. https://arxiv.org/abs/2311.09227Seger, E., N. Dreksler, R. Moulange, et al. 2023. “Open-Sourcing Highly Capable Foundation Models: An Evaluation of Risks, Benefits, and Alternative Methods for Pursuing Open-source Objectives”. In arXiv. Preprint, September 29. https://arxiv.org/abs/2311.09227.Seger, E., et al. “Open-Sourcing Highly Capable Foundation Models: An Evaluation of Risks, Benefits, and Alternative Methods for Pursuing Open-source Objectives”. arXiv, 29 Sept. 2023, https://arxiv.org/abs/2311.09227.Seger, E. et al. Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives. arXiv Preprint at https://arxiv.org/abs/2311.09227 (2023).E. Seger et al., “Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives”, Sep. 29, 2023. [Online]. Available: https://arxiv.org/abs/2311.09227
Sener, O. & Koltun, V.(2018). Multi-Task Learning as Multi-Objective Optimization. arXiv.Sener, O., & Koltun, V. (2018). Multi-Task Learning as Multi-Objective Optimization. In arXiv. https://arxiv.org/abs/1810.04650Sener, O., and V. Koltun. 2018. “Multi-Task Learning as Multi-Objective Optimization”. In arXiv. Preprint, October 10. https://arxiv.org/abs/1810.04650.Sener, O., and V. Koltun. “Multi-Task Learning as Multi-Objective Optimization”. arXiv, 10 Oct. 2018, https://arxiv.org/abs/1810.04650.Sener, O. & Koltun, V. Multi-Task Learning as Multi-Objective Optimization. arXiv Preprint at https://arxiv.org/abs/1810.04650 (2018).O. Sener and V. Koltun, “Multi-Task Learning as Multi-Objective Optimization”, Oct. 10, 2018. [Online]. Available: https://arxiv.org/abs/1810.04650
Sevilla et al.(2024). Can AI scaling continue through 2030?. Epoch AI.Sevilla et al. (2024). Can AI scaling continue through 2030?. Epoch AI. https://epoch.ai/blog/can-ai-scaling-continue-through-2030Sevilla et al. 2024. “Can AI Scaling Continue Through 2030?”. Epoch AI. https://epoch.ai/blog/can-ai-scaling-continue-through-2030.Sevilla et al. “Can AI Scaling Continue Through 2030?”. Epoch AI, 2024, https://epoch.ai/blog/can-ai-scaling-continue-through-2030.Sevilla et al. Can AI scaling continue through 2030?. Epoch AI https://epoch.ai/blog/can-ai-scaling-continue-through-2030 (2024).Sevilla et al., “Can AI scaling continue through 2030?”, Epoch AI. [Online]. Available: https://epoch.ai/blog/can-ai-scaling-continue-through-2030
Shah, H., Tamuly, K., Raghunathan, A., Jain, P. & Netrapalli, P.(2020). The Pitfalls of Simplicity Bias in Neural Networks. arXiv.Shah, H., Tamuly, K., Raghunathan, A., Jain, P., & Netrapalli, P. (2020). The Pitfalls of Simplicity Bias in Neural Networks. In arXiv. https://arxiv.org/abs/2006.07710Shah, H., K. Tamuly, A. Raghunathan, P. Jain, and P. Netrapalli. 2020. “The Pitfalls of Simplicity Bias in Neural Networks”. In arXiv. Preprint, June 13. https://arxiv.org/abs/2006.07710.Shah, H., et al. “The Pitfalls of Simplicity Bias in Neural Networks”. arXiv, 13 June 2020, https://arxiv.org/abs/2006.07710.Shah, H., Tamuly, K., Raghunathan, A., Jain, P. & Netrapalli, P. The Pitfalls of Simplicity Bias in Neural Networks. arXiv Preprint at https://arxiv.org/abs/2006.07710 (2020).H. Shah, K. Tamuly, A. Raghunathan, P. Jain, and P. Netrapalli, “The Pitfalls of Simplicity Bias in Neural Networks”, Jun. 13, 2020. [Online]. Available: https://arxiv.org/abs/2006.07710
Shah, R. et al.(2025). An Approach to Technical AGI Safety and Security. arXiv.Shah, R., Irpan, A., Turner, A. M., Wang, A., Conmy, A., Lindner, D., Brown-Cohen, J., Ho, L., Nanda, N., Popa, R. A., Jain, R., Greig, R., Albanie, S., Emmons, S., Farquhar, S., Krier, S., Rajamanoharan, S., Bridgers, S., Ijitoye, T., … Dragan, A. (2025). An Approach to Technical AGI Safety and Security. In arXiv. https://arxiv.org/abs/2504.01849Shah, R., A. Irpan, A. M. Turner, et al. 2025. “An Approach to Technical AGI Safety and Security”. In arXiv. Preprint, April 2. https://arxiv.org/abs/2504.01849.Shah, R., et al. “An Approach to Technical AGI Safety and Security”. arXiv, 2 Apr. 2025, https://arxiv.org/abs/2504.01849.Shah, R. et al. An Approach to Technical AGI Safety and Security. arXiv Preprint at https://arxiv.org/abs/2504.01849 (2025).R. Shah et al., “An Approach to Technical AGI Safety and Security”, Apr. 02, 2025. [Online]. Available: https://arxiv.org/abs/2504.01849
Shane Legg & Marcus Hutter(2007). Universal Intelligence: A Definition of Machine Intelligence. arXiv.Shane Legg, & Marcus Hutter. (2007). Universal Intelligence: A Definition of Machine Intelligence. In arXiv. https://arxiv.org/abs/0712.3329Shane Legg, and Marcus Hutter. 2007. “Universal Intelligence: A Definition of Machine Intelligence”. In arXiv. Preprint, December 20. https://arxiv.org/abs/0712.3329.Shane Legg, and Marcus Hutter. “Universal Intelligence: A Definition of Machine Intelligence”. arXiv, 20 Dec. 2007, https://arxiv.org/abs/0712.3329.Shane Legg & Marcus Hutter. Universal Intelligence: A Definition of Machine Intelligence. arXiv Preprint at https://arxiv.org/abs/0712.3329 (2007).Shane Legg and Marcus Hutter, “Universal Intelligence: A Definition of Machine Intelligence”, Dec. 20, 2007. [Online]. Available: https://arxiv.org/abs/0712.3329
Shao, M. et al.(2024). NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security. arXiv.Shao, M., Jancheska, S., Udeshi, M., Dolan-Gavitt, B., Xi, H., Milner, K., Chen, B., Yin, M., Garg, S., Krishnamurthy, P., Khorrami, F., Karri, R., & Shafique, M. (2024). NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security. In arXiv. https://arxiv.org/abs/2406.05590Shao, M., S. Jancheska, M. Udeshi, et al. 2024. “NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security”. In arXiv. Preprint, June 8. https://arxiv.org/abs/2406.05590.Shao, M., et al. “NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security”. arXiv, 8 June 2024, https://arxiv.org/abs/2406.05590.Shao, M. et al. NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security. arXiv Preprint at https://arxiv.org/abs/2406.05590 (2024).M. Shao et al., “NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security”, Jun. 08, 2024. [Online]. Available: https://arxiv.org/abs/2406.05590
Sharkey et al.(2024). static1.squarespace.com/static/659…39547455/auditing_framework_web.pdf.Sharkey et al. (2024). Sharkey et al., 2024. https://static1.squarespace.com/static/6593e7097565990e65c886fd/t/65a6f1389754fc06cb9a7a14/1705439547455/auditing_framework_web.pdfSharkey et al. 2024. “Sharkey Et Al., 2024”. https://static1.squarespace.com/static/6593e7097565990e65c886fd/t/65a6f1389754fc06cb9a7a14/1705439547455/auditing_framework_web.pdf.Sharkey et al. Sharkey Et Al., 2024. 2024, https://static1.squarespace.com/static/6593e7097565990e65c886fd/t/65a6f1389754fc06cb9a7a14/1705439547455/auditing_framework_web.pdf.Sharkey et al. Sharkey et al., 2024. https://static1.squarespace.com/static/6593e7097565990e65c886fd/t/65a6f1389754fc06cb9a7a14/1705439547455/auditing_framework_web.pdf (2024).Sharkey et al., “Sharkey et al., 2024”. [Online]. Available: https://static1.squarespace.com/static/6593e7097565990e65c886fd/t/65a6f1389754fc06cb9a7a14/1705439547455/auditing_framework_web.pdf
Shavit, Y.(2023). What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv.Shavit, Y. (2023). What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. In arXiv. https://arxiv.org/abs/2303.11341Shavit, Y. 2023. “What Does It Take to Catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring”. In arXiv. Preprint, March 20. https://arxiv.org/abs/2303.11341.Shavit, Y. “What Does It Take to Catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring”. arXiv, 20 Mar. 2023, https://arxiv.org/abs/2303.11341.Shavit, Y. What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv Preprint at https://arxiv.org/abs/2303.11341 (2023).Y. Shavit, “What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring”, Mar. 20, 2023. [Online]. Available: https://arxiv.org/abs/2303.11341
Shayegani, E., Mamun, M. A. A., Fu, Y., Zaree, P., Dong, Y. & Abu-Ghazaleh, N.(2023). Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks. arXiv.Shayegani, E., Mamun, M. A. A., Fu, Y., Zaree, P., Dong, Y., & Abu-Ghazaleh, N. (2023). Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks. In arXiv. https://arxiv.org/abs/2310.10844Shayegani, E., M. A. A. Mamun, Y. Fu, P. Zaree, Y. Dong, and N. Abu-Ghazaleh. 2023. “Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks”. In arXiv. Preprint, October 16. https://arxiv.org/abs/2310.10844.Shayegani, E., et al. “Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks”. arXiv, 16 Oct. 2023, https://arxiv.org/abs/2310.10844.Shayegani, E. et al. Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks. arXiv Preprint at https://arxiv.org/abs/2310.10844 (2023).E. Shayegani, M. A. A. Mamun, Y. Fu, P. Zaree, Y. Dong, and N. Abu-Ghazaleh, “Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks”, Oct. 16, 2023. [Online]. Available: https://arxiv.org/abs/2310.10844
Shen, Y., Song, K., Tan, X., Li, D., Lu, W. & Zhuang, Y.(2023). HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. arXiv.Shen, Y., Song, K., Tan, X., Li, D., Lu, W., & Zhuang, Y. (2023). HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. In arXiv. https://arxiv.org/abs/2303.17580Shen, Y., K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang. 2023. “HuggingGPT: Solving AI Tasks with ChatGPT and Its Friends in Hugging Face”. In arXiv. Preprint, March 30. https://arxiv.org/abs/2303.17580.Shen, Y., et al. “HuggingGPT: Solving AI Tasks with ChatGPT and Its Friends in Hugging Face”. arXiv, 30 Mar. 2023, https://arxiv.org/abs/2303.17580.Shen, Y. et al. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. arXiv Preprint at https://arxiv.org/abs/2303.17580 (2023).Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face”, Mar. 30, 2023. [Online]. Available: https://arxiv.org/abs/2303.17580
Shevlane, T. & Dafoe, A.(2019). The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?. arXiv.Shevlane, T., & Dafoe, A. (2019). The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?. In arXiv. https://arxiv.org/abs/2001.00463Shevlane, T., and A. Dafoe. 2019. “The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?”. In arXiv. Preprint, December 27. https://arxiv.org/abs/2001.00463.Shevlane, T., and A. Dafoe. “The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?”. arXiv, 27 Dec. 2019, https://arxiv.org/abs/2001.00463.Shevlane, T. & Dafoe, A. The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?. arXiv Preprint at https://arxiv.org/abs/2001.00463 (2019).T. Shevlane and A. Dafoe, “The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?”, Dec. 27, 2019. [Online]. Available: https://arxiv.org/abs/2001.00463
Shevlane, T. et al.(2023). Model evaluation for extreme risks. arXiv.Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., Ho, L., Siddarth, D., Avin, S., Hawkins, W., Kim, B., Gabriel, I., Bolina, V., Clark, J., Bengio, Y., … Dafoe, A. (2023). Model evaluation for extreme risks. In arXiv. https://arxiv.org/abs/2305.15324Shevlane, T., S. Farquhar, B. Garfinkel, et al. 2023. “Model Evaluation for Extreme Risks”. In arXiv. Preprint, May 24. https://arxiv.org/abs/2305.15324.Shevlane, T., et al. “Model Evaluation for Extreme Risks”. arXiv, 24 May 2023, https://arxiv.org/abs/2305.15324.Shevlane, T. et al. Model evaluation for extreme risks. arXiv Preprint at https://arxiv.org/abs/2305.15324 (2023).T. Shevlane et al., “Model evaluation for extreme risks”, May 24, 2023. [Online]. Available: https://arxiv.org/abs/2305.15324
Shi, F. et al.(2022). Language Models are Multilingual Chain-of-Thought Reasoners. arXiv.Shi, F., Suzgun, M., Freitag, M., Wang, X., Srivats, S., Vosoughi, S., Chung, H. W., Tay, Y., Ruder, S., Zhou, D., Das, D., & Wei, J. (2022). Language Models are Multilingual Chain-of-Thought Reasoners. In arXiv. https://arxiv.org/abs/2210.03057Shi, F., M. Suzgun, M. Freitag, et al. 2022. “Language Models Are Multilingual Chain-of-Thought Reasoners”. In arXiv. Preprint, October 6. https://arxiv.org/abs/2210.03057.Shi, F., et al. “Language Models Are Multilingual Chain-of-Thought Reasoners”. arXiv, 6 Oct. 2022, https://arxiv.org/abs/2210.03057.Shi, F. et al. Language Models are Multilingual Chain-of-Thought Reasoners. arXiv Preprint at https://arxiv.org/abs/2210.03057 (2022).F. Shi et al., “Language Models are Multilingual Chain-of-Thought Reasoners”, Oct. 06, 2022. [Online]. Available: https://arxiv.org/abs/2210.03057
Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K. & Yao, S.(2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv.Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. In arXiv. https://arxiv.org/abs/2303.11366Shinn, N., F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao. 2023. “Reflexion: Language Agents with Verbal Reinforcement Learning”. In arXiv. Preprint, March 20. https://arxiv.org/abs/2303.11366.Shinn, N., et al. “Reflexion: Language Agents with Verbal Reinforcement Learning”. arXiv, 20 Mar. 2023, https://arxiv.org/abs/2303.11366.Shinn, N. et al. Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv Preprint at https://arxiv.org/abs/2303.11366 (2023).N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language Agents with Verbal Reinforcement Learning”, Mar. 20, 2023. [Online]. Available: https://arxiv.org/abs/2303.11366
Shokri, R., Stronati, M., Song, C. & Shmatikov, V.(2016). Membership Inference Attacks against Machine Learning Models. arXiv.Shokri, R., Stronati, M., Song, C., & Shmatikov, V. (2016). Membership Inference Attacks against Machine Learning Models. In arXiv. https://arxiv.org/abs/1610.05820Shokri, R., M. Stronati, C. Song, and V. Shmatikov. 2016. “Membership Inference Attacks Against Machine Learning Models”. In arXiv. Preprint, October 18. https://arxiv.org/abs/1610.05820.Shokri, R., et al. “Membership Inference Attacks Against Machine Learning Models”. arXiv, 18 Oct. 2016, https://arxiv.org/abs/1610.05820.Shokri, R., Stronati, M., Song, C. & Shmatikov, V. Membership Inference Attacks against Machine Learning Models. arXiv Preprint at https://arxiv.org/abs/1610.05820 (2016).R. Shokri, M. Stronati, C. Song, and V. Shmatikov, “Membership Inference Attacks against Machine Learning Models”, Oct. 18, 2016. [Online]. Available: https://arxiv.org/abs/1610.05820
SIMA team(2025). SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds.SIMA team. (2025, November 13). SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds. https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worldsSIMA team. 2025. “SIMA 2: An Agent That Plays, Reasons, and Learns With You in Virtual 3D Worlds”. November 13. https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds.SIMA team. SIMA 2: An Agent That Plays, Reasons, and Learns With You in Virtual 3D Worlds. 13 Nov. 2025, https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds.SIMA team. SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds. https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds (2025).SIMA team, “SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds”. [Online]. Available: https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds
Simmons-Edler, R., Badman, R., Longpre, S. & Rajan, K.(2024). AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research. arXiv.Simmons-Edler, R., Badman, R., Longpre, S., & Rajan, K. (2024). AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research. In arXiv. https://arxiv.org/abs/2405.01859Simmons-Edler, R., R. Badman, S. Longpre, and K. Rajan. 2024. “AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research”. In arXiv. Preprint, May 3. https://arxiv.org/abs/2405.01859.Simmons-Edler, R., et al. “AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research”. arXiv, 3 May 2024, https://arxiv.org/abs/2405.01859.Simmons-Edler, R., Badman, R., Longpre, S. & Rajan, K. AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research. arXiv Preprint at https://arxiv.org/abs/2405.01859 (2024).R. Simmons-Edler, R. Badman, S. Longpre, and K. Rajan, “AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research”, May 03, 2024. [Online]. Available: https://arxiv.org/abs/2405.01859
Skare et al.(2024). Artificial intelligence and wealth inequality: A comprehensi.Skare et al. (2024). Artificial intelligence and wealth inequality: A comprehensi. https://ideas.repec.org/a/eee/teinso/v79y2024ics0160791x24002677.htmlSkare et al. 2024. “Artificial Intelligence and Wealth Inequality: A Comprehensi”. https://ideas.repec.org/a/eee/teinso/v79y2024ics0160791x24002677.html.Skare et al. Artificial Intelligence and Wealth Inequality: A Comprehensi. 2024, https://ideas.repec.org/a/eee/teinso/v79y2024ics0160791x24002677.html.Skare et al. Artificial intelligence and wealth inequality: A comprehensi. https://ideas.repec.org/a/eee/teinso/v79y2024ics0160791x24002677.html (2024).Skare et al., “Artificial intelligence and wealth inequality: A comprehensi”. [Online]. Available: https://ideas.repec.org/a/eee/teinso/v79y2024ics0160791x24002677.html
Slattery, P. et al.(2024). The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence. arXiv.Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2024). The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence. In arXiv. https://doi.org/10.1016/j.patter.2026.101517Slattery, P., A. K. Saeri, E. A. C. Grundy, et al. 2024. “The AI Risk Repository: A Meta-review, Database, and Taxonomy of Risks from Artificial Intelligence”. In arXiv. Preprint, August 14. https://doi.org/10.1016/j.patter.2026.101517.Slattery, P., et al. “The AI Risk Repository: A Meta-review, Database, and Taxonomy of Risks from Artificial Intelligence”. arXiv, 14 Aug. 2024, https://doi.org/10.1016/j.patter.2026.101517.Slattery, P. et al. The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence. arXiv Preprint at https://doi.org/10.1016/j.patter.2026.101517 (2024).P. Slattery et al., “The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence”, Aug. 14, 2024. doi: 10.1016/j.patter.2026.101517.
Snyder et al.(2020). Measuring Cybersecurity and Cyber Resiliency.Snyder et al. (2020). Measuring Cybersecurity and Cyber Resiliency. Internet Archive (https://web.archive.org/web/20250507024856/https://www.rand.org/pubs/research_reports/RR2703.html). https://rand.org/pubs/research_reports/RR2703.htmlSnyder et al. 2020. “Measuring Cybersecurity and Cyber Resiliency”. Https://web.archive.org/web/20250507024856/https://www.rand.org/pubs/research_reports/RR2703.html. Internet Archive. https://rand.org/pubs/research_reports/RR2703.html.Snyder et al. Measuring Cybersecurity and Cyber Resiliency. 2020, Internet Archive, https://web.archive.org/web/20250507024856/https://www.rand.org/pubs/research_reports/RR2703.html, https://rand.org/pubs/research_reports/RR2703.html.Snyder et al. Measuring Cybersecurity and Cyber Resiliency. https://rand.org/pubs/research_reports/RR2703.html (2020).Snyder et al., “Measuring Cybersecurity and Cyber Resiliency”. Accessed: May 07, 2025. [Online]. Available: https://rand.org/pubs/research_reports/RR2703.html
So8res(2022). A central AI alignment problem: capabilities generalization, and the sharp left turn. AI Alignment Forum.So8res. (2022, June 15). A central AI alignment problem: capabilities generalization, and the sharp left turn. AI Alignment Forum. https://alignmentforum.org/posts/GNhMPAWcfBCASy8e6/a-central-ai-alignment-problem-capabilities-generalizationSo8res. 2022. “A Central AI Alignment Problem: Capabilities Generalization, and the Sharp Left Turn”. AI Alignment Forum, June 15. https://alignmentforum.org/posts/GNhMPAWcfBCASy8e6/a-central-ai-alignment-problem-capabilities-generalization.So8res. “A Central AI Alignment Problem: Capabilities Generalization, and the Sharp Left Turn”. AI Alignment Forum, 15 June 2022, https://alignmentforum.org/posts/GNhMPAWcfBCASy8e6/a-central-ai-alignment-problem-capabilities-generalization.So8res. A central AI alignment problem: capabilities generalization, and the sharp left turn. AI Alignment Forum https://alignmentforum.org/posts/GNhMPAWcfBCASy8e6/a-central-ai-alignment-problem-capabilities-generalization (2022).So8res, “A central AI alignment problem: capabilities generalization, and the sharp left turn”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/GNhMPAWcfBCASy8e6/a-central-ai-alignment-problem-capabilities-generalization
So8res(2023). Deep Deceptiveness. AI Alignment Forum.So8res. (2023, March 21). Deep Deceptiveness. AI Alignment Forum. https://alignmentforum.org/posts/XWwvwytieLtEWaFJX/deep-deceptivenessSo8res. 2023. “Deep Deceptiveness”. AI Alignment Forum, March 21. https://alignmentforum.org/posts/XWwvwytieLtEWaFJX/deep-deceptiveness.So8res. “Deep Deceptiveness”. AI Alignment Forum, 21 Mar. 2023, https://alignmentforum.org/posts/XWwvwytieLtEWaFJX/deep-deceptiveness.So8res. Deep Deceptiveness. AI Alignment Forum https://alignmentforum.org/posts/XWwvwytieLtEWaFJX/deep-deceptiveness (2023).So8res, “Deep Deceptiveness”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/XWwvwytieLtEWaFJX/deep-deceptiveness
Soares(2023). Ability to solve long-horizon tasks correlates with wanting things in the behaviorist sense - Machine Intelligence Research Institute.Soares. (2023, November 24). Ability to solve long-horizon tasks correlates with wanting things in the behaviorist sense - Machine Intelligence Research Institute. Machine Intelligence Research Institute. https://intelligence.org/2023/11/24/ability-to-solve-long-horizon-tasks-correlates-with-wanting-things-in-the-behaviorist-senseSoares. 2023. “Ability to Solve Long-horizon Tasks Correlates with Wanting Things in the Behaviorist Sense - Machine Intelligence Research Institute”. Machine Intelligence Research Institute, November 24. https://intelligence.org/2023/11/24/ability-to-solve-long-horizon-tasks-correlates-with-wanting-things-in-the-behaviorist-sense.Soares. “Ability to Solve Long-horizon Tasks Correlates with Wanting Things in the Behaviorist Sense - Machine Intelligence Research Institute”. Machine Intelligence Research Institute, 24 Nov. 2023, https://intelligence.org/2023/11/24/ability-to-solve-long-horizon-tasks-correlates-with-wanting-things-in-the-behaviorist-sense.Soares. Ability to solve long-horizon tasks correlates with wanting things in the behaviorist sense - Machine Intelligence Research Institute. Machine Intelligence Research Institute https://intelligence.org/2023/11/24/ability-to-solve-long-horizon-tasks-correlates-with-wanting-things-in-the-behaviorist-sense (2023).Soares, “Ability to solve long-horizon tasks correlates with wanting things in the behaviorist sense - Machine Intelligence Research Institute”, Machine Intelligence Research Institute. [Online]. Available: https://intelligence.org/2023/11/24/ability-to-solve-long-horizon-tasks-correlates-with-wanting-things-in-the-behaviorist-sense
Solaiman, I.(2023). The Gradient of Generative AI Release: Methods and Considerations. arXiv.Solaiman, I. (2023). The Gradient of Generative AI Release: Methods and Considerations. In arXiv. https://arxiv.org/abs/2302.04844Solaiman, I. 2023. “The Gradient of Generative AI Release: Methods and Considerations”. In arXiv. Preprint, February 5. https://arxiv.org/abs/2302.04844.Solaiman, I. “The Gradient of Generative AI Release: Methods and Considerations”. arXiv, 5 Feb. 2023, https://arxiv.org/abs/2302.04844.Solaiman, I. The Gradient of Generative AI Release: Methods and Considerations. arXiv Preprint at https://arxiv.org/abs/2302.04844 (2023).I. Solaiman, “The Gradient of Generative AI Release: Methods and Considerations”, Feb. 05, 2023. [Online]. Available: https://arxiv.org/abs/2302.04844
Solaiman, I. et al.(2019). Release Strategies and the Social Impacts of Language Models. arXiv.Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., Krueger, G., Kim, J. W., Kreps, S., McCain, M., Newhouse, A., Blazakis, J., McGuffie, K., & Wang, J. (2019). Release Strategies and the Social Impacts of Language Models. In arXiv. https://arxiv.org/abs/1908.09203Solaiman, I., M. Brundage, J. Clark, et al. 2019. “Release Strategies and the Social Impacts of Language Models”. In arXiv. Preprint, August 24. https://arxiv.org/abs/1908.09203.Solaiman, I., et al. “Release Strategies and the Social Impacts of Language Models”. arXiv, 24 Aug. 2019, https://arxiv.org/abs/1908.09203.Solaiman, I. et al. Release Strategies and the Social Impacts of Language Models. arXiv Preprint at https://arxiv.org/abs/1908.09203 (2019).I. Solaiman et al., “Release Strategies and the Social Impacts of Language Models”, Aug. 24, 2019. [Online]. Available: https://arxiv.org/abs/1908.09203
Solaiman, I. et al.(2023). Evaluating the Social Impact of Generative AI Systems in Systems and Society. arXiv.Solaiman, I., Talat, Z., Agnew, W., Ahmad, L., Baker, D., Blodgett, S. L., Chen, C., Daumé, H., Dodge, J., Duan, I., Evans, E., Friedrich, F., Ghosh, A., Gohar, U., Hooker, S., Jernite, Y., Kalluri, R., Lusoli, A., Leidinger, A., … Subramonian, A. (2023). Evaluating the Social Impact of Generative AI Systems in Systems and Society. In arXiv. https://doi.org/10.1093/oxfordhb/9780198940272.013.0025Solaiman, I., Z. Talat, W. Agnew, et al. 2023. “Evaluating the Social Impact of Generative AI Systems in Systems and Society”. In arXiv. Preprint, June 9. https://doi.org/10.1093/oxfordhb/9780198940272.013.0025.Solaiman, I., et al. “Evaluating the Social Impact of Generative AI Systems in Systems and Society”. arXiv, 9 June 2023, https://doi.org/10.1093/oxfordhb/9780198940272.013.0025.Solaiman, I. et al. Evaluating the Social Impact of Generative AI Systems in Systems and Society. arXiv Preprint at https://doi.org/10.1093/oxfordhb/9780198940272.013.0025 (2023).I. Solaiman et al., “Evaluating the Social Impact of Generative AI Systems in Systems and Society”, Jun. 09, 2023. doi: 10.1093/oxfordhb/9780198940272.013.0025.
Srivastava, A. et al.(2022). Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models. arXiv.Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Power, A., Ray, A., Warstadt, A., Kocurek, A. W., Safaya, A., Tazarv, A., … Wu, Z. (2022). Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models. In arXiv. https://arxiv.org/abs/2206.04615Srivastava, A., A. Rastogi, A. Rao, et al. 2022. “Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models”. In arXiv. Preprint, June 9. https://arxiv.org/abs/2206.04615.Srivastava, A., et al. “Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models”. arXiv, 9 June 2022, https://arxiv.org/abs/2206.04615.Srivastava, A. et al. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models. arXiv Preprint at https://arxiv.org/abs/2206.04615 (2022).A. Srivastava et al., “Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models”, Jun. 09, 2022. [Online]. Available: https://arxiv.org/abs/2206.04615
Stanford(2024). The 2024 AI Index Report. Stanford HAI.Stanford. (2024). The 2024 AI Index Report. Stanford HAI. https://hai.stanford.edu/ai-index/2024-ai-index-reportStanford. 2024. “The 2024 AI Index Report”. Stanford HAI. https://hai.stanford.edu/ai-index/2024-ai-index-report.Stanford. “The 2024 AI Index Report”. Stanford HAI, 2024, https://hai.stanford.edu/ai-index/2024-ai-index-report.Stanford. The 2024 AI Index Report. Stanford HAI https://hai.stanford.edu/ai-index/2024-ai-index-report (2024).Stanford, “The 2024 AI Index Report”, Stanford HAI. [Online]. Available: https://hai.stanford.edu/ai-index/2024-ai-index-report
Stanford HAI(2024). AI Index | Stanford HAI.Stanford HAI. (2024). AI Index | Stanford HAI. https://aiindex.stanford.edu/reportStanford HAI. 2024. “AI Index | Stanford HAI”. https://aiindex.stanford.edu/report.Stanford HAI. AI Index | Stanford HAI. 2024, https://aiindex.stanford.edu/report.Stanford HAI. AI Index | Stanford HAI. https://aiindex.stanford.edu/report (2024).Stanford HAI, “AI Index | Stanford HAI”. [Online]. Available: https://aiindex.stanford.edu/report
Stanford HAI(2025). The 2025 AI Index Report. Stanford HAI.Stanford HAI. (2025). The 2025 AI Index Report. Stanford HAI. https://hai.stanford.edu/ai-index/2025-ai-index-reportStanford HAI. 2025. “The 2025 AI Index Report”. Stanford HAI. https://hai.stanford.edu/ai-index/2025-ai-index-report.Stanford HAI. “The 2025 AI Index Report”. Stanford HAI, 2025, https://hai.stanford.edu/ai-index/2025-ai-index-report.Stanford HAI. The 2025 AI Index Report. Stanford HAI https://hai.stanford.edu/ai-index/2025-ai-index-report (2025).Stanford HAI, “The 2025 AI Index Report”, Stanford HAI. [Online]. Available: https://hai.stanford.edu/ai-index/2025-ai-index-report
Stanford Institute for Human-Centered Artificial Intelligence(2024). Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights.Stanford Institute for Human-Centered Artificial Intelligence. (2024). Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights (Docket No. 240216-0052). Stanford Institute for Human-Centered Artificial Intelligence. https://hai.stanford.edu/sites/default/files/2024-03/Response-NTIA-RFC-Open-Foundation-Models.pdfStanford Institute for Human-Centered Artificial Intelligence. 2024. Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights. Docket No. 240216-0052. Stanford Institute for Human-Centered Artificial Intelligence. https://hai.stanford.edu/sites/default/files/2024-03/Response-NTIA-RFC-Open-Foundation-Models.pdf.Stanford Institute for Human-Centered Artificial Intelligence. Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights. Docket No. 240216-0052, Stanford Institute for Human-Centered Artificial Intelligence, 27 Mar. 2024, https://hai.stanford.edu/sites/default/files/2024-03/Response-NTIA-RFC-Open-Foundation-Models.pdf.Stanford Institute for Human-Centered Artificial Intelligence. Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights. https://hai.stanford.edu/sites/default/files/2024-03/Response-NTIA-RFC-Open-Foundation-Models.pdf (2024).Stanford Institute for Human-Centered Artificial Intelligence, “Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights”, Stanford Institute for Human-Centered Artificial Intelligence, Docket No. 240216-0052, Mar. 2024. [Online]. Available: https://hai.stanford.edu/sites/default/files/2024-03/Response-NTIA-RFC-Open-Foundation-Models.pdf
Stanford Institute for Human-Centered Artificial Intelligence(2025). The AI Index 2025 Annual Report.Stanford Institute for Human-Centered Artificial Intelligence. (2025). The AI Index 2025 Annual Report. Stanford Institute for Human-Centered Artificial Intelligence. https://doi.org/10.48550/arXiv.2504.07139Stanford Institute for Human-Centered Artificial Intelligence. 2025. The AI Index 2025 Annual Report. Stanford Institute for Human-Centered Artificial Intelligence. https://doi.org/10.48550/arXiv.2504.07139.Stanford Institute for Human-Centered Artificial Intelligence. The AI Index 2025 Annual Report. Stanford Institute for Human-Centered Artificial Intelligence, Apr. 2025, https://doi.org/10.48550/arXiv.2504.07139.Stanford Institute for Human-Centered Artificial Intelligence. The AI Index 2025 Annual Report. https://hai.stanford.edu/assets/files/hai_ai_index_report_2025.pdf (2025) doi:10.48550/arXiv.2504.07139.Stanford Institute for Human-Centered Artificial Intelligence, “The AI Index 2025 Annual Report”, Stanford Institute for Human-Centered Artificial Intelligence, Stanford, CA, Apr. 2025. doi: 10.48550/arXiv.2504.07139.
State of AI Report(2025). State of AI Report 2025. State of AI Report.State of AI Report. (2025, October 9). State of AI Report 2025. State of AI Report. https://stateof.aiState of AI Report. 2025. “State of AI Report 2025”. State of AI Report, October 9. https://stateof.ai.State of AI Report. “State of AI Report 2025”. State of AI Report, 9 Oct. 2025, https://stateof.ai.State of AI Report. State of AI Report 2025. State of AI Report https://stateof.ai (2025).State of AI Report, “State of AI Report 2025”, State of AI Report. [Online]. Available: https://stateof.ai
Stephanie Lin, Jacob Hilton & Owain Evans(2021). TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv.Stephanie Lin, Jacob Hilton, & Owain Evans. (2021). TruthfulQA: Measuring How Models Mimic Human Falsehoods. In arXiv. https://arxiv.org/abs/2109.07958Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. “TruthfulQA: Measuring How Models Mimic Human Falsehoods”. In arXiv. Preprint, September 8. https://arxiv.org/abs/2109.07958.Stephanie Lin, et al. “TruthfulQA: Measuring How Models Mimic Human Falsehoods”. arXiv, 8 Sept. 2021, https://arxiv.org/abs/2109.07958.Stephanie Lin, Jacob Hilton & Owain Evans. TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv Preprint at https://arxiv.org/abs/2109.07958 (2021).Stephanie Lin, Jacob Hilton, and Owain Evans, “TruthfulQA: Measuring How Models Mimic Human Falsehoods”, Sep. 08, 2021. [Online]. Available: https://arxiv.org/abs/2109.07958
Stewart, A. J. & Plotkin, J. B.(2012). Extortion and cooperation in the Prisoner’s Dilemma. Proceedings of the National Academy of Sciences.Stewart, A. J., & Plotkin, J. B. (2012). Extortion and cooperation in the Prisoner’s Dilemma. Proceedings of the National Academy of Sciences. https://doi.org/10.1073/pnas.1208087109Stewart, A. J., and J. B. Plotkin. 2012. “Extortion and Cooperation in the Prisoner’s Dilemma”. Proceedings of the National Academy of Sciences, ahead of print, June 18. https://doi.org/10.1073/pnas.1208087109.Stewart, A. J., and J. B. Plotkin. “Extortion and Cooperation in the Prisoner’s Dilemma”. Proceedings of the National Academy of Sciences, June 2012, https://doi.org/10.1073/pnas.1208087109.Stewart, A. J. & Plotkin, J. B. Extortion and cooperation in the Prisoner’s Dilemma. Proceedings of the National Academy of Sciences https://doi.org/10.1073/pnas.1208087109 (2012) doi:10.1073/pnas.1208087109.A. J. Stewart and J. B. Plotkin, “Extortion and cooperation in the Prisoner’s Dilemma”, Proceedings of the National Academy of Sciences, Jun. 2012, doi: 10.1073/pnas.1208087109.
Stix, C. et al.(2025). AI Behind Closed Doors: a Primer on The Governance of Internal Deployment. arXiv.Stix, C., Pistillo, M., Sastry, G., Hobbhahn, M., Ortega, A., Balesni, M., Hallensleben, A., Goldowsky-Dill, N., & Sharkey, L. (2025). AI Behind Closed Doors: a Primer on The Governance of Internal Deployment. In arXiv. https://arxiv.org/abs/2504.12170Stix, C., M. Pistillo, G. Sastry, et al. 2025. “AI Behind Closed Doors: A Primer on The Governance of Internal Deployment”. In arXiv. Preprint, April 16. https://arxiv.org/abs/2504.12170.Stix, C., et al. “AI Behind Closed Doors: A Primer on The Governance of Internal Deployment”. arXiv, 16 Apr. 2025, https://arxiv.org/abs/2504.12170.Stix, C. et al. AI Behind Closed Doors: a Primer on The Governance of Internal Deployment. arXiv Preprint at https://arxiv.org/abs/2504.12170 (2025).C. Stix et al., “AI Behind Closed Doors: a Primer on The Governance of Internal Deployment”, Apr. 16, 2025. [Online]. Available: https://arxiv.org/abs/2504.12170
Sunishchal Dev & Marius Hobbhahn(2024). Improving Model-Written Evals for AI Safety Benchmarking. AI Alignment Forum.Sunishchal Dev, & Marius Hobbhahn. (2024, October 15). Improving Model-Written Evals for AI Safety Benchmarking. AI Alignment Forum. https://alignmentforum.org/posts/yxdHp2cZeQbZGREEN/improving-model-written-evals-for-ai-safety-benchmarkingSunishchal Dev, and Marius Hobbhahn. 2024. “Improving Model-Written Evals for AI Safety Benchmarking”. AI Alignment Forum, October 15. https://alignmentforum.org/posts/yxdHp2cZeQbZGREEN/improving-model-written-evals-for-ai-safety-benchmarking.Sunishchal Dev, and Marius Hobbhahn. “Improving Model-Written Evals for AI Safety Benchmarking”. AI Alignment Forum, 15 Oct. 2024, https://alignmentforum.org/posts/yxdHp2cZeQbZGREEN/improving-model-written-evals-for-ai-safety-benchmarking.Sunishchal Dev & Marius Hobbhahn. Improving Model-Written Evals for AI Safety Benchmarking. AI Alignment Forum https://alignmentforum.org/posts/yxdHp2cZeQbZGREEN/improving-model-written-evals-for-ai-safety-benchmarking (2024).Sunishchal Dev and Marius Hobbhahn, “Improving Model-Written Evals for AI Safety Benchmarking”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/yxdHp2cZeQbZGREEN/improving-model-written-evals-for-ai-safety-benchmarking
Sutton(2019). The Bitter Lesson.Sutton. (2019). The Bitter Lesson. Internet Archive (https://web.archive.org/web/20260919194030/http://www.incompleteideas.net/IncIdeas/BitterLesson.html). https://incompleteideas.net/IncIdeas/BitterLesson.htmlSutton. 2019. “The Bitter Lesson”. Https://web.archive.org/web/20260919194030/http://www.incompleteideas.net/IncIdeas/BitterLesson.html. Internet Archive. https://incompleteideas.net/IncIdeas/BitterLesson.html.Sutton. The Bitter Lesson. 2019, Internet Archive, https://web.archive.org/web/20260919194030/http://www.incompleteideas.net/IncIdeas/BitterLesson.html, https://incompleteideas.net/IncIdeas/BitterLesson.html.Sutton. The Bitter Lesson. https://incompleteideas.net/IncIdeas/BitterLesson.html (2019).Sutton, “The Bitter Lesson”. Accessed: Sep. 19, 2026. [Online]. Available: https://incompleteideas.net/IncIdeas/BitterLesson.html
Sutton, R. S.(2023). AI Succession. World Artificial Intelligence Conference.Sutton, R. S. (2023, July 7). AI Succession. In World Artificial Intelligence Conference [Keynote address]. https://incompleteideas.net/Talks/waic3.pdfSutton, R. S. 2023. “AI Succession”. In World Artificial Intelligence Conference, Keynote address. July 7. https://incompleteideas.net/Talks/waic3.pdf.Sutton, R. S. “AI Succession”. Keynote address. World Artificial Intelligence Conference, https://incompleteideas.net/Talks/waic3.pdf.Sutton, R. S. AI Succession. World Artificial Intelligence Conference (2023).R. S. Sutton, “AI Succession”, in World Artificial Intelligence Conference, Jul. 07, 2023. [Online]. Available: https://incompleteideas.net/Talks/waic3.pdf
Swenson & Chan(2024). Election disinformation takes a big leap with AI being used to deceive worldwide. AP News.Swenson & Chan. (2024, March 14). Election disinformation takes a big leap with AI being used to deceive worldwide. AP News. https://apnews.com/article/artificial-intelligence-elections-disinformation-chatgpt-bc283e7426402f0b4baa7df280a4c3fdSwenson & Chan. 2024. “Election Disinformation Takes a Big Leap with AI Being Used to Deceive Worldwide”. AP News, March 14. https://apnews.com/article/artificial-intelligence-elections-disinformation-chatgpt-bc283e7426402f0b4baa7df280a4c3fd.Swenson & Chan. “Election Disinformation Takes a Big Leap with AI Being Used to Deceive Worldwide”. AP News, 14 Mar. 2024, https://apnews.com/article/artificial-intelligence-elections-disinformation-chatgpt-bc283e7426402f0b4baa7df280a4c3fd.Swenson & Chan. Election disinformation takes a big leap with AI being used to deceive worldwide. AP News https://apnews.com/article/artificial-intelligence-elections-disinformation-chatgpt-bc283e7426402f0b4baa7df280a4c3fd (2024).Swenson & Chan, “Election disinformation takes a big leap with AI being used to deceive worldwide”, AP News. [Online]. Available: https://apnews.com/article/artificial-intelligence-elections-disinformation-chatgpt-bc283e7426402f0b4baa7df280a4c3fd
Takemoto, K.(2024). All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks. arXiv.Takemoto, K. (2024). All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks. In arXiv. https://doi.org/10.3390/app14093558Takemoto, K. 2024. “All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks”. In arXiv. Preprint, January 18. https://doi.org/10.3390/app14093558.Takemoto, K. “All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks”. arXiv, 18 Jan. 2024, https://doi.org/10.3390/app14093558.Takemoto, K. All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks. arXiv Preprint at https://doi.org/10.3390/app14093558 (2024).K. Takemoto, “All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks”, Jan. 18, 2024. doi: 10.3390/app14093558.
Tallberg et al.(2023). The Global Governance of Artificial Intelligence: Next Steps for Empirical and Normative Research.Tallberg et al. (2023). The Global Governance of Artificial Intelligence: Next Steps for Empirical and Normative Research. Internet Archive (https://web.archive.org/web/20240514203326/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4424123). https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4424123Tallberg et al. 2023. “The Global Governance of Artificial Intelligence: Next Steps for Empirical and Normative Research”. Https://web.archive.org/web/20240514203326/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4424123. Internet Archive. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4424123.Tallberg et al. The Global Governance of Artificial Intelligence: Next Steps for Empirical and Normative Research. 2023, Internet Archive, https://web.archive.org/web/20240514203326/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4424123, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4424123.Tallberg et al. The Global Governance of Artificial Intelligence: Next Steps for Empirical and Normative Research. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4424123 (2023).Tallberg et al., “The Global Governance of Artificial Intelligence: Next Steps for Empirical and Normative Research”. Accessed: May 14, 2024. [Online]. Available: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4424123
tamera(2022). Externalized reasoning oversight: a research direction for language model alignment. AI Alignment Forum.tamera. (2022, August 3). Externalized reasoning oversight: a research direction for language model alignment. AI Alignment Forum. https://alignmentforum.org/posts/FRRb6Gqem8k69ocbi/externalized-reasoning-oversight-a-research-direction-fortamera. 2022. “Externalized Reasoning Oversight: A Research Direction for Language Model Alignment”. AI Alignment Forum, August 3. https://alignmentforum.org/posts/FRRb6Gqem8k69ocbi/externalized-reasoning-oversight-a-research-direction-for.tamera. “Externalized Reasoning Oversight: A Research Direction for Language Model Alignment”. AI Alignment Forum, 3 Aug. 2022, https://alignmentforum.org/posts/FRRb6Gqem8k69ocbi/externalized-reasoning-oversight-a-research-direction-for.tamera. Externalized reasoning oversight: a research direction for language model alignment. AI Alignment Forum https://alignmentforum.org/posts/FRRb6Gqem8k69ocbi/externalized-reasoning-oversight-a-research-direction-for (2022).tamera, “Externalized reasoning oversight: a research direction for language model alignment”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/FRRb6Gqem8k69ocbi/externalized-reasoning-oversight-a-research-direction-for
Tamsin Leake & JuliaHP(2023). formalizing the QACI alignment formal-goal. AI Alignment Forum.Tamsin Leake, & JuliaHP. (2023, June 10). formalizing the QACI alignment formal-goal. AI Alignment Forum. https://alignmentforum.org/posts/MR5wJpE27ymE7M7iv/formalizing-the-qaci-alignment-formal-goalTamsin Leake, and JuliaHP. 2023. “Formalizing the QACI Alignment Formal-goal”. AI Alignment Forum, June 10. https://alignmentforum.org/posts/MR5wJpE27ymE7M7iv/formalizing-the-qaci-alignment-formal-goal.Tamsin Leake, and JuliaHP. “Formalizing the QACI Alignment Formal-goal”. AI Alignment Forum, 10 June 2023, https://alignmentforum.org/posts/MR5wJpE27ymE7M7iv/formalizing-the-qaci-alignment-formal-goal.Tamsin Leake & JuliaHP. formalizing the QACI alignment formal-goal. AI Alignment Forum https://alignmentforum.org/posts/MR5wJpE27ymE7M7iv/formalizing-the-qaci-alignment-formal-goal (2023).Tamsin Leake and JuliaHP, “formalizing the QACI alignment formal-goal”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/MR5wJpE27ymE7M7iv/formalizing-the-qaci-alignment-formal-goal
Tanzer, G., Suzgun, M., Visser, E., Jurafsky, D. & Melas-Kyriazi, L.(2023). A Benchmark for Learning to Translate a New Language from One Grammar Book. arXiv.Tanzer, G., Suzgun, M., Visser, E., Jurafsky, D., & Melas-Kyriazi, L. (2023). A Benchmark for Learning to Translate a New Language from One Grammar Book. In arXiv. https://arxiv.org/abs/2309.16575Tanzer, G., M. Suzgun, E. Visser, D. Jurafsky, and L. Melas-Kyriazi. 2023. “A Benchmark for Learning to Translate a New Language from One Grammar Book”. In arXiv. Preprint, September 28. https://arxiv.org/abs/2309.16575.Tanzer, G., et al. “A Benchmark for Learning to Translate a New Language from One Grammar Book”. arXiv, 28 Sept. 2023, https://arxiv.org/abs/2309.16575.Tanzer, G., Suzgun, M., Visser, E., Jurafsky, D. & Melas-Kyriazi, L. A Benchmark for Learning to Translate a New Language from One Grammar Book. arXiv Preprint at https://arxiv.org/abs/2309.16575 (2023).G. Tanzer, M. Suzgun, E. Visser, D. Jurafsky, and L. Melas-Kyriazi, “A Benchmark for Learning to Translate a New Language from One Grammar Book”, Sep. 28, 2023. [Online]. Available: https://arxiv.org/abs/2309.16575
Team, G. et al.(2023). Gemini: A Family of Highly Capable Multimodal Models. arXiv.Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., Silver, D., Johnson, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., Lillicrap, T., Lazaridou, A., … Vinyals, O. (2023). Gemini: A Family of Highly Capable Multimodal Models. In arXiv. https://arxiv.org/abs/2312.11805Team, G., R. Anil, S. Borgeaud, et al. 2023. “Gemini: A Family of Highly Capable Multimodal Models”. In arXiv. Preprint, December 19. https://arxiv.org/abs/2312.11805.Team, G., et al. “Gemini: A Family of Highly Capable Multimodal Models”. arXiv, 19 Dec. 2023, https://arxiv.org/abs/2312.11805.Team, G. et al. Gemini: A Family of Highly Capable Multimodal Models. arXiv Preprint at https://arxiv.org/abs/2312.11805 (2023).G. Team et al., “Gemini: A Family of Highly Capable Multimodal Models”, Dec. 19, 2023. [Online]. Available: https://arxiv.org/abs/2312.11805
TED(2017). The rise of the useless class. ideas.ted.com.TED. (2017, February 24). The rise of the useless class. Ideas.ted.com. https://ideas.ted.com/the-rise-of-the-useless-classTED. 2017. “The Rise of the Useless Class”. Ideas.ted.com, February 24. https://ideas.ted.com/the-rise-of-the-useless-class.TED. “The Rise of the Useless Class”. Ideas.ted.com, 24 Feb. 2017, https://ideas.ted.com/the-rise-of-the-useless-class.TED. The rise of the useless class. ideas.ted.com https://ideas.ted.com/the-rise-of-the-useless-class (2017).TED, “The rise of the useless class”, ideas.ted.com. [Online]. Available: https://ideas.ted.com/the-rise-of-the-useless-class
Tegmark & Omohundro(2023). Provably safe systems: the only path to controllable AGI. arXiv.Tegmark & Omohundro. (2023). Provably safe systems: the only path to controllable AGI. In arXiv. https://arxiv.org/abs/2309.01933Tegmark & Omohundro. 2023. “Provably Safe Systems: The Only Path to Controllable AGI”. In arXiv. Preprint, September 5. https://arxiv.org/abs/2309.01933.Tegmark & Omohundro. “Provably Safe Systems: The Only Path to Controllable AGI”. arXiv, 5 Sept. 2023, https://arxiv.org/abs/2309.01933.Tegmark & Omohundro. Provably safe systems: the only path to controllable AGI. arXiv Preprint at https://arxiv.org/abs/2309.01933 (2023).Tegmark & Omohundro, “Provably safe systems: the only path to controllable AGI”, Sep. 05, 2023. [Online]. Available: https://arxiv.org/abs/2309.01933
Tegmark(2017). Life 3.0 Book Summary by Max Tegmark.Tegmark. (2017). Life 3.0 Book Summary by Max Tegmark. https://shortform.com/summary/life-3-0-summary-max-tegmarkTegmark. 2017. “Life 3.0 Book Summary by Max Tegmark”. https://shortform.com/summary/life-3-0-summary-max-tegmark.Tegmark. Life 3.0 Book Summary by Max Tegmark. 2017, https://shortform.com/summary/life-3-0-summary-max-tegmark.Tegmark. Life 3.0 Book Summary by Max Tegmark. https://shortform.com/summary/life-3-0-summary-max-tegmark (2017).Tegmark, “Life 3.0 Book Summary by Max Tegmark”. [Online]. Available: https://shortform.com/summary/life-3-0-summary-max-tegmark
Tegmark(2023). The 'Don't Look Up' Thinking That Could Doom Us With AI. TIME.Tegmark. (2023, April 25). The 'Don't Look Up' Thinking That Could Doom Us With AI. TIME. https://time.com/6273743/thinking-that-could-doom-us-with-aiTegmark. 2023. “The 'Don't Look Up' Thinking That Could Doom Us With AI”. TIME, April 25. https://time.com/6273743/thinking-that-could-doom-us-with-ai.Tegmark. “The 'Don't Look Up' Thinking That Could Doom Us With AI”. TIME, 25 Apr. 2023, https://time.com/6273743/thinking-that-could-doom-us-with-ai.Tegmark. The 'Don't Look Up' Thinking That Could Doom Us With AI. TIME https://time.com/6273743/thinking-that-could-doom-us-with-ai (2023).Tegmark, “The 'Don't Look Up' Thinking That Could Doom Us With AI”, TIME. [Online]. Available: https://time.com/6273743/thinking-that-could-doom-us-with-ai
Thang Luong & Edward Lockhart(2025). Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad.Thang Luong, & Edward Lockhart. (2025, July 21). Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad. https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiadThang Luong, and Edward Lockhart. 2025. “Advanced Version of Gemini with Deep Think Officially Achieves Gold-medal Standard at the International Mathematical Olympiad”. July 21. https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad.Thang Luong, and Edward Lockhart. Advanced Version of Gemini with Deep Think Officially Achieves Gold-medal Standard at the International Mathematical Olympiad. 21 July 2025, https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad.Thang Luong & Edward Lockhart. Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad. https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad (2025).Thang Luong and Edward Lockhart, “Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad”. [Online]. Available: https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad
The Bulletin(2024). MIT researchers ordered and combined parts of the 1918 pandemic influenza virus. Did they expose a security flaw?. Bulletin of the Atomic Scientists.The Bulletin. (2024, June 3). MIT researchers ordered and combined parts of the 1918 pandemic influenza virus. Did they expose a security flaw?. Bulletin of the Atomic Scientists. https://thebulletin.org/2024/06/mit-researchers-ordered-and-combined-parts-of-the-1918-pandemic-influenza-virus-did-they-expose-a-security-flawThe Bulletin. 2024. “MIT Researchers Ordered and Combined Parts of the 1918 Pandemic Influenza Virus. Did They Expose a Security Flaw?”. Bulletin of the Atomic Scientists, June 3. https://thebulletin.org/2024/06/mit-researchers-ordered-and-combined-parts-of-the-1918-pandemic-influenza-virus-did-they-expose-a-security-flaw.The Bulletin. “MIT Researchers Ordered and Combined Parts of the 1918 Pandemic Influenza Virus. Did They Expose a Security Flaw?”. Bulletin of the Atomic Scientists, 3 June 2024, https://thebulletin.org/2024/06/mit-researchers-ordered-and-combined-parts-of-the-1918-pandemic-influenza-virus-did-they-expose-a-security-flaw.The Bulletin. MIT researchers ordered and combined parts of the 1918 pandemic influenza virus. Did they expose a security flaw?. Bulletin of the Atomic Scientists https://thebulletin.org/2024/06/mit-researchers-ordered-and-combined-parts-of-the-1918-pandemic-influenza-virus-did-they-expose-a-security-flaw (2024).The Bulletin, “MIT researchers ordered and combined parts of the 1918 pandemic influenza virus. Did they expose a security flaw?”, Bulletin of the Atomic Scientists. [Online]. Available: https://thebulletin.org/2024/06/mit-researchers-ordered-and-combined-parts-of-the-1918-pandemic-influenza-virus-did-they-expose-a-security-flaw
The Guardian(2024). Ilya: the AI scientist shaping the world. YouTube.The Guardian. (2024). Ilya: the AI scientist shaping the world [Video recording]. In YouTube. https://www.youtube.com/watch?v=9iqn1HhFJ6cThe Guardian. 2024. “Ilya: The AI Scientist Shaping the World”. YouTube. https://www.youtube.com/watch?v=9iqn1HhFJ6c.The Guardian. “Ilya: The AI Scientist Shaping the World”. YouTube, 2024, https://www.youtube.com/watch?v=9iqn1HhFJ6c.The Guardian. Ilya: The AI Scientist Shaping the World. YouTube (2024).The Guardian, Ilya: the AI scientist shaping the world, (2024). [Online Video]. Available: https://www.youtube.com/watch?v=9iqn1HhFJ6c
The Human Podcast(2025). Why I'm Hosting Debates on AI 'Doom' - Liron Shapira. YouTube.The Human Podcast. (2025). Why I'm Hosting Debates on AI 'Doom' - Liron Shapira [Video recording]. In YouTube. https://www.youtube.com/watch?v=0KmaotctziEThe Human Podcast. 2025. “Why I'm Hosting Debates on AI 'Doom' - Liron Shapira”. YouTube. https://www.youtube.com/watch?v=0KmaotctziE.The Human Podcast. “Why I'm Hosting Debates on AI 'Doom' - Liron Shapira”. YouTube, 2025, https://www.youtube.com/watch?v=0KmaotctziE.The Human Podcast. Why I'm Hosting Debates on AI 'Doom' - Liron Shapira. YouTube (2025).The Human Podcast, Why I'm Hosting Debates on AI 'Doom' - Liron Shapira, (2025). [Online Video]. Available: https://www.youtube.com/watch?v=0KmaotctziE
Tian, K., Mitchell, E., Yao, H., Manning, C. D. & Finn, C.(2023). Fine-tuning Language Models for Factuality. arXiv.Tian, K., Mitchell, E., Yao, H., Manning, C. D., & Finn, C. (2023). Fine-tuning Language Models for Factuality. In arXiv. https://arxiv.org/abs/2311.08401Tian, K., E. Mitchell, H. Yao, C. D. Manning, and C. Finn. 2023. “Fine-tuning Language Models for Factuality”. In arXiv. Preprint, November 14. https://arxiv.org/abs/2311.08401.Tian, K., et al. “Fine-tuning Language Models for Factuality”. arXiv, 14 Nov. 2023, https://arxiv.org/abs/2311.08401.Tian, K., Mitchell, E., Yao, H., Manning, C. D. & Finn, C. Fine-tuning Language Models for Factuality. arXiv Preprint at https://arxiv.org/abs/2311.08401 (2023).K. Tian, E. Mitchell, H. Yao, C. D. Manning, and C. Finn, “Fine-tuning Language Models for Factuality”, Nov. 14, 2023. [Online]. Available: https://arxiv.org/abs/2311.08401
Tihanyi, N., Ferrag, M. A., Jain, R., Bisztray, T. & Debbah, M.(2024). CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge. arXiv.Tihanyi, N., Ferrag, M. A., Jain, R., Bisztray, T., & Debbah, M. (2024). CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge. In arXiv. https://arxiv.org/abs/2402.07688Tihanyi, N., M. A. Ferrag, R. Jain, T. Bisztray, and M. Debbah. 2024. “CyberMetric: A Benchmark Dataset Based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge”. In arXiv. Preprint, February 12. https://arxiv.org/abs/2402.07688.Tihanyi, N., et al. “CyberMetric: A Benchmark Dataset Based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge”. arXiv, 12 Feb. 2024, https://arxiv.org/abs/2402.07688.Tihanyi, N., Ferrag, M. A., Jain, R., Bisztray, T. & Debbah, M. CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge. arXiv Preprint at https://arxiv.org/abs/2402.07688 (2024).N. Tihanyi, M. A. Ferrag, R. Jain, T. Bisztray, and M. Debbah, “CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge”, Feb. 12, 2024. [Online]. Available: https://arxiv.org/abs/2402.07688
TIME(2024). How Anthropic Designed Itself to Avoid OpenAI’s Mistakes. TIME.TIME. (2024, May 30). How Anthropic Designed Itself to Avoid OpenAI’s Mistakes. TIME. https://time.com/6983420/anthropic-structure-openai-incentivesTIME. 2024. “How Anthropic Designed Itself to Avoid OpenAI’s Mistakes”. TIME, May 30. https://time.com/6983420/anthropic-structure-openai-incentives.TIME. “How Anthropic Designed Itself to Avoid OpenAI’s Mistakes”. TIME, 30 May 2024, https://time.com/6983420/anthropic-structure-openai-incentives.TIME. How Anthropic Designed Itself to Avoid OpenAI’s Mistakes. TIME https://time.com/6983420/anthropic-structure-openai-incentives (2024).TIME, “How Anthropic Designed Itself to Avoid OpenAI’s Mistakes”, TIME. [Online]. Available: https://time.com/6983420/anthropic-structure-openai-incentives
Time Magazine(2023). Microsoft's AI chatbot, TayTweets, suffers another meltdown. CBC.Time Magazine. (2023). Microsoft's AI chatbot, TayTweets, suffers another meltdown. Internet Archive (https://web.archive.org/web/20250716211155/https://www.cbc.ca/radio/day6/episode-279-playing-ball-on-grass-vs-turf-taytweets-big-fail-narco-subs-fake-food-and-more-1.3514966/microsoft-s-ai-chatbot-taytweets-suffers-another-meltdown-1.3515046). CBC. https://cbc.ca/radio/day6/episode-279-playing-ball-on-grass-vs-turf-taytweets-big-fail-narco-subs-fake-food-and-more-1.3514966/microsoft-s-ai-chatbot-taytweets-suffers-another-meltdown-1.3515046Time Magazine. 2023. “Microsoft's AI Chatbot, TayTweets, Suffers Another Meltdown”. CBC. Https://web.archive.org/web/20250716211155/https://www.cbc.ca/radio/day6/episode-279-playing-ball-on-grass-vs-turf-taytweets-big-fail-narco-subs-fake-food-and-more-1.3514966/microsoft-s-ai-chatbot-taytweets-suffers-another-meltdown-1.3515046. Internet Archive. https://cbc.ca/radio/day6/episode-279-playing-ball-on-grass-vs-turf-taytweets-big-fail-narco-subs-fake-food-and-more-1.3514966/microsoft-s-ai-chatbot-taytweets-suffers-another-meltdown-1.3515046.Time Magazine. “Microsoft's AI Chatbot, TayTweets, Suffers Another Meltdown”. CBC, 2023, Internet Archive, https://web.archive.org/web/20250716211155/https://www.cbc.ca/radio/day6/episode-279-playing-ball-on-grass-vs-turf-taytweets-big-fail-narco-subs-fake-food-and-more-1.3514966/microsoft-s-ai-chatbot-taytweets-suffers-another-meltdown-1.3515046, https://cbc.ca/radio/day6/episode-279-playing-ball-on-grass-vs-turf-taytweets-big-fail-narco-subs-fake-food-and-more-1.3514966/microsoft-s-ai-chatbot-taytweets-suffers-another-meltdown-1.3515046.Time Magazine. Microsoft's AI chatbot, TayTweets, suffers another meltdown. CBC https://cbc.ca/radio/day6/episode-279-playing-ball-on-grass-vs-turf-taytweets-big-fail-narco-subs-fake-food-and-more-1.3514966/microsoft-s-ai-chatbot-taytweets-suffers-another-meltdown-1.3515046 (2023).Time Magazine, “Microsoft's AI chatbot, TayTweets, suffers another meltdown”, CBC. Accessed: Jul. 16, 2025. [Online]. Available: https://cbc.ca/radio/day6/episode-279-playing-ball-on-grass-vs-turf-taytweets-big-fail-narco-subs-fake-food-and-more-1.3514966/microsoft-s-ai-chatbot-taytweets-suffers-another-meltdown-1.3515046
Tom Davidson(2024). Takeoff speeds presentation at Anthropic. AI Alignment Forum.Tom Davidson. (2024, June 4). Takeoff speeds presentation at Anthropic. AI Alignment Forum. https://alignmentforum.org/posts/Nsmabb9fhpLuLdtLE/takeoff-speeds-presentation-at-anthropicTom Davidson. 2024. “Takeoff Speeds Presentation at Anthropic”. AI Alignment Forum, June 4. https://alignmentforum.org/posts/Nsmabb9fhpLuLdtLE/takeoff-speeds-presentation-at-anthropic.Tom Davidson. “Takeoff Speeds Presentation at Anthropic”. AI Alignment Forum, 4 June 2024, https://alignmentforum.org/posts/Nsmabb9fhpLuLdtLE/takeoff-speeds-presentation-at-anthropic.Tom Davidson. Takeoff speeds presentation at Anthropic. AI Alignment Forum https://alignmentforum.org/posts/Nsmabb9fhpLuLdtLE/takeoff-speeds-presentation-at-anthropic (2024).Tom Davidson, “Takeoff speeds presentation at Anthropic”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/Nsmabb9fhpLuLdtLE/takeoff-speeds-presentation-at-anthropic
Tom Davidson(2025). Human takeover might be worse than AI takeover. LessWrong.Tom Davidson. (2025, January 10). Human takeover might be worse than AI takeover. LessWrong. https://lesswrong.com/posts/FEcw6JQ8surwxvRfr/human-takeover-might-be-worse-than-ai-takeoverTom Davidson. 2025. “Human Takeover Might Be Worse Than AI Takeover”. LessWrong, January 10. https://lesswrong.com/posts/FEcw6JQ8surwxvRfr/human-takeover-might-be-worse-than-ai-takeover.Tom Davidson. “Human Takeover Might Be Worse Than AI Takeover”. LessWrong, 10 Jan. 2025, https://lesswrong.com/posts/FEcw6JQ8surwxvRfr/human-takeover-might-be-worse-than-ai-takeover.Tom Davidson. Human takeover might be worse than AI takeover. LessWrong https://lesswrong.com/posts/FEcw6JQ8surwxvRfr/human-takeover-might-be-worse-than-ai-takeover (2025).Tom Davidson, “Human takeover might be worse than AI takeover”, LessWrong. [Online]. Available: https://lesswrong.com/posts/FEcw6JQ8surwxvRfr/human-takeover-might-be-worse-than-ai-takeover
Tom Davidson, Lukas Finnveden & Rose Hadshar(2025). AI-Enabled Coups: How a Small Group Could Use AI to Seize Power.Tom Davidson, Lukas Finnveden, & Rose Hadshar. (2025, April 15). AI-Enabled Coups: How a Small Group Could Use AI to Seize Power. https://forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-powerTom Davidson, Lukas Finnveden, and Rose Hadshar. 2025. “AI-Enabled Coups: How a Small Group Could Use AI to Seize Power”. April 15. https://forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power.Tom Davidson, et al. AI-Enabled Coups: How a Small Group Could Use AI to Seize Power. 15 Apr. 2025, https://forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power.Tom Davidson, Lukas Finnveden & Rose Hadshar. AI-Enabled Coups: How a Small Group Could Use AI to Seize Power. https://forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power (2025).Tom Davidson, Lukas Finnveden, and Rose Hadshar, “AI-Enabled Coups: How a Small Group Could Use AI to Seize Power”. [Online]. Available: https://forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power
Tom Davidson, Lukas Finnveden & rosehadshar(2025). AI-enabled coups: a small group could use AI to seize power. LessWrong.Tom Davidson, Lukas Finnveden, & rosehadshar. (2025, April 16). AI-enabled coups: a small group could use AI to seize power. LessWrong. https://lesswrong.com/posts/6kBMqrK9bREuGsrnd/ai-enabled-coups-a-small-group-could-use-ai-to-seize-power-1Tom Davidson, Lukas Finnveden, and rosehadshar. 2025. “AI-enabled Coups: A Small Group Could Use AI to Seize Power”. LessWrong, April 16. https://lesswrong.com/posts/6kBMqrK9bREuGsrnd/ai-enabled-coups-a-small-group-could-use-ai-to-seize-power-1.Tom Davidson, et al. “AI-enabled Coups: A Small Group Could Use AI to Seize Power”. LessWrong, 16 Apr. 2025, https://lesswrong.com/posts/6kBMqrK9bREuGsrnd/ai-enabled-coups-a-small-group-could-use-ai-to-seize-power-1.Tom Davidson, Lukas Finnveden & rosehadshar. AI-enabled coups: a small group could use AI to seize power. LessWrong https://lesswrong.com/posts/6kBMqrK9bREuGsrnd/ai-enabled-coups-a-small-group-could-use-ai-to-seize-power-1 (2025).Tom Davidson, Lukas Finnveden, and rosehadshar, “AI-enabled coups: a small group could use AI to seize power”, LessWrong. [Online]. Available: https://lesswrong.com/posts/6kBMqrK9bREuGsrnd/ai-enabled-coups-a-small-group-could-use-ai-to-seize-power-1
Tong(2023). What happens when your AI chatbot stops loving you back?. Reuters.Tong. (2023). What happens when your AI chatbot stops loving you back?. Internet Archive (https://web.archive.org/web/20250224001320/https://www.reuters.com/technology/what-happens-when-your-ai-chatbot-stops-loving-you-back-2023-03-18/). Reuters. https://reuters.com/technology/what-happens-when-your-ai-chatbot-stops-loving-you-back-2023-03-18Tong. 2023. “What Happens When Your AI Chatbot Stops Loving You Back?”. Reuters. Https://web.archive.org/web/20250224001320/https://www.reuters.com/technology/what-happens-when-your-ai-chatbot-stops-loving-you-back-2023-03-18/. Internet Archive. https://reuters.com/technology/what-happens-when-your-ai-chatbot-stops-loving-you-back-2023-03-18.Tong. “What Happens When Your AI Chatbot Stops Loving You Back?”. Reuters, 2023, Internet Archive, https://web.archive.org/web/20250224001320/https://www.reuters.com/technology/what-happens-when-your-ai-chatbot-stops-loving-you-back-2023-03-18/, https://reuters.com/technology/what-happens-when-your-ai-chatbot-stops-loving-you-back-2023-03-18.Tong. What happens when your AI chatbot stops loving you back?. Reuters https://reuters.com/technology/what-happens-when-your-ai-chatbot-stops-loving-you-back-2023-03-18 (2023).Tong, “What happens when your AI chatbot stops loving you back?”, Reuters. Accessed: Feb. 24, 2025. [Online]. Available: https://reuters.com/technology/what-happens-when-your-ai-chatbot-stops-loving-you-back-2023-03-18
Trajano & Ang(2023). We Need to Prevent a Global AI Arms Race Now. RSIS_NTU.Trajano & Ang. (2023). We Need to Prevent a Global AI Arms Race Now. RSIS_NTU. https://rsis.edu.sg/rsis-publication/rsis/we-need-to-prevent-a-global-ai-arms-race-nowTrajano & Ang. 2023. “We Need to Prevent a Global AI Arms Race Now”. RSIS_NTU. https://rsis.edu.sg/rsis-publication/rsis/we-need-to-prevent-a-global-ai-arms-race-now.Trajano & Ang. “We Need to Prevent a Global AI Arms Race Now”. RSIS_NTU, 2023, https://rsis.edu.sg/rsis-publication/rsis/we-need-to-prevent-a-global-ai-arms-race-now.Trajano & Ang. We Need to Prevent a Global AI Arms Race Now. RSIS_NTU https://rsis.edu.sg/rsis-publication/rsis/we-need-to-prevent-a-global-ai-arms-race-now (2023).Trajano & Ang, “We Need to Prevent a Global AI Arms Race Now”, RSIS_NTU. [Online]. Available: https://rsis.edu.sg/rsis-publication/rsis/we-need-to-prevent-a-global-ai-arms-race-now
Trusilo, D.(2023). Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability. Journal of Military Ethics.Trusilo, D. (2023). Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability. Journal of Military Ethics. https://doi.org/10.1080/15027570.2023.2213985Trusilo, D. 2023. “Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability”. Journal of Military Ethics, ahead of print, January 2. https://doi.org/10.1080/15027570.2023.2213985.Trusilo, D. “Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability”. Journal of Military Ethics, Jan. 2023, https://doi.org/10.1080/15027570.2023.2213985.Trusilo, D. Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability. Journal of Military Ethics https://doi.org/10.1080/15027570.2023.2213985 (2023) doi:10.1080/15027570.2023.2213985.D. Trusilo, “Autonomous AI Systems in Conflict: Emergent Behavior and Its Impact on Predictability and Reliability”, Journal of Military Ethics, Jan. 2023, doi: 10.1080/15027570.2023.2213985.
Truth Initiative(2017). the 5 ways tobacco companies lied about the dangers of smoking cigarettes. Truth Initiative.Truth Initiative. (2017). the 5 ways tobacco companies lied about the dangers of smoking cigarettes. Truth Initiative. https://truthinitiative.org/research-resources/tobacco-prevention-efforts/5-ways-tobacco-companies-lied-about-dangers-smokingTruth Initiative. 2017. “The 5 Ways Tobacco Companies Lied About the Dangers of Smoking Cigarettes”. Truth Initiative. https://truthinitiative.org/research-resources/tobacco-prevention-efforts/5-ways-tobacco-companies-lied-about-dangers-smoking.Truth Initiative. “The 5 Ways Tobacco Companies Lied About the Dangers of Smoking Cigarettes”. Truth Initiative, 2017, https://truthinitiative.org/research-resources/tobacco-prevention-efforts/5-ways-tobacco-companies-lied-about-dangers-smoking.Truth Initiative. the 5 ways tobacco companies lied about the dangers of smoking cigarettes. Truth Initiative https://truthinitiative.org/research-resources/tobacco-prevention-efforts/5-ways-tobacco-companies-lied-about-dangers-smoking (2017).Truth Initiative, “the 5 ways tobacco companies lied about the dangers of smoking cigarettes”, Truth Initiative. [Online]. Available: https://truthinitiative.org/research-resources/tobacco-prevention-efforts/5-ways-tobacco-companies-lied-about-dangers-smoking
Tsoy, N. & Konstantinov, N.(2024). Simplicity Bias of Two-Layer Networks beyond Linearly Separable Data. arXiv.Tsoy, N., & Konstantinov, N. (2024). Simplicity Bias of Two-Layer Networks beyond Linearly Separable Data. In arXiv. https://arxiv.org/abs/2405.17299Tsoy, N., and N. Konstantinov. 2024. “Simplicity Bias of Two-Layer Networks Beyond Linearly Separable Data”. In arXiv. Preprint, May 27. https://arxiv.org/abs/2405.17299.Tsoy, N., and N. Konstantinov. “Simplicity Bias of Two-Layer Networks Beyond Linearly Separable Data”. arXiv, 27 May 2024, https://arxiv.org/abs/2405.17299.Tsoy, N. & Konstantinov, N. Simplicity Bias of Two-Layer Networks beyond Linearly Separable Data. arXiv Preprint at https://arxiv.org/abs/2405.17299 (2024).N. Tsoy and N. Konstantinov, “Simplicity Bias of Two-Layer Networks beyond Linearly Separable Data”, May 27, 2024. [Online]. Available: https://arxiv.org/abs/2405.17299
Turing(1950). I.—COMPUTING MACHINERY AND INTELLIGENCE. OUP Academic.Turing. (1950). I.—COMPUTING MACHINERY AND INTELLIGENCE. Internet Archive (https://web.archive.org/web/20241203212535/https://academic.oup.com/mind/article/LIX/236/433/986238?login=false). OUP Academic. https://academic.oup.com/mind/article/LIX/236/433/986238?login=falseTuring. 1950. “I.—COMPUTING MACHINERY AND INTELLIGENCE”. OUP Academic. Https://web.archive.org/web/20241203212535/https://academic.oup.com/mind/article/LIX/236/433/986238?login=false. Internet Archive. https://academic.oup.com/mind/article/LIX/236/433/986238?login=false.Turing. “I.—COMPUTING MACHINERY AND INTELLIGENCE”. OUP Academic, 1950, Internet Archive, https://web.archive.org/web/20241203212535/https://academic.oup.com/mind/article/LIX/236/433/986238?login=false, https://academic.oup.com/mind/article/LIX/236/433/986238?login=false.Turing. I.—COMPUTING MACHINERY AND INTELLIGENCE. OUP Academic https://academic.oup.com/mind/article/LIX/236/433/986238?login=false (1950).Turing, “I.—COMPUTING MACHINERY AND INTELLIGENCE”, OUP Academic. Accessed: Dec. 03, 2024. [Online]. Available: https://academic.oup.com/mind/article/LIX/236/433/986238?login=false
Turing(1951). Alan Turing. Wikiquote.Turing. (1951). Alan Turing. Wikiquote. https://en.wikiquote.org/wiki/Alan_TuringTuring. 1951. “Alan Turing”. Wikiquote. https://en.wikiquote.org/wiki/Alan_Turing.Turing. “Alan Turing”. Wikiquote, 1951, https://en.wikiquote.org/wiki/Alan_Turing.Turing. Alan Turing. Wikiquote https://en.wikiquote.org/wiki/Alan_Turing (1951).Turing, “Alan Turing”, Wikiquote. [Online]. Available: https://en.wikiquote.org/wiki/Alan_Turing
Turner, A. M. & Tadepalli, P.(2022). Parametrically Retargetable Decision-Makers Tend To Seek Power. arXiv.Turner, A. M., & Tadepalli, P. (2022). Parametrically Retargetable Decision-Makers Tend To Seek Power. In arXiv. https://arxiv.org/abs/2206.13477Turner, A. M., and P. Tadepalli. 2022. “Parametrically Retargetable Decision-Makers Tend To Seek Power”. In arXiv. Preprint, June 27. https://arxiv.org/abs/2206.13477.Turner, A. M., and P. Tadepalli. “Parametrically Retargetable Decision-Makers Tend To Seek Power”. arXiv, 27 June 2022, https://arxiv.org/abs/2206.13477.Turner, A. M. & Tadepalli, P. Parametrically Retargetable Decision-Makers Tend To Seek Power. arXiv Preprint at https://arxiv.org/abs/2206.13477 (2022).A. M. Turner and P. Tadepalli, “Parametrically Retargetable Decision-Makers Tend To Seek Power”, Jun. 27, 2022. [Online]. Available: https://arxiv.org/abs/2206.13477
Turner, A. M., Smith, L., Shah, R., Critch, A. & Tadepalli, P.(2019). Optimal Policies Tend to Seek Power. arXiv.Turner, A. M., Smith, L., Shah, R., Critch, A., & Tadepalli, P. (2019). Optimal Policies Tend to Seek Power. In arXiv. https://arxiv.org/abs/1912.01683Turner, A. M., L. Smith, R. Shah, A. Critch, and P. Tadepalli. 2019. “Optimal Policies Tend to Seek Power”. In arXiv. Preprint, December 3. https://arxiv.org/abs/1912.01683.Turner, A. M., et al. “Optimal Policies Tend to Seek Power”. arXiv, 3 Dec. 2019, https://arxiv.org/abs/1912.01683.Turner, A. M., Smith, L., Shah, R., Critch, A. & Tadepalli, P. Optimal Policies Tend to Seek Power. arXiv Preprint at https://arxiv.org/abs/1912.01683 (2019).A. M. Turner, L. Smith, R. Shah, A. Critch, and P. Tadepalli, “Optimal Policies Tend to Seek Power”, Dec. 03, 2019. [Online]. Available: https://arxiv.org/abs/1912.01683
TurnTrout(2022). Inner and outer alignment decompose one hard problem into two extremely hard problems. AI Alignment Forum.TurnTrout. (2022, December 2). Inner and outer alignment decompose one hard problem into two extremely hard problems. AI Alignment Forum. https://alignmentforum.org/posts/gHefoxiznGfsbiAu9/inner-and-outer-alignment-decompose-one-hard-problem-intoTurnTrout. 2022. “Inner and Outer Alignment Decompose One Hard Problem into Two Extremely Hard Problems”. AI Alignment Forum, December 2. https://alignmentforum.org/posts/gHefoxiznGfsbiAu9/inner-and-outer-alignment-decompose-one-hard-problem-into.TurnTrout. “Inner and Outer Alignment Decompose One Hard Problem into Two Extremely Hard Problems”. AI Alignment Forum, 2 Dec. 2022, https://alignmentforum.org/posts/gHefoxiznGfsbiAu9/inner-and-outer-alignment-decompose-one-hard-problem-into.TurnTrout. Inner and outer alignment decompose one hard problem into two extremely hard problems. AI Alignment Forum https://alignmentforum.org/posts/gHefoxiznGfsbiAu9/inner-and-outer-alignment-decompose-one-hard-problem-into (2022).TurnTrout, “Inner and outer alignment decompose one hard problem into two extremely hard problems”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/gHefoxiznGfsbiAu9/inner-and-outer-alignment-decompose-one-hard-problem-into
TurnTrout(2023). Comment on “TurnTrout's shortform feed”. LessWrong.TurnTrout. (2023, June 1). Comment on “TurnTrout's shortform feed”. LessWrong. https://lesswrong.com/posts/dqSwccGTWyBgxrR58/turntrout-s-shortform-feed?commentId=Sw89AxHGJ5j7E7ETfTurnTrout. 2023. “Comment on “TurnTrout's Shortform Feed””. LessWrong, June 1. https://lesswrong.com/posts/dqSwccGTWyBgxrR58/turntrout-s-shortform-feed?commentId=Sw89AxHGJ5j7E7ETf.TurnTrout. “Comment on “TurnTrout's Shortform Feed””. LessWrong, 1 June 2023, https://lesswrong.com/posts/dqSwccGTWyBgxrR58/turntrout-s-shortform-feed?commentId=Sw89AxHGJ5j7E7ETf.TurnTrout. Comment on “TurnTrout's shortform feed”. LessWrong https://lesswrong.com/posts/dqSwccGTWyBgxrR58/turntrout-s-shortform-feed?commentId=Sw89AxHGJ5j7E7ETf (2023).TurnTrout, “Comment on “TurnTrout's shortform feed””, LessWrong. [Online]. Available: https://lesswrong.com/posts/dqSwccGTWyBgxrR58/turntrout-s-shortform-feed?commentId=Sw89AxHGJ5j7E7ETf
Turpin, M., Michael, J., Perez, E. & Bowman, S. R.(2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. arXiv.Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. In arXiv. https://arxiv.org/abs/2305.04388Turpin, M., J. Michael, E. Perez, and S. R. Bowman. 2023. “Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting”. In arXiv. Preprint, May 7. https://arxiv.org/abs/2305.04388.Turpin, M., et al. “Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting”. arXiv, 7 May 2023, https://arxiv.org/abs/2305.04388.Turpin, M., Michael, J., Perez, E. & Bowman, S. R. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. arXiv Preprint at https://arxiv.org/abs/2305.04388 (2023).M. Turpin, J. Michael, E. Perez, and S. R. Bowman, “Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting”, May 07, 2023. [Online]. Available: https://arxiv.org/abs/2305.04388
U.S Defense Innovation Unit(2023). The Replicator Initiative.U.S Defense Innovation Unit. (2023). The Replicator Initiative. Internet Archive (https://web.archive.org/web/20260303150022/https://www.diu.mil/replicator). https://diu.mil/replicatorU.S Defense Innovation Unit. 2023. “The Replicator Initiative”. Https://web.archive.org/web/20260303150022/https://www.diu.mil/replicator. Internet Archive. https://diu.mil/replicator.U.S Defense Innovation Unit. The Replicator Initiative. 2023, Internet Archive, https://web.archive.org/web/20260303150022/https://www.diu.mil/replicator, https://diu.mil/replicator.U.S Defense Innovation Unit. The Replicator Initiative. https://diu.mil/replicator (2023).U.S Defense Innovation Unit, “The Replicator Initiative”. Accessed: Mar. 03, 2026. [Online]. Available: https://diu.mil/replicator
Uesato, J. et al.(2022). Solving math word problems with process- and outcome-based feedback. arXiv.Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., & Higgins, I. (2022). Solving math word problems with process- and outcome-based feedback. In arXiv. https://arxiv.org/abs/2211.14275Uesato, J., N. Kushman, R. Kumar, et al. 2022. “Solving Math Word Problems with Process- and Outcome-based Feedback”. In arXiv. Preprint, November 25. https://arxiv.org/abs/2211.14275.Uesato, J., et al. “Solving Math Word Problems with Process- and Outcome-based Feedback”. arXiv, 25 Nov. 2022, https://arxiv.org/abs/2211.14275.Uesato, J. et al. Solving math word problems with process- and outcome-based feedback. arXiv Preprint at https://arxiv.org/abs/2211.14275 (2022).J. Uesato et al., “Solving math word problems with process- and outcome-based feedback”, Nov. 25, 2022. [Online]. Available: https://arxiv.org/abs/2211.14275
UK AISI(2024). AI Safety Institute approach to evaluations. GOV.UK.UK AISI. (2024). AI Safety Institute approach to evaluations. GOV.UK. https://gov.uk/government/publications/ai-safety-institute-approach-to-evaluations/ai-safety-institute-approach-to-evaluationsUK AISI. 2024. “AI Safety Institute Approach to Evaluations”. GOV.UK. https://gov.uk/government/publications/ai-safety-institute-approach-to-evaluations/ai-safety-institute-approach-to-evaluations.UK AISI. “AI Safety Institute Approach to Evaluations”. GOV.UK, 2024, https://gov.uk/government/publications/ai-safety-institute-approach-to-evaluations/ai-safety-institute-approach-to-evaluations.UK AISI. AI Safety Institute approach to evaluations. GOV.UK https://gov.uk/government/publications/ai-safety-institute-approach-to-evaluations/ai-safety-institute-approach-to-evaluations (2024).UK AISI, “AI Safety Institute approach to evaluations”, GOV.UK. [Online]. Available: https://gov.uk/government/publications/ai-safety-institute-approach-to-evaluations/ai-safety-institute-approach-to-evaluations
University of Oxford(2024). Prof. Geoffrey Hinton - "Will digital intelligence replace biological intelligence?" Romanes Lecture. YouTube.University of Oxford. (2024). Prof. Geoffrey Hinton - "Will digital intelligence replace biological intelligence?" Romanes Lecture [Video recording]. In YouTube. https://www.youtube.com/watch?v=N1TEjTeQeg0University of Oxford. 2024. “Prof. Geoffrey Hinton - "Will Digital Intelligence Replace Biological Intelligence?" Romanes Lecture”. YouTube. https://www.youtube.com/watch?v=N1TEjTeQeg0.University of Oxford. “Prof. Geoffrey Hinton - "Will Digital Intelligence Replace Biological Intelligence?" Romanes Lecture”. YouTube, 2024, https://www.youtube.com/watch?v=N1TEjTeQeg0.University of Oxford. Prof. Geoffrey Hinton - "Will Digital Intelligence Replace Biological Intelligence?" Romanes Lecture. YouTube (2024).University of Oxford, Prof. Geoffrey Hinton - "Will digital intelligence replace biological intelligence?" Romanes Lecture, (2024). [Online Video]. Available: https://www.youtube.com/watch?v=N1TEjTeQeg0
Urbina, F., Lentzos, F., Invernizzi, C. & Ekins, S.(2022). Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence.Urbina, F., Lentzos, F., Invernizzi, C., & Ekins, S. (2022). Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence. https://doi.org/10.1038/s42256-022-00465-9Urbina, F., F. Lentzos, C. Invernizzi, and S. Ekins. 2022. “Dual Use of Artificial-intelligence-powered Drug Discovery”. Nature Machine Intelligence, ahead of print, March 7. https://doi.org/10.1038/s42256-022-00465-9.Urbina, F., et al. “Dual Use of Artificial-intelligence-powered Drug Discovery”. Nature Machine Intelligence, Mar. 2022, https://doi.org/10.1038/s42256-022-00465-9.Urbina, F., Lentzos, F., Invernizzi, C. & Ekins, S. Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence https://doi.org/10.1038/s42256-022-00465-9 (2022) doi:10.1038/s42256-022-00465-9.F. Urbina, F. Lentzos, C. Invernizzi, and S. Ekins, “Dual use of artificial-intelligence-powered drug discovery”, Nature Machine Intelligence, Mar. 2022, doi: 10.1038/s42256-022-00465-9.
US & UK AISI(2024). Pre-deployment evaluation of Anthropic’s upgraded Claude 3.5 Sonnet. AI Security Institute.US & UK AISI. (2024). Pre-deployment evaluation of Anthropic’s upgraded Claude 3.5 Sonnet. AI Security Institute. https://aisi.gov.uk/work/pre-deployment-evaluation-of-anthropics-upgraded-claude-3-5-sonnetUS & UK AISI. 2024. “Pre-deployment Evaluation of Anthropic’s Upgraded Claude 3.5 Sonnet”. AI Security Institute. https://aisi.gov.uk/work/pre-deployment-evaluation-of-anthropics-upgraded-claude-3-5-sonnet.US & UK AISI. “Pre-deployment Evaluation of Anthropic’s Upgraded Claude 3.5 Sonnet”. AI Security Institute, 2024, https://aisi.gov.uk/work/pre-deployment-evaluation-of-anthropics-upgraded-claude-3-5-sonnet.US & UK AISI. Pre-deployment evaluation of Anthropic’s upgraded Claude 3.5 Sonnet. AI Security Institute https://aisi.gov.uk/work/pre-deployment-evaluation-of-anthropics-upgraded-claude-3-5-sonnet (2024).US & UK AISI, “Pre-deployment evaluation of Anthropic’s upgraded Claude 3.5 Sonnet”, AI Security Institute. [Online]. Available: https://aisi.gov.uk/work/pre-deployment-evaluation-of-anthropics-upgraded-claude-3-5-sonnet
US AI Safety Institute & UK AI Safety Institute(2024). US AISI and UK AISI Joint Pre-Deployment Test: Anthropic's Claude 3.5 Sonnet (October 2024 Release).US AI Safety Institute, & UK AI Safety Institute. (2024). US AISI and UK AISI Joint Pre-Deployment Test: Anthropic's Claude 3.5 Sonnet (October 2024 Release). National Institute of Standards and Technology. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/673b689ec926d8d32e889a8e_UK-US-Testing-Report-Nov-19.pdfUS AI Safety Institute, and UK AI Safety Institute. 2024. US AISI and UK AISI Joint Pre-Deployment Test: Anthropic's Claude 3.5 Sonnet (October 2024 Release). National Institute of Standards and Technology. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/673b689ec926d8d32e889a8e_UK-US-Testing-Report-Nov-19.pdf.US AI Safety Institute, and UK AI Safety Institute. US AISI and UK AISI Joint Pre-Deployment Test: Anthropic's Claude 3.5 Sonnet (October 2024 Release). National Institute of Standards and Technology, 19 Nov. 2024, https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/673b689ec926d8d32e889a8e_UK-US-Testing-Report-Nov-19.pdf.US AI Safety Institute & UK AI Safety Institute. US AISI and UK AISI Joint Pre-Deployment Test: Anthropic's Claude 3.5 Sonnet (October 2024 Release). https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/673b689ec926d8d32e889a8e_UK-US-Testing-Report-Nov-19.pdf (2024).US AI Safety Institute and UK AI Safety Institute, “US AISI and UK AISI Joint Pre-Deployment Test: Anthropic's Claude 3.5 Sonnet (October 2024 Release)”, National Institute of Standards and Technology, Nov. 2024. [Online]. Available: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/673b689ec926d8d32e889a8e_UK-US-Testing-Report-Nov-19.pdf
Valle-Pérez, G., Camargo, C. Q. & Louis, A. A.(2018). Deep learning generalizes because the parameter-function map is biased towards simple functions. arXiv.Valle-Pérez, G., Camargo, C. Q., & Louis, A. A. (2018). Deep learning generalizes because the parameter-function map is biased towards simple functions. In arXiv. https://arxiv.org/abs/1805.08522Valle-Pérez, G., C. Q. Camargo, and A. A. Louis. 2018. “Deep Learning Generalizes Because the Parameter-function Map Is Biased Towards Simple Functions”. In arXiv. Preprint, May 22. https://arxiv.org/abs/1805.08522.Valle-Pérez, G., et al. “Deep Learning Generalizes Because the Parameter-function Map Is Biased Towards Simple Functions”. arXiv, 22 May 2018, https://arxiv.org/abs/1805.08522.Valle-Pérez, G., Camargo, C. Q. & Louis, A. A. Deep learning generalizes because the parameter-function map is biased towards simple functions. arXiv Preprint at https://arxiv.org/abs/1805.08522 (2018).G. Valle-Pérez, C. Q. Camargo, and A. A. Louis, “Deep learning generalizes because the parameter-function map is biased towards simple functions”, May 22, 2018. [Online]. Available: https://arxiv.org/abs/1805.08522
Vanessa Kosoy(2023). The Learning-Theoretic Agenda: Status 2023. AI Alignment Forum.Vanessa Kosoy. (2023, April 19). The Learning-Theoretic Agenda: Status 2023. AI Alignment Forum. https://alignmentforum.org/posts/ZwshvqiqCvXPsZEct/the-learning-theoretic-agenda-status-2023Vanessa Kosoy. 2023. “The Learning-Theoretic Agenda: Status 2023”. AI Alignment Forum, April 19. https://alignmentforum.org/posts/ZwshvqiqCvXPsZEct/the-learning-theoretic-agenda-status-2023.Vanessa Kosoy. “The Learning-Theoretic Agenda: Status 2023”. AI Alignment Forum, 19 Apr. 2023, https://alignmentforum.org/posts/ZwshvqiqCvXPsZEct/the-learning-theoretic-agenda-status-2023.Vanessa Kosoy. The Learning-Theoretic Agenda: Status 2023. AI Alignment Forum https://alignmentforum.org/posts/ZwshvqiqCvXPsZEct/the-learning-theoretic-agenda-status-2023 (2023).Vanessa Kosoy, “The Learning-Theoretic Agenda: Status 2023”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/ZwshvqiqCvXPsZEct/the-learning-theoretic-agenda-status-2023
Variengien & Martinet(2024). AI Safety Institutes: Can countries meet the challenge?.Variengien & Martinet. (2024). AI Safety Institutes: Can countries meet the challenge?. https://oecd.ai/en/wonk/ai-safety-institutes-challengeVariengien & Martinet. 2024. “AI Safety Institutes: Can Countries Meet the Challenge?”. https://oecd.ai/en/wonk/ai-safety-institutes-challenge.Variengien & Martinet. AI Safety Institutes: Can Countries Meet the Challenge?. 2024, https://oecd.ai/en/wonk/ai-safety-institutes-challenge.Variengien & Martinet. AI Safety Institutes: Can countries meet the challenge?. https://oecd.ai/en/wonk/ai-safety-institutes-challenge (2024).Variengien & Martinet, “AI Safety Institutes: Can countries meet the challenge?”. [Online]. Available: https://oecd.ai/en/wonk/ai-safety-institutes-challenge
Venkat Somala, Anson Ho & Séb Krier(2025). Three challenges facing compute-based AI policies.Venkat Somala, Anson Ho, & Séb Krier. (2025, September 11). Three challenges facing compute-based AI policies. https://epoch.ai/gradient-updates/three-issues-undermining-compute-based-ai-policiesVenkat Somala, Anson Ho, and Séb Krier. 2025. “Three Challenges Facing Compute-based AI Policies”. September 11. https://epoch.ai/gradient-updates/three-issues-undermining-compute-based-ai-policies.Venkat Somala, et al. Three Challenges Facing Compute-based AI Policies. 11 Sept. 2025, https://epoch.ai/gradient-updates/three-issues-undermining-compute-based-ai-policies.Venkat Somala, Anson Ho & Séb Krier. Three challenges facing compute-based AI policies. https://epoch.ai/gradient-updates/three-issues-undermining-compute-based-ai-policies (2025).Venkat Somala, Anson Ho, and Séb Krier, “Three challenges facing compute-based AI policies”. [Online]. Available: https://epoch.ai/gradient-updates/three-issues-undermining-compute-based-ai-policies
Veronika Blablová & Robi Rahman(2025). Why China isn’t about to leap ahead of the West on compute.Veronika Blablová, & Robi Rahman. (2025, July 26). Why China isn’t about to leap ahead of the West on compute. https://epoch.ai/gradient-updates/why-china-isnt-about-to-leap-ahead-of-the-west-on-computeVeronika Blablová, and Robi Rahman. 2025. “Why China Isn’t About to Leap Ahead of the West on Compute”. July 26. https://epoch.ai/gradient-updates/why-china-isnt-about-to-leap-ahead-of-the-west-on-compute.Veronika Blablová, and Robi Rahman. Why China Isn’t About to Leap Ahead of the West on Compute. 26 July 2025, https://epoch.ai/gradient-updates/why-china-isnt-about-to-leap-ahead-of-the-west-on-compute.Veronika Blablová & Robi Rahman. Why China isn’t about to leap ahead of the West on compute. https://epoch.ai/gradient-updates/why-china-isnt-about-to-leap-ahead-of-the-west-on-compute (2025).Veronika Blablová and Robi Rahman, “Why China isn’t about to leap ahead of the West on compute”. [Online]. Available: https://epoch.ai/gradient-updates/why-china-isnt-about-to-leap-ahead-of-the-west-on-compute
Vika(2023). When discussing AI risks, talk about capabilities, not intelligence. AI Alignment Forum.Vika. (2023, August 11). When discussing AI risks, talk about capabilities, not intelligence. AI Alignment Forum. https://alignmentforum.org/posts/JtuTQgp9Wnd6R6F5s/when-discussing-ai-risks-talk-about-capabilities-notVika. 2023. “When Discussing AI Risks, Talk About Capabilities, Not Intelligence”. AI Alignment Forum, August 11. https://alignmentforum.org/posts/JtuTQgp9Wnd6R6F5s/when-discussing-ai-risks-talk-about-capabilities-not.Vika. “When Discussing AI Risks, Talk About Capabilities, Not Intelligence”. AI Alignment Forum, 11 Aug. 2023, https://alignmentforum.org/posts/JtuTQgp9Wnd6R6F5s/when-discussing-ai-risks-talk-about-capabilities-not.Vika. When discussing AI risks, talk about capabilities, not intelligence. AI Alignment Forum https://alignmentforum.org/posts/JtuTQgp9Wnd6R6F5s/when-discussing-ai-risks-talk-about-capabilities-not (2023).Vika, “When discussing AI risks, talk about capabilities, not intelligence”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/JtuTQgp9Wnd6R6F5s/when-discussing-ai-risks-talk-about-capabilities-not
Villalobos et al.(2024). Will we run out of data to train large language models?. Epoch AI.Villalobos et al. (2024). Will we run out of data to train large language models?. Epoch AI. https://epoch.ai/blog/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-dataVillalobos et al. 2024. “Will We Run Out of Data to Train Large Language Models?”. Epoch AI. https://epoch.ai/blog/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data.Villalobos et al. “Will We Run Out of Data to Train Large Language Models?”. Epoch AI, 2024, https://epoch.ai/blog/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data.Villalobos et al. Will we run out of data to train large language models?. Epoch AI https://epoch.ai/blog/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data (2024).Villalobos et al., “Will we run out of data to train large language models?”, Epoch AI. [Online]. Available: https://epoch.ai/blog/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data
Villalobos, P., Ho, A., Sevilla, J., Besiroglu, T., Heim, L. & Hobbhahn, M.(2022). Will we run out of data? Limits of LLM scaling based on human-generated data. arXiv.Villalobos, P., Ho, A., Sevilla, J., Besiroglu, T., Heim, L., & Hobbhahn, M. (2022). Will we run out of data? Limits of LLM scaling based on human-generated data. In arXiv. https://arxiv.org/abs/2211.04325Villalobos, P., A. Ho, J. Sevilla, T. Besiroglu, L. Heim, and M. Hobbhahn. 2022. “Will We Run Out of Data? Limits of LLM Scaling Based on Human-generated Data”. In arXiv. Preprint, October 26. https://arxiv.org/abs/2211.04325.Villalobos, P., et al. “Will We Run Out of Data? Limits of LLM Scaling Based on Human-generated Data”. arXiv, 26 Oct. 2022, https://arxiv.org/abs/2211.04325.Villalobos, P. et al. Will we run out of data? Limits of LLM scaling based on human-generated data. arXiv Preprint at https://arxiv.org/abs/2211.04325 (2022).P. Villalobos, A. Ho, J. Sevilla, T. Besiroglu, L. Heim, and M. Hobbhahn, “Will we run out of data? Limits of LLM scaling based on human-generated data”, Oct. 26, 2022. [Online]. Available: https://arxiv.org/abs/2211.04325
Wan, S. et al.(2024). CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models. arXiv.Wan, S., Nikolaidis, C., Song, D., Molnar, D., Crnkovich, J., Grace, J., Bhatt, M., Chennabasappa, S., Whitman, S., Ding, S., Ionescu, V., Li, Y., & Saxe, J. (2024). CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models. In arXiv. https://arxiv.org/abs/2408.01605Wan, S., C. Nikolaidis, D. Song, et al. 2024. “CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models”. In arXiv. Preprint, August 2. https://arxiv.org/abs/2408.01605.Wan, S., et al. “CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models”. arXiv, 2 Aug. 2024, https://arxiv.org/abs/2408.01605.Wan, S. et al. CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models. arXiv Preprint at https://arxiv.org/abs/2408.01605 (2024).S. Wan et al., “CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models”, Aug. 02, 2024. [Online]. Available: https://arxiv.org/abs/2408.01605
Wang(2020). Journal of Artificial General Intelligence. Paradigm.Wang. (2020). Journal of Artificial General Intelligence. Paradigm. https://sciendo.com/issue/JAGI/11/2Wang. 2020. “Journal of Artificial General Intelligence”. Paradigm. https://sciendo.com/issue/JAGI/11/2.Wang. “Journal of Artificial General Intelligence”. Paradigm, 2020, https://sciendo.com/issue/JAGI/11/2.Wang. Journal of Artificial General Intelligence. Paradigm https://sciendo.com/issue/JAGI/11/2 (2020).Wang, “Journal of Artificial General Intelligence”, Paradigm. [Online]. Available: https://sciendo.com/issue/JAGI/11/2
Wang et al.(2022). Adversarial Policies Beat Superhuman Go AIs. arXiv.org.Wang et al. (2022). Adversarial Policies Beat Superhuman Go AIs. arXiv.org. https://www.arxiv.org/abs/2211.00241Wang et al. 2022. “Adversarial Policies Beat Superhuman Go AIs”. arXiv.org. https://www.arxiv.org/abs/2211.00241.Wang et al. “Adversarial Policies Beat Superhuman Go AIs”. arXiv.org, 2022, https://www.arxiv.org/abs/2211.00241.Wang et al. Adversarial Policies Beat Superhuman Go AIs. arXiv.org https://www.arxiv.org/abs/2211.00241 (2022).Wang et al., “Adversarial Policies Beat Superhuman Go AIs”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2211.00241
Wang, A. et al.(2019). SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems. arXiv.Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., & Bowman, S. R. (2019). SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems. In arXiv. https://arxiv.org/abs/1905.00537Wang, A., Y. Pruksachatkun, N. Nangia, et al. 2019. “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems”. In arXiv. Preprint, May 2. https://arxiv.org/abs/1905.00537.Wang, A., et al. “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems”. arXiv, 2 May 2019, https://arxiv.org/abs/1905.00537.Wang, A. et al. SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems. arXiv Preprint at https://arxiv.org/abs/1905.00537 (2019).A. Wang et al., “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems”, May 02, 2019. [Online]. Available: https://arxiv.org/abs/1905.00537
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O. & Bowman, S. R.(2018). GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. arXiv.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., & Bowman, S. R. (2018). GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In arXiv. https://arxiv.org/abs/1804.07461Wang, A., A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman. 2018. “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”. In arXiv. Preprint, April 20. https://arxiv.org/abs/1804.07461.Wang, A., et al. “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”. arXiv, 20 Apr. 2018, https://arxiv.org/abs/1804.07461.Wang, A. et al. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. arXiv Preprint at https://arxiv.org/abs/1804.07461 (2018).A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”, Apr. 20, 2018. [Online]. Available: https://arxiv.org/abs/1804.07461
Wang, G. et al.(2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. In arXiv. https://arxiv.org/abs/2305.16291Wang, G., Y. Xie, Y. Jiang, et al. 2023. “Voyager: An Open-Ended Embodied Agent with Large Language Models”. In arXiv. Preprint, May 25. https://arxiv.org/abs/2305.16291.Wang, G., et al. “Voyager: An Open-Ended Embodied Agent with Large Language Models”. arXiv, 25 May 2023, https://arxiv.org/abs/2305.16291.Wang, G. et al. Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv Preprint at https://arxiv.org/abs/2305.16291 (2023).G. Wang et al., “Voyager: An Open-Ended Embodied Agent with Large Language Models”, May 25, 2023. [Online]. Available: https://arxiv.org/abs/2305.16291
Wang, L. et al.(2023). A Survey on Large Language Model based Autonomous Agents. arXiv.Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z., & Wen, J.-R. (2023). A Survey on Large Language Model based Autonomous Agents. In arXiv. https://doi.org/10.1007/s11704-024-40231-1Wang, L., C. Ma, X. Feng, et al. 2023. “A Survey on Large Language Model Based Autonomous Agents”. In arXiv. Preprint, August 22. https://doi.org/10.1007/s11704-024-40231-1.Wang, L., et al. “A Survey on Large Language Model Based Autonomous Agents”. arXiv, 22 Aug. 2023, https://doi.org/10.1007/s11704-024-40231-1.Wang, L. et al. A Survey on Large Language Model based Autonomous Agents. arXiv Preprint at https://doi.org/10.1007/s11704-024-40231-1 (2023).L. Wang et al., “A Survey on Large Language Model based Autonomous Agents”, Aug. 22, 2023. doi: 10.1007/s11704-024-40231-1.
Wang, S., Liu, S., Ye, W., You, J. & Gao, Y.(2024). EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data. arXiv.Wang, S., Liu, S., Ye, W., You, J., & Gao, Y. (2024). EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data. In arXiv. https://arxiv.org/abs/2403.00564Wang, S., S. Liu, W. Ye, J. You, and Y. Gao. 2024. “EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data”. In arXiv. Preprint, March 1. https://arxiv.org/abs/2403.00564.Wang, S., et al. “EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data”. arXiv, 1 Mar. 2024, https://arxiv.org/abs/2403.00564.Wang, S., Liu, S., Ye, W., You, J. & Gao, Y. EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data. arXiv Preprint at https://arxiv.org/abs/2403.00564 (2024).S. Wang, S. Liu, W. Ye, J. You, and Y. Gao, “EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data”, Mar. 01, 2024. [Online]. Available: https://arxiv.org/abs/2403.00564
Wang, X. et al.(2022). Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., & Zhou, D. (2022). Self-Consistency Improves Chain of Thought Reasoning in Language Models. In arXiv. https://arxiv.org/abs/2203.11171Wang, X., J. Wei, D. Schuurmans, et al. 2022. “Self-Consistency Improves Chain of Thought Reasoning in Language Models”. In arXiv. Preprint, March 21. https://arxiv.org/abs/2203.11171.Wang, X., et al. “Self-Consistency Improves Chain of Thought Reasoning in Language Models”. arXiv, 21 Mar. 2022, https://arxiv.org/abs/2203.11171.Wang, X. et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv Preprint at https://arxiv.org/abs/2203.11171 (2022).X. Wang et al., “Self-Consistency Improves Chain of Thought Reasoning in Language Models”, Mar. 21, 2022. [Online]. Available: https://arxiv.org/abs/2203.11171
Wang, X. et al.(2023). InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews. arXiv.Wang, X., Xiao, Y., Huang, J.-T., Yuan, S., Xu, R., Guo, H., Tu, Q., Fei, Y., Leng, Z., Wang, W., Chen, J., Li, C., & Xiao, Y. (2023). InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews. In arXiv. https://arxiv.org/abs/2310.17976Wang, X., Y. Xiao, J.-T. Huang, et al. 2023. “InCharacter: Evaluating Personality Fidelity in Role-Playing Agents Through Psychological Interviews”. In arXiv. Preprint, October 27. https://arxiv.org/abs/2310.17976.Wang, X., et al. “InCharacter: Evaluating Personality Fidelity in Role-Playing Agents Through Psychological Interviews”. arXiv, 27 Oct. 2023, https://arxiv.org/abs/2310.17976.Wang, X. et al. InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews. arXiv Preprint at https://arxiv.org/abs/2310.17976 (2023).X. Wang et al., “InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews”, Oct. 27, 2023. [Online]. Available: https://arxiv.org/abs/2310.17976
Wasil, A. R., Clymer, J., Krueger, D., Dardaman, E., Campos, S. & Murphy, E. R.(2024). Affirmative safety: An approach to risk management for high-risk AI. arXiv.Wasil, A. R., Clymer, J., Krueger, D., Dardaman, E., Campos, S., & Murphy, E. R. (2024). Affirmative safety: An approach to risk management for high-risk AI. In arXiv. https://arxiv.org/abs/2406.15371Wasil, A. R., J. Clymer, D. Krueger, E. Dardaman, S. Campos, and E. R. Murphy. 2024. “Affirmative Safety: An Approach to Risk Management for High-risk AI”. In arXiv. Preprint, April 14. https://arxiv.org/abs/2406.15371.Wasil, A. R., et al. “Affirmative Safety: An Approach to Risk Management for High-risk AI”. arXiv, 14 Apr. 2024, https://arxiv.org/abs/2406.15371.Wasil, A. R. et al. Affirmative safety: An approach to risk management for high-risk AI. arXiv Preprint at https://arxiv.org/abs/2406.15371 (2024).A. R. Wasil, J. Clymer, D. Krueger, E. Dardaman, S. Campos, and E. R. Murphy, “Affirmative safety: An approach to risk management for high-risk AI”, Apr. 14, 2024. [Online]. Available: https://arxiv.org/abs/2406.15371
Wasil, A. R., Reed, T., Miller, J. W. & Barnett, P.(2024). Verification methods for international AI agreements. arXiv.Wasil, A. R., Reed, T., Miller, J. W., & Barnett, P. (2024). Verification methods for international AI agreements. In arXiv. https://arxiv.org/abs/2408.16074Wasil, A. R., T. Reed, J. W. Miller, and P. Barnett. 2024. “Verification Methods for International AI Agreements”. In arXiv. Preprint, August 28. https://arxiv.org/abs/2408.16074.Wasil, A. R., et al. “Verification Methods for International AI Agreements”. arXiv, 28 Aug. 2024, https://arxiv.org/abs/2408.16074.Wasil, A. R., Reed, T., Miller, J. W. & Barnett, P. Verification methods for international AI agreements. arXiv Preprint at https://arxiv.org/abs/2408.16074 (2024).A. R. Wasil, T. Reed, J. W. Miller, and P. Barnett, “Verification methods for international AI agreements”, Aug. 28, 2024. [Online]. Available: https://arxiv.org/abs/2408.16074
Wei Dai(2019). AGI will drastically increase economies of scale. AI Alignment Forum.Wei Dai. (2019, June 7). AGI will drastically increase economies of scale. AI Alignment Forum. https://alignmentforum.org/posts/Sn5NiiD5WBi4dLzaB/agi-will-drastically-increase-economies-of-scale-1Wei Dai. 2019. “AGI Will Drastically Increase Economies of Scale”. AI Alignment Forum, June 7. https://alignmentforum.org/posts/Sn5NiiD5WBi4dLzaB/agi-will-drastically-increase-economies-of-scale-1.Wei Dai. “AGI Will Drastically Increase Economies of Scale”. AI Alignment Forum, 7 June 2019, https://alignmentforum.org/posts/Sn5NiiD5WBi4dLzaB/agi-will-drastically-increase-economies-of-scale-1.Wei Dai. AGI will drastically increase economies of scale. AI Alignment Forum https://alignmentforum.org/posts/Sn5NiiD5WBi4dLzaB/agi-will-drastically-increase-economies-of-scale-1 (2019).Wei Dai, “AGI will drastically increase economies of scale”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/Sn5NiiD5WBi4dLzaB/agi-will-drastically-increase-economies-of-scale-1
Wei, J. et al.(2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., & Zhou, D. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In arXiv. https://arxiv.org/abs/2201.11903Wei, J., X. Wang, D. Schuurmans, et al. 2022. “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”. In arXiv. Preprint, January 28. https://arxiv.org/abs/2201.11903.Wei, J., et al. “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”. arXiv, 28 Jan. 2022, https://arxiv.org/abs/2201.11903.Wei, J. et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv Preprint at https://arxiv.org/abs/2201.11903 (2022).J. Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”, Jan. 28, 2022. [Online]. Available: https://arxiv.org/abs/2201.11903
Wei, J. et al.(2022). Emergent Abilities of Large Language Models. arXiv.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., & Fedus, W. (2022). Emergent Abilities of Large Language Models. In arXiv. https://arxiv.org/abs/2206.07682Wei, J., Y. Tay, R. Bommasani, et al. 2022. “Emergent Abilities of Large Language Models”. In arXiv. Preprint, June 15. https://arxiv.org/abs/2206.07682.Wei, J., et al. “Emergent Abilities of Large Language Models”. arXiv, 15 June 2022, https://arxiv.org/abs/2206.07682.Wei, J. et al. Emergent Abilities of Large Language Models. arXiv Preprint at https://arxiv.org/abs/2206.07682 (2022).J. Wei et al., “Emergent Abilities of Large Language Models”, Jun. 15, 2022. [Online]. Available: https://arxiv.org/abs/2206.07682
Weidinger, L. et al.(2021). Ethical and social risks of harm from Language Models. arXiv.Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., … Gabriel, I. (2021). Ethical and social risks of harm from Language Models. In arXiv. https://arxiv.org/abs/2112.04359Weidinger, L., J. Mellor, M. Rauh, et al. 2021. “Ethical and Social Risks of Harm from Language Models”. In arXiv. Preprint, December 8. https://arxiv.org/abs/2112.04359.Weidinger, L., et al. “Ethical and Social Risks of Harm from Language Models”. arXiv, 8 Dec. 2021, https://arxiv.org/abs/2112.04359.Weidinger, L. et al. Ethical and social risks of harm from Language Models. arXiv Preprint at https://arxiv.org/abs/2112.04359 (2021).L. Weidinger et al., “Ethical and social risks of harm from Language Models”, Dec. 08, 2021. [Online]. Available: https://arxiv.org/abs/2112.04359
Weidinger, L. et al.(2023). Sociotechnical Safety Evaluation of Generative AI Systems. arXiv.Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., & Isaac, W. (2023). Sociotechnical Safety Evaluation of Generative AI Systems. In arXiv. https://arxiv.org/abs/2310.11986Weidinger, L., M. Rauh, N. Marchal, et al. 2023. “Sociotechnical Safety Evaluation of Generative AI Systems”. In arXiv. Preprint, October 18. https://arxiv.org/abs/2310.11986.Weidinger, L., et al. “Sociotechnical Safety Evaluation of Generative AI Systems”. arXiv, 18 Oct. 2023, https://arxiv.org/abs/2310.11986.Weidinger, L. et al. Sociotechnical Safety Evaluation of Generative AI Systems. arXiv Preprint at https://arxiv.org/abs/2310.11986 (2023).L. Weidinger et al., “Sociotechnical Safety Evaluation of Generative AI Systems”, Oct. 18, 2023. [Online]. Available: https://arxiv.org/abs/2310.11986
Whitehouse(2025). Fact Sheet: President Donald J. Trump Takes Action to Enhance America’s AI Leadership. The White House.Whitehouse. (2025, January 23). Fact Sheet: President Donald J. Trump Takes Action to Enhance America’s AI Leadership. The White House. https://whitehouse.gov/fact-sheets/2025/01/fact-sheet-president-donald-j-trump-takes-action-to-enhance-americas-ai-leadershipWhitehouse. 2025. “Fact Sheet: President Donald J. Trump Takes Action to Enhance America’s AI Leadership”. The White House, January 23. https://whitehouse.gov/fact-sheets/2025/01/fact-sheet-president-donald-j-trump-takes-action-to-enhance-americas-ai-leadership.Whitehouse. “Fact Sheet: President Donald J. Trump Takes Action to Enhance America’s AI Leadership”. The White House, 23 Jan. 2025, https://whitehouse.gov/fact-sheets/2025/01/fact-sheet-president-donald-j-trump-takes-action-to-enhance-americas-ai-leadership.Whitehouse. Fact Sheet: President Donald J. Trump Takes Action to Enhance America’s AI Leadership. The White House https://whitehouse.gov/fact-sheets/2025/01/fact-sheet-president-donald-j-trump-takes-action-to-enhance-americas-ai-leadership (2025).Whitehouse, “Fact Sheet: President Donald J. Trump Takes Action to Enhance America’s AI Leadership”, The White House. [Online]. Available: https://whitehouse.gov/fact-sheets/2025/01/fact-sheet-president-donald-j-trump-takes-action-to-enhance-americas-ai-leadership
Whitehouse(2025). Preventing Woke AI in the Federal Government. The White House.Whitehouse. (2025, July 23). Preventing Woke AI in the Federal Government. The White House. https://whitehouse.gov/presidential-actions/2025/07/preventing-woke-ai-in-the-federal-governmentWhitehouse. 2025. “Preventing Woke AI in the Federal Government”. The White House, July 23. https://whitehouse.gov/presidential-actions/2025/07/preventing-woke-ai-in-the-federal-government.Whitehouse. “Preventing Woke AI in the Federal Government”. The White House, 23 July 2025, https://whitehouse.gov/presidential-actions/2025/07/preventing-woke-ai-in-the-federal-government.Whitehouse. Preventing Woke AI in the Federal Government. The White House https://whitehouse.gov/presidential-actions/2025/07/preventing-woke-ai-in-the-federal-government (2025).Whitehouse, “Preventing Woke AI in the Federal Government”, The White House. [Online]. Available: https://whitehouse.gov/presidential-actions/2025/07/preventing-woke-ai-in-the-federal-government
Wijk, H. et al.(2024). RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts. arXiv.Wijk, H., Lin, T., Becker, J., Jawhar, S., Parikh, N., Broadley, T., Chan, L., Chen, M., Clymer, J., Dhyani, J., Ericheva, E., Garcia, K., Goodrich, B., Jurkovic, N., Karnofsky, H., Kinniment, M., Lajko, A., Nix, S., Sato, L., … Barnes, E. (2024). RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts. In arXiv. https://arxiv.org/abs/2411.15114Wijk, H., T. Lin, J. Becker, et al. 2024. “RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents Against Human Experts”. In arXiv. Preprint, November 22. https://arxiv.org/abs/2411.15114.Wijk, H., et al. “RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents Against Human Experts”. arXiv, 22 Nov. 2024, https://arxiv.org/abs/2411.15114.Wijk, H. et al. RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts. arXiv Preprint at https://arxiv.org/abs/2411.15114 (2024).H. Wijk et al., “RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts”, Nov. 22, 2024. [Online]. Available: https://arxiv.org/abs/2411.15114
William MacAskill & Rose Hadshar(2025). Intelsat as a Model for International AGI Governance.William MacAskill, & Rose Hadshar. (2025, March 13). Intelsat as a Model for International AGI Governance. https://forethought.org/research/intelsat-as-a-model-for-international-agi-governanceWilliam MacAskill, and Rose Hadshar. 2025. “Intelsat as a Model for International AGI Governance”. March 13. https://forethought.org/research/intelsat-as-a-model-for-international-agi-governance.William MacAskill, and Rose Hadshar. Intelsat as a Model for International AGI Governance. 13 Mar. 2025, https://forethought.org/research/intelsat-as-a-model-for-international-agi-governance.William MacAskill & Rose Hadshar. Intelsat as a Model for International AGI Governance. https://forethought.org/research/intelsat-as-a-model-for-international-agi-governance (2025).William MacAskill and Rose Hadshar, “Intelsat as a Model for International AGI Governance”. [Online]. Available: https://forethought.org/research/intelsat-as-a-model-for-international-agi-governance
William MacAskill(2025). Better Futures.William MacAskill. (2025, August 3). Better Futures. https://forethought.org/research/better-futuresWilliam MacAskill. 2025. “Better Futures”. August 3. https://forethought.org/research/better-futures.William MacAskill. Better Futures. 3 Aug. 2025, https://forethought.org/research/better-futures.William MacAskill. Better Futures. https://forethought.org/research/better-futures (2025).William MacAskill, “Better Futures”. [Online]. Available: https://forethought.org/research/better-futures
Williams et al.(2025). Forecasting LLM-enabled Biorisk and the Efficacy of Safeguards. Forecasting Research Institute.Williams et al. (2025). Forecasting LLM-enabled Biorisk and the Efficacy of Safeguards. Forecasting Research Institute. https://forecastingresearch.org/ai-enabled-bioriskWilliams et al. 2025. “Forecasting LLM-enabled Biorisk and the Efficacy of Safeguards”. Forecasting Research Institute. https://forecastingresearch.org/ai-enabled-biorisk.Williams et al. “Forecasting LLM-enabled Biorisk and the Efficacy of Safeguards”. Forecasting Research Institute, 2025, https://forecastingresearch.org/ai-enabled-biorisk.Williams et al. Forecasting LLM-enabled Biorisk and the Efficacy of Safeguards. Forecasting Research Institute https://forecastingresearch.org/ai-enabled-biorisk (2025).Williams et al., “Forecasting LLM-enabled Biorisk and the Efficacy of Safeguards”, Forecasting Research Institute. [Online]. Available: https://forecastingresearch.org/ai-enabled-biorisk
Willis et al.(2024). Assessing Global Catastrophic and Existential Risks.Willis et al. (2024). Assessing Global Catastrophic and Existential Risks. Internet Archive (https://web.archive.org/web/20260514192225/https://www.rand.org/pubs/research_reports/RRA2981-1.html). https://rand.org/pubs/research_reports/RRA2981-1.htmlWillis et al. 2024. “Assessing Global Catastrophic and Existential Risks”. Https://web.archive.org/web/20260514192225/https://www.rand.org/pubs/research_reports/RRA2981-1.html. Internet Archive. https://rand.org/pubs/research_reports/RRA2981-1.html.Willis et al. Assessing Global Catastrophic and Existential Risks. 2024, Internet Archive, https://web.archive.org/web/20260514192225/https://www.rand.org/pubs/research_reports/RRA2981-1.html, https://rand.org/pubs/research_reports/RRA2981-1.html.Willis et al. Assessing Global Catastrophic and Existential Risks. https://rand.org/pubs/research_reports/RRA2981-1.html (2024).Willis et al., “Assessing Global Catastrophic and Existential Risks”. Accessed: May 14, 2026. [Online]. Available: https://rand.org/pubs/research_reports/RRA2981-1.html
WillPetillo, Sean Herrington, Spencer Ames, Adebayo Mubarak & Can Narin(2025). Case Studies in Simulators and Agents. LessWrong.WillPetillo, Sean Herrington, Spencer Ames, Adebayo Mubarak, & Can Narin. (2025, May 25). Case Studies in Simulators and Agents. LessWrong. https://lesswrong.com/posts/uJFC5WrcyTdat3Qcc/case-studies-in-simulators-and-agentsWillPetillo, Sean Herrington, Spencer Ames, Adebayo Mubarak, and Can Narin. 2025. “Case Studies in Simulators and Agents”. LessWrong, May 25. https://lesswrong.com/posts/uJFC5WrcyTdat3Qcc/case-studies-in-simulators-and-agents.WillPetillo, et al. “Case Studies in Simulators and Agents”. LessWrong, 25 May 2025, https://lesswrong.com/posts/uJFC5WrcyTdat3Qcc/case-studies-in-simulators-and-agents.WillPetillo, Sean Herrington, Spencer Ames, Adebayo Mubarak & Can Narin. Case Studies in Simulators and Agents. LessWrong https://lesswrong.com/posts/uJFC5WrcyTdat3Qcc/case-studies-in-simulators-and-agents (2025).WillPetillo, Sean Herrington, Spencer Ames, Adebayo Mubarak, and Can Narin, “Case Studies in Simulators and Agents”, LessWrong. [Online]. Available: https://lesswrong.com/posts/uJFC5WrcyTdat3Qcc/case-studies-in-simulators-and-agents
Wiseman & McClements(2025). How much economic growth from AI should we expect, how soon?.Wiseman & McClements. (2025). How much economic growth from AI should we expect, how soon?. https://inferencemagazine.substack.com/p/how-much-economic-growth-from-aiWiseman & McClements. 2025. How Much Economic Growth from AI Should We Expect, How Soon?. Edition. https://inferencemagazine.substack.com/p/how-much-economic-growth-from-ai.Wiseman & McClements. How Much Economic Growth from AI Should We Expect, How Soon?. 2025, https://inferencemagazine.substack.com/p/how-much-economic-growth-from-ai.Wiseman & McClements. How much economic growth from AI should we expect, how soon?. https://inferencemagazine.substack.com/p/how-much-economic-growth-from-ai (2025).Wiseman & McClements, “How much economic growth from AI should we expect, how soon?”. [Online]. Available: https://inferencemagazine.substack.com/p/how-much-economic-growth-from-ai
Wojton, H. M., Porter, D. J. & Dennis, J. W.(2020). Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review.Wojton, H. M., Porter, D. J., & Dennis, J. W. (2020). Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review (IDA Document NS-D-14331). Institute for Defense Analyses. https://testscience.org/wp-content/uploads/formidable/20/Autonomy-Lit-Review.pdfWojton, H. M., D. J. Porter, and J. W. Dennis. 2020. Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review. IDA Document NS-D-14331. Institute for Defense Analyses. https://testscience.org/wp-content/uploads/formidable/20/Autonomy-Lit-Review.pdf.Wojton, H. M., et al. Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review. IDA Document NS-D-14331, Institute for Defense Analyses, Sept. 2020, https://testscience.org/wp-content/uploads/formidable/20/Autonomy-Lit-Review.pdf.Wojton, H. M., Porter, D. J. & Dennis, J. W. Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review. https://testscience.org/wp-content/uploads/formidable/20/Autonomy-Lit-Review.pdf (2020).H. M. Wojton, D. J. Porter, and J. W. Dennis, “Test & Evaluation of AI-enabled and Autonomous Systems: A Literature Review”, Institute for Defense Analyses, IDA Document NS-D-14331, Sep. 2020. [Online]. Available: https://testscience.org/wp-content/uploads/formidable/20/Autonomy-Lit-Review.pdf
Wolf, Y., Wies, N., Avnery, O., Levine, Y. & Shashua, A.(2023). Fundamental Limitations of Alignment in Large Language Models. arXiv.Wolf, Y., Wies, N., Avnery, O., Levine, Y., & Shashua, A. (2023). Fundamental Limitations of Alignment in Large Language Models. In arXiv. https://arxiv.org/abs/2304.11082Wolf, Y., N. Wies, O. Avnery, Y. Levine, and A. Shashua. 2023. “Fundamental Limitations of Alignment in Large Language Models”. In arXiv. Preprint, April 19. https://arxiv.org/abs/2304.11082.Wolf, Y., et al. “Fundamental Limitations of Alignment in Large Language Models”. arXiv, 19 Apr. 2023, https://arxiv.org/abs/2304.11082.Wolf, Y., Wies, N., Avnery, O., Levine, Y. & Shashua, A. Fundamental Limitations of Alignment in Large Language Models. arXiv Preprint at https://arxiv.org/abs/2304.11082 (2023).Y. Wolf, N. Wies, O. Avnery, Y. Levine, and A. Shashua, “Fundamental Limitations of Alignment in Large Language Models”, Apr. 19, 2023. [Online]. Available: https://arxiv.org/abs/2304.11082
Wongkamjan, W. et al.(2024). More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play. arXiv.Wongkamjan, W., Gu, F., Wang, Y., Hermjakob, U., May, J., Stewart, B. M., Kummerfeld, J. K., Peskoff, D., & Boyd-Graber, J. L. (2024). More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play. In arXiv. https://doi.org/10.18653/v1/2024.acl-long.672Wongkamjan, W., F. Gu, Y. Wang, et al. 2024. “More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play”. In arXiv. Preprint, June 7. https://doi.org/10.18653/v1/2024.acl-long.672.Wongkamjan, W., et al. “More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play”. arXiv, 7 June 2024, https://doi.org/10.18653/v1/2024.acl-long.672.Wongkamjan, W. et al. More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play. arXiv Preprint at https://doi.org/10.18653/v1/2024.acl-long.672 (2024).W. Wongkamjan et al., “More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play”, Jun. 07, 2024. doi: 10.18653/v1/2024.acl-long.672.
Wooldridge(2021). A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books.Wooldridge. (2021). A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books. https://amazon.com/Brief-History-Artificial-Intelligence-Where/dp/1250770742Wooldridge. 2021. “A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books”. https://amazon.com/Brief-History-Artificial-Intelligence-Where/dp/1250770742.Wooldridge. A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books. 2021, https://amazon.com/Brief-History-Artificial-Intelligence-Where/dp/1250770742.Wooldridge. A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books. https://amazon.com/Brief-History-Artificial-Intelligence-Where/dp/1250770742 (2021).Wooldridge, “A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books”. [Online]. Available: https://amazon.com/Brief-History-Artificial-Intelligence-Where/dp/1250770742
Wooldridge(2024). AI’s simple solution to rail problems: stop all trains running. The Telegraph.Wooldridge. (2024). AI’s simple solution to rail problems: stop all trains running. Internet Archive (https://web.archive.org/web/20250725100053/https://www.telegraph.co.uk/news/2024/01/07/artificial-intelligence-train-problems/). The Telegraph. https://telegraph.co.uk/news/2024/01/07/artificial-intelligence-train-problemsWooldridge. 2024. “AI’s Simple Solution to Rail Problems: Stop All Trains Running”. The Telegraph. Https://web.archive.org/web/20250725100053/https://www.telegraph.co.uk/news/2024/01/07/artificial-intelligence-train-problems/. Internet Archive. https://telegraph.co.uk/news/2024/01/07/artificial-intelligence-train-problems.Wooldridge. “AI’s Simple Solution to Rail Problems: Stop All Trains Running”. The Telegraph, 2024, Internet Archive, https://web.archive.org/web/20250725100053/https://www.telegraph.co.uk/news/2024/01/07/artificial-intelligence-train-problems/, https://telegraph.co.uk/news/2024/01/07/artificial-intelligence-train-problems.Wooldridge. AI’s simple solution to rail problems: stop all trains running. The Telegraph https://telegraph.co.uk/news/2024/01/07/artificial-intelligence-train-problems (2024).Wooldridge, “AI’s simple solution to rail problems: stop all trains running”, The Telegraph. Accessed: Jul. 25, 2025. [Online]. Available: https://telegraph.co.uk/news/2024/01/07/artificial-intelligence-train-problems
World Economic Forum(2025). The Dawn of Artificial General Intelligence? | World Economic Forum Annual Meeting 2025. YouTube.World Economic Forum. (2025). The Dawn of Artificial General Intelligence? | World Economic Forum Annual Meeting 2025 [Video recording]. In YouTube. https://www.youtube.com/watch?v=Y1BUaLo67acWorld Economic Forum. 2025. “The Dawn of Artificial General Intelligence? | World Economic Forum Annual Meeting 2025”. YouTube. https://www.youtube.com/watch?v=Y1BUaLo67ac.World Economic Forum. “The Dawn of Artificial General Intelligence? | World Economic Forum Annual Meeting 2025”. YouTube, 2025, https://www.youtube.com/watch?v=Y1BUaLo67ac.World Economic Forum. The Dawn of Artificial General Intelligence? | World Economic Forum Annual Meeting 2025. YouTube (2025).World Economic Forum, The Dawn of Artificial General Intelligence? | World Economic Forum Annual Meeting 2025, (2025). [Online Video]. Available: https://www.youtube.com/watch?v=Y1BUaLo67ac
WorldCoin(2024). Proof of personhood: What it is and why it’s needed. World.WorldCoin. (2024). Proof of personhood: What it is and why it’s needed. World. https://world.org/blog/world/proof-of-personhood-what-it-is-why-its-neededWorldCoin. 2024. “Proof of Personhood: What It Is and Why It’s Needed”. World. https://world.org/blog/world/proof-of-personhood-what-it-is-why-its-needed.WorldCoin. “Proof of Personhood: What It Is and Why It’s Needed”. World, 2024, https://world.org/blog/world/proof-of-personhood-what-it-is-why-its-needed.WorldCoin. Proof of personhood: What it is and why it’s needed. World https://world.org/blog/world/proof-of-personhood-what-it-is-why-its-needed (2024).WorldCoin, “Proof of personhood: What it is and why it’s needed”, World. [Online]. Available: https://world.org/blog/world/proof-of-personhood-what-it-is-why-its-needed
Wu, J. et al.(2021). Recursively Summarizing Books with Human Feedback. arXiv.Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., & Christiano, P. (2021). Recursively Summarizing Books with Human Feedback. In arXiv. https://arxiv.org/abs/2109.10862Wu, J., L. Ouyang, D. M. Ziegler, et al. 2021. “Recursively Summarizing Books with Human Feedback”. In arXiv. Preprint, September 22. https://arxiv.org/abs/2109.10862.Wu, J., et al. “Recursively Summarizing Books with Human Feedback”. arXiv, 22 Sept. 2021, https://arxiv.org/abs/2109.10862.Wu, J. et al. Recursively Summarizing Books with Human Feedback. arXiv Preprint at https://arxiv.org/abs/2109.10862 (2021).J. Wu et al., “Recursively Summarizing Books with Human Feedback”, Sep. 22, 2021. [Online]. Available: https://arxiv.org/abs/2109.10862
Xiang(2023). 'He Would Still Be Here': Man Dies by Suicide After Talking with AI Chatbot, Widow Says. VICE.Xiang. (2023, March 30). 'He Would Still Be Here': Man Dies by Suicide After Talking with AI Chatbot, Widow Says. VICE. https://vice.com/en/article/pkadgm/man-dies-by-suicide-after-talking-with-ai-chatbot-widow-saysXiang. 2023. “'He Would Still Be Here': Man Dies by Suicide After Talking with AI Chatbot, Widow Says”. VICE, March 30. https://vice.com/en/article/pkadgm/man-dies-by-suicide-after-talking-with-ai-chatbot-widow-says.Xiang. “'He Would Still Be Here': Man Dies by Suicide After Talking with AI Chatbot, Widow Says”. VICE, 30 Mar. 2023, https://vice.com/en/article/pkadgm/man-dies-by-suicide-after-talking-with-ai-chatbot-widow-says.Xiang. 'He Would Still Be Here': Man Dies by Suicide After Talking with AI Chatbot, Widow Says. VICE https://vice.com/en/article/pkadgm/man-dies-by-suicide-after-talking-with-ai-chatbot-widow-says (2023).Xiang, “'He Would Still Be Here': Man Dies by Suicide After Talking with AI Chatbot, Widow Says”, VICE. [Online]. Available: https://vice.com/en/article/pkadgm/man-dies-by-suicide-after-talking-with-ai-chatbot-widow-says
Xiao, Y. & Wang, W. Y.(2021). On Hallucination and Predictive Uncertainty in Conditional Language Generation. arXiv.Xiao, Y., & Wang, W. Y. (2021). On Hallucination and Predictive Uncertainty in Conditional Language Generation. In arXiv. https://arxiv.org/abs/2103.15025Xiao, Y., and W. Y. Wang. 2021. “On Hallucination and Predictive Uncertainty in Conditional Language Generation”. In arXiv. Preprint, March 28. https://arxiv.org/abs/2103.15025.Xiao, Y., and W. Y. Wang. “On Hallucination and Predictive Uncertainty in Conditional Language Generation”. arXiv, 28 Mar. 2021, https://arxiv.org/abs/2103.15025.Xiao, Y. & Wang, W. Y. On Hallucination and Predictive Uncertainty in Conditional Language Generation. arXiv Preprint at https://arxiv.org/abs/2103.15025 (2021).Y. Xiao and W. Y. Wang, “On Hallucination and Predictive Uncertainty in Conditional Language Generation”, Mar. 28, 2021. [Online]. Available: https://arxiv.org/abs/2103.15025
Xu, F. et al.(2025). Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models. arXiv.Xu, F., Hao, Q., Zong, Z., Wang, J., Zhang, Y., Wang, J., Lan, X., Gong, J., Ouyang, T., Meng, F., Shao, C., Yan, Y., Yang, Q., Song, Y., Ren, S., Hu, X., Li, Y., Feng, J., Gao, C., & Li, Y. (2025). Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models. In arXiv. https://arxiv.org/abs/2501.09686Xu, F., Q. Hao, Z. Zong, et al. 2025. “Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models”. In arXiv. Preprint, January 16. https://arxiv.org/abs/2501.09686.Xu, F., et al. “Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models”. arXiv, 16 Jan. 2025, https://arxiv.org/abs/2501.09686.Xu, F. et al. Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models. arXiv Preprint at https://arxiv.org/abs/2501.09686 (2025).F. Xu et al., “Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models”, Jan. 16, 2025. [Online]. Available: https://arxiv.org/abs/2501.09686
Xu, H., Zhao, R., Zhu, L., Du, J. & He, Y.(2024). OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models. arXiv.Xu, H., Zhao, R., Zhu, L., Du, J., & He, Y. (2024). OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models. In arXiv. https://arxiv.org/abs/2402.06044Xu, H., R. Zhao, L. Zhu, J. Du, and Y. He. 2024. “OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models”. In arXiv. Preprint, February 8. https://arxiv.org/abs/2402.06044.Xu, H., et al. “OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models”. arXiv, 8 Feb. 2024, https://arxiv.org/abs/2402.06044.Xu, H., Zhao, R., Zhu, L., Du, J. & He, Y. OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models. arXiv Preprint at https://arxiv.org/abs/2402.06044 (2024).H. Xu, R. Zhao, L. Zhu, J. Du, and Y. He, “OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models”, Feb. 08, 2024. [Online]. Available: https://arxiv.org/abs/2402.06044
Xu, Y., Deng, B., Wang, J., Jing, Y., Pan, J. & He, S.(2022). High-resolution Face Swapping via Latent Semantics Disentanglement. arXiv.Xu, Y., Deng, B., Wang, J., Jing, Y., Pan, J., & He, S. (2022). High-resolution Face Swapping via Latent Semantics Disentanglement. In arXiv. https://arxiv.org/abs/2203.15958Xu, Y., B. Deng, J. Wang, Y. Jing, J. Pan, and S. He. 2022. “High-resolution Face Swapping via Latent Semantics Disentanglement”. In arXiv. Preprint, March 30. https://arxiv.org/abs/2203.15958.Xu, Y., et al. “High-resolution Face Swapping via Latent Semantics Disentanglement”. arXiv, 30 Mar. 2022, https://arxiv.org/abs/2203.15958.Xu, Y. et al. High-resolution Face Swapping via Latent Semantics Disentanglement. arXiv Preprint at https://arxiv.org/abs/2203.15958 (2022).Y. Xu, B. Deng, J. Wang, Y. Jing, J. Pan, and S. He, “High-resolution Face Swapping via Latent Semantics Disentanglement”, Mar. 30, 2022. [Online]. Available: https://arxiv.org/abs/2203.15958
Yafah Edelman & Anson Ho(2025). Compute scaling will slow down due to increasing lead times.Yafah Edelman, & Anson Ho. (2025, September 5). Compute scaling will slow down due to increasing lead times. https://epoch.ai/gradient-updates/compute-scaling-will-slow-down-due-to-increasing-lead-timesYafah Edelman, and Anson Ho. 2025. “Compute Scaling Will Slow down Due to Increasing Lead Times”. September 5. https://epoch.ai/gradient-updates/compute-scaling-will-slow-down-due-to-increasing-lead-times.Yafah Edelman, and Anson Ho. Compute Scaling Will Slow down Due to Increasing Lead Times. 5 Sept. 2025, https://epoch.ai/gradient-updates/compute-scaling-will-slow-down-due-to-increasing-lead-times.Yafah Edelman & Anson Ho. Compute scaling will slow down due to increasing lead times. https://epoch.ai/gradient-updates/compute-scaling-will-slow-down-due-to-increasing-lead-times (2025).Yafah Edelman and Anson Ho, “Compute scaling will slow down due to increasing lead times”. [Online]. Available: https://epoch.ai/gradient-updates/compute-scaling-will-slow-down-due-to-increasing-lead-times
Yampolskiy, R. V.(2024). AI: Unexplainable, Unpredictable, Uncontrollable.Yampolskiy, R. V. (2024). AI: Unexplainable, Unpredictable, Uncontrollable. CRC Press. https://books.google.se/books/about/AI.html?id=V3XsEAAAQBAJ&redir_esc=yYampolskiy, R. V. 2024. AI: Unexplainable, Unpredictable, Uncontrollable. CRC Press. https://books.google.se/books/about/AI.html?id=V3XsEAAAQBAJ&redir_esc=y.Yampolskiy, R. V. AI: Unexplainable, Unpredictable, Uncontrollable. CRC Press, 2024, https://books.google.se/books/about/AI.html?id=V3XsEAAAQBAJ&redir_esc=y.Yampolskiy, R. V. AI: Unexplainable, Unpredictable, Uncontrollable. (CRC Press, 2024).R. V. Yampolskiy, AI: Unexplainable, Unpredictable, Uncontrollable. CRC Press, 2024. [Online]. Available: https://books.google.se/books/about/AI.html?id=V3XsEAAAQBAJ&redir_esc=y
Yampolsky;(2024). Transcript for Roman Yampolskiy: Dangers of Superintelligent AI. Lex Fridman.Yampolsky;. (2024, June 3). Transcript for Roman Yampolskiy: Dangers of Superintelligent AI | Lex Fridman Podcast #431. Lex Fridman. https://lexfridman.com/roman-yampolskiy-transcriptYampolsky;. 2024. “Transcript for Roman Yampolskiy: Dangers of Superintelligent AI | Lex Fridman Podcast #431”. Lex Fridman, June 3. https://lexfridman.com/roman-yampolskiy-transcript.Yampolsky;. “Transcript for Roman Yampolskiy: Dangers of Superintelligent AI | Lex Fridman Podcast #431”. Lex Fridman, 3 June 2024, https://lexfridman.com/roman-yampolskiy-transcript.Yampolsky;. Transcript for Roman Yampolskiy: Dangers of Superintelligent AI | Lex Fridman Podcast #431. Lex Fridman https://lexfridman.com/roman-yampolskiy-transcript (2024).Yampolsky;, “Transcript for Roman Yampolskiy: Dangers of Superintelligent AI | Lex Fridman Podcast #431”, Lex Fridman. [Online]. Available: https://lexfridman.com/roman-yampolskiy-transcript
Yang et. al;(2022). Chain of Thought Imitation with Procedure Cloning. arXiv.org.Yang et. al;. (2022). Chain of Thought Imitation with Procedure Cloning. arXiv.org. https://www.arxiv.org/abs/2205.10816Yang et. al;. 2022. “Chain of Thought Imitation with Procedure Cloning”. arXiv.org. https://www.arxiv.org/abs/2205.10816.Yang et. al;. “Chain of Thought Imitation with Procedure Cloning”. arXiv.org, 2022, https://www.arxiv.org/abs/2205.10816.Yang et. al;. Chain of Thought Imitation with Procedure Cloning. arXiv.org https://www.arxiv.org/abs/2205.10816 (2022).Yang et. al;, “Chain of Thought Imitation with Procedure Cloning”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2205.10816
Yang, J., Prabhakar, A., Narasimhan, K. & Yao, S.(2023). InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. arXiv.Yang, J., Prabhakar, A., Narasimhan, K., & Yao, S. (2023). InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. In arXiv. https://arxiv.org/abs/2306.14898Yang, J., A. Prabhakar, K. Narasimhan, and S. Yao. 2023. “InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback”. In arXiv. Preprint, June 26. https://arxiv.org/abs/2306.14898.Yang, J., et al. “InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback”. arXiv, 26 June 2023, https://arxiv.org/abs/2306.14898.Yang, J., Prabhakar, A., Narasimhan, K. & Yao, S. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. arXiv Preprint at https://arxiv.org/abs/2306.14898 (2023).J. Yang, A. Prabhakar, K. Narasimhan, and S. Yao, “InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback”, Jun. 26, 2023. [Online]. Available: https://arxiv.org/abs/2306.14898
Yao, S. et al.(2023). Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv.org.Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., & Narasimhan, K. (2023). Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In arXiv.org. https://arxiv.org/abs/2305.10601Yao, S., D. Yu, J. Zhao, et al. 2023. “Tree of Thoughts: Deliberate Problem Solving with Large Language Models”. In arXiv.org. Preprint, May 17. https://arxiv.org/abs/2305.10601.Yao, S., et al. “Tree of Thoughts: Deliberate Problem Solving with Large Language Models”. arXiv.org, 17 May 2023, https://arxiv.org/abs/2305.10601.Yao, S. et al. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv.org Preprint at https://arxiv.org/abs/2305.10601 (2023).S. Yao et al., “Tree of Thoughts: Deliberate Problem Solving with Large Language Models”, May 17, 2023. [Online]. Available: https://arxiv.org/abs/2305.10601
Yap(2024). How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs). Gold Penguin.Yap. (2024, January 8). How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs). Gold Penguin. https://goldpenguin.org/blog/midjourney-v1-to-v6-evolutionYap. 2024. “How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs)”. Gold Penguin, January 8. https://goldpenguin.org/blog/midjourney-v1-to-v6-evolution.Yap. “How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs)”. Gold Penguin, 8 Jan. 2024, https://goldpenguin.org/blog/midjourney-v1-to-v6-evolution.Yap. How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs). Gold Penguin https://goldpenguin.org/blog/midjourney-v1-to-v6-evolution (2024).Yap, “How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs)”, Gold Penguin. [Online]. Available: https://goldpenguin.org/blog/midjourney-v1-to-v6-evolution
Ye, W., Liu, S., Kurutach, T., Abbeel, P. & Gao, Y.(2021). Mastering Atari Games with Limited Data. arXiv.Ye, W., Liu, S., Kurutach, T., Abbeel, P., & Gao, Y. (2021). Mastering Atari Games with Limited Data. In arXiv. https://arxiv.org/abs/2111.00210Ye, W., S. Liu, T. Kurutach, P. Abbeel, and Y. Gao. 2021. “Mastering Atari Games with Limited Data”. In arXiv. Preprint, October 30. https://arxiv.org/abs/2111.00210.Ye, W., et al. “Mastering Atari Games with Limited Data”. arXiv, 30 Oct. 2021, https://arxiv.org/abs/2111.00210.Ye, W., Liu, S., Kurutach, T., Abbeel, P. & Gao, Y. Mastering Atari Games with Limited Data. arXiv Preprint at https://arxiv.org/abs/2111.00210 (2021).W. Ye, S. Liu, T. Kurutach, P. Abbeel, and Y. Gao, “Mastering Atari Games with Limited Data”, Oct. 30, 2021. [Online]. Available: https://arxiv.org/abs/2111.00210
You et al.(2025). How much power will frontier AI training demand in 2030?. Epoch AI.You et al. (2025). How much power will frontier AI training demand in 2030?. Epoch AI. https://epoch.ai/blog/power-demands-of-frontier-ai-trainingYou et al. 2025. “How Much Power Will Frontier AI Training Demand in 2030?”. Epoch AI. https://epoch.ai/blog/power-demands-of-frontier-ai-training.You et al. “How Much Power Will Frontier AI Training Demand in 2030?”. Epoch AI, 2025, https://epoch.ai/blog/power-demands-of-frontier-ai-training.You et al. How much power will frontier AI training demand in 2030?. Epoch AI https://epoch.ai/blog/power-demands-of-frontier-ai-training (2025).You et al., “How much power will frontier AI training demand in 2030?”, Epoch AI. [Online]. Available: https://epoch.ai/blog/power-demands-of-frontier-ai-training
Yu, J. et al.(2022). Scaling Autoregressive Models for Content-Rich Text-to-Image Generation. arXiv.Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., Hutchinson, B., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., & Wu, Y. (2022). Scaling Autoregressive Models for Content-Rich Text-to-Image Generation. In arXiv. https://arxiv.org/abs/2206.10789Yu, J., Y. Xu, J. Y. Koh, et al. 2022. “Scaling Autoregressive Models for Content-Rich Text-to-Image Generation”. In arXiv. Preprint, June 22. https://arxiv.org/abs/2206.10789.Yu, J., et al. “Scaling Autoregressive Models for Content-Rich Text-to-Image Generation”. arXiv, 22 June 2022, https://arxiv.org/abs/2206.10789.Yu, J. et al. Scaling Autoregressive Models for Content-Rich Text-to-Image Generation. arXiv Preprint at https://arxiv.org/abs/2206.10789 (2022).J. Yu et al., “Scaling Autoregressive Models for Content-Rich Text-to-Image Generation”, Jun. 22, 2022. [Online]. Available: https://arxiv.org/abs/2206.10789
Yudkowsky(2002). The AI-Box Experiment:. Eliezer S. Yudkowsky.Yudkowsky. (2002). The AI-Box Experiment:. Eliezer S. Yudkowsky. https://yudkowsky.net/singularity/aiboxYudkowsky. 2002. “The AI-Box Experiment:”. Eliezer S. Yudkowsky. https://yudkowsky.net/singularity/aibox.Yudkowsky. “The AI-Box Experiment:”. Eliezer S. Yudkowsky, 2002, https://yudkowsky.net/singularity/aibox.Yudkowsky. The AI-Box Experiment:. Eliezer S. Yudkowsky https://yudkowsky.net/singularity/aibox (2002).Yudkowsky, “The AI-Box Experiment:”, Eliezer S. Yudkowsky. [Online]. Available: https://yudkowsky.net/singularity/aibox
Yudkowsky(2022). AGI Ruin: A List of Lethalities - Machine Intelligence Research Institute.Yudkowsky. (2022, June 10). AGI Ruin: A List of Lethalities - Machine Intelligence Research Institute. Machine Intelligence Research Institute. https://intelligence.org/2022/06/10/agi-ruinYudkowsky. 2022. “AGI Ruin: A List of Lethalities - Machine Intelligence Research Institute”. Machine Intelligence Research Institute, June 10. https://intelligence.org/2022/06/10/agi-ruin.Yudkowsky. “AGI Ruin: A List of Lethalities - Machine Intelligence Research Institute”. Machine Intelligence Research Institute, 10 June 2022, https://intelligence.org/2022/06/10/agi-ruin.Yudkowsky. AGI Ruin: A List of Lethalities - Machine Intelligence Research Institute. Machine Intelligence Research Institute https://intelligence.org/2022/06/10/agi-ruin (2022).Yudkowsky, “AGI Ruin: A List of Lethalities - Machine Intelligence Research Institute”, Machine Intelligence Research Institute. [Online]. Available: https://intelligence.org/2022/06/10/agi-ruin
Yudkowsky(2023). Pausing AI Developments Isn't Enough. We Need to Shut it All Down - Machine Intelligence Research Institute.Yudkowsky. (2023, April 7). Pausing AI Developments Isn't Enough. We Need to Shut it All Down - Machine Intelligence Research Institute. Machine Intelligence Research Institute. https://intelligence.org/2023/04/07/pausing-ai-developments-isnt-enough-we-need-to-shut-it-all-downYudkowsky. 2023. “Pausing AI Developments Isn't Enough. We Need to Shut It All Down - Machine Intelligence Research Institute”. Machine Intelligence Research Institute, April 7. https://intelligence.org/2023/04/07/pausing-ai-developments-isnt-enough-we-need-to-shut-it-all-down.Yudkowsky. “Pausing AI Developments Isn't Enough. We Need to Shut It All Down - Machine Intelligence Research Institute”. Machine Intelligence Research Institute, 7 Apr. 2023, https://intelligence.org/2023/04/07/pausing-ai-developments-isnt-enough-we-need-to-shut-it-all-down.Yudkowsky. Pausing AI Developments Isn't Enough. We Need to Shut it All Down - Machine Intelligence Research Institute. Machine Intelligence Research Institute https://intelligence.org/2023/04/07/pausing-ai-developments-isnt-enough-we-need-to-shut-it-all-down (2023).Yudkowsky, “Pausing AI Developments Isn't Enough. We Need to Shut it All Down - Machine Intelligence Research Institute”, Machine Intelligence Research Institute. [Online]. Available: https://intelligence.org/2023/04/07/pausing-ai-developments-isnt-enough-we-need-to-shut-it-all-down
Yudkowsky, E.(2004). Coherent Extrapolated Volition.Yudkowsky, E. (2004). Coherent Extrapolated Volition. The Singularity Institute. https://intelligence.org/files/CEV.pdfYudkowsky, E. 2004. Coherent Extrapolated Volition. The Singularity Institute. https://intelligence.org/files/CEV.pdf.Yudkowsky, E. Coherent Extrapolated Volition. The Singularity Institute, 2004, https://intelligence.org/files/CEV.pdf.Yudkowsky, E. Coherent Extrapolated Volition. https://intelligence.org/files/CEV.pdf (2004).E. Yudkowsky, “Coherent Extrapolated Volition”, The Singularity Institute, San Francisco, CA, 2004. [Online]. Available: https://intelligence.org/files/CEV.pdf
Yudkowsky, E.(2013). Intelligence Explosion Microeconomics.Yudkowsky, E. (2013). Intelligence Explosion Microeconomics (2013-1). Machine Intelligence Research Institute. https://intelligence.org/files/IEM.pdfYudkowsky, E. 2013. Intelligence Explosion Microeconomics. 2013-1. Machine Intelligence Research Institute. https://intelligence.org/files/IEM.pdf.Yudkowsky, E. Intelligence Explosion Microeconomics. 2013-1, Machine Intelligence Research Institute, 2013, https://intelligence.org/files/IEM.pdf.Yudkowsky, E. Intelligence Explosion Microeconomics. https://intelligence.org/files/IEM.pdf (2013).E. Yudkowsky, “Intelligence Explosion Microeconomics”, Machine Intelligence Research Institute, Berkeley, CA, 2013-1, 2013. [Online]. Available: https://intelligence.org/files/IEM.pdf
zac_kenton et al.(2022). Clarifying AI X-risk. LessWrong.zac_kenton, Rohin Shah, David Lindner, Vikrant Varma, Vika, Mary Phuong, Ramana Kumar, & Elliot Catt. (2022, November 1). Clarifying AI X-risk. LessWrong. https://lesswrong.com/posts/GctJD5oCDRxCspEaZ/clarifying-ai-x-riskzac_kenton, Rohin Shah, David Lindner, et al. 2022. “Clarifying AI X-risk”. LessWrong, November 1. https://lesswrong.com/posts/GctJD5oCDRxCspEaZ/clarifying-ai-x-risk.zac_kenton, et al. “Clarifying AI X-risk”. LessWrong, 1 Nov. 2022, https://lesswrong.com/posts/GctJD5oCDRxCspEaZ/clarifying-ai-x-risk.zac_kenton et al. Clarifying AI X-risk. LessWrong https://lesswrong.com/posts/GctJD5oCDRxCspEaZ/clarifying-ai-x-risk (2022).zac_kenton et al., “Clarifying AI X-risk”, LessWrong. [Online]. Available: https://lesswrong.com/posts/GctJD5oCDRxCspEaZ/clarifying-ai-x-risk
Zaidan, E. & Ibrahim, I. A.(2024). AI Governance in a Complex and Rapidly Changing Regulatory Landscape: A Global Perspective. Humanities and Social Sciences Communications.Zaidan, E., & Ibrahim, I. A. (2024). AI Governance in a Complex and Rapidly Changing Regulatory Landscape: A Global Perspective. Humanities and Social Sciences Communications, 11, 1121. https://doi.org/10.1057/s41599-024-03560-xZaidan, E., and I. A. Ibrahim. 2024. “AI Governance in a Complex and Rapidly Changing Regulatory Landscape: A Global Perspective”. Humanities and Social Sciences Communications 11 (September): 1121. https://doi.org/10.1057/s41599-024-03560-x.Zaidan, E., and I. A. Ibrahim. “AI Governance in a Complex and Rapidly Changing Regulatory Landscape: A Global Perspective”. Humanities and Social Sciences Communications, vol. 11, Sept. 2024, p. 1121, https://doi.org/10.1057/s41599-024-03560-x.Zaidan, E. & Ibrahim, I. A. AI Governance in a Complex and Rapidly Changing Regulatory Landscape: A Global Perspective. Humanities and Social Sciences Communications 11, 1121 (2024).E. Zaidan and I. A. Ibrahim, “AI Governance in a Complex and Rapidly Changing Regulatory Landscape: A Global Perspective”, Humanities and Social Sciences Communications, vol. 11, p. 1121, Sep. 2024, doi: 10.1057/s41599-024-03560-x.
Zellers, R., Bisk, Y., Schwartz, R. & Choi, Y.(2018). SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference. arXiv.Zellers, R., Bisk, Y., Schwartz, R., & Choi, Y. (2018). SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference. In arXiv. https://arxiv.org/abs/1808.05326Zellers, R., Y. Bisk, R. Schwartz, and Y. Choi. 2018. “SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference”. In arXiv. Preprint, August 16. https://arxiv.org/abs/1808.05326.Zellers, R., et al. “SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference”. arXiv, 16 Aug. 2018, https://arxiv.org/abs/1808.05326.Zellers, R., Bisk, Y., Schwartz, R. & Choi, Y. SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference. arXiv Preprint at https://arxiv.org/abs/1808.05326 (2018).R. Zellers, Y. Bisk, R. Schwartz, and Y. Choi, “SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference”, Aug. 16, 2018. [Online]. Available: https://arxiv.org/abs/1808.05326
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A. & Choi, Y.(2019). HellaSwag: Can a Machine Really Finish Your Sentence?. arXiv.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., & Choi, Y. (2019). HellaSwag: Can a Machine Really Finish Your Sentence?. In arXiv. https://arxiv.org/abs/1905.07830Zellers, R., A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. 2019. “HellaSwag: Can a Machine Really Finish Your Sentence?”. In arXiv. Preprint, May 19. https://arxiv.org/abs/1905.07830.Zellers, R., et al. “HellaSwag: Can a Machine Really Finish Your Sentence?”. arXiv, 19 May 2019, https://arxiv.org/abs/1905.07830.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A. & Choi, Y. HellaSwag: Can a Machine Really Finish Your Sentence?. arXiv Preprint at https://arxiv.org/abs/1905.07830 (2019).R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “HellaSwag: Can a Machine Really Finish Your Sentence?”, May 19, 2019. [Online]. Available: https://arxiv.org/abs/1905.07830
Zhang et al.(2025). A Three-Layered Framework: An AI Governance Guide for Global Policymakers.Zhang et al. (2025). A Three-Layered Framework: An AI Governance Guide for Global Policymakers. Internet Archive (https://web.archive.org/web/20250601042409/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5241351). https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5241351Zhang et al. 2025. “A Three-Layered Framework: An AI Governance Guide for Global Policymakers”. Https://web.archive.org/web/20250601042409/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5241351. Internet Archive. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5241351.Zhang et al. A Three-Layered Framework: An AI Governance Guide for Global Policymakers. 2025, Internet Archive, https://web.archive.org/web/20250601042409/https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5241351, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5241351.Zhang et al. A Three-Layered Framework: An AI Governance Guide for Global Policymakers. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5241351 (2025).Zhang et al., “A Three-Layered Framework: An AI Governance Guide for Global Policymakers”. Accessed: Jun. 01, 2025. [Online]. Available: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5241351
Zhang, B., Anderljung, M., Kahn, L., Dreksler, N., Horowitz, M. C. & Dafoe, A.(2021). Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers. arXiv.Zhang, B., Anderljung, M., Kahn, L., Dreksler, N., Horowitz, M. C., & Dafoe, A. (2021). Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers. In arXiv. https://arxiv.org/abs/2105.02117Zhang, B., M. Anderljung, L. Kahn, N. Dreksler, M. C. Horowitz, and A. Dafoe. 2021. “Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers”. In arXiv. Preprint, May 5. https://arxiv.org/abs/2105.02117.Zhang, B., et al. “Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers”. arXiv, 5 May 2021, https://arxiv.org/abs/2105.02117.Zhang, B. et al. Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers. arXiv Preprint at https://arxiv.org/abs/2105.02117 (2021).B. Zhang, M. Anderljung, L. Kahn, N. Dreksler, M. C. Horowitz, and A. Dafoe, “Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers”, May 05, 2021. [Online]. Available: https://arxiv.org/abs/2105.02117
Zhang, F. et al.(2024). HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks. arXiv.Zhang, F., Wu, L., Bai, H., Lin, G., Li, X., Yu, X., Wang, Y., Chen, B., & Keung, J. (2024). HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks. In arXiv. https://arxiv.org/abs/2410.12381Zhang, F., L. Wu, H. Bai, et al. 2024. “HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks”. In arXiv. Preprint, October 16. https://arxiv.org/abs/2410.12381.Zhang, F., et al. “HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks”. arXiv, 16 Oct. 2024, https://arxiv.org/abs/2410.12381.Zhang, F. et al. HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks. arXiv Preprint at https://arxiv.org/abs/2410.12381 (2024).F. Zhang et al., “HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks”, Oct. 16, 2024. [Online]. Available: https://arxiv.org/abs/2410.12381
Zhang, G., Yan, C., Ji, X., Zhang, T., Zhang, T. & Xu, W.(2017). DolphinAtack: Inaudible Voice Commands. arXiv.Zhang, G., Yan, C., Ji, X., Zhang, T., Zhang, T., & Xu, W. (2017). DolphinAtack: Inaudible Voice Commands. In arXiv. https://doi.org/10.1145/3133956.3134052Zhang, G., C. Yan, X. Ji, T. Zhang, T. Zhang, and W. Xu. 2017. “DolphinAtack: Inaudible Voice Commands”. In arXiv. Preprint, August 31. https://doi.org/10.1145/3133956.3134052.Zhang, G., et al. “DolphinAtack: Inaudible Voice Commands”. arXiv, 31 Aug. 2017, https://doi.org/10.1145/3133956.3134052.Zhang, G. et al. DolphinAtack: Inaudible Voice Commands. arXiv Preprint at https://doi.org/10.1145/3133956.3134052 (2017).G. Zhang, C. Yan, X. Ji, T. Zhang, T. Zhang, and W. Xu, “DolphinAtack: Inaudible Voice Commands”, Aug. 31, 2017. doi: 10.1145/3133956.3134052.
Zhang, Z., Bai, F., Gao, J. & Yang, Y.(2023). ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models. arXiv.Zhang, Z., Bai, F., Gao, J., & Yang, Y. (2023). ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models. In arXiv. https://arxiv.org/abs/2310.00378Zhang, Z., F. Bai, J. Gao, and Y. Yang. 2023. “ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models”. In arXiv. Preprint, September 30. https://arxiv.org/abs/2310.00378.Zhang, Z., et al. “ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models”. arXiv, 30 Sept. 2023, https://arxiv.org/abs/2310.00378.Zhang, Z., Bai, F., Gao, J. & Yang, Y. ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models. arXiv Preprint at https://arxiv.org/abs/2310.00378 (2023).Z. Zhang, F. Bai, J. Gao, and Y. Yang, “ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models”, Sep. 30, 2023. [Online]. Available: https://arxiv.org/abs/2310.00378
Zhao, M., Zhang, L., Ye, J., Lu, H., Yin, B. & Wang, X.(2024). Adversarial Training: A Survey. arXiv.Zhao, M., Zhang, L., Ye, J., Lu, H., Yin, B., & Wang, X. (2024). Adversarial Training: A Survey. In arXiv. https://arxiv.org/abs/2410.15042Zhao, M., L. Zhang, J. Ye, H. Lu, B. Yin, and X. Wang. 2024. “Adversarial Training: A Survey”. In arXiv. Preprint, October 19. https://arxiv.org/abs/2410.15042.Zhao, M., et al. “Adversarial Training: A Survey”. arXiv, 19 Oct. 2024, https://arxiv.org/abs/2410.15042.Zhao, M. et al. Adversarial Training: A Survey. arXiv Preprint at https://arxiv.org/abs/2410.15042 (2024).M. Zhao, L. Zhang, J. Ye, H. Lu, B. Yin, and X. Wang, “Adversarial Training: A Survey”, Oct. 19, 2024. [Online]. Available: https://arxiv.org/abs/2410.15042
Zhou, A., Yan, K., Shlapentokh-Rothman, M., Wang, H. & Wang, Y.(2023). Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. arXiv.Zhou, A., Yan, K., Shlapentokh-Rothman, M., Wang, H., & Wang, Y.-X. (2023). Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. In arXiv. https://arxiv.org/abs/2310.04406Zhou, A., K. Yan, M. Shlapentokh-Rothman, H. Wang, and Y.-X. Wang. 2023. “Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models”. In arXiv. Preprint, October 6. https://arxiv.org/abs/2310.04406.Zhou, A., et al. “Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models”. arXiv, 6 Oct. 2023, https://arxiv.org/abs/2310.04406.Zhou, A., Yan, K., Shlapentokh-Rothman, M., Wang, H. & Wang, Y.-X. Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. arXiv Preprint at https://arxiv.org/abs/2310.04406 (2023).A. Zhou, K. Yan, M. Shlapentokh-Rothman, H. Wang, and Y.-X. Wang, “Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models”, Oct. 06, 2023. [Online]. Available: https://arxiv.org/abs/2310.04406
Zhou, L. et al.(2023). Predictable Artificial Intelligence. arXiv.Zhou, L., Moreno-Casares, P. A., Martínez-Plumed, F., Burden, J., Burnell, R., Cheke, L., Ferri, C., Marcoci, A., Mehrbakhsh, B., Moros-Daval, Y., hÉigeartaigh, S. Ó., Rutar, D., Schellaert, W., Voudouris, K., & Hernández-Orallo, J. (2023). Predictable Artificial Intelligence. In arXiv. https://arxiv.org/abs/2310.06167Zhou, L., P. A. Moreno-Casares, F. Martínez-Plumed, et al. 2023. “Predictable Artificial Intelligence”. In arXiv. Preprint, October 9. https://arxiv.org/abs/2310.06167.Zhou, L., et al. “Predictable Artificial Intelligence”. arXiv, 9 Oct. 2023, https://arxiv.org/abs/2310.06167.Zhou, L. et al. Predictable Artificial Intelligence. arXiv Preprint at https://arxiv.org/abs/2310.06167 (2023).L. Zhou et al., “Predictable Artificial Intelligence”, Oct. 09, 2023. [Online]. Available: https://arxiv.org/abs/2310.06167
Zhu, H. et al.(2024). Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics. arXiv.Zhu, H., Huang, B., Zhang, S., Jordan, M., Jiao, J., Tian, Y., & Russell, S. (2024). Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics. In arXiv. https://arxiv.org/abs/2405.04669Zhu, H., B. Huang, S. Zhang, et al. 2024. “Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics”. In arXiv. Preprint, May 7. https://arxiv.org/abs/2405.04669.Zhu, H., et al. “Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics”. arXiv, 7 May 2024, https://arxiv.org/abs/2405.04669.Zhu, H. et al. Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics. arXiv Preprint at https://arxiv.org/abs/2405.04669 (2024).H. Zhu et al., “Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics”, May 07, 2024. [Online]. Available: https://arxiv.org/abs/2405.04669
Zhu, Y., Li, Q., Wang, J., Xu, C. & Sun, Z.(2021). One Shot Face Swapping on Megapixels. arXiv.Zhu, Y., Li, Q., Wang, J., Xu, C., & Sun, Z. (2021). One Shot Face Swapping on Megapixels. In arXiv. https://arxiv.org/abs/2105.04932Zhu, Y., Q. Li, J. Wang, C. Xu, and Z. Sun. 2021. “One Shot Face Swapping on Megapixels”. In arXiv. Preprint, May 11. https://arxiv.org/abs/2105.04932.Zhu, Y., et al. “One Shot Face Swapping on Megapixels”. arXiv, 11 May 2021, https://arxiv.org/abs/2105.04932.Zhu, Y., Li, Q., Wang, J., Xu, C. & Sun, Z. One Shot Face Swapping on Megapixels. arXiv Preprint at https://arxiv.org/abs/2105.04932 (2021).Y. Zhu, Q. Li, J. Wang, C. Xu, and Z. Sun, “One Shot Face Swapping on Megapixels”, May 11, 2021. [Online]. Available: https://arxiv.org/abs/2105.04932
Zhuo, T. Y. et al.(2024). BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. arXiv.Zhuo, T. Y., Vu, M. C., Chim, J., Hu, H., Yu, W., Widyasari, R., Yusuf, I. N. B., Zhan, H., He, J., Paul, I., Brunner, S., Gong, C., Hoang, T., Zebaze, A. R., Hong, X., Li, W.-D., Kaddour, J., Xu, M., Zhang, Z., … Werra, L. V. (2024). BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. In arXiv. https://arxiv.org/abs/2406.15877Zhuo, T. Y., M. C. Vu, J. Chim, et al. 2024. “BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions”. In arXiv. Preprint, June 22. https://arxiv.org/abs/2406.15877.Zhuo, T. Y., et al. “BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions”. arXiv, 22 June 2024, https://arxiv.org/abs/2406.15877.Zhuo, T. Y. et al. BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. arXiv Preprint at https://arxiv.org/abs/2406.15877 (2024).T. Y. Zhuo et al., “BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions”, Jun. 22, 2024. [Online]. Available: https://arxiv.org/abs/2406.15877
Ziegler, D. M. et al.(2019). Fine-Tuning Language Models from Human Preferences. arXiv.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., & Irving, G. (2019). Fine-Tuning Language Models from Human Preferences. In arXiv. https://arxiv.org/abs/1909.08593Ziegler, D. M., N. Stiennon, J. Wu, et al. 2019. “Fine-Tuning Language Models from Human Preferences”. In arXiv. Preprint, September 18. https://arxiv.org/abs/1909.08593.Ziegler, D. M., et al. “Fine-Tuning Language Models from Human Preferences”. arXiv, 18 Sept. 2019, https://arxiv.org/abs/1909.08593.Ziegler, D. M. et al. Fine-Tuning Language Models from Human Preferences. arXiv Preprint at https://arxiv.org/abs/1909.08593 (2019).D. M. Ziegler et al., “Fine-Tuning Language Models from Human Preferences”, Sep. 18, 2019. [Online]. Available: https://arxiv.org/abs/1909.08593
Žiga Avsec & Natasha Latysheva(2025). AlphaGenome: AI for better understanding the genome.Žiga Avsec, & Natasha Latysheva. (2025, June 25). AlphaGenome: AI for better understanding the genome. https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genomeŽiga Avsec, and Natasha Latysheva. 2025. “AlphaGenome: AI for Better Understanding the Genome”. June 25. https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome.Žiga Avsec, and Natasha Latysheva. AlphaGenome: AI for Better Understanding the Genome. 25 June 2025, https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome.Žiga Avsec & Natasha Latysheva. AlphaGenome: AI for better understanding the genome. https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome (2025).Žiga Avsec and Natasha Latysheva, “AlphaGenome: AI for better understanding the genome”. [Online]. Available: https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome
Zou, A. et al.(2023). Representation Engineering: A Top-Down Approach to AI Transparency. arXiv.Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., … Hendrycks, D. (2023). Representation Engineering: A Top-Down Approach to AI Transparency. In arXiv. https://arxiv.org/abs/2310.01405Zou, A., L. Phan, S. Chen, et al. 2023. “Representation Engineering: A Top-Down Approach to AI Transparency”. In arXiv. Preprint, October 2. https://arxiv.org/abs/2310.01405.Zou, A., et al. “Representation Engineering: A Top-Down Approach to AI Transparency”. arXiv, 2 Oct. 2023, https://arxiv.org/abs/2310.01405.Zou, A. et al. Representation Engineering: A Top-Down Approach to AI Transparency. arXiv Preprint at https://arxiv.org/abs/2310.01405 (2023).A. Zou et al., “Representation Engineering: A Top-Down Approach to AI Transparency”, Oct. 02, 2023. [Online]. Available: https://arxiv.org/abs/2310.01405
Zvi(2025). On MAIM and Superintelligence Strategy. LessWrong.Zvi. (2025, March 14). On MAIM and Superintelligence Strategy. LessWrong. https://lesswrong.com/posts/kYeHbXmW4Kppfkg5j/on-maim-and-superintelligence-strategyZvi. 2025. “On MAIM and Superintelligence Strategy”. LessWrong, March 14. https://lesswrong.com/posts/kYeHbXmW4Kppfkg5j/on-maim-and-superintelligence-strategy.Zvi. “On MAIM and Superintelligence Strategy”. LessWrong, 14 Mar. 2025, https://lesswrong.com/posts/kYeHbXmW4Kppfkg5j/on-maim-and-superintelligence-strategy.Zvi. On MAIM and Superintelligence Strategy. LessWrong https://lesswrong.com/posts/kYeHbXmW4Kppfkg5j/on-maim-and-superintelligence-strategy (2025).Zvi, “On MAIM and Superintelligence Strategy”, LessWrong. [Online]. Available: https://lesswrong.com/posts/kYeHbXmW4Kppfkg5j/on-maim-and-superintelligence-strategy
Zvi(2025). The Paris AI Anti-Safety Summit. LessWrong.Zvi. (2025, February 12). The Paris AI Anti-Safety Summit. LessWrong. https://lesswrong.com/posts/qYPHryHTNiJ2y6Fhi/the-paris-ai-anti-safety-summitZvi. 2025. “The Paris AI Anti-Safety Summit”. LessWrong, February 12. https://lesswrong.com/posts/qYPHryHTNiJ2y6Fhi/the-paris-ai-anti-safety-summit.Zvi. “The Paris AI Anti-Safety Summit”. LessWrong, 12 Feb. 2025, https://lesswrong.com/posts/qYPHryHTNiJ2y6Fhi/the-paris-ai-anti-safety-summit.Zvi. The Paris AI Anti-Safety Summit. LessWrong https://lesswrong.com/posts/qYPHryHTNiJ2y6Fhi/the-paris-ai-anti-safety-summit (2025).Zvi, “The Paris AI Anti-Safety Summit”, LessWrong. [Online]. Available: https://lesswrong.com/posts/qYPHryHTNiJ2y6Fhi/the-paris-ai-anti-safety-summit
Zwetsloot, R. et al.(2021). Skilled and Mobile: Survey Evidence of AI Researchers' Immigration Preferences. arXiv.Zwetsloot, R., Zhang, B., Dreksler, N., Kahn, L., Anderljung, M., Dafoe, A., & Horowitz, M. C. (2021). Skilled and Mobile: Survey Evidence of AI Researchers' Immigration Preferences. In arXiv. https://doi.org/10.1145/3461702.3462617Zwetsloot, R., B. Zhang, N. Dreksler, et al. 2021. “Skilled and Mobile: Survey Evidence of AI Researchers' Immigration Preferences”. In arXiv. Preprint, April 15. https://doi.org/10.1145/3461702.3462617.Zwetsloot, R., et al. “Skilled and Mobile: Survey Evidence of AI Researchers' Immigration Preferences”. arXiv, 15 Apr. 2021, https://doi.org/10.1145/3461702.3462617.Zwetsloot, R. et al. Skilled and Mobile: Survey Evidence of AI Researchers' Immigration Preferences. arXiv Preprint at https://doi.org/10.1145/3461702.3462617 (2021).R. Zwetsloot et al., “Skilled and Mobile: Survey Evidence of AI Researchers' Immigration Preferences”, Apr. 15, 2021. doi: 10.1145/3461702.3462617.