Strategies to prevent misuse often focus on controlling access to dangerous capabilities or implementing technical safeguards to limit harmful applications.
External Access Controls #
Access control strategies directly address the inherent tension between open-sourcing benefits and misuse risks. The AI industry has moved beyond binary discussions of "release" or "don't release"; instead, practitioners think in terms of a continuous gradient of access to models (Kapoor et al., 2024). The question of who gets access to a model sits on a range from fully closed (internal use only) to fully open (publicly available model weights with no restrictions).
Among these various access options, API-based deployment represents one of the most commonly used strategic middle grounds. When we discuss access controls in this section, we're primarily talking about mechanisms that create a controlled gateway to AI capabilities—most commonly through API-based deployment, where most of the model (code, weights, and data) remain fully closed, but access to model capabilities is partially open. In this arrangement, developers retain control over how their models are accessed and used. API-based controls maintain developer oversight, allowing continuous monitoring, updating of safety measures, and the ability to revoke access when necessary (Seger et al., 2023).
API-based deployment establishes a protective layer between users and model capabilities. Instead of downloading model code or weights, users interact with the model by sending requests to a server where the model runs, receiving only the generated outputs in return. This architecture enables developers to implement various safety mechanisms:
- Input/Output Filtering: Screening prompts for harmful content and filtering generated responses according to safety policies. For example, filters can detect and block attempted generation of CSAM or instructions for building weapons. This approach directly counters misuses like generating illegal content or dangerous instructions.
- Rate Limiting: Preventing large-scale misuse through usage caps. By restricting the volume of requests, these controls mitigate risks of automated abuse like generating thousands of deepfakes or spam messages (Liang et al., 2022).
- Usage Monitoring: Beyond controlling request volume, usage monitoring enables identity and background checks for malicious users (similar to know your customer (KYC) laws). For example, this allows regulatory oversight, prevents repeated attempts to circumvent safety filters, and also enables deeper access to highly trusted users (Egan & Heim, 2023).
- Usage Restrictions: Enforcing terms of service that prohibit harmful applications. Companies can restrict high-risk applications like bioweapon research or autonomous cyber operations through legal agreements backed by technical monitoring (Anderljung et al., 2023). When violations are detected, access can be revoked.
- On-the-fly Updates: Rapidly deploying improvements to safety systems without user action. Unlike open-sourced models where unsafe versions persist indefinitely, API-based models can be continually improved to address newly discovered vulnerabilities (Weidinger et al., 2023). This helps counter novel attack vectors like jailbreaking techniques.
Most systems that are too dangerous to open source are probably too dangerous to be trained at all, given the kind of practices that are common in labs today, where it's very plausible they'll leak, or very plausible they'll be stolen, or very plausible if they're available over an API, they could cause harm.
Centralized control raises questions about power dynamics in AI development. When developers maintain exclusive control over model capabilities, they make unilateral decisions about acceptable uses, appropriate content filters, and who receives access. This concentration of power stands in tension with the democratizing potential of more open approaches. The strategy of mitigating misuse by restricting access therefore creates a side effect of potential centralization and power concentration, which requires other technical and governance strategies to counterbalance.
The first step in the "Access Control" strategy is to identify which models are considered dangerous and which are not via model evaluations. Before deploying powerful models, developers (or third parties) should evaluate them for specific dangerous capabilities, such as the ability to assist in cyberattacks or bioweapon design. These evaluations inform decisions about deployment and necessary safeguards (Shevlane et al., 2023).
Red Teaming can help assess if the mitigations are sufficient. During red teaming, internal teams try to exploit weaknesses in the system to improve its security. They should test whether a hypothetical malicious user can get a sufficient amount of bits of advice from the model without getting caught. We go into much more detail on concepts like red teaming and model evaluations in the subsequent dedicated chapter to the topic.
Internal Access Controls #
Internal access controls protect model weights and algorithmic secrets. While external access controls regulate how users interact with AI systems through APIs and other interfaces, internal access controls focus on securing the model weights themselves. If model weights are exfiltrated, all external access controls become irrelevant, as the model can be deployed without any restrictions. Several risk models often assume catastrophic risk due to weight exfiltration and espionage (Aschenbrenner, 2024; Nevo et al., 2024; Kokotajlo et al., 2025). Research labs developing cutting-edge models should implement rigorous cybersecurity measures to protect AI systems against theft. This seems simple, but it's not, and protecting models from nation-state-level actors could require extraordinary effort (Ladish & Heim, 2022). In this section, we try to explore strategies to protect model weights and protect algorithmic insights from unauthorized access, theft, or misuse by insiders or external attackers.
Adequate protection requires a multi-layered defense spanning technical, organizational, and physical domains. As an example, think about a frontier AI lab that wants to protect its most advanced model: technical controls encrypt the weights and limit digital access; organizational controls restrict knowledge of the model architecture to a small team of vetted researchers; and physical controls ensure the compute infrastructure remains in secure facilities with restricted access. If any single layer fails—for instance, if the encryption is broken but the physical access restrictions remain—the model still maintains some protection. This defense-in-depth approach ensures that multiple security failures would need to co-occur for a successful exfiltration.
Technical Safeguards #
Beyond access control and instruction tuning techniques like reinforcement learning from human feedback (RLHF), researchers are developing techniques to build safety mechanisms directly into the models themselves or their deployment pipelines. This adds another layer of defense in preventing potential misuse. The reason this section is listed under access control methods is that the vast majority of the technical safeguards that we can put in place require the developers to maintain access control over models. If there is an entirely open source model, then technical safeguards cannot be guaranteed.
Circuit Breakers. Inspired by representation engineering, circuit breakers aim to detect and interrupt the internal activation patterns associated with harmful outputs as they form (Andy Zou et al., 2024). By "rerouting" these harmful representations (e.g., using Representation Rerouting with LoRRA), this technique can prevent the generation of toxic content, demonstrating robustness against unseen adversarial attacks while preserving model utility when the request is not harmful. This approach targets the model's intrinsic capacity for harm, making it potentially more robust than input/output filtering.
Machine “Unlearning” involves techniques to selectively remove specific knowledge or capabilities from a trained model without full retraining. Applications relevant to misuse prevention include removing knowledge about dangerous substances or weapons, erasing harmful biases, or removing jailbreak vulnerabilities. Some researchers think that the ability to selectively and robustly remove capabilities could end up being really valuable in a wide range of scenarios, as well as being tractable (Casper, 2023). Techniques range from gradient-based methods to parameter modification and model editing. However, challenges remain in ensuring complete and robust forgetting, avoiding catastrophic forgetting of useful knowledge, and scaling these methods efficiently.
Socio-technical Strategies #
The previous strategies focus on reducing risks from models that are not yet widely available, such as models capable of advanced cyberattacks or engineering pathogens. However, what about models that enable deep fakes, misinformation campaigns, or privacy violations? Many of these models are already widely accessible.
Unfortunately, it is already too easy to use open-source models to do things like creating sexualized images of people from a few photos of them. There is no purely technical solution to counter such problems. For example, adding defenses (like adversarial noise) to photos published online to make them unreadable by AI will probably not scale, and empirically, every type of defense has been bypassed by attacks in the literature of adversarial attacks.
The primary solution is to regulate and establish strict norms against this type of behavior. Some potential approaches (Control AI, 2024):
- Laws and penalties: Enact and enforce laws making it illegal to create and share non-consensual deep fake pornography or use AI for stalking, harassment, privacy violations, intellectual property, or misinformation. Impose significant penalties as a deterrent.
- Content moderation: Require online platforms to proactively detect and remove AI-generated problematic content, misinformation, and privacy-violating material. Hold platforms accountable for failure to moderate.
- Watermarking: Encourage or require "watermarking" of AI-generated content. Develop standards for digital provenance and authentication.
- Education and awareness: Launch public education campaigns about the risks of deep fakes, misinformation, and AI privacy threats. Teach people to be critical consumers of online content.
- Research: Support research into technical methods of detecting AI-generated content, identifying manipulated media, and preserving privacy from AI systems.
These elements can be combined with other strategies and layers to attain defense in depth. For instance, AI-powered systems can screen phone calls in real-time, analyzing voice patterns, call frequency, and conversational cues to identify likely scams and alert users or block calls (Neuralt, 2024). Chatbots like Daisy (Anna Desmarais, 2024) and services like Jolly Roger Telephone employ AI to engage scammers in lengthy, unproductive conversations, wasting their time and diverting them from potential victims. These represent practical, defense-oriented applications of AI against common forms of misuse. But this is only an early step, and it is far from being sufficient.
Ultimately, a combination of legal frameworks, platform policies, social norms, and technological tools will be needed to mitigate the risks posed by widely available AI models.
References
- Anderljung, M. et al. (2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv.Anderljung, M., Barnhart, J., Korinek, A., Leung, J., O'Keefe, C., Whittlestone, J., Avin, S., Brundage, M., Bullock, J., Cass-Beggs, D., Chang, B., Collins, T., Fist, T., Hadfield, G., Hayes, A., Ho, L., Hooker, S., Horvitz, E., Kolt, N., … Wolf, K. (2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. In arXiv. https://arxiv.org/abs/2307.03718Anderljung, M., J. Barnhart, A. Korinek, et al. 2023. “Frontier AI Regulation: Managing Emerging Risks to Public Safety”. In arXiv. Preprint, July 6. https://arxiv.org/abs/2307.03718.Anderljung, M., et al. “Frontier AI Regulation: Managing Emerging Risks to Public Safety”. arXiv, 6 July 2023, https://arxiv.org/abs/2307.03718.Anderljung, M. et al. Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv Preprint at https://arxiv.org/abs/2307.03718 (2023).M. Anderljung et al., “Frontier AI Regulation: Managing Emerging Risks to Public Safety”, Jul. 06, 2023. [Online]. Available: https://arxiv.org/abs/2307.03718
- Andy K. Zhang et al. (2024). Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models. arXiv.Andy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W. Lin, Eliot Jones, Gashon Hussein, Samantha Liu, Donovan Jasper, Pura Peetathawatchai, Ari Glenn, Vikram Sivashankar, Daniel Zamoshchin, Leo Glikbarg, Derek Askaryar, Mike Yang, Teddy Zhang, Rishi Alluri, … Percy Liang. (2024). Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models. In arXiv. https://arxiv.org/abs/2408.08926Andy K. Zhang, Neil Perry, Riya Dulepet, et al. 2024. “Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models”. In arXiv. Preprint, August 15. https://arxiv.org/abs/2408.08926.Andy K. Zhang, et al. “Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models”. arXiv, 15 Aug. 2024, https://arxiv.org/abs/2408.08926.Andy K. Zhang et al. Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models. arXiv Preprint at https://arxiv.org/abs/2408.08926 (2024).Andy K. Zhang et al., “Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models”, Aug. 15, 2024. [Online]. Available: https://arxiv.org/abs/2408.08926
- Andy Zou et al. (2024). Improving Alignment and Robustness with Circuit Breakers. arXiv.Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, & Dan Hendrycks. (2024). Improving Alignment and Robustness with Circuit Breakers. In arXiv. https://arxiv.org/abs/2406.04313Andy Zou, Long Phan, Justin Wang, et al. 2024. “Improving Alignment and Robustness with Circuit Breakers”. In arXiv. Preprint, June 6. https://arxiv.org/abs/2406.04313.Andy Zou, et al. “Improving Alignment and Robustness with Circuit Breakers”. arXiv, 6 June 2024, https://arxiv.org/abs/2406.04313.Andy Zou et al. Improving Alignment and Robustness with Circuit Breakers. arXiv Preprint at https://arxiv.org/abs/2406.04313 (2024).Andy Zou et al., “Improving Alignment and Robustness with Circuit Breakers”, Jun. 06, 2024. [Online]. Available: https://arxiv.org/abs/2406.04313
- Anna Desmarais (2024). Découvrez Daisy, le chatbot "mamie" qui fait perdre du temps aux fraudeurs au téléphone. euronews.Anna Desmarais. (2024, November 27). Découvrez Daisy, le chatbot "mamie" qui fait perdre du temps aux fraudeurs au téléphone. Euronews. https://fr.euronews.com/next/2024/03/08/decouvrez-daisy-le-chatbot-mamie-qui-fait-perdre-du-temps-aux-fraudeurs-au-telephoneAnna Desmarais. 2024. “Découvrez Daisy, Le Chatbot "mamie" Qui Fait Perdre Du Temps Aux Fraudeurs Au Téléphone”. Euronews, November 27. https://fr.euronews.com/next/2024/03/08/decouvrez-daisy-le-chatbot-mamie-qui-fait-perdre-du-temps-aux-fraudeurs-au-telephone.Anna Desmarais. “Découvrez Daisy, Le Chatbot "mamie" Qui Fait Perdre Du Temps Aux Fraudeurs Au Téléphone”. Euronews, 27 Nov. 2024, https://fr.euronews.com/next/2024/03/08/decouvrez-daisy-le-chatbot-mamie-qui-fait-perdre-du-temps-aux-fraudeurs-au-telephone.Anna Desmarais. Découvrez Daisy, le chatbot "mamie" qui fait perdre du temps aux fraudeurs au téléphone. euronews https://fr.euronews.com/next/2024/03/08/decouvrez-daisy-le-chatbot-mamie-qui-fait-perdre-du-temps-aux-fraudeurs-au-telephone (2024).Anna Desmarais, “Découvrez Daisy, le chatbot "mamie" qui fait perdre du temps aux fraudeurs au téléphone”, euronews. [Online]. Available: https://fr.euronews.com/next/2024/03/08/decouvrez-daisy-le-chatbot-mamie-qui-fait-perdre-du-temps-aux-fraudeurs-au-telephone
- Aschenbrenner (2024). Introduction. SITUATIONAL AWARENESS - The Decade Ahead.Aschenbrenner. (2024). Introduction. SITUATIONAL AWARENESS - The Decade Ahead. https://situational-awareness.aiAschenbrenner. 2024. “Introduction”. SITUATIONAL AWARENESS - The Decade Ahead. https://situational-awareness.ai.Aschenbrenner. “Introduction”. SITUATIONAL AWARENESS - The Decade Ahead, 2024, https://situational-awareness.ai.Aschenbrenner. Introduction. SITUATIONAL AWARENESS - The Decade Ahead https://situational-awareness.ai (2024).Aschenbrenner, “Introduction”, SITUATIONAL AWARENESS - The Decade Ahead. [Online]. Available: https://situational-awareness.ai
- Control AI (2024). Deepfakes Policy. ControlAI.Control AI. (2024). Deepfakes Policy. ControlAI. https://controlai.com/deepfakes/deepfakes-policyControl AI. 2024. “Deepfakes Policy”. ControlAI. https://controlai.com/deepfakes/deepfakes-policy.Control AI. “Deepfakes Policy”. ControlAI, 2024, https://controlai.com/deepfakes/deepfakes-policy.Control AI. Deepfakes Policy. ControlAI https://controlai.com/deepfakes/deepfakes-policy (2024).Control AI, “Deepfakes Policy”, ControlAI. [Online]. Available: https://controlai.com/deepfakes/deepfakes-policy
- Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland & Romeo Dean (2025). AI 2027.Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, & Romeo Dean. (2025). AI 2027. https://ai-2027.comDaniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, and Romeo Dean. 2025. “AI 2027”. https://ai-2027.com.Daniel Kokotajlo, et al. AI 2027. 2025, https://ai-2027.com.Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland & Romeo Dean. AI 2027. https://ai-2027.com (2025).Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, and Romeo Dean, “AI 2027”. [Online]. Available: https://ai-2027.com
- Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland & Romeo Dean (2025). Security Forecast. AI 2027.Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, & Romeo Dean. (2025). Security Forecast. AI 2027. https://ai-2027.com/research/security-forecastDaniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, and Romeo Dean. 2025. “Security Forecast”. AI 2027. https://ai-2027.com/research/security-forecast.Daniel Kokotajlo, et al. “Security Forecast”. AI 2027, 2025, https://ai-2027.com/research/security-forecast.Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland & Romeo Dean. Security Forecast. AI 2027 https://ai-2027.com/research/security-forecast (2025).Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, and Romeo Dean, “Security Forecast”, AI 2027. [Online]. Available: https://ai-2027.com/research/security-forecast
- DeepSeek (2025). GitHub - deepseek-ai/DeepSeek-V3. GitHub.DeepSeek. (2025). GitHub - deepseek-ai/DeepSeek-V3. GitHub. https://github.com/deepseek-ai/DeepSeek-V3DeepSeek. 2025. “GitHub - deepseek-ai/DeepSeek-V3”. GitHub. https://github.com/deepseek-ai/DeepSeek-V3.DeepSeek. “GitHub - deepseek-ai/DeepSeek-V3”. GitHub, 2025, https://github.com/deepseek-ai/DeepSeek-V3.DeepSeek. GitHub - deepseek-ai/DeepSeek-V3. GitHub https://github.com/deepseek-ai/DeepSeek-V3 (2025).DeepSeek, “GitHub - deepseek-ai/DeepSeek-V3”, GitHub. [Online]. Available: https://github.com/deepseek-ai/DeepSeek-V3
- Douillard et al (2023). AI Safety. Tigera – Creator of Calico.Douillard et al. (2023). AI Safety. Tigera – Creator of Calico. https://tigera.io/learn/guides/llm-security/ai-safetyDouillard et al. 2023. “AI Safety”. Tigera – Creator of Calico. https://tigera.io/learn/guides/llm-security/ai-safety.Douillard et al. “AI Safety”. Tigera – Creator of Calico, 2023, https://tigera.io/learn/guides/llm-security/ai-safety.Douillard et al. AI Safety. Tigera – Creator of Calico https://tigera.io/learn/guides/llm-security/ai-safety (2023).Douillard et al, “AI Safety”, Tigera – Creator of Calico. [Online]. Available: https://tigera.io/learn/guides/llm-security/ai-safety
- Douillard et al (2024). DiPaCo: Distributed Path Composition. arXiv.org.Douillard et al. (2024). DiPaCo: Distributed Path Composition. arXiv.org. https://www.arxiv.org/abs/2403.10616Douillard et al. 2024. “DiPaCo: Distributed Path Composition”. arXiv.org. https://www.arxiv.org/abs/2403.10616.Douillard et al. “DiPaCo: Distributed Path Composition”. arXiv.org, 2024, https://www.arxiv.org/abs/2403.10616.Douillard et al. DiPaCo: Distributed Path Composition. arXiv.org https://www.arxiv.org/abs/2403.10616 (2024).Douillard et al, “DiPaCo: Distributed Path Composition”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2403.10616
- Douillard, A. et al. (2023). DiLoCo: Distributed Low-Communication Training of Language Models. arXiv.Douillard, A., Feng, Q., Rusu, A. A., Chhaparia, R., Donchev, Y., Kuncoro, A., Ranzato, M., Szlam, A., & Shen, J. (2023). DiLoCo: Distributed Low-Communication Training of Language Models. In arXiv. https://arxiv.org/abs/2311.08105Douillard, A., Q. Feng, A. A. Rusu, et al. 2023. “DiLoCo: Distributed Low-Communication Training of Language Models”. In arXiv. Preprint, November 14. https://arxiv.org/abs/2311.08105.Douillard, A., et al. “DiLoCo: Distributed Low-Communication Training of Language Models”. arXiv, 14 Nov. 2023, https://arxiv.org/abs/2311.08105.Douillard, A. et al. DiLoCo: Distributed Low-Communication Training of Language Models. arXiv Preprint at https://arxiv.org/abs/2311.08105 (2023).A. Douillard et al., “DiLoCo: Distributed Low-Communication Training of Language Models”, Nov. 14, 2023. [Online]. Available: https://arxiv.org/abs/2311.08105
- Egan, J. & Heim, L. (2023). Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers. arXiv.Egan, J., & Heim, L. (2023). Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers. In arXiv. https://arxiv.org/abs/2310.13625Egan, J., and L. Heim. 2023. “Oversight for Frontier AI Through a Know-Your-Customer Scheme for Compute Providers”. In arXiv. Preprint, October 20. https://arxiv.org/abs/2310.13625.Egan, J., and L. Heim. “Oversight for Frontier AI Through a Know-Your-Customer Scheme for Compute Providers”. arXiv, 20 Oct. 2023, https://arxiv.org/abs/2310.13625.Egan, J. & Heim, L. Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers. arXiv Preprint at https://arxiv.org/abs/2310.13625 (2023).J. Egan and L. Heim, “Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers”, Oct. 20, 2023. [Online]. Available: https://arxiv.org/abs/2310.13625
- Eiras, F. et al. (2024). Near to Mid-term Risks and Opportunities of Open-Source Generative AI. arXiv.Eiras, F., Petrov, A., Vidgen, B., de Witt, C. S., Pizzati, F., Elkins, K., Mukhopadhyay, S., Bibi, A., Csaba, B., Steibel, F., Barez, F., Smith, G., Guadagni, G., Chun, J., Cabot, J., Imperial, J. M., Nolazco-Flores, J. A., Landay, L., Jackson, M., … Foerster, J. (2024). Near to Mid-term Risks and Opportunities of Open-Source Generative AI. In arXiv. https://arxiv.org/abs/2404.17047Eiras, F., A. Petrov, B. Vidgen, et al. 2024. “Near to Mid-term Risks and Opportunities of Open-Source Generative AI”. In arXiv. Preprint, April 25. https://arxiv.org/abs/2404.17047.Eiras, F., et al. “Near to Mid-term Risks and Opportunities of Open-Source Generative AI”. arXiv, 25 Apr. 2024, https://arxiv.org/abs/2404.17047.Eiras, F. et al. Near to Mid-term Risks and Opportunities of Open-Source Generative AI. arXiv Preprint at https://arxiv.org/abs/2404.17047 (2024).F. Eiras et al., “Near to Mid-term Risks and Opportunities of Open-Source Generative AI”, Apr. 25, 2024. [Online]. Available: https://arxiv.org/abs/2404.17047
- EpochAI (2025). Could decentralized training solve AI’s power problem?. Epoch AI.EpochAI. (2025). Could decentralized training solve AI’s power problem?. Epoch AI. https://epoch.ai/blog/could-decentralized-training-solve-ais-power-problemEpochAI. 2025. “Could Decentralized Training Solve AI’s Power Problem?”. Epoch AI. https://epoch.ai/blog/could-decentralized-training-solve-ais-power-problem.EpochAI. “Could Decentralized Training Solve AI’s Power Problem?”. Epoch AI, 2025, https://epoch.ai/blog/could-decentralized-training-solve-ais-power-problem.EpochAI. Could decentralized training solve AI’s power problem?. Epoch AI https://epoch.ai/blog/could-decentralized-training-solve-ais-power-problem (2025).EpochAI, “Could decentralized training solve AI’s power problem?”, Epoch AI. [Online]. Available: https://epoch.ai/blog/could-decentralized-training-solve-ais-power-problem
- Exfilbench (2025). ExfilBench | Exfiltration & Replication Benchmark. ExfilBench.Exfilbench. (2025). ExfilBench | Exfiltration & Replication Benchmark. Internet Archive (https://web.archive.org/web/20250712035010/https://www.exfilbench.com/). ExfilBench. https://exfilbench.comExfilbench. 2025. “ExfilBench | Exfiltration & Replication Benchmark”. ExfilBench. Https://web.archive.org/web/20250712035010/https://www.exfilbench.com/. Internet Archive. https://exfilbench.com.Exfilbench. “ExfilBench | Exfiltration & Replication Benchmark”. ExfilBench, 2025, Internet Archive, https://web.archive.org/web/20250712035010/https://www.exfilbench.com/, https://exfilbench.com.Exfilbench. ExfilBench | Exfiltration & Replication Benchmark. ExfilBench https://exfilbench.com (2025).Exfilbench, “ExfilBench | Exfiltration & Replication Benchmark”, ExfilBench. Accessed: Jul. 12, 2025. [Online]. Available: https://exfilbench.com
- Hendrycks et al. (2025). AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy.Hendrycks et al. (2025). AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy. https://nationalsecurity.ai/chapter/ai-is-pivotal-for-national-securityHendrycks et al. 2025. “AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy”. https://nationalsecurity.ai/chapter/ai-is-pivotal-for-national-security.Hendrycks et al. AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy. 2025, https://nationalsecurity.ai/chapter/ai-is-pivotal-for-national-security.Hendrycks et al. AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy. https://nationalsecurity.ai/chapter/ai-is-pivotal-for-national-security (2025).Hendrycks et al., “AI Is Pivotal for National Security — Chapter 3 of Superintelligence Strategy”. [Online]. Available: https://nationalsecurity.ai/chapter/ai-is-pivotal-for-national-security
- Jaghouar, S. et al. (2024). INTELLECT-1 Technical Report. arXiv.Jaghouar, S., Ong, J. M., Basra, M., Obeid, F., Straube, J., Keiblinger, M., Bakouch, E., Atkins, L., Panahi, M., Goddard, C., Ryabinin, M., & Hagemann, J. (2024). INTELLECT-1 Technical Report. In arXiv. https://arxiv.org/abs/2412.01152Jaghouar, S., J. M. Ong, M. Basra, et al. 2024. “INTELLECT-1 Technical Report”. In arXiv. Preprint, December 2. https://arxiv.org/abs/2412.01152.Jaghouar, S., et al. “INTELLECT-1 Technical Report”. arXiv, 2 Dec. 2024, https://arxiv.org/abs/2412.01152.Jaghouar, S. et al. INTELLECT-1 Technical Report. arXiv Preprint at https://arxiv.org/abs/2412.01152 (2024).S. Jaghouar et al., “INTELLECT-1 Technical Report”, Dec. 02, 2024. [Online]. Available: https://arxiv.org/abs/2412.01152
- Jaime Sevilla (2025). How far can decentralized training over the internet scale?.Jaime Sevilla. (2025, December 29). How far can decentralized training over the internet scale?. https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scaleJaime Sevilla. 2025. “How Far Can Decentralized Training over the Internet Scale?”. December 29. https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scale.Jaime Sevilla. How Far Can Decentralized Training over the Internet Scale?. 29 Dec. 2025, https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scale.Jaime Sevilla. How far can decentralized training over the internet scale?. https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scale (2025).Jaime Sevilla, “How far can decentralized training over the internet scale?”. [Online]. Available: https://epoch.ai/gradient-updates/how-far-can-decentralized-training-over-the-internet-scale
- Jeffrey Ladish & lennart (2022). Information security considerations for AI and the long term future. LessWrong.Jeffrey Ladish, & lennart. (2022, May 2). Information security considerations for AI and the long term future. LessWrong. https://lesswrong.com/posts/2oAxpRuadyjN2ERhe/information-security-considerations-for-ai-and-the-long-termJeffrey Ladish, and lennart. 2022. “Information Security Considerations for AI and the Long Term Future”. LessWrong, May 2. https://lesswrong.com/posts/2oAxpRuadyjN2ERhe/information-security-considerations-for-ai-and-the-long-term.Jeffrey Ladish, and lennart. “Information Security Considerations for AI and the Long Term Future”. LessWrong, 2 May 2022, https://lesswrong.com/posts/2oAxpRuadyjN2ERhe/information-security-considerations-for-ai-and-the-long-term.Jeffrey Ladish & lennart. Information security considerations for AI and the long term future. LessWrong https://lesswrong.com/posts/2oAxpRuadyjN2ERhe/information-security-considerations-for-ai-and-the-long-term (2022).Jeffrey Ladish and lennart, “Information security considerations for AI and the long term future”, LessWrong. [Online]. Available: https://lesswrong.com/posts/2oAxpRuadyjN2ERhe/information-security-considerations-for-ai-and-the-long-term
- Kapoor, S. et al. (2024). On the Societal Impact of Open Foundation Models. arXiv.Kapoor, S., Bommasani, R., Klyman, K., Longpre, S., Ramaswami, A., Cihon, P., Hopkins, A., Bankston, K., Biderman, S., Bogen, M., Chowdhury, R., Engler, A., Henderson, P., Jernite, Y., Lazar, S., Maffulli, S., Nelson, A., Pineau, J., Skowron, A., … Narayanan, A. (2024). On the Societal Impact of Open Foundation Models. In arXiv. https://arxiv.org/abs/2403.07918Kapoor, S., R. Bommasani, K. Klyman, et al. 2024. “On the Societal Impact of Open Foundation Models”. In arXiv. Preprint, February 27. https://arxiv.org/abs/2403.07918.Kapoor, S., et al. “On the Societal Impact of Open Foundation Models”. arXiv, 27 Feb. 2024, https://arxiv.org/abs/2403.07918.Kapoor, S. et al. On the Societal Impact of Open Foundation Models. arXiv Preprint at https://arxiv.org/abs/2403.07918 (2024).S. Kapoor et al., “On the Societal Impact of Open Foundation Models”, Feb. 27, 2024. [Online]. Available: https://arxiv.org/abs/2403.07918
- Kinniment, M. et al. (2023). Evaluating Language-Model Agents on Realistic Autonomous Tasks. arXiv.Kinniment, M., Sato, L. J. K., Du, H., Goodrich, B., Hasin, M., Chan, L., Miles, L. H., Lin, T. R., Wijk, H., Burget, J., Ho, A., Barnes, E., & Christiano, P. (2023). Evaluating Language-Model Agents on Realistic Autonomous Tasks. In arXiv. https://arxiv.org/abs/2312.11671Kinniment, M., L. J. K. Sato, H. Du, et al. 2023. “Evaluating Language-Model Agents on Realistic Autonomous Tasks”. In arXiv. Preprint, December 18. https://arxiv.org/abs/2312.11671.Kinniment, M., et al. “Evaluating Language-Model Agents on Realistic Autonomous Tasks”. arXiv, 18 Dec. 2023, https://arxiv.org/abs/2312.11671.Kinniment, M. et al. Evaluating Language-Model Agents on Realistic Autonomous Tasks. arXiv Preprint at https://arxiv.org/abs/2312.11671 (2023).M. Kinniment et al., “Evaluating Language-Model Agents on Realistic Autonomous Tasks”, Dec. 18, 2023. [Online]. Available: https://arxiv.org/abs/2312.11671
- Kleinberg, J. & Raghavan, M. (2021). Algorithmic Monoculture and Social Welfare. arXiv.Kleinberg, J., & Raghavan, M. (2021). Algorithmic Monoculture and Social Welfare. In arXiv. https://doi.org/10.1073/pnas.2018340118Kleinberg, J., and M. Raghavan. 2021. “Algorithmic Monoculture and Social Welfare”. In arXiv. Preprint, January 14. https://doi.org/10.1073/pnas.2018340118.Kleinberg, J., and M. Raghavan. “Algorithmic Monoculture and Social Welfare”. arXiv, 14 Jan. 2021, https://doi.org/10.1073/pnas.2018340118.Kleinberg, J. & Raghavan, M. Algorithmic Monoculture and Social Welfare. arXiv Preprint at https://doi.org/10.1073/pnas.2018340118 (2021).J. Kleinberg and M. Raghavan, “Algorithmic Monoculture and Social Welfare”, Jan. 14, 2021. doi: 10.1073/pnas.2018340118.
- Korbak, T., Clymer, J., Hilton, B., Shlegeris, B. & Irving, G. (2025). A sketch of an AI control safety case. arXiv.Korbak, T., Clymer, J., Hilton, B., Shlegeris, B., & Irving, G. (2025). A sketch of an AI control safety case. In arXiv. https://arxiv.org/abs/2501.17315Korbak, T., J. Clymer, B. Hilton, B. Shlegeris, and G. Irving. 2025. “A Sketch of an AI Control Safety Case”. In arXiv. Preprint, January 28. https://arxiv.org/abs/2501.17315.Korbak, T., et al. “A Sketch of an AI Control Safety Case”. arXiv, 28 Jan. 2025, https://arxiv.org/abs/2501.17315.Korbak, T., Clymer, J., Hilton, B., Shlegeris, B. & Irving, G. A sketch of an AI control safety case. arXiv Preprint at https://arxiv.org/abs/2501.17315 (2025).T. Korbak, J. Clymer, B. Hilton, B. Shlegeris, and G. Irving, “A sketch of an AI control safety case”, Jan. 28, 2025. [Online]. Available: https://arxiv.org/abs/2501.17315
- Leike (2023). Self-exfiltration is a key dangerous capability.Leike. (2023). Self-exfiltration is a key dangerous capability. https://aligned.substack.com/p/self-exfiltrationLeike. 2023. Self-exfiltration Is a Key Dangerous Capability. Edition. https://aligned.substack.com/p/self-exfiltration.Leike. Self-exfiltration Is a Key Dangerous Capability. 2023, https://aligned.substack.com/p/self-exfiltration.Leike. Self-exfiltration is a key dangerous capability. https://aligned.substack.com/p/self-exfiltration (2023).Leike, “Self-exfiltration is a key dangerous capability”. [Online]. Available: https://aligned.substack.com/p/self-exfiltration
- Lermen, S., Rogers-Smith, C. & Ladish, J. (2023). LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B. arXiv.Lermen, S., Rogers-Smith, C., & Ladish, J. (2023). LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B. In arXiv. https://arxiv.org/abs/2310.20624Lermen, S., C. Rogers-Smith, and J. Ladish. 2023. “LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B”. In arXiv. Preprint, October 31. https://arxiv.org/abs/2310.20624.Lermen, S., et al. “LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B”. arXiv, 31 Oct. 2023, https://arxiv.org/abs/2310.20624.Lermen, S., Rogers-Smith, C. & Ladish, J. LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B. arXiv Preprint at https://arxiv.org/abs/2310.20624 (2023).S. Lermen, C. Rogers-Smith, and J. Ladish, “LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B”, Oct. 31, 2023. [Online]. Available: https://arxiv.org/abs/2310.20624
- Liang et al. (2022). Stanford CRFM.Liang et al. (2022). Stanford CRFM. https://crfm.stanford.edu/2022/05/17/community-norms.htmlLiang et al. 2022. “Stanford CRFM”. https://crfm.stanford.edu/2022/05/17/community-norms.html.Liang et al. Stanford CRFM. 2022, https://crfm.stanford.edu/2022/05/17/community-norms.html.Liang et al. Stanford CRFM. https://crfm.stanford.edu/2022/05/17/community-norms.html (2022).Liang et al., “Stanford CRFM”. [Online]. Available: https://crfm.stanford.edu/2022/05/17/community-norms.html
- Liu (2024). Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science.Liu. (2024). Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science. Ken Ziyu Liu - Stanford Computer Science. https://ai.stanford.edu/~kzliu/blog/unlearningLiu. 2024. “Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science”. Ken Ziyu Liu - Stanford Computer Science. https://ai.stanford.edu/~kzliu/blog/unlearning.Liu. “Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science”. Ken Ziyu Liu - Stanford Computer Science, 2024, https://ai.stanford.edu/~kzliu/blog/unlearning.Liu. Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science. Ken Ziyu Liu - Stanford Computer Science https://ai.stanford.edu/~kzliu/blog/unlearning (2024).Liu, “Machine Unlearning in 2024 | Ken Ziyu Liu - Stanford Computer Science”, Ken Ziyu Liu - Stanford Computer Science. [Online]. Available: https://ai.stanford.edu/~kzliu/blog/unlearning
- METR (2024). Autonomy Evaluation Resources. METR.METR. (2024, March). Autonomy Evaluation Resources. METR. https://metr.github.io/autonomy-evals-guideMETR. 2024. “Autonomy Evaluation Resources”. METR, March. https://metr.github.io/autonomy-evals-guide.METR. “Autonomy Evaluation Resources”. METR, Mar. 2024, https://metr.github.io/autonomy-evals-guide.METR. Autonomy Evaluation Resources. METR https://metr.github.io/autonomy-evals-guide (2024).METR, “Autonomy Evaluation Resources”, METR. [Online]. Available: https://metr.github.io/autonomy-evals-guide
- Millidge (2025). Open source AI has been vital for alignment.Millidge. (2025). Open source AI has been vital for alignment. https://beren.io/2023-11-05-Open-source-AI-has-been-vital-for-alignmentMillidge. 2025. “Open Source AI Has Been Vital for Alignment”. https://beren.io/2023-11-05-Open-source-AI-has-been-vital-for-alignment.Millidge. Open Source AI Has Been Vital for Alignment. 2025, https://beren.io/2023-11-05-Open-source-AI-has-been-vital-for-alignment.Millidge. Open source AI has been vital for alignment. https://beren.io/2023-11-05-Open-source-AI-has-been-vital-for-alignment (2025).Millidge, “Open source AI has been vital for alignment”. [Online]. Available: https://beren.io/2023-11-05-Open-source-AI-has-been-vital-for-alignment
- Neuralt (2024). Solution / SCAMblock.Neuralt. (2024). Solution / SCAMblock. https://neuralt.com/news-insights/protect-your-subscribers-against-scam-calls-with-ai-powered-scamblockNeuralt. 2024. “Solution / SCAMblock”. https://neuralt.com/news-insights/protect-your-subscribers-against-scam-calls-with-ai-powered-scamblock.Neuralt. Solution / SCAMblock. 2024, https://neuralt.com/news-insights/protect-your-subscribers-against-scam-calls-with-ai-powered-scamblock.Neuralt. Solution / SCAMblock. https://neuralt.com/news-insights/protect-your-subscribers-against-scam-calls-with-ai-powered-scamblock (2024).Neuralt, “Solution / SCAMblock”. [Online]. Available: https://neuralt.com/news-insights/protect-your-subscribers-against-scam-calls-with-ai-powered-scamblock
- Nevo et al. (2024). How AI Labs Can Safeguard Model Weights.Nevo et al. (2024). How AI Labs Can Safeguard Model Weights. Internet Archive (https://web.archive.org/web/20260920224534/https://www.rand.org/pubs/research_reports/RRA2849-1.html). https://rand.org/pubs/research_reports/RRA2849-1.htmlNevo et al. 2024. “How AI Labs Can Safeguard Model Weights”. Https://web.archive.org/web/20260920224534/https://www.rand.org/pubs/research_reports/RRA2849-1.html. Internet Archive. https://rand.org/pubs/research_reports/RRA2849-1.html.Nevo et al. How AI Labs Can Safeguard Model Weights. 2024, Internet Archive, https://web.archive.org/web/20260920224534/https://www.rand.org/pubs/research_reports/RRA2849-1.html, https://rand.org/pubs/research_reports/RRA2849-1.html.Nevo et al. How AI Labs Can Safeguard Model Weights. https://rand.org/pubs/research_reports/RRA2849-1.html (2024).Nevo et al., “How AI Labs Can Safeguard Model Weights”. Accessed: Sep. 20, 2026. [Online]. Available: https://rand.org/pubs/research_reports/RRA2849-1.html
- Piper (2024). Should we make our most powerful AI models open source to all?. Vox.Piper. (2024, February 2). Should we make our most powerful AI models open source to all?. Vox. https://vox.com/future-perfect/2024/2/2/24058484/open-source-artificial-intelligence-ai-risk-meta-llama-2-chatgpt-openai-deepfakePiper. 2024. “Should We Make Our Most Powerful AI Models Open Source to All?”. Vox, February 2. https://vox.com/future-perfect/2024/2/2/24058484/open-source-artificial-intelligence-ai-risk-meta-llama-2-chatgpt-openai-deepfake.Piper. “Should We Make Our Most Powerful AI Models Open Source to All?”. Vox, 2 Feb. 2024, https://vox.com/future-perfect/2024/2/2/24058484/open-source-artificial-intelligence-ai-risk-meta-llama-2-chatgpt-openai-deepfake.Piper. Should we make our most powerful AI models open source to all?. Vox https://vox.com/future-perfect/2024/2/2/24058484/open-source-artificial-intelligence-ai-risk-meta-llama-2-chatgpt-openai-deepfake (2024).Piper, “Should we make our most powerful AI models open source to all?”, Vox. [Online]. Available: https://vox.com/future-perfect/2024/2/2/24058484/open-source-artificial-intelligence-ai-risk-meta-llama-2-chatgpt-openai-deepfake
- Prime Intellect Team et al. (2025). INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning. arXiv.Prime Intellect Team, Sami Jaghouar, Justus Mattern, Jack Min Ong, Jannik Straube, Manveer Basra, Aaron Pazdera, Kushal Thaman, Matthew Di Ferrante, Felix Gabriel, Fares Obeid, Kemal Erdem, Michael Keiblinger, & Johannes Hagemann. (2025). INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning. In arXiv. https://arxiv.org/abs/2505.07291Prime Intellect Team, Sami Jaghouar, Justus Mattern, et al. 2025. “INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning”. In arXiv. Preprint, May 12. https://arxiv.org/abs/2505.07291.Prime Intellect Team, et al. “INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning”. arXiv, 12 May 2025, https://arxiv.org/abs/2505.07291.Prime Intellect Team et al. INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning. arXiv Preprint at https://arxiv.org/abs/2505.07291 (2025).Prime Intellect Team et al., “INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning”, May 12, 2025. [Online]. Available: https://arxiv.org/abs/2505.07291
- Ren, R. et al. (2024). Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. arXiv.Ren, R., Basart, S., Khoja, A., Gatti, A., Phan, L., Yin, X., Mazeika, M., Pan, A., Mukobi, G., Kim, R. H., Fitz, S., & Hendrycks, D. (2024). Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. In arXiv. https://arxiv.org/abs/2407.21792Ren, R., S. Basart, A. Khoja, et al. 2024. “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”. In arXiv. Preprint, July 31. https://arxiv.org/abs/2407.21792.Ren, R., et al. “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”. arXiv, 31 July 2024, https://arxiv.org/abs/2407.21792.Ren, R. et al. Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. arXiv Preprint at https://arxiv.org/abs/2407.21792 (2024).R. Ren et al., “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”, Jul. 31, 2024. [Online]. Available: https://arxiv.org/abs/2407.21792
- Rishub Tamirisa et al. (2024). Tamper-Resistant Safeguards for Open-Weight LLMs. arXiv.Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, & Mantas Mazeika. (2024). Tamper-Resistant Safeguards for Open-Weight LLMs. In arXiv. https://arxiv.org/abs/2408.00761Rishub Tamirisa, Bhrugu Bharathi, Long Phan, et al. 2024. “Tamper-Resistant Safeguards for Open-Weight LLMs”. In arXiv. Preprint, August 1. https://arxiv.org/abs/2408.00761.Rishub Tamirisa, et al. “Tamper-Resistant Safeguards for Open-Weight LLMs”. arXiv, 1 Aug. 2024, https://arxiv.org/abs/2408.00761.Rishub Tamirisa et al. Tamper-Resistant Safeguards for Open-Weight LLMs. arXiv Preprint at https://arxiv.org/abs/2408.00761 (2024).Rishub Tamirisa et al., “Tamper-Resistant Safeguards for Open-Weight LLMs”, Aug. 01, 2024. [Online]. Available: https://arxiv.org/abs/2408.00761
- Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan & Fabien Roger (2023). AI Control: Improving Safety Despite Intentional Subversion. arXiv.Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, & Fabien Roger. (2023). AI Control: Improving Safety Despite Intentional Subversion. In arXiv. https://arxiv.org/abs/2312.06942Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. 2023. “AI Control: Improving Safety Despite Intentional Subversion”. In arXiv. Preprint, December 12. https://arxiv.org/abs/2312.06942.Ryan Greenblatt, et al. “AI Control: Improving Safety Despite Intentional Subversion”. arXiv, 12 Dec. 2023, https://arxiv.org/abs/2312.06942.Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan & Fabien Roger. AI Control: Improving Safety Despite Intentional Subversion. arXiv Preprint at https://arxiv.org/abs/2312.06942 (2023).Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger, “AI Control: Improving Safety Despite Intentional Subversion”, Dec. 12, 2023. [Online]. Available: https://arxiv.org/abs/2312.06942
- scasper (2023). Deep Forgetting & Unlearning for Safely-Scoped LLMs. AI Alignment Forum.scasper. (2023, December 5). Deep Forgetting & Unlearning for Safely-Scoped LLMs. AI Alignment Forum. https://alignmentforum.org/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llmsscasper. 2023. “Deep Forgetting & Unlearning for Safely-Scoped LLMs”. AI Alignment Forum, December 5. https://alignmentforum.org/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llms.scasper. “Deep Forgetting & Unlearning for Safely-Scoped LLMs”. AI Alignment Forum, 5 Dec. 2023, https://alignmentforum.org/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llms.scasper. Deep Forgetting & Unlearning for Safely-Scoped LLMs. AI Alignment Forum https://alignmentforum.org/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llms (2023).scasper, “Deep Forgetting & Unlearning for Safely-Scoped LLMs”, AI Alignment Forum. [Online]. Available: https://alignmentforum.org/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llms
- Seger, E. et al. (2023). Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives. arXiv.Seger, E., Dreksler, N., Moulange, R., Dardaman, E., Schuett, J., Wei, K., Winter, C., Arnold, M., hÉigeartaigh, S. Ó., Korinek, A., Anderljung, M., Bucknall, B., Chan, A., Stafford, E., Koessler, L., Ovadya, A., Garfinkel, B., Bluemke, E., Aird, M., … Gupta, A. (2023). Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives. In arXiv. https://arxiv.org/abs/2311.09227Seger, E., N. Dreksler, R. Moulange, et al. 2023. “Open-Sourcing Highly Capable Foundation Models: An Evaluation of Risks, Benefits, and Alternative Methods for Pursuing Open-source Objectives”. In arXiv. Preprint, September 29. https://arxiv.org/abs/2311.09227.Seger, E., et al. “Open-Sourcing Highly Capable Foundation Models: An Evaluation of Risks, Benefits, and Alternative Methods for Pursuing Open-source Objectives”. arXiv, 29 Sept. 2023, https://arxiv.org/abs/2311.09227.Seger, E. et al. Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives. arXiv Preprint at https://arxiv.org/abs/2311.09227 (2023).E. Seger et al., “Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives”, Sep. 29, 2023. [Online]. Available: https://arxiv.org/abs/2311.09227
- Shevlane, T. & Dafoe, A. (2019). The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?. arXiv.Shevlane, T., & Dafoe, A. (2019). The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?. In arXiv. https://arxiv.org/abs/2001.00463Shevlane, T., and A. Dafoe. 2019. “The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?”. In arXiv. Preprint, December 27. https://arxiv.org/abs/2001.00463.Shevlane, T., and A. Dafoe. “The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?”. arXiv, 27 Dec. 2019, https://arxiv.org/abs/2001.00463.Shevlane, T. & Dafoe, A. The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?. arXiv Preprint at https://arxiv.org/abs/2001.00463 (2019).T. Shevlane and A. Dafoe, “The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?”, Dec. 27, 2019. [Online]. Available: https://arxiv.org/abs/2001.00463
- Shevlane, T. et al. (2023). Model evaluation for extreme risks. arXiv.Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., Ho, L., Siddarth, D., Avin, S., Hawkins, W., Kim, B., Gabriel, I., Bolina, V., Clark, J., Bengio, Y., … Dafoe, A. (2023). Model evaluation for extreme risks. In arXiv. https://arxiv.org/abs/2305.15324Shevlane, T., S. Farquhar, B. Garfinkel, et al. 2023. “Model Evaluation for Extreme Risks”. In arXiv. Preprint, May 24. https://arxiv.org/abs/2305.15324.Shevlane, T., et al. “Model Evaluation for Extreme Risks”. arXiv, 24 May 2023, https://arxiv.org/abs/2305.15324.Shevlane, T. et al. Model evaluation for extreme risks. arXiv Preprint at https://arxiv.org/abs/2305.15324 (2023).T. Shevlane et al., “Model evaluation for extreme risks”, May 24, 2023. [Online]. Available: https://arxiv.org/abs/2305.15324
- Snyder et al. (2020). Measuring Cybersecurity and Cyber Resiliency.Snyder et al. (2020). Measuring Cybersecurity and Cyber Resiliency. Internet Archive (https://web.archive.org/web/20250507024856/https://www.rand.org/pubs/research_reports/RR2703.html). https://rand.org/pubs/research_reports/RR2703.htmlSnyder et al. 2020. “Measuring Cybersecurity and Cyber Resiliency”. Https://web.archive.org/web/20250507024856/https://www.rand.org/pubs/research_reports/RR2703.html. Internet Archive. https://rand.org/pubs/research_reports/RR2703.html.Snyder et al. Measuring Cybersecurity and Cyber Resiliency. 2020, Internet Archive, https://web.archive.org/web/20250507024856/https://www.rand.org/pubs/research_reports/RR2703.html, https://rand.org/pubs/research_reports/RR2703.html.Snyder et al. Measuring Cybersecurity and Cyber Resiliency. https://rand.org/pubs/research_reports/RR2703.html (2020).Snyder et al., “Measuring Cybersecurity and Cyber Resiliency”. Accessed: May 07, 2025. [Online]. Available: https://rand.org/pubs/research_reports/RR2703.html
- Solaiman, I. (2023). The Gradient of Generative AI Release: Methods and Considerations. arXiv.Solaiman, I. (2023). The Gradient of Generative AI Release: Methods and Considerations. In arXiv. https://arxiv.org/abs/2302.04844Solaiman, I. 2023. “The Gradient of Generative AI Release: Methods and Considerations”. In arXiv. Preprint, February 5. https://arxiv.org/abs/2302.04844.Solaiman, I. “The Gradient of Generative AI Release: Methods and Considerations”. arXiv, 5 Feb. 2023, https://arxiv.org/abs/2302.04844.Solaiman, I. The Gradient of Generative AI Release: Methods and Considerations. arXiv Preprint at https://arxiv.org/abs/2302.04844 (2023).I. Solaiman, “The Gradient of Generative AI Release: Methods and Considerations”, Feb. 05, 2023. [Online]. Available: https://arxiv.org/abs/2302.04844
- Solaiman, I. et al. (2019). Release Strategies and the Social Impacts of Language Models. arXiv.Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., Krueger, G., Kim, J. W., Kreps, S., McCain, M., Newhouse, A., Blazakis, J., McGuffie, K., & Wang, J. (2019). Release Strategies and the Social Impacts of Language Models. In arXiv. https://arxiv.org/abs/1908.09203Solaiman, I., M. Brundage, J. Clark, et al. 2019. “Release Strategies and the Social Impacts of Language Models”. In arXiv. Preprint, August 24. https://arxiv.org/abs/1908.09203.Solaiman, I., et al. “Release Strategies and the Social Impacts of Language Models”. arXiv, 24 Aug. 2019, https://arxiv.org/abs/1908.09203.Solaiman, I. et al. Release Strategies and the Social Impacts of Language Models. arXiv Preprint at https://arxiv.org/abs/1908.09203 (2019).I. Solaiman et al., “Release Strategies and the Social Impacts of Language Models”, Aug. 24, 2019. [Online]. Available: https://arxiv.org/abs/1908.09203
- Stanford Institute for Human-Centered Artificial Intelligence (2024). Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights.Stanford Institute for Human-Centered Artificial Intelligence. (2024). Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights (Docket No. 240216-0052). Stanford Institute for Human-Centered Artificial Intelligence. https://hai.stanford.edu/sites/default/files/2024-03/Response-NTIA-RFC-Open-Foundation-Models.pdfStanford Institute for Human-Centered Artificial Intelligence. 2024. Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights. Docket No. 240216-0052. Stanford Institute for Human-Centered Artificial Intelligence. https://hai.stanford.edu/sites/default/files/2024-03/Response-NTIA-RFC-Open-Foundation-Models.pdf.Stanford Institute for Human-Centered Artificial Intelligence. Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights. Docket No. 240216-0052, Stanford Institute for Human-Centered Artificial Intelligence, 27 Mar. 2024, https://hai.stanford.edu/sites/default/files/2024-03/Response-NTIA-RFC-Open-Foundation-Models.pdf.Stanford Institute for Human-Centered Artificial Intelligence. Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights. https://hai.stanford.edu/sites/default/files/2024-03/Response-NTIA-RFC-Open-Foundation-Models.pdf (2024).Stanford Institute for Human-Centered Artificial Intelligence, “Response to NTIA Request for Comment on Dual Use Foundation Artificial Intelligence Models With Widely Available Model Weights”, Stanford Institute for Human-Centered Artificial Intelligence, Docket No. 240216-0052, Mar. 2024. [Online]. Available: https://hai.stanford.edu/sites/default/files/2024-03/Response-NTIA-RFC-Open-Foundation-Models.pdf
- Tom Davidson (2025). Human takeover might be worse than AI takeover. LessWrong.Tom Davidson. (2025, January 10). Human takeover might be worse than AI takeover. LessWrong. https://lesswrong.com/posts/FEcw6JQ8surwxvRfr/human-takeover-might-be-worse-than-ai-takeoverTom Davidson. 2025. “Human Takeover Might Be Worse Than AI Takeover”. LessWrong, January 10. https://lesswrong.com/posts/FEcw6JQ8surwxvRfr/human-takeover-might-be-worse-than-ai-takeover.Tom Davidson. “Human Takeover Might Be Worse Than AI Takeover”. LessWrong, 10 Jan. 2025, https://lesswrong.com/posts/FEcw6JQ8surwxvRfr/human-takeover-might-be-worse-than-ai-takeover.Tom Davidson. Human takeover might be worse than AI takeover. LessWrong https://lesswrong.com/posts/FEcw6JQ8surwxvRfr/human-takeover-might-be-worse-than-ai-takeover (2025).Tom Davidson, “Human takeover might be worse than AI takeover”, LessWrong. [Online]. Available: https://lesswrong.com/posts/FEcw6JQ8surwxvRfr/human-takeover-might-be-worse-than-ai-takeover
- Tom Davidson, Lukas Finnveden & rosehadshar (2025). AI-enabled coups: a small group could use AI to seize power. LessWrong.Tom Davidson, Lukas Finnveden, & rosehadshar. (2025, April 16). AI-enabled coups: a small group could use AI to seize power. LessWrong. https://lesswrong.com/posts/6kBMqrK9bREuGsrnd/ai-enabled-coups-a-small-group-could-use-ai-to-seize-power-1Tom Davidson, Lukas Finnveden, and rosehadshar. 2025. “AI-enabled Coups: A Small Group Could Use AI to Seize Power”. LessWrong, April 16. https://lesswrong.com/posts/6kBMqrK9bREuGsrnd/ai-enabled-coups-a-small-group-could-use-ai-to-seize-power-1.Tom Davidson, et al. “AI-enabled Coups: A Small Group Could Use AI to Seize Power”. LessWrong, 16 Apr. 2025, https://lesswrong.com/posts/6kBMqrK9bREuGsrnd/ai-enabled-coups-a-small-group-could-use-ai-to-seize-power-1.Tom Davidson, Lukas Finnveden & rosehadshar. AI-enabled coups: a small group could use AI to seize power. LessWrong https://lesswrong.com/posts/6kBMqrK9bREuGsrnd/ai-enabled-coups-a-small-group-could-use-ai-to-seize-power-1 (2025).Tom Davidson, Lukas Finnveden, and rosehadshar, “AI-enabled coups: a small group could use AI to seize power”, LessWrong. [Online]. Available: https://lesswrong.com/posts/6kBMqrK9bREuGsrnd/ai-enabled-coups-a-small-group-could-use-ai-to-seize-power-1
- Weidinger, L. et al. (2023). Sociotechnical Safety Evaluation of Generative AI Systems. arXiv.Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., & Isaac, W. (2023). Sociotechnical Safety Evaluation of Generative AI Systems. In arXiv. https://arxiv.org/abs/2310.11986Weidinger, L., M. Rauh, N. Marchal, et al. 2023. “Sociotechnical Safety Evaluation of Generative AI Systems”. In arXiv. Preprint, October 18. https://arxiv.org/abs/2310.11986.Weidinger, L., et al. “Sociotechnical Safety Evaluation of Generative AI Systems”. arXiv, 18 Oct. 2023, https://arxiv.org/abs/2310.11986.Weidinger, L. et al. Sociotechnical Safety Evaluation of Generative AI Systems. arXiv Preprint at https://arxiv.org/abs/2310.11986 (2023).L. Weidinger et al., “Sociotechnical Safety Evaluation of Generative AI Systems”, Oct. 18, 2023. [Online]. Available: https://arxiv.org/abs/2310.11986
Was this section useful?
Thank you for your feedback
Your input helps improve the Atlas.