What role does data play in AI risks? Data fundamentally shapes what AI systems can do and how they behave. For frontier foundation models , training data influences both capabilities and alignment - what systems can do and how they do it. Low quality or harmful training data could lead to misaligned or dangerous models ("garbage in, garbage out"), while carefully curated datasets might help promote safer and more reliable behavior (Longpre et al., 2024; Marcucci et al., 2023).
How well does data meet our governance target criteria? Data as a governance target presents a mixed picture when evaluated against our key criteria. Let's look at each:
- Measurability: While we can measure raw quantities of data, assessing its quality, content, and potential implications is far more difficult. Unlike physical goods like semiconductors, data can be copied, modified, and transmitted in ways that are hard to track. This makes comprehensive measurement of data flows extremely challenging.
- Controllability: Data's non-rival nature means it can be copied and shared widely - once data exists, controlling its spread is very difficult. Even when data appears to be restricted, techniques like model distillation can extract information from trained models (Anderljung et al., 2023). However, there might still be some promising control points, particularly around original data collection and the initial training of foundation models .
- Meaningfulness: Data is particularly meaningful when it comes to AI development. The data used to train models directly shapes their capabilities and behaviors. Changes in training data can significantly impact model performance and safety. This makes data governance potentially powerful, but only if we can overcome the challenges of measurement and control.
What are the key data governance concerns? Several aspects of data require careful governance to promote safe AI development:
- Training data quality and safety is fundamental - low quality or harmful data can create unreliable or dangerous models. For instance, technical data about biological weapons in training sets could enable models to assist in their development (Anderljung et al., 2023).
- Data poisoning and security pose increasingly serious threats. Malicious actors could deliberately manipulate training data to create models that behave dangerously in specific situations while appearing safe during testing. This might involve injecting subtle patterns that only become apparent under certain conditions (Longpre et al., 2024).
- Data provenance and accountability help ensure we can trace where model behaviors come from. Without clear tracking of training data sources and their characteristics, it becomes extremely difficult to diagnose and fix problems when models exhibit concerning behaviors (Longpre et al., 2023).
- Consent and rights frameworks protect both data creators and users. Many current AI training practices operate in legal and ethical grey areas regarding data usage rights. Clear frameworks could help prevent unauthorized use while enabling legitimate innovation (Longpre et al., 2024).
- Bias and representation in training data directly impact model behavior. Skewed or unrepresentative datasets can lead to models that perform poorly or make harmful decisions for certain groups, potentially amplifying societal inequities at a massive scale (Reuel et al., 2024).
- Data access and sharing protocols shape who can develop powerful AI systems. Without governance around data access, we risk either overly concentrated power in a few actors with large datasets, or conversely, uncontrolled proliferation of potentially dangerous capabilities (Heim et al., 2024).
How does data governance fit into overall AI governance? Even with strong governance frameworks, alternative data sources or synthetic data generation could potentially circumvent restrictions. Additionally, many concerning capabilities might emerge from seemingly innocuous training data through unexpected interactions or emergent behaviors. While data governance remains important and worthy of deeper exploration, other governance targets may offer more direct governance over frontier AI development in the near term. This is why in the main text we focused primarily on compute governance, which provides more concrete control points through its physical and concentrated nature.
References
- Anderljung, M. et al. (2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv.Anderljung, M., Barnhart, J., Korinek, A., Leung, J., O'Keefe, C., Whittlestone, J., Avin, S., Brundage, M., Bullock, J., Cass-Beggs, D., Chang, B., Collins, T., Fist, T., Hadfield, G., Hayes, A., Ho, L., Hooker, S., Horvitz, E., Kolt, N., … Wolf, K. (2023). Frontier AI Regulation: Managing Emerging Risks to Public Safety. In arXiv. https://arxiv.org/abs/2307.03718Anderljung, M., J. Barnhart, A. Korinek, et al. 2023. “Frontier AI Regulation: Managing Emerging Risks to Public Safety”. In arXiv. Preprint, July 6. https://arxiv.org/abs/2307.03718.Anderljung, M., et al. “Frontier AI Regulation: Managing Emerging Risks to Public Safety”. arXiv, 6 July 2023, https://arxiv.org/abs/2307.03718.Anderljung, M. et al. Frontier AI Regulation: Managing Emerging Risks to Public Safety. arXiv Preprint at https://arxiv.org/abs/2307.03718 (2023).M. Anderljung et al., “Frontier AI Regulation: Managing Emerging Risks to Public Safety”, Jul. 06, 2023. [Online]. Available: https://arxiv.org/abs/2307.03718
- Heim, L. et al. (2024). Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation. arXiv.Heim, L., Fist, T., Egan, J., Huang, S., Zekany, S., Trager, R., Osborne, M. A., & Zilberman, N. (2024). Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation. In arXiv. https://arxiv.org/abs/2403.08501Heim, L., T. Fist, J. Egan, et al. 2024. “Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation”. In arXiv. Preprint, March 13. https://arxiv.org/abs/2403.08501.Heim, L., et al. “Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation”. arXiv, 13 Mar. 2024, https://arxiv.org/abs/2403.08501.Heim, L. et al. Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation. arXiv Preprint at https://arxiv.org/abs/2403.08501 (2024).L. Heim et al., “Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation”, Mar. 13, 2024. [Online]. Available: https://arxiv.org/abs/2403.08501
- Longpre, S. et al. (2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv.Longpre, S., Mahari, R., Chen, A., Obeng-Marnu, N., Sileo, D., Brannon, W., Muennighoff, N., Khazam, N., Kabbara, J., Perisetla, K., Wu, X., Shippole, E., Bollacker, K., Wu, T., Villa, L., Pentland, S., & Hooker, S. (2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. In arXiv. https://arxiv.org/abs/2310.16787Longpre, S., R. Mahari, A. Chen, et al. 2023. “The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI”. In arXiv. Preprint, October 25. https://arxiv.org/abs/2310.16787.Longpre, S., et al. “The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI”. arXiv, 25 Oct. 2023, https://arxiv.org/abs/2310.16787.Longpre, S. et al. The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv Preprint at https://arxiv.org/abs/2310.16787 (2023).S. Longpre et al., “The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI”, Oct. 25, 2023. [Online]. Available: https://arxiv.org/abs/2310.16787
- Longpre, S. et al. (2024). Consent in Crisis: The Rapid Decline of the AI Data Commons. arXiv.org.Longpre, S., Mahari, R., Lee, A., Lund, C., Oderinwale, H., Brannon, W., Saxena, N., Obeng-Marnu, N., South, T., Hunter, C., Klyman, K., Klamm, C., Schoelkopf, H., Singh, N., Cherep, M., Anis, A., Dinh, A., Chitongo, C., Yin, D., … Pentland, S. (2024). Consent in Crisis: The Rapid Decline of the AI Data Commons. In arXiv.org. https://arxiv.org/abs/2407.14933Longpre, S., R. Mahari, A. Lee, et al. 2024. “Consent in Crisis: The Rapid Decline of the AI Data Commons”. In arXiv.org. Preprint, July 20. https://arxiv.org/abs/2407.14933.Longpre, S., et al. “Consent in Crisis: The Rapid Decline of the AI Data Commons”. arXiv.org, 20 July 2024, https://arxiv.org/abs/2407.14933.Longpre, S. et al. Consent in Crisis: The Rapid Decline of the AI Data Commons. arXiv.org Preprint at https://arxiv.org/abs/2407.14933 (2024).S. Longpre et al., “Consent in Crisis: The Rapid Decline of the AI Data Commons”, Jul. 20, 2024. [Online]. Available: https://arxiv.org/abs/2407.14933
- Marcucci, S., Alarcon, N. G., Verhulst, S. G. & Wullhorst, E. (2023). Mapping and Comparing Data Governance Frameworks: A benchmarking exercise to inform global data governance deliberations. arXiv.Marcucci, S., Alarcon, N. G., Verhulst, S. G., & Wullhorst, E. (2023). Mapping and Comparing Data Governance Frameworks: A benchmarking exercise to inform global data governance deliberations. In arXiv. https://arxiv.org/abs/2302.13731Marcucci, S., N. G. Alarcon, S. G. Verhulst, and E. Wullhorst. 2023. “Mapping and Comparing Data Governance Frameworks: A Benchmarking Exercise to Inform Global Data Governance Deliberations”. In arXiv. Preprint, February 27. https://arxiv.org/abs/2302.13731.Marcucci, S., et al. “Mapping and Comparing Data Governance Frameworks: A Benchmarking Exercise to Inform Global Data Governance Deliberations”. arXiv, 27 Feb. 2023, https://arxiv.org/abs/2302.13731.Marcucci, S., Alarcon, N. G., Verhulst, S. G. & Wullhorst, E. Mapping and Comparing Data Governance Frameworks: A benchmarking exercise to inform global data governance deliberations. arXiv Preprint at https://arxiv.org/abs/2302.13731 (2023).S. Marcucci, N. G. Alarcon, S. G. Verhulst, and E. Wullhorst, “Mapping and Comparing Data Governance Frameworks: A benchmarking exercise to inform global data governance deliberations”, Feb. 27, 2023. [Online]. Available: https://arxiv.org/abs/2302.13731
- Reuel, A. et al. (2024). Open Problems in Technical AI Governance. arXiv.Reuel, A., Bucknall, B., Casper, S., Fist, T., Soder, L., Aarne, O., Hammond, L., Ibrahim, L., Chan, A., Wills, P., Anderljung, M., Garfinkel, B., Heim, L., Trask, A., Mukobi, G., Schaeffer, R., Baker, M., Hooker, S., Solaiman, I., … Trager, R. (2024). Open Problems in Technical AI Governance. In arXiv. https://arxiv.org/abs/2407.14981Reuel, A., B. Bucknall, S. Casper, et al. 2024. “Open Problems in Technical AI Governance”. In arXiv. Preprint, July 20. https://arxiv.org/abs/2407.14981.Reuel, A., et al. “Open Problems in Technical AI Governance”. arXiv, 20 July 2024, https://arxiv.org/abs/2407.14981.Reuel, A. et al. Open Problems in Technical AI Governance. arXiv Preprint at https://arxiv.org/abs/2407.14981 (2024).A. Reuel et al., “Open Problems in Technical AI Governance”, Jul. 20, 2024. [Online]. Available: https://arxiv.org/abs/2407.14981
Was this section useful?
Thank you for your feedback
Your input helps improve the Atlas.