AI models can write, reason, code, generate media, and control robots—often matching or surpassing expert human performance on specific tasks. In this first section, we’ll try to give you a sense of what AI is actually capable of doing as of late 2025. Numbers and graphs don't really make these capabilities tangible, so we will try to include as many videos, images, and examples to really help you get a sense of where AI currently stands. But in case you are interested, we also mention benchmark scores that measure progress quantitatively.
The trajectory matters more than a snapshot. This section can serve both as a quick history and as a snapshot of current capabilities. As you read through it, try to keep in mind how quickly we got here. Language generation went from coherent paragraphs to research assistants in a few years. Image generation went from laughable to professional-grade in a decade. Video generation seems to be following a similar path compressed into three years.
If you’re already familiar with AI or machine learning , some of these stories might be review for you, but hopefully everyone who reads will be able to take away at least one new thing from this section.
Games #
Game-playing AI is already at the superhuman level for many games. Comparing AI and humans at games has been a common theme through the last few decades with AI making continuous progress. IBM’s Deep blue defeated chess grandmaster in 1997 (IBM, 2026), IBM’s Watson won overwhelmingly at Jeopardy! in 2011 (IBM, 2026), and AlphaGo managed to beat the Go world champion Lee Sedol in 2016 (DeepMind, 2016). AlphaGo was a landmark moment, because this is a game with more possible positions than atoms in the observable universe. During game 2, AlphaGo played Move 37, a move that had a 1 in 10,000 chance of being used. Commentators initially thought it was a mistake, instead it was the move that several rounds later led to winning the game. The move demonstrated what many consider to be sparks of genuine creativity that diverged from centuries of human play that the model was trained on.
For the first time in the history of mankind, I saw something similar to an artificial intellect.
AI's superhuman game playing capability extends to video games. Machine learning techniques on simple Atari games in 2013 (Mnih et al., 2013) progressed to OpenAI Five defeating world champions at DOTA 2 in 2019 (OpenAI, 2019). That same year, DeepMind's AlphaStar beat professional esports players at StarCraft II (Google DeepMind, 2019). These are games with open ended real time environments, requiring thousands of rapid decisions and long-term planning. By 2020 the MuZero system played Atari games, Go, chess, and shogi without even being told the rules (Google DeepMind, 2020). 1 These are AIs that all play and learn autonomously without human intervention. In 2025, game playing AIs have evolved to open-ended environments across a huge variety of games. They carry the abilities learned from one game onto the next improving their performance over time (Google DeepMind, 2025).
Game playing AI is relatively narrow in what it can do. Despite this it is extremely impressive because of the strategic planning, pattern recognition, and adversarial thinking it displays. These same reasoning abilities—planning ahead, building strategies, adapting to feedback—that started with game playing now also apply to scientific research, mathematical proofs, and complex real-world problem-solving.
Text Generation #
Generating text can take language models far beyond simple conversations. You're probably familiar with ChatGPT. These types of language generation AIs are what we call large language models (LLMs). In 2018, early versions of these LMs could only write a few coherent paragraphs. But over a few short years, they have gotten a lot better. By 2025, GPT-5 gets 92.5% of questions right in domains ranging from highly complex STEM fields, international law, to nutrition and religion (as measured by the MMLU benchmark). Models like Claude by Anthropic, Grok by X, GLM-4.7 by Zhipu AI, and DeepSeek-v3.2 showcased similar levels of performance across various domains (ArtificialAnalysis, 2025; EpochAI, 2025).
Language models have provided a core around which we have seen many impressive capabilities emerge like - scientific research, reasoning, and software development. All of these capabilities stem from the same principle of generating language and gradually refining it.
Tool Use #
LLMs can intelligently use external tools, dramatically boosting performance. Language models in 2020 exhibited remarkable abilities to solve new tasks from instructions, but they used to struggle with basic functions like arithmetic. Instead of trying to get a single model to do everything, increasingly LLMs use external tools to achieve both capabilities (Schick et al., 2023; Qin et al., 2023). They recognize when they need a calculator, code interpreter, or up to date information from a search engine—and call these tools appropriately. 2 Tool use significantly improves model performance; for example, the OpenAI o3 model with external tools outperforms o3 alone by almost 5% on benchmarks like Humanities Last Exam (HLE) (EpochAI, 2025). In December 2025, at least 10,000 tool servers are operational, including meta tools like ‘a tool to search for tools’ to help LLMs find the exact one they need for the specific situation (Anthropic, 2025; Anthropic, 2025). This leads to a significant enough boost that companies now report benchmark performance separately - with and without tools.
Reasoning and Research #
Models now demonstrate multi-step reasoning by working through problems step-by-step. In addition to using tools, LLMs now show their reasoning, catch their own errors, and backtrack when needed. In late 2024, OpenAI introduced o1, the first "reasoning" model. These AIs allocate more effort per problem—trading "thinking time" for accuracy. The longer they think, the better their responses tend to get (OpenAI, 2024). Using these reasoning techniques, both OpenAI and Google DeepMind achieved gold-medal performance at the 2025 International Mathematical Olympiad (OpenAI, 2025; Google DeepMind, 2025). On FrontierMath—a test of research-level mathematics—GPT-5.2 solved 41% of tier 1-3 problems (EpochAI, 2026).1 In ML terms, ‘thinking time’ is called ‘inference’. It is the act of generating new tokens/words, so increasing thinking time is more formally called inference time scaling. Similar to tool use this also led to such a boost in performance that companies report high thinking time (high compute) results of benchmarks separately than no thinking time (low compute) results.
LLMs can help generate and evaluate scientific hypotheses. Combining techniques like letting AI think for longer, and tools like web-search or specialized AI models, we are starting to see research assistants. As one example, Google introduced AI co-scientist in 2025. The team used it to generate and evaluate proposals for repurposing drugs, identifying drug targets, and explaining antimicrobial resistance in real-world laboratories (Google DeepMind, 2025). Others have attempted to build a fully autonomous AI scientist, which generates novel research ideas, writes code, executes experiments, visualizes results, describes its findings by writing a full scientific paper, and then runs a simulated review process for evaluation (SakanaAI, 2024).
AI is transitioning to an active research collaborator across scientific domains. Instead of only using language models, companies are also using approaches similar to AlphaZero to create a whole range of specialized models for scientific domains. For example, Demis Hassabis & John Jumper were awarded the nobel prize in chemistry for their work on building AlphaFold (Google DeepMind, 2024). This is a model that helped solve the long outstanding protein folding problem (Google DeepMind, 2022), and its successors AlphaFold 2 and 3 continue to aid thousands of researchers in biology (DeepMind, 2024). Similarly, AlphaGenome is helping us better understand human DNA (Google DeepMind, 2025). AlphaEvolve is helping generate faster algorithms for machine learning (Google DeepMind, 2025), and AlphaChip helps design the semiconductors that run algorithms (Google DeepMind, 2024).
Beyond just mathematics and scientific research, AI models are also developing more abstract intellectual skills. LLMs have some level of metacognition, they can evaluate the validity of their own claims and predict which questions they will be able to answer correctly (Kadavath, 2022). They have some knowledge about their own selves and their limitations. Similarly, they display the ability to attribute mental states to themselves and others (theory of mind). This helps in predicting human behaviors and responses (Kosinski 2023; Xu et al., 2024). We are going to talk more about how we concretely define and measure things like intelligence, meta-cognition and so on in later sections.
Software Development #
Coding is evolving from autocomplete to collaborative software development. LLMs can generate text in any form, and one particular type of text that they are proving to be especially good at is generating code. When paired with reasoning capabilities, and tools LLMs read documentation, edit codebases spanning thousands of files, run tests, debug failures, and iterate until tests pass—with increasingly minimal human guidance. In 2025 systems like Claude Opus 4.5, Gemini 3 Pro, and GPT-5.2 implement features and entire applications increasingly independently (Anthropic, 2025; Google DeepMind, 2025; OpenAI, 2025). When tested against real GitHub issues from open-source projects— in 2024 AI systems (Claude 3 Opus) could solve just 15% problems, but by 2025, this had jumped to being able to solve 74% of issues (Tools + Claude 4 Opus) (SWE bench, 2025).
Vision: Images and Video #
Image generation progressed from unrecognizable noise to photorealistic scenes in under a decade. In 2014, Generative Adversarial Networks (GANs) produced grainy, low-resolution faces (Goodfellow et al., 2014). By 2023, models generated detailed images from complex text prompts. Models like Midjourney v7 create photorealistic scenes nearly indistinguishable from professional photography. Video generation is following a similar trajectory. AI generated videos and DeepFakes are getting increasingly indistinguishable from real videos.
Large Multimodal Models (LMMs) combine language and image understanding capabilities. In 2025, multimodal models answer questions about spatial relationships in images, read embedded text in complex scenes, and extract information from charts and diagrams. Image models released in 2025 handle generation and sophisticated editing across modalities. Systems like Gemini 3 Pro Image work text-to-image, image-to-image, and handle complex editing—changing lighting, style, or composition while maintaining coherence (Google, 2025).
Robotics #
Both AI and robotics are evolving, with robots giving AI physical embodiment in the real world. Robotics is combining LLMs and visual models, to create robot control models (e.g. RT-1 or RT-2). These robots can use techniques borrowed from language models—like breaking complex actions into step-by-step plans—to control robot manipulators (Google DeepMind, 2024). They managed to learn behaviors like opening cabinets, operating elevators, and cooking tasks through observing human demonstrations. Robots have demonstrated the ability to perform intricate manipulation: sautéing shrimp, storing heavy pots in cabinets, and rinsing pans (Fu et al., 2024).
Autonomous robots are moving from research labs into real-world industrial deployment at significant scale. In 2023, China installed 276,300 industrial robots (AI Index Report, 2025). These systems handle welding, parts assembly, materials handling, and quality inspection—tasks requiring precision but not necessarily advanced reasoning. In addition to industrial robots, warehouse robotics represents one of the most mature deployments—Amazon operates over 1 million robots across its fulfillment network, handling everything from inventory storage to package sorting (Amazon, 2025). Robots in warehouses and industry are able to handle packages, speed up inventory identification using machine vision, and autonomously unload shipping containers.
Footnotes
KataGo is a system which is based on techniques used by DeepMind's AlphaGo Zero and similarly superhuman in its game play. In 2022, researchers managed to demonstrate that despite being superhuman KataGo can be beaten by humans and demonstrates "surprising failure modes" of AI systems. This is the kind of thing that will be a repeated theme throughout our text (Wang and Gleave et al., 2022)
↩Standards like the Model Context Protocol (MCP) are formalizing how AI assistants connect to data repositories and development environments (Anthropic, 2024). The MCP protocol has been donated to the Linux foundation(Anthropic, 2025).
↩
References
- AI Digest (2023). How fast is AI improving?. AI Digest.AI Digest. (2023). How fast is AI improving?. AI Digest. https://theaidigest.org/progress-and-dangersAI Digest. 2023. “How Fast Is AI Improving?”. AI Digest. https://theaidigest.org/progress-and-dangers.AI Digest. “How Fast Is AI Improving?”. AI Digest, 2023, https://theaidigest.org/progress-and-dangers.AI Digest. How fast is AI improving?. AI Digest https://theaidigest.org/progress-and-dangers (2023).AI Digest, “How fast is AI improving?”, AI Digest. [Online]. Available: https://theaidigest.org/progress-and-dangers
- Amazon (2023). 4 cool facts about Hercules, the small-but-mighty robot in Amazon’s fulfillment centers. Amazon News.Amazon. (2023, September 7). 4 cool facts about Hercules, the small-but-mighty robot in Amazon’s fulfillment centers. Amazon News. https://aboutamazon.com/news/operations/amazon-hercules-robotAmazon. 2023. “4 Cool Facts About Hercules, the Small-but-mighty Robot in Amazon’s Fulfillment Centers”. Amazon News, September 7. https://aboutamazon.com/news/operations/amazon-hercules-robot.Amazon. “4 Cool Facts About Hercules, the Small-but-mighty Robot in Amazon’s Fulfillment Centers”. Amazon News, 7 Sept. 2023, https://aboutamazon.com/news/operations/amazon-hercules-robot.Amazon. 4 cool facts about Hercules, the small-but-mighty robot in Amazon’s fulfillment centers. Amazon News https://aboutamazon.com/news/operations/amazon-hercules-robot (2023).Amazon, “4 cool facts about Hercules, the small-but-mighty robot in Amazon’s fulfillment centers”, Amazon News. [Online]. Available: https://aboutamazon.com/news/operations/amazon-hercules-robot
- Amazon (2024). Amazon uses robots that sort, lift, and carry packages—see them in action. Amazon News.Amazon. (2024, October 9). Amazon uses robots that sort, lift, and carry packages—see them in action. Amazon News. https://aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-centerAmazon. 2024. “Amazon Uses Robots That Sort, Lift, and Carry Packages—see Them in Action”. Amazon News, October 9. https://aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-center.Amazon. “Amazon Uses Robots That Sort, Lift, and Carry Packages—see Them in Action”. Amazon News, 9 Oct. 2024, https://aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-center.Amazon. Amazon uses robots that sort, lift, and carry packages—see them in action. Amazon News https://aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-center (2024).Amazon, “Amazon uses robots that sort, lift, and carry packages—see them in action”, Amazon News. [Online]. Available: https://aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-center
- Anthropic (2024). Introducing the Model Context Protocol.Anthropic. (2024). Introducing the Model Context Protocol. https://anthropic.com/news/model-context-protocolAnthropic. 2024. “Introducing the Model Context Protocol”. https://anthropic.com/news/model-context-protocol.Anthropic. Introducing the Model Context Protocol. 2024, https://anthropic.com/news/model-context-protocol.Anthropic. Introducing the Model Context Protocol. https://anthropic.com/news/model-context-protocol (2024).Anthropic, “Introducing the Model Context Protocol”. [Online]. Available: https://anthropic.com/news/model-context-protocol
- Anthropic (2025). Claude Opus.Anthropic. (2025). Claude Opus. https://anthropic.com/claude/opusAnthropic. 2025. “Claude Opus”. https://anthropic.com/claude/opus.Anthropic. Claude Opus. 2025, https://anthropic.com/claude/opus.Anthropic. Claude Opus. https://anthropic.com/claude/opus (2025).Anthropic, “Claude Opus”. [Online]. Available: https://anthropic.com/claude/opus
- Anthropic (2025). Donating MCP to the Agentic AI Foundation.Anthropic. (2025). Donating MCP to the Agentic AI Foundation. https://anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundationAnthropic. 2025. “Donating MCP to the Agentic AI Foundation”. https://anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation.Anthropic. Donating MCP to the Agentic AI Foundation. 2025, https://anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation.Anthropic. Donating MCP to the Agentic AI Foundation. https://anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation (2025).Anthropic, “Donating MCP to the Agentic AI Foundation”. [Online]. Available: https://anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation
- Anthropic (2025). Introducing advanced tool use on the Claude Developer Platform.Anthropic. (2025). Introducing advanced tool use on the Claude Developer Platform. https://anthropic.com/engineering/advanced-tool-useAnthropic. 2025. “Introducing Advanced Tool Use on the Claude Developer Platform”. https://anthropic.com/engineering/advanced-tool-use.Anthropic. Introducing Advanced Tool Use on the Claude Developer Platform. 2025, https://anthropic.com/engineering/advanced-tool-use.Anthropic. Introducing advanced tool use on the Claude Developer Platform. https://anthropic.com/engineering/advanced-tool-use (2025).Anthropic, “Introducing advanced tool use on the Claude Developer Platform”. [Online]. Available: https://anthropic.com/engineering/advanced-tool-use
- ARC-AGI (2024). ARC-AGI-1. ARC Prize.ARC-AGI. (2024). ARC-AGI-1. ARC Prize. https://arcprize.org/arc-agi/1ARC-AGI. 2024. “ARC-AGI-1”. ARC Prize. https://arcprize.org/arc-agi/1.ARC-AGI. “ARC-AGI-1”. ARC Prize, 2024, https://arcprize.org/arc-agi/1.ARC-AGI. ARC-AGI-1. ARC Prize https://arcprize.org/arc-agi/1 (2024).ARC-AGI, “ARC-AGI-1”, ARC Prize. [Online]. Available: https://arcprize.org/arc-agi/1
- ArtificialAnalysis (2025). Artificial Analysis Intelligence Index v4.3.2. Artificial Analysis.ArtificialAnalysis. (2025). Artificial Analysis Intelligence Index v4.3.2. Artificial Analysis. https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-indexArtificialAnalysis. 2025. “Artificial Analysis Intelligence Index V4.3.2”. Artificial Analysis. https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index.ArtificialAnalysis. “Artificial Analysis Intelligence Index V4.3.2”. Artificial Analysis, 2025, https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index.ArtificialAnalysis. Artificial Analysis Intelligence Index v4.3.2. Artificial Analysis https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index (2025).ArtificialAnalysis, “Artificial Analysis Intelligence Index v4.3.2”, Artificial Analysis. [Online]. Available: https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index
- Aschenbrenner (2024). I. From GPT-4 to AGI: Counting the OOMs. SITUATIONAL AWARENESS - The Decade Ahead.Aschenbrenner. (2024, May 29). I. From GPT-4 to AGI: Counting the OOMs. SITUATIONAL AWARENESS - The Decade Ahead. https://situational-awareness.ai/from-gpt-4-to-agiAschenbrenner. 2024. “I. From GPT-4 to AGI: Counting the OOMs”. SITUATIONAL AWARENESS - The Decade Ahead, May 29. https://situational-awareness.ai/from-gpt-4-to-agi.Aschenbrenner. “I. From GPT-4 to AGI: Counting the OOMs”. SITUATIONAL AWARENESS - The Decade Ahead, 29 May 2024, https://situational-awareness.ai/from-gpt-4-to-agi.Aschenbrenner. I. From GPT-4 to AGI: Counting the OOMs. SITUATIONAL AWARENESS - The Decade Ahead https://situational-awareness.ai/from-gpt-4-to-agi (2024).Aschenbrenner, “I. From GPT-4 to AGI: Counting the OOMs”, SITUATIONAL AWARENESS - The Decade Ahead. [Online]. Available: https://situational-awareness.ai/from-gpt-4-to-agi
- Bakhtin, A. et al. (2022). Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning. arXiv.Bakhtin, A., Wu, D. J., Lerer, A., Gray, J., Jacob, A. P., Farina, G., Miller, A. H., & Brown, N. (2022). Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning. In arXiv. https://arxiv.org/abs/2210.05492Bakhtin, A., D. J. Wu, A. Lerer, et al. 2022. “Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning”. In arXiv. Preprint, October 11. https://arxiv.org/abs/2210.05492.Bakhtin, A., et al. “Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning”. arXiv, 11 Oct. 2022, https://arxiv.org/abs/2210.05492.Bakhtin, A. et al. Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning. arXiv Preprint at https://arxiv.org/abs/2210.05492 (2022).A. Bakhtin et al., “Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning”, Oct. 11, 2022. [Online]. Available: https://arxiv.org/abs/2210.05492
- Boston Dynamics (2024). Atlas Goes Hands On. Boston Dynamics.Boston Dynamics. (2024). Atlas Goes Hands On. Boston Dynamics. https://bostondynamics.com/video/atlas-goes-hands-onBoston Dynamics. 2024. “Atlas Goes Hands On”. Boston Dynamics. https://bostondynamics.com/video/atlas-goes-hands-on.Boston Dynamics. “Atlas Goes Hands On”. Boston Dynamics, 2024, https://bostondynamics.com/video/atlas-goes-hands-on.Boston Dynamics. Atlas Goes Hands On. Boston Dynamics https://bostondynamics.com/video/atlas-goes-hands-on (2024).Boston Dynamics, “Atlas Goes Hands On”, Boston Dynamics. [Online]. Available: https://bostondynamics.com/video/atlas-goes-hands-on
- Boston Dynamics (2024). Stretch - Mobile Warehouse Robots. Boston Dynamics.Boston Dynamics. (2024). Stretch - Mobile Warehouse Robots. Boston Dynamics. https://bostondynamics.com/products/stretchBoston Dynamics. 2024. “Stretch - Mobile Warehouse Robots”. Boston Dynamics. https://bostondynamics.com/products/stretch.Boston Dynamics. “Stretch - Mobile Warehouse Robots”. Boston Dynamics, 2024, https://bostondynamics.com/products/stretch.Boston Dynamics. Stretch - Mobile Warehouse Robots. Boston Dynamics https://bostondynamics.com/products/stretch (2024).Boston Dynamics, “Stretch - Mobile Warehouse Robots”, Boston Dynamics. [Online]. Available: https://bostondynamics.com/products/stretch
- Bubeck, S. et al. (2023). Sparks of Artificial General Intelligence: Early experiments with GPT-4. arXiv.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., & Zhang, Y. (2023). Sparks of Artificial General Intelligence: Early experiments with GPT-4. In arXiv. https://arxiv.org/abs/2303.12712Bubeck, S., V. Chandrasekaran, R. Eldan, et al. 2023. “Sparks of Artificial General Intelligence: Early Experiments with GPT-4”. In arXiv. Preprint, March 22. https://arxiv.org/abs/2303.12712.Bubeck, S., et al. “Sparks of Artificial General Intelligence: Early Experiments with GPT-4”. arXiv, 22 Mar. 2023, https://arxiv.org/abs/2303.12712.Bubeck, S. et al. Sparks of Artificial General Intelligence: Early experiments with GPT-4. arXiv Preprint at https://arxiv.org/abs/2303.12712 (2023).S. Bubeck et al., “Sparks of Artificial General Intelligence: Early experiments with GPT-4”, Mar. 22, 2023. [Online]. Available: https://arxiv.org/abs/2303.12712
- Chollet, F., Knoop, M., Kamradt, G. & Landers, B. (2024). ARC Prize 2024: Technical Report. arXiv.Chollet, F., Knoop, M., Kamradt, G., & Landers, B. (2024). ARC Prize 2024: Technical Report. In arXiv. https://arxiv.org/abs/2412.04604Chollet, F., M. Knoop, G. Kamradt, and B. Landers. 2024. “ARC Prize 2024: Technical Report”. In arXiv. Preprint, December 5. https://arxiv.org/abs/2412.04604.Chollet, F., et al. “ARC Prize 2024: Technical Report”. arXiv, 5 Dec. 2024, https://arxiv.org/abs/2412.04604.Chollet, F., Knoop, M., Kamradt, G. & Landers, B. ARC Prize 2024: Technical Report. arXiv Preprint at https://arxiv.org/abs/2412.04604 (2024).F. Chollet, M. Knoop, G. Kamradt, and B. Landers, “ARC Prize 2024: Technical Report”, Dec. 05, 2024. [Online]. Available: https://arxiv.org/abs/2412.04604
- DeepMind (2016). AlphaGo. Google DeepMind.DeepMind. (2016). AlphaGo. Google DeepMind. https://deepmind.com/research/highlighted-research/alphagoDeepMind. 2016. “AlphaGo”. Google DeepMind. https://deepmind.com/research/highlighted-research/alphago.DeepMind. “AlphaGo”. Google DeepMind, 2016, https://deepmind.com/research/highlighted-research/alphago.DeepMind. AlphaGo. Google DeepMind https://deepmind.com/research/highlighted-research/alphago (2016).DeepMind, “AlphaGo”, Google DeepMind. [Online]. Available: https://deepmind.com/research/highlighted-research/alphago
- DeepMind (2024). How AlphaChip transformed computer chip design. Google DeepMind.DeepMind. (2024). How AlphaChip transformed computer chip design. Google DeepMind. https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-designDeepMind. 2024. “How AlphaChip Transformed Computer Chip Design”. Google DeepMind. https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-design.DeepMind. “How AlphaChip Transformed Computer Chip Design”. Google DeepMind, 2024, https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-design.DeepMind. How AlphaChip transformed computer chip design. Google DeepMind https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-design (2024).DeepMind, “How AlphaChip transformed computer chip design”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-design
- DeepMind (2025). AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. Google DeepMind.DeepMind. (2025). AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. Google DeepMind. https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithmsDeepMind. 2025. “AlphaEvolve: A Gemini-powered Coding Agent for Designing Advanced Algorithms”. Google DeepMind. https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms.DeepMind. “AlphaEvolve: A Gemini-powered Coding Agent for Designing Advanced Algorithms”. Google DeepMind, 2025, https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms.DeepMind. AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. Google DeepMind https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms (2025).DeepMind, “AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms
- EpochAI (2024). FrontierMath: LLM Benchmark for Advanced AI Math Reasoning. Epoch AI.EpochAI. (2024). FrontierMath: LLM Benchmark for Advanced AI Math Reasoning. Epoch AI. https://epoch.ai/frontiermathEpochAI. 2024. “FrontierMath: LLM Benchmark for Advanced AI Math Reasoning”. Epoch AI. https://epoch.ai/frontiermath.EpochAI. “FrontierMath: LLM Benchmark for Advanced AI Math Reasoning”. Epoch AI, 2024, https://epoch.ai/frontiermath.EpochAI. FrontierMath: LLM Benchmark for Advanced AI Math Reasoning. Epoch AI https://epoch.ai/frontiermath (2024).EpochAI, “FrontierMath: LLM Benchmark for Advanced AI Math Reasoning”, Epoch AI. [Online]. Available: https://epoch.ai/frontiermath
- EpochAI (2025). Epoch Capabilities Index. Epoch AI.EpochAI. (2025). Epoch Capabilities Index. Epoch AI. https://epoch.ai/benchmarks/eciEpochAI. 2025. “Epoch Capabilities Index”. Epoch AI. https://epoch.ai/benchmarks/eci.EpochAI. “Epoch Capabilities Index”. Epoch AI, 2025, https://epoch.ai/benchmarks/eci.EpochAI. Epoch Capabilities Index. Epoch AI https://epoch.ai/benchmarks/eci (2025).EpochAI, “Epoch Capabilities Index”, Epoch AI. [Online]. Available: https://epoch.ai/benchmarks/eci
- EpochAI (2025). FrontierMath Sample Problems. Epoch AI.EpochAI. (2025). FrontierMath Sample Problems. Epoch AI. https://epoch.ai/frontiermath/benchmark-problemsEpochAI. 2025. “FrontierMath Sample Problems”. Epoch AI. https://epoch.ai/frontiermath/benchmark-problems.EpochAI. “FrontierMath Sample Problems”. Epoch AI, 2025, https://epoch.ai/frontiermath/benchmark-problems.EpochAI. FrontierMath Sample Problems. Epoch AI https://epoch.ai/frontiermath/benchmark-problems (2025).EpochAI, “FrontierMath Sample Problems”, Epoch AI. [Online]. Available: https://epoch.ai/frontiermath/benchmark-problems
- Fu, Z., Zhao, T. Z. & Finn, C. (2024). Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. arXiv.Fu, Z., Zhao, T. Z., & Finn, C. (2024). Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. In arXiv. https://arxiv.org/abs/2401.02117Fu, Z., T. Z. Zhao, and C. Finn. 2024. “Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation”. In arXiv. Preprint, January 4. https://arxiv.org/abs/2401.02117.Fu, Z., et al. “Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation”. arXiv, 4 Jan. 2024, https://arxiv.org/abs/2401.02117.Fu, Z., Zhao, T. Z. & Finn, C. Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. arXiv Preprint at https://arxiv.org/abs/2401.02117 (2024).Z. Fu, T. Z. Zhao, and C. Finn, “Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation”, Jan. 04, 2024. [Online]. Available: https://arxiv.org/abs/2401.02117
- Goodfellow, I. J. et al. (2014). Generative Adversarial Networks. arXiv.Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative Adversarial Networks. In arXiv. https://arxiv.org/abs/1406.2661Goodfellow, I. J., J. Pouget-Abadie, M. Mirza, et al. 2014. “Generative Adversarial Networks”. In arXiv. Preprint, June 10. https://arxiv.org/abs/1406.2661.Goodfellow, I. J., et al. “Generative Adversarial Networks”. arXiv, 10 June 2014, https://arxiv.org/abs/1406.2661.Goodfellow, I. J. et al. Generative Adversarial Networks. arXiv Preprint at https://arxiv.org/abs/1406.2661 (2014).I. J. Goodfellow et al., “Generative Adversarial Networks”, Jun. 10, 2014. [Online]. Available: https://arxiv.org/abs/1406.2661
- Google (2025). Introducing Nano Banana Pro. Google.Google. (2025, November 20). Introducing Nano Banana Pro. Google. https://blog.google/innovation-and-ai/products/nano-banana-proGoogle. 2025. “Introducing Nano Banana Pro”. Google, November 20. https://blog.google/innovation-and-ai/products/nano-banana-pro.Google. “Introducing Nano Banana Pro”. Google, 20 Nov. 2025, https://blog.google/innovation-and-ai/products/nano-banana-pro.Google. Introducing Nano Banana Pro. Google https://blog.google/innovation-and-ai/products/nano-banana-pro (2025).Google, “Introducing Nano Banana Pro”, Google. [Online]. Available: https://blog.google/innovation-and-ai/products/nano-banana-pro
- Google DeepMind (2019). AlphaStar: Mastering the real-time strategy game StarCraft II. Google DeepMind.Google DeepMind. (2019). AlphaStar: Mastering the real-time strategy game StarCraft II. Google DeepMind. https://deepmind.google/discover/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-iiGoogle DeepMind. 2019. “AlphaStar: Mastering the Real-time Strategy Game StarCraft II”. Google DeepMind. https://deepmind.google/discover/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-ii.Google DeepMind. “AlphaStar: Mastering the Real-time Strategy Game StarCraft II”. Google DeepMind, 2019, https://deepmind.google/discover/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-ii.Google DeepMind. AlphaStar: Mastering the real-time strategy game StarCraft II. Google DeepMind https://deepmind.google/discover/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-ii (2019).Google DeepMind, “AlphaStar: Mastering the real-time strategy game StarCraft II”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-ii
- Google DeepMind (2020). AlphaFold: a solution to a 50-year-old grand challenge in biology.Google DeepMind. (2020, November 30). AlphaFold: a solution to a 50-year-old grand challenge in biology. https://deepmind.google/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biologyGoogle DeepMind. 2020. “AlphaFold: A Solution to a 50-year-old Grand Challenge in Biology”. November 30. https://deepmind.google/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology.Google DeepMind. AlphaFold: A Solution to a 50-year-old Grand Challenge in Biology. 30 Nov. 2020, https://deepmind.google/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology.Google DeepMind. AlphaFold: a solution to a 50-year-old grand challenge in biology. https://deepmind.google/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology (2020).Google DeepMind, “AlphaFold: a solution to a 50-year-old grand challenge in biology”. [Online]. Available: https://deepmind.google/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology
- Google DeepMind (2024). AlphaFold.Google DeepMind. (2024). AlphaFold. https://deepmind.google/science/alphafoldGoogle DeepMind. 2024. “AlphaFold”. https://deepmind.google/science/alphafold.Google DeepMind. AlphaFold. 2024, https://deepmind.google/science/alphafold.Google DeepMind. AlphaFold. https://deepmind.google/science/alphafold (2024).Google DeepMind, “AlphaFold”. [Online]. Available: https://deepmind.google/science/alphafold
- Google DeepMind (2024). Demis Hassabis & John Jumper awarded Nobel Prize in Chemistry.Google DeepMind. (2024, October 9). Demis Hassabis & John Jumper awarded Nobel Prize in Chemistry. https://deepmind.google/blog/demis-hassabis-john-jumper-awarded-nobel-prize-in-chemistryGoogle DeepMind. 2024. “Demis Hassabis & John Jumper Awarded Nobel Prize in Chemistry”. October 9. https://deepmind.google/blog/demis-hassabis-john-jumper-awarded-nobel-prize-in-chemistry.Google DeepMind. Demis Hassabis & John Jumper Awarded Nobel Prize in Chemistry. 9 Oct. 2024, https://deepmind.google/blog/demis-hassabis-john-jumper-awarded-nobel-prize-in-chemistry.Google DeepMind. Demis Hassabis & John Jumper awarded Nobel Prize in Chemistry. https://deepmind.google/blog/demis-hassabis-john-jumper-awarded-nobel-prize-in-chemistry (2024).Google DeepMind, “Demis Hassabis & John Jumper awarded Nobel Prize in Chemistry”. [Online]. Available: https://deepmind.google/blog/demis-hassabis-john-jumper-awarded-nobel-prize-in-chemistry
- Google DeepMind (2025). Gemini.Google DeepMind. (2025). Gemini. https://deepmind.google/models/geminiGoogle DeepMind. 2025. “Gemini”. https://deepmind.google/models/gemini.Google DeepMind. Gemini. 2025, https://deepmind.google/models/gemini.Google DeepMind. Gemini. https://deepmind.google/models/gemini (2025).Google DeepMind, “Gemini”. [Online]. Available: https://deepmind.google/models/gemini
- Google DeepMind (2025). Gemini 3.1 Pro.Google DeepMind. (2025). Gemini 3.1 Pro. https://deepmind.google/models/gemini/proGoogle DeepMind. 2025. “Gemini 3.1 Pro”. https://deepmind.google/models/gemini/pro.Google DeepMind. Gemini 3.1 Pro. 2025, https://deepmind.google/models/gemini/pro.Google DeepMind. Gemini 3.1 Pro. https://deepmind.google/models/gemini/pro (2025).Google DeepMind, “Gemini 3.1 Pro”. [Online]. Available: https://deepmind.google/models/gemini/pro
- Gottweis, J. et al. (2025). Accelerating scientific discovery with Co-Scientist. arXiv.Gottweis, J., Weng, W.-H., Daryin, A., Tu, T., Sirkovic, P., Myaskovsky, A., Glowaty, G., Weissenberger, F., Orlandi, A., Popovici, D., Palepu, A., Rong, K., Tanno, R., Saab, K., Zhang, F., Blum, J., Carroll, A., Kulkarni, K., Tomasev, N., … Natarajan, V. (2025). Accelerating scientific discovery with Co-Scientist. In arXiv. https://doi.org/10.1038/s41586-026-10644-yGottweis, J., W.-H. Weng, A. Daryin, et al. 2025. “Accelerating Scientific Discovery with Co-Scientist”. In arXiv. Preprint, February 26. https://doi.org/10.1038/s41586-026-10644-y.Gottweis, J., et al. “Accelerating Scientific Discovery with Co-Scientist”. arXiv, 26 Feb. 2025, https://doi.org/10.1038/s41586-026-10644-y.Gottweis, J. et al. Accelerating scientific discovery with Co-Scientist. arXiv Preprint at https://doi.org/10.1038/s41586-026-10644-y (2025).J. Gottweis et al., “Accelerating scientific discovery with Co-Scientist”, Feb. 26, 2025. doi: 10.1038/s41586-026-10644-y.
- IBM (2025). Deep Blue. IBM.IBM. (2025, November 10). Deep Blue. IBM. https://ibm.com/history/deep-blueIBM. 2025. “Deep Blue”. IBM, November 10. https://ibm.com/history/deep-blue.IBM. “Deep Blue”. IBM, 10 Nov. 2025, https://ibm.com/history/deep-blue.IBM. Deep Blue. IBM https://ibm.com/history/deep-blue (2025).IBM, “Deep Blue”, IBM. [Online]. Available: https://ibm.com/history/deep-blue
- IBM (2025). Watson, Jeopardy! champion. IBM.IBM. (2025, November 10). Watson, Jeopardy! champion. IBM. https://ibm.com/history/watson-jeopardyIBM. 2025. “Watson, Jeopardy! Champion”. IBM, November 10. https://ibm.com/history/watson-jeopardy.IBM. “Watson, Jeopardy! Champion”. IBM, 10 Nov. 2025, https://ibm.com/history/watson-jeopardy.IBM. Watson, Jeopardy! champion. IBM https://ibm.com/history/watson-jeopardy (2025).IBM, “Watson, Jeopardy! champion”, IBM. [Online]. Available: https://ibm.com/history/watson-jeopardy
- Jack Parker-Holder & Shlomi Fruchter (2025). Genie 3: A new frontier for world models.Jack Parker-Holder, & Shlomi Fruchter. (2025, August 5). Genie 3: A new frontier for world models. https://deepmind.google/blog/genie-3-a-new-frontier-for-world-modelsJack Parker-Holder, and Shlomi Fruchter. 2025. “Genie 3: A New Frontier for World Models”. August 5. https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models.Jack Parker-Holder, and Shlomi Fruchter. Genie 3: A New Frontier for World Models. 5 Aug. 2025, https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models.Jack Parker-Holder & Shlomi Fruchter. Genie 3: A new frontier for world models. https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models (2025).Jack Parker-Holder and Shlomi Fruchter, “Genie 3: A new frontier for world models”. [Online]. Available: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models
- Julian Schrittwieser et al. (2020). MuZero: Mastering Go, chess, shogi and Atari without rules.Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, & David Silver. (2020, December 23). MuZero: Mastering Go, chess, shogi and Atari without rules. https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rulesJulian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, et al. 2020. “MuZero: Mastering Go, Chess, Shogi and Atari Without Rules”. December 23. https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules.Julian Schrittwieser, et al. MuZero: Mastering Go, Chess, Shogi and Atari Without Rules. 23 Dec. 2020, https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules.Julian Schrittwieser et al. MuZero: Mastering Go, chess, shogi and Atari without rules. https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules (2020).Julian Schrittwieser et al., “MuZero: Mastering Go, chess, shogi and Atari without rules”. [Online]. Available: https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules
- Kadavath, S. et al. (2022). Language Models (Mostly) Know What They Know. arXiv.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., … Kaplan, J. (2022). Language Models (Mostly) Know What They Know. In arXiv. https://arxiv.org/abs/2207.05221Kadavath, S., T. Conerly, A. Askell, et al. 2022. “Language Models (Mostly) Know What They Know”. In arXiv. Preprint, July 11. https://arxiv.org/abs/2207.05221.Kadavath, S., et al. “Language Models (Mostly) Know What They Know”. arXiv, 11 July 2022, https://arxiv.org/abs/2207.05221.Kadavath, S. et al. Language Models (Mostly) Know What They Know. arXiv Preprint at https://arxiv.org/abs/2207.05221 (2022).S. Kadavath et al., “Language Models (Mostly) Know What They Know”, Jul. 11, 2022. [Online]. Available: https://arxiv.org/abs/2207.05221
- Kosinski, M. (2023). Evaluating Large Language Models in Theory of Mind Tasks. arXiv.Kosinski, M. (2023). Evaluating Large Language Models in Theory of Mind Tasks. In arXiv. https://doi.org/10.1073/pnas.2405460121Kosinski, M. 2023. “Evaluating Large Language Models in Theory of Mind Tasks”. In arXiv. Preprint, February 4. https://doi.org/10.1073/pnas.2405460121.Kosinski, M. “Evaluating Large Language Models in Theory of Mind Tasks”. arXiv, 4 Feb. 2023, https://doi.org/10.1073/pnas.2405460121.Kosinski, M. Evaluating Large Language Models in Theory of Mind Tasks. arXiv Preprint at https://doi.org/10.1073/pnas.2405460121 (2023).M. Kosinski, “Evaluating Large Language Models in Theory of Mind Tasks”, Feb. 04, 2023. doi: 10.1073/pnas.2405460121.
- Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J. & Ha, D. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv.Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., & Ha, D. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. In arXiv. https://arxiv.org/abs/2408.06292Lu, C., C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha. 2024. “The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery”. In arXiv. Preprint, August 12. https://arxiv.org/abs/2408.06292.Lu, C., et al. “The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery”. arXiv, 12 Aug. 2024, https://arxiv.org/abs/2408.06292.Lu, C. et al. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv Preprint at https://arxiv.org/abs/2408.06292 (2024).C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha, “The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery”, Aug. 12, 2024. [Online]. Available: https://arxiv.org/abs/2408.06292
- Maslej, N. et al. (2025). Artificial Intelligence Index Report 2025. arXiv.Maslej, N., Fattorini, L., Perrault, R., Gil, Y., Parli, V., Kariuki, N., Capstick, E., Reuel, A., Brynjolfsson, E., Etchemendy, J., Ligett, K., Lyons, T., Manyika, J., Niebles, J. C., Shoham, Y., Wald, R., Walsh, T., Hamrah, A., Santarlasci, L., … Oak, S. (2025). Artificial Intelligence Index Report 2025. In arXiv. https://arxiv.org/abs/2504.07139Maslej, N., L. Fattorini, R. Perrault, et al. 2025. “Artificial Intelligence Index Report 2025”. In arXiv. Preprint, April 8. https://arxiv.org/abs/2504.07139.Maslej, N., et al. “Artificial Intelligence Index Report 2025”. arXiv, 8 Apr. 2025, https://arxiv.org/abs/2504.07139.Maslej, N. et al. Artificial Intelligence Index Report 2025. arXiv Preprint at https://arxiv.org/abs/2504.07139 (2025).N. Maslej et al., “Artificial Intelligence Index Report 2025”, Apr. 08, 2025. [Online]. Available: https://arxiv.org/abs/2504.07139
- Mnih, V. et al. (2013). Playing Atari with Deep Reinforcement Learning. arXiv.Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., & Riedmiller, M. (2013). Playing Atari with Deep Reinforcement Learning. In arXiv. https://arxiv.org/abs/1312.5602Mnih, V., K. Kavukcuoglu, D. Silver, et al. 2013. “Playing Atari with Deep Reinforcement Learning”. In arXiv. Preprint, December 19. https://arxiv.org/abs/1312.5602.Mnih, V., et al. “Playing Atari with Deep Reinforcement Learning”. arXiv, 19 Dec. 2013, https://arxiv.org/abs/1312.5602.Mnih, V. et al. Playing Atari with Deep Reinforcement Learning. arXiv Preprint at https://arxiv.org/abs/1312.5602 (2013).V. Mnih et al., “Playing Atari with Deep Reinforcement Learning”, Dec. 19, 2013. [Online]. Available: https://arxiv.org/abs/1312.5602
- OpenAI (2019). OpenAI Five defeats Dota 2 world champions.OpenAI. (2019). OpenAI Five defeats Dota 2 world champions. Internet Archive (https://web.archive.org/web/20240425003107/https://openai.com/research/openai-five-defeats-dota-2-world-champions). https://openai.com/research/openai-five-defeats-dota-2-world-championsOpenAI. 2019. “OpenAI Five Defeats Dota 2 World Champions”. Https://web.archive.org/web/20240425003107/https://openai.com/research/openai-five-defeats-dota-2-world-champions. Internet Archive. https://openai.com/research/openai-five-defeats-dota-2-world-champions.OpenAI. OpenAI Five Defeats Dota 2 World Champions. 2019, Internet Archive, https://web.archive.org/web/20240425003107/https://openai.com/research/openai-five-defeats-dota-2-world-champions, https://openai.com/research/openai-five-defeats-dota-2-world-champions.OpenAI. OpenAI Five defeats Dota 2 world champions. https://openai.com/research/openai-five-defeats-dota-2-world-champions (2019).OpenAI, “OpenAI Five defeats Dota 2 world champions”. Accessed: Apr. 25, 2024. [Online]. Available: https://openai.com/research/openai-five-defeats-dota-2-world-champions
- OpenAI (2024). Introducing OpenAI o1. OpenAI.OpenAI. (2024). Introducing OpenAI o1. OpenAI. https://openai.com/o1OpenAI. 2024. “Introducing OpenAI O1”. OpenAI. https://openai.com/o1.OpenAI. “Introducing OpenAI O1”. OpenAI, 2024, https://openai.com/o1.OpenAI. Introducing OpenAI o1. OpenAI https://openai.com/o1 (2024).OpenAI, “Introducing OpenAI o1”, OpenAI. [Online]. Available: https://openai.com/o1
- OpenAI (2025). Introducing GPT-5.2. OpenAI.OpenAI. (2025). Introducing GPT-5.2. OpenAI. https://openai.com/index/introducing-gpt-5-2OpenAI. 2025. “Introducing GPT-5.2”. OpenAI. https://openai.com/index/introducing-gpt-5-2.OpenAI. “Introducing GPT-5.2”. OpenAI, 2025, https://openai.com/index/introducing-gpt-5-2.OpenAI. Introducing GPT-5.2. OpenAI https://openai.com/index/introducing-gpt-5-2 (2025).OpenAI, “Introducing GPT-5.2”, OpenAI. [Online]. Available: https://openai.com/index/introducing-gpt-5-2
- OpenAI (2025). OpenAI (@OpenAI) on X. X (formerly Twitter).OpenAI. (2025, July 19). OpenAI (@OpenAI) on X. X (formerly Twitter). https://x.com/OpenAI/status/1946594928945148246OpenAI. 2025. “OpenAI (@OpenAI) on X”. X (formerly Twitter), July 19. https://x.com/OpenAI/status/1946594928945148246.OpenAI. “OpenAI (@OpenAI) on X”. X (formerly Twitter), 19 July 2025, https://x.com/OpenAI/status/1946594928945148246.OpenAI. OpenAI (@OpenAI) on X. X (formerly Twitter) https://x.com/OpenAI/status/1946594928945148246 (2025).OpenAI, “OpenAI (@OpenAI) on X”, X (formerly Twitter). [Online]. Available: https://x.com/OpenAI/status/1946594928945148246
- OpenAI (2025). Sora 2 is here. OpenAI.OpenAI. (2025). Sora 2 is here. OpenAI. https://openai.com/index/sora-2OpenAI. 2025. “Sora 2 Is Here”. OpenAI. https://openai.com/index/sora-2.OpenAI. “Sora 2 Is Here”. OpenAI, 2025, https://openai.com/index/sora-2.OpenAI. Sora 2 is here. OpenAI https://openai.com/index/sora-2 (2025).OpenAI, “Sora 2 is here”, OpenAI. [Online]. Available: https://openai.com/index/sora-2
- OpenAI et al. (2023). GPT-4 Technical Report. arXiv.OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., … Zoph, B. (2023). GPT-4 Technical Report. In arXiv. https://arxiv.org/abs/2303.08774OpenAI, J. Achiam, S. Adler, et al. 2023. “GPT-4 Technical Report”. In arXiv. Preprint, March 15. https://arxiv.org/abs/2303.08774.OpenAI, et al. “GPT-4 Technical Report”. arXiv, 15 Mar. 2023, https://arxiv.org/abs/2303.08774.OpenAI et al. GPT-4 Technical Report. arXiv Preprint at https://arxiv.org/abs/2303.08774 (2023).OpenAI et al., “GPT-4 Technical Report”, Mar. 15, 2023. [Online]. Available: https://arxiv.org/abs/2303.08774
- Qin, Y. et al. (2023). ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv.Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., Zhao, S., Hong, L., Tian, R., Xie, R., Zhou, J., Gerstein, M., Li, D., Liu, Z., & Sun, M. (2023). ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. In arXiv. https://arxiv.org/abs/2307.16789Qin, Y., S. Liang, Y. Ye, et al. 2023. “ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs”. In arXiv. Preprint, July 31. https://arxiv.org/abs/2307.16789.Qin, Y., et al. “ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs”. arXiv, 31 July 2023, https://arxiv.org/abs/2307.16789.Qin, Y. et al. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv Preprint at https://arxiv.org/abs/2307.16789 (2023).Y. Qin et al., “ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs”, Jul. 31, 2023. [Online]. Available: https://arxiv.org/abs/2307.16789
- Robotics team (2024). Shaping the future of advanced robotics.Robotics team. (2024, January 4). Shaping the future of advanced robotics. https://deepmind.google/blog/shaping-the-future-of-advanced-roboticsRobotics team. 2024. “Shaping the Future of Advanced Robotics”. January 4. https://deepmind.google/blog/shaping-the-future-of-advanced-robotics.Robotics team. Shaping the Future of Advanced Robotics. 4 Jan. 2024, https://deepmind.google/blog/shaping-the-future-of-advanced-robotics.Robotics team. Shaping the future of advanced robotics. https://deepmind.google/blog/shaping-the-future-of-advanced-robotics (2024).Robotics team, “Shaping the future of advanced robotics”. [Online]. Available: https://deepmind.google/blog/shaping-the-future-of-advanced-robotics
- Schick, T. et al. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv.Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. In arXiv. https://arxiv.org/abs/2302.04761Schick, T., J. Dwivedi-Yu, R. Dessì, et al. 2023. “Toolformer: Language Models Can Teach Themselves to Use Tools”. In arXiv. Preprint, February 9. https://arxiv.org/abs/2302.04761.Schick, T., et al. “Toolformer: Language Models Can Teach Themselves to Use Tools”. arXiv, 9 Feb. 2023, https://arxiv.org/abs/2302.04761.Schick, T. et al. Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv Preprint at https://arxiv.org/abs/2302.04761 (2023).T. Schick et al., “Toolformer: Language Models Can Teach Themselves to Use Tools”, Feb. 09, 2023. [Online]. Available: https://arxiv.org/abs/2302.04761
- SIMA team (2025). SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds.SIMA team. (2025, November 13). SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds. https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worldsSIMA team. 2025. “SIMA 2: An Agent That Plays, Reasons, and Learns With You in Virtual 3D Worlds”. November 13. https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds.SIMA team. SIMA 2: An Agent That Plays, Reasons, and Learns With You in Virtual 3D Worlds. 13 Nov. 2025, https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds.SIMA team. SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds. https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds (2025).SIMA team, “SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds”. [Online]. Available: https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds
- SWE bench (2025). SWE-bench Leaderboards.SWE bench. (2025). SWE-bench Leaderboards. https://swebench.comSWE bench. 2025. “SWE-bench Leaderboards”. https://swebench.com.SWE bench. SWE-bench Leaderboards. 2025, https://swebench.com.SWE bench. SWE-bench Leaderboards. https://swebench.com (2025).SWE bench, “SWE-bench Leaderboards”. [Online]. Available: https://swebench.com
- Thang Luong & Edward Lockhart (2025). Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad.Thang Luong, & Edward Lockhart. (2025, July 21). Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad. https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiadThang Luong, and Edward Lockhart. 2025. “Advanced Version of Gemini with Deep Think Officially Achieves Gold-medal Standard at the International Mathematical Olympiad”. July 21. https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad.Thang Luong, and Edward Lockhart. Advanced Version of Gemini with Deep Think Officially Achieves Gold-medal Standard at the International Mathematical Olympiad. 21 July 2025, https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad.Thang Luong & Edward Lockhart. Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad. https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad (2025).Thang Luong and Edward Lockhart, “Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad”. [Online]. Available: https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad
- Venkat Somala, Anson Ho & Séb Krier (2025). Three challenges facing compute-based AI policies.Venkat Somala, Anson Ho, & Séb Krier. (2025, September 11). Three challenges facing compute-based AI policies. https://epoch.ai/gradient-updates/three-issues-undermining-compute-based-ai-policiesVenkat Somala, Anson Ho, and Séb Krier. 2025. “Three Challenges Facing Compute-based AI Policies”. September 11. https://epoch.ai/gradient-updates/three-issues-undermining-compute-based-ai-policies.Venkat Somala, et al. Three Challenges Facing Compute-based AI Policies. 11 Sept. 2025, https://epoch.ai/gradient-updates/three-issues-undermining-compute-based-ai-policies.Venkat Somala, Anson Ho & Séb Krier. Three challenges facing compute-based AI policies. https://epoch.ai/gradient-updates/three-issues-undermining-compute-based-ai-policies (2025).Venkat Somala, Anson Ho, and Séb Krier, “Three challenges facing compute-based AI policies”. [Online]. Available: https://epoch.ai/gradient-updates/three-issues-undermining-compute-based-ai-policies
- Wang et al. (2022). Adversarial Policies Beat Superhuman Go AIs. arXiv.org.Wang et al. (2022). Adversarial Policies Beat Superhuman Go AIs. arXiv.org. https://www.arxiv.org/abs/2211.00241Wang et al. 2022. “Adversarial Policies Beat Superhuman Go AIs”. arXiv.org. https://www.arxiv.org/abs/2211.00241.Wang et al. “Adversarial Policies Beat Superhuman Go AIs”. arXiv.org, 2022, https://www.arxiv.org/abs/2211.00241.Wang et al. Adversarial Policies Beat Superhuman Go AIs. arXiv.org https://www.arxiv.org/abs/2211.00241 (2022).Wang et al., “Adversarial Policies Beat Superhuman Go AIs”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2211.00241
- Wang, G. et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. In arXiv. https://arxiv.org/abs/2305.16291Wang, G., Y. Xie, Y. Jiang, et al. 2023. “Voyager: An Open-Ended Embodied Agent with Large Language Models”. In arXiv. Preprint, May 25. https://arxiv.org/abs/2305.16291.Wang, G., et al. “Voyager: An Open-Ended Embodied Agent with Large Language Models”. arXiv, 25 May 2023, https://arxiv.org/abs/2305.16291.Wang, G. et al. Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv Preprint at https://arxiv.org/abs/2305.16291 (2023).G. Wang et al., “Voyager: An Open-Ended Embodied Agent with Large Language Models”, May 25, 2023. [Online]. Available: https://arxiv.org/abs/2305.16291
- Wooldridge (2021). A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books.Wooldridge. (2021). A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books. https://amazon.com/Brief-History-Artificial-Intelligence-Where/dp/1250770742Wooldridge. 2021. “A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books”. https://amazon.com/Brief-History-Artificial-Intelligence-Where/dp/1250770742.Wooldridge. A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books. 2021, https://amazon.com/Brief-History-Artificial-Intelligence-Where/dp/1250770742.Wooldridge. A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books. https://amazon.com/Brief-History-Artificial-Intelligence-Where/dp/1250770742 (2021).Wooldridge, “A Brief History of Artificial Intelligence: What It Is, Where We Are, and Where We Are Going: Wooldridge, Michael: 9781250770745: Amazon.com: Books”. [Online]. Available: https://amazon.com/Brief-History-Artificial-Intelligence-Where/dp/1250770742
- Xu, H., Zhao, R., Zhu, L., Du, J. & He, Y. (2024). OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models. arXiv.Xu, H., Zhao, R., Zhu, L., Du, J., & He, Y. (2024). OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models. In arXiv. https://arxiv.org/abs/2402.06044Xu, H., R. Zhao, L. Zhu, J. Du, and Y. He. 2024. “OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models”. In arXiv. Preprint, February 8. https://arxiv.org/abs/2402.06044.Xu, H., et al. “OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models”. arXiv, 8 Feb. 2024, https://arxiv.org/abs/2402.06044.Xu, H., Zhao, R., Zhu, L., Du, J. & He, Y. OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models. arXiv Preprint at https://arxiv.org/abs/2402.06044 (2024).H. Xu, R. Zhao, L. Zhu, J. Du, and Y. He, “OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models”, Feb. 08, 2024. [Online]. Available: https://arxiv.org/abs/2402.06044
- Yap (2024). How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs). Gold Penguin.Yap. (2024, January 8). How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs). Gold Penguin. https://goldpenguin.org/blog/midjourney-v1-to-v6-evolutionYap. 2024. “How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs)”. Gold Penguin, January 8. https://goldpenguin.org/blog/midjourney-v1-to-v6-evolution.Yap. “How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs)”. Gold Penguin, 8 Jan. 2024, https://goldpenguin.org/blog/midjourney-v1-to-v6-evolution.Yap. How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs). Gold Penguin https://goldpenguin.org/blog/midjourney-v1-to-v6-evolution (2024).Yap, “How Midjourney Evolved Over Time (Comparing V1 to V7 Outputs)”, Gold Penguin. [Online]. Available: https://goldpenguin.org/blog/midjourney-v1-to-v6-evolution
- Žiga Avsec & Natasha Latysheva (2025). AlphaGenome: AI for better understanding the genome.Žiga Avsec, & Natasha Latysheva. (2025, June 25). AlphaGenome: AI for better understanding the genome. https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genomeŽiga Avsec, and Natasha Latysheva. 2025. “AlphaGenome: AI for Better Understanding the Genome”. June 25. https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome.Žiga Avsec, and Natasha Latysheva. AlphaGenome: AI for Better Understanding the Genome. 25 June 2025, https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome.Žiga Avsec & Natasha Latysheva. AlphaGenome: AI for better understanding the genome. https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome (2025).Žiga Avsec and Natasha Latysheva, “AlphaGenome: AI for better understanding the genome”. [Online]. Available: https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome
Was this section useful?
Thank you for your feedback
Your input helps improve the Atlas.