Current LLMs, although trained on abundant data, are still far from perfect.
Will these problems persist in future iterations, or will they disappear? This section examines the main criticisms of those models and tries to determine if they are valid even for future LLMs.
This kind of qualitative assessment is important to know whether LLMs represent the most likely route to AGI or not.
Empirically insufficiency? #
Can LLMs be creative? The creativity of LLMs is often debated, but there are clear indications that AI, in principle, is capable of creative processes in various ways:
- Autonomous Scientific Research: Recent advancements have shown that LLMs can indeed make novel discoveries. For example, a study by DeepMind demonstrated that an LLM "discovered new solutions for the cap set problem, a long-standing open problem in mathematics" (DeepMind, 2023) which was a favorite open problem of Terence Tao. This indicates that AI can not only understand existing knowledge but also contribute new insights in complex fields like mathematics.
- Autonomous Discovery: AI has the capability to rediscover human strategies and openings independently. AlphaGo, for example, rediscovered human Go strategies and openings through self-play (McGrath et al., 2021), without any human data input. This demonstrates an AI's ability to independently learn and innovate within established domains.
- Creative Optimization: AI can optimize in surprisingly creative ways. The phenomena of specification gaming, where AI finds unintended solutions to problems, illustrate this. Although this unpredictability poses its challenges, it also shows that AI systems can come up with novel, creative solutions that might not be immediately obvious or intuitive to human problem solvers. DeepMind's blog post on Specification Gaming illustrates this point vividly (Krakovna et al., 2020).
Aren’t LLMs just too slow at learning things? Arguments against transformer based language models often state that they are too sample inefficient, and that LLMs are extremely slow to learn new concepts when compared to humans. To increase performance in new tasks or situations, it’s often argued that LLMs require training on vast amounts of data — millions of times more than a human would need. However, there's a growing trend towards data efficiency, and an increasing belief that this can be significantly improved in future models.
EfficientZero is a reinforcement learning agent that surpasses median human performance on a set of 26 Atari games after just two hours of real-time experience per game (Ye et al., 2021; Wang et al., 2024). This is a considerable improvement over previous algorithms, showcasing the potential leaps in data efficiency. The promise here is not just more efficient learning but also the potential for rapid adaptation and proficiency in new tasks, akin to a child's learning speed. EfficientZero is not an LLM, but it shows that deep learning can sometimes be made efficient.
Scaling laws indicate that larger AIs tend to be more data efficient, requiring less data to reach the same level of performance as their smaller counterparts. Papers such as "Language Models are Few-Shot Learners" (Brown et al., 2020) and the evidence that larger models seem to take less data to reach the same level of performance (Kaplan et al., 2020), suggest that as models scale, they become more proficient with fewer examples. This trend points towards a future where AI might be able to rapidly adapt and learn from limited data, challenging the notion that AIs are inherently slow learners compared to humans.
Are LLMs robust to distributional shifts? While it is true that AI has not yet achieved maximal robustness, for example being able to perform perfectly after a change in distribution, there has been considerable progress:
- Robustness correlates with capabilities: Robustness is closely linked to the capabilities of AI models when AIs are trained on difficult tasks. For instance, there is a significant improvement in robustness and transfer learning from GPT-2 to GPT-4. In computer vision, recent models like Segment Anything (Kirillov et al., 2023) are far more robust and capable of transfer learning than their less capable predecessors. This progression isn't due to any mysterious factors but rather a result of scaling and improving upon existing architectures.
- Robustness is a continuum, and perfect robustness may be not necessary: Robustness in AI should not be viewed as a binary concept, but rather as existing on a continuum. This continuum is evident in the way AI models, like those in image classification, often surpass human performance in both capability and robustness (Korzekwa, 2022). However, it's important to recognize that no system is completely immune to challenges such as adversarial attacks. This is exemplified by advanced AIs like Katago in Go, which, despite being vulnerable to such attacks (Wang et al., 2022), still achieves a superhuman level of play. However, the quest for perfect robustness may not be essential to create capable transformative AI, as even systems with certain vulnerabilities can achieve superhuman levels of competence. However, while robustness may not be necessary to create capable AI, the creation of safe, aligned AI will have to solve the problem of misgeneralizing goals.
Shallow Understanding? #
Stochastic Parrots: Do AIs only memorize information without truly compressing it? There are two archetypal ways to represent information in an LLM: either memorize point by point, like a look-up table, or compress the information by only memorizing higher-level features, which we can then call "the world model". This is explained in the very important paper "Superposition, Memorization, and Double Descent" (Anthropic, 2023): it turns out that to store points, initially the model learns the position of all the points (pure memorization), then, if we increase the number of points, the model starts to compress this knowledge, and the model is now capable of generalization (and implements a simple model of the data).
Unfortunately, too few people understand the distinction between memorization and understanding. It's not some lofty question like ‘does the system have an internal world model?’, it's a very pragmatic behavior distinction: ‘is the system capable of broad generalization, or is it limited to local generalization?’
AI is capable of compressing information, often in a relevant manner. For example, when examining the representations of words representing colors in LLMs like "red" and "blue", the structure formed by all the embeddings of those colors creates the correct color circle (This uses a nonlinear projection such as a T-distributed stochastic neighbor embedding (T-SNE) to project from high-dimensional space to the 2D plane). Other examples of world models are presented in a paper called "Eight Things to Know about Large Language Models" (Bowman, 2023).
Of course, there are other domains where AI resembles more of a look-up table, but it is a spectrum, and each case should be examined individually. For example, for "factual association," the paper "Locating and Editing Factual Associations in GPT" shows that the underlying data structure for GPT-2 is more of a look-up table (Meng et al., 2023), but the paper "Emergent Linear Representations in World Models of Self-Supervised Sequence Models" demonstrates that a small GPT is capable of learning a compressed world model of OthelloGpt. (Nanda et al., 2023) There are more examples in the section dedicated to world models in the paper "Eight Things to Know about Large Language Models" (Bowman, 2023).
It’s clear that LLMs are compressing their representations at least a bit. Many examples of impressive capabilities are presented in the work "The Stochastic Parrot Hypothesis is debatable for the last generation of LLMs", which shows that it cannot be purely a memorization. (Feuillade-Montixi & Peigné, 2023)
Will LLMs Inevitably Hallucinate?
LLMs are prone to "hallucinate," a term used to describe the generation of content that is nonsensical or factually incorrect in response to certain prompts. This issue, highlighted in studies such as "On Faithfulness and Factuality in Abstractive Summarization" by Maynez et al. (Maynez et al., 2020) and "TruthfulQA: Measuring How Models Mimic Human Falsehoods" by Lin et al. (Lin et al., 2022), poses a significant challenge. However, it's important to see that these challenges are anticipated due to the training setup and can be mitigated:
- Inherent Bias in Source Texts: One of the fundamental reasons LLMs may produce untrue content is training data , which may not always be entirely factual or unbiased. In essence, LLMs are reflecting the diverse and sometimes contradictory nature of their training data . In this context, LLMs are constantly 'hallucinating', but occasionally, these hallucinations align with our perception of reality.
- Strategies to Enhance Factual Accuracy: The tendency of LLMs to generate hallucinations can be significantly diminished using various techniques. See the box below for a breakdown of those.
- Larger models can be more truthful than smaller ones. This is the case with TruthfulQA. OpenAI reports that GPT-4 is 40% more accurate and factually consistent than its predecessor.
Structural inadequacy? #
Are LLMs missing System 2? System 1 and System 2 are terms popularized by economist Daniel Kahneman in his book "Thinking, Fast and Slow," describing the two different ways our brains form thoughts and make decisions. System 1 is fast, automatic, and intuitive; it's the part of our thinking that handles everyday decisions and judgments without much effort or conscious deliberation. For example, when you recognize a face or understand simple sentences, you're typically using System 1. On the other hand, System 2 is slower, more deliberative, and more logical. It takes over when you're solving a complex problem, making a conscious choice, or focusing on a difficult task. It requires more energy and is more controlled, handling tasks such as planning for the future, checking the validity of a complex argument, or any activity that requires deep focus. Together, these systems interact and influence how we think, make judgments, and decide, highlighting the complexity of human thought and behavior.
A key concern is whether LLMs are able to emulate System 2 processes, which involve slower, more deliberate, and logical thinking. Some theoretical arguments about the depth limit in transformers show that they are provably incapable of internally dividing large integers (Delétang et al., 2023). However, this is not what we observe in practice: GPT-4 is capable of detailing some calculations step-by-step and obtaining the expected result through a chain of thought or via the usage of tools like a code interpreter.
Emerging Metacognition. Emerging functions in LLMs, like the Reflexion technique (Shinn et al., 2023), allow these models to retrospectively analyze and improve their answers. It is possible to ask the LLM to take a step back, question the correctness of its previous actions, and consider ways to improve the previous answer. This greatly enhances the capabilities of GPT-4, enhancing its capabilities and aligning them more closely with human System 2 operations. Note that this technique is emergent and does not work well with previous models.
These results suggest a blurring of the lines between these two systems. System 2 processes may be essentially an assembly of multiple System 1 processes, appearing slower due to involving more steps and interactions with slower forms of memory. This perspective is paralleled in how language models operate, with each step in a System 1 process akin to a constant time execution step in models like GPT. Although these models struggle with intentionally orchestrating these steps to solve complex problems, breaking down tasks into smaller steps (Least-to-most prompting) or prompting them for incremental reasoning (Chain-of-Thought (CoT) prompting) significantly improves their performance.
Are LLMs missing an internal world model? The notion of a "world model" in AI need not be confined to explicit encoding within an architecture. Contrary to approaches like H-JEPA (LeCun, 2022), which advocate for an explicit world model to enhance AI training, there's growing evidence that a world model can be effectively implicit. This concept is particularly evident in reinforcement learning (RL), where the distinction between model-based and model-free RL can be somewhat misleading. Even in model-free RL, algorithms often implicitly encode a form of a world model that is crucial for optimal performance.
- Time and geographical coordinates: Research on Llama-2 models reveals how these models can represent spatial and temporal information (Gurney & Tegmark, 2024). LLMs like Llama-2 models encode approximate real-world coordinates and historical timelines of cities. Key findings include the gradual emergence of geographical representations across model layers, the linearity of these representations, and the models' robustness to different prompts. Significantly, the study shows that the models are not just passively processing this information but actively learning the global geometry of space and time.
- Board representation: In the paper "Emergent Linear Representations in World Models of Self-Supervised Sequence Models" (Nanda et al., 2023), the author presents significant findings on the nature of representations in AI models. The paper delves into how the Othello-GPT model, trained to predict legal moves in the game of Othello, develops an emergent world representation of the game board! Contrary to previous beliefs that this representation was non-linear, he demonstrates that it is, in fact, linear. He discovers that the model represents board states not in terms of black or white pieces, but as "my color" or "their color," aligning with the model's perspective of playing both sides. This work sheds light on the potential of AI models to develop complex, yet linear, world representations through simple objectives like next-token prediction.
- Other examples are presented in the paper: "Eight Things to know about LLMs". (Bowman, 2023)
Can LLMs learn continuously, and have long term memory? Continual learning and the effective management of long-term memory represent significant challenges in the field of AI in general.
Catastrophic Forgetting. A crucial obstacle in this area is catastrophic forgetting, a phenomenon where a neural network , upon learning new information, tends to entirely forget previously learned information. This issue is an important focus of ongoing research, aiming to develop AI systems that can retain and build upon their knowledge over time. For example, suppose we train an AI on an Atari game. At the end of the second training, the AI has most likely forgotten how to play the first game. This is an example of catastrophic forgetting.
But now suppose we train a large AI on many ATARI games, simultaneously, and even add some Internet text and some robotic tasks. This can just work. For example, the AI GATO illustrates this training process and exemplifies what we call the blessing of scale, which is that what is impossible in small regimes can become possible in large regimes.
Other techniques are being developed to solve long-term memory, for example, Scaffolding-based approaches have also been employed for achieving long-term memory and continual learning in AI. Scaffolding in AI refers to the use of hard-coded wrappers explicitly programmed structures by humans that involve a for loop to query continuously the model:
- LangChain addresses these challenges by creating extensive memory banks. LangChain is a Python library that allows LLM to retrieve and utilize information from large datasets, essentially providing a way for AI to access a vast repository of knowledge and use this information to construct more informed responses. However, this approach may not be the most elegant due to its reliance on external data sources and complex retrieval mechanisms. A potentially more seamless and integrated solution could involve utilizing the neural network 's weights as dynamic memory, constantly evolving and updating based on the tasks performed by the network.
- Voyager: A remarkable example of a scaffolding-based long-term memory is the AI Voyager, an AI system developed under the "AutoGPT" paradigm. This system is notable for its ability to engage in continuous learning within a 3D game environment like Minecraft. In a single game session, AI Voyager demonstrates the capacity to learn basic controls, achieve initial goals such as resource acquisition, and eventually advance to more complex behaviors, including combat with enemies and crafting tools for gathering sophisticated resources. This demonstrates a significant stride in LLM's ability to learn continually and manage long-term memory within dynamic environments.
It should be noted that scaffold-based long-term memory is not considered an elegant solution, and purists would prefer to use the system's own weights as long-term memory.
Planning
Planning is an area that AIs currently struggle with, but there is significant progress. Some paradigms, such as those based on scaffolding, enable task decomposition and breaking down objectives into smaller, more achievable sub-objectives.
Furthermore, the paper "Voyager: An Open-Ended Embodied Agent with Large Language Models" demonstrates that it is possible to use GPT-4 for planning in Natural language in Minecraft (Wang et al., 2023).
Differences with the brain #
It appears that there are several points of convergence between the LLMs and the linguistic cortex:
- Behavioral similarities. LLMs show a close comparison to human linguistic abilities and the linguistic cortex (Canell, 2022). These models have excelled in mastering syntax and a significant portion of semantics in human language. Of course, today, they still lag in aspects such as long-term memory, coherence, and general reasoning - faculties that in humans depend on various brain regions like the hippocampus and prefrontal cortex, but we explained in the last sections that those problems may be solvable.
- Convergence in internal Representations: LLMs have a representation that converges with scale toward the brain representation. This is supported by the study, "Brains and algorithms partially converge in natural language processing." (Caucheteux & King, 2022) Additional insights can be found in the works "The Brain as a Universal Learning Machine" (Canell, 2015) and "Brain Efficiency: Much More than You Wanted to Know." (Canell, 2022) At comparable learning stages, LLMs and the linguistic cortex develop similar or equivalent feature representations. In some evaluations, advanced LLMs have been able to predict 100% of the explainable neural variance, as detailed by Schrimpf, Martin, et al. in "The neural architecture of language: Integrative modeling converges on predictive processing." (Schrimpf et al., 2021)
- Scale is also important in primates. The principal architectural difference between human and other primate brains seems to be the number of neurons rather than anything else, as demonstrated in various studies. (Houzel, 2012; Pearson et al., 2023; Charvet, 2021).
Further reasons to continue scaling LLMs #
Following are some reasons to believe that labs will continue to scale LLMs.
Scaling Laws on LLM implies further qualitative improvements. The scaling laws might not initially appear impressive. However, linking these quantitative measures can translate to a qualitative improvement in algorithm quality. An algorithm that achieves near-perfect loss, though, is one that necessarily comprehends all subtleties, and displays enormous adaptability. The fact that the scaling laws are not bending is very significant and means that we can make the model a qualitatively better reasoner.
From simple correlations to understanding. During a training run, GPTs go from basic correlations to deeper and deeper understanding. Initially, the model merely establishes connections between successive words. Gradually, it develops an understanding of grammar and semantics, creating links between sentences and subsequently between paragraphs. Eventually, GPT masters the nuances of writing style.
Text completion is probably an AI-complete test (Wikipedia, 2022).
Current LLMs have only as many parameters as small mammals have synapses, no wonder they are still imperfect. Models like GPT-4, though very big compared to other models, should be noted for their relatively modest scale compared to the size of a human brain. To illustrate, the largest GPT-3 model has a similar number of parameters to the synapses of a hedgehog. We don't really know how many parameters GPT-4 has, but if it is the same size as PALM, which has 512 B parameters, then GPT-4 has only as many parameters as a chinchilla has synapses. In contrast, the human neocortex contains about 140 trillion synapses, which is over 200 times more synapses than a chinchilla. For a more in-depth discussion on this comparison, see the related discussion here. For a discussion of the number of parameters necessary to emulate a synapse, see the discussion on biological anchors.
GPT-4 is still orders of magnitude cheaper than other big science projects.: Despite the high costs associated with training large models, the significant leaps in AI capabilities provided by scaling justify these costs. For example, GPT-4 is expensive compared to other ML models. It is said to cost 50M in training. But the Manhattan Project cost 25B, which is 500 times more without accounting for inflation, and achieving Human-level intelligence, may be more economically important than achieving the nuclear bomb.
Collectively, these points support the idea that AGI can be achieved by only scaling current algorithms.
References
- Anthropic (2023). Superposition, Memorization, and Double Descent.Anthropic. (2023). Superposition, Memorization, and Double Descent. https://transformer-circuits.pub/2023/toy-double-descent/index.htmlAnthropic. 2023. “Superposition, Memorization, and Double Descent”. https://transformer-circuits.pub/2023/toy-double-descent/index.html.Anthropic. Superposition, Memorization, and Double Descent. 2023, https://transformer-circuits.pub/2023/toy-double-descent/index.html.Anthropic. Superposition, Memorization, and Double Descent. https://transformer-circuits.pub/2023/toy-double-descent/index.html (2023).Anthropic, “Superposition, Memorization, and Double Descent”. [Online]. Available: https://transformer-circuits.pub/2023/toy-double-descent/index.html
- Bowman, S. R. (2023). Eight Things to Know about Large Language Models. arXiv.Bowman, S. R. (2023). Eight Things to Know about Large Language Models. In arXiv. https://arxiv.org/abs/2304.00612Bowman, S. R. 2023. “Eight Things to Know About Large Language Models”. In arXiv. Preprint, April 2. https://arxiv.org/abs/2304.00612.Bowman, S. R. “Eight Things to Know About Large Language Models”. arXiv, 2 Apr. 2023, https://arxiv.org/abs/2304.00612.Bowman, S. R. Eight Things to Know about Large Language Models. arXiv Preprint at https://arxiv.org/abs/2304.00612 (2023).S. R. Bowman, “Eight Things to Know about Large Language Models”, Apr. 02, 2023. [Online]. Available: https://arxiv.org/abs/2304.00612
- Brown, T. B. et al. (2020). Language Models are Few-Shot Learners. arXiv.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language Models are Few-Shot Learners. In arXiv. https://arxiv.org/abs/2005.14165Brown, T. B., B. Mann, N. Ryder, et al. 2020. “Language Models Are Few-Shot Learners”. In arXiv. Preprint, May 28. https://arxiv.org/abs/2005.14165.Brown, T. B., et al. “Language Models Are Few-Shot Learners”. arXiv, 28 May 2020, https://arxiv.org/abs/2005.14165.Brown, T. B. et al. Language Models are Few-Shot Learners. arXiv Preprint at https://arxiv.org/abs/2005.14165 (2020).T. B. Brown et al., “Language Models are Few-Shot Learners”, May 28, 2020. [Online]. Available: https://arxiv.org/abs/2005.14165
- Caucheteux, C. & King, J. (2022). Brains and algorithms partially converge in natural language processing. Communications Biology.Caucheteux, C., & King, J.-R. (2022). Brains and algorithms partially converge in natural language processing. Communications Biology, 5, 134. https://doi.org/10.1038/s42003-022-03036-1Caucheteux, C., and J.-R. King. 2022. “Brains and Algorithms Partially Converge in Natural Language Processing”. Communications Biology 5 (February): 134. https://doi.org/10.1038/s42003-022-03036-1.Caucheteux, C., and J.-R. King. “Brains and Algorithms Partially Converge in Natural Language Processing”. Communications Biology, vol. 5, Feb. 2022, p. 134, https://doi.org/10.1038/s42003-022-03036-1.Caucheteux, C. & King, J.-R. Brains and algorithms partially converge in natural language processing. Communications Biology 5, 134 (2022).C. Caucheteux and J.-R. King, “Brains and algorithms partially converge in natural language processing”, Communications Biology, vol. 5, p. 134, Feb. 2022, doi: 10.1038/s42003-022-03036-1.
- Charvet, C. J. (2021). Cutting across structural and transcriptomic scales translates time across the lifespan in humans and chimpanzees. Proceedings of the Royal Society B: Biological Sciences.Charvet, C. J. (2021). Cutting across structural and transcriptomic scales translates time across the lifespan in humans and chimpanzees. Proceedings of the Royal Society B: Biological Sciences. https://doi.org/10.1098/rspb.2020.2987Charvet, C. J. 2021. “Cutting Across Structural and Transcriptomic Scales Translates Time Across the Lifespan in Humans and Chimpanzees”. Proceedings of the Royal Society B: Biological Sciences, ahead of print, February 10. https://doi.org/10.1098/rspb.2020.2987.Charvet, C. J. “Cutting Across Structural and Transcriptomic Scales Translates Time Across the Lifespan in Humans and Chimpanzees”. Proceedings of the Royal Society B: Biological Sciences, Feb. 2021, https://doi.org/10.1098/rspb.2020.2987.Charvet, C. J. Cutting across structural and transcriptomic scales translates time across the lifespan in humans and chimpanzees. Proceedings of the Royal Society B: Biological Sciences https://doi.org/10.1098/rspb.2020.2987 (2021) doi:10.1098/rspb.2020.2987.C. J. Charvet, “Cutting across structural and transcriptomic scales translates time across the lifespan in humans and chimpanzees”, Proceedings of the Royal Society B: Biological Sciences, Feb. 2021, doi: 10.1098/rspb.2020.2987.
- Chollet (2023). François Chollet (@fchollet) on X. X (formerly Twitter).Chollet. (2023, December 16). François Chollet (@fchollet) on X. X (formerly Twitter). https://x.com/fchollet/status/1736079054313574578?s=20Chollet. 2023. “François Chollet (@fchollet) on X”. X (formerly Twitter), December 16. https://x.com/fchollet/status/1736079054313574578?s=20.Chollet. “François Chollet (@fchollet) on X”. X (formerly Twitter), 16 Dec. 2023, https://x.com/fchollet/status/1736079054313574578?s=20.Chollet. François Chollet (@fchollet) on X. X (formerly Twitter) https://x.com/fchollet/status/1736079054313574578?s=20 (2023).Chollet, “François Chollet (@fchollet) on X”, X (formerly Twitter). [Online]. Available: https://x.com/fchollet/status/1736079054313574578?s=20
- Creswell, A., Shanahan, M. & Higgins, I. (2022). Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning. arXiv.Creswell, A., Shanahan, M., & Higgins, I. (2022). Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning. In arXiv. https://arxiv.org/abs/2205.09712Creswell, A., M. Shanahan, and I. Higgins. 2022. “Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning”. In arXiv. Preprint, May 19. https://arxiv.org/abs/2205.09712.Creswell, A., et al. “Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning”. arXiv, 19 May 2022, https://arxiv.org/abs/2205.09712.Creswell, A., Shanahan, M. & Higgins, I. Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning. arXiv Preprint at https://arxiv.org/abs/2205.09712 (2022).A. Creswell, M. Shanahan, and I. Higgins, “Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning”, May 19, 2022. [Online]. Available: https://arxiv.org/abs/2205.09712
- DeepMind (2023). FunSearch: Making new discoveries in mathematical sciences using Large Language Models. Google DeepMind.DeepMind. (2023). FunSearch: Making new discoveries in mathematical sciences using Large Language Models. Google DeepMind. https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-modelsDeepMind. 2023. “FunSearch: Making New Discoveries in Mathematical Sciences Using Large Language Models”. Google DeepMind. https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models.DeepMind. “FunSearch: Making New Discoveries in Mathematical Sciences Using Large Language Models”. Google DeepMind, 2023, https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models.DeepMind. FunSearch: Making new discoveries in mathematical sciences using Large Language Models. Google DeepMind https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models (2023).DeepMind, “FunSearch: Making new discoveries in mathematical sciences using Large Language Models”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models
- Delétang, G. et al. (2022). Neural Networks and the Chomsky Hierarchy. arXiv.Delétang, G., Ruoss, A., Grau-Moya, J., Genewein, T., Wenliang, L. K., Catt, E., Cundy, C., Hutter, M., Legg, S., Veness, J., & Ortega, P. A. (2022). Neural Networks and the Chomsky Hierarchy. In arXiv. https://arxiv.org/abs/2207.02098Delétang, G., A. Ruoss, J. Grau-Moya, et al. 2022. “Neural Networks and the Chomsky Hierarchy”. In arXiv. Preprint, July 5. https://arxiv.org/abs/2207.02098.Delétang, G., et al. “Neural Networks and the Chomsky Hierarchy”. arXiv, 5 July 2022, https://arxiv.org/abs/2207.02098.Delétang, G. et al. Neural Networks and the Chomsky Hierarchy. arXiv Preprint at https://arxiv.org/abs/2207.02098 (2022).G. Delétang et al., “Neural Networks and the Chomsky Hierarchy”, Jul. 05, 2022. [Online]. Available: https://arxiv.org/abs/2207.02098
- Fluri, L., Paleka, D. & Tramèr, F. (2023). Evaluating Superhuman Models with Consistency Checks. arXiv.Fluri, L., Paleka, D., & Tramèr, F. (2023). Evaluating Superhuman Models with Consistency Checks. In arXiv. https://arxiv.org/abs/2306.09983Fluri, L., D. Paleka, and F. Tramèr. 2023. “Evaluating Superhuman Models with Consistency Checks”. In arXiv. Preprint, June 16. https://arxiv.org/abs/2306.09983.Fluri, L., et al. “Evaluating Superhuman Models with Consistency Checks”. arXiv, 16 June 2023, https://arxiv.org/abs/2306.09983.Fluri, L., Paleka, D. & Tramèr, F. Evaluating Superhuman Models with Consistency Checks. arXiv Preprint at https://arxiv.org/abs/2306.09983 (2023).L. Fluri, D. Paleka, and F. Tramèr, “Evaluating Superhuman Models with Consistency Checks”, Jun. 16, 2023. [Online]. Available: https://arxiv.org/abs/2306.09983
- Gurnee, W. & Tegmark, M. (2023). Language Models Represent Space and Time. arXiv.Gurnee, W., & Tegmark, M. (2023). Language Models Represent Space and Time. In arXiv. https://arxiv.org/abs/2310.02207Gurnee, W., and M. Tegmark. 2023. “Language Models Represent Space and Time”. In arXiv. Preprint, October 3. https://arxiv.org/abs/2310.02207.Gurnee, W., and M. Tegmark. “Language Models Represent Space and Time”. arXiv, 3 Oct. 2023, https://arxiv.org/abs/2310.02207.Gurnee, W. & Tegmark, M. Language Models Represent Space and Time. arXiv Preprint at https://arxiv.org/abs/2310.02207 (2023).W. Gurnee and M. Tegmark, “Language Models Represent Space and Time”, Oct. 03, 2023. [Online]. Available: https://arxiv.org/abs/2310.02207
- Herculano-Houzel, S. (2012). The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences.Herculano-Houzel, S. (2012). The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences. https://doi.org/10.1073/pnas.1201895109Herculano-Houzel, S. 2012. “The Remarkable, yet Not Extraordinary, Human Brain as a Scaled-up Primate Brain and Its Associated Cost”. Proceedings of the National Academy of Sciences, ahead of print, June 22. https://doi.org/10.1073/pnas.1201895109.Herculano-Houzel, S. “The Remarkable, yet Not Extraordinary, Human Brain as a Scaled-up Primate Brain and Its Associated Cost”. Proceedings of the National Academy of Sciences, June 2012, https://doi.org/10.1073/pnas.1201895109.Herculano-Houzel, S. The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences https://doi.org/10.1073/pnas.1201895109 (2012) doi:10.1073/pnas.1201895109.S. Herculano-Houzel, “The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost”, Proceedings of the National Academy of Sciences, Jun. 2012, doi: 10.1073/pnas.1201895109.
- jacob_cannell (2015). The Brain as a Universal Learning Machine. LessWrong.jacob_cannell. (2015, June 24). The Brain as a Universal Learning Machine. LessWrong. https://lesswrong.com/posts/9Yc7Pp7szcjPgPsjf/the-brain-as-a-universal-learning-machinejacob_cannell. 2015. “The Brain as a Universal Learning Machine”. LessWrong, June 24. https://lesswrong.com/posts/9Yc7Pp7szcjPgPsjf/the-brain-as-a-universal-learning-machine.jacob_cannell. “The Brain as a Universal Learning Machine”. LessWrong, 24 June 2015, https://lesswrong.com/posts/9Yc7Pp7szcjPgPsjf/the-brain-as-a-universal-learning-machine.jacob_cannell. The Brain as a Universal Learning Machine. LessWrong https://lesswrong.com/posts/9Yc7Pp7szcjPgPsjf/the-brain-as-a-universal-learning-machine (2015).jacob_cannell, “The Brain as a Universal Learning Machine”, LessWrong. [Online]. Available: https://lesswrong.com/posts/9Yc7Pp7szcjPgPsjf/the-brain-as-a-universal-learning-machine
- jacob_cannell (2022). AI Timelines via Cumulative Optimization Power: Less Long, More Short. LessWrong.jacob_cannell. (2022, October 6). AI Timelines via Cumulative Optimization Power: Less Long, More Short. LessWrong. https://lesswrong.com/posts/3nMpdmt8LrzxQnkGp/ai-timelines-via-cumulative-optimization-power-less-longjacob_cannell. 2022. “AI Timelines via Cumulative Optimization Power: Less Long, More Short”. LessWrong, October 6. https://lesswrong.com/posts/3nMpdmt8LrzxQnkGp/ai-timelines-via-cumulative-optimization-power-less-long.jacob_cannell. “AI Timelines via Cumulative Optimization Power: Less Long, More Short”. LessWrong, 6 Oct. 2022, https://lesswrong.com/posts/3nMpdmt8LrzxQnkGp/ai-timelines-via-cumulative-optimization-power-less-long.jacob_cannell. AI Timelines via Cumulative Optimization Power: Less Long, More Short. LessWrong https://lesswrong.com/posts/3nMpdmt8LrzxQnkGp/ai-timelines-via-cumulative-optimization-power-less-long (2022).jacob_cannell, “AI Timelines via Cumulative Optimization Power: Less Long, More Short”, LessWrong. [Online]. Available: https://lesswrong.com/posts/3nMpdmt8LrzxQnkGp/ai-timelines-via-cumulative-optimization-power-less-long
- jacob_cannell (2022). Brain Efficiency: Much More than You Wanted to Know. LessWrong.jacob_cannell. (2022, January 6). Brain Efficiency: Much More than You Wanted to Know. LessWrong. https://lesswrong.com/posts/xwBuoE9p8GE7RAuhd/brain-efficiency-much-more-than-you-wanted-to-knowjacob_cannell. 2022. “Brain Efficiency: Much More Than You Wanted to Know”. LessWrong, January 6. https://lesswrong.com/posts/xwBuoE9p8GE7RAuhd/brain-efficiency-much-more-than-you-wanted-to-know.jacob_cannell. “Brain Efficiency: Much More Than You Wanted to Know”. LessWrong, 6 Jan. 2022, https://lesswrong.com/posts/xwBuoE9p8GE7RAuhd/brain-efficiency-much-more-than-you-wanted-to-know.jacob_cannell. Brain Efficiency: Much More than You Wanted to Know. LessWrong https://lesswrong.com/posts/xwBuoE9p8GE7RAuhd/brain-efficiency-much-more-than-you-wanted-to-know (2022).jacob_cannell, “Brain Efficiency: Much More than You Wanted to Know”, LessWrong. [Online]. Available: https://lesswrong.com/posts/xwBuoE9p8GE7RAuhd/brain-efficiency-much-more-than-you-wanted-to-know
- Kadavath, S. et al. (2022). Language Models (Mostly) Know What They Know. arXiv.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., … Kaplan, J. (2022). Language Models (Mostly) Know What They Know. In arXiv. https://arxiv.org/abs/2207.05221Kadavath, S., T. Conerly, A. Askell, et al. 2022. “Language Models (Mostly) Know What They Know”. In arXiv. Preprint, July 11. https://arxiv.org/abs/2207.05221.Kadavath, S., et al. “Language Models (Mostly) Know What They Know”. arXiv, 11 July 2022, https://arxiv.org/abs/2207.05221.Kadavath, S. et al. Language Models (Mostly) Know What They Know. arXiv Preprint at https://arxiv.org/abs/2207.05221 (2022).S. Kadavath et al., “Language Models (Mostly) Know What They Know”, Jul. 11, 2022. [Online]. Available: https://arxiv.org/abs/2207.05221
- Kaplan, J. et al. (2020). Scaling Laws for Neural Language Models. arXiv.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling Laws for Neural Language Models. In arXiv. https://arxiv.org/abs/2001.08361Kaplan, J., S. McCandlish, T. Henighan, et al. 2020. “Scaling Laws for Neural Language Models”. In arXiv. Preprint, January 23. https://arxiv.org/abs/2001.08361.Kaplan, J., et al. “Scaling Laws for Neural Language Models”. arXiv, 23 Jan. 2020, https://arxiv.org/abs/2001.08361.Kaplan, J. et al. Scaling Laws for Neural Language Models. arXiv Preprint at https://arxiv.org/abs/2001.08361 (2020).J. Kaplan et al., “Scaling Laws for Neural Language Models”, Jan. 23, 2020. [Online]. Available: https://arxiv.org/abs/2001.08361
- Kirillov, A. et al. (2023). Segment Anything. arXiv.Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment Anything. In arXiv. https://arxiv.org/abs/2304.02643Kirillov, A., E. Mintun, N. Ravi, et al. 2023. “Segment Anything”. In arXiv. Preprint, April 5. https://arxiv.org/abs/2304.02643.Kirillov, A., et al. “Segment Anything”. arXiv, 5 Apr. 2023, https://arxiv.org/abs/2304.02643.Kirillov, A. et al. Segment Anything. arXiv Preprint at https://arxiv.org/abs/2304.02643 (2023).A. Kirillov et al., “Segment Anything”, Apr. 05, 2023. [Online]. Available: https://arxiv.org/abs/2304.02643
- Korzekwa (2020). Time for AI to cross the human performance range in ImageNet image classification. AI Impacts.Korzekwa. (2020, October 19). Time for AI to cross the human performance range in ImageNet image classification. AI Impacts. https://aiimpacts.org/time-for-ai-to-cross-the-human-performance-range-in-imagenet-image-classificationKorzekwa. 2020. “Time for AI to Cross the Human Performance Range in ImageNet Image Classification”. AI Impacts, October 19. https://aiimpacts.org/time-for-ai-to-cross-the-human-performance-range-in-imagenet-image-classification.Korzekwa. “Time for AI to Cross the Human Performance Range in ImageNet Image Classification”. AI Impacts, 19 Oct. 2020, https://aiimpacts.org/time-for-ai-to-cross-the-human-performance-range-in-imagenet-image-classification.Korzekwa. Time for AI to cross the human performance range in ImageNet image classification. AI Impacts https://aiimpacts.org/time-for-ai-to-cross-the-human-performance-range-in-imagenet-image-classification (2020).Korzekwa, “Time for AI to cross the human performance range in ImageNet image classification”, AI Impacts. [Online]. Available: https://aiimpacts.org/time-for-ai-to-cross-the-human-performance-range-in-imagenet-image-classification
- Krakovna et al. (2020). Specification gaming: the flip side of AI ingenuity. Google DeepMind.Krakovna et al. (2020). Specification gaming: the flip side of AI ingenuity. Google DeepMind. https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuityKrakovna et al. 2020. “Specification Gaming: The Flip Side of AI Ingenuity”. Google DeepMind. https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity.Krakovna et al. “Specification Gaming: The Flip Side of AI Ingenuity”. Google DeepMind, 2020, https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity.Krakovna et al. Specification gaming: the flip side of AI ingenuity. Google DeepMind https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity (2020).Krakovna et al., “Specification gaming: the flip side of AI ingenuity”, Google DeepMind. [Online]. Available: https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity
- LeCun, Y. (2022). A Path Towards Autonomous Machine Intelligence. OpenReview.LeCun, Y. (2022). A Path Towards Autonomous Machine Intelligence. In OpenReview (Version 0.9.2). https://openreview.net/pdf?id=BZ5a1r-kVsfLeCun, Y. 2022. “A Path Towards Autonomous Machine Intelligence”. In OpenReview, version 0.9.2. Preprint, June 27. https://openreview.net/pdf?id=BZ5a1r-kVsf.LeCun, Y. “A Path Towards Autonomous Machine Intelligence”. OpenReview, Version 0.9.2, 27 June 2022, https://openreview.net/pdf?id=BZ5a1r-kVsf.LeCun, Y. A Path Towards Autonomous Machine Intelligence. OpenReview Preprint at https://openreview.net/pdf?id=BZ5a1r-kVsf (2022).Y. LeCun, “A Path Towards Autonomous Machine Intelligence”, Jun. 27, 2022. [Online]. Available: https://openreview.net/pdf?id=BZ5a1r-kVsf
- Lightman, H. et al. (2023). Let's Verify Step by Step. arXiv.Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., & Cobbe, K. (2023). Let's Verify Step by Step. In arXiv. https://arxiv.org/abs/2305.20050Lightman, H., V. Kosaraju, Y. Burda, et al. 2023. “Let's Verify Step by Step”. In arXiv. Preprint, May 31. https://arxiv.org/abs/2305.20050.Lightman, H., et al. “Let's Verify Step by Step”. arXiv, 31 May 2023, https://arxiv.org/abs/2305.20050.Lightman, H. et al. Let's Verify Step by Step. arXiv Preprint at https://arxiv.org/abs/2305.20050 (2023).H. Lightman et al., “Let's Verify Step by Step”, May 31, 2023. [Online]. Available: https://arxiv.org/abs/2305.20050
- Maynez, J., Narayan, S., Bohnet, B. & McDonald, R. (2020). On Faithfulness and Factuality in Abstractive Summarization. arXiv.Maynez, J., Narayan, S., Bohnet, B., & McDonald, R. (2020). On Faithfulness and Factuality in Abstractive Summarization. In arXiv. https://arxiv.org/abs/2005.00661Maynez, J., S. Narayan, B. Bohnet, and R. McDonald. 2020. “On Faithfulness and Factuality in Abstractive Summarization”. In arXiv. Preprint, May 2. https://arxiv.org/abs/2005.00661.Maynez, J., et al. “On Faithfulness and Factuality in Abstractive Summarization”. arXiv, 2 May 2020, https://arxiv.org/abs/2005.00661.Maynez, J., Narayan, S., Bohnet, B. & McDonald, R. On Faithfulness and Factuality in Abstractive Summarization. arXiv Preprint at https://arxiv.org/abs/2005.00661 (2020).J. Maynez, S. Narayan, B. Bohnet, and R. McDonald, “On Faithfulness and Factuality in Abstractive Summarization”, May 02, 2020. [Online]. Available: https://arxiv.org/abs/2005.00661
- McGrath, T. et al. (2021). Acquisition of Chess Knowledge in AlphaZero. arXiv.McGrath, T., Kapishnikov, A., Tomašev, N., Pearce, A., Hassabis, D., Kim, B., Paquet, U., & Kramnik, V. (2021). Acquisition of Chess Knowledge in AlphaZero. In arXiv. https://doi.org/10.1073/pnas.2206625119McGrath, T., A. Kapishnikov, N. Tomašev, et al. 2021. “Acquisition of Chess Knowledge in AlphaZero”. In arXiv. Preprint, November 17. https://doi.org/10.1073/pnas.2206625119.McGrath, T., et al. “Acquisition of Chess Knowledge in AlphaZero”. arXiv, 17 Nov. 2021, https://doi.org/10.1073/pnas.2206625119.McGrath, T. et al. Acquisition of Chess Knowledge in AlphaZero. arXiv Preprint at https://doi.org/10.1073/pnas.2206625119 (2021).T. McGrath et al., “Acquisition of Chess Knowledge in AlphaZero”, Nov. 17, 2021. doi: 10.1073/pnas.2206625119.
- Meng, K., Bau, D., Andonian, A. & Belinkov, Y. (2022). Locating and Editing Factual Associations in GPT. arXiv.Meng, K., Bau, D., Andonian, A., & Belinkov, Y. (2022). Locating and Editing Factual Associations in GPT. In arXiv. https://arxiv.org/abs/2202.05262Meng, K., D. Bau, A. Andonian, and Y. Belinkov. 2022. “Locating and Editing Factual Associations in GPT”. In arXiv. Preprint, February 10. https://arxiv.org/abs/2202.05262.Meng, K., et al. “Locating and Editing Factual Associations in GPT”. arXiv, 10 Feb. 2022, https://arxiv.org/abs/2202.05262.Meng, K., Bau, D., Andonian, A. & Belinkov, Y. Locating and Editing Factual Associations in GPT. arXiv Preprint at https://arxiv.org/abs/2202.05262 (2022).K. Meng, D. Bau, A. Andonian, and Y. Belinkov, “Locating and Editing Factual Associations in GPT”, Feb. 10, 2022. [Online]. Available: https://arxiv.org/abs/2202.05262
- Nanda, N., Lee, A. & Wattenberg, M. (2023). Emergent Linear Representations in World Models of Self-Supervised Sequence Models. arXiv.Nanda, N., Lee, A., & Wattenberg, M. (2023). Emergent Linear Representations in World Models of Self-Supervised Sequence Models. In arXiv. https://arxiv.org/abs/2309.00941Nanda, N., A. Lee, and M. Wattenberg. 2023. “Emergent Linear Representations in World Models of Self-Supervised Sequence Models”. In arXiv. Preprint, September 2. https://arxiv.org/abs/2309.00941.Nanda, N., et al. “Emergent Linear Representations in World Models of Self-Supervised Sequence Models”. arXiv, 2 Sept. 2023, https://arxiv.org/abs/2309.00941.Nanda, N., Lee, A. & Wattenberg, M. Emergent Linear Representations in World Models of Self-Supervised Sequence Models. arXiv Preprint at https://arxiv.org/abs/2309.00941 (2023).N. Nanda, A. Lee, and M. Wattenberg, “Emergent Linear Representations in World Models of Self-Supervised Sequence Models”, Sep. 02, 2023. [Online]. Available: https://arxiv.org/abs/2309.00941
- Pearson, A., Bruner, E. & Polly, P. D. (2023). Updated imaging and phylogenetic comparative methods reassess relative temporal lobe size in anthropoids and modern humans. American Journal of Biological Anthropology.Pearson, A., Bruner, E., & Polly, P. D. (2023). Updated imaging and phylogenetic comparative methods reassess relative temporal lobe size in anthropoids and modern humans. American Journal of Biological Anthropology. https://doi.org/10.1002/ajpa.24712Pearson, A., E. Bruner, and P. D. Polly. 2023. “Updated Imaging and Phylogenetic Comparative Methods Reassess Relative Temporal Lobe Size in Anthropoids and Modern Humans”. American Journal of Biological Anthropology, ahead of print, February 15. https://doi.org/10.1002/ajpa.24712.Pearson, A., et al. “Updated Imaging and Phylogenetic Comparative Methods Reassess Relative Temporal Lobe Size in Anthropoids and Modern Humans”. American Journal of Biological Anthropology, Feb. 2023, https://doi.org/10.1002/ajpa.24712.Pearson, A., Bruner, E. & Polly, P. D. Updated imaging and phylogenetic comparative methods reassess relative temporal lobe size in anthropoids and modern humans. American Journal of Biological Anthropology https://doi.org/10.1002/ajpa.24712 (2023) doi:10.1002/ajpa.24712.A. Pearson, E. Bruner, and P. D. Polly, “Updated imaging and phylogenetic comparative methods reassess relative temporal lobe size in anthropoids and modern humans”, American Journal of Biological Anthropology, Feb. 2023, doi: 10.1002/ajpa.24712.
- Quentin FEUILLADE--MONTIXI & Pierre Peigné (2023). The Stochastic Parrot Hypothesis is debatable for the last generation of LLMs. LessWrong.Quentin FEUILLADE--MONTIXI, & Pierre Peigné. (2023, November 7). The Stochastic Parrot Hypothesis is debatable for the last generation of LLMs. LessWrong. https://lesswrong.com/posts/HxRjHq3QG8vcYy4yy/the-stochastic-parrot-hypothesis-is-debatable-for-the-lastQuentin FEUILLADE--MONTIXI, and Pierre Peigné. 2023. “The Stochastic Parrot Hypothesis Is Debatable for the Last Generation of LLMs”. LessWrong, November 7. https://lesswrong.com/posts/HxRjHq3QG8vcYy4yy/the-stochastic-parrot-hypothesis-is-debatable-for-the-last.Quentin FEUILLADE--MONTIXI, and Pierre Peigné. “The Stochastic Parrot Hypothesis Is Debatable for the Last Generation of LLMs”. LessWrong, 7 Nov. 2023, https://lesswrong.com/posts/HxRjHq3QG8vcYy4yy/the-stochastic-parrot-hypothesis-is-debatable-for-the-last.Quentin FEUILLADE--MONTIXI & Pierre Peigné. The Stochastic Parrot Hypothesis is debatable for the last generation of LLMs. LessWrong https://lesswrong.com/posts/HxRjHq3QG8vcYy4yy/the-stochastic-parrot-hypothesis-is-debatable-for-the-last (2023).Quentin FEUILLADE--MONTIXI and Pierre Peigné, “The Stochastic Parrot Hypothesis is debatable for the last generation of LLMs”, LessWrong. [Online]. Available: https://lesswrong.com/posts/HxRjHq3QG8vcYy4yy/the-stochastic-parrot-hypothesis-is-debatable-for-the-last
- Rae, J. W. et al. (2021). Scaling Language Models: Methods, Analysis & Insights from Training Gopher. arXiv.Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., … Irving, G. (2021). Scaling Language Models: Methods, Analysis & Insights from Training Gopher. In arXiv. https://arxiv.org/abs/2112.11446Rae, J. W., S. Borgeaud, T. Cai, et al. 2021. “Scaling Language Models: Methods, Analysis & Insights from Training Gopher”. In arXiv. Preprint, December 8. https://arxiv.org/abs/2112.11446.Rae, J. W., et al. “Scaling Language Models: Methods, Analysis & Insights from Training Gopher”. arXiv, 8 Dec. 2021, https://arxiv.org/abs/2112.11446.Rae, J. W. et al. Scaling Language Models: Methods, Analysis & Insights from Training Gopher. arXiv Preprint at https://arxiv.org/abs/2112.11446 (2021).J. W. Rae et al., “Scaling Language Models: Methods, Analysis & Insights from Training Gopher”, Dec. 08, 2021. [Online]. Available: https://arxiv.org/abs/2112.11446
- Schrimpf et al. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. PNAS.Schrimpf et al. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. Internet Archive (https://web.archive.org/web/20220221135833/https://www.pnas.org/content/118/45/e2105646118). PNAS. https://pnas.org/content/118/45/e2105646118Schrimpf et al. 2021. “The Neural Architecture of Language: Integrative Modeling Converges on Predictive Processing”. PNAS. Https://web.archive.org/web/20220221135833/https://www.pnas.org/content/118/45/e2105646118. Internet Archive. https://pnas.org/content/118/45/e2105646118.Schrimpf et al. “The Neural Architecture of Language: Integrative Modeling Converges on Predictive Processing”. PNAS, 2021, Internet Archive, https://web.archive.org/web/20220221135833/https://www.pnas.org/content/118/45/e2105646118, https://pnas.org/content/118/45/e2105646118.Schrimpf et al. The neural architecture of language: Integrative modeling converges on predictive processing. PNAS https://pnas.org/content/118/45/e2105646118 (2021).Schrimpf et al., “The neural architecture of language: Integrative modeling converges on predictive processing”, PNAS. Accessed: Feb. 21, 2022. [Online]. Available: https://pnas.org/content/118/45/e2105646118
- Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K. & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv.Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. In arXiv. https://arxiv.org/abs/2303.11366Shinn, N., F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao. 2023. “Reflexion: Language Agents with Verbal Reinforcement Learning”. In arXiv. Preprint, March 20. https://arxiv.org/abs/2303.11366.Shinn, N., et al. “Reflexion: Language Agents with Verbal Reinforcement Learning”. arXiv, 20 Mar. 2023, https://arxiv.org/abs/2303.11366.Shinn, N. et al. Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv Preprint at https://arxiv.org/abs/2303.11366 (2023).N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language Agents with Verbal Reinforcement Learning”, Mar. 20, 2023. [Online]. Available: https://arxiv.org/abs/2303.11366
- Stephanie Lin, Jacob Hilton & Owain Evans (2021). TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv.Stephanie Lin, Jacob Hilton, & Owain Evans. (2021). TruthfulQA: Measuring How Models Mimic Human Falsehoods. In arXiv. https://arxiv.org/abs/2109.07958Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. “TruthfulQA: Measuring How Models Mimic Human Falsehoods”. In arXiv. Preprint, September 8. https://arxiv.org/abs/2109.07958.Stephanie Lin, et al. “TruthfulQA: Measuring How Models Mimic Human Falsehoods”. arXiv, 8 Sept. 2021, https://arxiv.org/abs/2109.07958.Stephanie Lin, Jacob Hilton & Owain Evans. TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv Preprint at https://arxiv.org/abs/2109.07958 (2021).Stephanie Lin, Jacob Hilton, and Owain Evans, “TruthfulQA: Measuring How Models Mimic Human Falsehoods”, Sep. 08, 2021. [Online]. Available: https://arxiv.org/abs/2109.07958
- Tian, K., Mitchell, E., Yao, H., Manning, C. D. & Finn, C. (2023). Fine-tuning Language Models for Factuality. arXiv.Tian, K., Mitchell, E., Yao, H., Manning, C. D., & Finn, C. (2023). Fine-tuning Language Models for Factuality. In arXiv. https://arxiv.org/abs/2311.08401Tian, K., E. Mitchell, H. Yao, C. D. Manning, and C. Finn. 2023. “Fine-tuning Language Models for Factuality”. In arXiv. Preprint, November 14. https://arxiv.org/abs/2311.08401.Tian, K., et al. “Fine-tuning Language Models for Factuality”. arXiv, 14 Nov. 2023, https://arxiv.org/abs/2311.08401.Tian, K., Mitchell, E., Yao, H., Manning, C. D. & Finn, C. Fine-tuning Language Models for Factuality. arXiv Preprint at https://arxiv.org/abs/2311.08401 (2023).K. Tian, E. Mitchell, H. Yao, C. D. Manning, and C. Finn, “Fine-tuning Language Models for Factuality”, Nov. 14, 2023. [Online]. Available: https://arxiv.org/abs/2311.08401
- Wang et al. (2022). Adversarial Policies Beat Superhuman Go AIs. arXiv.org.Wang et al. (2022). Adversarial Policies Beat Superhuman Go AIs. arXiv.org. https://www.arxiv.org/abs/2211.00241Wang et al. 2022. “Adversarial Policies Beat Superhuman Go AIs”. arXiv.org. https://www.arxiv.org/abs/2211.00241.Wang et al. “Adversarial Policies Beat Superhuman Go AIs”. arXiv.org, 2022, https://www.arxiv.org/abs/2211.00241.Wang et al. Adversarial Policies Beat Superhuman Go AIs. arXiv.org https://www.arxiv.org/abs/2211.00241 (2022).Wang et al., “Adversarial Policies Beat Superhuman Go AIs”, arXiv.org. [Online]. Available: https://www.arxiv.org/abs/2211.00241
- Wang, G. et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. In arXiv. https://arxiv.org/abs/2305.16291Wang, G., Y. Xie, Y. Jiang, et al. 2023. “Voyager: An Open-Ended Embodied Agent with Large Language Models”. In arXiv. Preprint, May 25. https://arxiv.org/abs/2305.16291.Wang, G., et al. “Voyager: An Open-Ended Embodied Agent with Large Language Models”. arXiv, 25 May 2023, https://arxiv.org/abs/2305.16291.Wang, G. et al. Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv Preprint at https://arxiv.org/abs/2305.16291 (2023).G. Wang et al., “Voyager: An Open-Ended Embodied Agent with Large Language Models”, May 25, 2023. [Online]. Available: https://arxiv.org/abs/2305.16291
- Wang, S., Liu, S., Ye, W., You, J. & Gao, Y. (2024). EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data. arXiv.Wang, S., Liu, S., Ye, W., You, J., & Gao, Y. (2024). EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data. In arXiv. https://arxiv.org/abs/2403.00564Wang, S., S. Liu, W. Ye, J. You, and Y. Gao. 2024. “EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data”. In arXiv. Preprint, March 1. https://arxiv.org/abs/2403.00564.Wang, S., et al. “EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data”. arXiv, 1 Mar. 2024, https://arxiv.org/abs/2403.00564.Wang, S., Liu, S., Ye, W., You, J. & Gao, Y. EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data. arXiv Preprint at https://arxiv.org/abs/2403.00564 (2024).S. Wang, S. Liu, W. Ye, J. You, and Y. Gao, “EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data”, Mar. 01, 2024. [Online]. Available: https://arxiv.org/abs/2403.00564
- Wikipedia (2022). AI-complete. Wikipedia.Wikipedia. (2022). AI-complete. Wikipedia. https://en.wikipedia.org/wiki/AI-completeWikipedia. 2022. “AI-complete”. Wikipedia. https://en.wikipedia.org/wiki/AI-complete.Wikipedia. “AI-complete”. Wikipedia, 2022, https://en.wikipedia.org/wiki/AI-complete.Wikipedia. AI-complete. Wikipedia https://en.wikipedia.org/wiki/AI-complete (2022).Wikipedia, “AI-complete”, Wikipedia. [Online]. Available: https://en.wikipedia.org/wiki/AI-complete
- Ye, W., Liu, S., Kurutach, T., Abbeel, P. & Gao, Y. (2021). Mastering Atari Games with Limited Data. arXiv.Ye, W., Liu, S., Kurutach, T., Abbeel, P., & Gao, Y. (2021). Mastering Atari Games with Limited Data. In arXiv. https://arxiv.org/abs/2111.00210Ye, W., S. Liu, T. Kurutach, P. Abbeel, and Y. Gao. 2021. “Mastering Atari Games with Limited Data”. In arXiv. Preprint, October 30. https://arxiv.org/abs/2111.00210.Ye, W., et al. “Mastering Atari Games with Limited Data”. arXiv, 30 Oct. 2021, https://arxiv.org/abs/2111.00210.Ye, W., Liu, S., Kurutach, T., Abbeel, P. & Gao, Y. Mastering Atari Games with Limited Data. arXiv Preprint at https://arxiv.org/abs/2111.00210 (2021).W. Ye, S. Liu, T. Kurutach, P. Abbeel, and Y. Gao, “Mastering Atari Games with Limited Data”, Oct. 30, 2021. [Online]. Available: https://arxiv.org/abs/2111.00210
Was this section useful?
Thank you for your feedback
Your input helps improve the Atlas.