Artificial general intelligence is a hypothetical artificial system whose competence spans essentially the full range of cognitive tasks humans perform, rather than being confined to a domain it was built for. The term functions mainly by contrast with narrow AI: a chess engine, a protein-structure predictor and a fraud-detection model are each superhuman within their scope and useless outside it. Whether any existing system has begun to cross that boundary is one of the most consequential open disputes in science, and it is unresolved partly because the boundary has never been defined in a way that would settle it.
Defining the target
There is no accepted definition. Legg and Hutter, surveying the problem, assembled more than seventy published definitions of intelligence and extracted a common core: intelligence measures an agent's ability to achieve goals in a wide range of environments.1 That formulation is precise enough to be formalised and too abstract to adjudicate any actual system.
Working definitions split into three families. Task-based definitions specify jobs a system must do — the Turing test is the ancestor, and modern versions substitute economically valuable work or a benchmark suite. Capability-based definitions specify cognitive functions such as transfer learning, long-horizon planning, or acquiring new skills without retraining. Autonomy-based definitions require operation over extended periods without human scaffolding. Nick Bostrom sidesteps the threshold in Superintelligence, defining superintelligence as intellect greatly exceeding human performance in virtually all domains and treating human-level capability as a waypoint.
A framework from Google DeepMind researchers replaces the binary with a grid, scoring systems on performance (from "emerging" through "competent", "expert", "virtuoso" and "superhuman") crossed with generality.2 On that scheme, mid-2020s systems are plausibly "emerging AGI" and clearly superhuman narrow AI in several domains — more informative than either "AGI is here" or "AGI is not here".
Terminology"AGI" is used inconsistently even within single organisations: human-level competence across tasks, full automation of remote work, or the ability to do novel science. These are different claims with different timelines, and arguments equivocate between them.
Origins of the term
Alan Turing's 1950 paper posed machine thinking in behavioural terms and proposed the imitation game to sidestep definitional argument.3 The 1956 Dartmouth proposal assumed generality as the default goal, expecting machines to use language, form abstractions and improve themselves within a summer.
The narrower ambitions of expert systems and, later, statistical machine learning displaced that goal for decades. The phrase "artificial general intelligence" appears in a 1997 paper by Mark Gubrud on the security implications of advanced technology, and was independently adopted in the early 2000s by Shane Legg, Ben Goertzel and Peter Voss, becoming established with a 2007 edited volume of that title. The phrase then distinguished a small community from a mainstream that regarded the goal as disreputable; several of the largest laboratories now state it as their objective.
Measuring progress
Benchmark performance is the field's main evidence, and it has an unusual failure mode: benchmarks saturate faster than they can be built. Knowledge tests, grade-school mathematics sets and function-level coding tests that were discriminating in the early 2020s reached ceiling within a few years. Their harder successors — graduate-level science, research mathematics, adversarially collected expert questions, repository-scale software tasks — have compressed in turn.
Saturation is weaker evidence than it appears. Contamination is hard to exclude when training corpora are web-scale and benchmarks are published; benchmarks measure what is easy to score, biasing them toward short checkable answers; and Goodhart's law applies with force when the benchmark is also the optimisation target for the labs reporting the scores.
The most-cited contamination-resistant test is François Chollet's Abstraction and Reasoning Corpus, built around tasks requiring a novel rule to be inferred from a handful of examples.4 It resisted large language models for years. In late 2024 a preview of a model trained to spend far more computation at inference time reported private-set scores in the range of typical human performance, at costs orders of magnitude above ordinary use; a harder successor set was released in 2025 on which scores fell sharply again. The gap closes, then is redefined, and which of those moves is more informative remains unclear.
The case for nearness
Three arguments carry most of the weight.
Scaling. Empirical relationships between model scale, data and loss have held over several orders of magnitude, and frontier training compute grew at roughly six-month doubling times through the 2010s and early 2020s.5 If capability tracks loss and loss tracks compute, extrapolation implies continued rapid gains. The inferential structure is the one examined in Accelerating change, and its best-known long-range form is Ray Kurzweil's curve-fitting. The addition of inference-time computation as a second scaling axis from 2024 onward reopened headroom that pretraining scaling alone appeared to be exhausting.
Generality already observed. A single model trained on next-token prediction can write code, translate, summarise clinical notes, prove competition mathematics and control a browser. That breadth was not designed in, and it is the strongest evidence against the older view that generality requires explicit architecture — the position Richard Sutton called the bitter lesson.
Expert opinion has moved. Large surveys of published AI researchers have shifted median estimates for human-level machine intelligence substantially earlier across successive editions, with the aggregate 50% point falling around mid-century in the most recent.6 Such forecasts have a poor track record in both directions, but the direction of movement is itself data.
The case against
Jaggedness. Capability is not a scalar. Systems that solve olympiad problems fail at tasks a child handles, and the failures are not obviously less "general" than the successes. Studies of professional use find a frontier uneven in ways users cannot predict — a different situation from a uniformly sub-human system approaching parity.7
Reliability rather than capability is the binding constraint. A system correct 95% of the time is not 95% of the way to automating a task if errors are uncorrelated with confidence and expensive to detect. Long-horizon agentic work compounds this: per-step reliability must be very high for hundred-step tasks to succeed.
Missing components. Critics including Gary Marcus and Melanie Mitchell argue that continual learning, causal models, compositional generalisation and grounded understanding are absent rather than underdeveloped, and that benchmark progress measures interpolation over a vast training distribution.8 Moravec's observation still holds that sensorimotor competence, not abstract reasoning, is the hard part, and robotics has advanced far more slowly than text models.
The definition problem cuts both ways. If no test would convince skeptics and no failure would convince proponents, the disagreement is not empirical, and several participants on both sides have said so.
Overhang and discontinuity
The capability overhang argument holds that deployed capability lags latent capability. A trained model's abilities are elicited by prompting, scaffolding, tool access and fine-tuning, none of which requires new compute, and jumps in agentic performance from scaffolding alone support the claim's general shape. Its policy relevance is that it undermines the assumption of a smooth, observable approach to any threshold: if a large gap exists between what systems can do and what they are seen doing, warning time is shorter than deployment curves suggest. That premise underwrites both proposals to sequence safety work ahead of capability work and precautionary pre-deployment evaluation regimes.
The related compute overhang argument, prominent in Existential risk discussion, holds that pausing capability research accumulates hardware enabling a faster jump when it resumes. Both are structural arguments, and neither has been quantified convincingly.
Why AGI timelines matter here
Almost every long-horizon claim on this wiki has an AI term in it. Rejuvenation biology is bottlenecked on a system with more interacting variables than experiments can resolve, and protein structure prediction has already changed what is tractable there; companies such as Retro Biosciences have made model-assisted protein engineering an explicit programme. The limit of that transfer is visible in AI drug discovery, where models now generate candidate molecules far faster than medicinal chemists did and no drug produced that way has been approved. Connectomics and Whole brain emulation are limited by segmentation and simulation, both compute-bound. Neural decoding improves with better sequence models. Conversely, Dual-use research of concern risk rises if design tools lower the expertise barrier for pathogens.
The Technological singularity argument depends entirely on this article's subject: an intelligence explosion requires systems capable of improving themselves, which is a strictly stronger condition than AGI. Intelligence amplification and Human–AI merger describe the alternative in which capability grows through human–machine teams rather than autonomous systems, and Mind uploading would require both AGI-scale compute and neuroscience that does not exist. Whether such a system would be a subject as well as an agent is a separate question, treated in Machine consciousness; the Substrate independence premise it shares with most Posthuman forecasting is assumed far more often than it is argued.
Forecasts are not findingsNo timeline for AGI is well supported. Survey medians, scaling extrapolations and insider statements are all weak evidence, and the field's forecasting record — including confident predictions of imminence in the 1960s and of impossibility in the 1990s — should discount any specific date, including recent ones.
Outlook
Governance has moved faster than definition. The EU's AI Act imposes obligations on general-purpose models by capability thresholds measured in training compute, several states have established safety institutes with evaluation mandates, and an international expert report on advanced AI risk was published in 2025. Every one of these instruments must operationalise "general" and "capable" without a scientific definition of either, and they have done so with proxies — training FLOP, benchmark scores, red-team results — that are known to be poor. Whether a capability threshold can be specified precisely enough to regulate before anyone can specify it precisely enough to measure is the practical form of the definitional problem, and it now binds.
See also
- Technological singularity
- Machine consciousness
- Whole brain emulation
- Human–AI merger
- Existential risk
- Differential technological development
- Accelerating change
- Future of humanity
References
Footnotes
-
paperLegg, S. and Hutter, M. "Universal Intelligence: A Definition of Machine Intelligence." Minds and Machines, 2007. ↩
-
paperMorris, M. R. et al. "Levels of AGI for Operationalizing Progress on the Path to AGI." Proceedings of the International Conference on Machine Learning, 2024. ↩
-
paperTuring, A. M. "Computing Machinery and Intelligence." Mind, 1950. ↩
-
preprintChollet, F. "On the Measure of Intelligence." arXiv preprint, 2019. ↩
-
paperSevilla, J. et al. "Compute Trends Across Three Eras of Machine Learning." International Joint Conference on Neural Networks, 2022. ↩
-
preprintGrace, K. et al. "Thousands of AI Authors on the Future of AI." arXiv preprint, 2024.↩A survey of authors who published at major machine-learning venues; it records opinion, and successive editions of the series have moved sharply.
-
preprintDell'Acqua, F. et al. "Navigating the Jagged Technological Frontier." Harvard Business School working paper, 2023.↩A working paper reporting a field experiment with consultants at one firm, so the frontier it maps is specific to that task set.
-
preprintMitchell, M. "Why AI is Harder Than We Think." arXiv preprint, 2021. ↩