Rich Bowen
Working essay · October 2026

What Would Count as AGI?

World peace, medical breakthroughs, human work, and the problem of defining general intelligence

Rich Bowen · Intelligence, capability & human judgment

The argument

Artificial general intelligence cannot be defined by one dramatic achievement or by economic disruption alone. World peace is not simply an intelligence problem, medical breakthroughs can come from narrow systems, and replacing workers depends on institutions as well as capability. A useful definition must combine breadth, depth, learning, reliability, and autonomy across unfamiliar tasks.

A finish line that moves

Artificial general intelligence is discussed as though everyone is waiting for the same event. One laboratory will announce it. A benchmark will be crossed. A system will solve a problem that humans could not, and the argument will end.

But AGI has no agreed finish line. OpenAI's charter defines it economically, as highly autonomous systems that outperform humans at most economically valuable work. Google DeepMind researchers have proposed levels based on both the breadth of tasks a system can perform and the depth of its performance. More recent work has tried to divide general intelligence into measurable cognitive abilities rather than treating it as a single score. [1] [2] [3]

These definitions overlap, but they are not interchangeable. A system could transform an industry without understanding much outside it. Another could reason across many subjects while remaining unreliable in long tasks. A third could possess impressive capabilities but require constant human direction.

Before asking when AGI will arrive, we should decide what kind of claim the term is supposed to make.

World peace is not a benchmark

One possible test is civilizational: call a system generally intelligent when it solves a problem such as war, poverty, or climate change. The appeal is obvious. Intelligence should produce outcomes that matter.

World peace, however, is not an equation waiting for a sufficiently clever answer. Conflict involves incompatible interests, identity, fear, power, history, resources, and moral disagreement. An AI might model negotiations, identify possible agreements, or forecast consequences better than any human team. It could still be unable to make governments or populations accept the result.

If leaders reject a workable peace plan, that does not show that the system lacks intelligence. If they follow its plan, the outcome may reflect political authority and trust as much as reasoning. Success depends on humans and institutions that sit outside the model.

World peace would be evidence of extraordinary influence. It would not be a clean test of general intelligence.

Breakthroughs can be narrow

Scientific and medical breakthroughs offer a second candidate. An AI that discovers a treatment, explains a disease mechanism, or resolves a major scientific problem may appear to have crossed into general intelligence.

Such achievements deserve to be taken seriously, but a breakthrough can result from narrow excellence. AlphaFold transformed protein structure prediction and gave researchers a tool of enormous scientific value. Its success demonstrated exceptional capability in a defined domain, not competence across the full range of human thought. [4]

The distinction is easy to miss because narrow systems can exceed every human in their specialty. A chess engine is superhuman at chess. A protein model may solve a challenge that occupied biologists for decades. Neither result, by itself, shows that the system could learn constitutional law, repair a plumbing failure, plan an unfamiliar experiment, or recognize that its instructions are based on a mistaken assumption.

A medical breakthrough may be more valuable than general intelligence. Value and generality are different properties.

Work is an incomplete measure

Economic definitions have a practical advantage. Work contains thousands of tasks that can be observed, compared, and priced. If a system can perform most of them at or above human level, calling it general seems reasonable.

Yet replacing human activity is not the same as reproducing human intelligence. Jobs are bundles of tasks shaped by wages, regulation, liability, customer preference, and organizational design. A capable system may not replace a worker if deployment is expensive or legally restricted. A less capable system may replace one if a company redesigns the job around what the system can do.

Automation also differs from augmentation. Current usage data already distinguishes cases in which AI completes work with little input from cases in which people and AI collaborate. The balance varies by occupation and task. [5]

An economy could be transformed before any system becomes fully general. Conversely, a broadly capable intelligence might be deliberately prevented from replacing people. Labor impact tells us what institutions did with a technology, not only what the technology understood.

The dimensions that matter

A better definition begins with several dimensions rather than one symbolic achievement.

Breadth asks whether the system can operate across domains rather than within a carefully bounded specialty. It should move among language, quantitative reasoning, planning, interpretation, and practical problem solving.

Depth asks how well it performs. Shallow familiarity with many subjects is not enough. General intelligence should include skilled performance on a meaningful range of difficult tasks.

Learning and adaptation ask whether the system can acquire a new skill, recognize unfamiliar conditions, and transfer what it learned without being rebuilt for every case.

Reliability asks whether success persists beyond a demonstration. A system that produces brilliance alongside confident fabrication cannot safely carry broad responsibility.

Autonomy asks how much direction is required. Can it formulate intermediate steps, notice failure, seek missing information, and recover? Autonomy should not be confused with wisdom. It measures independence of action, not the quality of the goals pursued.

Grounding asks whether the system can connect abstract reasoning to the world in which consequences occur. This need not require a human-shaped body, but it does require feedback, context, and an ability to distinguish a plausible description from a successful action.

A working definition

I would call a system AGI when it can reliably learn and perform a broad range of consequential cognitive tasks, across unfamiliar domains, at the level of a skilled human, with enough autonomy to pursue multi-step goals without task-specific retraining or continuous human correction.

Every phrase is doing work.

“Reliably” excludes isolated demonstrations. “Learn and perform” requires adaptation rather than stored fluency alone. “Broad range” separates general intelligence from a collection of narrow triumphs. “Consequential” prevents trivial games from dominating the measurement. “Unfamiliar domains” tests transfer. “Skilled human” sets a demanding but intelligible comparison. “Enough autonomy” recognizes that intelligence becomes socially important when it can sustain action over time.

This definition describes a threshold, but AGI would still come in degrees. Breadth and depth could continue to increase. A system might be generally capable at a competent level before becoming expert or superhuman across most domains. The language of levels is more useful than pretending that one morning the world contains no AGI and that afternoon it does.

What AGI would not prove

Even a system meeting this definition would not settle every larger question attached to AGI.

It would not prove consciousness. Capability and subjective experience are different claims. It would not prove moral wisdom. Intelligence can identify means without supplying ends. It would not prove benevolence, political legitimacy, or a right to make decisions for humanity.

It would also not guarantee world-changing outcomes. Scientific progress depends on laboratories, data, institutions, and implementation. Peace depends on consent and power. Economic abundance depends on how resources and gains are distributed.

Calling a system AGI should describe what it can do. It should not quietly authorize what it is allowed to do.

Why the definition matters

Definitions determine what companies announce, what investors reward, what governments regulate, and what the public fears. If AGI means economic replacement, labor disruption becomes the central evidence. If it means broad cognition, evaluation must test transfer, learning, and reliability. If it means world-changing benefit, the term may never be reached because social problems do not yield to intelligence alone.

A vague definition is useful for mythology. It lets every new system be described as nearly general and every failure dismissed as one more missing feature. A measurable definition makes the claim accountable.

The question is not whether an AI can produce one miracle. It is whether the same system can enter unfamiliar territory, learn what matters, perform competently, understand when it is failing, and continue without a human quietly doing the general part on its behalf.

Medical breakthroughs would be evidence of power. Replacing human tasks would be evidence of economic reach. Contributing to peace would be evidence of influence and perhaps judgment. AGI would require something more basic beneath all three: a durable ability to learn and reason across the boundaries that normally separate one kind of intelligence from another.

Calling a system AGI should describe what it can do. It should not quietly authorize what it is allowed to do.

Sources & note

This essay proposes a working definition for discussion. AGI remains a disputed concept, and no single framework or benchmark has established a universally accepted threshold.

  1. OpenAI. OpenAI Charter.
  2. Morris, M. R., et al. (2023). Levels of AGI for Operationalizing Progress on the Path to AGI.
  3. Google DeepMind (2026). Measuring progress toward AGI: A cognitive framework.
  4. Jumper, J., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583-589.
  5. Anthropic (2025). Economic Index: AI's role in the U.S. and global economy.

Read The Artificial Divine →

← All essays