Artificial Intelligence

Is Artificial General Intelligence possible?

The honest answer is that nobody knows, and the reasons people give for their position are more interesting than the position. Here are the four camps, and what would move each one.

Four positions on artificial general intelligence plotted by what each treats as the main obstacle
The disagreement is about the obstacle, not the date. Positions are placed by what each treats as the binding constraint.
In short

There is no scientific consensus, and no agreed test that would settle it. The four serious positions disagree less about the timeline than about the obstacle: whether more compute and data are sufficient, whether a capability such as causal reasoning is missing by design, whether intelligence requires acting on a physical world, or whether the question is too ill-defined to have an answer.

Key takeaways
  • There is no agreed test. Without one, "is AGI close?" cannot be settled by evidence, which is why the debate does not converge.
  • The camps disagree about the obstacle, not the timeline. Scaling, a missing ingredient, embodiment, or definitional confusion.
  • Benchmark saturation is weak evidence. Once a test is popular enough to optimise against, it stops measuring generalisation.
  • The useful question is falsifiability. Not which camp you are in, but which observation would move you out of it.

Why the question resists an answer

Most scientific questions can in principle be settled by an observation. "Is AGI possible?" currently cannot, and the reason is not that the evidence is thin. It is that the term has no operational definition the field agrees on.

"General" is doing all the work and is never pinned down. General across which tasks? At what level — median human, expert human, any human who has ever lived? Under what constraints of time, energy and supervision? Different answers make the question trivially true, obviously false, or unanswerable, and participants in the argument frequently hold different answers without noticing.

So the productive move is not to answer the question. It is to identify what each serious position treats as the binding constraint, and what would count against it.

Position one: scale is sufficient

The strongest empirical case. Over roughly a decade, capability on a wide range of tasks has improved smoothly and predictably with training compute, dataset size and parameter count. The relationships are regular enough to be fitted and extrapolated. Capabilities that were absent at one scale — multi-step arithmetic, translation between language pairs never explicitly paired in training — appeared at larger ones without any change of method.

The argument is straightforward: the trend has held across several orders of magnitude and every predicted wall has been passed. Absent a specific reason to expect a discontinuity, expect it to continue.

What would refute it: a capability that stays flat across two further orders of magnitude of training compute while everything around it improves. Not a temporary plateau — a genuine ceiling that more of the same does not lift.

Position two: an ingredient is missing

This camp accepts the scaling data and denies its extrapolation. The claim is that current systems lack something specific — reliable causal reasoning, grounding of symbols in referents, persistent memory that updates with experience — and that no quantity of the existing recipe supplies a component the architecture does not have.

The evidence offered is qualitative but persistent: systems that answer sophisticated questions correctly and then fail simple variants that a person would treat as the same problem. On this reading, the failures are not gaps in coverage to be filled with more data. They are the signature of interpolation over a training distribution rather than reasoning about a structure.

What would refute it: scale alone producing reliable causal inference on genuinely novel problems — where "novel" means demonstrably outside the training distribution, which is the part that is hard to establish and where much of this argument actually happens.

Position three: intelligence needs a body

An older tradition, and the one most often dismissed without being read. The claim is that understanding is not a property of a symbol system at all; it develops from acting on a world and paying a cost for being wrong. A system trained on text has encountered descriptions of consequences and never a consequence.

The supporting observation is that human competence is grounded in sensorimotor experience in ways that show up everywhere in language — a fact that a text-only system must reconstruct from correlations rather than from having a body.

What would refute it: a text-only system that transfers reliably to novel physical tasks without physical training. Robotics results trained on internet-scale video are the live test of this, and the evidence is currently genuinely mixed.

Position four: the question is ill-posed

The position least represented in public argument and hardest to dismiss. It holds that "general intelligence" is not a coherent single quantity — that human cognition is a collection of specialised capacities with a shared substrate, and that expecting one threshold to be crossed misdescribes what is being built.

On this reading, systems will keep getting more capable at more things, some sooner and some later, and there will be no moment anyone can point at. The argument about whether AGI arrived will be settled by convention rather than by measurement — the way "is a virus alive" was.

What would refute it: a benchmark that the other three camps accept in advance as decisive. That none has been agreed after decades of the field is the position's main evidence.

What to do with a question you cannot settle

Two things, and both are more useful than a forecast.

First, notice when a claim about AGI is doing work that a claim about a specific capability would do better. "Will AGI arrive by 2030" is unanswerable. "Will a system reliably plan a multi-step laboratory protocol and recover from its own errors" is a question with an answer, a timeline, and consequences for what you should learn.

Second, hold the position you hold with its refutation attached. That is what separates a considered view from an alignment. If you cannot name the observation that would change your mind, the thing you have is not a belief about the world.

Sources & further reading

  1. Kaplan, J. et al. — Scaling Laws for Neural Language Models (2020). The empirical basis of the scaling position. Link →
  2. Marcus, G. — The Next Decade in AI: Four Steps Towards Robust Artificial Intelligence (2020). A statement of the missing-ingredient case. Link →
  3. Brooks, R. — Intelligence Without Representation. Artificial Intelligence 47 (1991). The embodiment argument in its original form. Link →
  4. Chollet, F. — On the Measure of Intelligence (2019). Argues the field measures skill where it should measure generalisation. Link →

Common questions

Do AI researchers agree on whether AGI is possible?

No. Surveys of published researchers return wide, unstable distributions on timelines and disagreement on whether the question is well-formed. Reported medians shift substantially year to year, which is itself informative.

Have current models passed tests that were supposed to indicate general intelligence?

Several, including the Turing test under some framings. Each time, the field concluded the test measured less than intended — which is a reasonable response, and also the pattern that makes benchmark evidence weak here.

Is scaling still producing gains?

Yes, though the input has shifted. Straightforward increases in pre-training data hit practical limits, and much of the recent improvement has come from data quality, post-training methods and inference-time computation rather than raw parameter count.

Does it matter to me which camp is right?

Less than it appears. Every camp expects capable, useful, unreliable systems for the foreseeable future. The skill of judging when a system's output can be trusted is required under all four.

Nanoschool AI Desk

Artificial intelligence editorial team · Reviewed by Nanoschool faculty

Covers machine learning for people who intend to build with it rather than read about it: what the architectures actually do, what the benchmarks actually measure, and which claims survive contact with a dataset you did not choose.

Hi! Need help? Chat with NSTC ✨