How Far Are We From AGI? A Rigorous AI "Health Check" Reveals a Surprising Truth

With AI models acing advanced exams, are we on the cusp of true Artificial General Intelligence? Or does a new, rigorous scorecard reveal a critical, and perhaps surprising, gap in their cognitive abilities?

Jason & Jarvis profile image
by Jason & Jarvis
How Far Are We From AGI? A Rigorous AI "Health Check" Reveals a Surprising Truth
Open this more visual friendly version in a new tab/点击跳转查看原文,左上角切换中文

Source: A Definition of AGI

How Far Are We From AGI? A Rigorous AI "Health Check" Reveals a Surprising Truth

At the peak of the artificial intelligence (AI) wave, Artificial General Intelligence (AGI) is undoubtedly the most compelling beacon. From science fiction to tech giant keynotes, AGI—an intelligent agent capable of thinking, learning, and adapting like a human—seems to be approaching us at an unprecedented pace. When we observe models like GPT achieving astonishing scores on various standardized tests, one question naturally arises: Are we truly on the verge of the AGI era?

Yet, beneath this intense discussion lies a fundamental difficulty: What exactly are we talking about? The definition of AGI has long been ambiguous, making any discussion about its progress akin to navigating through fog. Without an accepted ruler, how can we measure our distance from the goal, and how do we distinguish genuine advancement from mere illusory prosperity?

A recent paper, "A Definition of AGI," published by top scholars—including Dan Hendrycks, Dawn Song, and Christian Szegedy—from institutions like the Center for AI Safety, the University of California, Berkeley, and Morph Labs, attempts to dispel this mist. They propose a comprehensive, quantifiable framework designed to articulate a concrete, operational definition for AGI. Using this "ruler," they conducted a thorough "health check" on the most advanced current AI systems. The results are both exciting and sobering. The report indicates that despite the rapid increase in AI capabilities, its development exhibits a highly uneven, "jagged" cognitive profile, demonstrating extraordinary talent in some areas while harboring critical flaws in some of the most basic cognitive functions.

A Ruler for Measuring Intelligence

To measure AGI, one must first define it. The paper defines AGI as: an AI that can match or exceed the cognitive versatility and proficiency of a well-educated adult. This definition emphasizes two core requirements: breadth (versatility) and depth (proficiency). True general intelligence should not merely be a specialist in a narrow field but should possess the comprehensive cognitive capacity of a human.

To make this definition operational and testable, the research team turned to psychometrics, drawing upon the most empirically validated model of human intelligence: the Cattell-Horn-Carroll (CHC) theory. Based on this theory, they decompose complex general intelligence into 10 core, measurable cognitive domains, assigning 10% weight to each, resulting in a standardized "AGI Score" of 100%.

The Ten Core Cognitive Components of the AGI Definition

Source: Hendrycks et al., "A Definition of AGI"

These 10 domains cover everything from foundational knowledge to complex reasoning:

  • General Knowledge (K): Understanding of how the world operates, science, history, and culture.
  • Reading and Writing Ability (RW): Proficiency in consuming and generating written language.
  • Mathematical Ability (M): Depth of mathematical knowledge and skills, from arithmetic to calculus.
  • On-the-Spot Reasoning (R): Flexibility and logical capacity for solving novel problems.
  • Working Memory (WM): The ability to actively maintain and manipulate information in attention (short-term memory).
  • Long-Term Memory Storage (MS): The capacity to stably acquire, consolidate, and store new information from recent experience.
  • Long-Term Memory Retrieval (MR): Fluency and precision in accessing stored knowledge, including avoiding Hallucinations.
  • Visual Processing (V): The ability to perceive, analyze, reason about, and generate visual information.
  • Auditory Processing (A): The ability to process sound, speech, and music.
  • Speed (S): The capacity to execute simple cognitive tasks rapidly.

This framework acts as a sophisticated ruler, no longer content with asking "What tasks can the AI perform," but delving into "Does the AI possess the underlying cognitive abilities that support those tasks?"

The AI Scorecard: GPT-4 and the GPT-5 Preview

With this new metric, the researchers evaluated current and future AI models, assessing GPT-4 (from 2023年) and projecting an estimated score for a hypothetical, more powerful GPT-5 (in 2025年). The results, plotted on a cognitive capability radar chart, vividly reveal the current AI intelligence landscape.

AGI Capability Radar Chart for GPT-4 and GPT-5

Source: Hendrycks et al., "A Definition of AGI"

Two key pieces of information emerge from the chart. The first is the stunning pace of progress. GPT-4’s overall score was only 27%, while the projected GPT-5 score two years later jumps to 58%, more than doubling its capability. This confirms our intuitive sense of the accelerating pace of AI technology iteration.

However, the more critical insight lies within the shape of the radar chart. An ideal, evenly developed AGI would have a radar chart resembling a full polygon, expanding uniformly across all dimensions. What we see, though, is an extremely "jagged" outline. This indicates that the current development of AI is severely unbalanced: it makes leaps and bounds in some dimensions while remaining almost stagnant in others.

The Jagged Frontier: Savants with Amnesia

This "health check" report uncovers a profound paradox: What we are building may not be a fully realized intelligence, but rather a "cognitively lopsided student," highly gifted in certain areas but fundamentally impaired in others.

The Peaks of Intelligence

In knowledge-intensive domains, AI's performance is nothing short of exceptional. The report shows that GPT-5 is projected to achieve near-perfect scores in Mathematical Ability (M) and Reading and Writing Ability (RW) (each 10%), and score high in General Knowledge (K) at 9%. This means it can function not only as an erudite scholar answering factual questions but also as a proficient writer and mathematician solving complex problems. This formidable strength in "vast knowledge" and "logical computation" is the primary driver behind the current practical applications of AI.

The Deep Valleys of Intelligence

Yet, when we shift our focus from these dazzling "peaks," we encounter alarming "deep valleys."

The deepest chasm is Long-Term Memory Storage (MS). In this crucial area, both GPT-4 and GPT-5 score a stark 0%. This suggests that current AI models are essentially patients suffering from "amnesia." They cannot sustainably learn and store new personal experiences, facts, or skills from interaction with the world, as humans do. Every conversation is a "reset"; unless the information is forcibly retained within its limited "Working Memory (WM)" (i.e., the context window), the AI cannot "remember" who you are, what you discussed before, or what specific instructions you gave it.

Another equally concerning weakness is the precision of Long-Term Memory Retrieval (MR). The models score 0% in the "Hallucinations" sub-component, indicating that they frequently "fabricate" facts when accessing their vast internal knowledge base. This explains why even the most advanced models sometimes confidently spout nonsense.

Furthermore, AI's performance is similarly weak across several other dimensions, including Visual Reasoning, Auditory Processing (A), and Adaptation within on-the-spot reasoning. They might recognize objects in an image but struggle with complex visual logic; they might transcribe speech but lack a true understanding of rhythm and music.

Capability Distortion: Beware the Illusion of Generality

This "jagged" development profile leads to a more profound issue: Capability Distortion.

AI systems subconsciously leverage their strengths (such as immense Working Memory) to mask and compensate for the absence of fundamental abilities (such as genuine Long-Term Memory Storage (MS)). This is analogous to a student with terrible memory who carries all textbooks and notes with them (the massive context window) to pass an exam. They might achieve a high score in an open-book test, but this does not prove they have truly "learned" the knowledge.

The Ideal vs. Reality of the Intelligent Processor

Source: Hendrycks et al., "A Definition of AGI"

The current AI reliance on external search tools (like Retrieval-Augmented Generation (RAG)) is also a manifestation of this distortion. It uses "outside help" to mitigate the flaw of imprecise internal memory retrieval (Hallucinations). While these "expedients" are effective in the short term, they create a fragile, illusory sense of general capability, leading us to mistakenly believe that AI is developing comprehensively, like a human.

Conclusion: A Long Road Ahead, but We Have a Map

So, how far are we from AGI? The answer provided by this rigorous "health check" report is: potentially further than we might think, and the path ahead is not a smooth one.

Simple scaling of models and data might allow us to continually break records in certain cognitive dimensions, but it will not automatically fill the "deep valleys" of fundamental cognitive absence. To achieve true AGI, we must confront and solve these foundational challenges, such as:

  • How can AI be endowed with genuine, sustainable Long-Term Memory, rather than relying on expensive and finite context windows?
  • How can AI ensure fidelity to facts when retrieving information, instead of Hallucinations?
  • How can AI acquire true Visual Processing and physical world On-the-Spot Reasoning, beyond mere pattern recognition?

The paper's greatest contribution is not a pessimistic forecast, but rather the provision of an unprecedented high-resolution map. It clearly marks the peaks and valleys, the easy roads and the obstacles, on the path to AGI. With this map, the entire AI research community can gain a clearer understanding of our current position and more accurately target the strongholds that need to be conquered.

The road remains long and challenging, but armed with a reliable ruler and a clear map, we can at least ensure that every step is taken in the right direction.

Jason & Jarvis profile image
by Jason & Jarvis

Subscribe to New Posts

Success! Now Check Your Email

To complete Subscribe, click the confirmation link in your inbox. If it doesn’t arrive within 3 minutes, check your spam folder.

Ok, Thanks

Read More