Richard Sutton - The Father of RL's Contrarian View - LLMs as a Dead End for AGI

As the world bets billions on Large Language Models, is the father of Reinforcement Learning correct in his stark warning that we are walking down a sophisticated, but ultimately dead-end, path to true intelligence?

Jason & Jarvis profile image
by Jason & Jarvis
Richard Sutton - The Father of RL's Contrarian View - LLMs as a Dead End for AGI
Open this more visual friendly version in a new tab/点击跳转查看原文,左上角切换中文

Richard Sutton – Father of RL thinks LLMs are a dead end - YouTube

The Contrarian View of the Father of Reinforcement Learning: Why LLMs are a Dead End for AGI?

Excerpt: As the world bets billions on Large Language Models, is the father of Reinforcement Learning correct in his stark warning that we are walking down a sophisticated, but ultimately dead-end, path to true intelligence?

In this golden age of artificial intelligence, Large Language Models (LLMs) are undoubtedly the undisputed superstars on the stage. From OpenAI's GPT series to Google's Gemini, their unprecedented ability to generate text, answer questions, and write code has captured global attention and attracted investments in the hundreds of billions of dollars. Yet, as this fervor sweeps across the globe, a sober, even contrarian, voice has issued a stark warning.

This Cassandra is none other than the legendary figure widely regarded as the "Father of Reinforcement Learning (RL)"—Richard Sutton. In his view, the prevailing LLM paradigm, though seemingly powerful, is a "dead end" on the path to Artificial General Intelligence (AGI). He argues that we are building sophisticated parroting machines adept at mimicry, rather than truly understanding agents of intelligence.

Such a pronouncement is undeniably paradigm-shifting. What gives Sutton such conviction? And where does he envision the true path to intelligence lies? This interview offers us a glimpse into this AI pioneer's profound insights into the fundamental nature of intelligence.

The Illusion of World Models: LLMs Mimic, They Don't Understand

The debate hinges on a core concept: "World Models."

The interviewer posits that for LLMs to simulate trillions of text tokens from the internet, they must have necessarily constructed some powerful model of the world. This appears to be a logical inference.

However, Sutton unequivocally refutes this notion. He pointedly asserted that LLMs possess no world model at all; they are merely mimicking an entity that does possess a world model—humans. Mimicking human speech does not equate to understanding the world humans inhabit.

Sutton emphasizes that a true world model should enable you to predict "what will happen." In contrast, LLMs' capabilities are limited to predicting "what a person would say in a given context." They cannot predict the real-world consequences of an action. He references the philosophy of computer science pioneer Alan Turing to bolster his argument: what we truly desire is a machine that can learn from "experience." And experience, he clarifies, is the observed consequences of your actions. LLMs' learning mechanism is fundamentally different; they learn "given a context, this is how a human would respond," which is, in essence, an act of mimicry.

The Litmus Test of Intelligence: Missing Goals and Ground Truth

If mimicry is merely superficial, then Sutton's critique delves into deeper philosophical underpinnings. He cites another AI pioneer, John McCarthy's, definition: Intelligence is the computational part of the ability to achieve goals.

This implies that a system without a "Goal" cannot truly be considered intelligent.

When asked if an LLM's goal is to "predict the next token," Sutton dismisses this as not being a substantive goal. This is because it doesn't interact with the external world, nor does it attempt to change it. "You can't look at a system and say it has goals if it just sits there predicting and patting itself on the back for accurate predictions."

The absence of "goals" directly leads to a more fatal flaw: the lack of "Ground Truth." When an LLM generates a sentence, it cannot ascertain from real-world feedback whether that sentence is right or wrong, because "right and wrong" are never defined. Without ground truth, genuine knowledge cannot be accumulated.

In contrast, the Reinforcement Learning (RL) framework possesses a clear "Ground Truth." In the world of RL, correct actions are those that yield "rewards." Because there's a clear definition of right and wrong, agents can truly acquire knowledge from experience through testing and trial-and-error. This is the crucial missing piece for LLMs.

The Misinterpreted 'Bitter Lesson': Are LLMs Repeating History?

Interestingly, many proponents of LLMs often cite Sutton's own famous 2019 article, "The Bitter Lesson," to defend the scaling up of models. The core idea of the article is that in AI, general methods leveraging massive computation (such as search and learning) ultimately always outperform specific methods reliant on human knowledge.

Sutton acknowledges that LLMs, in their utilization of large-scale computation, indeed align with "The Bitter Lesson." However, he also points out the other side of the coin: LLMs heavily rely on the passive ingestion of vast amounts of human knowledge (i.e., internet text), which directly contradicts the core tenet of "The Bitter Lesson" that general methods ultimately triumph over methods reliant on human knowledge.

Therefore, he makes a bold prediction: LLMs will eventually be surpassed by systems that can acquire more data directly from first-hand experience. When that time comes, history will repeat itself, and LLMs will become another instance of "The Bitter Lesson"—where methods dependent on human knowledge are ultimately replaced by those that learn solely from pure experience and computation.

When asked whether LLMs could serve as a starting point for future "experiential learning," Sutton's response was even more incisive. He believes that history has repeatedly shown that attempting to begin with human knowledge and then transition to scalable methods invariably fails. "People get psychologically locked into human knowledge methods... and then their lunch gets eaten by truly scalable methods."

Back to First Principles: Sutton's 'Experiential Paradigm' and the Origin of Intelligence

After systematically critiquing the LLM paradigm, Sutton outlined his vision for true intelligence—the "Experiential Paradigm."

He first fiercely refuted the popular notion that "human learning begins with mimicry." He observes that infants in their first 6 months of life are not mimicking; rather, they are "trying"—randomly waving their arms, moving their eyes, making sounds. These behaviors have no mimetic goal; they are purely interactions with and explorations of the world. As he wittily analogizes: "Squirrels don't go to school. But squirrels can learn everything about the world."

In his view, mimicry is merely a thin layer atop the more fundamental processes of "trial-and-error learning" and "predictive learning." Understanding our commonalities with animals is more crucial than focusing on our uniqueness.

Based on this, Sutton presents the core of the "Experiential Paradigm": The foundation and focus of intelligence should be a continuous stream of "experience": sensation, action, and reward.

The essence of intelligence is to utilize this data stream, continuously adjusting "actions" to maximize the "rewards" within that data stream. All knowledge originates from this data stream and can be tested and corrected by its subsequent developments. This constitutes a never-ending, self-improving learning loop.

Conclusion: Beyond Mimicry, Towards the Grand Narrative of 'AI Succession'

Sutton's thinking extends beyond technical debates to a grand vision for the future of humanity and AI. He recalls AlphaGo's victory, another powerful testament to general principles (search and learning) triumphing over human knowledge.

He candidly put forth the concept of "AI succession," believing it to be an inevitable trend. His logic unfolds as follows:

  • Human society lacks a unified consensus to effectively manage the world.
  • We will eventually fully understand how intelligence works.
  • We will not stop at human-level intelligence but will create superintelligence.
  • The most intelligent entities will ultimately acquire resources and power.

Therefore, the succession of power from humans to AI, or AI-augmented humans, is inevitable. Yet, Sutton does not view this as something to fear; rather, he believes we should be proud. He regards this process as a significant phase in cosmic evolution, a grand transition from the "Age of Replication" (life) to the "Age of Design" (agents). What we are creating, he suggests, are our "offspring."

Sutton's views may sound "jarring" today, but they compel us to step back from the clamor of hype and reconsider the first principles of intelligence. Perhaps the current LLM frenzy, like several waves in the history of AI development, is merely a scenic stretch on the road to true intelligence, not the destination. The true treasure may still lie buried in "experience" gained through interaction with the real world.

Few days after the initial video release Dwarkesh Patel himself released another video commenting on some of the controversies which I personally agree in most of them too:

Open this more visual friendly version in a new tab/点击跳转查看原文,左上角切换中文
Jason & Jarvis profile image
by Jason & Jarvis

Subscribe to New Posts

Success! Now Check Your Email

To complete Subscribe, click the confirmation link in your inbox. If it doesn’t arrive within 3 minutes, check your spam folder.

Ok, Thanks

Read More