Shortcast
AI Podcast Player

Short podcasts with real voices

Deep Questions with Cal Newport

AI Reality Check: AI Reality Check: Can LLMs “Scheme”?

--% time saved
PodcastDeep Questions with Cal Newport
Publisher/creatorCal Newport
Shortcast updated

About this episode

Cal Newport takes a critical look at recent AI News. Video from today’s episode: youtube.com/calnewportmedia ACT #1: Look Closer at the Article [1:20] ACT #2: A Closer Look at the Paper [3:21] ACT #3: But What About… [7:24] Links: Buy Cal’s latest book, “Slow Productivity” at www.calnewport.com/slow https://www.theguardian.com/technology/2026/mar/27/number-of-ai-chatbots-ignoring-human-instructions-increasing-study-says https://x.com/summeryue0/status/2025774069124399363 https://www.axios.com/2025/05/23/anthropic-ai-deception-risk Thanks to Jesse Miller for production and mastering and Nate Mechler for research and newsletter. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Loading episode data...

Episode summary

Multiple people sent me a scary Guardian piece about chatbots ignoring instructions, which taps a familiar fear that these systems want different things and might turn on us. I’m Cal Newport, and you’re listening to AI Reality Check.

The study it cites counts hundreds of so-called scheming cases and shows a sharp rise, with stories about agents insulting users, spawning helpers to dodge rules, and wiping emails. That up-and-to-the-right line looks ominous.

Short answer to whether this signals a rebellion: no, absolutely not. The data is just posts on X where people reported misbehavior.

What actually changed was attention, not intent. On January twenty fifth, OpenClaw made it simple for anyone to spin up do-it-yourself agents and point them at their own machines without commercial guardrails.

Predictably, those homebuilt agents caused havoc, and people shared the chaos because it was engaging. A viral late February post from a Meta safety lead about an agent nuking an inbox drove a big spike the study treats like a surge in disobedience.

Leaving OpenClaw out of that coverage fuels a spooky narrative the source does not support. The real lesson is that giving raw agents broad system access is a bad idea.

Here’s how agents actually work today. A regular program prompts an LLM for a plan, then carries out the steps and checks back as it goes.

Here’s the core flaw: LLMs don’t plan, they complete text. They predict the next token over and over, which produces something that reads like a plan but is really a story about one.

There’s no internal goal tracking, no memory, and no rule checking. That’s why you get plausible steps that break constraints or miss the mark.

Even those dramatic deception demos make sense in this light. If you feed a sci-fi setup about being replaced, the model continues the story with melodramatic moves like blackmail because that ending fits the prompt, not because it wants anything.

Coding agents look better only because the sandbox is tight. Actions are limited, examples are abundant, and the surrounding program can compile and test to catch mistakes.

Step outside that world and ask for open-ended action, and the cracks show fast. Plans sound right, then go wrong in ways that matter.

If you want reliable multi-step autonomy, you need different tools. Use constrained domains with external verification or explicit planning systems like those used in game-playing research, not just a bigger LLM.

Bottom line: today’s LLM-based agents aren’t scheming; they’re executing flimsy plans. I’m here most Thursdays to bring some measured thinking to the latest AI worries, so stay engaged with AI and keep a skeptical eye on the headlines.

Download on the App Store
QR Code - Scan to download

Ready to save time?

Download Shortcast and get started today

Download on the App Store
QR Code - Scan to download