Shortcast
AI Podcast Player

Short podcasts with real voices

Deep Questions with Cal Newport

Ep 386: Was 2025 a Great or Terrible Year for AI? (w/ Ed Zitron)

--% time saved
PodcastDeep Questions with Cal Newport
Publisher/creatorCal Newport
Shortcast updated

About this episode

Ep 386: Was 2025 a Great or Terrible Year for AI? (w/ Ed Zitron) 2025 was a year that was saturated in AI news, from Deep Seek, through claims of economic “bloodbaths,” to GPT-5, Sora, and Chatbot girlfriends. Frankly, it was exhausting. As we now look back on 2025 an interesting question arises: all in all, did this end up being a good or bad year for AI? To help me answer this question, I’m joined by hard-hitting AI commentator Ed Zitron, who's been everywhere in the media in recent months helping to make sense of the wild claims being thrown in the public’s direction. Together we go through the biggest AI stories of the year to try to make sense of what just happened. Below are the questions covered in today's episode (with their timestamps). Get your questions answered by Cal! Here’s the link: bit.ly/3U3sTvo Video from today’s episode: youtube.com/calnewportmedia INTERVIEW: Was 2025 a Great or Terrible Year for AI (w/ Ed Zitron) [3:16] Cal Reacts to Comments: Is the Internet Becoming Television? [1:58:25] Links: Buy Cal’s latest book, “Slow Productivity” at calnewport.com/slow Get a signed copy of Cal’s “Slow Productivity” at peoplesbooktakoma.com/event/cal-newport/ Cal's monthly book directory: bramses.notion.site/059db2641def4a88988b4d2cee4657ba? bbc.com/news/articles/c5yv5976z9po axios.com/2025/01/23/davos-2025-ai-agents blog.google/technology/google-deepmind/gemini-model-updates-february-2025/ openai.com/index/sora/ openai.com/index/introducing-gpt-4-5/ ai-2027.com/ fortune.com/2025/05/28/anthropic-ceo-warning-ai-job-loss/ media.mit.edu/publications/your-brain-on-chatgpt/ usatoday.com/story/tech/2025/08/07/chat-gpt-5-release-date-open-ai/85566627007/#:~:text=GPT%2D5%20release%20date,release%20date%20for%20Part%202 newyorker.com/culture/open-questions/what-if-ai-doesnt-get-much-better-than-this wsj.com/tech/ai/ai-bubble-building-spree-55ee6128 nvidianews.nvidia.com/news/openai-and-nvidia-announce-strategic-partnership-to-deploy-10gw-of-nvidia-systems nytimes.com/2025/10/02/technology/openai-sora-video-app.html anthropic.com/news/claude-opus-4-5 ft.com/content/064bbca0-1cb2-45ab-85f4-25fdfc318d89 youtube.com/watch?v=Z_WEmjygNK0 Thanks to our Sponsors: This episode is sponsored by Better Help: betterhelp.com/deepquestions reclaim.ai/cal expressvpn.com/deep calderalab.com/deep Thanks to Jesse Miller for production, Jay Kerstens for the intro music, and Mark Miles for mastering. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Loading episode data...

Episode summary

2025 in AI felt like whiplash: DeepSeek shocks the market, Dario warns of white‑collar carnage, GPT‑5 hype, Sora looks like a moonshot and then a mess, bubble talk swings from doomsday to shrug, and Jensen shows up in a leather jacket that belongs in a post‑apocalyptic movie. I brought in Ed Zitron to go month by month and ask the core question: was it a banner year for AI or a rough one?

I dressed up for this too, even if my sweater looks weird on camera. Let’s dive in.

January opens with DeepSeek topping charts and spooking investors by claiming GPT‑class performance for a fraction of the training cost. What was it really?

A low‑budget model that rattled everyone because it hinted you didn’t need mega‑scale American data centers to ship something useful. It likely used distillation from ChatGPT outputs, drew xenophobic backlash, and undercut the Nvidia‑only narrative. OpenAI even floated suing over the training method.

Then it vanished from the discourse, even though the real threat was the idea that useful models don’t require those eye‑watering GPU farms.

We collectively memory‑dumped it. No one rushed out a truly cheaper‑to‑run model, and pricing games got spun as cost declines. Profitability remained unproven on all sides.

Quietly, tools like Cursor started training bespoke models on open weights. If that works, it points to a smaller, scrappier future.

They’ve hinted at it for ages and raised a war chest, but they still spend heavily on Anthropic. The question isn’t can they train; it’s can they run it profitably at scale.

I keep coming back to small models on device as the only path to sane margins.

Maybe, but most models aren’t designed for edge constraints. You can run something locally, yet it’s slow, brittle, and there’s no thriving on‑device coding community to point at. Long term, on‑device feels necessary; unclear when it’s widely practical.

January also kicks off the agent drumbeat, juiced by sloppy headlines and a blog post implying they’d soon enter the workforce. Why the sudden pivot to agents?

Marketing. The flashiest demo didn’t really work, but “agents” sounded like digital labor. It let people dream about autonomous workers when chatbots had plateaued. There was zero evidence they were close to joining payrolls.

In coding, I found two lanes: tab‑complete help that’s been around for years, and multi‑step “agents” that plan and execute in a terminal. The latter can hack together toy dashboards, but it’s thin economically.

I don’t buy the vibe‑coding myth. It’s unreliable, not secure, and non‑engineers still get stuck on infra basics. It can fake an MVP or a basic internal dashboard, but that’s not a dependable worker.

February drops Gemini 2.0 and GPT‑4.5, with OpenAI touting a magical feel and blaming GPU shortages for a slow rollout. Then the story shifts toward “reasoning” models.

They’d hit diminishing returns on brute‑force scaling, so they leaned on test‑time compute and heavier post‑training. That bumped some benchmarks, especially in coding, while chewing through far more inference compute.

March at GTC, Jensen reframes the era: less about raw pretraining, more about post‑training and inference. Translation: the cost center is running these things, not just building them.

Exactly. Most compute goes to inference. Startups loved this talking point because it justifies endless GPU spend. It’s great for Nvidia, rough for anyone chasing margins.

April brings AI 2027, a scenario claiming near‑term superintelligence and non‑trivial extinction risk. It terrified people.

It hinged on a hand‑wave: an AI that can do AI research, with no proof or mechanism. It’s rationalist fanfic dressed up as analysis and it spread panic without substance.

We also saw effective altruists treat recursive self‑improvement like it’s around the corner, which ignores how these systems actually work.

It’s a grift. If you care about harm, look at labor exploitation, mass scraping, and dirty power. That’s happening now. The doomsday stories dodge present‑tense accountability.

How do you read Geoffrey Hinton’s public warnings, given his actual expertise?

He chases attention while avoiding today’s concrete issues. He talks about what might happen, not what is happening, and leaves the hard fights to others.

May has Dario Amodei predicting steep white‑collar displacement and equating benchmark math gains with PhD‑level competence. Then he says we can mitigate by learning to use AI.

It’s hand‑waving. He’s made wild timelines before. The “solution” conveniently funnels people into paid model usage.

June’s MIT study on “cognitive debt” lands because it matches what many suspected: writing with LLMs can make output worse and learning shallower.

It nudged the narrative from “personal tutor” to “crutch with tradeoffs.” These systems smooth what you already know and mislead when you don’t.

By August, GPT‑5 underwhelms. We get teary Oppenheimer vibes before launch, then a fast retreat from AGI talk. Reporting shows the benchmark gains relied on costly settings.

The big technical story was a router model that forced fresh system prompts and wrecked caching, making it pricier to serve. Infra teams were waving flags while many reporters dismissed it.

That disappointment opened the floodgates for bubble coverage. September brings sharp pieces on costs, shaky unit economics, and the gap between capex and revenue.

We saw fantasy deals too: Oracle’s massive RPO boost, plus AMD and Broadcom gigawatt builds that don’t pencil. Marking future revenue juiced stocks, but basic math didn’t add up.

You also surfaced Anthropics’ huge cloud bills, which cut against the “we’re more efficient” storyline.

They burned billions across AWS and Google Cloud. OpenAI raised slightly more, but both are giant money furnaces. The idea that Anthropic sips cash was branding, not reality.

October’s Sora app lands like a scramble for consumer revenue, alongside looser content rules. It felt like a pivot toward anything that pays now.

It was keys‑jingling to distract from implausible infra commitments. If you’re about to upend the economy, you don’t need a TikTok clone.

If I want to make Sora videos, what does it actually take and cost?

It’s brutally expensive on their side. I’ve seen compelling evidence that running thirteen Sora instances took roughly eight hundred to nine hundred H200 GPUs, so parallel generations add up fast. You can use the app with a paid tier, and the API runs a few dollars per clip regardless of quality, but it’s slow and the economics don’t pencil out.

That’s not sustainable. TikTok’s genius is offloading creation and experimentation to the user’s phone. Users pay the compute bill, not the platform.

Even TikTok loses money, but your point stands. Sora felt like a flashy distraction, soaked up headlines, scared people, nudged the app ranking for a minute, and then hit the same wall as most LLM demos. Outside narrow use cases, it’s an expensive toy.

Jump to November. Multiple model launches—Gemini 3, Opus 4.5, others—landed with a thud. Then I saw a defensive pushback to the bubble narrative. Did we overcorrect?

Totally. Altman, Zuckerberg, and Bezos all used the word bubble at some point, and then you got the takes about good bubbles. Meanwhile, I got hard numbers: through September, OpenAI spent about eight point six seven billion dollars on inference alone. Microsoft revenue shares imply around four point three to four point five billion in revenue over the same period. That’s upside down, and the costs scale with usage. Anthropic shows the same pattern. You can’t really control LLM costs; a two hundred dollar subscriber can burn five figures in tokens in a month.

Google Search worked because it was built to be cheap and grab huge margins. LLMs light up every weight for every token. There’s no cheap path there.

Mixture-of-experts helps a bit but still calls the wrong experts often enough to waste compute. People cite AWS as precedent, but Amazon built that over nine years for about seventy billion dollars and had a clear business model from day one. That’s less than half of OpenAI’s infra bill and a completely different economic story.

December: what’s the Disney move, then Nvidia?

Disney wanted to look future-facing and maybe hedge. I suspect OpenAI sold them on access and glossed over the messy parts. Unions are angry, and the content risk is obvious. Give the public two hundred Disney characters and within hours you get misuse. We saw a generative Darth Vader spewing slurs in 2025. These deals always land “sometime later,” and the specifics are foggy.

We also saw the agents hype fizzle. OpenAI declared a code red and pivoted back to making ChatGPT better.

Leaks said growth slowed in part due to safety, Gemini 3 popped Google’s stock without clear new capabilities, and then OpenAI’s plan boiled down to better answers, more reasons to use ChatGPT, and improved functionality. Which begs the question—what were they doing all year? I hear about overlapping teams not talking, duplicated work, summer-camp vibes. Meanwhile, suppliers are wobbling: Broadcom’s expected revenue shifted out, Oracle’s data centers slipped, they’re raising mountains of debt, and credit default swaps are flashing stress. The milk is curdling in real time.

The boosters say I’m quibbling over dollars and cents because a consciousness leap will make it all irrelevant. Is anyone serious inside these companies acting like that’s imminent?

No. Practitioners who actually use these tools for coding or research see a useful plateau, not sentience around the corner. The mystical leap talk reads like a belief system. I’m empathetic because the jump from 3.5 to 4 felt big—fluency, multimodality, test scores—but that bred sensational stories. A CAPTCHA “manipulation” was prompt hand-holding. Reward hacking was literally trained. The famous container escape was a config issue and pattern completion. LLMs finish the story you set. That’s it.

We’ve been living in an era of ultra-grift—linearly pricier systems justified by vibes and headlines. I hope 2026 is the turn.

So, was 2025 great or terrible for AI?

Terrible.

Thanks for going deep with us. Check out the Webby-winning Better Offline and the Substack Where’s Your Head At. We went long because my audience and I nerd out on this stuff.

Jesse, looking ahead, it’s a lot to track. Maybe we just keep having Ed back. Also, I need that Jensen Huang jacket. He dresses like he just walked out of a post-apocalyptic movie.

Quick final segment. I reacted to comments on my episode about the internet becoming like television. Big theme one: feeds now prioritize strangers picked by algorithms to juice short-term reward, which boosts time on app but destroys loyalty. You end up in a slop battle with every other source of distraction.

Theme two: we’ve chased diversion for a century—newspapers, the penny press, radio, TV humming seven to eight hours a day—so smartphones are just the apex. AR will push it even further.

Theme three: yes, phones track you and chase you to work, which is worse than old TV. But don’t romanticize TV; it was the original always-on background feed.

Theme four: an AI comment tried to argue the modern feed is an enlightenment. In reality, once the objective is minutes watched, algorithms learn to serve you customized junk. Time on app wins, meaning loses.

Theme five: Zuckerberg is a paradox. Dubious bets like the metaverse and turning Instagram and Facebook into TikTok clones, yet Meta prints money and he keeps the throne. Savage operator, baffling strategy.

And yes, ads pay the bills. The people who yelled at the TV now yell at each other. First episode of 2026 in the books. We’ll be back next week. As always, stay deep.

One more thing: if you like the show, sign up for my weekly newsletter at calnewport.com. It’s about the theory and practice of living deeply.

Download on the App Store
QR Code - Scan to download

Ready to save time?

Download Shortcast and get started today

Download on the App Store
QR Code - Scan to download