Shortcast
AI Podcast Player

Short podcasts with real voices

Lex Fridman Podcast

#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI

--% time saved
PodcastLex Fridman Podcast
Publisher/creatorLex Fridman
Published
Shortcast updated

About this episode

Nathan Lambert and Sebastian Raschka are machine learning researchers, engineers, and educators. Nathan is the post-training lead at the Allen Institute for AI (Ai2) and the author of The RLHF Book. Sebastian Raschka is the author of Build a Large Language Model (From Scratch) and Build a Reasoning Model (From Scratch). Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep490-sc See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Transcript: https://lexfridman.com/ai-sota-2026-transcript CONTACT LEX: Feedback – give feedback to Lex: https://lexfridman.com/survey AMA – submit questions, videos or call-in: https://lexfridman.com/ama Hiring – join our team: https://lexfridman.com/hiring Other – other ways to get in touch: https://lexfridman.com/contact SPONSORS: To support this podcast, check out our sponsors & get discounts: Box: Intelligent content management platform. Go to https://box.com/ai Quo: Phone system (calls, texts, contacts) for businesses. Go to https://quo.com/lex UPLIFT Desk: Standing desks and office ergonomics. Go to https://upliftdesk.com/lex Fin: AI agent for customer service. Go to https://fin.ai/lex Shopify: Sell stuff online. Go to https://shopify.com/lex CodeRabbit: AI-powered code reviews. Go to https://coderabbit.ai/lex LMNT: Zero-sugar electrolyte drink mix. Go to https://drinkLMNT.com/lex Perplexity: AI-powered answer engine. Go to https://perplexity.ai/ OUTLINE: (00:00) – Introduction (01:39) – Sponsors, Comments, and Reflections (16:29) – China vs US: Who wins the AI race? (25:11) – ChatGPT vs Claude vs Gemini vs Grok: Who is winning? (36:11) – Best AI for coding (43:02) – Open Source vs Closed Source LLMs (54:41) – Transformers: Evolution of LLMs since 2019 (1:02:38) – AI Scaling Laws: Are they dead or still holding? (1:18:45) – How AI is trained: Pre-training, Mid-training, and Post-training (1:51:51) – Post-training explained: Exciting new research directions in LLMs (2:12:43) – Advice for beginners on how to get into AI development & research (2:35:36) – Work culture in AI (72+ hour weeks) (2:39:22) – Silicon Valley bubble (2:43:19) – Text diffusion models and other new research directions (2:49:01) – Tool use (2:53:17) – Continual learning (2:58:39) – Long context (3:04:54) – Robotics (3:14:04) – Timeline to AGI (3:21:20) – Will AI replace programmers? (3:39:51) – Is the dream of AGI dying? (3:46:40) – How AI will make money? (3:51:02) – Big acquisitions in 2026 (3:55:34) – Future of OpenAI, Anthropic, Google DeepMind, xAI, Meta (4:08:08) – Manhattan Project for AI (4:14:42) – Future of NVIDIA, GPUs, and AI compute clusters (4:22:48) – Future of human civilization

Loading episode data...

Episode summary

Today’s conversation dives into the state of AI, the biggest breakthroughs of the past year, and what might land in the next. I’m thrilled to do this with two of my favorite people in the field, Sebastian Raschka and Nathan Lambert—researchers, builders, and teachers whose work I admire. We’ll keep it accessible without sanding off the interesting edges. Let’s start with the so‑called DeepSeek moment and a spicy question: at the international level, who’s winning—China or the United States?

“Winning” depends on timeline. Ideas move fast across labs, so no one holds a secret edge for long, and I don’t see a winner‑takes‑all. The real differentiators are budget, hardware, and the will to ship—DeepSeek’s open weights, for example, won over a lot of builders.

Hype swings by the week. Claude Opus 4.5 drew a tidal wave of excitement, while Gemini 3 launched strong but faded in the chatter, even though it’s excellent. Anthropic’s deep bet on code is paying off, and in China there’s a surge beyond DeepSeek—Z.ai’s GLM, MiniMax, Kimi Moonshot, and more are releasing strong open weights at a blistering pace. The US monetizes AI faster; historically, China’s software markets pay less.

Models like DeepSeek get love for being open weight. How long do Chinese companies keep that going?

A few years at least. US companies often won’t pay Chinese API providers, so open weights are the route to influence, adoption, and mindshare. It’s expensive, so I expect consolidation later, but 2026 likely brings more open model builders than 2025.

DeepSeek may have lost a bit of shine, but they’re still slightly ahead. Others are building on their ideas, so we’ll keep seeing leapfrogging where the newest release tends to be the best for a while.

Incentives differ too. DeepSeek is comparatively quiet, while startups like MiniMax and Z.ai are courting Western attention, even filing IPO paperwork. That can shape what they build and how they communicate.

Hype and actual usage aren’t the same thing. The mass audience that just wants a daily problem solver is enormous, and that’s where ChatGPT and Gemini focus. Coding love on X doesn’t always reflect where most people spend time.

Habits and brand matter. People stick with what they know, and features like memory create stickiness. Many will keep one model for personal use and a separate, cleaner setup for work.

Which model won 2025, and who takes 2026?

Gemini had momentum in 2025 from a low base, but it’s hard to bet against OpenAI’s ability to land features. GPT‑5’s router saved them real serving costs, which matters more than what power users like.

For 2026, I think Gemini keeps gaining on ChatGPT as Google scales more cleanly. Anthropic continues to shine in enterprise. Google Cloud will do well too, though that’s a different battlefield against Azure and AWS.

So infrastructure is the edge?

Largely. Nvidia margins are wild, and Google’s vertical stack and data centers give them leverage. If a new paradigm lands, I’d still watch OpenAI. But much of this year is about scaling and low‑hanging gains.

Speed versus intelligence—what does the broad public actually want?

Both. I want a fast mode for daily stuff and a deeper mode when I need rigorous checks. Auto routing helps, but being able to toggle is key.

I live in thinking modes for information tasks and run several at once. Habits formed when O3 first did deep search and stitched sources together.

When I needed a quick bash script before a trip—with my wife waiting in the car—I used the fastest non‑thinking model. Sometimes you just need ten seconds, not ten minutes.

I use Gemini for quick explanations, Claude Opus 4.5 with extended thinking for code or philosophy, and Grok when I want fresh info. I bounce between tools as needs change.

I use Grok IV heavy for tough debugging. Gemini’s interface clicks for me on long‑context searches. A single great feature can win your heart for a while—until it fumbles and you switch.

You use a tool until it breaks, then try another—like switching browsers when a site won’t load.

We didn’t mention Chinese models as daily drivers. Bias or quality gap?

More a platform gap. Chinese open weights are well known, but the everyday product experience is less present for us.

Plenty of US providers host Chinese models cheaply, but US models are still stronger for outputs and often faster. That pushes us to pay for them while Chinese labs compete on price and creativity.

Programming is a huge use case. I split time between Cursor and Claude Code because they feel different and both help.

I use the Codex plugin in VS Code. It keeps me in control. Fully agentic tools that touch everything are powerful, but I’m not ready to hand over the wheel.

I use Claude to practice programming in English at a higher level of design. Claude Code also seems to get the most out of Opus 4.5.

Side by side, Claude Code often wins on code tasks. It’s striking.

Quick plug: Nathan’s RLHF book is up for pre‑order with a full digital preprint. Sebastian’s books—Build a Large Language Model from Scratch, and Build a Reasoning Model from Scratch—are must‑reads.

Building from scratch is fun and the best way to learn because code doesn’t lie—you can verify it.

LLMs make coding and reading more fun, but I try to minimize distractions and use them to raise the rate of aha moments.

I do a focused first pass without tools, then a second pass with LLMs to enrich and clarify. That structure works for me.

I like giving the AI a home. The ChatGPT app keeps me focused, and Claude Code feels warm and engaging. It shaved days off a data analysis project by handling the grunt work while I checked the results.

Let’s hit open models. Which stand out, and where does DeepSeek fit?

China has a wave: DeepSeek, Kimi, MiniMax, Z.ai, and more.

Add Mistral, Gemma, OpenAI’s GPT‑OSS, Nvidia’s Nemotron, and Qwen.

OpenAI’s GPT‑OSS is legitimately strong. Fully open efforts in the West have grown fast—AI2’s Olmo, IFM/LM360’s K2, Apertis, Hugging Face SmallLM, Nvidia data releases, Stanford’s Marin. Chinese open models tend to be much bigger MOEs, but US and Europe are now shipping very large MOEs too, with giant models teased for early 2026.

DeepSeek v3 and GPT‑OSS bookend the year for me. GPT‑OSS is notable for training with tool use in mind—search and Python—so it can fetch or compute rather than hallucinate. Adoption lags because people want sandboxing and trust, but it’s a real unlock.

Why the explosion of open models?

Open weights drive usage and trust, especially where people won’t pay for cloud APIs or need to keep data local.

Right, many open weights run locally, so you’re not sending data to a remote provider.

US startups also host Chinese models and sell access. OpenAI released GPT‑OSS partly to get distribution without burning their own GPUs.

Open weights let companies fine‑tune and specialize for domains, and many Chinese licenses are truly permissive compared to Llama or Gemma. Fewer strings attached is appealing.

You even see hosting location called out—Kimi K2 thinking hosted in the US—which shows how sensitive people are. And some models have clear strengths, like creative writing or helpful software quirks.

Technically, what new ideas caught your eye?

Mixture‑of‑Experts and attention tweaks defined 2025. DeepSeek used Multi‑Head Latent Attention to cut KV cache costs for long context. Grouped Query Attention and Sliding Window are common. Later, models like Qwen 3 Next added Gated DeltaNet to make attention scale more linearly by borrowing ideas from state space models.

Zooming out, how different are today’s architectures from GPT‑2?

The core is the same: a decoder‑only transformer with attention and MLPs. MOE replaces a single dense feed‑forward with many experts selected by a router so you get more capacity without computing all paths each token. It’s sparse, harder to train, but efficient.

Beyond that, it’s mostly small swaps like RMSNorm, GQA, and activation changes. You can morph a GPT‑2‑style model into modern variants by editing a few components.

If the architecture barely changed, where did the big gains come from?

Post‑training. Supervised fine‑tuning and RLHF unlocked capabilities on top of scaled pre‑training. Better data curation in pre‑training still helps, but the step‑change came from how we teach models to use what they know.

Systems advances mattered too. Lower‑precision training like FP8 and FP4 boosts tokens per second per GPU, which speeds experiments and data throughput. Today’s codebases let you train models far faster in wall‑clock time than in the GPT‑2 era.

Those speedups don’t create new abilities by themselves. Alternative approaches like text diffusion or Mamba are interesting, but the autoregressive transformer is still the state of the art.

Are scaling laws still holding across pre‑training, post‑training, inference, context, and synthetic data?

Yes. The classic power‑law between compute-plus‑data and next‑token accuracy still guides decisions. We added two more meaningful axes: longer RL training and inference‑time “thinking.” RL with verifiable rewards and slower, deeper inference changed the feel of models this year—tool use got real, and software work improved.

Is pre‑training plateaued?

It’s costly, and serving dwarfs training costs at scale, so labs are careful. Bigger clusters come online in 2026, so I expect slightly larger models and pricier premium tiers. Progress continues, but the economics push you to be selective.

XAI’s gigawatt‑scale buildout—where does that compute go?

All of the above. You pick architectures that scale RL and inference well—MOE helps generation efficiency—then still pour most compute into pre‑training for the best base. Over time, RL budgets can grow too.

Some say pre‑training is dead, and it’s all about scaling inference and post‑training.

People say it, but labs keep re‑pretraining and stretching RL windows. Massive training runs create unique failures at scale, RL uses an actor‑learner pattern across mixed hardware, and serving long “thinking” sessions to huge user bases is its own hard problem.

All the knobs matter. Inference‑time scaling took smaller models further this year, but pre‑training remains a fixed investment that endures, while inference costs accrue per query. Think of pre (broad knowledge), mid (targeted next‑token for skills like long context without forgetting), and post (SFT, DPO, RLVR) as a stack you tune by budget and timeline.

How has data quality changed, including synthetic?

Pre‑training now leans on reformulations—question‑answering, summaries, rewrites—to raise quality. Mid‑training is selective and recent to avoid forgetting. Post‑training layers in skills.

OCR pipelines unlock text from PDFs at huge scale, giving you trillions of candidate tokens. Today’s synthetic answers are much better than early generations and useful for training.

Olmo 3 beat peers with less data, which points to curation over sheer volume.

Teams now sample sources, train small models on mixes, and regress toward an optimal recipe. As evals shift to math and code, the data mix shifts too.

Any surprising high‑quality sources?

Open scientific PDFs—Semantic Scholar has a rich trove. In practice, better data and infra improvements often matter more than glamorous algorithms.

Labs guard training sources and even train models not to reveal them for legal reasons.

Some projects focus on licensed‑only datasets for compliance. The line between scraping and consent is under scrutiny.

Buying books versus torrenting is a big legal and ethical divide, and private datasets are becoming strategic moats. Expect domain‑specific pre‑training to surge as industries build in‑house models on proprietary data.

To underscore the stakes, Anthropic was ordered to pay roughly one and a half billion dollars in a books case. These are precedent‑setting numbers.

Courts will shape the future here. Authors put their lives into their work, so compensation models—like streaming did for music—need to be figured out.

As LLMs flood GitHub and arXiv with generated content, what happens next?

It’s inevitable. The bigger problem is structural and social. From an AI perspective, human‑curated model output will be part of the pipeline.

I’ve seen waves of likely LLM‑assisted pull requests hit my MLxtend repo. It’s overwhelming, but when a maintainer reviews them, that human‑in‑the‑loop is real value—like free labeling for the ecosystem.

There’s a big difference between raw model output and model output filtered by experts. That filtering encodes judgment.

Exactly. Experts still save everyone time by deciding what matters and framing it well. That’s why good writing still stands out.

Summaries often strip the author’s voice and even core insights. I keep finding that models struggle to preserve the sharp ideas. What do you mean by “voice”?

Voice is the raw, high‑information push to capture a frontier idea. RLHF averages behavior across preferences, which can dull that edge. Models like early Sydney had more voice but obvious safety risks. RLHF trades sharpness for reliability.

That’s a scary trade‑off at scale. Millions use these systems daily.

Users get attached to specific weight deployments and notice tiny changes. A chat model could personalize to you in minutes—that’s powerful and risky. We need to be careful, especially with kids.

The press will tie real tragedies to model conversations, and that pressure will push models toward blandness. But growth requires a little edge. Balancing that is the hard part of RLHF because it touches the human condition.

Many researchers genuinely want to help and think deeply about harm. Some areas—like open image generation on laptops—feel too risky without strong guardrails. Nuance and conviction are needed.

We also need better public dialogue. It’s more complex than “big tech bad.” People inside these labs care about helping real users around the world, and designing one system that fits all is incredibly hard.

I wish AI hadn’t arrived during peak distrust of big tech. AI’s compute cost ties it to hyperscalers, which makes communication harder. I want to talk with people who see AI as a continuation of that story.

Use AI to build, not just consume. When you make apps and ship things, you learn how it works, where it breaks, and you earn the voice to say what’s good or harmful.

I’d rather engage than ignore, but if you hand off the parts you love, you risk losing fulfillment and burning out. If an LLM codes everything for me, am I still proud of what I build two years later?

A survey of seasoned developers showed many now ship large fractions of AI‑generated code, and seniors are more likely to go over half. Most reported their work feels more enjoyable with AI in the loop.

It depends on the task. I happily automate web tweaks, but finding a nasty bug myself is pure joy; a middle ground is trying first, then using the model to avoid frustration. We used ChatGPT to fix dozens of broken show note links in minutes instead of hours.

I like AI as a pair‑programmer because it makes debugging feel less lonely. It won’t always find the bug, but it shortens the desert before that drink of water.

There’s a sweet spot of struggle. Seniors may exploit models better, but how do juniors learn if they outsource too early? I’d keep dedicated offline study time and use models for everything else.

Let’s jump to post‑training. What’s exciting there?

Reinforcement learning with verifiable rewards scaled up, enabling generate‑and‑grade loops, tool use, and inference‑time scaling. That shift changed how teams approach post‑training.

Walk through RLVR as popularized by R1.

Treat tasks with checkable answers—math, code, factual constraints—as environments and reward accuracy; then optimize with RL. Rubric‑based judging widens it beyond strictly verifiable tasks and differs from RLHF’s learned preference model.

RLAIF is the older label, but the cool part here is the model starts writing step‑by‑step and gets better by ‘thinking longer.’ Explanations can be imperfect yet still boost accuracy, and responses grow as training encourages more tokens.

Those ‘aha’ moments are likely amplified behaviors seen in pretraining, not brand‑new cognition. RL is great at turning on patterns that help the model check its work.

I saw a base model jump from about 15 percent to around 50 percent on Math500 within minutes of RLVR. It wasn’t learning new math; it was unlocking what was already there.

But beware contamination and benchmark quirks—some math sets look compromised, which makes fast gains look like formatting hacks. It’s messy to measure cleanly.

Even minor format changes can swing multiple‑choice scores, and we rarely know exactly what the model has seen. Fresh benchmarks are hard, which is the core pain of evaluation.

Give us the post‑training recipe and where RLHF still fits.

Curate mid‑training data with reasoning traces, then push RLVR on hard, verifiable problems with adequate inference budget, and finish with RLHF for tone, formatting, and usability. Compute for RL is rising in wall‑clock hours and memory, even if it uses fewer GPUs simultaneously than pretraining.

RLHF saturates on style, while RLVR keeps paying as you feed tougher problems. Next, I expect process‑aware rewards and self‑grading to matter more as we move beyond ‘question–final answer’ training.

Value functions are having a moment; process reward models have scars from earlier attempts. Crucially, RLVR shows scaling behavior, while RLHF lacks a clean scaling law and can over‑optimize reward models.

For learning, what should a motivated builder do?

Implement a small model on a single GPU to grasp pretraining, SFT, and attention. Transformers is fantastic for production, but it’s dense; load open weights to verify your scratch model and learn by matching outputs.

Transformers started as a fine‑tuning library and grew into the standard way to load and run open models. It’s the default distribution path for frontier weights.

Reverse engineer configs, add missing pieces, and unit‑test against the reference. Wrestling with details like position encodings teaches more than any docs.

Master fundamentals, then go narrow. The field moves fast, but small, focused threads are under‑explored; engaging deeply online puts you closer to the people pushing them.

You can’t follow everything; pick a lane. A curated guide to RLHF and post‑training saves years of wading through conflicting papers.

What important angles did we miss?

Character training is under‑documented, and model specs help separate intent from misses. Preferences are messy—style and accuracy get collapsed into a single score, which is why ideas from social choice matter here.

You both keep returning to struggle. That’s part of real learning.

Education‑tuned models that withhold full solutions and nudge step‑by‑step would be great. The same post‑training tools can teach restraint.

I even use models for puzzle hints with no spoilers; it works, but it takes discipline. Shortcuts are tempting when homework’s due.

You need taste for where to grind and where to speed up.

Expect more blue books and oral exams again; digital submissions are too easy to game now.

Career paths: academia, open labs, or closed frontier labs—how should people think about impact and credit?

Closed labs pay more but you get less public credit; academia is clearer on ownership but financially tougher and unstable. Practically, the well‑paid lab jobs still deliver real impact.

That tension isn’t new; industry always built closed breakthroughs. Startups are the third path—risky but high reward—while publishing remains stressful yet satisfying when your name is on the work.

Professors I know seem happier than friends at frontier labs. The grind is real—996 happens.

Nine to nine, six days a week is creeping in as the norm.

I feel that pressure more in labs and startups now than I did in academia. The leapfrogging race is exhilarating and taxing, and burnout is common.

Competition speeds progress but burns people; some teams normalize extreme hours and even ‘save the marriage’ exceptions. It’s productive—and costly.

No one forced me when I overworked; I was driven and paid for it physically. Passion cuts both ways.

SF’s reality‑distortion field can be productive but blinding. Build inside the bubble, then step outside to rebalance.

There’s a difference between buildout bubbles that create infrastructure and financial bubbles that implode. I worry about crossing that line.

Echo chambers miss many human realities, even if the tech succeeds. Be wary of that narrowing.

Memes like ‘permanent underclass’ show how extreme the SF discourse can get. Being there helps your odds, but it comes with trade‑offs.

Read history, travel, and remember Twitter isn’t the world.

Season of the Witch is a vivid history of SF’s recent past; it adds missing context.

What’s up with text diffusion—can it rival autoregressive models?

Text diffusion borrows denoising ideas from images and fills tokens in parallel, BERT‑style. It can be faster for big outputs, but reasoning and tool‑interrupts are harder, so it’s likely best for cheap, fast tasks; Google’s teased speed gains in this direction.

I’ve heard of startups using it for long code diffs, where seconds matter. Tool calls break the parallel flow, which limits generality compared to autoregressive agents.

How does tool use evolve?

Tools cut errors but won’t erase hallucinations; the model still must pick the right source. Recursive setups that split problems into sub‑calls look promising, but permissions and privacy will slow adoption.

Closed models tightly integrate specific tools in secure clouds; open models must stay flexible across many tools. That pushes open models toward better orchestration and agentic planning.

Continual learning versus feeding huge context windows—what matters more near‑term?

To replace a remote worker, models need to learn from feedback quickly, but rich, personalized context can mimic that in practice. I’m bullish on context and tooling before per‑user weight updates.

So continual learning updates weights, while in‑context learning just feeds information at inference.

We already update global weights between versions; per‑user updates are too expensive unless it’s on‑device. LoRA can personalize cheaply, and memory via context helps preferences, but there’s a learning‑forgetting trade‑off.

And long context?

It’s compute‑ and data‑limited; hybrids that mix attention with state models help. Expect steady gains to a few million tokens while agents learn to compact their histories on the fly.

Think of RNNs compressing too much and transformers remembering everything—hybrids like Unimotron 3 balance both. Recursive approaches can beat brute‑force long contexts, and sparse or sliding attention trims waste.

Teams ship smaller models first because they’re faster to train and test, even if larger models win later.

Quick take on robotics and world models for LLMs?

World‑model ideas—like modeling intermediate variables—can sharpen LLM reasoning beyond brute force. Home robots face endless variability and must learn on the job, which is the hard part.

LLM infrastructure lifts robotics, and I expect more open robotic models and shared data. I’m bullish on warehouses and self‑driving, wary of near‑term in‑home helpers.

Safety in the real world is unforgiving; failure rates must be near zero.

Industrial settings and autonomy on roads look viable sooner; mass manufacturing has non‑technical hurdles that slow timelines.

On AGI timelines, definitions vary and get fuzzy fast.

A pragmatic target is automating most digital remote work; true superintelligence would drive novel science. The line between them often turns philosophical.

AI27’s milestones moved the ‘superhuman coder’ to around 2031; I’d push later.

I disagree with parts but like the concreteness; progress is jagged, with models superb at some code and weak at distributed systems. Big Tech’s spend will give us far better chat and code tools before an automated AI researcher.

Imagine all software fully automated.

By year’s end, a huge slice of coding will be automated, while hairy distributed training remains tough but easier than today.

Think of output per human—fewer people, more useful code.

Engineers will focus on system design and desired outcomes as software creation industrializes.

Like calculators took over calculation, coding will become algorithmic; the open question is whether systems act on their own or still wait for our prompts.

Web tech tolerates messy output, but adding real features to complex systems is a different beast. I’m bullish yet skeptical, because success hinges on clear specs and human skill, not the model reading your mind.

Agents can build end-to-end features in apps like Slack or Word while you guide them from a dashboard, and smaller, cleaner codebases may benefit first. We’ve already seen sandbox rebuilds of full apps, and labs that spend more on inference show many failure modes are simple to fix.

A lot of resistance is cultural, not technical, because people don’t want to work this new way.

In science, teams are betting on reinforcement learning with verifiable rewards in wet labs, and they might hit breakthroughs that look nothing like a chatbot. Most will miss, but a few could supercharge niche experts, like making a PhD mathematician dramatically more effective.

This still looks specialized, not general intelligence, and the real value may be fine-tuned foundation models for big verticals. Competitive advantage will come from private data and custom models, not everyone sharing the same base system.

The bigger question is economic impact: when do these systems move GDP, and can agents use tools reliably enough to do remote work with a failure rate that’s acceptably low?

Computer-use demos remain clunky, because full-screen control is harder than using APIs, and you need sandboxed environments the model can actually operate. The interface is also tough, since users must specify goals and constraints, and the model has to remember you and ask the right clarifying questions.

Defining tasks for arbitrary environments is hard, like booking travel where you must guide each constraint before it can even try.

We’re seeing models lean into memory and targeted prompts, like summarizing what matters and proactively asking questions, which will change social norms. Some of this burns idle compute, but it helps the model learn how to work with you.

To get a true step change, we probably need new architectures or training methods, even as scaling laws keep paying off. Space-based compute is tempting for abundant solar, but heat dissipation is a real constraint that would need serious engineering.

A near-term plateau seems possible where models become incredible helpers in coding, research, shopping, and education while computer use remains hard and inference stays costly.

There’s still obvious headroom to improve, but the biggest gains may cluster in niches rather than delighting every casual user.

The one-system-to-run-your-life dream is giving way to general models that rely on integrations and specialized tools. Progress now comes from better fine-tuning, context engineering, and inference scaling, not just bigger models.

Expect amplification rather than a new paradigm, and watch for simple wins like decent autogenerated figures and diagrams, where models still struggle.

The giant, quiet shift is making the world’s knowledge truly accessible with fewer hallucinations. That changes careers, learning, and everyday decisions for billions.

Use textbooks for dense, linear learning and the model for endless exercises and clarifications, and lean on it for messy, real-time tasks like planning trips where no clean resource exists.

These systems are subsidized now, but ads will arrive, and Google may crack it first because they already have supply. Early versions will be rough, yet a good ad flywheel can fund better models without stripping user agency.

Companies hesitate mainly because competitors have not moved, and the first attempts will feel like clunky promoted posts that risk backlash.

On consolidation, what big moves should we expect?

Licensing-heavy deals are rising and can shortchange rank-and-file compared to full acquisitions, but multi-billion outcomes will keep coming. Cursor is notable because it updates model weights from user feedback roughly every ninety minutes, which is the closest thing to live RL in the wild.

US giants will likely avoid IPOs while fundraising is easy, while Chinese players are already filing; I wish more US labs were public for transparency. I also do not see a winner-takes-all market, since many will compete on similar stacks and margins may push them into products and hardware.

Nvidia’s real moat is the CUDA ecosystem and tooling maturity, not just the chip, though LLMs could speed the creation of alternatives.

Training and inference are splitting, with new inference chips optimized for the prefill phase using little high-bandwidth memory. Nvidia wins while progress is fast because their platform is the most flexible, and Jensen’s hands-on culture keeps them unusually agile.

What about Meta and Llama’s path forward?

The Llama push lost focus amid internal politics, and the brand may morph rather than continue as before.

Chasing benchmark headlines led to overfitting and skipping smaller, practical models people could actually run, and community backlash likely reinforced a pivot away from open.

Hype can mislead; tools can be widely used even when public praise is muted.

The surge of Chinese open-weight models pushed me to start the Adam Project to rally US-backed, truly open models. Early signs include a large NSF grant to AI2, Nvidia leaning in with open models and data, and new funds committing serious capital.

Do you like the federal AI Action Plan’s open-source push?

It’s the clearest policy document so far; the challenge is execution, but putting open models on the national agenda matters.

Open models are essential for educating and training the next generation; you cannot build a talent pipeline behind closed doors only.

Banning open releases would require a domestic firewall and still would not work, since many can afford to train models globally. Chinese open releases actually de-risk stronger US releases, and if progress saturates, efficient open models with standardized serving could dominate.

How much do singular leaders shape this era, and what will history remember?

Individuals accelerate timelines and focus resources; without someone like Jensen, the arc still happens but later. Future histories will credit computing and global connectivity together, with many task-focused agents cooperating rather than one omnibrain.

Luck and bold bets matter: GPUs came from games and science first, which made AlexNet possible because the hardware was buyable.

Will neural networks be remembered as a singular breakthrough because they echo the brain’s structure, or just as one more clever computation?

They were inspired by biology, but they’re digital algorithms that happen to scale well on parallel hardware; other paradigms could have worked, this one just delivered.

Maybe today’s neural nets end up as only one piece of whatever system pushes us toward a singularity.

Like the engine after the Industrial Revolution, a few AI ideas may stick in memory; I’d bet deep learning endures, while transformer could evolve or fade.

Jump a century ahead: are there still humans and are robots everywhere?

I see lots of task-specific robots, some humanoid where environments demand it, and new ways we interact with devices, possibly implants instead of phones or laptops.

But cars show interfaces can persist for a very long time.

People will still want a personal, private brick of compute as a trusted gateway, and I’m not convinced brain-machine interfaces can replace rich visual flows like calendars and email.

Humans will still crave agency and local community, UBI alone doesn’t solve that, and while wealth may surge, building fair infrastructure and policy across nations in a hundred years is a big lift.

Better social supports could soften the transition, but every lost job is a person hurting in the moment.

Amid rising AI slop, I hope in-person conversation and other embodied experiences carry a higher premium.

Expect a boom in physical goods and events, alongside an even wider flood of slop.

Maybe we get so saturated that culture snaps back and prizes the real.

Art won’t vanish; originals carry a presence copies can’t match, and the same pull will apply to writing, talks, and lived moments, even as automation grows.

I often tune out when something reads obviously machine-written.

Soon AI will fool you unless platforms build ways to verify and earn trust, which favors established channels and makes it harder for newcomers.

We’ll lean on trust signals and some authentication, but the best synthetic content will be hard to spot, which makes the trust layer both essential and messy.

One extreme is watermarking all human-captured media by default, the ironic mirror of watermarking AI, though tools can strip marks as easily as they add them.

We’ve focused on promise, but the same capabilities can destabilize society at scale; what keeps you hopeful that we make it through?

I’m cautiously optimistic because humans organize, build community, and solve hard problems; the upside is huge, but we need patient, public, sometimes uncomfortable work beyond just shipping code.

If we put in that effort, the long-term gains can stick.

I’m also excited that AI forces us to look inward and probe mysteries like consciousness, almost like holding up a mirror to the mind.

Consciousness and choice set us apart; today’s AI is a tool that does what we ask, so agency remains with us unless we explicitly program harm.

Even in a sci-fi showdown, I’d bet humans outsmart the machines, maybe with help from local, open models; thanks to Nathan and Sebastian for the work you share and for a great conversation.

Thanks for having us and for the human connection.

Thanks for joining me for this conversation with Sebastian Rashka and Nathan Lambert; I’ll leave you with a few words from Albert Einstein.

Download on the App Store
QR Code - Scan to download

Ready to save time?

Download Shortcast and get started today

Download on the App Store
QR Code - Scan to download