Shortcast
AI Podcast Player

Short podcasts with real voices

Y Combinator Startup Podcast

Scaling and the Road to Human-Level AI | Anthropic Co-founder Jared Kaplan

--% time saved
PodcastY Combinator Startup Podcast
Publisher/creatorY Combinator
Published
Shortcast updated

About this episode

Jared Kaplan on June 16th, 2025 at AI Startup School in San Francisco.Jared Kaplan started out as a theoretical physicist chasing questions about the universe. Then he helped uncover one of AI’s most surprising truths: that intelligence scales in a predictable, almost physical way.That insight became foundational to the modern era of large language models—and led him to co-found Anthropic.In this talk, he walks through how that discovery reshaped the path to human-level AI, what it means for future models like Claude, and why even the dumbest questions can lead to the biggest breakthroughs. He reflects on memory, oversight, and what’s left to solve as models grow smarter—and longer-horizon tasks come within reach.

Loading episode data...

Episode summary

Hey everyone, I’m Jared Kaplan. I started as a theoretical physicist chasing sci‑fi questions from my mom’s influence—could we build faster‑than‑light drives, is the universe deterministic, do we have free will? Over time I got frustrated with slow progress in physics and skeptical of AI—then I saw the data. Modern AI has two phases: pre‑training to predict the next word across vast human text and multimodal data, then reinforcement learning to shape behavior toward helpful, honest, harmless outcomes. The surprising part is how clean the scaling laws are in both phases—predictable gains across many orders of magnitude in data, parameters, and compute, like straight lines in physics. RL shows similar trends; even simple games like Hex revealed linear Elo gains with more compute. Capabilities expand along two axes: flexibility across modalities, and the time horizon of tasks. A team found the task horizon is doubling roughly every seven months. That hints we’re on a smooth path toward models that can take on projects spanning days, weeks, and eventually organizational‑scale work. What’s left to unlock broadly human‑level AI? Organizational knowledge so models operate with deep context, durable memory to carry state across long tasks, and better oversight to train on nuanced, fuzzy goals—not just crisp code tests. Also keep pushing from text to multimodal to robotics. My advice: build things that almost work—models are improving fast. Use AI to integrate AI. And find domains where adoption can go from zero to one quickly. Alright, let’s bring Diana up for a chat.

Awesome talk. With Claude 4 now out, how does this change what is possible over the next twelve months as releases compound?

I hope it is not twelve months until the next big step. With Claude 4 we focused on better agentic behavior—especially for coding—stronger supervision so it follows directions and improves code quality, and practical memory so it can persist work across many context windows by saving and retrieving files. Expect steady, incremental improvements—scaling looks like a smooth curve toward broadly human‑level capability.

Any alpha features people here will love in the new APIs?

Memory. I’m most excited about unlocking longer‑horizon tasks so Claude can collaborate on larger and larger chunks of work. Yes—imprecise, but for software tasks that is about right. AI can be brilliant and also make basic errors; its judgment is closer to its generation ability than ours is, so human review still matters for the hardest work.

We saw teams shift from co‑pilots needing human approval to end‑to‑end automation. How do you see builders here leaning into that?

It depends on what success rate is acceptable—some tasks are fine at seventy to eighty percent, others need ninety‑nine point nine percent. The frontier is more fun where partial correctness is useful, and reliability is rising. Near‑term, human‑in‑the‑loop will dominate advanced tasks; long‑term, more will be fully automated.

Paint your picture of human‑AI collaboration from here.

We already see it in biomed research with the right orchestration. Think breadth and depth: models ingest civilization‑scale knowledge in pre‑training, so they can synthesize across biology, psychology, and history in ways no single expert can, while we keep improving on deep, hard problems like math and coding. Scaling predicts the trend; exact implementation paths are hard to forecast. Any role in front of a computer working with data: finance, heavy spreadsheet work, and law—though regulation is tougher. The huge lever is integrating AI into existing businesses. Like electricity, the win is not swapping a steam engine for a motor; it is redesigning the factory. Use AI to accelerate AI integration.

How did your physics training shape your AI research?

By forcing me to ask big, dumb, precise questions and make trends quantitative. Is learning exponential or a power law? What does it mean to truly move the needle? The holy grail is a better slope on the scaling law—more capability per unit compute. Getting precise lets you know if you are actually beating the trend.

Any physics heuristics that proved directly useful?

Large‑matrix approximations matter because networks are huge. But mostly, the field is young; naive questions often beat fancy techniques. Many basics—like interpretability—are still open. Interpretability feels more like neuroscience or biology. The advantage in AI is we can measure everything—every unit and connection—so reverse‑engineering is far more tractable.

What would convince you the scaling curves are changing?

I use scaling to diagnose training. When it looks broken, it is usually our fault—architecture, bottlenecks, precision issues. It would take a lot to convince me the empirical laws stopped working rather than us training wrong.

As compute gets scarce, how far down the precision ladder would you go—FP4, ternary?

We chase frontier capability while driving efficiency. We are seeing roughly three to ten times yearly gains from algorithms and inference efficiency. Lower precision will be a major lever—joke being we will get computers back to binary. In a steady state, AI could be extremely cheap, but if capability keeps leaping, we will keep prioritizing intelligence over, say, FP2 perfection.

Classic Jevons paradox: better intelligence drives more demand.

Totally. And I suspect much of the value sits at the frontier because long‑horizon, end‑to‑end tasks are far more convenient than stitching many tiny steps with weaker models—though great integrators can extract value across the spectrum.

What should early‑career folks here master to stay relevant?

Learn how these models actually work, get great at leveraging and integrating them, and build at the frontier. Let’s take audience questions.

You showed linear scaling trends, yet the time‑saved curve looked exponential. Why the switch?

Great question—I do not know for sure. My mental model is self‑correction: modest improvements in noticing and fixing mistakes can double how far you get before failing, which compounds horizon length. That could look exponential even if underlying capability scales smoothly. The empirical trend is the anchor; theory can catch up.

To extend horizons, do we just scale data and verification signals—easy in coding, harder elsewhere—or is there a better approach?

Worst‑case, we grind: build ever more complex, long‑horizon tasks and train with RL. Given the value, people will do it. Better is AI supervising AI with granular feedback so we are not waiting years for an end‑to‑end signal. Practically, we mix AI‑generated tasks and human‑crafted ones. As models improve, we lean more on AI—even as the frontier keeps rising. Thanks, everyone.

Download on the App Store
QR Code - Scan to download

Ready to save time?

Download Shortcast and get started today

Download on the App Store
QR Code - Scan to download