About this episode
ARC-AGI is redefining how to measure progress on the path to AGI - focusing on reasoning, generalization, and adaptability instead of memorization or scale. During this month's NeurIPS 2025 conference, YC's Diana Hu sat down with ARC Prize Foundation President Greg Kamradt to find out why most AI benchmarks fail, how ARC-AGI reveals the limits of today’s models, and why measuring intelligence may be harder than building it.
Listen to the original episode
Episode summary
Thrilled to welcome Greg Kamrat, President of the Art Prize, joining us here in beautiful San Diego. What does the foundation actually do?
We’re a nonprofit with a strong point of view: pull forward progress toward systems that generalize like people, not just memorize or scale.
François Chollet frames intelligence as learning new skills efficiently, not acing harder tests; what should founders chasing MMLU take from that?
Intelligence is about the capacity to pick up new skills, and ARC was built to test exactly that; unlike ever-harder academic quizzes, ARC tasks are solvable by regular people, which exposed how poorly early base models handled novelty.
For context, large language models before 2024 struggled on ARC, right?
Yes; the base GPT‑4 without explicit reasoning was around four to five percent, then models like O1 jumped to about twenty‑one percent, signaling a real shift in reasoning.
Big labs now include ARC‑AGI in their model releases.
We’re glad frontier teams report ARC results; OpenAI, XAI, Google, and Anthropic have all shared performance, which helps the community see what matters.
What’s going well, and what gives you pause?
Adoption is great, but we watch for vanity score‑keeping; our mission is to catalyze open research and small teams focused on generalization, not just leaderboards.
Common false positives you see when teams ship AI products?
Overfitting with bespoke reinforcement learning environments can look like progress but rarely transfers; ARC uses a hidden test set and rewards systems that generalize without training on the target setup.
Give us the arc of ARC‑AGI: versions one, two, and what three changes.
ARC 1 launched in 2019 with hundreds of tasks and the original measure-of-intelligence framing; ARC 2 arrived in 2025 as a deeper static suite; ARC 3, coming next, is interactive with roughly one hundred fifty game-like environments where you must infer the goal with no instructions.
We validate every ARC 3 environment with everyday people and exclude anything regular folks can’t solve, so if humans clear it and AI can’t, we know there’s a missing capability.
There’s growing interest in evaluating models in human terms, beyond accuracy; how close are we to that?
Wall‑clock time mostly reflects compute, so we focus on data and energy; ARC 3 will measure the number of actions taken to solve a task and compare that to human action counts, blocking brute‑force strategies like the old Atari runs that spammed millions of frames.
Magic wand: a team scores one hundred percent on ARC‑AGI tomorrow; what should the world update?
Solving ARC is necessary evidence of broad generalization but not sufficient for AGI; beating ARC 3 would be the strongest signal yet, and we’d want to study the system, map its failure modes, and keep guiding the field toward a credible AGI declaration.
Great place to end. Thanks for joining us, Greg.