About this episode
(0:00) Intro live from Nvidia GTC (0:37) CoreWeave CEO, Michael Intrator (32:58) Perplexity CEO, Aravind Srinivas (1:07:11) Mistral CEO, Arthur Mensch (1:18:57) IREN CEO, Daniel Roberts Our episode is sponsored by the New York Stock Exchange - a modern marketplace and exchange for building the future. It all happens at the NYSE - https://nyse.com Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@theallinpod Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg Intro Video Credit: https://x.com/TheZachEffect
Listen to the original episode
Episode summary
Live at NVIDIA’s GTC, I’m sitting down with four powerhouse AI CEOs—let’s dive right in.
Michael, you were early to the GPU wave and built what some called a new kind of cloud; how did CoreWeave get there so fast, and what did you build?
We started as hedge‑fund nerds using GPUs for crypto, then climbed into CGI rendering, batch science, and finally neural nets—donating early A100s to the EleutherAI community was our tuition for learning large‑scale parallel compute.
So you focused the business on specialized AI infrastructure instead of a do‑everything cloud?
Yes, we built a purpose‑built layer that lives above NVIDIA’s stack and beneath the models, obsessing over software, observability, and operations to deliver clusters that can train world‑changing models at scale.
When did model builders start knocking, and what are you seeing now on inference?
Eleuther was first, Inflection was our first big commercial win, and then we expanded into the major model labs; today inference is exploding—it’s how AI investments monetize—and cutting‑edge GPUs train first, then move into experiments and long‑lived inference.
There’s debate about GPU obsolescence—what’s the real shelf life and how do you secure supply?
Our average contract runs five years and we depreciate over six, while Ampere pricing even appreciated this year because new entrants and use cases keep buying; NVIDIA allocates on a first‑in, first‑out basis, so we show up with POs and build for clients who fit our long‑term model.
Walk us through the financing that lets you scale so fast.
We package each project into a single box—customer contract, GPUs, and data‑center agreements—so revenue waterfalls to pay power, principal, and interest before returns flow back; that structure helped us raise about $35 billion in 18 months, drive payback to roughly two and a half years on five‑year deals, and cut our cost of capital by around 600 basis points.
How deep is demand, and what’s constraining growth right now?
Demand has been relentless for years, and the choke points extend beyond GPUs to memory, networking, optics, and power; memory’s cyclical under‑investment collided with AI’s surge, but capitalism’s boom‑bust cycles lay foundations—like fiber did for YouTube—while token costs keep plunging and AI lowers the barrier for anyone with a good idea.
Aravind, Perplexity’s been all about accuracy and live knowledge; what is Perplexity Computer and how do you handle local access and trust?
Computer is an AI that becomes your computer, orchestrating many specialized models and sub‑agents like instruments in a symphony; for trust, our Personal Computer syncs with a local Mac mini so sensitive orchestration can run on your hardware, while long or heavy tasks can, with permission, shift to your private server.
Where do local models and the operating system fit, and what’s unique about your browser?
AI becomes the OS by taking objectives and coordinating tools, files, and connectors, often running on stable, customizable Linux you can access from your phone; our Comment browser lets the agent natively control web tasks, so automations that still live in the browser execute alongside everything else in Computer.
How’s enterprise adoption and pricing, and how do you compete with giants while routing across models?
We serve tens of millions of consumers and thousands of companies, with enterprise tiers at about $40 and $400 per user per month, and every dollar we earn carries positive gross margins because we route smartly and avoid bloated contexts; we’re model‑agnostic—GPT, Claude, Gemini, Llama, DeepSeek, Kimi, Qwen, and more—and auto‑route by task, while Model Council compares outputs and highlights agreements and differences.
You ship fast—any examples of agents creating bespoke tools, and what unlocked the orchestration leap?
Speed is our moat, helped by AI coding so even non‑engineers can iterate through Slack; Computer now drafts our board memos and press briefs, and with advances like Claude Opus 4.5 plus sandboxed tools, models manage long, multi‑step workflows by pulling only the context they need.
What’s next, and how do you think about job displacement?
We’re pushing toward near‑autonomous small businesses—run ads, integrate Stripe, ship features, and handle support while you focus on ideas; some roles will shift, but this is a big unlock for personal agency and entrepreneurship if people lean in and learn the tools.
Mobile and APIs: can I run Computer on the go and get cooperative access to sites?
Computer already runs in our app, and Comment’s server‑side browsing powers web automations; we’re eager to partner on official APIs so users can authorize well‑behaved agents and even pay for access—a win for platforms and customers.
Arthur, big news with NVIDIA—what are you building and why open source, especially for verticals?
We’re training the next frontier models with NVIDIA to release the best open models, then specializing them through products like Forge and Studio for domains like finance, engineering, science, and government; open lets customers cut costs, deploy anywhere, and go deeper by modifying both the model and the orchestration to match their proprietary workflows.
How do you fine‑tune securely, what’s the role of synthetic data, and what did the recent agent wave signal for enterprises?
We ship a portable platform to the customer’s infrastructure, forward‑deploy scientists to work with subject experts, and transfer the know‑how so they can retrain in place with no data flowing back; synthetic data warms up and compresses, but human signals are essential, and while agent autonomy is exciting, enterprises need governance, observability, role‑based access, and deterministic gates to run mission‑critical processes safely.
Daniel, you and your brother built data centers from the Bitcoin era—how did you pivot to AI and how big is the build today?
We used Bitcoin mining to bootstrap large‑scale sites, then began swapping in AI chips as demand took off; we develop the land, permits, and grid ourselves, with a 750‑megawatt flagship in Texas and about 4.5 gigawatts of power lined up, and a $9.7 billion Microsoft deal that’s only around five percent of our capacity.
What’s the main constraint now, and what does building at that scale mean for workforce and communities?
Across the industry power is tight, but for us the hurdle is time to compute—thousands of skilled workers, stressed supply chains, and constant problem‑solving—so we hire locally first, invest in the community, retrain workforces around legacy electrical infrastructure, and partner with schools as wages rise for in‑demand trades.
AI is chewing through talent, and energy is the other headline—Texas has wind, solar, gas, oil, even nuclear in the mix. How do you think about powering this surge?
We designed for sustainability from day one and have run on 100 percent renewables since launch, using hydro in British Columbia and wind and solar in West Texas. That region has about 45 to 50 gigawatts of renewables but only around twelve gigawatts of transmission, so we put data centers at the source and turn excess power into compute that moves at fiber speed.
Distance makes moving power expensive, so you follow the turbines and panels—got it. How do you handle lulls without giant on‑site batteries?
Utilities manage intermittency and value those scarce grid interconnects, which is why they’re so hard to get. Once connected, we get reliable, around‑the‑clock power, and the utility can balance with other sources if needed.
So they smooth it on their side, even if you’re contracted to renewables—makes sense. With rumors of some teams tapping the brakes, is demand actually cooling or still red‑hot?
It’s full throttle. The constraint is time to deliver compute, and there are effectively no idle GPUs anywhere.
Nvidia says software and networking will slash token costs many times over—does that bend the curve down or fuel even more use?
Faster, cheaper outputs drive more usage; it’s Jevons paradox. If image generation drops from minutes to five or ten seconds, people will create far more.
Custom silicon is coming from the hyperscalers—has it truly landed in data centers, or does Nvidia still set the pace?
Many chips are seeking homes, but Nvidia has a large head start with mature tools and standards. Following their roadmap is the safest path for scale right now, while alternatives continue to mature.
On the other end, power users want monster desktops to run open models locally—flashback to hacker roots. Is that real or just a fad?
It’s real. Software breakthroughs are empowering everyone from home tinkerers to mega‑clusters, and agents, autonomy, and robotics will compound that demand at every level.
Let’s talk nuclear—new designs are cleaner and safer. Are you leaning into that?
We have to. Big nuclear and modular reactors will take roughly a decade, so now is the time for policy, capital, and planning.
Do you have sites near nuclear yet, and what happens if small modular reactors can sit beside data centers?
Not yet, but we’re tracking it closely, and colocated SMRs would unlock more clean generation, strengthen US competitiveness, and enable more distributed compute—another demand flywheel.
Inside the buildings, the network fabric is changing fast. How do you think about that backbone?
Treat the whole building like one machine. The cabling, hop count, and latency between GPUs—whether over InfiniBand or Ethernet—directly set cluster performance, so every millisecond matters.
What about pushing data centers into space—visionary or premature?
He’s called a lot right, but today launch costs and radiation make it extremely hard. That’s never stopped him, though.
Last one—can you get data out of a desert site without latency headaches, or does fiber become the choke point?
That’s a myth; Texas is laced with buried fiber. Our round‑trip to Dallas is about six milliseconds, which users don’t notice.
Love it—continued success. You’re hiring, right?
We’ve got around 129 open roles.
Fantastic. Thanks for hanging with us here at All‑In at GTC.