Shortcast
AI Podcast Player

Short podcasts with real voices

Lenny's Podcast: Product | Career | Growth

An AI state of the union: We’ve passed the inflection point, dark factories are coming, and automation timelines | Simon Willison

--% time saved
PodcastLenny's Podcast: Product | Career | Growth
Publisher/creatorLenny Rachitsky
Shortcast updated

About this episode

Simon Willison is a prolific independent software developer, a blogger, and one of the most visible and trusted voices on the impact AI is having on builders. He co-created Django, the web framework that powers Instagram, Pinterest, and tens of thousands of other websites. He coined the term “prompt injection,” popularized the terms “AI slop” and “agentic engineering,” and has built over 100 open source projects, including Datasette, a data analysis tool used by investigative journalists worldwide. What makes Simon unique is that he’s made the leap from traditional software engineering to AI-native development more fully and visibly than almost anyone—and he’s been documenting everything he learns in real time on his blog, SimonWillison.net . In our in-depth conversation, Simon shares: 1. Why November 2025 was the inflection point when AI coding agents crossed from “mostly works” to “actually works” 2. How Simon writes 95% of his code from his phone now and why he’s mentally exhausted by 11 a.m. 3. Why mid-career engineers (not juniors) are most at risk right now 4. The three agentic engineering patterns Simon uses daily (red/green TDD, templates, hoarding) 5. The next leap: the “dark factory” pattern where nobody writes or reviews code and AI does its own QA 6. Why prompt injection is an unsolved security problem and the “lethal trifecta” that will likely lead to an AI Challenger disaster 7. Why the pelican riding a bicycle became the unofficial benchmark for AI model quality — Brought to you by: WorkOS —Modern identity platform for B2B SaaS, free up to 1 million MAUs Vanta —automate compliance, manage risk, and accelerate trust with AI — Episode transcript: https://www.lennysnewsletter.com/p/an-ai-state-of-the-union — Archive of all Lenny's Podcast transcripts: https://www.dropbox.com/scl/fo/yxi4s2w998p1gvtpu4193/AMdNPR8AOw0lMklwtnC0TrQ?rlkey=j06x0nipoti519e0xgm23zsn9&st=ahz0fj11&dl=0 — Where to find Simon Willison: • X: https://x.com/simonw • LinkedIn: https://www.linkedin.com/in/simonwillison • Website: https://simonwillison.net • Agentic Engineering Patterns: https://simonwillison.net/guides/agentic-engineering-patterns — Where to find Lenny: • Newsletter: https://www.lennysnewsletter.com • X: https://twitter.com/lennysan • LinkedIn: https://www.linkedin.com/in/lennyrachitsky/ — In this episode, we cover: (00:00) Introduction to Simon Willison (02:40) The November 2025 inflection point (08:01) What’s possible now with AI coding (10:42) Vibe coding vs. agentic engineering (13:57) The dark-factory pattern (20:41) Where bottlenecks have shifted (23:36) Where human brains will continue to be valuable (25:32) Defending of software engineers (29:12) Why experienced engineers get better results (30:48) Advice for avoiding the permanent underclass (33:52) Leaning into AI to amplify your skills (35:12) Why Simon says he’s working harder than ever (37:23) The market for pre-2022 human-written code (40:01) Prediction: 50% of engineers writing 95% AI code by the end of 2026 (44:34) The impact of cheap code (48:27) Simon’s AI stack (54:08) Using AI for research (55:12) The pelican-riding-a-bicycle benchmark (59:01) The inherent ridiculousness of AI (1:00:52) Hoarding things you know how to do (1:08:21) Red/green TDD pattern for better AI code (1:14:43) Starting projects with good templates (1:16:31) The lethal trifecta and prompt injection (1:21:53) Why 97% effectiveness is a failing grade (1:25:19) The normalization of deviance (1:28:32) OpenClaw: the security nightmare everyone is looking past (1:34:22) What’s next for Simon (1:36:47) Zero-deliverable consulting (1:38:05) Good news about Kakapo parrots — References: https://www.lennysnewsletter.com/p/an-ai-state-of-the-union — Production and marketing by https://penname.co/ . For inquiries about sponsoring the podcast, email [email protected] . — Lenny may be an investor in the companies discussed. To hear more, visit www.lennysnewsletter.com

Loading episode data...

Episode summary

I woke up to a new reality where agents run the loop, I ship massive amounts of code without typing most of it, and I can prompt from my phone while walking the dog, so I set a wild resolution to take on more.

It’s a funny twist: AI boosts output, yet the most AI‑pilled folks feel busier than ever.

Running multiple agents well soaks up decades of engineering judgment, and when I spin up four in parallel I’m mentally spent before lunch.

You’ve warned we’re headed for a Challenger‑style AI failure because we’re normalizing risky behavior.

Each safe launch breeds overconfidence, and I expect a high‑profile incident will snap us back to responsible use.

My guest is Simon Wilson—co‑creator of Django, originator of terms like prompt injection, AI slop, and agentic engineering, builder of the dataset tool, and a rare voice who shares what he learns while building; give us the November inflection and what’s possible now.

In 2025 the big labs treated code as the killer app, stacked models with reasoning, and in November crossed a reliability line where agents usually deliver working software, which jolted engineers who returned from holidays to find they could describe a tool and get a solid build.

You’re coding on your phone—what’s next beyond that?

Vibe coding lets anyone describe a thing and play with it without reading code, which is liberating for personal tools but risky for others, while agentic engineering is the pro discipline of guiding test‑running agents toward production quality.

The frontier is dark‑factory building, where you don’t even read the code yet still enforce standards and safety at scale.

So what does this factory actually do, and who steers it?

First no one types, then no one reads: StrongDM showed a swarm of agent testers acting like users in a simulated Slack and Jira, hammering the product around the clock so quality emerges without manual review, even if it burns serious tokens.

Couldn’t the labs just bundle that?

You can wire agents to browsers and tests, but the leap is redesigning process, and meanwhile agents are getting credible at security work—labs even run restricted models that helped find real Firefox bugs—though unvetted AI vuln reports still waste maintainers’ time.

The build and review middle is filling in fast, so the gap is deciding what to make; can AI help act like a PM?

Because builds take hours not weeks, I spin up multiple prototype directions, then I need real people to try them since simulated users do not surface the right problems.

Where do human brains still shine?

Models are great for first‑wave brainstorming and weird mashups that spark ideas, while humans judge, combine, and pick what’s worth pursuing.

How does this reshape careers—do 10x veterans get even stronger, and what about newcomers?

Agents amplify expertise and warp effort estimates, interns ramp shockingly fast, and the squeezed segment is mid‑level folks who lack both deep patterning and the newcomer’s onboarding lift.

What’s your advice to the middle so they don’t get left behind?

Lean in to learn faster, use AI to climb steeper hills, and double down on your own agency because models don’t have it.

Ambition matters now; leaders like Jensen argue companies need bigger visions to match the new leverage.

I flipped my resolution to take on more, and while it may cost focus, it’s been energizing.

Still, people feel fried even as output rises.

I do have more hours, but the cognitive load is heavier, so good teams will watch expectations or burn out their best builders.

The flip side is joy—people are burning down decade‑old side‑project lists.

That speed also elevates craft, because I only trust code I’ve lived with, so we need proof of real usage not just pretty tests, and there’s even a market now for pre‑AI “handwritten” repos.

When will half of engineers let AI write almost all their code?

Given cultural differences I’ll say approximately 95 percent AI‑authored code could be common by year’s end for many, with the real challenge being skillful use rather than model quality.

Openings for engineers, PMs, and even recruiters are climbing again despite layoffs.

Hiring signals are noisy because AI now writes both resumes and job posts, so we need better reads on what’s truly changing.

Walk us into agentic engineering—the key skill shift you’re writing about.

Typing stopped being the bottleneck, so the work is designing, validating, and keeping quality high while using cheap prototypes to choose direction without piling up slop.

What’s in your stack right now?

I mostly use Claude Code for Web linked to GitHub so I can run YOLO mode safely from my phone, I also lean on GPT 5.4 which is excellent and cheaper, I disable sticky memories and even ported them across models with a prompt, I do research through model‑integrated search, and I only generate images for fun.

And the pelican benchmark?

My SVG pelican‑on‑a‑bicycle test became a weirdly reliable proxy for overall capability, with better models drawing better birds and leaving delightful code comments, and I keep backup animals in case labs overfit because, honestly, the absurdity is half the joy.

Quick detour about how hard it is to even sketch a bicycle, then back on track: you’ve got a pattern called hoarding what you know—what does that look like in practice?

It’s career fuel: stockpile small wins and failed experiments so you can remix them on demand, and AI makes that library grow fast with cheap prototypes and code you can revisit.

So you’re saving bite‑size tools and research notes—where does this live and how public do you make it?

I default to public on GitHub for durability and discoverability, with repos of tiny HTML and JavaScript utilities plus AI‑run research write‑ups, and I keep private stashes and a mountain of personal notes for the rest.

How do you actually use that archive—do you feed it to the model or point it to sources as needed?

Both: I ask the model to read specific repos, stitch code together, and search across my corpus; for example, it merged a PDF viewer with in‑browser OCR to process documents page by page.

Let’s hit testing: you champion red slash green TDD and starting by running the test—why is that essential with coding agents?

Agents must execute code, and tests force them to do it; a growing suite guards against regressions and, counterintuitively, speeds you up far more than skipping tests ever will.

I’ll tell the agent to use red slash green TDD—write the test, watch it fail, then make it pass—and that tiny prompt reliably upgrades quality without a wall of instructions.

And for folks worried AI code is brittle, practices like this are why it can be trusted in the loop.

I now tolerate large test suites because agents happily maintain the boring parts, and that cheap safety net is worth it.

Any other pattern before we pivot?

Start new projects with a slim template—a single test and your preferred layout—because agents mirror whatever patterns you give them, and I publish starter skeletons for common stacks.

Onto security: you coined prompt injection and later the lethal trifecta; why is this so risky, and what do you expect to happen?

Prompt injection is an application‑level flaw where outside text can overrule your instructions, and the lethal trifecta appears when an agent has private data, a channel for hostile prompts, and a way to send data back—break any one leg or you’re exposed.

Why can’t we just tell the model to ignore tricks?

Rule‑based filtering tops out around the high nineties, which is still failure territory, so the saner approach is reducing blast radius and resisting the slow slide into risky normal—until a Challenger‑style wake‑up event finally hits.

Are we getting closer to a fix, or is this permanently messy?

Detection scores keep rising, but without a real proof they only lull us into false confidence, and I don’t yet see a clean, complete solution.

Best mitigations you’ve seen?

The Camel approach splits a privileged agent from a sandboxed one, tracks tainted inputs, and only pings a human for high‑risk actions; it’s complex but promising.

And when agents control machines in the real world, the stakes jump.

Exactly—and that’s the nightmare scenario.

Last on security: OpenClaw exploded in months but had notable gaps; what’s your read?

It’s the exact high‑risk assistant everyone craves—email access, action tools—and despite real incidents, demand was so strong people endured painful setup; the big prize is a version that’s genuinely safe, and while I admire the project, I only poke at it inside a container on a dedicated box—think digital pet in its own tank.

The vibes are undeniable, and that personality is why copycats keep appearing.

I love that “claws” became a category and plan to build a simple one myself; also, the Spider‑Man Doc Ock reference fits a little too well.

What’s next for you—projects, goals, and the book?

Day to day I ship open‑source tools for data journalists, using AI as an untrusted aide to extract and structure facts from messy sources, and I’d love to see my software be a small part of a Pulitzer‑winning investigation; the not‑a‑book keeps rolling, my blog now pays modestly, and I do zero‑deliverable, hour‑long consulting routed through partners.

How should folks reach you if they want that hour?

I’ll keep that behind a small puzzle—if you really want it, you’ll find me.

Anything to leave us with?

Joyful wildlife news: New Zealand’s flightless kākāpō are finally having a strong breeding season thanks to rimu trees fruiting again, with dozens of new chicks and live nest cams to watch.

Best note to end on; thanks for listening—subscribe on your favorite app, drop a rating or review, and find past episodes at Lenny’s podcast dot com.

Download on the App Store
QR Code - Scan to download

Ready to save time?

Download Shortcast and get started today

Download on the App Store
QR Code - Scan to download