About this episode
Aishwarya Naresh Reganti and Kiriti Badam have helped build and launch more than 50 enterprise AI products across companies like OpenAI, Google, Amazon, and Databricks. Based on these experiences, they’ve developed a small set of best practices for building and scaling successful AI products. The goal of this conversation is to save you and your team a lot of pain and suffering. We discuss: 1. Two key ways AI products differ from traditional software, and why that fundamentally changes how they should be built 2. Common patterns and anti-patterns in companies that build strong AI products versus those that struggle 3. A framework they developed from real-world experience to iteratively build AI products that create a flywheel of improvement 4. Why obsessing about customer trust and reliability is an underrated driver of successful AI products 5. Why evals aren’t a cure-all, and the most common misconceptions people have about them 6. The skills that matter most for builders in the AI era — Brought to you by: Merge —The fastest way to ship 220+ integrations: https://merge.dev/lenny Strella —The AI-powered customer research platform: https://strella.io/lenny Brex —The banking solution for startups: https://www.brex.com/product/business-account?ref_code=bmk_dp_brand1H25_ln_new_fs — Transcript: https://www.lennysnewsletter.com/p/what-openai-and-google-engineers-learned — My biggest takeaways (for paid newsletter subscribers): https://www.lennysnewsletter.com/i/183007822/referenced — Get 15% off Aishwarya and Kiriti’s Maven course, Building Agentic AI Applications with a Problem-First Approach , using this link: https://bit.ly/3V5XJFp — Where to find Aishwarya Naresh Reganti: • LinkedIn: https://www.linkedin.com/in/areganti • GitHub: https://github.com/aishwaryanr/awesome-generative-ai-guide • X: https://x.com/aish_reganti — Where to find Kiriti Badam: • LinkedIn: https://www.linkedin.com/in/sai-kiriti-badam • X: https://x.com/kiritibadam — Where to find Lenny: • Newsletter: https://www.lennysnewsletter.com • X: https://twitter.com/lennysan • LinkedIn: https://www.linkedin.com/in/lennyrachitsky/ — In this episode, we cover: (00:00) Introduction to Aishwarya and Kiriti (05:03) Challenges in AI product development (07:36) Key differences between AI and traditional software (13:19) Building AI products: start small and scale (15:23) The importance of human control in AI systems (22:38) Avoiding prompt injection and jailbreaking (25:18) Patterns for successful AI product development (33:20) The debate on evals and production monitoring (41:27) Codex team’s approach to evals and customer feedback (45:41) Continuous calibration, continuous development (CC/CD) framework (58:07) Emerging patterns and calibration (01:01:24) Overhyped and under-hyped AI concepts (01:05:17) The future of AI (01:08:41) Skills and best practices for building AI products (01:14:04) Lightning round and final thoughts — Referenced: • LevelUp Labs: https://levelup-labs.ai/ • Why your AI product needs a different development lifecycle: https://www.lennysnewsletter.com/p/why-your-ai-product-needs-a-different • Booking.com : https://www.booking.com • Research paper on agents in production (by Matei Zaharia’s lab): https://arxiv.org/pdf/2512.04123 • Matei Zaharia’s research on Google Scholar: https://scholar.google.com/citations?user=I1EvjZsAAAAJ&hl=en • The coming AI security crisis (and what to do about it) | Sander Schulhoff: https://www.lennysnewsletter.com/p/the-coming-ai-security-crisis • Gajen Kandiah on LinkedIn: https://www.linkedin.com/in/gajenkandiah • Rackspace: https://www.rackspace.com • The AI-native startup: 5 products, 7-figure revenue, 100% AI-written code | Dan Shipper (co-founder/CEO of Every): https://www.lennysnewsletter.com/p/inside-every-dan-shipper • Semantic Diffusion: https://martinfowler.com/bliki/SemanticDiffusion.html • LMArena: https://lmarena.ai • Artificial Analysis: https://artificialanalysis.ai/leaderboards/providers • Why humans are AI’s biggest bottleneck (and what’s coming in 2026) | Alexander Embiricos (OpenAI Codex Product Lead): https://www.lennysnewsletter.com/p/why-humans-are-ais-biggest-bottleneck • Airline held liable for its chatbot giving passenger bad advice—what this means for travellers: https://www.bbc.com/travel/article/20240222-air-canada-chatbot-misinformation-what-travellers-should-know • Demis Hassabis on LinkedIn: https://www.linkedin.com/in/demishassabis • We replaced our sales team with 20 AI agents—here’s what happened | Jason Lemkin (SaaStr): https://www.lennysnewsletter.com/p/we-replaced-our-sales-team-with-20-ai-agents • Socrates’s quote: https://en.wikipedia.org/wiki/The_unexamined_life_is_not_worth_living • Noah Smith’s newsletter: https://www.noahpinion.blog • Silicon Valley on HBO Max: https://www.hbomax.com/shows/silicon-valley/b4583939-e39f-4b5c-822d-5b6cc186172d • Clair Obscur: Expedition 33: https://store.steampowered.com/app/1903340/Clair_Obscur_Expedition_33/ • Wisprflow: https://wisprflow.ai • Raycast: https://www.raycast.com • Steve Jobs’s quote: https://www.goodreads.com/quotes/463176-you-can-t-connect-the-dots-looking-forward-you-can-only — Recommended books: • When Breath Becomes Air : https://www.amazon.com/When-Breath-Becomes-Paul-Kalanithi/dp/081298840X • The Three-Body Problem : https://www.amazon.com/Three-Body-Problem-Cixin-Liu/dp/0765382032 • A Fire Upon the Deep : https://www.amazon.com/Fire-Upon-Deep-Zones-Thought/dp/0812515285 — Production and marketing by https://penname.co/ . For inquiries about sponsoring the podcast, email [email protected] . — Lenny may be an investor in the companies discussed. To hear more, visit www.lennysnewsletter.com
Episode summary
We worked on a guest post together and the core idea stuck with me: building AI products is not the same as building traditional software. Today I’m joined by Aishwarya and Kiriti to share what they’ve learned shipping dozens of AI deployments, with one goal for this episode—save you time, pain, and false starts; so, on the ground, what’s going well and what’s not?
This year feels different—skepticism is way down and teams are rethinking flows instead of just slapping chat on their data. Execution is still messy though, because the AI lifecycle breaks old handoffs and forces tighter, shared feedback loops across PMs, engineers, and data folks.
Let’s hit the two big differences you’ve been emphasizing that change how we build.
First, non‑determinism on both sides: users express intent in countless ways, and models respond probabilistically, so neither inputs nor outputs are predictable. Second, autonomy versus control: the more decision‑making you hand to an agent, the more governance you give up, so the system must earn trust before you expand its freedom.
Start small and train up, like preparing for a big hike by doing shorter trails first. Begin with low‑risk tasks under strong human control, then increase autonomy only as you gain confidence.
Can you walk through a concrete example of that progression?
Take support: begin with AI suggestions for human agents to accept or edit, learn where it fails, then let it reply directly for common cases, and only later add actions like refunds or creating tickets. Stacking complexity slowly keeps it controllable.
Think of it as behavior calibration. Constrain autonomy where risk is high, like pre‑authorizing simple tests while keeping invasive procedures human‑reviewed, and log human actions to power a steady improvement loop without breaking trust.
You’ve shown similar step‑ups for coding assistants and marketing tools, and the throughline is that user inputs and model outputs both vary, so you build confidence in stages.
The wild part is also the good part—natural language lowers friction and feels human. The hard part is mapping that fluidity back to reliable outcomes.
And when teams jump straight to full autonomy, they often stall out or label the whole approach a failure.
Problem‑first thinking keeps you grounded. Tight scope with limited autonomy clarifies what you’re actually solving and reduces the blast radius when things go wrong.
Reliability is the top blocker—roughly seventy five percent of enterprises cite it—and that’s why most wins so far live in productivity helpers, not end‑to‑end automation.
We also touched on adversarial risks like prompt injection and jailbreaks, which only get scarier as agents act on the world.
Security will loom large as these systems go mainstream, since non‑deterministic behavior widens the attack surface for injected instructions.
I’m still optimistic—adoption is early, and with sensible human checkpoints you can focus on process gains while containing risk.
What working patterns do you see in teams that consistently ship successful AI products?
It’s a triangle of leadership, culture, and technical craft. Leaders must rebuild instincts by getting hands‑on—one CEO I worked with blocked early mornings for AI deep dives—while culture should empower subject matter experts rather than scare them, and teams need a surgical grasp of workflows, mixing deterministic code with ML, building data flywheels, and ignoring one‑click agent hype because enterprise reality takes months.
On that note, leaders who live in the tools seem to drive the biggest gains; let’s talk eVals, since that’s been a hot debate—do they solve it or are they overrated?
It’s a false choice. Offline evals encode your product judgment, while production monitoring surfaces real failures via explicit and implicit signals like regenerations or feature toggles; you need both because each catches different classes of issues.
Also, the word “evals” now covers everything from expert error notes to PM checklists to model benchmarks, which confuses teams; focus on building actionable feedback loops, and remember LLM judges are brittle in complex domains, so sometimes user signals plus quick fixes beat elaborate judges.
How does the Codex team handle this balance in practice?
We use guardrail evals for core behaviors and lean hard on customer feedback, A/B tests, and hands‑on trials, because coding agents are highly customizable and no static set covers all integrations. We ship, watch real usage, and iterate quickly.
Let’s unpack your continuous calibration, continuous development framework that ties all this together.
We created CCCD after painful end‑to‑end agent projects, including one we had to shut down, and seeing incidents like a refund policy hallucination. Continuous development means scoping capability, curating examples, aligning on desired behavior, setting metrics, and deploying, while continuous calibration means inspecting traces, spotting new error patterns, adding or refining metrics when needed, and fixing issues; start with low autonomy and human‑in‑the‑loop, then progress from routing to copilot drafts to autonomous resolution, logging edits to fuel the flywheel, and constrain autonomy by action count or by topic risk.
How do you know when it’s time to step up autonomy?
Advance when surprises taper off. Expect to recalibrate after model changes or shifts in user behavior—for example, underwriters moved from simple lookups to deep historical questions, which required a very different system.
What feels misunderstood or underappreciated right now?
Multi‑agent setups are often misapplied; supervised orchestrators with sub‑agents can work, but peer‑to‑peer chatter is hard to control in production. Coding agents are still underused outside tech hubs and will unlock big gains.
Chasing tools is overrated; what’s underrated is deep problem understanding, strong product taste, and thoughtful design.
Paint the next year or two—what’s coming by the end of twenty twenty six?
Proactive background agents will plug into the real places work happens and nudge you with completed steps or timely prompts, like waking up to reviewed PRs or triaged tickets.
Richer multimodal experiences will get us closer to human conversational richness and unlock value in messy inputs like handwritten forms and ugly PDFs.
For builders who want to level up, what single skill should they invest in?
Cultivate taste, judgment, and ownership—shipping is cheap, insight is scarce—and act with agency, like a teammate who replaced a pricey tool with a simple internal app to rethink our process.
Add persistence to that; your willingness to endure the messy iterations becomes a moat, because the hard‑won knowledge compounds.
Final reminder: obsess over customers and data flows. Most of the real work is mapping workflows and edge cases, not chasing model novelties.
Lightning round time—quick picks and favorites.
Book: When Breath Becomes Air for its lived wisdom; show: rewatch Silicon Valley, it’s timeless; product: WhisperFlow’s concept‑aware dictation is surprisingly useful; motto: be bold enough to try even when the odds look bad.
Book: the Three‑Body trilogy for its sweeping perspective; game: Expedition 33; tools: Recast for speed and caffeinate to keep long runs alive; motto: you connect the dots only in hindsight.
What I admire about Kiriti is his calm, grounded judgment, the way he anticipates pitfalls, and his kindness; he lets the work speak.
What I admire about Ash is her gift for teaching—she makes complex ideas feel simple and empowering, and she pairs that with relentless follow‑through.
Find my writing on LinkedIn, grab our free GitHub resource list, and check out our Maven course on building enterprise AI products; we also share lots of free sessions.
I’m also on LinkedIn, and my DMs are open if you want to talk about coding agents or tricky product use cases.
Thanks for listening; subscribe on your favorite app and consider leaving a rating so others can discover the show.