Rootly

Humans of Reliability

Behind every reliable software system, there are people working hard to keep it online.  Humans of Reliability is a series that spotlights the engineers, leaders, and innovators at the heart of incident management and system reliability. Through candid conversations, we explore the challenges, lessons, and personal journeys of those navigating complex technical landscapes to ensure the systems we rely on run smoothly.  From unforgettable incident stories to favorite tools, workflows, and hobbies, Humans of Reliability uncovers the human side of technology—offering insights and inspiration for...

Author

Rootly

Category

Technology

Podcast website

rootly.com

Latest episode

Jul 6, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Stop shipping features for a quarter, it might save you w/ Monday.com co-founder 06.07.2026

When incidents pile up fast enough, every part of the company bleeds: support is fielding angry customers, AEs are on apology calls, and engineering is burning cycles on retrospectives instead of shipping. For Eran Kampf (VP of Engineering at Twingate ) where the product is the network, that was the moment he made a call most engineering leaders won't: stop all feature work for a quarter and...

"Hey Claude, where's my database?": How an AI agent nuked production w/ Alexey Grigorev 16.06.2026

"Hey Claude, where's my database?" The answer came back: "Oh, sorry, your database is gone." Alexey Grigorev , founder of DataTalks. Club and creator of the Zoomcamp courses that have taught data and ML engineering to over 100,000 people, knows Terraform well enough to cover it in his curriculum. That didn't stop a chain of small, reasonable-sounding decisions from en...

Every pilot is ready for engine failure: are your engineers? w/ Hamed Silatani (Uptime Labs) 28.05.2026

Every pilot who's never had an engine failure is still ready for one. The same can't be said for most software engineers facing their first major incident.  Hamed Silatani , co-founder and CEO of Uptime Labs , and former Head of Reliability Engineering at IG Group, has spent two decades watching engineers learn incident response the hard way: alone, under pressure, with no training.  A m...

LLM Observability: Lessons From MLOps w/ Maria Vechtomova (Cauchy) 14.05.2026

For nine years, Maria Vechtomova was shouting about monitoring. Nobody cared, until LLMs arrived.  As co-founder of Cauchy, Databricks MVP, and one of the most followed voices in MLOps, Maria has watched the field evolve from hand-built experiment trackers to today's flood of observability tools, and her central claim might surprise you: globally, nothing has changed.  The fundamentals are th...

The Golden Hour: Why the First 15 Minutes of an Incident Decide Everything w/ Gandhi M. N. Kumar (Twillio) 28.04.2026

Most incident response advice focuses on tools, alerts, and post-mortems. Gandhi Mathi Nathan Kumar , Principal Incident Commander at Twilio , with 14 years running calls that have pulled in up to 100 responders, argues the work that actually matters happens in the first 15 minutes.  In this episode, Gandhi walks through what he calls the golden hour: the window where you decide whether you know w...

From 600 to 6,000: Federating Incident Response w/ Cliff Snyder (ex-LinkedIn SRE) 22.04.2026

A centralized SRE team of 600 engineers as the first line of defense for every incident works - until the business asks you to spread that responsibility across 6,000.  Cliff Snyder, senior SRE at Multimedia and a decade-long veteran of LinkedIn's SRE org, walks through the 18-month project to democratize incident response: replacing a patchwork of Jira, Google Docs, and in-house tooling with...

AI Didn't Change the Game, It Just Exposed Your Bottlenecks w/ Ganesh Datta (CTO, Cortex) 09.04.2026

Every engineering org says they want to improve reliability — but most can't even agree on what "good" looks like. Ganesh Datta, Co-Founder and CTO of Cortex, has spent the better part of a decade helping companies confront that gap.  In this conversation, Ganesh makes the case that platform engineering and SRE are solving the same human problem — earning adoption through influence,...

Fear, Identity & Flaky Tests: AI in Reliability w/ Dana Lawson (CTO, Netlify) 31.03.2026

The self-healing systems that SREs have dreamed about for a decade aren't a distant promise anymore — they're already being built, and the biggest barrier left is cultural.  Dana Lawson, CTO at Netlify, has spent over 25 years in the trenches of developer infrastructure, from sysadmin roots to running the platform that powers 5% of the internet.  In this episode, Dana makes the case that...

The Incident You Never Had: Deterministic Simulations w/ Will Wilson (Antithesis CEO) 17.03.2026

Most reliability engineering happens after something breaks. Will Wilson thinks that's the wrong place to be. As co-founder and CEO of Antithesis, the autonomous testing platform that just raised $105M in a Series A led by Jane Street, Will has spent years building the infrastructure to catch failure modes before they ever reach production. His starting point is uncomfortable: the testing pra...

Burnout Doesn't Ask Permission: Recognizing, Recovering, and Rebuilding w/ Stephen Townsend 04.03.2026

Burnout doesn't announce itself. For Stephen Townsend, SRE team lead and host of the Slight Reliability podcast, it crept in over months of mounting pressure on a massive transformation program, and announced itself overnight with an inability to sleep. In this episode, Stephen shares his personal burnout story with rare honesty: the physical symptoms he dismissed, the org structure that left...

Code Is Cheap, Reliability Isn’t: Owning Production in the AI era w/ Swizec Teller 16.02.2026

Code has never been easier to write. With AI copilots and agentic coding tools, spinning up features feels almost effortless. But production systems don’t run on vibes, they run on reliability. In this episode of Humans of Reliability , Swizec Teller , author of the bestselling Scaling Fast , argues that writing code was never the hard part. The hard part is making it run. We talk about the “Feynm...

Democratizing Reliability: Empowering Non-Devs with Dileshni Jayasinghe (commonsku) 14.01.2026

Many companies don’t invest in incident management until something goes wrong. commonsku took a different path. In this episode of Humans of Reliability, Sylvain sits down with Dileshni Jayasingha, VP of Technology at commonsku, to talk about what it really takes to introduce incident management in a mature, profitable SaaS that had never formalized it. From rolling out observability and incident...

99%+ Accuracy on a Moving Target: Model Deprecation and Reliability with Tomás Hernando Koffman (Not Diamond) 22.12.2025

Shipping systems powered by LLMs would be hard enough if the models stayed the same. But in reality, they don’t. Models get updated and deprecated at a pace traditional software wouldn’t. All while teams are still expected to hit reliability targets that look a lot like traditional SLAs. In this episode, Tomás Hernando Koffman, Co-founder of Not Diamond, breaks down what it really takes to reach 9...

The Reality of GenAI in Production with Eduardo Ordax (AWS) 12.12.2025

GenAI demos are easy. Production is where everything breaks. In this episode, Eduardo Ordax , Principal GTM GenAI at AWS, breaks down what actually stops companies from shipping reliable AI systems, and why the real blockers have little to do with technology. We dig into the culture problems that slow down adoption, the endless cycle of experimentation, and why non-deterministic LLMs don’t fit nea...

It’s Never Different This Time: LLM Reliability Without the Hype with Julien Simon 19.11.2025

In this episode, Julien Simon, longtime voice in the open-source ML world, reminds us that even in the era of GenAI, reliability fundamentals haven’t changed. Julien breaks down why calling “the same model” from different providers can produce wildly different results, how deployment choices introduce hidden variability, and why reliability teams need to think of LLM systems as distributed systems...

You Can’t Fix What You Don’t Measure: Observability in the Age of AI with Conor Bronsdon 05.11.2025

Only 50% of companies monitor their ML systems . Building observability for AI is not simple: it goes beyond 200 OK pings. In this episode, Sylvain Kalache sits down with Conor Brondsdon ( Galileo ) to unpack why observability, monitoring, and human feedback are the missing links to make large language model (LLM) reliable in production. Conor dives into the shift from traditional test-driven deve...

The End of “Good Code”? AI, Throughput, and Reliability with CircleCI CTO Rob Zuber 10.09.2025

Is “good code” still the right measure of engineering success in an AI-driven world? In this episode of Humans of Reliability , Rob Zuber , CircleCI CTO, joins Sylvain to explore how coding assistants are reshaping developer workflows and changing what teams value. Rob shares what he’s seeing across CircleCI’s customer base: a clear boost in throughput, new bottlenecks shifting from code creation...

Frontline Reliability: Protecting User Journeys with SLOs with Shery Brauner (Razor, ex-Zalando) 20.08.2025

What does it really take to move from firefighting incidents to building reliability at scale? In this episode of Humans of Reliability , Shery Brauner (Razor, ex-Zalando) shares her unique journey from frontend and backend engineering to leading site reliability practices. She explains why protecting the user journey is the key to effective incident management, how SLOs cut through noisy alerts,...

Balancing Reliability at the Crypto-Finance Frontier with Brian Shaw (Uphold) 03.07.2025

Sylvain Kalache sits down with Brian Shaw, Senior Engineering Leader at Uphold , to explore the reliability challenges that arise when operating at the intersection of traditional finance and crypto markets. Brian shares how unexpected market events can create massive traffic spikes, how their platform architecture and Kubernetes setup help them stay resilient, and why Uphold's transparency a...

Command Under Pressure: David Owczarek on Incident Leadership and Human-Centered Reliability 17.06.2025

Incident response is as much about people as it is about systems. In this episode, David Owczarek, a veteran engineer leader and seasoned incident commander, joins Silvan Kalache to unpack the human dynamics behind effective reliability leadership. Drawing on experiences across startups and global enterprises, David shares what really matters when everything breaks, including: – How incident respo...

AI at the Frontlines of Healthcare Reliability with Ryan Lockard (CVS Health) 30.05.2025

AI is transforming reliability work—from reactive firefighting to proactive engineering. In this episode, Ryan Lockard, VP of Platform Engineering and AI Enablement at CVS Health, joins Sylvain Kalache to break down how AI is showing up on the frontlines of healthcare infrastructure and operations. From LLM copilots to cultural shifts in ownership, Ryan walks us through: How AI tools help troubles...

Trust Is the Product: Building Reliable Billing in the AI Era with Cosmo Wolfe (Metronome) 26.05.2025

In this episode, we sit down with Cosmo Wolfe , Head of Technology at Metronome , to unpack how reliability, trust, and architecture intersect in one of the most critical and overlooked parts of the AI product stack: billing. As AI workloads introduce unpredictable usage patterns and nontraditional pricing models—from token-based to outcome-based—companies are navigating a new frontier of customer...

The Golden Path to Nowhere: When Platforms Undermine Reliability with Chase Roberts (Northflank) 14.05.2025

Internal platforms promise speed, consistency, and scale — but what happens when they become a distraction? In this episode, Chase Roberts, COO at Northflank , joins Sylvain Kalache to examine the quiet ways platforms erode developer experience when not planned carefully.  From abandoned golden paths to shadow deployments and brittle YAML pipelines, Chase walks us through:  Why early PaaS got deve...

AI can boost developer productivity, if used right, with Justin Reock, Deputy CTO at DX 30.04.2025

In this episode of Humans of Reliability , we sit down with Justin Reock , Deputy CTO at DX , to unpack the real impact of generative AI on developer productivity. Drawing from early data in DX’s GenAI Impact Report, he explains why time savings alone don’t tell the full story and why the real value might lie in shifting cognitive load toward meaningful work.  We also explore how traditional produ...

Why Reliability in the AI Era Starts with the Network with Marino Wijay 17.04.2025

In this episode, we explore how networking has shaped reliability as we know it. Marino Wijay cloud networking expert and Staff Solutions Architect at Kong shares how his journey began not as an SRE, but with cables, routers, and switches. Marino explains the evolution of the fabric holding systems together through virtualization, and how software-defined networking, which is now a key element to...

Listen to the Humans of Reliability podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.