Fexingo

The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering

Business EN ↓ 104 episodes

Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and orga...

Author

Fexingo

Category

Business

Podcast website

www.fexingo.com

Latest episode

Jul 11, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

How Slack Cut Alert Noise by 90 Percent 22.05.2026

SRE teams drown in alerts—most of them false positives. Lucas and Luna break down how Slack’s reliability team redesigned their alerting pipeline to cut noise by 90 percent without missing critical incidents. They walk through the specific techniques: tiered severity classification, suppression rules based on dependency graphs, and a weekly 'alert review' meeting that treats every notification as...

Why Error Budgets Changed How SRE Teams Sleep at Night 21.05.2026

Episode 3 of The Site Reliability Podcast with Fexingo dives into the concept of error budgets — the SRE mechanism that turns uptime targets into engineering velocity. Lucas and Luna unpack how Google introduced error budgets to solve the tension between reliability and feature releases, using real examples from a major streaming service and a payments API. Lucas explains the math behind 99.9% vs...

Incident Response Playbooks That Actually Work 21.05.2026

Lucas and Luna dive into what makes an incident response playbook effective versus one that just sits in a wiki. Lucas walks through a real example from Stripe's 2023 outage where the playbook itself became the bottleneck. They discuss the anatomy of a good playbook, why most fail, and how to test them without waiting for a real incident. Luna pushes back on whether playbooks create a checklist me...

The Cost of a Second Downtime at a Major Bank 19.05.2026

On the premiere of The Site Reliability Podcast, Lucas and Luna examine a specific incident: a 47-second outage at a major US bank in early 2026 that halted digital payments for four hours. They break down what actually happened inside the SRE war room, how the bank's incident response failed at the human level, and why that half-minute of downtime cost an estimated $3.2 million in direct losses....

Listen to the The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.