ThousandEyes
The Internet Report
This is The Internet Report, a podcast uncovering what’s working and what’s breaking on the Internet—and why. Tune in to hear ThousandEyes’ Internet experts dig into some of the most interesting outage events from the past couple weeks, discussing what went awry—was it the Internet, or an application issue? Plus, learn about the latest trends in ISP outages, cloud network outages, collaboration network outages, and more.
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Meta, LinkedIn, and Comcast Outages, Explained | Pulse Update 16.03.2024 17:49
Over a two-day period this past week, major social media platforms—Meta’s Facebook and Instagram, LinkedIn, and Discord—all experienced disruptions. In the same timeframe, Comcast was also impacted by an outage that affected access to specific services and applications. Meta experienced issues with its log-in process, Discord navigated unexpectedly high load volumes, Comcast dealt with 100% packet...
AT&T Outage and Disruptions at Google Cloud, Front, and More | Pulse Update 04.03.2024 16:26
Load is a fundamental but, at times, challenging variable for networks and operations teams to handle. In the past few weeks, ThousandEyes saw various load-related problems impact organizations including Google Cloud, Front, several Australian banks, and Minnesota State University Moorhead. Tune in to learn more about what happened during these incidents, as well as hear our commentary on the rece...
Square Outage, Data Center Issues & Planning for Resiliency | Pulse Update 17.02.2024 17:10
When outages happen, it’s what you do next that matters. It’s important to have a backup plan in place that you can quickly activate to minimize the impact of an incident. Over the past two weeks, companies initiated a range of resiliency actions, including asking customers to use alternate authentication methods (or to avoid logging out of a service), setting up a new contact center to re-establi...
Security, Great Digital Experiences & Why Visibility Matters 10.02.2024 16:53
The ThousandEyes Internet Intelligence team joins us from Cisco Live in Amsterdam, talking about a major theme from the event—security. Tune in to hear their thoughts on how visibility can help companies in their security efforts, the sovereignty of data in flight, and why you don’t have to choose between security and performance. ——— CHAPTERS 00:00 Intro 01:09 Evolving Security Landscape 04:53 Se...
Understanding the Microsoft Teams & Azure Disruptions | Pulse Update 03.02.2024 16:48
What happened during the recent Microsoft Teams and Azure disruptions? Go under the hood of these incidents and also explore other recent disruptions in this week’s Pulse Update. CHAPTERS - 01:03 Network issue leads to Microsoft Teams service disruption - 04:09 Azure Resource Manager exhausts capacity, causing service issues - 06:20 Oracle Cloud experiences network outage - 09:56 Jira users encoun...
Unpacking Recent ChatGPT Issues & Other Outage News | Pulse Update 20.01.2024 24:26
What caused recent dips in performance for OpenAI’s ChatGPT? Tune in to hear The Internet Report team unpack this and other recent disruptions, including a hack that led to an outage at the Spanish branch of the Orange mobile network, and a blip for customers of the cloud services provider DigitalOcean. They’ll also cover the outage trends they’re seeing in 2024 so far and how extreme cold weather...
2023 Internet Outage Trends & the New Outage Landscape | Pulse Update 13.01.2024 11:06
As they launch into 2024, organizations are facing a different outage landscape than they had at the start of 2023. The past year saw increases in cloud service provider (CSP) outages, application outages, and the percentage of U.S.-centric outages—all of which point to an evolution in the way outages happen and the need for different strategies to minimize the impact of disruptions. In this episo...
Insights From the Ghosts of NetOps Past, Present, and Future 21.12.2023 21:32
As 2023 comes to a close, in the spirit of Dickens’ holiday classic “A Christmas Carol,” let’s reflect on the valuable insights left by the ghosts of network operations teams past, present, and yet to come. Tune in to hear host Mike Hicks (Principal Solutions Analyst at ThousandEyes) discuss lessons from the NetOps teams of the past, the current state of NetOps, and what the future might hol...
Peering Issues, Internet Resilience, and Cloud Outage News | Pulse Update 12.12.2023 13:52
Recent changes appeared to trigger a series of events for two peering points internationally—with very different impacts. Tune in to learn more about these incidents, why they differed, and the lessons they leave. Mike Hicks, Principal Solutions Analyst at ThousandEyes, will also cover the latest outage numbers and explore other recent incidents, including an Oracle Cloud outage and a duo of disru...
Scaling To Meet the Black Friday Demand: Tips for IT Teams 22.11.2023 15:02
As companies gear up for Black Friday, The Internet Report team shares some best practices for delivering great customer experiences and minimizing downtime during one of the retail industry’s biggest days of the year. Mike Hicks, Principal Solutions Analyst at ThousandEyes, will cover some helpful case studies of Black Fridays that experienced some hiccups and what you can do to guard again...
Understanding the Recent Workday and Cloudflare Outages | Pulse Update 14.11.2023 32:11
Backend-related incidents have been a recurring theme in outages across 2023, caused by everything from data center issues and hardware mishaps to failures at common (shared) services. Recently, we saw two examples of these backend issues when data center power problems led to outages at both Cloudflare and Workday. Tune in to hear more about what happened at Cloudflare and Workday, as well as our...
Halloween Special: Ghosts of Outages Past 01.11.2023 43:36
This Halloween, The Internet Report team is sharing some of their most thrilling (and chilling) networking tales. Pull up a chair (and a big bowl of your favorite Halloween candy) to hear what happened—and important lessons learned. ——— CHAPTERS 00:00 Intro 01:40 Haunting obstacles with a dynamic routing protocol that thwarted crew changes on an oil platform 10:00 A spooky code base rollout that u...
Insights From Outages at Citibank, DBS, and Other News | Pulse Update 30.10.2023 24:18
In recent weeks, back-end infrastructure work and other backend-related issues impacted various online and consumer banking services, including DBS and Citibank in Singapore. Simple front-facing customer experiences that we’ve become accustomed to today can often mask considerable complexity on the backend. The service delivery chain of technologies powering the front end often comprises a mix of...
Talking Data Freshness + Slack, Cloudflare, and Google Outages | Pulse Update 17.10.2023 30:47
Outages and degradations can happen when underlying data isn’t fresh enough. In recent weeks, stale data may have contributed to incidents at both Slack and Cloudflare. Slack began experiencing issues when, by our best guess, its app stopped trusting the freshness of the data in the cache; and, separately, Cloudflare’s 1.1.1.1 DNS resolver ran into some issues related to stale root zone data...
Internet Outages: Why One Small Link Can Break the Whole Chain | Pulse Update 02.10.2023 22:25
Providing great digital experiences relies on a complex service delivery chain. The past few weeks brought multiple reminders that the root cause of cloud and app disruptions often comes down to one single link in this chain. While the component at issue may appear small, if it’s not functioning normally, the consequences can be significant. Additionally, the impact of a malfunctioning “link...
Data Center Disruptions, Square Down, and More News | Pulse Update 15.09.2023 33:13
In a world that operates at “hyperscale,” the potential for hyperscale-sized problems is also very real. The measure of a good provider—and a well-engineered system—is how well they handle these anomalous conditions and minimize disruption. During recent weeks, some of these hyperscale-sized outages hit, including data center-focused disruptions that impacted companies like Square, Oracle OCI, Net...
Disruptions at Slack and X + Thoughts on “Take Twos” | Pulse Update 02.09.2023 21:15
An outage occurs, a change is rolled back, and everything stabilizes. But what happens when the change is attempted a second time? These second tries often go much more smoothly. While another outage might still occur during this “take two,” the impact is usually far less severe. The engineering team has learned from what went wrong the first time and is ready to stop at the first hint of trouble....
An August Slack Outage and Why Context Matters | Pulse Update 21.08.2023 34:23
Context matters when working on a distributed web-based application or service where everything is linked and dependent on each part functioning correctly. It’s all too easy for one team to make a change that unexpectedly affects something another team is working on. Or the combined impact of both changes may also accidentally break something. To avoid such mishaps, teams should cut back on...
SharePoint Outage and Security Certificate Considerations | Pulse Update 05.08.2023 26:30
In an end-to-end service delivery chain, isolated changes can have broad consequences. This played out recently when an erroneous SSL certificate change at Microsoft appeared to cause a SharePoint Online and OneDrive for Business outage. While this incident definitely underscores the importance of valid security certificates, it’s also a reminder of what can happen when even one component in an en...
Azure Disruption, Meta App Issues, and Navigating Edge Cases | Pulse Update 21.07.2023 19:15
Let’s face it. Not every contingency can be planned for. Sometimes an outlier scenario pops up and causes an unexpected outage or disruption. Over the past few weeks, multiple companies appeared to be impacted by such edge cases: Azure; GitLab; and Meta’s WhatsApp, Facebook, Instagram, and Threads—its newest addition. Tune into the latest Pulse Update episode to learn more about what happened duri...
A Front Door, But No House: Explaining Application Outages | Pulse Update 10.07.2023 17:32
The application opens, but users encounter errors when they try to do anything—what gives? It’s the curious case of the disappearing backend. Discover why application issues often show up like this, with the service reachable but unresponsive beyond rendering a basic landing page, and sometimes an accompanying error message. In this episode, hosts Mike Hicks and Brian Tobia discuss this comm...
Application Outages Up in 2023—What to Know | Pulse Update 28.06.2023 21:53
Though network outages are still far more common, application outages seem to be increasing in 2023—and having bigger impacts. Tune in to learn more about this trend and dive into incidents at Okta and Instagram. Host Mike Hicks will also explore other outage trends from the first half of the year in this special episode reflecting on the state of the Internet in 2023 thus far. To learn more...
Is Spring Cleaning Causing an Outage Spike? | Pulse Update 10.06.2023 27:21
For three consecutive years, there appears to have been a spike in outages and degradations in May. A potential “spring cleaning effect” may explain why. Tune in to learn more about this possible trend and explore what happened during recent incidents at Twitter; Microsoft 365; Slack; Instagram; Apple’s iMessage; and subscription-based streaming service, Max (formerly known as HBO Max). Afte...
How Outages Can Impact Distributed Dev Teams | Pulse Update 26.05.2023 26:56
Tune in to explore ways that outages can impact distributed software development teams and what companies can learn from recent incidents at GitHub, Google Cloud, and Apple. To learn more, check out these links: Internet Report: Pulse Update Blog: https://www.thousandeyes.com/blog/internet-report-pulse-update-outages-and-distributed-dev-teams?utm_source=transistor&utm_medium=referral&...
Redundancy in the Cloud Era: Two Case Studies | Pulse Update 15.05.2023 26:42
When it comes to your technology strategy, it's a good idea to have more than one way to access every resource—just in case. As IT environments have changed, so has the thinking around the right approaches to achieve this desired redundancy. Two recent incidents at Google Cloud and Microsoft 365 reinforce the importance of redundancy—and the need for evolving strategies to meet this goal. To learn...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.