The Pulse: Firebase’s global outage & poor response
The Pulse: Firebase’s global outage & poor responseThe Firebase iOS SDK crashed after a backend change, crashing all apps which used Firebase analytics for 2-6 hours. Google didn’t update the status page, but did offer a postmortem 4 days later.
Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from last week’s issue of The Pulse. Full subscribers received the article below seven days ago. If you’ve been forwarded this email, you can subscribe here. Firebase – built by Google – has had a nasty outage this week with shockingly poor incident management at odds with how Google itself usually deals with high-severity incidents. The outage started on Tuesday (29 Sep) at 5:41pm (PDT), when iOS apps using the Firebase SDK started to crash upon first opening; every iOS app that uses the Firebase SDK with analytics enabled was affected in this way. Developers of affected apps opened a GitHub ticket, in the absence of much else to do. On the ticket, the message “it’s crashing for me too!” was oft-repeated. Devs reporting their apps crashing. Source: GitHub 6:51pm (PDT): acknowledgement. An hour and ten minutes after the crashes started, an engineer on the Firebase team acknowledged that they were aware of the outage. Just over an hour into the incident, the Firebase team became aware of the outage. Source: GitHub It’s unclear if the Firebase team was alerted via this ticket with 100+ comments by devs, or if Google’s own monitoring tool showed the issue. I asked Google/Firebase two days ago and haven’t had a response. Not having anything better to do than wait for Google to resolve the issue, the memes began: Memes while waiting More memes Others attempted to help the Firebase team by pinpointing the potential issue. Indeed, before a Google engineer acknowledged the incident, an external developer found the root cause at 6:37pm PDT; it was a zero-length entry that was crashing the SDK: Given the flags are shipped by the backend, the offending change was a backend one, and the easiest resolution would be to roll it back, which the community practically begged Google to do: Frustrating: Understanding the problem and how to solve it, but nothing to do but post. Source: GitHub Here’s a neat summary of the incident from another dev: Summarizing the incident better than any Google dev ever did. Source: GitHub 7:24pm (PDT): rollback starting. An hour-and-a-half into the incident, the Firebase team started rolling back the offending backend change: Finally – the rollback started! Source: GitHub 8:16pm (PDT): rollback complete. And the rollback completed ~50 minutes later: Rollback complete, minus the caching problem. Source: GitHub Software engineer, Nick Cooke, on the Firebase team posted a summary with more accurate timestamps: Source: GitHub What we can deduce from this:
Incident management basicsThe Firebase team itself closed the outage with a short report effectively saying that there had been an outage, but they’d resolved it now, so thanks for your patience and have a nice day. This handling of a high-impact incident is absolutely not typical of Google, the company that coined the term ‘Site Reliability Engineer’ and wrote the SRE book. For one, Firebase never bothered updating its status page. Oddly enough, the official Firebase status page showed all systems green – despite the acknowledgement of the outage. Indeed, during it and afterward, they didn’t update the status page to indicate the lengthy outage: A global outage was never recorded on the status page. Source: Firebase But status pages exist for good reasons, including:
It’s worth asking: if an outage that takes down most (or all?) iOS apps using Firebase doesn’t warrant an update to the status page, then what does! Google published a postmortem four days later, answering questions on how the outage happened. On Friday, 2 October, Google published a postmortem on the Firebase blog. It was a configuration change that crashed so many iOS apps. From the postmortem:
In the postmortem, Google noted that engineers were alerted to the outage through both GitHub reports coming from external developers, as well as their internal monitoring. It took another hour to pinpoint the cause being a legacy configuration flag cleanup. Firebase says they have no way to update their status page for client-side outages. In the postmortem, Google explained that there is no place to indicate client-side outages on their dashboard (emphasis mine):
It’s good to see Google not dropping the ball fully, and recognizing that both their dashboards and their incident management process need improvement. It’s fair to ask though: why did only iOS crash, and not Android? Firebase’s Android SDK seems to be hardened more than iOS, as the feature flag removal did not crash Android devices. Especially that now, with AI, it’s easier than ever to compare iOS and Android implementations to ensure they are identical – and it’s what Shopify has been doing during their native rewrite – could it have been a missed opportunity for Google to audit the differences between the iOS and Android SDKs? To me, not having an action item here feels like a missed opportunity. Still, this is a good reminder to anyone and everyone shipping iOS and Android apps: aim to harden them, and when possible, run tests with malformed payloads, then fix crashes those payloads cause. Déjà vu: the 2020 Facebook SDK crashThe last time there was a similar crash was in 2020, with Facebook. That May, apps such as Spotify, TikTok, Pinterest, and others also started to suddenly crash due to the Facebook SDK crashing all apps using it. Back then too, devs followed along on a GitHub ticket and they also found that bug: a value that should have been a dictionary but was a boolean: What caused the 2020 Facebook crash. Source: GitHub Then as now, there was banter by devs being made to wait for a fix: One of the memes from the 2020 crash. Source: GitHub And requests to not move fast and break things any more: A plea for prioritizing reliability in the future. Source: GitHub Making light of the situation: Apps that did not initialize the SDK unconditionally upon startup should not have crashed – but most did Source: GitHub And also anticipating the resolution: Some more memes on the GitHub issue In the end, Facebook reverted the backend change, but shared even less than the bare minimum details from Google this time. This is all we know about that 2020 outage that was arguably more wide-ranging than the Firebase one: All that Facebook shared about their global outage I wonder if some people think that public-facing incident management is no longer important or valuable, even for developer-facing products. I’m not shocked that Facebook/Meta never bothered to communicate much about their outage because dev tools are not part of the DNA there. But with Firebase, I am surprised that more than a week later, the postmortem is still not visible on the Firebase status page. And maybe this is Google “shipping their org chart” playing out, live. The outage technically was caused by Google Analytics (who made the feature flag change), but is the responsibility of the Firebase SDK (whose iOS SDK was not hardened enough to deal with this new payload). The outage itself was buried inside a Google Ads dashboard (!!) which suggests that whatever team is seen responsible for the outage is inside the Google Ads organization. In the end, despite the Firebase team committing to “improving status dashboard latency and coverage,” last week, those teams are in no hurry to carry out this work. AI agents might be making lots of work more efficient, but following up on action items seems to move at the same snail pace at Google, as it did pre-AI! Read the full issue of The Pulse this is from, or check out this week’s The Pulse. This week’s issue covers:
You’re on the free list for The Pragmatic Engineer. For the full experience, become a paying subscriber. Many readers expense this newsletter within their company’s training/learning/development budget. If you have such a budget, here’s an email you could send to your manager. This post is public, so feel free to share and forward it. If you enjoyed this post, you might enjoy my book, The Software Engineer's Guidebook: navigating senior, tech lead, staff and principal positions at tech companies and startups.
|


















Comments
Post a Comment