← BibliothèqueToutes les vidéos
Tech.Rocks Summit 2023
Managing a massive scale incident
- Alexis Lê-Quôc (CTO et co-fondateur, Datadog)
Tech.Rocks Summit 2023 · 7 décembre 2023 · 34 min · en anglais
Résumé
Le 8 mars 2023, Datadog a subi une panne mondiale de très grande ampleur. Alexis Lê-Quôc explique ce qui a déclenché l'incident et pourquoi le rétablissement a demandé un effort aussi important. Il partage les enseignements tirés, la conduite de la réponse à incident, qui a coordonné plus de 500 ingénieurs pendant plus de deux jours, et la manière dont Datadog a bâti une organisation d'ingénierie capable d'y parvenir avec un minimum d'héroïsme.
L’essentiel
Le retour sur un incident majeur chez Datadog, de sa mécanique technique à l’organisation de la réponse.
Pour préparer une équipe à gérer un incident et améliorer la coopération pendant la résolution.
Les idées clés
- Désigner un responsable d’incident, puis répartir les rôles : communication, chantiers de résolution, direction et relation client. Alexis Lê-Quôc précise que le responsable n’est pas nécessairement la personne la plus senior dans la pièce. à 21:31
- Garder une démarche sans recherche de coupable, quelle que soit la gravité. Le point d’attention devient ce qui s’est passé, comment le réparer et ce qu’on en apprend. à 19:20
- Enfin, s’entraîner : Alexis présente la pratique répétée comme le moyen trouvé pour améliorer une réponse qui reste imparfaite. à 27:54
Questions pour votre équipe
- Qui coordonnera l’incident, la communication et les chantiers techniques ?
- Comment préserver la concentration des personnes qui interviennent ?
- Quel exercice pouvons-nous organiser avant le prochain incident ?
Il s’agit d’un incident dans une infrastructure particulière. Les causes techniques et la réponse décrites ne forment pas une procédure universelle.
Chapitres
Summary
On 8 March 2023, Datadog experienced a massive global outage. Alexis Lê-Quôc explains what triggered the incident and why recovering from it took such an effort. He shares the lessons learned, how the incident response was run, coordinating more than 500 engineers over more than two days, and how Datadog built an engineering organisation capable of this with minimal heroism.
Thèmes : Cloud, infra & ops
Transcript complet
Transcription automatique, à relire : les noms propres peuvent être mal orthographiés.
Is everyone doing well? Has everyone found their neighbour again? Yes? Shall we start? This will be the last part of our first day. There is a lovely surprise at the end. So stay with us, don't leave. Make yourselves comfortable: it's going to be fantastic. But first, what better way to embody the no-bullshit spirit than by sharing a major failure that hit Datadog on March 8? And the way they recovered deserves our applause, as you will see. Please welcome Alexis Lê-Quôc, CTO of Datadog. Dear Alexis, [unclear words]. And let me point out that this talk will be in English. We need to have the micro open, please.
Ah, there's Vanessa from the Tech.Rocks team, at last. Thank you, Vanessa. Safe. OK, cool. OK, so I'm going to talk to you about what happened to us. I think the title was The World Blew Up. We're all OK. So we had a... We had an issue as, you know, in last March, March 8th. But before I get into that, and I want to talk to you a little bit about Datadog. So for those who don't know Datadog, we're a SaaS provider of observability and cloud security. We're based in New York City, primarily with offices in Paris, fairly sizable. We started about 13 years ago in 2010 in New York, two people, and now we're over 5,000 employees, lots of customers. Now, I want to situate a little bit the context for what happened in last March. So in terms of just to size a little bit, we get just a lot of data from customers, tens of trillions of points every single day, millions of hosts reporting every day.
10 seconds or so, lots of containers reporting. This is a massive data pipeline, real-time data pipeline. And so to run that, we need about several hundreds of thousands of machines. We run entirely in the cloud. We've chosen Kubernetes to power all this. It's since about 2018. And we're truly multi-cloud. We run the exact same products and code in the three main cloud providers, being Azure, AWS, and GCP. So that's the data log in a slide. We grow rapidly. But March 8th, as I was mentioning, we had to publish this on our status page. What happened really is we were down for about 14 to 16 hours. It took us over 48 hours to restore all service, all data, everything.
So after 14 hours, service restarted, but we had to go and backfill, and that took about 48 hours. And what I want to talk to you about is what happened, why it happened, and how we responded to it. And back then, we were about 5,000 people. So you imagine massive site, everything goes down, and you have 5,000 people who have to respond and bring everything back up. Last thing I would say is if you're a customer and you're in I'm sorry for what happened that day. I hope that I can share some of the lessons learned. Okay, so this is a very quick, you don't really need to understand the details, but one of the main pain points with Kubernetes is networking. And so because you have a host, in this case we call that a node, so a node, and then you have pods that run workloads inside it, and then obviously these pods are networked with each other sometimes and with the rest of the world. And there's some fairly intricate machinery at play here, just to make sure that pods can talk to one another and to the rest of the world.
So the way we've done that, because we have to run multi-cloud, because we have to run... Clouds have differences, so we run our own Kubernetes, if you will. We chose to basically have a fairly delicate interplay between node networking and pod networking. So we used something called Cilium for that. And the way it works is fundamentally there's a set of IP roles that determine which packets, when they come, where they go, when they leave the pod, where they go, and so on and so forth. And that's all interconnected. And that usually works really well. And that scales to really hundreds of thousands of nodes and millions and millions of containers. And on March 8th, what happened is part of the picture just wet up in flame. So specifically, and I'll come back to it, there's this component of pretty much all modern Linux machines called SystemD that just restarted.
And what it did, it basically nuked the network of a machine. So in this case, you could see the little flames. And the impact was actually fairly quickly felt. So this is the network traffic of one of our data centers. It happens to be a European one. And you can see before that little square, you know, sort of normal, everything's, there's no unit, but, you know, it's a lot of data. And then fairly quickly drops to almost zero. And the fairly quickly is March 8th, between 6 and 7 a.m., actually around 7 a.m. UTC. And so everything went silent, almost silent, for some time until we were able to figure out what happened, recover, and then scale everything back up and then sell. So that's what happened. I want to differentiate a little bit just in terms of technical details. In GCP and Azure, so in case you're interested and you run multi-cloud, what happened is the pods lose connectivity, so the nodes lose connectivity, so just went silent, as I just showed you.
But the nodes actually, the instances, the VMs, if you will, they actually stick around. about two hours to realize that at the onset of an incident. They stick around and what we realized, and somebody tried, and I can't, you know, I don't remember exactly, but somebody must have tried, okay, well, it's still here, let's reboot it. So we rebooted one, and it came back normally. And so what then the response on GCP and Azure was to basically reboot tens of thousands or hundreds of thousands of nodes. And so, you know, not a small feat by itself, but that fixed the thing. In AWS, it interestingly was a different response from the cloud provider. There is no good one between GCP, Azure, or AWS. The response that the cloud provider provided to a network failure, both of them are valid. Niblis is a little bit different in the sense that the pods disappear, they lose connectivity, then the nodes at the same time disappear.
And because we run on Niblis everything in autoscaling groups, the logic for autoscaling groups is to try to effectively ping the host, the node, just to make sure is it alive or not. And if it's not alive, if somehow the node is stuck, then the autoscaling group zaps the node and restarts a new one automatically. So that's great because that's usually what you want. Now the problem is, there's a couple of problems. One is you could see Here, between about 6 and 7, we start to recycle 50% of our fleet. 50% of our fleet, so that's maybe 75 or 100,000 nodes. That put a lot of stress on AWS because we didn't tell them, hey, at 6 a.m., we're going to blow everything up, and you have to be ready to restart. And so luckily we work with them, and they sort of remove limits so we can respin things. The other problem is, or I should say, the benefit when you auto-restart things is if you have a bunch of stateless services, great, because as soon as they lose connectivity, the auto-scaling group
starts a new one and service resumes. The problem is when you use local storage. And here, we use local storage on AWS for a number of reasons. So you can see here for Kafka, Zookeeper, Foundation, DB, Cassandra, and so on. Part of it is because local SSDs to this day, I believe, are still probably the best price point between performance, sort of the ratio of performance and cost. If you need a lot of throughput, it's hard to beat local SSDs. And remember, we need a lot of throughput because we get a ton of data coming in. And so we need to persist that data really quickly. So the issue when you recycle a node that has local SSDs is that local SSDs get thrown away effectively. They get reallocated to some other customer. Maybe it's another node of ours, maybe it's some... someone else's. And so what the problem that it then creates is for distributed databases, you lose quorum, you lose data, and that happened a lot.
And so at least temporarily. In the end, we were able to recover almost the entirety of the data. But nonetheless, in the moment, it's like, oh, well, the database just went bye-bye, and now we have to recreate it. So now, this is a high level what happened, the interplay of networking and the effect that it had on the various cloud providers and how the cloud providers reacted to that kind of failure. Why it happened is kind of interesting. So I can tell you it's my fault. It's literally my fault. Because in 2013, when the company was about 15 people, What I thought was a great idea is that I'd turn on, this is like an old Git commit for a chef recipe. I'd say, hey, why don't we try to auto-upgrade packages to respond to critical security issues? That way we won't fall behind. And it sounded like a good idea. And actually, it worked beautifully for seven years, for a long time. But for the first seven years, we were pretty much always up to date because every time a package had a critical security issue, an upgrade would apply and it would restart and we'd be safe.
Where it gets interesting is, In 2020, there's a, and I didn't see the PR at the time, but there's a pull request on SystemD that says, hey, You know, when SystemD restarts, SystemD NetworkD, the component of SystemD that controls NetworkD, when it restarts, it flushes foreign addresses, it flushes a bunch of stuff. So why don't we flush the IP rules as well, just to be consistent, clean? And that gets accepted, and in V248, that's what it does. When it restarts, it starts blank slate, and so you have to reapply everything. The next version, 249, they realized that this was probably not a... This is effectively a breaking change. And so what they added, what the maintainers added, is an option, which you can tell SystemD, if the option is present, SystemD will preserve the existing IP roles when it restarts. So it won't restart from a blank slate. Now the problem is, in 2020, I didn't know that existed, or no one in the company knew that this landed.
So we didn't add the option, because we didn't know it existed. And so again, 2020 to 2023, nothing bad happens. Actually, good things happen because critical secure updates continue to get applied automatically. And at that time, between 2020 and 2023, there's 2022. In 2022, Ubuntu, which is the operating system we run across the board, releases a long-term support version. And generally speaking, we want to stay pretty current. So we go from... We decided to go from 2004 LTS to 2204 LTS, just to remain generally compliant. We don't want to run out-of-date packages of software and so on and so forth. So everything is motivated by really good, at least it seemed at the time to be good reasons to do it. And interestingly enough, so between, we started in November 2022, we started to upgrade from 2004 to 2204.
And by the time March 8 rolled in, we're about 90% done. Of course, the irony is if we had decided not to do any update, none of this would have happened. If we had kept 2004, where it was and decided, you know what, we'll wait on 2204. None of this would have happened. But the problem, of course, if we had an upgrade to 2204, is we'd have a big push as it gets closer to the end of the long-term support, really. So there's always a trade-off. So we say, okay, let's get started. And from November through March, as we roll up, and you can see it's a gradual upgrade. We don't flip a switch all at once. Everything works fine. We don't have any major issues. So we feel pretty confident, yeah, this is going to be totally fine. So why didn't we see the problem before March 8, 2023? And the core reason is, well, I've mentioned some of the reasons, but ultimately we never restart processes ourselves.
Like our release processes, our deployment processes never restart processes. We just kill a node, spawn a new one with new images or new processes, and that's how we do it. So effectively, we had never tried to restart SystemD. And we would never try except that package that I installed almost 10 years to the day before March 8, which does exactly that. It installs a new package and it restarts. And when it installs that new version of SystemD, it did not automatically add that special flag that preserves the state. It just omitted it. So we restart and everything goes crazy. And why it was so concentrated is because the default rule of these unattended upgrades that patch critical secure updates, the default happens to be at 6 a.m. With a 60-minute jitter, well, positive jitter. And so, which means, which is why the issue is starting to happen around 6 a.m. From 6 a.m. To 7 a.m. Everything starts to, all the updates start to be applied, and the system D starts to restart, and the nodes start to lose connectivity.
So in summary, March 7, the patch is uploaded by Ubuntu. It takes after 6 a.m. UTC, so it takes an entire day for the update to be applied on the machines. And the network D restarts, flushes IP rules, Kubernetes goes completely dark, all the nodes go dark. We estimated about 60% of the nodes got disconnected from the network, partially because, you know, why not 100%? We hadn't completely upgraded, and then there were some other conditions that came into play. The interesting, and the other thing I wanted to point you out is, in GCP Azure, when that happened, all we needed is a reboot. Because the reboot actually starts with a blank slate and then applies the right rules, the restart does not do that. On AWS, the nodes are terminated, which means that we have to rebuild a lot of stuff. And so what also that meant is initially we thought GCP and Azure were much, the impact was much worse, the recovery we thought would be much harder, and it turned out to be the opposite.
Okay, so I've talked about technical piece, you know, what happened, why it happened. Now I want to talk about the humans in the loop, our response. The first, this is the initial reaction. I was like, oh, shit. And this is some internal graphs, and you can see at some point they stopped publishing because everything is really down. But this is internal graphs that show, and the red is bad, and it goes from 0 to 100%. And it's almost, you know, basically shortly after 7 a.m. UTC, everything goes 100% bad, almost 100% bad. And then we lose our internal monitoring. This is a Slack channel. It's hard to read, maybe a little bit, but it starts, hey, it looks like we have some problems. Let me check. It's like, uh-oh, everything's down. Okay, let's start a Zoom so we can all figure it out. And then, you know, let's get started. And so upgrade to Sev1. Sev1 is the highest criticality for our incidents. Second reaction, quickly, okay, now we know what to do, let's get to work.
And you have a lot of people running towards the problem, asking, what do I need to do to fix this? The interesting thing, which you may have missed on this, is at least 400. So we estimated, we counted in this case, baseline is always 500 people in response. Sometimes we had about 800 people for 48 hours. So we had to figure out how to get 500 to 800 people productive, fixing the problem for 48 hours straight. There was very little sleep, at least from my point, not a ton. But we managed to go through, and I'm not going to tell you why. Basically, the good incident response we've discovered is two things. It's something you can train for and something that you need a good culture for. You need both things to make it work. It actually happens that we've had at the very least 10,000 drills in the past 13 years. Because we've had literally about probably 15,000 incidents of all kinds.
And why we've been able to operate a reasonably successful business with 10,000 to 15,000 incidents in 13 years is that we built from the get-go this culture. You build it, you run it, you own it. So all the developers, most developers carry a pager. I mean, they carry a phone. They don't have a physical pager. And that means that whenever there's a problem, they don't know if it's bad or if it's tiny. They just know what to do. They go and say, okay, let me check what I need to do. We've run blameless incidents since the start first. And also, they are blameless regardless of severity. SEV 5 is somebody's going to do an update and want to say, hey, heads up, I'm going to update something. I'm creating an incident because it's a good way to signal to the rest of the engineering team something's going to happen. SEV 1 is what we saw, it's like everything's down. And so regardless of the severity, we keep it blameless. I don't care who, I care what happened and how we fix it and then what we learn.
We use out-of-band monitoring, which means that when we go down, we have external stuff that tells us, hey, you're down, which is very useful. We also use our own incident app, which I think gets people used to responding to an incident. We have, again, important as well, the same incident process, whether it's SEV 5 or SEV 1. It's all fairly, and it took us years to really fine-tune that. But the good thing is you don't have special incidents, or it's just an incident. And maybe we started at SEV, it's happened that we started at SEV 2, meaning it's an outage, it's serious, but not the end of the world. And then we realized, you know what, it's not as bad, let's downgrade to SEV 4. And we have this constantly. But we do have a sort of, for SEV2 and SEV1, a special rotation of people. So it's whenever anybody in the company can say there's an incident and it's a severity one or two, and somebody somewhere around the world is going to get paged, and that person has
sort of the experience, the training to deal with fairly complicated things or fairly high-impact things. But eventually, so we started with, you know, I was probably the number one, and then we added a couple of people, and now it's about 30 people who are trained, certified to operate at that sort of level of complexity, pressure, and that's very useful. Not for the same stuff I used to be on call 10 years ago or 30 years ago, but I'm on call nonetheless. And so are all the execs. Now, how we organize is, we didn't invent this. This is, the playbook is, the disaster response in the US follows generally this. You have an incident commander. It doesn't have to be the highest, the sort of most senior person in the room. It's just whomever happens to pick up the page on an incident command rotation, that's the incident commander. That person, basically, their job is to make sure we manage to close the incident. And for this type of complex events, we need to spin up a comms lead, work stream leads, exec lead, customer liaison.
So there's specific roles that we can say, okay, you're going to be, the incident commander says, hey, I need a comms lead. And we have people who can be comms lead who have the experience to say, okay, either the pager determines who it is, or if the person happens to be in that Zoom, I'll take, I'm it. So now, if you have anything around communication, I'm the person who effectively drives the whole thing. And then responders. The scale, though, is pretty tough. So for each work stream, we spawn a Slack channel. And there are about 70 or 80, so it got pretty crazy pretty quickly. We discovered some interesting things, like you can have no more than 100 editors on a Google Doc. And so we had to, like, oh, shit, nobody can edit. So we tell on Zoom, everybody close your tab. And then we reopen in edit mode and we send a view mode on Slack. It's interesting. We had a lot of people. This was one of the core ways to share context. Remember, this thing goes on for 48 hours.
So you can't just wait for it. Somebody is not going to stay around for 48 hours on Zoom. That's not possible. So we need to sort of memorialize that so that people can pick it up. Other things, of course, is we have 80 plus Slack channels. You can't read 80 channels at once, and you can't copy-paste over 80 channels. So what we do is we get all the work stream leads in a specific Slack channel, and this is how we try to aggregate the information and redispatch. And there are some responders that know the impact across different work streams, and they'll say, oh, you know what, this is going to impact so-and-so, so let me just share that message across. So it's fairly organic. Now you can see this is the scale of machines we restarted. We restarted about 60% of machines once we figured out what's going on after maybe two or three hours. And then there's a bit of a lull because all that coordination complexity made it so that it took us another maybe six to eight hours to figure out, oh, you know what, now we understand the problem.
We know exactly what we have to do. These are the blue bars. The blue bars we need to restart as well. If we had perfect communication, these blue bars would be all the way to the left, because we know exactly what to do, but it's not the case. What we discovered too is that we have training, we have rehearsals, we have incidents of various severity, but effectively a lot of stuff we didn't do. We had never done before. So this problem solving at all levels, at the response level, at the communication level, we have to invent as we go. And that's okay. That's totally okay. And we don't get it right all the time. Sometimes we try something and it's like, ah, this doesn't work to communicate effectively, so we have to do something else. One thing we figured out is we had to designate somebody to be the roving troubleshooter. And that person was going work stream to work stream to get a sense, is everything working? Are they blocked, or are they sort of spinning their wheels, or are they actually making progress?
Quick shout out to support. This is the multiple of support tickets we got during the incident. So one is everything's normal. We get always a baseline of support tickets. And you can see that about two hours after the start, it looks like we go to 30x the number of, and then 40x the number of tickets and customer interaction. So a lot of work to help you, if you're a customer, to tell you what's going on, what to expect, what we know, what we don't know. Exec support, so this is maybe... This is the VP of infrastructure, Dada Dog. He didn't wear that costume that day, but this is his signature picture. You know, a lot of the... A lot of the value was to stay calm, like handle customer messaging and really resist the sort of pressure from customer, which is totally understandable, and just project that image of calm. And also make sure that we don't burn people out.
Some lessons learned, I think, to sum it up. One, and this is the, you know, 10 years ago, this is what happened. Complex systems grow in surprising behaviors, you know, that can only be discovered much later when conditions change, or it's impossible to predict. Which means that having an efficient response and an effective response is much better than being to predict failures. Because my belief is you can't. Things are going to break. The only way for things to not break is to turn everything off. I think an efficient response, an effective response requires creativity, which itself means that people need to be focused, which itself means that people need to feel like there's nobody breathing down their neck. They don't have to be scared. They just have to focus on the task at hand. A blameless response, what it buys you, it actually creates a space for that psychological safety. So people don't feel like, oh, there's an incident, my God, am I going to lose my job? Which means that they don't mind jumping onto an incident because they know this is the job and doing that is
the job, which means that when something bad happens, they're ready to run into the fire. And that's what we want. Lastly, I think, or almost lastly, engineering execs, particularly in this group, has to be able to withstand market or customer pressure. I've been shouted at on these that day, and I deserved it, but nonetheless, it's not fun. So I can't turn around and start shouting at people. That's bad. That's going to undermine the whole response. And lastly, maybe most importantly, so you're not, however your incident response is, it's not perfect, neither is ours. The only way we found to make it better is to practice. You know, practice, practice, and practice again. Thank you. Thank you, Alexis. You're welcome. Ok. We can be bilingual for the question if you want.
It can be in French or English. That's the wonderful thing about it. There we go. We have five short minutes for questions. Who would like to ask Alexis a question about this experience? There we go. Do you have the microphone? Yes. Great. French? Yes. Let's go. Thank you for the presentation. You explained the cause and how the team reacted while everything was on fire. How long did it take you to find the cause? Was that done separately from incident management? We found it. It took us about two hours at the beginning. Because at first, when everything was down, we didn't know what was happening. I actually brought in the security team, because it hit all the data centres across three different cloud providers at the same time. There are no connections between the data centres, nothing. So we were thinking: this is impossible, what's going on? Has someone managed to hack everything?
And we found that wasn't the case. After two hours of investigation, someone noticed something strange had restarted at that moment, and tried starting... a new machine, tried restarting systemd, and realised that was what was causing the crash. Then we had to work out that for GCP [provider name unclear], we only needed to reboot; for AWS, it would be more complicated. All the details and diagrams came afterwards. But because the problem was so deep in the infrastructure, it had to be fixed before we could bring everything back up. Otherwise, we didn't know exactly what order to restart everything in, or whether that would be enough. So we spent quite a while finding the cause, to be reasonably sure that the fix would solve the problem. Then we could start the long, difficult recovery, knowing we were building on solid foundations.
Right. You said it dropped all the configurations? Yes. How did you restore them? From your images? No. Usually, a reboot brings things back correctly, because the boot sequence includes systemd, where there's nothing initially, followed by scripts that start up and add things back: Cilium and so on. For GCP and Azure, rebooting everything was enough because all the volumes were still there. On AWS, when it restarts, it boots correctly and networking is there. But the volumes are blank. There is nothing left on them. So that's where we had to... We had to improvise a little. Because we had never done... Before that incident, we had never... We run quite a few game days. We try out things we think might happen. But we had never imagined this one. We have a second question, with two hands raised on the right. Hello, thank you for sharing this very detailed account.
It shows an extremely broad failure domain. With hindsight, what architectural changes did you decide to make to avoid this kind of problem in the future? Obviously, we stopped unattended upgrades. That's over. Since that day, it's finished. Gone. Because that was the only channel we had missed. All the other update channels never make changes in place. They start something new, then stop the old one. That's it. Other things we have done include... This gets a little technical: we have tens or hundreds of Kubernetes clusters that are themselves managed by other Kubernetes clusters. We are now decoupling them: the clusters that manage the other clusters will be managed by someone else, namely the cloud provider. We hadn't done that before. When we started with Kubernetes in 2018, I think the largest cluster on the most advanced service at the time [name unclear]
was limited to 100 nodes. And we easily had tens of thousands. How were we going to do that? We couldn't use that service. That's why we built everything ourselves. That has advantages and disadvantages. One disadvantage was that we exposed ourselves to a risk without realising it at the time. There is a clear tension: the same infrastructure everywhere is great to manage, but creates a kind of monoculture. That makes us vulnerable to errors like this. Thank you very much, Alexis. I'm sorry, our time is up: we had barely ten seconds left. That was fascinating. Thank you especially for being transparent about what happened. I think everyone appreciated you sharing it. Thank you, Alexis. Thank you.
