Tech.Rocks Summit 2024

Les modèles de résilience

Tech.Rocks Summit 2024 · 2 décembre 2024 · 53 min · en anglais

Résumé

En informatique, la résilience est souvent comprise de manière très restrictive, centrée sur des aspects purement techniques comme les modèles de défaillance isolés, la redondance ou l'autoréparation, au détriment d'autres facteurs essentiels. Sam Newman explore deux modèles pour comprendre ce qui fait la résilience : il interprète les « Quatre concepts pour la résilience » de David Woods dans le contexte des systèmes logiciels, avec des conseils pratiques pour les organisations, et introduit les systèmes sociotechniques pour montrer l'imbrication des aspects sociaux et techniques de la production logicielle.

L’essentiel

Sam Newman propose de voir la résilience comme une manière dont un système se comporte, pas comme une propriété technique, et la lit à travers trois modèles : les quatre concepts de David Woods, une grille d’analyse de la résilience et un modèle hexagonal des systèmes sociotechniques.

Pour élargir une discussion sur la fiabilité au-delà de la redondance et des SLO, et y inclure les processus, la culture et les objectifs de l’organisation.

Les idées clés

  1. La résilience n’est pas quelque chose qu’on possède ni une propriété binaire : c’est une caractéristique de la manière dont un système fonctionne, qui doit maintenir les opérations indispensables face à l’attendu comme à l’inattendu, quitte à ajuster son fonctionnement. Sam Newman décrit le passage de la gestion de la sécurité, qui cherche à empêcher les événements négatifs, à l’ingénierie de la résilience, qui cherche à ce que le plus de choses possible se passent bien. à 4:10
  2. Les quatre concepts de David Woods : robustesse (absorber les problèmes connus, sachant qu’ajouter de la complexité pour être plus robuste augmente la surface de ce qui peut mal tourner), rebond (se rétablir après l’incident), extensibilité (faire face à l’imprévu, ce qui relève surtout des personnes et des processus) et adaptabilité durable (apprendre en continu, bien au-delà d’un post-mortem de principe). Il juge les SLO utiles mais insuffisants, car la capacité à anticiper et à apprendre ne se mesure pas objectivement. à 11:45
  3. Les systèmes logiciels sont des systèmes sociotechniques : technologie, infrastructure, processus, personnes, culture et objectifs s’influencent mutuellement. Dans la panne de Telstra qu’il analyse, un basculement manuel reposait sur une seule personne, et la réaction publique de la direction, qui a désigné un coupable, illustre selon lui une culture du blâme qui décourage de signaler les problèmes. à 30:38

Questions pour votre équipe

Il s’agit d’une keynote conceptuelle, tirée d’un livre en cours d’écriture : les modèles sont présentés comme des outils de réflexion, et l’intervenant dit lui-même ne pas être convaincu par la grille d’analyse. Les exemples (aviation, rail, Telstra) sont racontés à titre d’illustration, sans analyse chiffrée, et l’affaire Telstra repose sur son propre récit.

Chapitres

  1. Présentation
  2. Définir la résilience
  3. De la gestion de la sécurité à l’ingénierie de la résilience
  4. Les quatre concepts de David Woods
  5. Mesurer la résilience : grille d’analyse et SLO
  6. Les systèmes sociotechniques
  7. Telstra, déraillement et Grenfell : des exemples
  8. Questions de la salle

Summary

In IT, resilience is often understood very narrowly, focusing on purely technical aspects such as isolated failure modes, redundancy and self-healing systems, while other critical factors are overlooked. Sam Newman explores two models for understanding what shapes resilience: he interprets David Woods' “Four Concepts for Resilience” in the context of software systems, with practical advice for organisations, and introduces sociotechnical systems to show how the social and technical sides of software delivery are intertwined.

Thèmes : Architecture & développement

Transcript complet

Transcription automatique, à relire : les noms propres peuvent être mal orthographiés.

Comment ça va ? Ah, parce que là on rentre dans la phase digestive. Il ne va pas falloir mollir, je compte sur vous, à minimum. D'accord? Ah, j'ai un peu de ma bouche d'oreille. Ça vous gêne quand ça fait du bruit dans ma bouche d'oreille ou pas? Oui, ça te gêne Philippe, je vais l'enlever. Et vraiment, je l'enlève pour toi. Comme ça, on ne va pas être dérangé du tout par ce que je vais raconter. Parce que là, je vais vous parler de Roméo et Juliette. Roméo, Juliette, qui a vu la pièce? Qui a vu le film? Ah, la culture geek, j'adore. Bref, ce n'est pas le visage, c'est le cœur qui fait la beauté. C'est exactement ce que disait Roméo à Juliette. Et si nous appliquions cette sagesse à la résilience, nous comprenons que ce n'est pas seulement la structure d'un système qui fait sa robustesse. Mais la connexion profonde entre ces composants, qu'ils soient techniques ou...

Merci, il y en a quatre qui suivent, ou humains. Et c'est Sam Newman qui va nous guider au-delà des seules solutions techniques pour nous montrer que la véritable résilience, et je prends mon café parce que j'aurai besoin de résilience tout à l'heure avec ce café, réside dans l'équilibre subtil entre l'humain justement et la machine. Et je vous demande d'accueillir... Sam Newman, qui est technologiste chez Sam Newman et Associates Limited. Sam, the floor is yours. We are so happy to have you with us. Thank you. So I hope it will be a comfortable moment for you. Lovely. Please raise your hands. I'll be there. If you need me, I can be back. I'll give the secret signal if I'm in trouble. Yes, like that. Please help us out, me. Okay? Thank you so much. Thank you. Thank you so much for inviting me here. This is an amazing venue, by the way. It's not every day you get to perform on stage in this forum. I think the weirdest place before I've spoken was at a cinema.

This is slightly less daunting in terms of the size of the screen. I also do not speak French. I recognize French words when it relates to food, and that is as far as it goes. So I am going to apologize in advance. This will be in my native tongue of English, which I speak moderately well. You can always find me later on if there are any translation issues. I'm going to be talking to you today about the nature of resilience and taking you along for the journey I've been going on myself for the last couple of years. I'm in the process of writing a book, and I like writing books. It's one of the things I enjoy doing. One of the challenges with this particular book is that I want to start with the small things and the big things as well. This book is all about helping people build resilient systems. I have spent a month of my time writing seven books. thousand words about how to tune a timeout. These are important things for developers to know, but they are also very small things in the bigger picture of resilience. To talk today is about those bigger topics which I've been exploring as part of writing this particular book.

Lots of work and research, and I'm going to be sharing some of those things I've been looking into with you today. We're going to be exploring three models for how we think about the nature of resilience. Some of these models work better than others. We're going to be looking at Woods'four concepts for resilience engineering, Elliot Holnangle's resilience analysis grid, and also a hexagonal system for looking at socio-technical systems. And for those of you who know my previous work, you know that I love a hexagon. We'll also be talking about what modern-day philosopher kings Chumbawamba think about resiliency as well. Some of you might already know the answer to that particular question. We are, though, going to start off by trying to define what resilience is. It means different things to different people. This is my talk, and this will be my definition for what resilience means. I'm actually going to be taking somebody else's definition, though. The first thing to say, though, is that resilience is not a thing you have. It's not something you own. It's not like having a pile of money and saying, I am rich, I have money.

Resilience isn't a binary property. It's not like you throw a switch and say, I'm resilient, now I'm not resilient. That's not how it works. Resilience is a characteristic of how A system performs. It's a quality. It's closer in some ways to a quality attribute than anything else. And that changes how we think about resilience. This is my favorite definition at the moment. It's quite a wordy definition. I'm going to break it all down, and this is going to give us some jumping-off points for looking at the topic of resilience. A system is said to perform in a manner that is resilient when it sustains required operations under both expected unexpected conditions by adjusting its functioning prior to, during, and following events. There's a lot going on here, so let's break it all apart, talk about these constituent bits, because this is wordy for English, and many of you, English is not your first language. This took me a while to parse, so let's parse it together.

So let's talk about this idea of a system sustaining required operations. When we think about a system being resilient, we have to define the things that need to happen. What are the important things that need to continue operating when all other things fail? When something goes wrong, what still needs to happen? We also, though, need to be prepared for both the expected and the unexpected. Systems which are only prepared to deal with the expected end up becoming brittle. So we've got to be prepared for both the known and the unknown. But we also have to accept that for any system to be resilient, it has to be okay to adjust how it functions when something goes wrong, and that this is acceptable. We then have this interesting trade-off between wanting to sustain operations to make sure things keep working, but also accepting that sometimes those things need to be traded off. We're going to be exploring all of these ideas throughout the talk. But you could maybe distill down that more sort of wordy definition of what resiliency is to being able to say, we want to be able to ensure certain things remain true when things go wrong.

We think about a plane. I actually got the train here. One of the benefits of living in the UK and this being in Paris is I can get the train. But I may have flown here. We have two engines on a plane. You could have taken off on holiday somewhere. Something could have happened to one of those engines. This is, of course, unfortunate. Something has gone wrong. The engine has failed. What are the required operations that we need to maintain in that event? Well, making sure that no one dies and the plane still lands. But we accept that we might need to change our plans. Maybe now we're not going to Malaga to sign ourselves on the Cotterdale crime. Instead, we're going to have to land back at the airport we took off from. There's some trade-offs happening here. Against this backdrop, we also have this shift from traditional safety management towards what we now refer to as resilience engineering. With traditional safety management, the focus is very much on trying to stop bad events happening in order to prevent bad outcomes.

So if you think about a computer, for example, what are the bad events that we might try and stop happening? Well, with a computer, what can fail on a computer? We don't want a computer to fail. So what fails? Moving parts, power supplies, spinning platter hard drives. If those are the things that can fail, we're going to add redundancy. So if one power supply fails, the computer keeps on. operating. This is us trying to eliminate that bad event, which is the computer failing. There's nothing wrong with this, but it can be quite limiting. We're spending all of our time trying to stop something bad happening. The problem is that systems are messy, they're more complex than that. They involve many moving parts, not just in our software, but in terms of the human beings being human. And this adds a degree of complexity that means that traditional safety management can be overly limiting. The shift towards resiliency engineering is saying that not only should we be doing those things from traditional safety management,

But we should also be thinking more in terms of making sure that as many things go right as possible. So rather than stopping bad things from happening and our focus being there, instead our focus should be on making sure that as many good things happen as is possible. So this is really about us being able to operate even if bad things occur, rather than trying to eliminate all bad things. If we accept the premise that a computer can fail, and that why we might want to spend some energy in ensuring that computers don't fail, it is also inevitable that computers will fail. So if we want to keep making sure that as many things as possible keep going right, then we want to build systems that don't care when a computer fails. And that then leads us into interesting places. This is a real server from Google. This is about a 15-year-old server. You see some interesting things about this server. You see it has no case on it. It is run like this in the rack. One of the things that's most likely to fail things is moving parts.

We see the power supply and the two spinning platter hard drives here secured with Velcro. There's no redundancy here. That's not even a RAID system. If a power supply fails, that machine fails. If a hard drive fails, that machine fails. Google recognized that and instead focused on ensuring they build systems that don't care. And what they did instead was made it as easy as possible to get that machine up and running again. The reason Velcro is used to secure these things is because an engineer will go around in the morning, ripping out the old power supplies that have failed, throwing them in a bin, and taping on a new one and turning the machine back on. This is actually an example of one of their racks. This particular rack is in the Silicon Valley Computer Museum. There's also one, I think, that was donated to the Science Museum in the UK. This was worked for, they had tens of thousands of these things running, and it was fine. We can be looking at these three models as a way of exploring different aspects of this idea of resiliency engineering. I think two of the models work pretty well.

The third one, I'm still on the fence about. I'll share my thoughts when we get there. We also have to remember that we are looking at models, they are abstractions, they hide information, and we have to accept that all models are wrong. As statistician George Box once said, some models are useful. Well, you can decide how useful you think these models are at the end. The first model we're going to look at, which is in many ways the piece of work which started my journey down this line, and this is work done by David Woods. This is a seminal paper in the area of resilience engineering called Four Concepts for Resiliency. The full paper is difficult to read. David Woods is looking at resilience engineering in the whole. He's not narrowly looking at computer-based systems. He's thinking about everything from corporations to air travel to hospital surgeries to entire biological systems. As a result, the paper is quite generic and quite academic in its language and, I would say, quite hard to understand. I have attempted to take these ideas and distill them down into things that make sense for the software-based distributed systems that most of us nowadays are building.

We're going to be looking at four concepts in this model. Robustness, rebound, graceful extensibility, and sustained adaptability. Let's start at the beginning. These are all, I think, actually quite sensible ideas. We'll start with robustness. When Woods talks about a system being robust, he's referring to our system's ability to absorb known adverse events. So we have a think about what could go wrong, and we are prepared for those things to go wrong. This is how well can our system handle issues that we expect. The planes that we fly on have two engines. Why do they have two engines? So that if one engine fails, that is something we can expect. Maybe it's a bird strike, a mechanical failure. If one engine fails, we've got a good chance of still being able to land the plane. It is an expected issue, and that is built into the system. When we think about a service-based architecture, it would be common to have multiple instances of a service behind a load balancer.

And often the reason we are doing this is because we are predicting the fact that an individual instance might fail, perhaps due to a hardware failure, and we are thinking about that in advance and tolerating that. that failure because we can still support the requests coming in from the remaining nodes. We expect a machine to fail, therefore we build things into our system to anticipate that failure. There's an interesting almost contradiction or challenge here. And as Woods points out, expanding a system's ability to handle some additional perturbations, these are those error conditions that we might predict, that increases the system's vulnerability in other ways to other kinds of events. Distilling it down, Often, we want to make our system more robust. To do that, we add complexity. And by adding complexity, we increase the surface area of things that can then go wrong. Cue a very simple stack diagram of a Kubernetes architecture.

Much more complicated architecture, often implemented on the grounds that it's going to make things more robust, and yet just gives us more things that could potentially go wrong. Let's look at rebound. This is our ability to recover once a traumatic event has occurred. And this again is one area where resilience engineering goes beyond this idea of traditional safety management. Safety management was all about robustness. It wasn't about the rebound. So how quickly can we recover when a traumatic event occurs? What can we do when something goes wrong? How quickly can we get back up and running? That plane that couldn't take off, those Holidaymakers didn't get to go to Malaga. How are we going to recover? Well, firstly, let's land the plane. Step one, tick, done. Now how quickly can we rebook those passengers? How quickly can we get that plane serviced and get the issue fixed? Provide support to the air crew who are probably going to be traumatized after a near-death experience. Is that built into your system? If not, if you have a bad ability to rebound, you're constantly going to find yourself on the back foot when bad things occur.

And we all know that bad things will occur. How quickly can we bring back a replacement node when one fails? If a service has been offline for a period of time, how quickly can a system recover its state? In this case here, maybe picking up and resynchronizing our inventory state after it goes offline. Let's look at the third concept from Woods. Resilience as graceful extensibility. I think there's a lot of words going on here. I think we could just simplify it down to extensibility, right? And this is really, though, I think, how well do we handle the unexpected? What can we do when things we do not expect occurs? How do we handle those situations? This has nothing to do with software. This has precious little to do with hardware. This is all about us, how we react, our processes, and everything else. We cannot anticipate every eventuality. Surprising things will happen. There will be a global pandemic we could not predict.

A volcano may erupt under a glacier, grounding travel. A country might elect the same ill-fitting leader twice. I'm not mentioning anything about your election. I'm not getting involved with that one. But I feel far enough away from the US that I can pass comments, right? How did that happen twice? Surprising things happen. Are we prepared for that? By the way, it turned out that everybody did anticipate the Spanish Inquisition because they actually sent... a letter 30 days in advance of turning up, which is an interesting little tidbit I found while researching this talk. Let's talk about the fourth concept of resilience from Wood's model, this idea of resilience as sustained adaptability. Again, wordy titles. This is our ability to adapt our system over longer timescales to continue to evolve and make the system itself more resilient. This is our ability to learn and to adapt our system.

It's one of my favorite quotes that I think distills down this concept. Sustained adaptability takes the concept of adaptability one step further. It calls for continuous learning and transformation, using crises as windows of opportunity to evolve and innovate. Sometimes people simplify this down to this as, maybe we should have a post-mortem, which is the most tokenistic side of this. This is about adapting your entire system to say, our system is about learning. If you fail to learn, all you're ever going to do is continue repeating the same patterns and be even less able to deal with what happens next. Again, this is one facet that a lot of brittle systems lack. They may have been stable for a known set of issues and outcomes, but never adapted. So, robustness, rebound, graceful extensibility, and sustained adaptability. These all give us a bit of a... framing to help us understand the breadth of what resilience can mean. And so although Woods'work is often quite generic, as it should be given he's looking at the field of resilience engineering in the whole, from the point of view of our software-based systems, this already starts helping us understand there's more to the world than this.

This then takes us on to another topic, which is thinking about how we measure how resilient we are. And this, I've got to be honest with you, is that part of the talk where I have to say, I'm not sure I'm happy with this model, but I'm going to show it anyway, because I think it gives us an interesting lens through which to look at these things. This is where we're going to look at Eric Kolnagel's resilience analysis grid. And Eric's work has been fundamental in a lot of the resilience engineering work. It's great stuff, but I'm not sure his work necessarily applies into our work as easily as it does elsewhere. So if we accept that resilience is how a system performs, then it follows that we need to find some way to measure how resilient we are. And so in this paper, Eric was defining a grid where we would analyze our system's ability to be resilient on four criteria, different to Woods'criteria, annoyingly. Not the same. So he looked at how good are we at responding. How good are we at monitoring so that we can find out that things might be going wrong before?

This gives us a bit of a sense about our anticipation. How well we can learn and how well we can kind of have that foresight. So monitor is kind of more knowing what's going on. That's more insight today. Anticipation is more taking that information and looking at the longer-term trends. Actually, the original paper is very approachable, even if you don't know all the verbiage that often comes from resilience engineering. And this model is actually used by systems all over the world to help analyze their internal performance. The issue is the way we assess these things. We're looking here at the resilience of the system as a whole, all the moving parts. When we look at our third model, a system is quite a big thing. So the way that Eric suggests we analyze how resilient we are is by lots of questions, basically. These are just summaries of some of the questions that he might expect you to ask. This is just from one of those quadrants. A lot of this is about talking to people. A lot of this is about subjectivity.

His suggestion is that we use this analysis not to say, how good are you now, but to say, how good are you going? This is about trends. This is quite important with any measurement which includes some subjective measure. It is about trends and comparison, not about absolute comparisons with other people or other parts of the organization. This is used in the Australian Nuclear Research Association, uses this model for looking after its nuclear power station. It only has the one. There's reasons we'll go into later, maybe. This is heavyweight stuff that makes sense in safety-critical systems, but we are not building safety-critical systems. What we're used to in terms of measuring resiliency is maybe the Google SRE-centric view of the service level objective, the SLO. Some people live their lives by SLOs. And the SLO is great because it is an objective measure. It is a number. We can all agree on what the number is, and we can then say what number is a good number and what number is a bad number.

I feel good about myself. I feel bad about myself. And we've distilled the... world down into this. We have simplified our measurements down. And look, there are some things which are quite easy to measure, but the SLOs that we often use to decide how well we're going as a team or how well is our system going primarily focus on the non-messy stuff which is specific to a very narrow aspect of our technology. We're focusing on things like latency, availability, data loss, things which can be measured. We can all agree on how long a millimeter, a centimeter, or a meter is. We can say, yes, that's a meter. We can then use that shared measurement to measure something and agree, yes, that is 1.5 meters in length. That is objective measurement. When defining an SLO, the argument comes down to what is the acceptable threshold, not how to measure or anything else. It's easy, objective, and often something you can automate and give you nice little green boxes on your dashboards, which gives you that false sense of security.

Just as that shift from safety management to resiliency engineering doesn't mean you should ignore safety management, I don't think these measurements are bad. I think the mistake is thinking that they are in some way sufficient. If we accept that to be resilient, we have to get good at anticipating the unexpected and about learning, How do you measure that? How do you measure your ability to learn? How do you measure your ability to anticipate and do so in an objective way on which we can all agree? I don't think that's possible. And so Eric's work is all about the soft side of these services. Now, I have seen people trying to use these models inside sort of modern IT organizations. Fred Herbert used the analysis grid at Honeycomb looking at on-call health. The idea is that you take Eric's grid and adapt it within your own sphere. And so he took a quite simplified version of this and was using it as a qualitative individual measurement, asking people who have been on call to reflect on their experiences, tracking that information over time.

And so I think in the small, there's something there. But as I said, I've yet to find that kind of middle ground between the very kind of quite heavyweight and continuing assessments that are required when using Eric's grid and safety critical systems to the kinds of things that we actually have the appetite for in a corporate environment where, at best, we're doing some halfway decent SLOs. Now, interestingly, I do think the world of developer experience has started to come to grips with the challenges around measurement in an interesting way. There's often been this, sort of recently, there's a fixation on the DORA metrics that came from the State of DevOps reports, those four metrics. The state of of DevOps support said were useful to track in terms of understanding if a system was performing correctly or achieving high performance. The problem is that that became overly reductive. Organizations started becoming fixated on those four measurements. I would go and ask people, so how is your software development process going?

And they said, we're tracking the four metrics. And I said, how is your software development going? We're tracking the four metrics. What does that mean? And they said, well, these are the four metrics you've got to track. Okay, we need to have another conversation, don't we? Because that's not what the world is about. And so in the world of developer productivity and developer experience, we've got things like the space framework, which is saying you probably need to look at a bunch of different things to really understand the extent to which your developers are having a good time. So while the DORA metrics might be a small part of it, there's actually probably five different areas that you might want to be analyzing, which gives you a mix of quantitative and qualitative measurements. The space framework doesn't say use all of these metrics. It says for each one of these spaces in the space model, you kind of want to pick one or two of them. I would love for this kind of work to be done for looking at how we measure our resiliency. Unfortunately, I think so much of the conversations are dominated by that sort of SRE-centric, SLO, SLI, SLA view.

As a result, there isn't a space for anything else. We've got the kind of safety-critical world over here and the we-are-building-software-over-here world with not much in between. Hopefully that will change. This is why this model for me, I'm still struggling to gather it, to deal with it a little bit. Let's come on to a model that I feel is a bit more successful in helping us think about resiliency. And this is really more a model for how we think about what IE system is, and specifically what are the systems that we ourselves are a part of. A computer is simple. It is a... Electronic device with zeros and ones with known inputs will give us known outputs, as long as we exclude the possibility of the odd cosmic ray flipping a bit here and there. So we assume that these are basically simple things, and we kid ourselves that we can often break these things down to mathematical concepts, and in the small, that is kind of correct. The computer itself is simple, but we're not dealing with systems that consist of a single computer. We are building systems that consist of multiple computers, actually multiple services.

Each one of these services with multiple instances. We've got lots and lots of computers. And when you go from one computer to lots of computers, interesting things start to happen. Least of which, how do you run all the cables between all those machines? It's the networks, it's all the complexity that emerges from a more complex technology platform. But we've also got other things to deal with, like the inherent nature of a distributed system. We might be dealing with the vagaries of public internet and public clouds, or private clouds, or power outages and floods, and the worst element of all, people. These are the people that can really screw you up, aren't they? These are the little gummy bits in the machines that are kind of necessary. There's lots of techies out there that would love to pretend that people don't exist. But here's the thing, without people we wouldn't need the software. And the moment you add people in it adds complexity. So maybe we need to understand the role that people play in the systems we build. Systems are...

Messy. Fundamentally, the kind of distributed software-based systems that we are building nowadays are what we describe as a type of socio-technical system. It's a big word, let's break it down. Socio, people, technical. This is kind of in the old-fashioned sense of technology, which basically means knowledge and doing. So this is the, as we'll see, we'll break this down even more. This is about how we bring people together with ways of doing things and understanding how all this fits together. Socio-technical systems was first pioneered by looking at how coal mining was done during World War II. A lot of work done by the Tavistock Institute in the UK. You may have heard of the Tavistock Institute. Weirdly, they were part of a conspiracy theory that they had created the Beatles during the 60s in an effort to bring down Western civilization. There were crazy ideas out there. All I can say is if that was the goal, they're taking a very long time to get through.

But nonetheless, interesting work. Socio-technical systems. But again, a topic I struggled with. It is such a big field. And again, I'm like somebody who likes pictures. I have all kinds of issues with huge chunks of text. I like a picture. I like a model to help me make sense of the world. And I wanted this because I understood we are building socio-technical systems. If we want to make them resilient, we have to understand the system as a whole. We can't just look at the code part of that in isolation. When we think about air travel, we can look at the safety of an engine. Make sure this engine survives bird strikes, it handles debris in the way that we expect. We can try and focus our energy on that. But when we think about flight safety, it's also down to the air crews, how well they are trained. Not just the people flying the plane, but the cabin crew as well, who are often an important part. of any sort of safety systems?

How well trained are the crew who maintain the plane? What about the systems that manage air traffic control and the people that work as air traffic controllers, assuming they're not on strike? All of these make up our systems, our flight travel systems. It involves multiple parts, multiple constituent parts. And I was very thankful to find this model. I struggled to actually find the original source for this, but I found this paper from the Leeds Business School talking about the socio-technical systems theory and their overview, and they've kind of distilled it down into this hexagonal model. So, when we think about a socio-technical system, we think in terms of the technology. So in the case of our software-based systems, this would be the code that we write. We then have the infrastructure that we rely on. This could be the public cloud, the private cloud, infrastructure that might be provided to us, the physical machines, the networking, the power. We've then got the processes.

How are those people organized to do their work? The people, the human beings in the system itself, the culture that those people work within, and the goals of the organization driving those people. And we can consider any socio-technical system in these senses. An important thing we learn from our study of social technical systems is that these things very much interact and influence one another. We can think of the socio stuff on the side, the techni on the other side. And immediately once I found this, it started helping me make sense of previous situations I'd been in, understanding the interplay between various different forces. Many years ago, I lived in Australia. I love living in Australia. It's a nice place. It's a big country. Population, about half of that of France. All of Europe can fit inside Australia with room to spare. It's very spread out. Telstra, who were the main telco, they were previously the state-owned monopoly, had a massive outage while I was there in Australia.

And it made the news. Fixed phone telephony went down, mobile telephony went down, internet went down, all of it. And in a country where a lot of rural communities rely on this for connections, this was a big deal. I actually wrote a blog post about this because I was sitting in a coffee shop and I was bored. That's why I do most of my writing. But mostly I did it because of the way that Telstra responded to this. They actually came out quite quickly with a statement about what had happened. And this caused me to write a blog post that got me in a little bit of trouble. I'll explain why in a minute. The CEO of Telstra came out with remarkable speed, had seen the issue, and told us about what had happened. I'm going to read out parts of their statement. We took that node down. Unfortunately, the individual that was managing the issue did not follow the correct procedure. And he reconnected the customers. to the malfunctioning node rather than transferring them to the nine other redundant nodes that he should have transferred people to. Now, in English, we have lots of nice idioms, lots of concepts.

This is what we would call putting the boot in. This is the COO of a massive company publicly saying, this dude screwed up. I love the fact they took the time to emphasize that this idiot could have sent them to the nine other nodes that were working perfectly fine, and they failed to do so. Nine, not eight, nine. They had nine chances to do this right, and they didn't do this right. The employee responsible didn't follow procedures, and clearly that's not a good thing. Very not a good thing. But I wouldn't want to preempt the proper investigation. Oh, no. I wouldn't want to prejudge anything by already deciding what had happened. And we'll figure out what the right response is when we've had a chance to dig into the detail. I think the right response here was to blame an individual as quickly as possible in a national newspaper. I had friends that worked at Telstra, employed tens of thousands of people. It did not take long for everyone to know who this person was, who had screwed it up for everybody.

And I kind of love this detail as well. We wouldn't want to preempt the proper investigation. I think what you've done here is preempt the proper investigation, haven't you? Because you've come out and told people what went wrong. So that's... That was bad enough, I thought. The response, a quick response, not a problem. This response, bad, because you are immediately blaming people. You're also highlighting a significant flaw in a system. So I then started sort of subsequently thinking, how do I put this in a context? And so I kind of, earlier this year, started to map this back to this model. So we had a node failure. It wasn't exactly clear to me if it was a hardware or software-based failure, but it's somewhere in this space. We've then got a process issue. Firstly, when a node fails, there's apparently a manual process to divert traffic to a non-failing node. So that doesn't feel great, a manual process. Moreover, it's a process which relies entirely on a single person not making a mistake.

Any system that relies on a human being never making a mistake is inherently flawed. So that's not great. We've then clearly got a blame culture from the top. The way the COO reacted in the aftermath was not uncommon, and I heard from people at Telstra that there was a lot of blame that used to flow around in the organization. Now, if you've got an organization which is constantly blaming people when things go wrong, do you think that encourages people to bring things forward when they make mistakes? Oh, I made a mistake and this thing happened. Maybe we should fix it. No, no, no, because in a blame environment, I'm not going to mention anything, am I? I'm going to keep my head down. I'm going to cover my ass. I'm not going to share information because the COO might out me in a national newspaper. And if I'm not raising issues, if I'm not telling you what's going on because of that blame culture, that's going to have an in-play into our goals. Are we going to be looking to fix these problems? No, because we don't even know they exist.

Now, interestingly, despite this outage being caused entirely by one single individual, the COO, who at the time was... being groomed to take over as CEO of Telstra, left their job within a year because Telstra suffered multiple further outages. All I can assume is that they kept putting that one idiot individual back in charge of things again and again and again, because that must be what the cause was, right? Now, there are other forces at play here. We can think about your system in this light, and it is very, very useful, but we do also have to accept the external forces that operate on this as well. You might have external stakeholders, your customers, your shareholders. You might have regulations and things to deal with. You might have circumstances outside your control, like a global financial crash, all of which can influence what is happening to your system. These are things that are largely outside of your control. In the case of Telstra, one of these things was the press. The press were asking what the hell's going on.

My blog post got picked up by a journalist I know. He interviewed me, put it in the paper. I then got in trouble when I found out the company I worked for at that time was still doing a lot of work for Telstra. And apparently that was suboptimal to call the Telstra COO out in a newspaper. It was all smoothed over, but that desire to come out quickly with a statement was driven both on some of those external forces. And again, coming out quickly with a statement is not the issue. The issue is coming out with that statement. It speaks to significant issues inside the organization. I also found some further work from this hexagonal model, including this model, which is really interesting, which involved Chris Clegg, who's pioneered a lot of work around socio-technical systems. And this was an attempt to use this model to determine in advance or predict whether or not failures may happen. I think their prediction model did not work, but nonetheless, this paper is stocked full of a whole load of examples of failures using a socio-technical model to help us understand what happened.

One example would be the Grey Rig derailment, in which 30 people were seriously injured. One person was unfortunately killed in this accident. It could have been a lot worse, but it was still bad. So they highlighted that there were some infrastructure issues. There were literally physical rail issues that caused the derailment. They identified a bunch of problems around the technology. Like the system that was tracking these defects wasn't able to spot that this kind of defect was reoccurring. And that was being missed in the software. There's a whole lot of process issues. Like people weren't going out rigorously enough checking these fasteners on the rail. They also found out that it was culturally accepted to just not go through proper safety procedures, which doesn't seem great for a railway. And those things started getting really interesting for me, and then they started looking at the goals of the particular rail network at the time. There'd been a big shift towards modernizing the rail. And so the focus from a corporate level was on modernization, and that had squeezed out time for regular maintenance activity, which kind of fit into that culture.

We don't have time for maintenance, therefore it's okay to skip maintenance activities. So that kind of led directly into people just not doing things that were part of their job. Because they didn't do it because they didn't have time to do it, and they never got in trouble for not doing it because they were doing the things they were told. to do, and it all spiraled. And this really helped me start to break these complex systems down. And I think it's important to understand these are the systems we're building. You may not be building a rail network. Hopefully you're not. You should be taking advice from me over how to build a rail network. But this gives us an insight as to the way these things can be interconnected. And again, this article is well worth a read. And pictures and seeing how things are connected can be powerful. This is Grenfell Tower, which fortunately caught fire, and 70 people died at the scene. This is right around the corner from where I used to live. I used to walk past this regularly. I moved out before the fire occurred. There was a fire in a refrigerator that was then spread through flammable cladding. I mean, that's the simplest answer and also a flawed answer in terms of what caused it.

That's not enough to say what caused it. This picture came out at the conclusion of the second part of the investigation. These are all the different bodies that were involved in some way, shape, or form with Grenfell Tower. These arrows are one body blaming another body for what they did wrong. This is what's called the web of blame. I suspect this thing will be picked over for years to come. There is power from models. There is power from abstractions. Systems are complicated. This doesn't mean you shouldn't do your timeouts. This doesn't mean you shouldn't think about load shedding and back pressure, and that you shouldn't try and stop one machine from falling over and taking out your own service. That stuff is all useful. and sensible to do. But to be resilient, that by itself is not sufficient. Which brings us to the philosopher king's chumbawamba. While doing this research, I put a call out on Mastodon. Twitter is a cesspit. So I put on a Mastodon, what does resilience mean to you?

And I got back actually now my new favorite definition. Good friend of mine, Bruce Derling, just sent back one word, chumbawamba. And I immediately knew what he meant. Chumbawamba was big here, of course, tub thumping. It's a drinking song. So in the context of that, this loses a bit of value. But I get knocked down, but I get up again. You're never going to keep me down. This highlights that shift from safety management to resilience engineering. Safety management, we're never getting knocked down. Resilience engineering, oh, we're going to get knocked down, but we're going to get back up again. And this, for me, is resilience as well. Maybe in the simplest possible form. Thank you so much for your time. You can find the slides over on my website. Thank you, Sam. Thank you. This is a brilliant demonstration of the fact that we have to be humble just in front of the uncertainty of the world sometimes. Absolutely. I think so, if you agree.

So we have 10 minutes for a Q&A session. Fantastic. I'm sure that a lot of people have many questions to ask you. So you know the code. You have to, yes, raise your hand. Je ne sais pas si Amélie est là. Amélie, are you here? We have to get a microphone to the person. Il faudrait juste vous lever pour qu'on vous voit bien. The mic is running. It's running. It's arriving, yes. One, two, three. This is a collective thing. This is it. We're all working together. Brilliant. Two mics. You're lucky. Hi, thank you so much for the presentation. I just want to ask a question about the external factors. You were mentioning the system itself as a closed system, which should be resilient enough to react to factors that come from inside, but also from outside.

How do you, in your hexagonal design or model, do you integrate the external factors that could be against the system itself? This is about using your understanding of the context. So I'll give you a concrete example. I did work with a Norwegian government agency about five years ago. And the Norwegian government had just decided that it was going to be okay to use public cloud, but only if that public cloud was in-country. That was a decision made by the government to okay this as a government agency. So they now accept, okay, there's a possibility we can use public cloud, and actually doing so would improve the quality of service we could give to, it was a part of the welfare agency there. But they acknowledge what could happen. What happens to governments? Elections. Elections happen, people change, the public winds of opinion change. That's an external force. They can't influence, nor should a government agency be influencing the election. So they had that go-ahead, but they also recognized that that could change. And so when building the system, they had to take into account that they cannot assume that thing is a given.

So for me, it's about understanding and engaging with those stakeholders. Your ability to influence them is limited, right? But your ability, but that doesn't mean you think they pretend they don't exist. So this is about anticipation. I mentioned that sort of model of thinking ahead, what could go wrong here. So I think it's understanding those forces that come in and understanding what can you do to engage and anticipate. So in the example of the agency, they went live on Azure because it was running in Norway at the time. But they made a decision to build their system with an abstraction so that they could move it to a different vendor or move it back on-prem. Net result was they ran entirely in Kubernetes, which added a lot of complexity, actually. It would have been better going straight to the Azure services. But they were paying a cost to hedge the risk against having to move later on. So that gives you an example about how you can play that out. Honestly, if you speak to any good project manager like this, and I'm married to one, so it helps, any project manager would go, oh, yeah, that's risk management. And there's a great book called Waltzing with Bears, which I can thoroughly recommend by DeMarco and Lister, which I think is a great insight to how to think in terms of risk management.

And that is all about how we deal with uncertainty. And I think that's a great book to read. Even though it's an old book now, it's a good book to read if you want to understand about risk management in a software-based system anyway. Thank you, Sam. There are another burning question coming from the people. No. Anyone else? So you have a lot of reference to give to the book club of Tech.Rocks Summit. Yes. If you don't mind. Just do the book club? Is that what I'm supposed to say? There is a book club. There is a book club. You should participate in the book club. I don't... Who is in charge of the book club? Oh, we might have a question over there. There's a question? Oh, we might have a question over there on the side. Two questions, yeah. Yeah. How do you sell resiliency to management? So I think this is a more generic question, which is how would you sell your management on anything? As developers, as technologists, our view of the world is like this.

Management's world of view is like this. We don't share the same context. And the problem comes about when we try and convince management to do something, we don't understand their context. What are they worried about? And we often use the wrong language to communicate. So if you start going to your manager saying we need to track SLOs, what's an SLO? Now I've got to explain what an SLO is. So the way to think about it is I need to sell management. What does management care about? What are their concerns? What are their worries? Have you looked at their risk register, for example? They will have one. If not, I'm not sure what they're doing for their jobs, but they will have one. So for me, whenever you're trying to sell anybody on anything, is what is it they want and how can I help them? So if you're seeing an issue where you're worried about your system failing in some way, shape, or form, Are those things that your managers care about? That's what it all comes down to. What do they want? What is their context? Get on the same page about that. And so I actually think, honestly, in that situation, one of the best things that you can do is grab your manager for a coffee and say, what are you worried about at the moment?

What's going on? Start to hear what's going on in their world. Start to hear the language they care about as well. That's all, all of that's going to help. And then ultimately, you know, you want to put your things you're trying to convince them of in context that they'll understand. Talking about concurrencies and latencies and uptimes, eh, may not help so much. Talking about customers not being able to do things, maybe you find out what the cost would be of you being down for an hour, for example. And say, look, if we're down for an hour, that's going to cost us like $30,000 an hour or whatever it might be, putting it in terms that make sense to them. But it all comes back down to, do you even care about the same things? Do you even care about the same things? If you don't care about the same things and you don't want the same outcomes, you can't sell them on anything because you're going in a different direction. Thank you. There was another question to the right, no? No. Any more questions?

Thank you. It was very interesting. Thank you. About culture, because it's very linked to resilience, how do we change the culture of people and the company? It's honestly, it's about desire. It's very difficult for an individual to decide to change culture and for an individual to change culture. So I think if you want to change culture, it does come down to getting a bunch of people that want to do it with you. And it's probably going to be a lot of small things. So I think it's about saying, I'm seeing a problem. We need to change something about our culture. Does anyone else agree? It's a thing called the Cotter change model. It's very old-fashioned change management, but it kind of works. It's got a lovely phrase. You need to find a coalition of the willing, which sounds like we're going to invade somewhere in the Middle East, but we're not. It's about saying you see a problem, right? Something's wrong.

Next thing you need is a sense of urgency. We need to make this change happen. Then you need people with you that want to make that change happen. And then it's about working together. And it's probably lots of little small things. I was involved to help try and roll out a culture of testing inside a large dot com. And there was basically a little subject interest group that were interested in testing. They thought, no one's doing any testing. What was this manual testing going on? How do you make that change happen? So what they did was they grew the group, the community. They got more people involved. They got them going out doing different things. And one of the people in this group, a colleague of mine called Joe, and Joe said, well, we need to reach people with this information. Where can we find people? Where can we capture people's intent? How can we get something brief to explain what's going on, to share the problem and share the solution? So you come up with these little one-page, like, A4 sheets. They were US. because the US does things weirdly, little sheets, and he thought, well, just a complete summary of how to do something small, make it easy to automate your tests.

People are time poor, where can we get them sat down? The toilet. So he got these things up on the back of toilet doors, and so when people's attention was otherwise you know, engaged, there was a little one pager. Now that's the fun story that comes out of it, but there were loads of little things that they did. It started from, there's a problem, and there's a bunch of people that are with me that want to fix it, and now you're going out and you're doing lots and lots of little things. They did that themselves, and eventually they got buy-in. From on high, who then was able to bring more resources in, bring more people in to help transform it. And then a really interesting thing came about, which was they've just been going for a period of time. We need a culture of testing. We need to get good at testing. Test automation is the way we want to go. And we've pushed some people, and some people have come with us. And now we need to make that big step. And at that point, they got the buy-in from more senior management, who then took the decision to basically get rid of every single manual test in the entire organization, basically overnight. And say they're going, they pulled away that sort of last little support that people had.

That only happened because there was three years worth of build up to that. Lots and lots of little small things. If you want to hear some of the stories about how that cultural change is made, a guy called Mike Bland wrote up his experience of being in a group called the Test Mercenaries. It's not as bad as it sounds, but he goes through all the little things they did as a little group. But hopefully that's useful. Thank you so much, Sam. Time is up, I'm afraid. Thank you for this brilliant demonstration. Thank you. We can go by this way. Do I have everything? I've got my things. Thank you.