Tech.Rocks Summit 2021

The Unicorn Project, and The Structure and Dynamics of Great Organizations

Tech.Rocks Summit 2021 · 9 décembre 2021 · 43 min · en anglais

Résumé

Après 22 ans passés à étudier les organisations technologiques les plus performantes, l'auteur de « The Phoenix Project » présente « The Unicorn Project » : ce qu'il a appris des leaders tech qui adoptent les principes et pratiques DevOps dans de grandes organisations complexes, souvent face à des orthodoxies bien ancrées. Il détaille les cinq idéaux (localité et simplicité ; focus, flux et joie ; amélioration du travail quotidien ; sécurité psychologique ; focus client), illustrés d'études de cas, et explique pourquoi le DevOps restera selon lui l'une des forces économiques majeures des décennies à venir.

Summary

After 22 years studying high-performing technology organisations, the author of “The Phoenix Project” presents “The Unicorn Project”: what he learned from technology leaders adopting DevOps principles and patterns in large, complex organisations, often against deeply entrenched orthodoxies. He describes the Five Ideals (locality and simplicity; focus, flow and joy; improvement of daily work; psychological safety; customer focus), with case studies, and explains why he believes DevOps will be one of the most potent economic forces for decades to come.

Thèmes : Cloud, infra & ops · Management & organisation

Page du Tech.Rocks Summit 2021

Transcript complet

Transcription automatique, à relire : les noms propres peuvent être mal orthographiés.

Jen Kim, auteure de l'ouvrage The Phoenix Project, The Novel About IT, DevOps and Helping Your Business Win, et chercheur, va venir partager dans quelques secondes pourquoi il croit plus que jamais que DevOps sera l'une des forces économiques les plus puissantes pour les décennies à venir. Venez découvrir ses arguments juste tout de suite. Hey there, I'm Jean-Laurent Morlon, the VP Engineering of Docker, and I'm super excited to be there today. Look, I've been working with Dimitri Bailely, one of the Tech.Rocks co-founders, on a list of books that we want to publish for engineering leaders. And while doing so, I was curating the list with them, and we were like, oh, one author was standing out, right? And I'm so very excited to be introducing the talk of this author today. That's Gene Kim. So you probably heard some of this book, the Phoenix Project, Accelerate, the Unicorn Project, which is going to go into details today. And I don't know, like people like talking about focus and flow in 2021, but at the same time, talking about psychological safety and joy can't be wrong, right?

So anyway, you're not here to listen to me. I'm super excited to be introducing Gene King now. So Gene, take it away. Jean-Laurent, thank you so much. It's a delight to be here today. I've been studying high-performing technology organizations since 1999. That was a journey that started back when I was the CTO and founder of a company called Trip. wire in the information security and compliance space. And our goal was to study these high performing technology organizations that simultaneously had the best project due date performance in development, that had the best operational stability and reliability in operations, as well as the best posture of security and compliance. And so we wanted to understand how did those amazing organizations make their good to great transformation, so we could understand how other organizations could replicate those amazing outcomes. And so as you can imagine, in a 22 year journey, there are many surprises, but the biggest surprise was how it took me into the middle of the DevOps movement, which I think is so urgent and important. The last time that any industries have been disrupted to the extent that all our industries are being disrupted today was likely manufacturing in the 1980s, when it was revolutionized through the application of the lean principles.

And I think that's exactly what DevOps is. Take those same lean principles, apply them to the technology value stream that we work in every day, and you end up with these emergent patterns that allow organizations to do tens, hundreds, or even hundreds of thousands of deployments per day, while preserving world-class reliability, security, and stability. And so that's something that I certainly didn't think possible, say, 15 years ago. And yet, I think we take that as something we expect these days. And so the Phoenix project came out in 2013 that Jean-Laurent mentioned. And so I cannot tell you just how much I've learned since then. So what I'd like to do in the next 40 minutes or so is share with you what the top learnings for me have been. And so let me start by sharing what my definition of DevOps is. It is to what extent can we quickly and safely deliver value to our customers through architectural, technical and cultural norms that allows us to increase our ability to deliver application services quickly and safely while enabling rapid experimentation and innovation so that we can deliver value in the fastest possible way to our customers. And we do this without sacrificing security, reliability, and stability.

And so why do we care about that? It is so that we can win in the marketplace. So I love that definition that went into the DevOps handbook in 2016 because we don't actually say what DevOps is. We're just describing what the outcomes that we aspire to are. But as much as I love that definition, there's a definition that I love even more and doesn't come from me. It comes from my friend John Smart, who led the ways of working at Barclays, a bank founded in the year 1635, which predates the invention of paper cash. And his definition is simply this. It is better value, sooner, safer, and happier. So I love this definition for many reasons, but one of them is that it's a shorter definition than what I gave you, and yet just as accurate. And secondly, I love this definition because it's so difficult to object to. Not even your biggest DevOps skeptic would say that they want less value later with more danger and more misery. So I think it does a wonderful job in articulating and selling what we're trying to do within the DevOps community. So I mentioned how much I've learned since 2013, since the Phoenix project came out. And here are some of the top problems I think still very much remain.

One is the absence of understanding of all the invisible structures required to truly unleash developer productivity. The second is this orthogonal problem of how do we get data to where it needs to go. So in the DevOps community, we rightly recognize that it took too much time and effort and danger to get code to where it needed to go, which is in production. So customers are saying thank you. But there's other problem around data, which is often data is trapped in systems of records, systems engagement, where it may take weeks, months, or even quarters to get it to where it needs to go, which is in the hands of people who make decisions. And so somewhere between 30 to 50% of people in organizations use or manipulate data in their daily work. This is even a bigger problem than what DevOps set out to solve. There's often very strong opposition to support these newer ways of working. And there's ambiguity in terms of what we need from our leaders to support these type of transformations. Because not all of us are lucky to have someone like Jean Laurent leading their engineering organizations. It's not just engineering, it's the entire organization. And so these are the areas that I want to explore in what became the Unicorn Project, which came out in 2019.

And so in the Phoenix Project, we had the three ways, we had the four types of work, and the Unicorn Project, we have the five ideals, which are intended to describe concepts that I think are important for us to get from here to there. So before I go into that, I just want to make a little detour to talk about how DevOps is urgent and important. And the reason is, is that without something like DevOps leads to a sense of powerlessness that we are trapped in the system that preordains failure, that is actually getting worse over time, that preordains horrendous outcomes. And this affects us whether we are in infrastructure and operations. Whether we are in development, whether we are a product owner, whether we are in information security or compliance, or whether we are an architect. And by the way, I'm so thrilled that, if I understand correctly, Gregor Hope is speaking to you in this conference as well. I'm a huge fan of his work. And so DevOps helps with all of us in the enablement of our objectives, regardless of what part of the value stream we play in.

So I had also made this claim that DevOps creates business value. And this is based on the work that went into the Accelerate book that I did with Dr. Nicole Forsgren and Jez Humble. And so this is a cross-population study that spanned over 36,000 respondents. So the cross-population study is the same instrument that the medical community used to determine the link between early morbidity and mortality and smoking. So our goal was to understand what factors create high performance. What we found from 2013 to 2019 is decisive. We found that high performance exists and they massively outperform the non-high performing peers. So as measured by what? We know that high performers are deploying more frequently. High performers are deploying multiple times per day, two orders of magnitude more frequently than their peers. We know that high performers, they can do deployments more quickly. So they can go from code committed through integration, through testing, through deployment, to customers saying thank you in one hour or less. That is two orders of magnitude more quickly than their peers.

When they do a deployment, high performers are seven times more likely to succeed without causing a server outage, a service impairment, a security breach, or a compliance failure. And when something goes wrong, because Murphy's Law does exist, high performers can fix those issues in one hour or less. They have a mean time to restore service that is three orders of magnitude faster than their peers. And so for six years in a row, we found that these high performers existed and that the only way way that you can get these sort of amazing reliability profiles is to do smaller deployments more frequently. So over the years, we looked at other dimensions of quality as well. We know that high performers, because they are integrating information security objectives into everyone's daily work, they are only spending one half the amount of time remediating security issues. Starting in 2015, we looked at not just technology performance, but organizational performance. We found that these same high performers were twice as likely to exceed profitability, market share, and productivity goals. And for not-for-profits, say military agencies, government agencies, not-for-profits, they were twice as likely to achieve organizational and mission goals, regardless of how they chose to measure it, whether it was customer satisfaction, quantity, or quality.

And so for me, what this means is that if mission achievement requires work that we do in the technology value stream, DevOps principles and patterns help with the achievement of those missions. We found that in high performers, employees were twice as likely to recommend their organizations as a great place to work. So this is as measured by the employee net promoter score. And so to me, what this describes is what the opposite of technical debt is. To what degree can we safely, quickly, reliably, and securely achieve all the goals, dreams, and aspirations of the organizations that we serve? And so the first of the ideals that I think is so important to get from here to there is locality and simplicity. And so, so much of this was informed by me learning about a technology, how it was created and how it died at Etsy. So certainly Etsy is a great place to learn about technology. one of the e-commerce giants. It went public a couple of years ago, but it was not always a great organization. And this is essentially a story about how teams of engineers work together to create value for customers. So in 2008, this is in the dark days of Etsy.

This is when almost every engineer would quit after every retail holiday season because they knew they couldn't make it another year. But even in 2008, Etsy engineering knew that they had this problem. And it was that two teams would have to work together to create value for customers. would work in the front end. In their case, it was in PHP. And then they would have to work with the DBAs on the back end. In their case, it was in stored procedures inside of Postgres. And so it meant that in order to ship a feature, these two teams would have to communicate and coordinate and schedule together, prioritize together, and even worse yet, deploy together. And so this was a big problem. And so what they did about it was create something called Sprouter. This is short for Stored Procedure Router. And their goal was to create something in the middle that would allow the devs and DBAs to work independently and then meet in the middle inside of Sprouter. And as Ian Malpas, one of their senior engineers, said, this required a degree of synchronization and coordination that was rarely achieved to the point where almost every deployment became a mini outage.

And so if you're trying to do multiple deployments per day, this is terrible. So the act of going from two teams having to communicate and coordinate to three teams having to communicate and coordinate, where the code deployment lead times went way down, and the quality of the outcomes went way down. And so this is not good. And so as part of the great rebirth of engineering culture at Etsy, one of their many goals was to kill Sprouter. Their goal was to try to make sure that every developer would be able to deploy value to customers by themselves with no dependence on a back-end team. And so they use things like object relational models to do that. And what they found was in every part of the Etsy property where they killed Sprouter, suddenly code deployment lead times went way down. And the quality outcomes went way up because there was no need to communicate or coordinate with anybody else. And so in my mind, this is one of the best examples of Conway's Law, which says that if you, to paraphrase, if you have a four-pass compiler, I'm sorry, if you have four teams working on a compiler, you will get a four-pass compiler.

And so consider how bad it got when you had two teams having to communicate and coordinate going to three, and how great it went as you went from three teams having to communicate going down to one. And so in most organizations, it's not just three teams having to communicate and coordinate. Consider a large complex deployment where you initiate a ticket to initiate a deployment on the left, and you may have 20, 40, 60 teams all having to work together to get a happy customer on the right. You might have middleware teams, environment teams. They have to be configured properly. User accounts. You might have to have scores of changes. You might have firewall rule changes, security reviews, change approval boards. It doesn't take a lot to go wrong before we are looking at code deployment lead times measured in weeks, months, or quarters. This is the exact opposite of the architecture that we want. What we found in the state of DevOps research, one of the biggest surprises for me, was that the top predictor performance was architecture. So it's measured by what? So to what extent can teams make large-scale changes to their parts of the system without permission from anyone else outside of their team?

To what extent can they do their work within the team without a lot of fine-grained communication and coordination with people outside of their team? To what extent can that team deploy and release their service on demand, independent of services it may depend upon? And I love this last one. To what extent can they do their own testing on demand without the use of a scarce integrated testing environment, which there are never enough of? They're never cleaned up, which actually jeopardizes the actual test objectives. And so if all those things are true, then you should be able to do normal deployments during normal business hours with negligible downtime. And by the way, when it comes to that integrated test environment, that is what couples us to the rest of the enterprise. And so just to explain why it was such a big surprise to me, when I was at Tripwire from 1997 to 2010, we were trained that it was always safe to ignore architects, especially chief architects, because everyone knew that architects, the only thing they did was come out of their ivory tower once a year and then publish a PowerPoint slide or a Visio diagram and email it to everybody. And then they would go back to their ivory tower, never to be seen for another year.

And so this was just a polite way of saying that architects did not impact how daily work was performed. And so if that were true, then it is certainly not true now, because what this finding shows is that there is nothing that impacts more the daily work of developers than architecture. So, the first ideal in the Phoenix project would be the bus factor. It's measured by how many people need to be hit by a bus or win the lottery and leave before the project, service, or entire organization is in grave jeopardy. And so in the Phoenix project, the bus factor was one. In other words, if Brent got hit by a bus, or won the lottery and left, then suddenly no outage could be fixed, no complex work could be performed, because all the knowledge was in Brent's head. So we don't want a small bus factor, we want a large bus factor. We don't want to be reliant on one individual, we would far rather be reliant on a team, or better yet, a team of teams. So we want a large bus factor. So in the unicorn project, the corresponding metric would be the lunch factor. As measured by, in order to get something important done, done, how many people do we need to take out the lunch?

Is it the Amazon ideal of the two pizza team where every team can independently develop, deploy a value to customers independent of everyone else? No larger can be fed than two pizzas? Or do we have to feed everyone in the building? So consider that large complex deployment I had showed you before. We may have a lunch factor of two, three hundred people and we may have to feed them for multiple days, right? Far too high driving up our lunch factor. Or consider if we are a development team and we have to implement a feature, but we have a dependency on 43 different teams. Now we have to take out 43 different people at the lunch, explain what we're trying to do, why it's important, what we need from them. And if any one of the 43 people say no, then suddenly we can't get done what needs to get done. So driving up our lunch factor. Lower is better. So in the ideal, anyone can implement what they need to by looking at one file, one module, one application, make all the needed changes there and be done. Not ideal is that in order to implement our changes, we have to change all the files, all the modules, all the applications because the functionality is smeared across that entire surface area, driving up our lunch factor.

It's not just about implementation. Ideal is also about testing. Ideal is that we can not only implement our changes in one place, but we can test them in one place. Ideally, isolated from every other component. So that is the notion of composability. And by the way, Docker Compose is all about isolating components from each other. Not ideal is that for us to get any assurance, that our component will work as designed in production, we have to test it in the presence of every other component, thus drawing us back into those scarce integrated test environments, thus coupling us to the entire enterprise, driving up our launch factors. It's not just about expertise and capabilities, it's about authority. In the ideal, to do what the customer wants, we can make all the decisions within our team. Not ideal is to do what the customer wants. We have to escalate everything up two levels, up two, over two, and then down two. Or visually depicted, we have to go up two, over two, and then down two.

So one of my favorite stories of coupling and cohesion is the book Team of Teams. And so this is a book written by General Stanley McChrystal, Chris Fussell, and David Silverman. And this is the amazing story of the Joint Special Forces Task Force, whose mission was to dismantle the enemy terrorist networks in Iraq in 2004. And what they found is initially they couldn't, despite a size advantage and a technology advantage, because they couldn't move and make decisions fast enough. And so this is an amazing story about how they pushed decision-making to the edges and allowed mid-level leaders to traverse and trade resources across this very vast network and was eventually able to achieve their mission of dismantling the terrorist networks. So one of my favorite books I've read in the last 10 years, and I think will resonate with everyone in the DevOps community. So that was the first. First ideal, locality and simplicity. The second ideal is about focus, flow, and joy that Jean Laurent mentioned. And so much of this was informed by me learning a functional programming language called Clojure.

That either runs on the JVM or gets transpiled into JavaScript or even can run on the CLR. And so this is one of the most technically difficult things I've ever had to learn in my career, but it was also one of the most rewarding. And so when I say it was rewarding, it even changed how I thought about myself. So I, for over 20 years, have self-identified primarily as an ops person. This is despite getting my graduate degree in compiler design back in 1993 at the University of Arizona. And yet... I always gravitated towards ops because it was my observation that it was where ops were. Ops was where all the saves are made. It was ops who saved us from terrible developers who didn't care about quality and pushed to production anyway. It was ops who was actually protecting our data and our applications because there certainly weren't those ineffective security people as depicted in the Phoenix project. And yet I've changed my mind. After I learned Clojure in 2016, I now self-identify not as an ops person. But as a developer. And I think the reason is that you can build so much with so little these days because of all the miracles that technology affords, thanks to things like containers and open source.

And it's so fun. And so the famous French philosopher, Claude Lévi-Strauss, he would say of certain tools, is it a good tool to think with? And I think there are so many things that come from functional programming that give us better tools to think with. And so these concepts came from programming languages, but these are such good tools to think with, is that we're finding them in infrastructure and operations as well. So things like immutability, things like composability. And so if you look at Docker, that is fundamentally all about immutability. That is the reason why if we want to change a container, we can't. If we want it to persist, we have to create a whole new one. So Kubernetes applies that not just to Kubernetes. the component level, it applies it at the system level. Whenever we are using things like Apache Kafka, someone's thinking about an immutable data model where we are not allowed to change the past. We can now do have immutability in things like databases, even if we certainly find them in something as fundamental as version control. The reason we would get yelled at when we rewrite the commit history is because we are attempting to rewrite the past, which is something that we shouldn't be able to do.

So in the ideal, when we are using these better tools to think with, our best time and energy is focused on solving the business problem and we're having fun. Not ideal is when we are using not good tools to think with is that all our time is spent solving problems that we don't even want to solve, like writing correct YAML files or trying to figure out how to escape spaces inside of file names inside of make files or writing bash scripts. And so maybe the most peculiar thing that has happened to me since learning Clojure in 2016 is that there are all these things I used to enjoy doing for 10, 20, maybe even 30 years that I now hate doing. I now hate doing anything outside of my application. I've become one of those developers. I hate connecting anything to anything else. I hate updating dependencies because everything breaks. I hate secrets management because I can never remember how to do it. And I'm the person who will occasionally check in secrets into the repositories, causing all sorts of problems. I hate authentication authorization. I hate data masking, bash files, patching YAML files.

I am the person who cannot figure out why my cloud costs are so high. And so this is not to diminish any of these things, especially when it comes to security, because we know those things are as important as the feature I'm trying to build. It's just not fun for me anymore. And this is why I'm so excited about development platforms. Because so many of these things we can now get in the platforms that developers use in their daily work, whether it's monitoring, deployment, environment creation, security scans, orchestration, database provisioning, all these things we can get self-service and on-demand. We don't have to open up a ticket and bug someone for weeks, begging them for them to do work on our behalf. Instead, we can get them with immediacy. And fast feedback, which are the conditions that allow us to have focus and flow in our work, which also allows the conditions that allow us to have joy in our work. And so flow has a lot of connotations in the lean community and DevOps, but the connotation I love most of all comes from Dr. Mihaly Csikszentmihalyi. So he is a a cognitive scientist.

He gave the best TED talk of all time called Flow, the Secret to Happiness. He wrote this amazing book called Flow, the Psychology of Optimal Experience. And he describes flow as the state of mind when we are engrossed in the work that we are doing, where we lose track of time, or we may even lose sense of self, that transcendental experience we have when we are immersed in work that challenges us, that we get gratification out of, that creates value for others. And so platforms allow us to have the sense of flow, which in my mind, experience increases a sense of productivity by orders of magnitude. And so before I leave this section, I want to share one metric that I think is so important for every technology leader, which is the code deployment lead time. In other words, how quickly can we go from code committed into version control through integration, through testing, through deployment, so that customers are getting value? And so you may ask, why do we start the lead climb clock at version control? Why don't we start it earlier when a feature goes into implementation in development or when an idea is first conceived?

In other words, that point of ideation. And the reason is that at the point of which change I put into version control, it is the dividing line between two very different parts of the value stream. Everything to the left of code commit is design and development. And so those are inherently highly experimental processes where the work takes a long time. It doesn't take minutes or hours. It may take weeks, months, or quarters. We may never have done it before, so we have no ability to estimate. So this is a period of incredible uncertainty. In fact, everything is about experimentation, whether it's A-B testing, whether it's testing whether we are taking users down a journey they want to go on. And that is the nature of design and development work. However, everything to the right of code committed into version control is the exact opposite. That is product build, test, and deploy, otherwise known as product delivery. There we want the exact opposite characteristics. We want builds, tests, and deploys to happen quickly, repeatedly. In fact, we want it to happen all the time after every commit into version control.

We want it to happen the exact same way every time, ideally entirely automated. So what is amazing about this metric is that code deployment lead time simultaneously predicts the effectiveness of product build, test, and deploy, but also predicts how quickly are we giving developers feedback on their work. So if I make a mistake and check into version control, do I find out in minutes or hours because of automated testing, or do I find out in nine months when I'm doing integration? testing or when I do a deployment into production and customers tell us. But it's not just about learning from mistakes, but it's also about learning from customers. So my favorite quote on this comes from Scott Cook, the founder of Intuit. And he said a decade ago that for the TurboTax property, so this is in the United States, how most people file the taxes. They did 165 production experiments during the peak three months of the tax filing season. And so when I first heard this, I thought this was crazy. Because, for example, in retailing, we are so afraid of the holiday outage that we have a change freeze from October 30th to January 15th.

And so why would these people be willing to make changes when it is the most dangerous? And the reason is revealed in the next paragraph. Is that the business result was that they were able to increase the conversion rate of their customer acquisition funnel by 50%. And so it was worth it, right? If you didn't acquire the customer this year, you could lose them to a competitor where you might not have another chance to acquire them back. And employees loved it because now their ideas are making it to market. And so we now live in an age where in order to win in the marketplace requires us to out-learn the competition. So it turns out there's one question you can ask that can predict every metric I've talked about, every technical practice and architectural practice and culture norm. And it's by asking this one very simple question. On a scale of 1 to 7, to what degree do we fear doing deployments? One is we have no fear at all. We just did one. Seven is we have existential fear of doing deployments, which is why if we could wave a magic wand, the next deployment we would do is never again.

So first ideal, locality and simplicity. Second ideal is focus, flow, and joy. And the third ideal is improvement of daily work. And so this shows up in the Phoenix project where we state that improvement of daily work may be even more important than daily work itself. So not ideal is twaddy. In other words, the way we've always done it. So obviously that is not so good. Ideal is MTBTT. In other words, make tomorrow better than today. So this is Google SRE principle number two. And in my mind, this is almost poetic. Who would not want to make tomorrow better than today? And when it comes to leaders, it is our job to make tomorrow better than today. And this gets to my next point, which is that greatness is never free. Greatness is a decision, and in our work, it so much comes down to what is our philosophy around technical debt. So I'm going to tell you a story of where technical debt comes from and what we must do about it. So imagine a point in time when we want to be first to market. Or imagine a point in time where we would be happy to be last to market because we're not even in the game yet.

We need to get in the game. And so under those conditions, this is when it's all about features. And we are willing to do whatever it takes to get those features to market. We are willing to build up technical debt, take more risks. knowing that this will drive down quality and increase the number of defects that make it into production. But the story does not end there. When you fast forward in time, you find invariably that the feature rate goes down, and the amount of time that we spend working on defects go up, maybe even above 100%. And these are exactly the conditions where defects dominate daily work. This is when site reliability tanks. This is when we go slower and slower, customers leave, morale plunges, and engineers leave because everything that was once easy is now difficult or maybe even impossible. So my friend John Cutler from the product community, he said exactly on Twitter, he said, case in point, in 2015, a certain reference feature would take 15 to 30 days. Three years later, the same class of feature now takes 10 times longer.

And this happens even though we're adding more and more developers. And so make no mistake, technical debt kills companies. So the second book I will recommend to you is this amazing book called Transforming Nokia by Risto Salasma. And some of you may be wondering, what can we learn from someone who oversaw the decimation of 95% of the market cap of Nokia? And my claim would be a lot. Risto Slosma was the founder of F-Secure, and when he was invited to join the board of Nokia in 2008, he thought this would be the pinnacle of his career, joining this dream team of a board. But it is this unflinching look about his own inadequacies as he is unable to deal with a very domineering board chair. But my favorite line in the book describes when he learned in 2010 from the VP of strategy at Nokia that the build times for the Symbian operating system would take two days. He said it felt like being hit in the head by a sledgehammer because he knew that if NG... Engineer would require two days to determine whether their change worked or would have to be redone, then all their hopes, dreams, and aspirations that resided on Symbian OS was an illusion.

It was a lie. And so that's actually what drove them to Windows Mobile, which arguably did not treat them so well either. But that was actually a far better bet than staying on Symbian OS. So arguably, Nokia didn't make it. However, this was the point that every company tech giant has faced, whether it's eBay, Microsoft, Google, Amazon, Twitter, LinkedIn, Etsy, all of them almost died because of technical debt. But they made a very different decision. And they all decided that they had to solve the technical debt issue in order for them to regain agility. And I think one of the best examples was Microsoft in 2002. So many of you will remember that 2002 was the summer of worms. This is when almost every Microsoft product was being mown down by security vulnerabilities. Code Red, SQL Slammer, NIMDA. And so the company was potentially going to go out of business. And this is what Bill Gates, then CEO, decided as well. And so he put almost every Microsoft product on a feature freeze for nearly a year.

He wrote this memo that he distributed inside of Microsoft and also sent outside the company describing the trustworthy computing initiative. But my favorite line in this letter is when he says, if a developer ever has to choose between working on a feature or fixing a security defect, fix the security defect because the survival of the company depends upon it. And so every Microsoft product stopped working on features, whether it was SQL Server, Windows, Office. And this is what arguably saved the company. So all those companies, Amazon, Google, Etsy, they all made the same decision. We stopped working on features so we can pay down technical debt, so we can increase quality and drive down defects, maybe not down to zero, but something that we can actually sustain over time. So we can reinvest in architecture and re-enable developer productivity, often by orders of magnitude. And this is indeed what has enabled every one of these companies to succeed. And it's not just CEOs, even though this is a CEO-level issue. Marty Kagan, who wrote this amazing book called Inspired, How to Create Products That Customers Love, he has trained generations of product owners on how to build great products.

And for nearly 20 years, he's trained every product owner to say, take 20% of all cycles off the table. They are not there for product owners to spend. They are there for engineers to use however they best see fit to fix problematic areas of code, to re-architect, to automate, to do whatever it takes to make sure that we can always safely ship features quickly and reliably. So where did Marty Kagan learn this lesson? At eBay in 2001. Because during the two years that he was at eBay, he did not ship one major feature because every engineer was just trying to keep the site up. So his lesson is, the only way that you can not get in that situation is pay technical debt down as you go. If you don't pay your 20% tax, you will inevitably pay 100% tax. So the third ideal, in the ideal, 3-5% of developers are dedicated to improving developer productivity. Google famously has over 1,500 of their best engineers working on improving developer productivity.

That's over a billion dollars annual spend. Microsoft likely has multiples more than that. Not ideal is that the only people working on automation, CI pipelines, and dev platforms are the summer interns and people not good enough to be real developers. So in the tech giants, it's the other way around. They put their best engineers working on dev productivity. So, such a Nadella in a town hall a year and a half ago said in a town hall meeting, if a developer ever has to choose between working on a feature or working on dev productivity, always work on dev productivity because this is how we can use compounding interest in our favor as opposed to technical debt, which is very much working against us. So this gets to the fourth ideal of psychological safety. And so in researching the Unicorn Project, it was so rewarding to revisit the work of Project Aristotle and Project Oxygen. So this is an amazing study. It is this quest that they had to understand what made the best teams the best.

And what they found after benchmarking over 60,000 respondents, over 250 teams across six years, was that the top factor was psychological safety. As measured by to what degree do members on a team feel safe to take risks and say what they really think without fear of feeling insecure, embarrassed, ridiculed, or maybe even punished. And so this is a higher factor than dependability, structure and clarity of work, meaning work, or even impact of work. And so my area of passion since 2014 has been studying not DevOps in the tech giants, the Facebooks, Amazon, Netflix, Googles, and Microsoft. Instead, it is in large, complex organizations that have been around for decades or even centuries. And since 2014, we've held 14. events where we've captured over a thousand case studies from technology leaders who are showing how they are using DevOps principles and patterns to win in the marketplace. And it is, I'm so delighted that we have leaders from almost every industry vertical. I will say that we don't have enough companies from France.

We had a leadership team from Orange SA present, but if you are leading an engineering organization from a large complex organization, we would live to hear from you. And so Jean-Laurent knows how to reach me. But we don't hear from just technology leaders. We also try to understand what is the top obstacles in the way. So one of my favorite presentations was in 2019, where we had heard that so many organizations struggling to implement DevOps because their auditors or compliance team wouldn't let them. So we had people from each one of the big four auditors present on how DevOps is not just possible to do, but in an audible, secure, and compliant way, but how they need to, because they want the customers to still be around in 10 years. One of my favorite presentations was Scott Havens. Then he was Director of Software Engineering at Walmart, the world's largest company, describing how he was able to re-architect the inventory management systems for the world's largest supply chain.

And reduce the API calls required to do item availability lookup from 23 deeply nested API calls, which thus forced 23 teams to have five nines availability and be highly reliable. redundant down to two so not what not just more reliable but also immensely cheaper to uh to deliver and safer as well another one of my favorite presentations was from american airlines where maya liebman uh evp and cio presented with ross clanton managing director and chief architect how they were delivering value in a cheaper faster and safer way but my favorite part of the presentation was how doug parker ceo of american airlines was asked whose job is it is it anyway to deliver value sooner safer and happier and he laughed he said it is not technology's job we have to hold every business leader accountable for delivering value more quickly regardless of whether that customer is internal or external Another one of my favorite presentations was Fannie Mae.

They are an organization with a balance sheet of $4 trillion. Kimberly Johnson, EVP and Chief Operating Officer, co-presented with the Chief Information Security Officer, who then described how they were creating paved roads so that every development team could get to production not only more quickly and reliably, but also more securely. So one of the best markers of culture is the Western organizational topology model, who I got to interview for four hours, who also presented at DevOps Enterprise. And what he found in his journey in reactor operations, healthcare, was that in the worst performing organizations, the top markers of the worst performers had these characteristics. Information was hidden. Responsibility and bridging between teams were discouraged. We cover failures and new ideas are crushed. Whereas in the highest performers, he found these characteristics. We seek information. Messengers are trained to tell bad news. We share responsibilities because we know that InfoSec is not just InfoSec's job. Just like uptime and availability aren't just ops'job, they are everybody's job.

And when failures happen, it causes a genuine sense of inquiry and new ideas are welcomed. So in the DevOps community, we love talking about blameless postmortems, where when something goes wrong, we truly seek to understand the timeline and the mental models that caused the actions to happen so we can better prevent, detect, and correct in the future. We love talking about things like chaos monkeys, where we deliberately inject faults into the production environment to see if we are as resilient as we think we are. But none of this is possible without psychological safety. Which gets us to the last ideal of customer focus. So this is one of my top professional aha moments in my career. And I want to share with you where I learned it. So it is 2019 in January. I'm in Detroit, Michigan, visiting the CEO of Copyware, Chris O'Malley, the famously resurgent vendor. And I had learned so much from him over the years. And I'm there with my friend, Dr. McKirsten, who wrote the book Project to Product, CEO of Tasktop. And we're walking to the campus. And I look down at the agenda and I see that the first thing on the agenda is a data center tour and I feel immediately embarrassed.

I turn to Mick and I say, I'm so sorry. I thought today would be so great. I don't know what we're going to learn by seeing someone's data centers, to see their Halon extinguishers, which I've seen too many of already in my career. But what we saw in the data center blew me away. Because what we saw was... An empty data center, about 45,000 square feet of empty data center space. So you'll see on the floor these green outlines, like a murder scene in a TV show, where the server racks used to be. In the middle of the outlines you will see a tombstone, a placard that says what business process used to run there and how much money did they save by either getting rid of it or moving it to a SaaS service. So you'll see things like EMEA financials after they unified it, on-prem email, desktop backup systems, password reset systems. You see 17 as a sign that says 17 tons of obsolete equipment removed and recycled sent to a better place. So in my mind, this is one of the best examples of moving from context to core.

So this comes from the amazing book Zone to Win by Dr. Geoffrey Moore. And so this is one of the best, that picture was one of the best examples of this, where they were able to move $8 million of context into core. So core is defined as the core competencies of the organization that create lasting, durable business advantage that our customers are willing to pay us money for. And so all those things, $8 million of stuff that was removed was context, often mission critical, but customers don't care about them. Customers are not willing to pay us extra money just because we have world-class payroll systems. So that last picture demonstrated $8 million of context that they were able to reinvest in R&D, which customers absolutely do value. So not ideal is that functional silo managers prioritize their silo goals over the most important business goals. Whereas ideal is that functional silo managers, and for that matter, everybody can look at it. the work they're doing in any given moment and ask unflinchingly, does this create lasting, durable business advantage that our customers value, that they're willing to pay us money for?

And if the answer is no, we should be able to unflinchingly ask, is this work something that we should be doing at all? Is this something that we should be moving to a vendor where it is their core competency and we are more than happy to pay us money to get access for that? So why do I think this is important? It is because the world is changing very fast. It is not big beating the small anymore. Instead, it is fast beating the slow. So the five ideals represent five important concepts that I think are so important to get us from here to there. And I hope this is something that you find relevant and useful in your own missions that help you win in the marketplace. So if you're interested in this presentation, Jean-Laurent and crew can provide that to you. But if you're interested in links to all of the DevOps Enterprise presentations and excerpts to everything I've written, links to the podcast that include an interview I did with Dr. Ron Westham, just send an email to realgenekim at sendyourslides.com with the subject line of DevOps, and you will get an automated response within a couple of minutes.

So with that, Jean-Laurent, I'd love to turn it back over to you. And again, it was so great to be in this forum of technology leaders. Thank you very much, Gene, again for today. It doesn't matter how great you believe your engineering organization is today. I'm sure there is a ton of tips that you can take out of the five ideals that Gene went through today. So again, Gene, thank you very much for spending the time. That's it. Thank you so much.