← BibliothèqueToutes les vidéos
Tech.Rocks Summit 2024
Faire preuve de résilience
- Neal Ford (Director / Software Architect, Thoughtworks, Inc)
- Rachel Laycock (CTO, Thoughtworks, Inc)
Tech.Rocks Summit 2024 · 3 décembre 2024 · 53 min · en anglais
Résumé
Comment les organisations peuvent-elles atteindre la résilience, c'est-à-dire la capacité à se transformer sans se briser sous le raz-de-marée du changement ? Neal Ford et Rachel Laycock montrent comment créer des systèmes résilients grâce à une architecture évolutive, et comment gérer des changements tectoniques comme l'IA générative dans l'écosystème du développement logiciel.
L’essentiel
Neal Ford et Rachel Laycock (Thoughtworks) abordent la résilience sous deux angles : l’architecture, en décomposant la résilience en caractéristiques mesurables, et les équipes, organisées selon la loi de Conway et Team Topologies, avec les pratiques d’ingénierie comme ciment.
Pour discuter de la manière de rendre un système et une organisation capables d’absorber les changements (réorganisations, acquisitions, IA générative) sans se briser.
Les idées clés
- Rendre la résilience mesurable. Selon Neal Ford, la résilience est une capacité composite qu’aucune mesure unique ne capture : il faut la décomposer en caractéristiques mesurables (élasticité, capacité de reprise, intégrité des données). Les « fitness functions », sortes de tests unitaires pour les capacités de l’architecture, permettent par exemple d’empêcher les dépendances cycliques qui dégradent peu à peu un code. à 2:17
- Concevoir les équipes pour concevoir l’architecture. Rachel Laycock rappelle la loi de Conway et recommande l’approche Team Topologies : des équipes petites et durables, qui se font confiance, un découpage du domaine qui réduit au minimum les communications entre équipes, et des objectifs mesurés au niveau de l’équipe plutôt qu’individuels. Elle décrit quatre types d’équipes : stream, plateforme, enabling et sous-système. à 25:51
- Les pratiques d’ingénierie tiennent l’ensemble. Intégration continue, livraison continue et TDD sont les « sensible defaults » de Thoughtworks ; un bon signe de résilience est de pouvoir passer un correctif urgent en production comme n’importe quel changement, sans hotfix. En réponse à la salle, Neal Ford estime que le plus important n’est pas qui lance une réorganisation mais qui la soutient quand la pression du calendrier arrive. à 36:27
Questions pour votre équipe
- Quelles caractéristiques de résilience de notre système pouvons-nous mesurer objectivement ?
- Notre découpage en équipes limite-t-il les communications entre elles, et qui soutiendra ce choix sous pression ?
- Quel est aujourd’hui notre chemin le plus rapide pour passer un correctif urgent en production ?
Il s’agit d’une keynote de deux intervenants de Thoughtworks, société de conseil : les exemples clients sont anonymes et sans chiffres de résultats, et les références citées (Technology Radar, livre sur les architectures évolutives) sont des publications de Thoughtworks ou de ses membres.
Chapitres
Summary
How can organisations achieve resilience, the ability to transform without breaking under a tidal wave of change? Neal Ford and Rachel Laycock show how to build resilient systems through evolutionary architecture, and how to handle tectonic shifts such as generative AI in the software development ecosystem.
Thèmes : Architecture & développement
Transcript complet
Transcription automatique, à relire : les noms propres peuvent être mal orthographiés.
Nous allons démarrer par notre premier talk de la journée 2. Nous allons accueillir deux géants de la technologie pour un talk digne des meilleurs thrillers numériques, Rachel Laycock et Neal Ford. Alors préparez-vous à plonger dans l'univers de la résilience avec Resilience, où ils nous dévoileront comment les organisations peuvent plier sans rompre, un peu comme nous avons pu le voir hier, face à une onde de choc du changement. Donc je vous demande de faire un triomphe d'applaudissements à Rachel Laycock, CTO de Thoughtworks, et de Neal Ford, directeur software architecte de Thoughtworks aussi. Welcome, Rachel. Welcome, Neal. The floor is yours. Thank you very much. Thank you very much. Enjoy. Welcome, everyone, to day two of the Tech Rocks Summit. And Rachel and I today are going to talk about resiliency, but we're going to talk about it from two different standpoints, two different perspectives on resilience and resiliency.
One of my favorite ways to start a talk like this is with a dictionary definition, since we're supposed to talk about resiliency. Here are a couple of definitions of resilience. The ability of a substance or object to spring back into shape or elasticity. And the second definition is the capacity to withstand or recover quickly from difficulties or toughness. And I'm going to look at each of these definitions individually, starting with the first one. And we are going to talk about this combination of factors between architecture and teams. And let's talk about architecture first. The first thing we have to talk about is what does resilience mean to a software architect? In fact, yesterday, Sam Newman did a keynote about this very topic, about resiliency in software architecture. What we really want to be able to do is to measure and see how resilient our system is. But there's a problem with that is how do you measure resilience at the software architecture level?
Well, you can't. And part of the reason you can't is because it's what we refer to as a composite capability in your architecture. So we think about structural design, we split that into behavior, which is the thing your system is doing, but the capabilities are the kinds of things it can support, and resilience is a capability that we want to be able to support. And we really like to be able to measure these things, but there is no single measure for resilience. And so what we do with composite characteristics like this is we break it apart. And see, is this made up of things that we can actually measure? It turns out it is. Resilience is composed of things like elasticity and recoverability and dataman system integrity. Once you get to the point where you can define objectively what your goal for resilience is, now you have a way of actually addressing that from an architectural standpoint.
Elasticity is a great example of this, measuring the burst of users that show up to a site all of a sudden. And in fact, this is part of our definition of resilience, as we saw earlier, the ability to be elastic. But I want to talk for a second about This seems like a made-up word. A nexus, of course, is the joining of things, but a nexus with a little hat on it is actually a combination of nexuses. So it's like a nexus of nexus. And I want to talk for a second about the many intersections of architecture that encompass resilience. Because it turns out software architecture is unique in that it intersects with all these different parts of our ecosystem. For example, software architecture intersects with things like implementation. Obviously, if we design an architecture and then groups of developers actually implement that architecture for us.
But software architecture also intersects with things like infrastructure. Particularly, when you look at architectures like microservices, the infrastructure component is a huge part of that architecture. When you look at things like engineering practices, we'll talk a little bit about this specifically later, things like how we use version control, which may not seem like it impacts architecture, but it turns out that it does, and in fact we can measure some of those impacts if we are clever about it. Team topologies. Rachel is going to spend a bit of time talking about teams and how that works and resilience and architecture of course has an impact on teams and we'll describe how in just a bit. One of the great revelations of things like microservices is how much data influences our decision in software architecture and vice versa. There's an obvious intersection between architecture and data topologies.
But there's also an intersection between system integration, not just at the application level, but the entire system level. And in fact, beyond just system integration, we have intersections with things like the enterprise, which is the entire organization. There's strategy as part of the enterprise and part of that strategy is your software architecture and the kinds of things that it can support if you want your enterprise to grow or to merge. We'll talk a little bit more about that in just a second as well. There's also an intersection of software architecture and the business environment. If we want to move faster to market, we can design architectures that allow us to do that more easily. So there's a tight relationship between those things. And of course, we can't have a keynote at a conference in 2024 without mentioning generative AI. And we'll mention it some more, but there's an obvious intersection here.
In fact, that intersection is quite interesting and we'll look a little more in detail of that toward the end of the talk today. All of these things intersect with architecture and make this sort of nexus around software architecture. And here's a great example of that. I'm the part-time host of the ThoughtWorks Technology Podcast, and we were lucky enough to have Dave Farley as one of our guests. Dave is well known for a couple of things. He co-authored the Continuous Delivery book, but he's also really well known for This concept of mechanical sympathy, he designed this architecture called LMAX that allowed six million transactions per second on a single Java thread by understanding the mechanics of how the system executes things. My podcast co-host Mike Mason asked Dave a really clever question that does mechanical sympathy mean anything in the cloud now?
Because do we really care about the details about how it executes in places like the cloud? And he made a brilliant observation that I'll share with you now. This is Dave Farley's idea. Why do we normalize data? Besides the fact that Peter Code told us to in what the 1950s or 60s. Why do we do that? Well, it turns out that since time began, CPU is cheap and storage is expensive. And so we use the CPU to break the data up so we can store it the most efficient way And then you use the CPU to put it back together. But in the cloud, the opposite is true. It's only when you move data around and manipulate it that you incur cost. Data at rest is essentially free. And so when you start thinking about architecture for the cloud, you can stop normalizing your data and instead duplicate it and shard it everywhere because it turns out that's more efficient and more resilient.
To changes because it's already there and you're not having to piece it all together. And that's a great example of the intersection of data topologies and architecture and infrastructure. Oh, and if you don't think your data team is going to be worried about the fact that your architects have decided to stop normalizing data, You don't live in the real world, and so that's going to affect teams as well. So decisions like this affect all the different parts of your ecosystem. That's what we're talking about with these intersections. And part of this idea of these intersections is that once we can concretely identify those intersections, we can start talking about this concept of architecture as code. And actually codify our design decisions into architecture. And that's really the realm of this book that I wrote several years ago with Rebecca Parsons, Pat Kwa, and Promod Sadalge, Evolutionary Architectures. This is really the mechanics of how you build resiliency in software architecture, particularly around this concept of guided change.
And this is the idea of an architectural fitness function, which is sort of like a unit test, but for architectural capabilities rather than behavior in your system. Here's a great example of an architectural fitness function. As an architect, what you want to try to avoid are cyclic dependencies between components. That's where one component talks to another. which talks to a third one, which talks back to the original one. This is an anti-pattern because I can't easily reuse one of those components without taking all of them with me. And if I let this keep spreading around the whole code base, before long my code base is going to look terrible. But how do you prevent that from happening? This is a tool called ArcUnit that allows you to write code that prevents these kind of cyclic dependencies from happening in your code base. This is a way to prevent the gradual decay of the internal quality of your code base by writing these kinds of checks.
So fitness functions are a variety of things. They might be monitors. They might be chaos engineering. They might be unit tests. They might be metrics. There might be lots of other mechanisms that we're only now seeing come to light. Yeah, and I can give an example where actually having that kind of fitness function would have helped an organization I worked with a few years ago, which was a very, very large financial organization that wanted to introduce the concept of continuous delivery. So this was actually pre-microservices, and we'll talk a bit more about where microservices came from. But ultimately, in order to try and get something on a pipeline and get it to be able to deploy faster and more consistently, at the time they had lots and lots of issues with testing and integration and everything just taking weeks and weeks. sometimes months and spending a weekend trying to deploy and then not being able to and having to roll back. And this, when we turned up, this had been going on for weeks and weeks and weeks.
And when we started to look at the code base, we couldn't run normal dependency graphs against the code base because it was just too complex and it wouldn't work. And it was also too cyclic. And that was part of the problem. They'd really created this huge, huge ball of mud. And when I went to talk to the architects about how we could redesign the system and told them how the system currently worked, I'm not kidding you, they said, we didn't design it like that. And I said, well, no way, nobody would. But they didn't have any testing in place that would ensure that that would not happen going forward. So this would have been really, really useful. And it's a really important tool to introduce it because cyclic dependencies can be so problematic, especially when you want to be able to move faster. Yeah, there's this concept. When I was a kid, I heard this as bit rot, but Wikipedia seems to think it's system rot. It's this idea that any piece of software that's old must be terrible inside because it's been slowly decaying inside, and it's exactly this kind of stuff that causes that decay.
Software doesn't actually rust. But it seems like it does if you're not careful. You can build all sorts of interesting fitness functions. This is one that checks structural alignment to see, to make sure layers are not talking to one another or allowed to talk to one another. So you can build some really sophisticated tests including one for mono repos versus repo per service. This is one of those age-old questions that does not have a correct answer because there are trade-offs on either side. But let's say you've chosen one. What's the big danger in a monorepo? It's cheating on dependencies. And so you can build a fitness function. This is the pseudocode for it. But you can translate this into actual code that checks to make sure that you don't actually cheat on those dependencies if you're using a monorepo. So this takes the engineering practice you want, but then builds a safety net around the way you use it to make sure that the things that would cause resilience to go down, which is coupling, you can prevent by writing fitness functions to do that.
So as we mentioned, this is a composite characteristic, but what we really want to do is learn how to decompose these things into objective measures. That's the whole game in software architecture because that leads you toward the idea of software engineering. And before we go further into that, Rachel is going to talk about the other aspect of resilience, which is about teams. Yeah, so these days it's very hard to talk about software architecture and even resilience without talking about teams. Does anyone want to shout out why? We've been talking about this for a long time. It's because of Conway's Law, which I'll... summarize as, you know, essentially the architecture of your system will look like the communication structures of your organization. And how this played out in the past, and I know organizations still have a lot of these systems, is that they would create specialized teams, right? You'd have a specialized front-end team, a specialized back-end team, database people, infrastructure.
And those specialized folks would keep the context of those specializations and those parts of the system. in their head or in their own sets of documentation. And what that created was essentially layered architectures. And in fact, we even used to design it like that. And we used to create these huge monoliths with huge layered architectures. There's a huge problem with that. Does anyone want to shout out one of the biggest problems with those kinds of layered systems, architectures, and specialized teams? You're a very quiet group this morning. It's not much of a shout-out kind of a crowd. No, it's not a shout-out crowd. I'm not in the U.S. Anymore, so I'm used to people shouting at me. See, now I've got to laugh. I was told the French don't laugh at jokes. So anyway, so one of the big problems is it reduces flow, right? It creates a lot of bottlenecks. And in my experience as a developer, I was a back-end developer. That's where I started my career. And when you're trying to get things into production and maybe you're under a lot of time pressure, you would finish your part.
You're like, I'm done, right? I'm sat around waiting for the DBAs to do their part, the testers to do their part, the infrastructure folks to deploy. And when there's time pressure, sitting around waiting doesn't look good. So what do you do? You just pick up the next thing. Well, then, of course, you get bugs and issues that come up in testing or in production or in the way that you've normalized your database. And you would then have to respond to those, which means you have task switching, which has its own issues. And so what we started to... to introduce, and I can't believe it's like 10 years, was microservices. The idea of breaking things down into components that allows you to be able to deploy them, but also introducing the idea of cross-functional teams, teams that could build it and run with the thing that they had. So the challenge, sorry, I went the wrong way. Coming back to resilience, like the big challenges this creates from a resilience perspective, and I just talked about deployment as one aspect, but once something's actually in production, there are always issues and there are always challenges.
And another way that teams were often organized around software and sometimes still are is this concept of projects, right? So they're very short-lived. We gather together all the requirements that need tackling. And then once we've deployed it, the team kind of moves on to the next thing, which means that there's not a lot of prioritization and maybe not a lot of funding into whatever technical debt or issues arise. In fact, they usually end up being patches, which is where you start to get this entropy in the system. So what we're really looking for in resilience, as Neal said, is the ability to recover quickly from difficulties. And the trouble is, is when we talk about that, is that difficulties are almost just an everyday occurrence when it comes to software. Because we aren't building, like we borrowed concepts with that approach. that I just described, which was classic waterfall, we borrowed concepts from manufacturing and from engineering, which are actually pretty stable parts of the world. As Neal said, you build a bridge, it rusts, but the terrain in which the bridge is on doesn't change.
And the requirements of the bridge don't constantly change. And there aren't new, you know, you don't have to keep moving the bridge to a different place to capture a new market opportunity, right? You have to build a new one. So the metaphor very quickly breaks down that we essentially borrowed from. But there's some really everyday... I guess, difficulties and challenges that I've certainly experienced over my 20-year career, and I'm certain that you could probably name even more. And the first one is this schedule pressure, which I talked about a little bit earlier, where when you're under a lot of schedule pressure, it doesn't look good. To be sitting and waiting, right? And it's not good if there's not flow through the system and the ability to move things through the system quickly. This is where you'll get hot fixes, which also introduce risk and challenges. Another very, very common difficulty which, you know, from a business perspective, isn't considered a difficulty, is this, like, new business direction, right? Maybe you want to enter a new market.
Maybe you want to acquire something. Maybe you just, you're... You want to get new features into production. But sometimes a new business direction isn't a choice. I don't know about you, but certainly the clients and certainly at ThoughtWorks over the last two years, we've had to make a lot of headcount reductions. It's been really, really tough in the industry. And that's essentially a direction that impacts the technical team as well. And when you've got specialized teams or individuals specifically owning pieces of software, that creates risk within the system if you have to let people go. Or in order to adapt to the reduction in headcount, often we go through the organizations, which then disrupts the team's ability to form, to storm, to norm, and to get to a performing place. So you're creating all this disruption in the system, and it's not necessarily by choice. Another very common difficulty, let's say the business is doing very well, and I have an example of this
that I worked with about a year ago of an organization that was booking, a travel and experience booking organization, and they were doing great and they wanted to go IPO, and as part of that they were basically swallowing up and acquiring lots and lots of smaller companies. What does that mean from a technical perspective and from the resilience of the system? Well, suddenly you've got lots of systems doing the same thing. Now this can happen through acquisitions, it can also happen through the constant big bang rewrites, which is another big challenge. But essentially now you've got more systems to take care of, more technical debt, and you've got to make some decisions about what you keep and what's good and what's not. Because sometimes when you acquire, actually the software that you've acquired is better than the software that you've got in-house. So what do you do about that? And then the last, I guess, difficulty that's very, very common is the ecosystem change. So as Neal said, you can't stand up on a stage and talk about, we talk about anything without mentioning Gen AI. That's a big shift in our ecosystem. I think there was a huge report by Goldman Sachs on Gen AI.
They were absolutely scathing about, you know, products that include Gen AI. or how many of them had gotten to production. And I've seen metrics like 4%, 7%, but it's definitely under 10%. But where I am seeing immediate impact is how we build software and how we understand software. And we've created a tool that helps you understand mainframes, as an example. Like 70% of the world's systems still run on mainframes. Talk about a resilience problem. Is anyone even learning that at university anymore? I certainly didn't. It was outdated when I was at university. But that's an example of an ecosystem change. But another one that we went through over the last 10 years is the move to cloud, as Neal said. And this promise of a reduction in the total cost of ownership of the software. But of course, if you just migrate things straight to the cloud, it turned out actually the costs went up. And so now organizations are redesigning the system to think about how they really take advantage of the cloud, or maybe because of new regulatory concerns or the cost, they're going back on-prem.
And these are the shifting sands that I'm talking about. So again, going back to that metaphor of leveraging manufacturing or engineering, you know, you don't build a bridge and then the terrain underneath you changes. But when you build software, that's exactly what happens. And in fact, abstractions are constantly being introduced where you might have built a feature in your system that now is just part of the software that you're leveraging. So you want to replace it because it's better to use something that's off the shelf, that's a commodity piece of functionality versus something that you build yourself. So these quote-unquote difficulties are actually just everyday occurrences almost within building software. So how do you handle that and how do you build resilience? And another example actually is we don't do software enablement anymore. We used to for years at ThoughtWorks, we would sit side by side with teams and help them understand our engineering practices, right? So whether it's test-driven development or pair programming or continuous delivery or continuous integration, we used to do a lot of that kind of work
to help them really leverage the practices and get to a place where they could build software repeatedly, faster, get it into production, and feel confident in the quality of that code. And it would all go great until one of those difficulties came along. Now, it would all go great except for the fact that it was twice as expensive, because now not only are you building software, you're also training people and helping them learn along the way. It got super, super expensive. And then as soon as one of these difficulties came along, People would just drop it. It was not sticky. So how do you, even when you recognize that you need to introduce better practices, like the fitness functions and stuff that Neal was talking about, to create more resilience in the team, but how do you prioritize that and make sure it's sticky? Well. One of the things that we're saying you shouldn't do anymore is silo for efficiency. And so what we did in the past is we deconstructed the systems across these specialist teams, which creates
issues, it creates slowness, it's not really leaning into flow, which again was another thing that really happened in our industry about 10 years ago, is this introduction of thinking about things from a flow perspective, and how do you get things into production quickly, effectively, and in a way that you can trust. And so we don't want to de- construct things anymore, we know that taking agile approaches across functional team where you've got the product owner, the front end, the back end, whoever you need is inside the team, that's great. But it created a new problem for us, which is, well, how do you scale that? And like anything in software, we kind of went down the the wrong approach before we went down the right. Now, I don't have all day to stand up here and tell you what I think about the Scaled Agile framework and SAFE, but I will say that when I've seen organizations implement it, what it starts to feel very much like is waterfall again, or a new version of waterfall, where we've gone leaning into the process and we start to lose some of the benefits of Agile along the way. So that didn't really work for us.
And a lot of organizations that we've worked with who've implemented it are now rolling it back. So what's the alternative? Well, I actually can't remember exactly when this happened, but what we recognized at ThoughtWorks is that Conway's Law is a law, right? It is what it is. You will have the architecture of your communication structures. And so in order to create scaled agile, you need to lean into that. And we called it the inverse Conway's Maneuver at ThoughtWorks, but somebody else came up with a better name a couple of years later, which is Team Topologies, which is really, it's a great book and I highly recommend it. And people have heard about the Spotify model. They essentially implemented that. And then people... tried to just copy the Spotify model for their organization, which didn't work. But what Team Topologies did is it took the core concepts out and said, okay, these are the concepts that you need to think about in your system in order to create the speed, efficiency, effectiveness, and safety that Spotify has.
And it's a really, really interesting approach. And actually, when I read the book myself, I was like, oh, these are the things that I've been doing with clients without having a name for it. And so now I have names for it, which is great. But so I'm going to talk a little bit about how Team Topologies helps you pull together the, lean into the problem of communication structures within your organization, basically looking at your architecture. And one of the traps that I fell into with this, by the way, is mixing up org structure with communication structure. Org structure is very different. Most org structures are very hierarchical. Even in a matrix structure, they're still fairly hierarchical. And if you think that your communication structure is the same as your org structure, think about this concept for a second. In many organizations, you have a CTO organization with the developers sitting under them, and you might have a chief product officer and an organization sitting under them.
Now, if we were to follow the idea that the communication structure is the same as the org structure, then in order to make any decisions, we'd just keep escalating. Right, which, I mean, escalation has its use in an organization, but you don't want to do it for every single decision. You would drive your CDO and your CPO insane with that. And if you were doing that on purpose, I would consider that weaponizing the org structure. The real communication structure is happening across and up and down across every level of the organization. And so what you need to do when you're thinking about the team design is take that into account and just assume that the org structure is just one representation. And in fact, I would say that similar to the architecture, once you've actually written it down, it's almost straight away out of date. I've been through several reorgs over the last few years and helped many clients with it. And almost as soon as you're done with the structure, you think, okay, I've got it. Everybody's accountabilities are clear. And then people try and do work. And like 50 questions appear. Well, who's responsible for this? Who's accountable for that? How does this get done? And so it's a constantly evolving thing, much like the architecture.
But what does team topologies do to help address that? Well, it takes a systems thinking approach. And it takes into account that ultimately what we're trying to do is improve flow. We're trying to get the software into production. And so it has a few core concepts, one of which is you need small teams. Cognitive load is a huge problem. It's actually probably part of the reason why we tended to put individuals in specialized role to own specific parts of the system. That has its own challenges of, you know, the bus factor or somebody winning the lottery or just getting another job and then suddenly you've lost that knowledge which also decreases the resilience in your system. The other thing, what you really want to do is have small, long-lived teams that really understand their part of the system. And then you want to have trust within those teams, because teams that trust each other are creative, they work well together, they get the job done. So that's a really important piece. The other thing is you need to rest.
communications across teams. So once you can see what those communications are, you need to break up the domain in a way that there is as little communication across teams as possible. Because the more communications across teams, the more risk, the slower things go. You do not, I've fallen into this trap myself of thinking, oh, the more communication, the better. But more communication is more cognitive load, more things to take into account, more edge cases, the more people you get in the room, the more ideas you get, which is great in the design and discovery phase, not great when you're trying to execute and get things done. You need to make decisions and move forward. And then the other thing that's really important when you introduce the team topologies concept is moving away from individual incentives. So nearly every time, before team topologies was a thing, when I was working with clients, I would nearly always end up in a room with HR. Because you're starting to mess with people's job descriptions, people's expectations, people's measures of success. But what you want is something that measures the whole team.
And so then they're all incentivized to work together and get the job done. So when you take a team topologies type approach, how does that help you with the resilience aspect, with things like schedule pressure? Well, a team that's focused on flow, a team that's... Cohesive and not coupled to many other teams, is able to get things into production And so when you have schedule pressure, you don't have to make crazy patches. You just get the job done. It also helps with new business direction. And even, in fact, with acquisitions. Because suddenly, if you've got a team that owns a specific domain, and this is one of the key concepts of team topologies, they really understand their specific domain, their complex domain or complicated domain. When you acquire a new piece of software or an organization that comes with new pieces of software, that specific domain team can own that domain in the new software as well.
So then they can make decisions because they're the experts on what you keep and what you get rid of and how you integrate. But I will say, when you make headcount reductions, that's going to impact no matter what, right? It's going to reduce trust from not necessarily within the team, but the team's trust of the organization. And that's not easy to come back from. And it's just one of the challenges that a lot of us are dealing with in the last few years. And, you know, there's only so much you can do about that. So that will impact resilience. And so just a quick last thing I want to say about team topologies is there's the structure of it is there's basically four core team types. The most important one, or the main one, I guess, is stream teams, right? These domain teams, these people that really understand their domain and they have a cross-functional team that delivers to that. They moved away from the idea of calling them products and feature teams, one, because that team might be building a service, which isn't a product, maybe it's a feature.
But they also might be domain-specific stream teams within a platform. Right? And that's actually one of the other teams that you'll see emerge is this idea of a platform team. Now, it's very different. It's not the same as a DevOps team, which is actually more of an enabling team, which I'll get to in a minute, and a short-lived team. It's a team that's building out things that you do all the time, right? You deploy to production. If you're in a highly regulated environment, you need to take into account security. You need to take into account compliance. You just want that built in, and everybody can just absorb that. But when you think about a platform team, you have to take it again, this very product-centric approach so that people want to use the platform. You need to know what the platform's roadmap is. If people want to be able to contribute back to the platform, you need to have an inner source type of approach. So that was the second team. The third team. Is the enabling team. And this is actually the short-lived team. And this is the team that I've essentially been on as a consultant over and over again, right, when we were doing enabling organizations to introduce continuous delivery, enabling organizations to do microservices.
And those tend to be short-lived teams. And one very large media company that I worked with in New York, we did this for about a year. They had, I think, 20 to 30 teams. I got around about 10 of them over that year. So you basically spent a little bit of time with each team. help them use the platform, help them understand the ways of working, and then you really want them to own it and take it forward. And then we help them introduce their own enabling teams inside the organization so that as new technologies and new practices came in, that team can help introduce it to the whole organization instead of 20 teams trying to, you know, everybody trying to learn how to do, you know, Gen AI, for example, or how to do... How to introduce new tools into the system, how to review tools into the system. So having an enabling team that takes that time to do the research, do some work, work with some of the stream teams, and then roll it out is also an important concept. And then the last one is this subsystem team.
So this is the one that's very rare or may not exist at all, where you have a very, very specialist skill set. And I think a mainframe is probably one of the ones that is very common in organizations because they do tend to be specialized. There's not that many of them. There's not that many people that understand the system. And so that's an example where you might have a very specialist skill set, but you want to avoid having too many of those because of all the issues and risks that it creates that I talked about earlier. So the thing that brings it all together, right? So now you've got, you're redesigning your system, and the way to redesign your system is to redesign your teams. Because one of the pieces of research that came out since we've been talking about Conway's Law the last 10 years is that if your architecture and your org communication structures are at odds with each other, guess which wins? Org structure, every time. Org communication structure, every time, is the thing that will win. So if you want to design an architecture that's resilient, that's highly cohesive, that's decoupled, then that's actually how you have to design the teams.
But the things that bring it all together is the engineering. And at ThoughtWorks, we call these sensible defaults because people are always asking us to publish the ThoughtWorks methodology as we were going around with clients and doing things like team topologies before it had a name. What's the ThoughtWorks Agile? What's the ThoughtWorks whatever? We were like, there isn't a set methodology. This is kind of where the safe framework falls down. There are a set of practices that we see by using over and over again creates flow. Gives you the ability to get things into production safely and sense what's happening in the market and make changes and create that flow constantly in the system. And examples of those types of practices is continuous integration, it is continuous delivery, it's test-driven development, all of those things I've mentioned so far. So, when we're thinking about resilience and about issues that might come up into production, when we think about how to organize our teams and the things that Neal talked about around architecture
and engineering practices, when we should be able to answer the question of, like, what's the quickest way to get an emergency fix into production, when we're doing this well, One of the measures you'll have is that you can just check it into version control, just like any other change. If you can do that, you don't have to hotfix, you don't have to patch, then you've built resilience into your system. And engineering practices is the glue that really holds those two things together. And again, if you want to incorporate a new trend like AI into your organization, it comes back to these good engineering practices. And of course, going back to team topologies, the enabling teams we talked about earlier. Neal, you want to talk about the radar? Yes. So I'll wrap up and talk a little bit about the Tall Works Technology radar. This is a publication we put out twice a year, and we just recently went through one of these exercises. And I wanted to point out this intersection between Gen AI and architecture in particular.
You'll notice that, so a little bit of terminology here, assess means that we're evaluating it. Trial means that we've used it in production. Adopt means that it's a sort of a default, makes really good sense. And whole means you shouldn't start anything new in it. That's our blips. And then the gray ones mean that they were purposely kicked off by the proposer. These are curated technologies from our project teams, and we get together as a group and decide which ones get published on this radar. And our last radar, 38% of the blips had to do with AI. And this is a visualization our colleague Birgitta made that highlights all of these AI-related blips. And you can see the blue are all the ones that made it on trial in our radar. The more important thing for us as we talk about architecture and AI. It's really the intersection of LLMs are big giant black boxes to us. And so what's more interesting from our standpoint are all the categories of ways this may impact our ability to do software development.
These are all the categories that Birgitta identified. And you can pull some of these out as being very pertinent, like RAG, which is one of the ways that we customize the behavior of a large language model. And you see there are a bunch of tools here that are related to RAG. But we also have vector data. databases and you see a number of those are in trial. Evals and guardrails are the ways that we validate what our AI is producing for us. Structured outputs allow us to take the output and do useful things with it. And then finally, observability, of course, we care about that as architects. But then agentic AIs are ways that we can take behavior or the output and turn it into behavior. This is where the interesting thing for us in architecture and AI is, is how do you incorporate these things into software development and make it effective? Well, Rachel's already given you the answer to that. It's good engineering practices. And this is the secret to all of these new things and the secret to architectural resilience and institutional resilience is having good engineering practices that allow you to support
brand new things like AI and other disruptive technologies that may come along. So resilience requires holistic relationship between technology and teams, sound engineering practices, and organizational support for resilience. That's our time. Thank you very much, and hope you enjoyed it. Thank you. Thank you. This is the encore, we're on the stage. Q&A. We forgot there was Q&A. Like in a vaudeville, you know? Yes. There is people inside, people outside. And I was just waiting for you. I think it was a fantastic show for you because everybody applauded.
So congrats. Let's start for our Q&A session. If you don't mind, I'm sure that many people want to ask you many things. So who is going to be the first for that? Okay, so like yesterday, thank you so much to give the mic. Would you please stand up because it's for the camera because you're going to be a star. It's the César ceremony, don't forget. Yes. Thanks a lot for your presentation. In terms of changing team topologies, which part of the company should be the motor? The management, the CTO, the HR? Who is the most... Who is the best person to do it, to launch the change in topology and by the way in architecture? Well, it's a great question. Because every org has slightly different structures, right? So some of them have a CIO, some of them have a CTO, some of them have a chief product officer, some of them are joint CTO and product officer, some of them have a chief digital officer.
Honestly, to me, it could be any of those leaders. But if you are in a place where different teams are reporting to different leaders, like up to the product organization or up to the technology organization, you actually need them both to sign up to it and introduce it because there's so much crossover there. And similarly, like I've seen organizations where architecture might report to a different organization or infrastructure, and the database organization might report to a different org structure. That's fine. I know people constantly keep changing these things and shifting these roles. It doesn't really matter as long as the leaders that need to support it are actually on board. If they don't, that's part of the challenge. So that's sometimes why organizations will consolidate all under one leader and have that leader drive the change. I see that a lot because it removes that friction. Because what I'm talking about is when you don't have that, when you need to get three, four, five different leaders to agree to introduce this big wholesale change, that just creates more friction.
So there's not like a one answer to that. I think it's just like if either or any could, but they need to get... their partners and peers to buy into it as well. So in the great tradition of these Q&A sessions, I'm going to take your question and actually answer a different question, but the question that I think you should have asked, which is not who starts the thing, but who supports it when the going gets rough? Because starting new initiatives is easy, but keeping it in place when schedule pressure shows up. This is, Rachel and I both have experience of working on these enablements, and you get all this beautiful new structure in place, and everything's humming and it's working, and then one kilogram of schedule pressure shows up, and everything just snaps back to the old habits. It's really depressing because you've spent all this time and effort just to see it all melt away. And that's exactly team reorganization. It's a great initiative, but the real test of it is when schedule pressure shows up, do you stick by it or do you snap back to what it was?
So I think whoever supports it is in many ways more important than whoever starts it. Which is why it's easier if it rolls into one leader. Yep. Okay, do we have another question? Amélie, yes, thank you so much. Thank you for coming. I'm great fun. You mentioned Conway's Law. For my take, Conway's Law is a bit incorrect, in fact. Yeah, I want to hear your answer on that. You say, okay, Conway's Law mirrors our architecture, our communication mirrors the organization, but in fact it's worse. It mirrors the whole history of an organization. Not the one that we see today. And often we miss that part. It would be almost easy if it was a real mirror.
And why don't we often forget this historical mess that we have to face? I don't know if we forget it or we choose to forget it. The legacy is most of the system. Most of the work that we do is brownfield work. You know, there's the old adage, you read code more than you write code. So I think it's just a tendency to want to build things new. It's almost like human psychology of like, well, we did it wrong then. This time, this time we're going to get it right. Trauma response. Yeah, a bit of a trauma response. But it's really important not to forget it, right? Because when you're reorganizing, let's say you reorganize leveraging the team topologies method, those domain experts, those stream teams need to hold the legacy stuff as well as the new stuff. Because if they don't, they don't hold the whole domain, right? They're missing parts of it. And maybe that would create too big of a team. So you might have two small teams that then communicate really strongly with each other.
But you do need to take into account both. Because you're right. I don't know if that was Conway's law actually is incorrect or it just shows like the path of where things were. But yeah, it's definitely worse. And I think this is actually why I think what's happening with Gen AI, one of the interesting things about it is the ability to understand code. So I predict that in a few years, all the IDEs will just be able to understand any code base, because that's half the problem is like, do you understand it? Do you know what it's doing? And then it gives you the ability to make decisions about whether you want to rewrite or re-architect any of those pieces. But you cannot forget about the old stuff, that's for sure. Well, but, and there's a great example of this that I see all the time where you're accidentally driving the wrong incentive. And one of the classic plays. are database schemas in a large organization. Because you look at that and think, oh, well, that's some relationship between the entities, but it's not. It's every relationship that's ever been laid on top of each other. Because DBAs were taught many years ago, never change anything.
Because you might break something downstream that's coupled to that, even though it shouldn't be, but you get blamed if it breaks, and so never change anything. Just bolt more joins onto it to add to it, and before long, it's incomprehensible. Go look at your enterprise database schema and ask yourself this question. If we started from scratch, would it look anything like this? And the answer is almost always no, but that's exactly what you're talking about is the fossilized remains of every solution that we've ever had sort of layered on top of each other. Which is why it's important to keep on top of that both code and databases and refactor and shepherd those things over time. They don't have to turn into a big ugly mess, but they often do. We have 200 left. Any more questions? Another one. Quatrième rang, juste là. Thank you so much. Thanks for your talk. It's always interesting to have an explanation on what we see every day.
I have a question regarding the good engineering practices that you highlighted at the end of the presentation. How do you manage no code and low code with that? Because I really see good engineering practices on that. And how do you do for resilience on that? Low code, so this is something I wrote about in the Building of Lucia Architecture's book. Low code, no code is a zombie that comes out of the grave every five or six years, and all the villagers have to get together with torches and pitchforks and beat it back down into submission. It's always promised as one of the greatest things ever, but if it was, that's all anybody would be writing code in. The problem with those environments is that you, not that you shouldn't use them, but you should apply what we call the last 10% trap. And we learned this many years ago. I worked for a consulting company that built in a variety of platforms, and one of them was Microsoft Access. And we realized that every Access project started as a success and ended in total failure.
And we wanted to understand why that only happened to Access projects. And we codified the last 10% trap. First 80% of what you want to get done in Access is super fast and easy, and it's amazing. The next 10% you want is possible but difficult, because now you're trying to get it to do something it wasn't really designed to do. The problem is that last 10% is impossible, and users always want 100% of what they want, which is why we don't use, remember the 4GLs? Those of you who have gray hair in the audience, remember the fourth generation languages, PowerBuilder, DBase, Clipper, FoxPro? Those all faded away because they all suffered from the last 10% trap. And that's exactly what low-code, no-code does, is gives you a really fast path to almost build a complete piece of software, but then you can't get all the way there. So the way to evaluate low-code, no-code is not just completely discount it, but treat it pragmatically. So when you pick a new platform, the first thing you build typically is Hello World. But if you can't get Hello World to work in low-code, no-code, run away.
What you want to find is where is it going to break? What are the edges of its capabilities? And can I live with that? in this organization because the trade-off is much faster development, but I'm going to hit this wall. Is it okay if I stop at that wall or do I need some fundamental capability it can't provide? That's the proper way to think about that new category of tools that come along. It's not magic. It's just software. And I will add that many of our clients are now rolling it back because it's so problematic. And I think we're going to fall into a similar trap with Gen AI and people just generating code. It can generate code, but it can't architect a good system because architecture is more of an art than a science, in spite of our best efforts to try and turn it into a science. Because when you start talking about team design, you start talking about humans. And humans are weird and wonderful and special and odd. And you need to think about that when you're designing your system. And so, you know, we can't just generate things.
Like, I think we can generate things that are repeatable and that we do the same over and over. But anything that's got any kind of complex domain logic requires a human. I think that's a great takeaway from our keynote. Humans are weird. Humans are weird. Let's take that as a conclusion. I think it's perfect. Thank you so much, Neal and Rachel. We are really proud. Thank you. Proud to have you. Thank you so much. Thank you.
