Tech.Rocks Summit 2020

Beyond the Spotify model: using Team Topologies for fast flow

Tech.Rocks Summit 2020 · 10 décembre 2020 · 26 min · en anglais

Résumé

Des systèmes logiciels modernes et connectés au cloud demandent d'organiser les équipes d'une certaine manière. En tenant compte de la loi de Conway, le talk propose d'aligner la structure des équipes sur l'architecture logicielle visée, en favorisant ou en limitant la communication et la collaboration pour obtenir les meilleurs résultats.

Summary

Effective, modern, cloud-connected software systems require teams to be organised in certain ways. Taking Conway's Law into account, the talk looks at matching team structures to the required software architecture, enabling or restricting communication and collaboration for the best outcomes.

Thèmes : Architecture & développement

Page du Tech.Rocks Summit 2020

Transcript complet

Transcription automatique, à relire : les noms propres peuvent être mal orthographiés.

Grâce à celui que nous allons accueillir maintenant, je ne verrai plus jamais l'application Spotify comme avant. Il s'agit de Matthew Skelton, qui est consultant en organisation et surtout qui est le co-auteur du livre à succès Team Topologies. Alors, mode confinement oblige, la conférence de Matthew va avoir lieu en mode distanciel, mais vous allez en savoir beaucoup plus sur ce que sont les Team Topologies. Merci Elsa. First, a little introduction in French. Bonjour à toutes et à tous. Je suis ravi d'être ici aujourd'hui. Merci Jérémy, Hervé, Dimitri et toute l'équipe de Tech.Rocks pour m'avoir invité à cette conférence. Malheureusement, je ne peux pas continuer en français car j'ai presque tout oublié. Alors, maintenant en anglais. So hello everyone, it's good to be here. Thank you to the team for inviting me to Tech Rocks. My name is Matthew Skelton. I'm the co-author of this book, Team Topologies. And I'd like to share with you today some ideas about why we need to move beyond this Spotify model when we're thinking about organization design for modern software delivery.

My background is as a software developer and an architect and increasingly involved in helping organizations shape their teams and practices for effective flow of change in software delivery. I'm the founder of Conflux in the UK. In this talk, we'll look at these four things. We'll look at what we mean by the Spotify model. What is that? We'll look at some limitations in that model. And then we'll explore some ideas in team topologies and how these overcome the limitations of Spotify model. Towards the end of the talk, we've got a few ideas about getting started. Is this what we mean by the Spotify model? He's a model, he's listening to Spotify. No, maybe it's this person here. She's a model, she's listening to Spotify. No, what we mean is the Spotify model of team design for software delivery. It became popularized and famous after an article written in 2012 by some consultants who were then working at Spotify.

And this diagram is quite famous. It describes the kind of four different kinds of groupings they've got in the organization, a squad, a tribe, a chapter, and a guild. Squad is a semi-autonomous delivery team with responsibility for a particular slice of the product or particular application. A tribe is like a family of squads working on related things. A chapter is kind of line management within a tribe. And a guild is a cross-tribe grouping with a particular interest. And these ideas have actually helped many different organizations around the world. So it's worth looking at where the value is in this existing kind of Spotify model, how this has helped some organizations. So it does actually help to encourage a flow of change. Because these semi-autonomous squads are responsible for building and running part of the software.

There's no handoff to another team. There's no handing over to another team to run it or to do testing or something. They own all of that kind of everything relating to that part of the software. So that helps to encourage a rapid flow of change. It also, the model also helps to establish and clarify team responsibilities. And that's because we've had to think about what is this squad responsible for? Exactly which application or set of services is the responsibility of this squad? And that actually forces us to think or encourages us to think about the kind of responsibilities at a software architecture level as well. So there's some use, there's some value in there. It also promotes good kinds of team collaboration. So not all collaboration is good from my perspective. I'll help to explain why a little bit later in the talk. But in the Spotify model, they talked about where they had groups of teams working on similar things would be sitting closely together in the office.

And groups and teams that were working on very different things would be sitting further apart. So they were making it easier for teams that needed to work together to be physically close. And it helped to plan and budget for cross-team enabling activities like talks and these guilds or communities of practice, this kind of thing. So there is some value in... in the Spotify model. The problem is there are big gaps. There are many things that it does not do. It was never designed to be a kind of model for software delivery. It was just an article that some consultants wrote back in 2012. But many people kind of adopted what they thought were the key messages in this model. They copied the names of the teams. They copied the kind of team kind of structure and things and thinking that would help them. And as I'm sure you were, lots of organizations have found that just copying the names of the teams and some of the obvious patterns that were described doesn't actually make their software delivery work any quicker.

So let's look at some limitations in this model. And the first warning really was actually written in the article by the authors, Knibber and Ivasen. They said, this article is only a snapshot of our current way of working, a journey in progress, not a journey completed. By the time you read this, things have already changed. So they put a warning in the original article saying basically don't go ahead and copy this. But of course, lots of organizations did copy it. So Spotify actually had to run a talk, had to give a talk at a conference. They gave several talks. One particular talk was at a conference called Spark the Change 2016, where they said there is no Spotify model. Please don't copy this. But of course, organizations are still copying this model. I discovered an organization, a large organization, in the UK where I live, had decided to change their entire kind of software development process based on the Spotify model. And that was just a couple of weeks ago. So this is still happening. So let's look at some gaps in this model, things that we do need to address that this original article simply doesn't talk about, for good reasons.

I mean, the article was not intending to describe a fully fledged, a full kind of software development approach at all. It was just an observation about how some things were happening at Spotify at the time. One of the things that is missing from the Spotify article is some clues about the size of software and the cognitive load on teams. So we'll look at this a little bit later. But what's increasingly clear these days is we have to consider the size of software that we're asking teams to build. If the software is too large, the teams will not be able to understand it properly and own it, and therefore we cannot get a fast flow of change. We also need to have some clues or what we call heuristics for Conway's Law. Conway's Law is this kind of socio-technical mirroring that affects When we're building a system, the communication pathways in the organization will tend to be reflected in the software architecture that we're building, or any kind of system architecture, really.

So we need some kind of clues that tell us how close we are to having this kind of mirroring in place and is the architecture that we're likely to build what we need? Do we need to change the shape of the architecture or the organization? We also need some patterns for team interactions in any kind of complex system like an organization. The behavior is usually a result of the interactions between different bits, not the behavior of individual parts itself. So we really need... to think about these kind of interactions between lots of different teams in the organization. And finally, I think we need to look at some signals that tell us it's now time to change and evolve our organization. Because particularly with modern organizations building software, we've got regulatory changes. We've got compliance changes. We've got trading relationships that are changing between different geographies in the world.

Things like the COVID-19 pandemic, changing how we need to kind of interact with our customers, suppliers, how we need to talk together with our own organization, and so on. So the pressures on modern organizations are increasing. To the point where an organization needs to be able to adapt and change many times a year to deal with these external and internal pressures and influences. So at a minimum, we need to address these four things. I think there's many additional things we need to address, but in terms of the talk that we're doing today, It's at least these four things. The size of software and the cognitive load, the results from the size of the software. We need to have some clues or heuristics for Conway's law. We need some patterns for team interactions. And we need to... Listen to signals that tell us it's time to change and evolve the organization. So team topologies has some patterns and ideas in there that help us to do some of these things.

Before we get into the detail, let's have a look at this word topology. It's not very common or very, very often used. It comes from a Greek word that means place, where something sits, where something exists. So in the context of team topologies, we can say this is about the way in which constituent teams are interrelated or arranged. So teams, how teams in the organization are interrelated or arranged. When we were preparing to write the Team Topologies book, we were doing research over the previous five years across lots of different industry sectors. We looked at the academic research across many of the different dimensions that we talk about in the book. So we looked at more than 50 peer-reviewed journal articles from academia and industrial research. We want to make sure that the ideas have a really strong foundation.

We ourselves, so my co-author and I, we did consulting work for many different organizations, over 30 organizations around the world, lots of different places, so UK, Europe, United States, China, India, and so on, many different kind of industry sectors as well. So we had some good first-hand experience of what it's like to work in particular teams in different situations and seeing the dynamics in different organizations. And in the book, we've got more than 12 case studies as well from different organizations doing different things. The origins of the book were actually in a blog post that I wrote in 2013, where I was exploring different kinds of ways in which development and operations could work together. So this then turned into a website called DevOps Topologies, which became quite well known and quite well used by lots of organizations around the world. So this is kind of an early version of some of the ideas that are now in the team topologies book.

But these patterns are quite basic, quite simple. And in the team topologies book, we've gone well beyond these early patterns. But the early patterns were still useful to some organizations. For example, Netflix used these early patterns. So Philip Fisher Ogden, who's a director of engineering at Netflix, said, thanks for your insightful articulations of DevOps topologies. They inspired many discussions and helps us to think about what model Netflix teams could be or are using. The patterns were also used by companies like Condé Nast, the big publishing house. Crystal Hirshhorn was until recently director of engineering at Condé Nast, and she said, your topological models resonated extremely well on both the dev and the ops sides. I like the balanced arguments, the different perspectives for each pattern. So these are just two examples of many organizations that have found those original DevOps topologies patterns really useful. So we were inspired to take the ideas much, much further and include many different dimensions to make it much more kind of comprehensive and really address the real needs of organizations and teams building modern software systems.

The book was published in September 2019 by IT Revolution Press. That's the home of books like Accelerate, DevOps Handbook, Phoenix Project, Project to Product. It's a whole great family of books. Basically, any book from IT Revolution Press is essential for your bookshelf. And for a period of time shortly after launch, Team Topology's book was actually first and fourth place in one of the Amazon rankings, which was great to see. And it's been selling extremely well in the last year or so. The book has been described as innovative tools and concepts for structuring the next generation digital operating model. That was a quotation from Charles Betts, who's principal analyst at Forrester Research. So these ideas are really kind of important strategically for modern organizations. I'm going to share four concepts or kind of ideas from Team Topologies now, and we'll see how these help to address some of the things that are missing in the Spotify model.

So team first thinking, Conway's law, team interactions, and sensing for evolution. So the first thing that we saw before, the first missing dimension that we saw before was the size of software and the cognitive load on teams. And we address this really by taking a team-first approach within team topologies. The team is the means of delivery. The team is the smallest unit in the organization. We do not think about individuals in terms of a unit of delivery or in terms of a means of delivery. The team is the thing that is doing the work. And there's good reasons for that because we can have high performing teams are incredibly more effective than just a group of individuals. We've seen this in plenty of research settings and so on that demonstrate how a group of people performing as a team, really performing as a team, is much more effective, much more high performing.

So this is our starting point. It also means that we need to design for team cognitive load. If our team is the means of delivery, then we need to make sure that the team does not have too much to deal with. If we get to a point where we exceed the cognitive load on the team, then things will slow down. The team will not be able to move as quickly. They'll not be able to have a nice fast flow of change. So that means that the size of software is now related to the cognitive load or cognitive capacity of a team. Some teams may have more experienced people, so that group can take on more complicated software potentially. But there's still a limit because there's a limit to the size of a team. A team has a limit of about eight people. Beyond that, the trust starts to reduce inside that team. Generally speaking, we want to have very high trust inside the team and therefore a very rapid flow of change because the team will be able to trust each other to do the right thing for the team and therefore we can have a nice quick flow of change.

This also means that when we're thinking about software architecture, we need to choose boundaries that help with team ownership. So we're looking for, it's a different set of principles for thinking about kind of architecture of systems. We're trying to optimize team ownership, we're trying to optimize for fast flow of change. And part of that is to make sure that the team can actually own the software that they're building. Because if it's too large or it's only a small part or the practices mean that teams feel like it's not their software, we don't get that fast flow of change. We also need to think about the physical and digital workspace. If we're in an office, we expect to have a kind of team shaped space in the office where that team can work. And feel like it's their part of the space. If we're working remotely or semi-remotely will have a digital space, which is something like Slack or Microsoft Teams or something like a chat room and online tooling and things. Again, how could the team feel like the digital workspace is their space to be able to work and collaborate and own the software systems that they're working on?

So this team first thinking is very familiar to some organizations, but it's quite new to a lot of organizations. It feels very different. So these are important things to bring out and to consider. I mentioned before that I think we need heuristics for Conway's law. As a common law is this property that typically if we're building systems, the architecture of the system that we're building will reflect the communication patterns, pathways in the organization, typically speaking. And so we need some clues for the kind of natural or expected design that we occur. If our organization pathways are kind of one particular shape, then we'd expect the architecture to be a similar shape. If the communication pathways are a different shape, we'd expect the architecture to be a different shape. And so this gives us a sense of, well, let's look at the communication pathways now. What's the architecture?

Are we fighting against this kind of natural tendency, this natural mirroring effect? And it's not really something that... We can really avoid, we can put lots of effort to try and avoid it, but it seems like more useful to kind of go with the flow, if you like, go with this tendency. And so what we can do is use something called reverse Conway to avoid the worst effects of this kind of socio-technical mirroring. If we want to have software in a particular shape, then we can change our organization communication pathways to mirror the kind of architecture we want. And there's a better chance that we're going to achieve this kind of software architecture that we think is needed from a business perspective. And fundamentally, Conway's law suggests that there's actually a constraint on the kind of solutions that we can find. So it's termed a constraint on the solution search space. And this means that if our organization organization is set up with particular kind of communication pathways, it means that we will never find

some architectures, we'll never find certain solutions because we don't have the right kind of internal connections and communications that are going to allow us to find the solution that could actually be really, really useful. So this is a strategic problem potentially for organizations building and relying on software systems as a key part of the viability of that organization. So there's a whole set of things we need to think about around Conway's Law. I mentioned before that we need to think about patterns for team interactions as a key aspect of how to think about and how to listen to what's happening in this complex system of our organization. And in the Team Topologies book, we define three interaction modes, three team interaction modes. And we think these three team interaction modes are the only three ways in which teams need to interact, should interact when building software. The first interaction mode is called collaboration. This is two teams working together for a defined period of time to achieve a specific outcome.

It's a very, very narrow definition of collaboration. It's very deliberate. Collaboration is expensive, it's costly, it increases cognitive load between the two teams, but it can help to deliver something new, it can help us to discover something, help us to find where a boundary should be, for example, help us to find where the right kind of responsibility boundaries should be for a new API or a new service or we're discovering something about the technology. Exactly where should the responsibilities lie? Collaboration can really help us to do that. But we should expect this to only be for a few days or a few weeks, not many months and certainly not years. Generally speaking, we're expecting to collaborate for a short period of time, and then we've now discovered where a boundary should be and we can have a kind of friction-free interaction across this boundary. When we get to that friction-free interaction, that's called X as a service.

One team is providing something and one team is consuming something. And there's very little need to talk to each other, really, because we've got a good boundary. And that interaction like that with low communication interaction could continue indefinitely. It could continue permanently in the future. If we found the right boundary, if we found a good boundary, and we'll come back to that later in the talk. The final team interaction mode that we define is called facilitating. This is where we've got one team helping another team to bridge a capability gap or to discover something, to work with a new technology perhaps or a new practice to move from one database to another database, something like this. Where one team doesn't quite have all the skills that are needed or capabilities, and another team of experts typically is going to help that team bridge that gap. So we've got these three different kind of team interaction modes.

And it helps, we've seen that this really helps organizations think about the kind of way in which different teams can work together to make these kind of interactions much more explicit and to choose one to decide, okay, this is the right interaction mode now because we've got this purpose. Or actually now we need to change our interaction mode because we're doing something different. In the book, we also define four fundamental team typologies, four different types of team. And we think these are the only kinds of team needed for modern software delivery. The most important is the streamlined team in yellow you see at the top there. This is a team that has end-to-end responsibility. For building and running a slice of the software system, a piece of the software system. It might be an application, a service, an entire website. The key thing there is they have all the skills needed and there is no handover from starting to build it to the software running in production.

That team looks after everything. So there's a nice fast flow of change because the size of the software that that team is looking after relates to the cognitive load of the team. So we've not exceeded the team cognitive load. They have all the skills needed. And they're not going to pass the work to someone else. So they own it all the way through, and then it runs in production, and they get feedback from production to help them improve that system. So that's a streamlined team. That's a kind of fundamental starting point. That's the starting point that enables us to have a rapid flow of change in the organization. Then there's three supporting types of teams. An enabling team, which is a team of experts who help streamline teams typically to bridge a capability gap, to shift from one practice to another or adopt something new. An enabling team does not own any part of the software system. They're just expertise, mentoring, and kind of teaching, if you like, teaching other teams how to do things.

As a complicated subsystem team, which is only used if we've got a part of the system which needs very detailed, very specialist knowledge of usually something around mathematics or kind of complicated processing or very, very complicated logic. We only have a complicated subsystem team if it's not reasonable to have those skills inside a streamlined team. So if it's very difficult to find those skills, for example, or if we don't want to have to hire many, many, many, many people with a particular specialist skill into several streamlined teams, it might be the right thing to have a complicated subsystem team. But generally, we try and avoid it because it can often block a flow of change. Sometimes it's the right thing to do. And we also have the concept of a platform team. And actually, generally speaking, a platform team is actually a platform group. And inside the platform, there's multiple other kinds of teams, including streamlined teams, complicated subsystem teams, enabling teams.

And other platforms, in fact. So the same patterns exist at multiple different levels of the organization, multiple different level kind of magnification levels in the organization. And we represent it like this. We represent the teams in the organization like this. There's always a flow of change from left to the right. We always see that in these diagrams. That's why the diagrams look very wide or long. We're explicitly representing a flow of change. So in this diagram here, we've got three streamlined teams. We've got one enabling team in purple on the right-hand side. We've got a complicated subsystem in red at the left-hand side, and the platform is represented at the bottom. At the moment, the top two streamlined teams are having some kind of facilitation from an enabling team. Enabling team is the expert helping them do something like move from one database to another. And the second and third streamlined teams are using a complicated subsystem.

So that shows the teams at this point in time. This is what they look like. It becomes important then to understand the kinds of interactions that these teams are having. So going from the left of the diagram, The complicated subsystem team is providing that complicated subsystem as a service to the second and third teams. So the second and third extreme aligned teams should have very little communication really with the complicated subsystem team because it's been provided as a service. We have a nice clean boundary, a nice clear API, and those teams can just use it. All three streamlined teams use aspects of the platform as a service. So again, there's very little need for lots of communication. We found nice boundaries and we simply use aspects of the platform at those APIs, self-service for the streamlined teams. Then you can see at the bottom right, the third streamlined team is actually collaborating with the platform.

That means they're working closely with the platform. If they were in the office, they're sitting next to each other. Engineers are sitting at the same computer. If we're working remotely, we're doing lots of screen sharing or working closely together to define where a new boundary should be, a new service should be, to define perhaps how a new metrics or logging service works. And we might be. be doing that for maybe two or three weeks or four weeks or something like that until we've discovered how this new thing should work. So diagrams like this from team topologies are always a snapshot in time, always like a photograph, a point in time. It's not the permanent representation of how the organization should be working. That's a really important point. And then at the top right of the diagram there, you can see that the enabling team has this facilitating interaction. So we're thinking about different kinds of teams, but also what kind of interaction is happening now with which team? So we've got different kinds of teams and different kinds of interactions at different times, depending on what the teams need to achieve.

So finally, the final thing that was missing in the Spotify model, from my point of view, is signals that tell us it's time to change and evolve. We call these triggers. figures for change and evolution. So crucially, this enables the organization to kind of sense, it has to listen to its environment. Crucially, not all teams in the organization look the same. This is kind of important. It allows us to sense things that are happening inside the organization and between the different teams. Typically speaking, we're expecting to collaborate. In other words, discover things, discover where the boundary should be or discover that the boundary is in the wrong place or discover we don't have the right skills or whatever. We're working, collaborating, discovering something, then potentially after we've discovered the right place for this capability, typically we would push it into a platform, make it available for multiple teams to use. Not always, sometimes we'll actually take something out of a platform, but often the pattern is we discover, find where the boundary should be, then make it available from the platform.

Now, crucially, when we have this combination of four types of team and three team interaction modes, it allows us to use awkwardness or difficulty in the interactions between teams as signals that tell us that something is wrong, that something is missing, a capability is missing, a boundary is in the wrong place or something else. And so instead of teams kind of fighting or becoming frustrated when working together, we can actually Listen to these signals and if there's frustration, if there's awkwardness, if so, let me give you an example. If we if if one team, if we two teams expect to have an X as a service relationship interaction. We think we should have a nice clear boundary here. But actually, if the consuming team is having to speak every day in detail to the team that's providing that service, then

that doesn't feel like a proper service relationship. We shouldn't need that amount of communication across that boundary. So something's wrong. Is the boundary in the wrong place? Maybe the boundary was defined 18 months ago in the past. Maybe it was perfect, a good place to have that boundary in the past. But now the technology has moved on and changed, and we need to shift the responsibilities of this thing. And that's why there's all this kind of communication happening. We can listen to these awkward signals and tell us it's time to evolve. It's time to change something. And crucially, it allows us to kind of evolve the organization with the changing ecosystem. So the business ecosystem inside the organization, but also the wider ecosystem around it. Partners, regulatory compliance, government, trade, this kind of thing allows us to evolve this organization because we're listening. For signals that tell us it's time to change, listening for these triggers that tell us it's time to change. So how can we get started with with a team topologies approach.

Starting point is to be explicit about cognitive load. So how well can a team as a unit understand the systems they own and develop? Just ask them, just ask them a question, literally, how well do you understand the systems that you work on, a scale of one to five? And if the team is saying, well, actually, we don't really understand these systems, there's too much to think about, that's a problem. That is working against a fast flow of change. If the cognitive load on the team is too high, it's working against a fast flow of change. So you can just run that as a survey internally. Consider if we need to push some things into a platform. If the cognitive load currently on the streamlined teams is too high. But also ask, are skills or capabilities missing? Maybe these streamlined teams need additional skills and capabilities inside the team. Maybe we need skills around Internet of Things, IoT, or we need skills around data processing or machine learning or something like this inside each one of these teams.

It's not always the right answer to push something into a platform. Sometimes we need to change the skills and capabilities in streamlined teams as well. We can also think, use, we use Conway's law and look at very big, what we call mismatches, very big differences between the communication pathways in the organization and the shape of the software architecture. If there's very big differences, then it's it's likely that we're kind of pushing against Conway's law. And we need to look for opportunities to kind of align the organizational communication and the kind of software architecture that we want. So think about things that could be easily changed and start there. So it might be just a small part of the system that could be easily changed first and get some familiarity with doing that reverse Conway maneuver. We can also think about the team interactions. So ask. Ask your organization or ask part of your organization, what would change if you adopted these three team interaction modes, collaboration, X as a service, and facilitating?

How would teams react and behave? What additional capabilities would that give you? What additional signals would that give you? And then finally, you can think about what we call thinnest viable platform. So how is your platform currently defined? Make sure you actually document it, write it down, describe what facilities the platform actually provides. And then ask, well, what's the very smallest, what's the thinnest platform that could actually work? Because often in the past, organizations have just built huge platforms, which are far too big and don't have the right kind of purpose, looking to have the very smallest platform possible with a strong focus on the developer experience or the user experience for engineers are using the platform. So these four things here are good places to get started. Explicit cognitive load, large Conway mismatches, team interactions, and thinnest viable platform. Let's quickly have a look back at what we've seen today.

So the Spotify model has helped some organizations in the past few years. It was never designed as really as a model, but some ideas in there were clearly useful. But there's quite a number of things that are really missing in there. The size of software and cognitive load on teams. Heuristics or clues for Conway's law, the patterns for team interactions, which are really important in a kind of complex system like an organization, and we need to have some triggers for change and evolution. We've seen these four things in this talk, team first thinking, we've looked at the implications of Conway's law, We've seen these team interaction modes and how we use them with the four different types of team. And then explored a little bit about how we kind of use these to sense and evolve our organization. So as I mentioned, the book is available from basically all good bookstores around the world. If you go to teamtopologies.com book, you will be able to get a link to your preferred bookseller of choice.

We've got a newsletter, so if you want to sign up, go to teamtopologies.com. We do have some fully remote friendly training. We've been running that since April 2020. And we've also We've got a whole load of free resources online as well. If you go to teamtopologies.com slash resources, you'll find lots of stuff there. We've also got some free templates, open source, Creative Commons templates, GitHub. If you just search for Team Topologies, you'll find those things there. We've got some resources like these, which are... We've got a PDF version, but we've also got a physical printed version of some modeling shapes. So you can actually use these pieces of paper or card on a desk, on a tabletop, and model your organization. It's actually the quickest way to model different kind of organization designs and approaches using these cards. So if you go to teamtopologies.com slash shop, you'll find those there. We'd love your feedback. You'll find us online at Team Topologies. You can send us an email, info at teamtopologies.com.

So thank you very much. It's been great to be part of Tech.Rocks, and I'm looking forward to the questions and answer session next. Thank you.