← BibliothèqueToutes les vidéos
Tech.Rocks Summit 2024
Stratégies clés pour construire des plateformes de développement interne robustes
- Daniel Bryant (Platform Engineer / Product Marketer, Syntasso)
Tech.Rocks Summit 2024 · 2 décembre 2024 · 36 min · en anglais
Résumé
Dans un paysage technique en évolution rapide, la résilience des plateformes de développement interne (IDP) est primordiale. À partir d'une série d'études de cas, Daniel Bryant montre comment concevoir et maintenir des plateformes qui permettent aux équipes d'aller plus vite, garantissent l'application des politiques et des workflows, et passent d'un cluster à plusieurs. Parmi les exemples : une place de marché biface qui migre vers une plateforme Mesos pour tenir le Black Friday, et une banque britannique qui passe d'approbations par plusieurs comités à une conformité et une gouvernance automatisées sur Kubernetes.
L’essentiel
Daniel Bryant (Syntasso) défend l’idée qu’une plateforme interne bien conçue rend la livraison logicielle plus résiliente, et l’illustre par trois cas : un éditeur britannique de logiciels de gestion, la place de marché Not on the High Street et la banque NatWest.
Pour concevoir ou relancer une plateforme interne de développement et discuter de la collaboration entre équipes plateforme, développement, QA et ops.
Les idées clés
- La plateforme est un produit interne dont les clients sont les développeurs. Des organisations qui avaient bâti un « AWS interne » sans parler aux développeurs ne voyaient personne l’utiliser. Chez l’éditeur de logiciels, trois personnes pour des centaines de développeurs ont proposé un « golden path » via un portail, avec la possibilité d’en sortir en assurant soi-même l’astreinte. à 7:29
- Réunir dev, QA et ops autour d’objectifs communs. Chez Not on the High Street, le site tombait chaque Black Friday ; l’équipe ops disposait des logs Apache de l’année précédente, qui ont été nettoyés, rejoués et amplifiés contre le nouveau site. Les KPI individuels des équipes ont laissé place à des KPI déclinés à partir d’un objectif commun. à 20:36
- La résilience ne se décrète pas. Chez NatWest, obtenir un environnement prenait 11 mois ; la banque a créé une équipe « platform as a product » et une équipe chargée de faciliter les contributions, en s’appuyant sur l’inner source pour ouvrir la plateforme aux autres équipes. à 23:37
Questions pour votre équipe
- Quand avons-nous demandé pour la dernière fois aux développeurs ce qu’ils attendent de notre plateforme ?
- Quelles données réelles (logs, trafic passé) pourraient rendre nos tests de charge plus réalistes ?
- Comment d’autres équipes pourraient-elles contribuer à notre plateforme ?
Talk de 25 minutes, qui reste à un niveau général. L’intervenant travaille chez Syntasso, qui développe Kratix, framework open source de plateforme utilisé dans le premier cas ; il a travaillé ou travaille encore sur les trois cas (comme consultant pour Not on the High Street). Le premier cas est anonyme et les résultats sont peu chiffrés.
Chapitres
Summary
In today's fast-moving technical landscape, the resilience of internal developer platforms (IDPs) is paramount. Drawing on a series of case studies, Daniel Bryant shows how to design and maintain platforms that help teams move faster, enforce policies and workflows, and scale from one cluster to many. Examples include a two-sided marketplace that built and migrated to a Mesos-based platform to run reliably through Black Friday sales, and a UK bank that moved from multi-board change approvals to automated compliance and governance on a Kubernetes-based platform.
Thèmes : Cloud, infra & ops
Transcript complet
Transcription automatique, à relire : les noms propres peuvent être mal orthographiés.
Alors, ça va ? Ça va ou pas? Come on. Si la résilience devenait l'ossature même des plateformes de développement. C'est Daniel Bryant, Product Manager de Syntasso, qui nous invite à découvrir comment, tel un dramaturge, il construit des infrastructures capables de braver les tempêtes numériques. C'est à travers des récits inspirants, Une place de marché affrontant les furies du Black Friday, c'était avant-hier, ou une banque automatisant la gouvernance, il révèle les clés pour faire de la résilience la trame maîtresse de nos systèmes. Il n'a rien à voir avec Suzette, même si son homologue français s'appelle Danny Briand. Je ne pouvais pas résister à la faire, excusez-moi. Voilà, donc accueillons sous vos applaudissements nourris Daniel Bryant in English, sir. Acte 1, scène 2. Hello, Daniel.
Nice to meet you. Nice to meet you, too. The floor is yours. Thank you very much. I'll grab the microphone here. You have the remote. Super. I'll say bonjour, but that will be the limit of my French. Apologies. And I will try and talk slowly today as well, because one of the things folks often say is I do talk quite fast. But it's great to be here. Right, let's get started. So I'm going to talk to you today about building resilient platforms. Now, everyone I'm talking to at the moment, and granted, I go to a lot of platform conferences, but everyone I'm talking to is building platforms. Whether they sort of know it or not, they are constructing a platform to deliver software. So I'm going to talk today about building in resilience to the foundations of that platform, both the people side. And also the technology side as well. Both are vitally important. Hopefully, we all know that. The socio-technical systems, and I'm going to talk about building resilience into these systems. I like to start my talks with a high-level TLDR, kind of key takeaways, if you will.
My pitch today is a well-designed and implemented platform enables resilience throughout the software delivery lifecycle. I'm going to say simplifying platform interactions reduces errors. As much as I love DevOps, one of the things I've seen is DevOps is a you build it, you run it mentality. And at scale, you have very different build-its and very different run-its. Very hard to be resilient when you've got multiple ways of doing things across a very large org. There we go. Resilience testing on the first slide. Always a good look. I'm going to argue that graceful degradation supports business goals. Now, I'm going to briefly touch on policy, but today, most importantly, I'm going to be looking at fault tolerance. It's really important to think about, you know, clearly like Black Friday just gone. I'll mention that. It's really important to design for scale. To test the scale, and also when stuff falls over, inevitably, make sure it's resilient about coming back up.
I'm going to argue that shared responsibility promotes resilience. I think throughout my 20-year software development career, what I've seen is inherently people like handing stuff off. They like going, well, I've done my bit. It's your bit now. Dev, ops, kind of classic, right? But there's many examples of this. But I'm going to argue that we really need to collaborate and share the responsibility across the platform and across the software architecture too. And I'm going to flash this meme up on stage now. I'm going to break it down. It's XKCD, and I've modified it very slightly. You can spot where I've modified it. But what I want you to take away from this meme at the moment is if you're not careful, one small little bit of your platform that's not very strong can take down the whole thing. Very briefly, this is me, at Daniel Bryant UK on most of the interwebs. I'm on Blue Sky more than I am Twitter, X these days. I started my career about 20 years ago as a developer, primarily in Java, a bit of Go and JavaScript and Ruby along the way as well.
And I moved into QA roles and then ultimately architecture roles and then built a bunch of platforms on Mesos, Cloud Foundry, Heroku, and of course Kubernetes of late, right? I've done leadership roles, I've done tech lead, I've done CTO roles in a couple of gigs as well. I love learning and I love teaching, hence thank you very much for the invite today. I've written a few books for friends along the way, and I've worked on a whole bunch of open source projects like the OpenJDK back in the day, the open source Java, and more recently, Cratics, which is an open source framework for building platforms. So, let's define the what of platforms. Platforms is some of the... an overloaded term, and it means many things. When I talk about platforms, I like this definition from Evan Botcher. And I'll pull out a few things here. So digital platform is a foundation of self-service APIs, tools, services, knowledge, support, which are arranged in a compelling internal product.
Autonomous delivery teams can make use of the platform of the product to deliver product features at a higher pace with reduced coordination. There's a lot of things to like here. Excuse the slightly strange jump there, but a lot of things like, I like the notion of self-service. We as developers just want to get our job done, right? We want to ship value to end users. I like the focus on tools and on knowledge and support. Socio-technical systems, right? We're doing people and processes and technology. And I like this focus on building internal products. We are actually building a product to help folks ship the products to our end users. We've got to think about the internal users as much as we do the external users. And it's all about shipping at a higher pace with reduced coordination. Not no coordination, I've argued for that already, but reduced coordination, reduced handoffs. We're trying to optimize for fast flow of value from idea to shipping code, delivering value in production.
Now, if you're building a platform, you're doing some form of platform engineering. The Gartner folks have arrived in the space, and no one gets fired for buying Gartner, as we like to say, right? It's good to see the Gartner folks are in this space. They're focusing, again, on improving developer experience, self-service capabilities, automating infrastructure operations. This is what they're saying. If you're building a platform, You are platform engineering. And I've got a talk at KubeCon a couple of weeks ago where I go into more detail about how to build a platform, the three layers that Gartner have got on this diagram, and I go into sort of more of a guide for platform engineering for software architects. So if you're an architect or a CTO, check out on YouTube that talk. That will go into even more depth into the mechanics of the platforms. Now that we're aligned on platforms, I want to talk about building these resilient foundations. And my hypothesis for today is that a well-designed and implemented platform does enable this resilience, like I said earlier on. Simplifying platform interactions reduces the errors. And a big thing is being developer-centric.
Too many times when I was doing platform rescues about 10 years ago, I was going into big organizations, banks, e-commerce places in London, and they were building out internal clouds. Internal AWS often was the thing. I'd say, you know, your platform's breaking, it's not resilient, it's not being adopted, why is this? And they'd say, well, I don't know. We've built out an internal AWS, but none of our developers use it. And I'd sort of be puzzled, and I'm like, what do your developers want? Have you talked to your developers? And they're like, no, no, no, of course we haven't talked to developers. We know what we're building, right? We're building an internal cloud. You've got to speak to your customers. If you're building a platform, your customers are the developers, are the QA, are the folks internally, right? I talked a little bit about graceful degradation supporting business goals. And here it's building in policy and fault tolerance to the platform. I like the ideas behind shift left.
I see a lot of folks saying you go shift left. Think about security, observability, reliability, all the good resilience things, right? Think about it earlier in the software delivery lifecycle. But the danger of shifting left is you're kind of dumping left. Often developers can't keep up with all the things they've been charged with. They've got to think about security, observability, all these things. It's a lot going on, right? We need to make it easy to do the right thing. Bake a lot of this stuff into the platform to make it easy for folks. Shared responsibility promotes resilience. This is about strong collaboration and governance. In big organizations, you have to weigh up the centralization versus the decentralization tax, right? And for me, again, baking this stuff in policy as code, things like Kyverno and OPA, Open Policy Agent, many good things are available rather than us trying to do it manually. To be resilient, we want to put as much as we can into the platform.
And this is the meme I showed you earlier, right? This is the XKD classic meme. It talked about some of the JavaScript projects that we see, some of the infrastructure projects we see, saying it's very tempting to think we've got this very fancy modern digital infrastructure, but it's actually resting on this unmaintained project down in the bottom of the stack, right? And if that falls over or gets compromised, it doesn't matter if you're building sort of like resilient systems on top, you're going to have a hard time. And in the opening slide, I changed, if you look, subtly changed to your modern digital apps and services. Some platform component, some random person is maintaining, thankfully, in your org. I see this all too many times, right, where people don't think about the whole stack of value to trying to be resilient. You've got to weave in the architecture of the software and the architecture of the platform. They are symbiotic. They go together. So our case studies are the next five minutes apiece, roughly. I'm going to look at a large software company headquartered in the UK.
They focus on providing business management software across the world. They're actually expanding rapidly. I work for this company. I'm working with them now. And they've been in business for many years, and they're really scaling out. So they're looking to improve resilience for their developers by reducing cognitive load. I'm going to talk about Not on the High Street. They're a two-sided e-commerce marketplace, kind of like eBay, Etsy, that kind of thing. I worked on their platform 10 years ago, so pre-Kubernetes pretty much, and they were focusing on building resilient tech for Black Friday, which is convenient because that was literally last Friday. We're in Cyber Monday today, right? They were struggling maintaining resilience, maintaining uptime during these events where everyone went to their website and tried to buy stuff. And I'm going to also look, if I can, at the third case study. Fingers crossed. We're not working. Oh, dear. Let's try two case studies by the look of it. Not sure what's...
The batteries? Battery, yeah. Yeah. Let me think. Cool. I'm not an expert, but I... Ah, there we go. It's because of you. Magic touch. Thank you. Sorry, everyone. NetWest, the British banking and insurance company based in Scotland. I'm working with them now, and they're trying to promote resiliency through continuous improvement of their platform. They've got thousands of developers, and they're struggling with resilience across the actual SDLC in terms of reliably delivering software at pace and with a certain level of quality, too. So, let's focus on our first case study, the large UK business management software company. Unfortunately, I can't name them. I'm still working with them, and we're working on a case study, so stay tuned to my social media. Hopefully, the name will eventually come out, but I'll walk you through what's going on there. Classic meme, of course, on here as well. So the context, the business is growing rapidly. They are looking to scale software delivery across their teams and across their geographies. Resilience for them means being able to deliver at pace with that level of quality, that level of safety, and with the scale as well.
The challenge they've got, they do a lot of mergers and acquisitions. They're a private equity company, so they're doing tuck-in acquisitions all the time. And of course, that leads to very diverse range of tech. They've got... Java apps, .NET, TypeScript, Ruby. They're on VMs. They're in the cloud. They're all manner of different kinds of things. They want to go to more cloud-native technologies. They're going to the KubeCon conferences. They're looking at reInvent, which is on the moment, right? But this is where that meme comes from. They were just overwhelmed. And as they're spread across the world, different teams were spinning up different technologies, different bits of the platform, and it was leading to cognitive overload for developers. They couldn't be resilient because they couldn't ship code, right? They were all getting confused on how to deliver value rapidly to production. There is only a small platform team. This may ring true to many of you. They've got hundreds of developers, but three people in the platform team trying to standardize across the whole organization.
And really, they can only be consultants to some degree. They can't be too dictatorial because there's only three of them, right? So what do they do? They worked with us, actually, and they built a platform using Golden Paths, and I'll explain that in just a second. Kubernetes and Kratix was the underlying framework, and the developers were interacting with an abstraction rather than Kubernetes itself. So, Kratix is basically a framework to build platforms that goes on top of Kubernetes. It's very opinionated in the workflow, but not very opinionated in the technologies you put together to make that work. Workflow. So if you've used Terraform, and if you use similar kind of vibe, Terraform is very opinionated on TF plan, TF apply. It uses directed acyclic graphs to build infrastructure. But the actual configuration, the stanzas, are cloud specific. So opinionated on the workflow, but not so opinionated on the actual technologies. And that's where Kratix comes in. You can build these abstractions.
So they actually spun up a sort of DIY project where they... Used backstage, very common portal technology at the moment. They used backstage where developers could go on, click a button, get a GitHub repo. They'd say, I want a new Java app. It would then fire up a new Java app, a skeleton, a template, scaffolding. It would wire it into the platform for observability, for security, and it would give them back a URL to clone their GitHub repo, and then the portal would offer a series of buttons to deploy the code that was currently in the repo and then actually access the application. So developers were interacting with the abstractions, not Kubernetes itself. Now, the platform team allowed the engineers, the developers, to pull the escape handle and actually start creating their own bits of the golden path. So the golden path was the best way, the right way, the easiest way. To deliver code to production the most resilient way.
Another word for a golden path is a paved path. So you can go off the paved path, do your own thing, but you are then on call for that. If you follow the golden path for Team Created, they're on first-line support. If something goes wrong, they'll help you kind of stand things back up. But it allowed that flexibility of some developers didn't even know they were deploying onto Kubernetes. They got their repo, they did some coding, they pushed a button in a portal, and it was running with best practices in production. This was great for them. Some of the more advanced developers, they still had to follow the Kratix workflows, but they could customize them. And then they were on call for maintaining that. And that was a trade-off they were happy to make. Some learnings. APIs, abstractions, and automation reduce developer cognitive load. You've got to get it right. And the classic kind of thing I say is it's a Goldilocks thing. Not too hot, not too cold, just right. I've worked on some platforms that had too much magic back in my Java days, even Ruby on Rails.
Whereas when stuff went wrong, it was really hard for me to understand what went wrong because I didn't know what was behind the magic. I didn't know the trick. I couldn't understand exactly what was going on. So you want to get the APIs and the abstractions just right for your developers, and that is primarily through talking to them. A large portion of developers in this company just wanted a simple abstraction, a PAS-like abstraction, and therefore this abstraction they created was perfect. 10%, 20% wanted to do their own thing, and they gave them the option to do that. A hackathon was the way they got this into a company. They basically created environment vending through Kubernetes and Kratix. And developers were sort of like environment vending, getting hold of a provisioned Kubernetes cluster and deploying stuff onto it. And that was the way they brought everyone together, a guiding coalition, to get on board with this new platform initiative. And this is still underway. We're still learning a lot here, but this is the way APIs and abstraction, getting everyone together was really important here.
Moving on to Not on the High Street. So, Not on the High Street, the context there was they were struggling to meet Black Friday demand. The application was not resilient. It would fall over, typically midday on Friday, when too many people, when the US woke up, too many people started having the website, and the website would fall over. It was not resilient. They had migrated to the cloud, but it was a straight lift and shift. They literally took their technology, their processes from the in-house tin, and put it onto AWS. There was no re-architecting, no changing of processes. They wanted to automate scaling and DRBC, disaster recovery, but they never had time. The challenge, the software delivery teams were yet again experiencing cognitive overload. This was 10 years ago, pre-Kubernetes, pre-CNCF. You can imagine what it's like now, right? But this was back in the day. There were so many technologies out there, they weren't sure how to automate many of these things. When stuff fell over, it required the ops teams to literally log into the servers, kick things, restart things.
All the knowledge for resilience was locked in these static runbooks, literally Confluence wiki pages. Yeah, I think I might do. How strange. It's not like it. To push and grin stronger. Yeah. Now we've gone. Oh, how strange. Can I go back? It's okay. Yeah, we're good. Yes. Perfect. Thank you. So, yeah, knowledge was locked in static runbooks. And one thing that when I went in and I tried to help these folks as a consultant, I noticed that ops were not involved in any way in testing. So developers were kind of like trying to simulate Black Friday and doing a bad job, to be honest, because every time the site fell over. But yet ops, who had a lot of great knowledge, were not involved. in any of the testing. So what do we do? Platform solution, we modularize the architecture and the platform in harmony. We were reading Sam Newman's book, and Sam's going to speak this afternoon. We did microservices. Again, that was just emerging back then.
And we worked with the ops team to make sure we modularize the architecture of the software. And the platform. We actually ended up using Mesos, Apache Mesos, from Twitter, with some CLI tools, Jenkins, and we used RunDeck as a very simple portal for developers to log in, create an application, get it deployed, that kind of thing. And the real magic here was load testing with previous year's data. I mentioned that every year it fell over, every year developers tried to do load testing and they failed. So what I did, what my team did, we got the whole company together, the whole dev team together, QA, dev, ops, and we started running through the scenarios and the challenges. And the developers were like, we can't find realistic ways of testing things. And suddenly someone on the ops team was like, hey, I've got the Apache web logs from last year. I can literally show you every request that was made on the site last Black Friday. We were like, boom, brilliant. We need to sanitize that to remove personal data potentially, but we could literally take that Apache log, clean it, and replay it against the site, the new site, to see if it was going to stand up.
And we could also tweak and sort of like scale up the data. So the Apache logs had... realistic traffic patterns, we could pull out the key personas, replay it against the site, and then magnify that traffic to make sure it was going to hold up to the new goals. This was a game changer. Again, getting those three teams in the room, QA, Dev, and Ops, new solutions popped up. We ran the tests. We fixed things that broke. The database in particular was always a problem, so we preemptively scaled that up, making it more resilient. Happy days. The learnings here were forming a guiding coalition with clear goals and KPIs. Really, really important. I'm sure to the room of leaders here, this sounds really obvious to you, but it's all too easy to get stuck in the mindset or stuck in the silo sometimes and not bring everyone together with high-level goals. In this company, each team had their individual KPIs, but they were not aligned in actually delivering a good experience for Black Friday. We brought up the KPIs, said we want to hit this performance, given this load on the web server, and then cascaded the KPIs rather than having individual KPIs in the teams.
We automated a bunch of the failover, like real simple stuff, but it was really powerful, like using Terraform, Bash to restart servers, to preemptively scale up and down, magic stuff. And a big thing that enabled a mindset shift for the platform team was more resilient than ops people. We brought the platform team into the whole dev team, and we gave them an identity. That made them more resilient, made the technology more resilient. They felt part of the team, right? And this is really a key thing, bringing folks together. Moving on to our final case study, NatWest Group. So this is an ongoing project I'm working on. Chris Plank did a great talk at Kubernetes Community Day UK in London about a month ago now. So you can check out Chris's talk. But he talked about building resilience by wanting to decrease time to market. NatWest is a very traditional bank, a very successful bank, but they're being sort of disrupted a little bit by Starling, Monzo, Revolut, all the challenger banks.
They wanted to, the platform team or the team within the organization responsible for the platform wanted to increase contributions, get more buy-in, build resilient processes by not dictating what should be done. The organization has thousands of developers and then a reasonable sized platform team, but they realized resilience could not be dictated. They had to allow folks to contribute. The challenge for them, developer environment provisioning was a limiting factor. It used to take them 11 months to get a new environment. There was lots of policy checks that had to be done, all in the name of resilience, right? Security, uptime, guarantees, these things. But it used to take 11 months to get a Kubernetes environment for their developers to work on. So they were always preemptively provisioning, and then 11 months later it would arrive and it would be out of date. Lots of manual processes and patterns. Good intentions there for sure, right? There was lots of really good patterns, but they were all manually done rather than automation.
And there was limited resources. Poor Chris and his team were trying to keep the lights on and trying to build a new platform. What did they do? Platform solution. So they implemented a platform as a product. They stood up a platform as a product team. That was literally the name. They stood up an enabling team to facilitate contributions into the platform. If you're familiar with team topologies, that's two of your teams already, right? Platform team, enabling team. You've got the stream-aligned teams, the subsystem teams. Team topologies makes a recurring sort of theme throughout this presentation underlying. Some learnings. A composable platform did enable the flexibility. They were able to make a more resilient platform, and it's still ongoing, but by allowing these contributions and automating these manual processes. And the framework they actually used was inner sourcing. Folks are familiar, there is innersourcing.org, I believe it is. The PayPal folks talked about this. It's about doing open source principles internally. And you imagine NatWest is a huge company, and there's like applying like requests for comments and how you merge pull requests in, all the good stuff we've learned from open source and building resilient software, you can bring internally to your org.
So Chris and his team are still working on this. Check out his presentation to learn more about it. But again, he was bringing people together, clear roles, clear responsibilities, a shared goal across the whole organization to make sure they're building resilient platforms. Right, final slide, and I've got like 18 seconds left, so we're looking pretty good, I think, on this. I hope you're going to take away that well-designed platforms can enable resilience through the SDLC. In 25 minutes, it's only a high-level overview, but hopefully I've given you some thinking points to take away. For me, APIs, abstraction, and automation are key, but you really have to chat to the the developers, talk to the developers to understand what they want. Platform is a product you're building for your customers who are developers. Software architecture and platform architecture are symbiotic. Bring those teams together, right? Too many times I've seen platforms fail because they didn't know who their customers were, but also I've seen software architectures not take advantage of the platform, resilient features within the platform.
And this guiding coalition was so powerful as me as a consultant getting the right people in the room together, sometimes miracles happened, right? Giving folks an identity, you're part of the team, you're a platform team, you're not ops folks anymore, you're a platform team, you're building a platform for us as developers to get our code shipped. And lastly, inner sourcing is really powerful for pooling knowledge and increasing resilience across the org at all levels of the SDLC. Enabling folks to feel empowered to contribute code, ideas in a structured way is really, really powerful for improving resilience across the SDLC. And at that point, I say thank you very much for your time. I appreciate it. Thank you so much, Daniel. Thank you so much. So now we have eight minutes for a Q&A session, if you agree, if you don't mind. No, perfect. Happy. Do we have any people to raise hands to ask questions to Daniel? It's the first time I can see the audience. It's great. Oh, a first one.
Amélie. Would you please, two questions. In the same rows. Fantastic. Thank you. Hello. You're talking about involving developers into your platform and how to prioritize it, I would say. How do you organize to achieve that? Because if you have hundreds of developers, It's hard to say, okay, this developer needs that, but maybe the other one needs something else, and how to prioritize between them. Yeah, thanks for your question. It is difficult. So just to paraphrase, if you've got many, many developers, how do you weigh up all the contributions? I'm a big fan of using surveys to get general consensus. There's companies that can help you do this, like DX is very, no affiliation, but I just like what they're doing. DX help you create surveys to figure out what the blockers are in your organization, what developers want you to prioritize. And if you send that out to all the developers, the ones that care will usually respond back, and you can often identify patterns.
And then it's a case of what do you think gives the biggest bang for the buck or biggest bang for the euro, right? And I do think at the start you have to optimize for the 80%, the Pareto rule, right, 80-20. Unless it's a really strong, powerful business use case for those 20%, go for the 80. Ship that out. Work with some of the folks you identified in the survey that are really bought into this thing. Fix their problem, and then rinse and repeat. Do another survey with folks, and gradually work your way through. That's the best way I've found. Surveys are an underused tool, I think. And if you send it out and publicize this stuff, the folks that really care usually suddenly pop up on your radar. And then once you've done a couple of iterations, people really start dialing in. If I want to change something, I need to get involved with this initiative. Thank you, Daniel. Another question? Yes. Thank you for your presentation. How did you measure the adoption of maybe the key factors and the value of implementing platform engineering?
Great question. Thank you very much. So I've got a whole talk on this. If you look at my KubeCon talk and also my KCDUK talk, I go into this in much more depth about measuring platform. But again, you can look at sort of leading and lagging indicators, but lagging indicators are literally how many projects do you have taking advantage of the platform. A key thing there is how many folks do you see come back to your platform? I see many folks roll out a platform, seen some success, but then no one else uses it in the future. They go to Vercel or they go to like Heroku or something, right? So definitely measure that over time. But I literally look at how many projects are we running on the platform? What percentage of workloads are we running on the platform? These kind of things. And then I definitely think you want to measure happiness or do developers actually enjoy working with the platform? Again, surveys are a good way of doing this. There's a really good book that was published a couple of weeks ago by Camille Fournier and Ian Noland.
There's a whole chapter around how to balance business metrics with technical metrics. I read that on the flight over to Salt Lake City for KubeCon a couple of weeks ago, and that book's got an even better answer than I've given there. There's a whole chapter. on product health metrics, business impact metrics, and technical guardrail metrics. Because you always want to have a mix of the three of those. Product health is the health of the platform, folks adopting it. Impact is business impact. Are we reducing lead time? Are we reducing incidents, for example? And then guardrails, they're like, what's the change fail percentage? Dora metrics, that kind of stuff as well. But you need a nice mix of those three metrics. And Camille's book really switched it on to me. I was like, this is a great mental model for creating metrics for platform success. Any more questions? Yes. Thank you. Hi, thank you very much for your talk. Thank you. You mentioned that you want to bake as much as possible into the platform, and you also talk about standardization. But from a systems point of view, resilience comes from diversity, not from uniformization.
So how do you strike a right balance to get this resilience in this question of diversity versus standardization and uniformization? Yeah, that's a great question. You know, as I was doing my rehearsals for this last week, I thought someone might ask this. It's a great question because I totally get it, right, in terms of if you are all deploying onto one platform and then there's a bug, a critical bug for that platform, everyone's affected, right? There's no diversity there. So it's always trade-offs, right? And the consultant in me wants to say, you know, it depends. That's what we always say, right? Sadly, you can't bill for that, but that is what it is. But I think you have to weigh up the business priorities. Some folks really want a homogenized developer experience. And the trade-off there, if something goes wrong, higher impact, right? Some folks are like, we want to optimize for resilience. You know, we're a nuclear reactor-like shop, right? Very different mindset there. So understanding, are you trying to sort of beat your competitors in e-commerce, or are you creating, you know, mission-critical infrastructure things?
And then it's always about those trade-offs. And being very clear in the trade-offs, documenting them, using things like architecture decision records. You can do like platform decision records. We chose... to add some diversity into this because we value system resilience more than we do developer experience resilience, as an example. But always about trade-offs, great question. It's not an easy answer, to be completely honest. But I'll be able to chat more afterwards if you want to do that. Yes. So another one? Yes. Just for one minute. Okay, thank you for your presentation. Is making your platform resilient make you think differently for the security that just shift left the security and make a security champion? How do you handle these security things? So I guess you're asking about sort of how much security can you bake into the platform? More that security is always seen as a stopper for developers and ops platforms for pushing elements in the platform.
What's the balance there, kind of trade-offs? Yeah, so I always like doing threat modeling. So actually, my mastering API architecture with Jim Goff and Matt, we talked a lot about that. Threat modeling is a great way to get developers in the mindset of, like, how my systems potentially could be compromised. So I think development teams should do that as much as platform teams should do that. And then I like platform teams to make it easy for me to do the right thing as a developer. And often I've sketched out in teams I've run, where does the line of responsibility lie? It's tricky because you're all ultimately responsible for the whole thing. Like if we get breached, that's everyone's security problem. But just knowing roughly in terms of like the app developers are responsible for the dependencies they've got bringing into the application, right? They're running security scans on these things. But the platform team can make it easy for them to do that. Using Sneak or some other kind of tool, right? So developers say, hey, I'm responsible for these dependencies, but platform team, can you make it easier for me to do that? And then the same thing, right?
And platform team are responsible for the infrastructure, but they might want to get some tools to sort of make it easy for them to understand these things. But I guess with my seven seconds left, I'd say be super clear on who's responsible for what, but it's a dialogue constantly. Can developers get something from the platform team? And then can the platform team make asks of developers for certain guarantees they're going to provide in their application to sort of have collaboration across the whole teams for security goals. But come and chat to me afterwards if you want to know more. It's a great question. Yeah, but maybe let's continue the discussion during the break, if you don't mind. Thank you so much, Daniel. Thanks, Bill. For your enthusiasm. Thank you.
