Masterclass Tech.Rocks

Kubernetes, what else?

Masterclass Tech.Rocks · 16 octobre 2020 · 60 min · en anglais

Résumé

Masterclass sur Kubernetes organisée avec SFEIR, un événement exclusif réservé aux membres Tech.Rocks Core, avec Kelsey Hightower, Developer & Open Source Advocate chez Google Cloud. Tout au long de sa carrière dans la tech, Kelsey Hightower a porté toutes les casquettes et occupé des postes de direction axés sur la réalisation de projets et la livraison de logiciels. Fervent défenseur de l'open source, il s'attache à construire des outils simples qui font rêver les gens. Quand il n'écrit pas de code Go, il anime des ateliers techniques allant de la programmation à l'administration système. Au programme : - témoignage et retour sur le parcours technique de Kelsey - entretien La session de questions-réponses en petits groupes et de networking qui a suivi n'apparaît pas dans cette vidéo.

Summary

Masterclass on Kubernetes organised with SFEIR, an exclusive event for Tech.Rocks Core members, featuring Kelsey Hightower, Developer & Open Source Advocate at Google Cloud. Throughout his career in tech, Kelsey Hightower has worn every possible hat and held leadership roles focused on delivering projects and shipping software. A passionate open source advocate, he focuses on building simple tools that make people dream. When he is not writing Go code, he runs technical workshops ranging from programming to system administration. On the agenda: - Kelsey's testimony and a look back at his technical career - an interview The small-group Q&A and networking session that followed is not included in this video.

Thèmes : Cloud, infra & ops

Transcript complet

Transcription automatique, à relire : les noms propres peuvent être mal orthographiés.

So hello, Kelsey. So it's a great pleasure for me to welcome you on Tech.Rocks stage. So how are you today? I'm good. Yeah, I'm here pretty early, but the sun is bright, so I'm off to a good start. I think you are from Portland. Yeah, I'm close. Camas, Washington. So right above Portland, Oregon. Okay, great. What time is it for you? Oh, it's like 9 a.m. Okay. That's pretty early. Okay, for us it's 6 p.m. Okay, so could you tell us a little bit about yourself and what you are doing now? Oh, wow. There's a lot. I mean, you know, I consider myself a technologist. I mean, I'm at Google Cloud, so I'm working on a wide range of things from security, but a lot of it centers in this cloud native space. So what comes to mind for most people is like Kubernetes.

You have a lot of the things with networking and service mesh, authentication, security, those kind of things. So I kind of stay in that sweet spot in terms of Google Cloud, and I'm currently working on some contributions. I'm writing a new book on service mesh in general and open source contributions. So continuing that theme. Great. And how was your journey to get where you are now? Oh, I mean, for most people, you think about how do you get into tech. For me, I didn't go through any CS courses or anything like that. I'm kind of self-taught. You know, I bought a book on, I think, Linux way earlier, like in my early 20s, just to start to get my feet wet with this whole Unix space. And like most people, I started out doing bash scripting. And there's a long history there, but I've opened the computer store and did that for a while until I got my first job inside of a Google data center doing system administration work. So think large server farms, you're automating switches, you're automating servers and provisioning.

And throughout my career, I've worked plenty of jobs in terms of like finance, large enterprises, small startups, a lot of open source. Source. And now I'm in the cloud. So I've just been learning along the way. Okay, great. So technology cycle is shorter and shorter. So what is your secret sauce to stay up to date and to prepare for the next evolution? Well, so far, I think learning the fundamentals. I think a lot of people skip that. For example, when someone sees something new like Kubernetes, they think new, got to change your mindset, cloud native. But to me, what I see is a continuation of like all the things I learned from the Linux world. You can package your application. Back then, we used RPMs. Today, people say containers. We run processes. We give them permissions. They bind to ports. And then there's metadata produced by our binaries, right? So how much memory, how much CPU they're using, the metrics, requests coming in and out.

So when I see these new things, I always try to ask myself, how does it relate to the fundamentals? How does it relate to the previous way we did things? And I tend to find like 80% similarities. The APIs may change, the name may change. So I usually stick with the fundamentals. So that means when I learn something, I try to go super deep and understand how it works at the very low level. That way I'm equipped to learn the next evolution when it comes down the pipe. So that's been my key, staying on top of things. Then I try things with my own hands. I just don't read the blogs and go with that. I say, hey. If this is a new thing, let me download it. Let me try to build an app with it. Let me try to understand it. Let me look at the source code. And then if I feel like I understand it, I can put it back on the shelf. So that's how I stay up to date on these things. So, I mean, how many hours by week you try new stuff and things like that? I would probably say I spend about 20 hours a week. So I work up pretty early. For example, last night I was up since 2 a.m.

Contributing to Open Policy Agent, right? The guy there, you know, the idea there was, you know, there's this new security framework people are using for authorization. And I was like, well, how does this work? So I want to run this inside of Google Cloud. I noticed that they didn't have native GCP authentication, meaning using the metadata service. So I spent maybe four or five hours, added support for that. So wrote a little bit of code, wrote some documentation, wrote some unit tests, and then I'm able to create the tutorials that I share with people. So again, Going super deep, learning how it works. I think I spend about 20 hours a week either writing code or trying to make sure that I understand how all these things fit together. Yeah, it's quite a lot. Okay, so now let's talk about Kubernetes and Kubernetes for CTO. So can you explain what is Kubernetes to a CTO audience? And what is absolute minimum to know about Kubernetes?

Yeah, if you're coming at this from the CTO perspective and you're in charge of the technology vision, right, there's multiple types of CTOs, right? There are some people who are like leading the engineering team, trying to help design what gets built, what gets bought and how that technology is leveraged by the business. Sometimes you're going to have a visionary CTO about where the company should be going. And maybe there's a VP of engineering that handles a lot of the day to day responsibilities. But I think when it comes to a CTO, you have to ask yourself, are you aware of all the technology that's available to your business? I think that's step one. So if you think about 2020 in terms of writing custom applications. And then hosting those in a reliable way. So any application being built these days, you have this expectation that it could run globally. It's always on, it's highly available, it can auto scale. All of these things that we've been learning over the last 15, 20 years, The new checkpoint for all of those features can be found in Kubernetes.

So if you're a CTO and you say, I want my team to build resiliency into the platform, I want my team to have a framework for doing security, you know, previous policies and the ones that are yet to come, you can go and build that from scratch, right? So if you only have a limited set of team members that can work on things, do you really want them building all of this from the ground up, right? You wouldn't go build your own operating system from scratch. You go leverage something like Windows or Linux to go do that. So I I think on the compute side, when you think about what Kubernetes is, I remember when like Apple iPhone came out and now you have iOS. You now have a much better SDK to start building mobile applications than you did back in the day. So I look at Kubernetes is very similar to what's happening in the mobile space. Can the platform get a lot smarter? Have all the patterns built in in terms of thinking about security, auto scaling and so forth. And then when I write my application tier or I go build my part of the platform, I don't have to start from scratch.

So I think a CTO, whether you use Kubernetes or not, You have to be aware of what's available and does it make sense to leverage that for your particular team? And if you're paying attention to where the industry is going, this style of infrastructure is starting to become popular across all mainstream cloud providers. You even see traditional vendors like VMware starting to provide this service as well. So it may become commonplace, just like Linux was 20 years prior. Okay, so for example, I'm a CTO, I'm starting my startup. My first step is to set up a Kubernetes cluster. Should I wait for a time where I need it? So at what time, when I should start to think about Kubernetes? So let's walk through some scenarios. So if I was a CTO and let's say I was building a healthcare service, right? So the first thing I would ask myself is, how much infrastructure do I need to start, right?

Because this is a startup. So maybe I focus on my APIs and my user experience. Then I would probably say serverless is super interesting. So I work over at Google Cloud, and there's a way to take your containers. So maybe you want to have a portable packaging format. I can write my web applications using any popular framework, whether it's Node.js and Express or just GoLang in the standard library. You probably can get off the ground by just starting with serverless. Manage services for your networking. You can use something like Cloudflare for DDoS. You can use a managed database to store your data. And for a lot of people, the first six months of just trying to get that first customer and finding a product, you know, market fit for your product. I think a lot of times service is a great way to get off the ground. You don't overinvest in infrastructure that you're going to have to probably blow away anyway. So you just have some clean abstractions using managed services. Then once you start to get into more... You know, diverse workloads like background jobs where you're processing lots of data. So let's say you're now one, your first customer, and now they're sending you tons of data.

And again, most serverless platforms can handle that well, but maybe you need a GPU to process the data in some way. Maybe you have some algorithms that you want to apply on top of this. You're getting into machine learning. Then the flexibility of Kubernetes being able to run not just stateless web applications like most passes can, but also the ability to do some of these head heavyweight jobs. If you're a voice over IP and you need to do some telephony work, then network performance is going to be critical. Again, having something like Kubernetes that allows you to model, but I wouldn't recommend you go and build Kubernetes from scratch and go hire a Kubernetes team. There are enough managed services like Kube, for example, if you're using virtual machines, in the cloud, then you're probably using a cloud provider's managed hypervisor, even though you're getting a VM. So there's no reason to try to manage the underlying infrastructure as a startup. I would look for managed services, no matter what cloud provider you pick, they're going to have something for you there. Okay, so I have a startup. Now I need a cluster, maybe with 100 nodes, even if it's managed.

What is the typical team to take care about this cluster of 100 nodes? So. When you think about the team size, so some things that you're going to get from just a managed nature, it's not going to be zero overhead or zero ops, as some people like to call it. Having a provider take care of things like the Kubernetes version you should be running, taking care of patches, that's going to be handled by your provider. So Azure has a good service. Amazon has a good service. So does Google, Alibaba, you name it. But then what's left for your team to do? So this is what we start to think about kind of delegated responsibility. Your team is still going to need to be knowledgeable about security concerns. The way you configure that cluster could open up some security issues. ramifications that you have to be aware of. You're still going to need someone thinking about your policies. So even if you have a Kubernetes clusters, who's going to create those container images and keep them patched? Who's going to actually build the job specifications that say, I need three copies of this app running with this memory, with this CPU, making sure that you have enough resources to run the app properly.

You're also going to need some form of site reliability, SRE as it's called. So when your startup makes a promise to a customer that the app is going to be up for three nines of availability, then you're going to have to have a team that's making sure that's true. That means you're doing things like metrics, you're having your logs, you're being able to mediate issues. So I think it's still going to need an ops footprint, but they won't be working on. Deploying and patching servers, you're going to need that team focused on things like observability and managing the application tier as it meets the infrastructure. Great. So what happens if the clusters still grow? I mean, we are at 100, now it's 1,000. Do we need to increase this team? If it's managed, it's... You know, how we scale up. Yeah, so I get it. So in Google, we try very hard to make sure that SRE as a practice doesn't necessarily require the more machines you add, the more people you need to add, right?

So in this case is that Kubernetes provides a nice set of abstracts. is that even when you add from 100 to 1,000 machines, you've already kind of paid the price in terms of scale horizontally. I think when you start talking about growing skill sets in the team is when you start to leverage things like server mesh, right? Like Istio, Linkerd, that's when you start to think about, do I have the right networking skill set to leverage these things? Or when you start to get into multiple regions, right? So having everything in one cluster, you can probably get away with the same type of skill set. Of course, they need to learn a little bit about Kubernetes and how to manage it. But once you start to have multiple clusters across multiple regions, now you're starting to get into some hard areas of distributed computing. You're going to have to think about how you split your data. Do you partition the data? What type of databases are you using? How do you reconcile those kind of things? So while I don't think the number of Kubernetes administrators will have to kind of increase with the cluster size, I do think once you start to do this kind of global scale service, you're going to probably need some new skill sets that you don't currently have.

So just learning Kubernetes, even though that's going to let you scale horizontally. It's probably not enough to think about these other disciplines around storage, security, and networking. Okay, great. So some people say that Kubernetes is a kind of OS operating system for cloud. Do you agree with that? Well, I agree with it in the sense that when you think about a Linux server and you want to write to a file, you just kind of send data to a file. You have no idea of all the hard work the kernel underneath does to actually write data to a hard drive and deal with file systems and inodes and so forth. And when you want to run a process on Linux, again, you have things like systemd, where you can write a systemd unit file and say, hey, run this application for us. So when you think about a user land on top of a kernel, Kubernetes has a lot of those characteristics, right? It has an opinion about packaging an application. On Red Hat, we used to do RPMs. The Java world, you do WAR files. But in Kubernetes, you bring a container, so you have that component. Then you also have this kind of configuration language that It says, I want to run an application in this way.

And once you make that declarative statement and the app were to crash for some reason, then Kubernetes starts to exhibit kernel-like behavior where it will restart the process. It will babysit the process. If you have some security policies where you say, I only want these applications to talk to this other application, just like when we used to use IP tables on a local server, you can do something very similar into the Kubernetes world. So it does feel, it starts to feel like at some level, it provides these OS-like things inside of the user land to allow you to just operate at a very high level. Okay, great. So we agree that Kubernetes is a kind of OS like Linux. In the Linux world, we have many distributions. Is it possible that in some days, the kind of Kubernetes distribution like the Red Hat one, the SUSE one, etc., there could be a Kubernetes GCP, Kubernetes

Amazon, Kubernetes Azure? Do you think it's possible or not? Yeah, that happened three years ago, right? So we're three years into this. So when you think about like the raw Linux kernel, almost no one uses just Linux by itself, right? You're using GNU Linux, you have Bash, you have some user land, you have a bunch of utilities, LSGREP, you name it. So most people have never used Linux by itself. in a long time. And then you start to think about distribution. So how should you package an application on Linux? Well, if you go with Red Hat, it's going to be RPM. If you go with Ubuntu or Debian, it's going to be kind of a dev or snap package based. So in Kubernetes, I think one thing that is kind of universal between the distributions, we kind of all agree that it should be a container image as the way we distribute things that run into this operating system, to use your analogy. But the user land can be different. So when we talk about network policies, you can choose to use a different component like cumulus networks. You can use just IP tables, the standard thing that comes out of the box.

When you think about the interface, right, you can go from bash shell to Z shell on Linux if you want a different user interface. In Kubernetes, you don't have to write YAML files for all your configs. You can use something like a Helm chart. You can change the way you interact. You can use the command line, kubectl, to do a lot of things. Or you can use a web UI or something like OpenShift that adds even more things on top of Kubernetes where it feels like a PaaS, just like Android takes Linux and turns it into a mobile operating system. Now it becomes a platform-specific thing. And you've seen this with Kubernetes. There are companies that are building CICD platforms on top of Kubernetes to make a whole new platform using that base component. Okay, great. Thanks a lot for this. So now we will talk about architecture around Kubernetes. So there's two visions at the moment. Some people say that we need only one big cluster that will welcome all our applications.

And some people says we need multiple clusters, maybe clusters for developers, clusters for production. Even some people says I need one cluster for each of my teams. So what is your recommendation there? Well, so, I mean, I think people already know the answer if you were to think about the fundamentals. Let's say I can just run all of my production on one big Linux server, right? Just give me a big machine, get 200 CPUs there, two terabytes of RAM, put my database, put my app, and it should be fine, right? Everyone just uses one big server. You could do that, but we don't. Because we don't want the one control plane. Let's think about it in terms of Linux, right? You have this kernel. If that kernel were to have a bug or were to reboot, then we lose all of our applications. So let's go into the Kubernetes analogy. If you have one big cluster, even if you try to spread it over multiple availability zones, you kind of have one kernel managing everything. So let's think about a bug. Let's say you update from Kubernetes 1.13 to 1.15, and then maybe some fields were deprecated.

There is a chance that none of your applications will restart. because you're using some outdated config and now your entire cluster is down. So when you think about it, the blast radius for Kubernetes is really controlled and determined by that one control plane. It's not about having the API server go down because you have redundancy and you're running multiple copies. Usually it's going to be a configuration mistake, meaning you deleted the wrong app in this one cluster. So now that app is down, all of these horizontally scaled containers, they're gone in an instance. So to remove that, just like in the Linux world, we went from one server to two just to have some higher availability. The same thing we do with Flex Storage. You have more than one spinning disk. to have highly available. So in the Kubernetes world, depending on what your needs are, if you're telling customers that you can survive one zone going down, then there's no way you can do it with one Kubernetes control plane. You're going to have to have another control plane running or give you the ability to say that this cluster went down or maybe we made a configuration mistake.

We've updated the wrong app. Now it's down. Well, now you can go to your load balancer tier and just have traffic go over to the other cluster, even though this version of Kubernetes is bad. So I think the pragmatic approach here is this. What blast radius do you want? And we know through fundamentals that you have to limit your brass radius usually at the control plane level. It doesn't matter how many nodes you have. It doesn't matter how highly available that that control plane is. You're just going to have a highly available outage. Okay, so just to fix the idea, how many nodes we could have now in one cluster? And what is the ambition? You know, I think recently we published that you can have 15,000 nodes in a cluster. And then some people say, why would you even do this? This makes no sense. This is dangerous. Because really the limitation of Kube is not necessarily just the number of nodes you have, right? That's kind of a thing that used to be true back in the day. It comes down to the number of objects that Kubernetes needs to manage.

So for those that need to understand what an object is, if you have one node in your cluster, Every set of containers and configuration and secrets that you manage, all of these things are represented as objects. You want to run a cron job, you create an object for that. Now, when those objects get assigned to a machine, then they have to be help checked. There's logs that are collected. The agent that runs on the worker node has to call the API server and say, hey, is this object up to date? Should I be running another container? Should I restart this container? All that puts a tax on the API server. So if you have a small API server, even if you have a small set of nodes, you're not going to be able to handle all of the objects. So when you hear 15,000 nodes, you're not going to get 100 different objects on all 15,000 of those nodes. So the use case there is, let's say I'm doing machine learning, where that entire machine is going to be saturated probably by one workload. It's going to use the entire GPU. It's going to use all the memory and all the CPU. So in that case, I'm not going to benefit from bin packing, you know, placing multiple workloads.

on a single server. So in that case, I'm going to have to take those VMs or machines underneath and scale them horizontally and do more of a one-to-one pairing to utilize the resources. So that's the case for the extreme large clusters. But then if you have a mixed workload, what some customers will do is they'll have some machines that are one-to-one for the big machine learning, but you'll have another node pool that are more for the web apps, more for the cron jobs, and those can be mixed and matched neatly. So it's not about cluster size per se. It's about the number of objects under managed. And do you have enough bandwidth in your control plane to handle all of those secrets updates and rotations of certificates, et cetera? Okay, great. Thanks. So at first, Kubernetes was built to run mainly compute stuff. And now we are also able to run databases like Cassandra.

So we are now able to run storage stuff on top of Kubernetes. So beyond the technical challenge, is it a good idea to run, for example, Cassandra on top of Kubernetes? So this was less about a good idea or best practice. This is more like since day one, Kubernetes in theory could always run any compute job. Because if you think about what Kubernetes is doing, it's taking an app and it container, a tarball, pulls it to a server, and runs it. That's very equivalent to taking a binary and just copying it to a Linux server. There's not a fundamental difference here. So then you ask yourself, what can Kubernetes do to help me run a database? Well, what does Linux help you do to run a database? Not very much, but somehow we're able to run a database on a Linux server. So now we have to talk about responsibilities. So for a stateless app, Kubernetes is going to do very nice things like if the machine crash is going to run the container on another machine.

It can do that because the contract between your app and the infrastructure is pretty small. Now, when we get to a database, we tend to add another component to this contract. If it's a single machine database, like I'm just running a big MySQL instance, then the primary contract is going to be with the storage. Right. I need a data volume to store my data. In Kubernetes, this doesn't have to be hard. You can take. one node in your cluster if you really wanted to, you can mount the volume just like you do in Linux. You actually can mount the volume before Kubernetes is even installed. If you need two terabytes of data, just mount the volume, mount it to slash var lib data. And then when Kubernetes runs the container, you can just tell Kubernetes, use the volume that's already there. There's no need to do all of this additional orchestration if you don't want to. So in theory, you can do it already. Now what people are looking for is something like, wow, when I use a managed service, it does everything. It does snapshot of the data.

It backs up the data. It upgrades the database. That stuff is really, really, really hard. And there's lots of control planes and workflows that go to do that. So when you look at Kubernetes, it's very attractive because it has all the machinery at a base layer to build that kind of automation, all the metadata about what's running, how to update a container. You can do all of those things. But what Kubernetes does in 2020. It will help you with things like I can create a volume and mount it to a VM for you before the database lands. All right. That might be a little help. But when it comes to things like replication and cluster management, so you start to have multiple nodes, Kubernetes doesn't quite know how to deal with like updating a config file to tell you where your replica is. That's out of scope for Kubernetes. So databases like CockroachDB, MongoDB. These have built-in replication components that can then leverage the metadata from Kubernetes to automate updating of the configuration, update and automate the failing and recovering of a cluster of these database components.

So it's more about Kubernetes will meet you maybe halfway, and then you're going to have to bring in some additional help. And that additional help, most people call it like an operator. I'm using the Kafka operator, right? But those are really just saying. This is the things that Kubernetes doesn't do and everything in this operator is about closing the gap in the delta between running a fully automated database and what Kubernetes provides. Okay, great. So, do you think it's... It's good to run the storage part and the compute part of an application in the same cluster. Or do we need to have a cluster for storage and a cluster for compute? I want to talk about this at two levels. We already do this today. If you have a cloud account and you pick a zone or a region, you have some VMs in this VPC that run the web app, and then you have some other machines in the same VPC that run the database. So we've already been doing this for 20 years because those machines are effectively in the same cluster.

Now the question comes down to at the second layer, should the same management tool manage both? Yes. So this is where we get back to the shared control plane. Now, if my web app goes. I can probably spin those up pretty quickly. So meaning if I have a bug in Kubernetes, let's say I'm updating Kubernetes because my stateless web apps want to use some new fancy Kubernetes feature that my database doesn't need. Now what you're doing is you're increasing the risk of an outage on the data component for the needs of the web app component. So using the same control plane that's actively managing may put you into additional risk. So then you would ask yourself, well, if the data is super critical. And I want to decouple the lifecycle of the management tool away from the data. So then what you will see people do, which is very common, like when you use a managed service like S3 for your object store, that's a decoupled control plane from your compute control plane, right? So it's always been a good idea if we can decouple these things if we're worried about that causing an outage.

So for me, if I had very sensitive data and I wanted to limit various things that could happen, I would probably consider having a separate cluster in the same VPC. The VM still can talk to each other as if you had no Kubernetes. Don't worry about that part. It's more about I have a separate control plane that I may over here. I may stick to Kubernetes 1.16, whereas on this cluster, I may use the latest Kubernetes 1.18, and I just reduce the amount of churn over here because the life cycles are different. So it's just really about pragmatism and reducing the chances of an outage. By getting a little bit of separation. Yeah, great. Thanks a lot. And what's new in the world of Kubernetes? What could we expect in the next quarters? I mean, the way I look at it as a person who looks at fundamentals and pragmatism, I don't expect a whole lot ever from Kubernetes going forward, right? The whole thing now is about making things stable, making sure that we have the right hooks for security, making sure we get a little bit better.

With our storage and back end. So you'll start to see a lot of our work going from beta to stable. That stuff takes a lot of time to roll out because we have to get it right. People will be doing upgrades in their cluster. So from Kubernetes as a core, It's fairly complete, right? Stateful sets are now GA. If you think about the networking, so you still have to see some evolution on the networking side. So when Kubernetes first came out, it had a very simple networking model that worked for a lot of the kind of three-tier web apps, right? I deployed my app. I need a little bit of integration with the load balancer. We call that ingress. But now that people are starting to play around with service mesh, you're starting to see the ingress spec. You know, change over time, because now that we're looking at it, we need more data. Maybe we should not try to have one ingress object trying to control Google load balancer, Amazon load balancer, and Nginx, right? Maybe that's not a good fit. Maybe we should have a better separation so we can attack the unique needs of the various platforms. We've been through this before, right? Puppet, Chef, and Ansible, I remember when those started out, you know, they try to abstract away an app and then try to do magic under the covers about what SSH meant for Ubuntu versus what SSH meant for Red Hat.

They actually have different package names with some of those operating systems. So it turns out you really can't make a universal standard for every load balancer. So you're going to see Kubernetes think about what does version two of the ingress networking object look like. And then when we start to have more service mesh features, you see this stuff from Microsoft. They did the whole service mesh interface, which is maybe we can bring a few of the service mesh components that people like from Istio and Linkerd, and maybe they should be available in Kubernetes by default. Things like TLS mutual auth, basic authorization, and then we'll get a new set of objects in Kubernetes. But again, those things aren't quite core. They're kind of like these third-party things that are easy to install. Okay, great. So now let's talk about Kubernetes in the Google world. So I know that Google has a product, the name is Anthos. Is it a solution to watch? Can you explain us what is Anthos?

Yeah, so it depends on what level you want to come at it at. So some people come to GCP and they say, wow, this GKE thing is really great. Fully managed Kubernetes. You manage the worker nodes. Maybe you should just explain what is GKE first. Got it. So GKE is Google's Kubernetes engine. So, you know, there's open source Kubernetes and we take those binaries and we do all the integration work with Google Load Balancer, Google IAM, with all the VMs, storage, networking, GPUs. So just to make sure that it's easy to make sure that we can use all the capabilities of GCP when you're using Kubernetes. And then we have a separate control plane. So when you want to create clusters, let's say I want 25 clusters and this cluster needs to be in this zone. This cluster needs to be in that region. Oh, and this cluster needs to have GPUs, but the other ones don't. So GKE is a control plane to actually manage the lifecycle of Kubernetes itself. So that means we can horizontally scale worker nodes.

We can upgrade them in place. There's so many things we can do in terms of cluster management that lives outside of core Kubernetes. So that's GKE. It's a managed Kubernetes offering. Kind of like that Kubernetes distro thing we were talking about before. So then people say, well, how do I get GKE on-premise inside of my data center? Well, in order to do that, we call that Anthos. So as more and more GCP features people want on-prem, like, for example, we announce a part of BitQuery being available in other places. When we say other places, we need to have a solid base infrastructure to make sure we can run BitQuery like we do internally at Google. We need infrastructure to make sure that we can run GKE like we do in GCE. So Anthos is more of this umbrella project to say all of these things you like about GCP, if you want to see them in other places, they're probably going to be done under the Anthos umbrella. So when you're in GCP, you might be just using GKE. When you're on-prem on top of VMware, then you're using Anthos, which is then controlling and programming VMware to give you a GKE experience, right?

And the same is true for metrics and logs and all the services we plan to roll on top of Anthos going forward. Okay, so Google is kind of... leader in the Kubernetes world. So is there something that Google will release in the next few months that you are excited about? Yeah, so when we say leader, it's that I believe we bet really big on Kubernetes since day one, right? Created the project, we released it to open source, and we've been stewards of that community ever since, welcoming contributions from the Red Hats and the VMwares and the... Microsoft of the world. So there's this nice collaboration happening here. But when you go to GCP, you can feel that we're all in, like the console, the native integration. And then you see things like using the Kubernetes API to do things like manage GCP resources. Like if I want to create a spanner database, which is our multi-region SQL database that has consistency.

I can now create like a Kubernetes config, give it to the control plane, and then it will manage the database for me. So we're taking this Kubernetes API definitions and layering it on top of other parts of GCP. The other things we've been doing things is like vertical auto scaling. So everyone is familiar with the baseline horizontal auto scaling. Maybe you're having a lot of CPU from your containers. And then Kubernetes can watch that and then scale horizontally. But what we're also doing is a thing around vertical autoscaling. So when we see that you're running out of memory, maybe you got an out of memory exception. When that container crashed, we can, based on your policy, we can update that particular deployment spec to have a new set of memory requirements and then redeploy. So what we're doing here is we've already taken care of horizontal node autoscaling. We can do things like if you give us a new workload and no nodes exist, we can auto provision those nodes by just looking at the workload and using some policy that you do when you create the cluster.

So adding things like vertical pod autoscaling, really hard problem. That's a part of GKE. Things like binary authorization, this concept of end-to-end, when you build your containers, when you scan the container image, and then when it's deployed, you can do all of these policy checks. So a lot of the things that we're doing in the Kubernetes space is making sure that you can have end-to-end policy management. So in GKE's world, we're trying to just blend those worlds, for example. When it comes to service mesh, and we talked about the complexity of multiple clusters, we want to give people new concepts like environments. So imagine having 100 clusters, but 30 of them are in production and 70 of them are in dev and staging. Well, what we want to do is make it super easy to use all of these service mesh concepts, all of these configuration management concepts that you see with like GitOps approach, but give you new high level of abstraction. So you can say these 30 servers are in this environment, and then we will automatically add in the service mesh components. We will automatically do things like TLS encryption and give you a way to easily manage multiple clusters like you do with one cluster.

So the things I'm excited about is all of our work making it easy to adopt multiple clusters, not just in GCP. But even on-prem. So if you look at our load balancer, it can now send traffic outside of GCP. If you have any experience with most cloud providers, they will not send traffic outside of their VPCs or their products. We've done a lot of work to make sure that you can actually have traffic come into GCP. Some of that traffic can go into your in-cloud Kubernetes or on cloud. on-prem by targeting a private set of IPs going through your VPC routes and direct connections. So I'm excited about all this multi-cluster work. Okay, great. So you talk a lot about service mesh. There was a kind of battle around service mesh. Google make a great product, issue, but Google didn't give issue to the CNCF. So first, my question, my first question is why?

And the second question is, could it hurt? The global ecosystem because service mesh is very important in a Kubernetes cluster. So why do we, I think we need, we will need a kind of default implementation of service mesh. It could have been easier, but it's not in front of the key. So let's walk through this real quick. So a lot of people get hung up on the foundation component. Google helps start the CNCF, right? People need to remember that component. And then the community aspects of the Istio. That control plane, there are lots of contributors. The IP has clear legal guidance now with the foundation they started for trademarks. So if you like Red Hat, you can ship Istio today just like Red Hat does. If you're like Amazon, you can ship Envoy, which is not part of Istio natively, right?

That comes from CNCF. So Istio depends on Envoy, which is in the CNCF. So the real workhorse in most service meshes is going to be this kind of proxy component. And for Istio, that's Envoy. So Envoy is already across the board. Now what we're talking about is what control plane do you want to use to configure Envoy? Right. So the Istio product, IBM actually is a founder of the project. So it's not. Google only thing. Actually, if you go back in history, IBM started the project with an existing product they had. So you already have a world of contributions there. From a licensing standpoint, it's Apache 2. Linux is GPL. You can take any of these components, just like Hatchicorp Terraform, right? Terraform is not a part of a foundation. People use it just fine. So there's that. And then if you think about it, a lot of the world's most successful things are not necessarily in the foundation. Like, think about it. Most of the cloud-native products that you see in CEF, what language are they written in?

They're written in a Go programming language. The Go programming language is not in the foundation at all. And it's fine. It will be okay. I think the real thing here is in the community is if Google starts having bad behavior, meaning rejecting bug fixes, rejecting feature work that people are willing to roll up their sleeves and do the work, that's when you have to hold them accountable from the community. So the way I look at it is for the last four years, They've done all the right things in terms of source code and collaboration. So what foundation it goes to, look, the community can debate that one endlessly. What foundation it belongs to. OpenStack was in a foundation. I don't know if that made it super successful or not. So I don't know if foundation is a key to success, but you have to hold them accountable on the source code at the contributor level, how it's governed. I'm going to let people decide. I can't speak for everyone else, but I think it's been okay so far. We just have to be really pragmatic about this. It's not if it's in CNCF, that's the only thing that counts. You know how many software that people depend on that are not in CNCF?

So we can't have that be the baseline of whether we can use software or not. So I just think we have to be careful. The foundation does offer lots of benefits, but I think some of these projects don't necessarily need only CNCF to achieve those things. Okay, thanks for this clear answer. So last question before we take a question from the audience. Do you think Kubernetes is a real solution for green computing? I mean, there's so many more components to green computing. If you think about it, if you're super serious, then you're going to be looking at the system calls you make. You're going to be looking at things at lower levels, like turning off machines. Like when Google, we wrote a paper about how we really get to like power efficiencies. So I think having a scheduler is a key component. You really want to automate the good behavior. If you're not using something, you want to turn it off if you can. So like serverless is a good concept here. Shared infrastructure, scale to zero when it's not being used.

That's just like turning off the lights in your house when you're not using that particular room. So does Kubernetes provide a foundation to do something like that? Yeah, it's a great solution. It's probably better than people trying to build it from scratch, doing it inefficiently and not benefiting from the collective effort here. But will Kubernetes alone solve green computing? I'm not sure about that. But I do think it's going to actually make it easier than not having something like Kubernetes. But you have to look past just Kube. Are you making efficient system calls? Are you using bloated frameworks that are just eating up so much memory doing literally nothing? Right. That stuff is extremely wasteful when you don't need it. So I think there's really things you can do up and down. But having a scheduler is a fundamental thing that makes it easier to continue to have, you know, regulate your resource usages. Okay, great, Casey. So thanks a lot for this first part. Now we will switch to the, we will welcome a question from the audience.

Kelsey, could you see the questions in the, I can see the question. Oh, I was looking in the general chat. So I guess it's in French something. Oh, yeah. It's in the, yeah. Yeah. All right. So I can see the general chat. So I can tell questions from here. Maybe you can ask a question and answer directly. Yeah, I can do that. So I'm just going to go with the highest voted questions here. So I'm just going to read the question, now give the answer. Yeah, great. All right, so the first question is, you know, how hard is it to transform my VM-based legacy application to a Kubernetes one? And this is one of the things I used to do a lot in the early days. So when I used to contribute to Kubernetes, I used to just take a lot of existing apps like Jira, written in Java, needs lots of memory. lots of configuration options, et cetera. And I used to just try to run them on Kubernetes just to see if I can take existing apps and run them. And a lot of the work in the Kubernetes space was just so we could do that. That's why you can mount configuration files.

So if you have an app that needs a whole directory of XML files and all of that, you can do that. You don't have to do 12-factor. And do everything with environment variables. Kubernetes tries to make it easy to replicate what you're doing on the VM. Also, don't lose the fact That you're still just running on a VM. It's not like you're running in some brand new kernel with a new everything. You're just running with a few new restrictions, meaning you're going to restrict access to the file system. You may limit how much CPU. So let's say you have an existing app that's making you money. It survived the test of time. Good job. For a lot of people, they've been running those applications as root, doing whatever on the system, so they're not quite sure how much of the operating system that the app is using. So when you go to containers and Kubernetes for the first time, you start to learn a lot about your app. You say, oh, wow, I didn't know my app needed that part of the file system. So then you find yourself mounting things back in that are turned off by default. Maybe your container isn't running as root anymore and it doesn't work because it needed to run as root.

Well, you can go fix that in the app if you have to. But to be honest, some people are just running those containers as root in the same way that they're running them on the VM. So what you have to look at is when you go through this trial and error process, I recommend people don't make any code changes first. First thing to do is say, hey, what Kubernetes features do I need to turn off in order to make my app run unmodified in Kubernetes? The nice thing about Kube is when you make those decisions, you'll see it in your deployment YAML files like, oh, you're mounting in these extra paths from the file system. Oh, you're running these Linux capabilities. Oh, you need to run as this user. Once you've done all that, All that. It's going to take a little time. Now you know it's like, okay, these are the things I need to run this unmodified app. Maybe you don't like running this root, but then you need to teach your app how not to require root to run, and then you can change that in Kubernetes. So get a nice clean baseline. Don't overthink it. If you need config files, don't try to rewrite your app to use environment variables.

You don't have to. You can have Kubernetes map in those configuration files by using something like a config map or a secret. So I recommend most people start with a clean baseline, take an iterative approach, keep either removing features from Kubernetes because you can turn things off you don't need in terms of when you deploy, and then mount in things as you need them to get to that clean baseline. So that's the route I would take, depending on how well you know your application and its contract with the underlying operating system. Is how much better you're going to do with Kubernetes. And don't be afraid to ask people. When you ask, say, hey, this is what I'm doing on my VM. What do I need to do inside of Kubernetes? I don't want to make a lot of changes to make it work there. Then you'll see what you can change over time. So that's a really great question. It's doable because it's just Linux. Again, next question. When will KH achieve write nothing, deploy nowhere? Haha, this is a great joke. So I have a framework called NoCode. And the story behind this joke is that I think people start writing code way too fast before understanding the problem.

And then sometimes they just create another mess and another mess. We've been looking for the holy grail. I want to write the least amount of code as possible in order to get my application up and running. And this is where things like serverless platforms. And when you think about Kubernetes, when you combine Kubernetes with like a service mesh and other little utilities around your app, then we have a chance to simplify the app and get more from the platform. Right. We've gone through this evolutionary loop with application servers. But that's the idea. Either we have big apps with lots of client libraries or we have smart platforms with very tiny, slim applications. So will Kubernetes ever get there? I don't know. But I think people will definitely try. All right. So we'll go to the next question. What is the performance overhead of Kubernetes? Is it relevant for low latency applications that needs response time around a few milliseconds? So this is kind of the foundation of the previous question. So there's an overhead of Kubernetes in terms of resources that it uses.

And if you really want to get into the low-level details of the kernel, if you have a Kubernetes agent that's like checking and pulling containers, it's going to be doing some IO. It's going to be taking some CPU cycles. If you have something like FluidD that's handling your logging, again, it's going to be doing some IO. It's going to be taking CPU cycles. Those little utilities contribute to the overhead, including the Kubernetes agent itself, and maybe something like Docker or RunC, depending on your container runtime. So those collection of things will need some amount of CPU. So that's going to be overhead in terms of operating system resources required to even just run the platform. But then once your application is pulled from a container instance image and then it's run on a node, you're back into normal. Kernel process land. So for most people, if you're using a modern kernel, you're already using container technology. You're just running in the default namespace. So the kernel already has this level of accounting unless you compile those flags.

So by default, most people are already running under a kernel that has container features. So if you don't explicitly do something like Docker does, where you create a new network namespace, a new mount namespace, and some of these C groups to protect the process and to create some isolation, then you just run in the default namespace. So you're already incurring some overhead no matter what you do. So when you think about it. Let's say I was running some low latency networking application. One thing that I've done in Kubernetes is I might decide to run those applications with a flag called network equals host. So when you think about network equals host, you're saying I don't want to be part of Kubernetes overlay network. I don't want to go through IP tables. I don't want any of these extra networking features Kubernetes has. I want raw performance. So in that case, you still can package in a container. You can still use Kubernetes to schedule that container to some machines, but then you can tell Kubernetes, do not do any of the fancy container networking. I want to just go very low to the low-level OS and just bind to the port in the default network namespace.

Then you're going to be back into familiar territory. And depending on your cloud provider, like in GCP, we have a thing called IP aliases or VPC native networking. And what that does is even when I'm using Kubernetes, you know, subnets like, you know, IP per pod. What we do then is we make sure that that IP that's assigned to that container is a first class IP, meaning. We no longer need to go through IP tables when we're coming from a load balancer, and that will help you preserve some of that network performance that you're talking about here. So, again, you're basically telling Kubernetes to run your process on Linux. And if you don't want some of the tradeoff for conveniences, just turn those things off. If you find that they're getting in the way. All right. Do you always need to start K8s the hard way? Absolutely not. So this is a reference to some documentation that I wrote years ago, because I think part of having a successful infrastructure platform is knowledge, not just a good installation script, not just a good managed service.

You actually need to have the administrators of that platform be knowledgeable about how those components fit together. So even though GKE, Google Cloud does a great job of managing Kubernetes for you, if there's ever a problem, you may have to troubleshoot it. So it really helps to understand how the kubelet and docker and all the components fit together. So I think the hard way is good education for anyone that's going to be using a Kubernetes and may have to troubleshoot it at some point in the future. Education can really save you there. But for a lot of people, there's probably going to be no need to really try to build it from the ground up, just like Linux, right? We go get a distro or now you buy a laptop where it's already pre-installed and does automatic updates, right? I'm talking to you now on a Chromebook that's running Linux, and I've never patched it from scratch before. So, you know, Just keep that in mind where we are. All right, so the next question. How would you handle the dev lifecycle from local case stack all the way to prod with deployment stack?

Service mesh deployment artifacts. So this is where it gets super tricky. And I think for years, people have been trying to create production on their laptop. Right. So back in the day when when everyone was doing hypervisors, people running Vagrant on their laptop just to run the services like, oh, I'm going to spin up 10 VM so I can mimic production. The truth is you will never mimic production. You can get close, I guess, but you will never mimic production. So now you have to ask yourself, what is the true interface required to do development? So I'll tell you how I work, and I'll tell you how some other people work. The way I work is I don't use containers at all on my laptop. Never, ever do I use containers. If I want to run a database, I'll run like a, you know, I have a Chromebook here, so I have a thing called Kistini where I can run MySQL or Postgres, and then I use multiple virtual databases if I need to do that. I write in Golang, so this is a little bit easier for me in terms of the dependencies. No GS helps here, but it's a little bit more challenging when you talk about Ruby and Python if you're not using something like virtual F.

But since I write in Golang, when I'm writing my services, I actually just use SystemD to run them in the background, meaning since I do leverage service mesh or serverless platforms quite a bit, I really focus on just the binaries, making API. calls to each other. Now, when it's time to go to staging, that's where I start to leverage Kubernetes. And staging, I will let the CICD system build the container because that's what's required to go to staging. That's when I start to bring in the Kubernetes config files. I can bring in service mesh to add all the security that I don't have natively in my app. I'm relying on a sidecar to do that. So what most people would do is try to make staging reflective as much as possible of production, even though you'll never be a 100% match. But you can say, well, I'm using Istio in staging. I'm using all of these Kubernetes features in staging. So when I'm developing, I'm making sure my unit tests are great. I'm making sure the apps will talk to each other when they talk natively directly to the exposed sockets and making sure the configuration files are clean.

But then when it's time to go into the staging environment, there is no way to get around integration testing. So that's where I start to do all of this Kubernetes. So when I think about integration tests, which are different than unit tests, that's when I'm starting to make sure that my app runs in a container, runs package, runs with the right Kubernetes configuration. I work all that out in staging. And then those artifacts, those are the things I promote up to production. I can't really promote things from my laptop because Minikube is not going to be configured like a Kubernetes cluster, even in staging. So for me, I've been able to survive all these transitions from Vagrant to Docker to now Kubernetes by just staying with the fundamentals. If I write an app and it has HTTP port, I just hit it with curl. I don't need to do all of this container stuff, just start doing development. So that's just how I work. And some people go and try to put Kubernetes on everyone's laptop and install all the components and hope for the best. You have to do what's right for you. But to me, I like fundamentals. I like CICD to do integration testing.

And then I promote from there. That's kind of my advice there. There's no one way for everything. I just wanted to share how I do it. All right, next question. How do you implement visibility and monitoring, not only on the hardware part, but on the process and the deployment. So I think when you're in the cloud or using some hypervisor, it's going to give you a lot of underlying metrics and performance about how the hardware is behaving. So that's good. And then your operating system is going to give a little bit more visibility. The kernel is going to expose a lot of features and you can get things from there. We've been doing that for years. But now we're starting to get to like what can Kubernetes add to the process visibility. So Kubernetes itself has lots of monitoring endpoints. It implements like cloud native standards such as Prometheus slash metrics. You hit that. It's going to give you all these stats around latency, how long it takes to schedule a container, how much memory your Kubernetes component is using. That's great. We kind of standardized there. In terms of logging, again, we try to standardize with structured logs. So when Kubernetes is doing anything, it's always putting out these logs, okay?

And then once you start to get to the app level, so Kubernetes is going to give you a lot of metadata for the process, right? This process has this IP. This process is running on this machine. This process is using this much memory. So Kubernetes is going to give you kind of half the battle by harvesting data it can get from the kernel and mixing it with its own metadata, right? So that's great. But then for your app layer, you have to start thinking about health checks, right? So slash health. Then I can use any metrics or monitoring tool that can ping the health endpoint and make sure that that backend is healthy. And then I can propagate that up to other tools. Again, you can use Prometheus. Prometheus inside of your own app. So you can import the Prometheus libraries, you know, collect metrics. And then whenever something comes by to collect those metrics, you can implement your own slash metrics endpoint. So now you can have visibility up and down the stack. And then using something like your container ID that you get from Kubernetes, you can start to correlate. All of this data between all of the components in the stack.

So now that we have all of this metadata, now we have all of these standard interfaces, we now are able to produce things that can give us end-to-end visibility. So whether you're using Stackdriver or Datadog, and there's some new startups out there like Pixie Labs that are actually taking data from the kernel via eBPF, you can now start to correlate all this data and paint any dashboard graph that you want as long as you kind of know what you're looking for. So I can't think of a better time than now to be able to do this. Okay. All right. Casey, we will need to switch to the networking part. So this global session will end and now everybody has to go to his table and then you will have five minutes by table and somebody will take you from table to table. Thanks, Kelsey. It was very nice for me to meet you and to be able to have this chat with you. See you. Awesome. Thank you. Maybe in Paris one day.