[00:00.000 --> 00:00.880] Anybody? [00:01.180 --> 00:01.300] Nice. [00:01.420 --> 00:02.120] Did you perform? [00:02.800 --> 00:03.880] How was it? [00:10.040 --> 00:12.340] Bring your talents tonight, okay? [00:13.880 --> 00:15.740] And just mute your phones. [00:15.920 --> 00:17.300] These are the basic ones for today. [00:17.700 --> 00:18.120] Enjoy. [00:18.540 --> 00:19.960] You know, ask questions. [00:20.980 --> 00:26.160] So today's talk is, or this morning is, An Engineer's Guide to Linux Kernel Upgrades. [00:26.400 --> 00:28.080] We have Ignat Korchagin. [00:28.140 --> 00:29.200] He's going to present. [00:29.200 --> 00:31.480] And then we want to leave some time for questions. [00:31.680 --> 00:33.720] Join the Matrix Chats if you haven't already. [00:34.120 --> 00:40.360] The cool thing about it is that you can keep the conversation going as long, you know, once the conference ends. [00:40.600 --> 00:42.960] So give him a hand and let's get started. [00:48.870 --> 00:49.490] Good morning. [00:50.470 --> 00:56.230] Thank you for finding the strength to come to my lecture today after yesterday's celebrations. [00:57.050 --> 00:58.390] My name is Ignat. [00:58.470 --> 00:59.730] I'm from Cloudflare. [00:59.950 --> 01:03.090] And today we're going to talk about Linux Kernel Upgrades. [01:03.090 --> 01:05.510] Mostly in production systems. [01:06.090 --> 01:12.330] And hopefully after this talk you will know more ins and outs of Linux Kernel upgrades. [01:12.530 --> 01:14.370] How they are different from software upgrades. [01:14.550 --> 01:21.730] And you will have some tips on how to safely and often upgrade your kernel in production. [01:23.270 --> 01:23.990] Okay. [01:24.270 --> 01:25.830] A little bit about myself first. [01:26.190 --> 01:28.570] So I run the Linux team at Cloudflare. [01:29.150 --> 01:32.570] I'm passionate about system security and performance. [01:32.570 --> 01:34.510] And I enjoy low-level programming. [01:35.390 --> 01:41.050] Kernel, device drivers, boot loaders, and other scary low-level stuff in C. [01:43.170 --> 01:43.750] Okay. [01:44.090 --> 01:48.530] But before we dive in, let's do a small show of hands to wake up. [01:48.990 --> 01:50.830] Like, imagine you're working. [01:50.830 --> 01:52.810] And what would you do in this case? [01:53.010 --> 01:54.510] You get a notification. [01:55.170 --> 01:56.150] Upgrades available. [01:56.830 --> 01:59.570] Who will press the button to apply immediately? [02:01.150 --> 02:01.690] Okay. [02:02.370 --> 02:02.730] Yes. [02:03.150 --> 02:03.250] Yeah. [02:03.610 --> 02:03.890] Fine. [02:04.390 --> 02:05.630] Who will postpone? [02:06.310 --> 02:06.850] Or... [02:06.850 --> 02:08.230] Oh, wow. [02:08.450 --> 02:08.610] Wow. [02:10.110 --> 02:12.210] And this is just your laptops, right? [02:13.210 --> 02:15.570] So, like, the next question. [02:15.710 --> 02:20.830] Like, of all the people who would apply immediately, would you do differently if it was a production system? [02:22.770 --> 02:24.490] Oh, who would not do differently? [02:24.590 --> 02:25.790] Who would still press apply? [02:27.450 --> 02:28.250] Good job. [02:30.110 --> 02:30.590] Right. [02:30.790 --> 02:31.110] So, yeah. [02:31.250 --> 02:36.010] For production system upgrades are a little bit, like, very pain points, right? [02:36.170 --> 02:38.210] So, you actually have two options. [02:38.370 --> 02:40.510] Either remind me later or don't do it at all. [02:42.970 --> 02:45.090] And this is natural, right? [02:45.370 --> 02:48.610] So, human beings are conservative. [02:48.610 --> 02:56.550] And, like, we perceive, especially software engineers, like, we perceive things, changing things as a threat, right? [02:56.750 --> 03:01.510] So, if it works, like, why should we touch it in the first place, right? [03:01.810 --> 03:07.830] But the reality is, like, we perceive different software we may perceive with different levels of threat. [03:08.390 --> 03:13.150] So, with regular software upgrades, yeah, they're bad, right? [03:13.250 --> 03:16.050] They're, like, a threat, a risk to your production system. [03:16.050 --> 03:23.850] And, like, if it goes bad, it kind of, like, will somewhat get you in trouble, right? [03:24.570 --> 03:28.950] But, so, upgrades are monsters, but they're not really that scary. [03:29.110 --> 03:32.710] Yeah, they're ugly, annoying, but, like, we can deal with them, right? [03:33.850 --> 03:37.550] When we mentioned Linux kernel upgrades, so the perception is this. [03:37.550 --> 03:43.650] So, it's, like, an all-destroying monster which can, like, distract the whole planet in, like, five minutes. [03:44.450 --> 03:53.310] And this is, again, this is natural because we kind of know how to deal with bad software upgrades, right? [03:53.490 --> 04:02.310] So, imagine in this case we upgraded a service and, like, let's say if it just keeps crashing, we can roll it back fast. [04:02.310 --> 04:02.910] It's okay. [04:03.110 --> 04:07.950] But let's say it kind of works, but once in a while it crashes, right? [04:08.130 --> 04:09.790] And we know how to deal with that. [04:09.910 --> 04:11.150] So, it's not an end of the world. [04:11.350 --> 04:17.870] So, if you use, like, some kind of service manager, you can tell it, like, please monitor my service and restart it once in a while. [04:19.530 --> 04:22.110] And, yeah, and the job is done, right? [04:22.110 --> 04:26.110] Well, it's not done, but you're kind of, you're not breaking stuff. [04:26.430 --> 04:33.970] You're kind of operating in a degraded mode, but it gives you, like, time to more thoroughly debug things and fix it, right? [04:36.430 --> 04:40.910] But when the Linux kernel crashes, right, everything is down. [04:41.070 --> 04:45.870] You don't have your system, it's not working, and everything is bad, right? [04:45.870 --> 04:52.490] And, therefore, like, nobody likes and nobody wants to risk kernel upgrade. [04:55.640 --> 05:01.420] So, people naturally try to avoid that, especially in production systems, right? [05:01.940 --> 05:08.200] But if you don't do that, you're really missing out, right? [05:08.460 --> 05:15.200] And let's talk about what are the risks of not applying software updates and kernel upgrades in particular. [05:17.000 --> 05:24.920] Well, the first and the most obvious things, your bugs are not getting fixed, right? [05:25.080 --> 05:29.680] Like, let's assume all good intentions from all software developers. [05:29.900 --> 05:32.840] People do not release updates just because they want to. [05:33.460 --> 05:37.540] Yeah, they introduce new features, new code, but as well as, like, fix a lot of bugs. [05:37.740 --> 05:47.180] And, therefore, it is important to keep up with updates, with software updates and Linux kernel updates, in particular, because you want these bugs to be fixed. [05:48.540 --> 05:50.240] And here is some data. [05:50.640 --> 05:58.400] So, Cloudflare now usually follows the latest Linux kernel stable long-term release branch, which is currently 5.15. [05:59.040 --> 06:06.500] And at the time of compiling this presentation, there were 55 bug fix releases in the 5.15 branch. [06:06.500 --> 06:11.820] I will talk about release branches and bug fix releases later in this presentation. [06:12.080 --> 06:16.960] But so far, there are, like, 55 releases in the branch, which was just bug fixes, right? [06:18.200 --> 06:25.880] And this is the data of number of commits, therefore, bug fixes in each release, right? [06:25.880 --> 06:35.800] And it is, out of 55 releases, we have 29 releases with more than 100 commits. [06:36.020 --> 06:39.280] So, somewhat, with more than 100 bug fixes. [06:40.300 --> 06:44.120] And, by the way, these releases happen roughly every week, right? [06:44.740 --> 06:48.200] So, 10 of them have more than 200 commits. [06:48.200 --> 06:56.240] So, and for these high bars here, four releases had 600 commits in one release. [06:56.440 --> 07:08.900] So, imagine if you're not applying a weekly kernel bug fix release, you might be missing out on at least, at least 100 bug fixes into your production systems. [07:10.720 --> 07:16.420] Well, the second thing is you're also missing out on various performance improvements. [07:17.760 --> 07:21.180] This is, again, an example from Cloudflare production systems. [07:23.380 --> 07:28.720] When I talk about performance improvements, I talk here about in a wider sense. [07:29.000 --> 07:34.860] So, like, it's not only about speed, but performance improvements means better resource utilization. [07:34.860 --> 07:42.620] And this was the case for Cloudflare when we migrated from a 5.4 kernel to 5.10 kernel. [07:42.620 --> 07:46.600] So, we, of course, we didn't upgrade everything at once. [07:46.780 --> 07:48.900] We did a limited deployment to compare. [07:49.240 --> 07:57.500] And, like, just upgrading the kernel actually saved us, like, around 5 gigabyte of RAM per server. [07:58.260 --> 08:04.880] Because nice folks from Facebook, like, optimized the memory management system in one of the major kernel releases. [08:04.880 --> 08:06.680] And we just got it for free, right? [08:07.160 --> 08:17.920] And, like, in Cloudflare scales, where we have, like, around 300 data centers across the world, 5 gigabyte of RAM per server is a massive saving. [08:21.990 --> 08:29.190] Also, if you're not applying the releases, you have the accumulating change delta problem, right? [08:29.190 --> 08:34.110] So, this is kind of like the same data I presented several slides before. [08:34.410 --> 08:39.570] Number of commits per release, but from a different view. [08:39.710 --> 08:43.430] So, this is a total commits per release since release zero. [08:43.690 --> 08:48.610] So, in this graph, like, release one has some X amount of commits. [08:49.650 --> 08:54.830] Release two shows you number of commits from release one plus release two and so on. [08:54.830 --> 09:02.010] So, it's just a different view of the data, but it kind of allows you to measure the commit delta, right? [09:02.510 --> 09:15.830] So, let's say you're currently running on 5.15.16 version, and you're considering to upgrade to 5.15.32 version, right? [09:16.330 --> 09:17.770] Like, 32 version is the latest. [09:17.910 --> 09:24.630] So, your kind of commit delta is 2,196 commits, okay? [09:27.430 --> 09:37.450] And we... it would be natural to assume that the number of commits you are accepting in production is proportional to your risk, right? [09:37.650 --> 09:47.410] And it's... yeah, the more changes you apply, the more risk there is that something will break after the upgrade. [09:47.410 --> 09:54.970] So, let's say, for whatever reasons, you want to postpone the upgrade, right? [09:55.190 --> 09:56.150] And you wait. [09:56.710 --> 10:00.050] You wait twice as long as you originally intended. [10:00.330 --> 10:12.610] And once you figure out that you're ready to upgrade, you're now upgrading from 5.15.16 to 5.15.48, because you waited twice as long. [10:12.790 --> 10:20.190] And now your commit change delta is 5,436 commits in this case, right? [10:20.930 --> 10:28.610] So, now we can calculate, like, the difference or the relative change commit delta. [10:28.610 --> 10:34.810] And in this particular case, it will be almost 2.5, right? [10:35.030 --> 10:53.650] And this is an important number because it will show you, in this particular case, that for a 2x delay, you get 2.5 higher risk of a breaking change, right? [10:53.650 --> 11:01.890] So, your risk of delaying the upgrade grows faster than the time you delay, which is very interesting. [11:02.130 --> 11:09.890] So, therefore, like, small regular releases and keeping your change delta small gives you lower risk. [11:12.470 --> 11:17.350] When you're not applying updates, you're missing out on security vulnerabilities, right? [11:17.350 --> 11:19.470] And this is from bug fixes. [11:19.650 --> 11:24.030] These are another types of fixes which are introduced in every kernel weekly release. [11:24.930 --> 11:29.710] This is, again, data from the 5.15 kernel stable branch. [11:30.490 --> 11:33.830] I didn't have the data for the latest releases. [11:34.110 --> 11:37.010] The data is available after .54. [11:39.050 --> 11:52.590] And the important message here is, out of 54 releases in this graph, 40 have at least one CVE patched, right? [11:52.810 --> 11:57.370] So, almost every release has at least one CVE patched. [11:57.470 --> 12:00.830] And three of those had more than 10 CVEs patched. [12:00.830 --> 12:09.890] Now, imagine if you are not applying a particular kernel upgrade, you have CVEs not patched, which is... [12:09.890 --> 12:13.190] And these are known public CVEs, by the way. [12:13.350 --> 12:16.990] So, these are not some kind of zero days and some research, right? [12:17.550 --> 12:24.290] And on top of risking your security, you have also compliance risks, right? [12:24.290 --> 12:34.990] So, once the fix has been published, if your production system is compliant to something, you will likely have this requirement. [12:35.250 --> 12:51.590] I mean, I took the example from the PCI DSS certification, but, like, other compliance systems have similar requirements, that for a known patch, you have a limited timeframe where you need to deploy to production. [12:51.590 --> 12:56.610] And, like, for PCI DSS, for example, it's for critical components. [12:56.810 --> 13:01.450] And Linux kernel most likely is a critical component because it's an operating system, right? [13:01.510 --> 13:02.750] Your production operating system. [13:03.050 --> 13:09.610] And, like, for PCI DSS, you will have only, like, one month since the patch was introduced to release it to your production. [13:13.200 --> 13:16.200] Who here didn't heard about Equifax? [13:19.020 --> 13:19.560] Right? [13:19.560 --> 13:25.740] So, Equifax didn't patch a known vulnerability, and it got exploited. [13:25.740 --> 13:32.060] And it got exploited with very severe consequences, financial, for business, and everything. [13:32.480 --> 13:32.580] Right? [13:33.700 --> 13:42.560] So, remember, every weekly kernel release has at least one CVE patch, then your clock starts ticking when it gets released, right? [13:42.560 --> 13:58.340] And, therefore, you know, like, I remember, like, 10 or 20 years ago when you go to the sysadmin forums online, and, like, people, like, various sysadmin boasting, like, oh, my uptime is two years, my uptime is six years. [13:58.440 --> 14:00.180] So, this is not great anymore. [14:00.420 --> 14:06.100] Like, if your uptime is more than 30 days, most likely you're vulnerable and you're breaking compliance, right? [14:10.410 --> 14:11.050] Okay. [14:11.470 --> 14:32.130] So, now that we talked about the risks of not applying that break, let's talk about the common anti-patterns, which, like, I've encountered in my own experience working in Cloudflare and other companies as well as I've seen in other even big companies which manage production systems. [14:32.730 --> 14:40.250] So, how do they approach Linux kernel upgrades and why they are wrong, right? [14:42.370 --> 14:47.830] So, yeah, in many companies you have, like, some kind of SRE organization or production engineering. [14:47.830 --> 14:50.430] Like, in Facebook, they're responsible for production. [14:50.430 --> 14:58.750] And oftentimes, it's a team which, like, manages the Linux kernel distribution is a different team. [14:58.910 --> 15:03.190] So, you have to negotiate the upgrade with the production engineering team. [15:03.510 --> 15:11.130] And the problem is they try to apply the common patterns they have for any regular software to the Linux kernel, right? [15:11.330 --> 15:19.590] So, remember, Linux kernel is released, bugfix released weekly, and they said, like, okay, we need to upgrade every week. [15:19.590 --> 15:25.470] And, like, every time I come to them, they would ask, okay, but do we need to upgrade? [15:25.590 --> 15:27.730] Like, have you reviewed the changelog? [15:27.830 --> 15:31.210] Which things from the changelog are actually applicable to us? [15:31.310 --> 15:32.990] Can you justify the upgrade? [15:34.390 --> 15:37.110] And for the Linux kernel, it's actually not possible. [15:37.110 --> 15:39.550] We're going back to this graph, right? [15:39.830 --> 15:49.390] So, more than a half of the weekly releases have more than 100 commits, 100 changes, right? [15:49.690 --> 15:57.550] So, and you just expect us, like, a small kernel team to review all of them, like, continuously. [15:57.910 --> 16:01.730] Like, we'll be doing just that and, like, not anything else. [16:01.970 --> 16:14.610] And, moreover, because of the sheer volume of the commits and changes coming into the release ground, there is a very high chance there is something from that 100, or even 600 is really applicable to your system. [16:14.770 --> 16:16.150] So, you don't have to review it. [16:16.310 --> 16:17.630] You just have to take it. [16:17.850 --> 16:20.990] And, like, most likely, there will be something that you need, right? [16:24.190 --> 16:40.610] When you come with an upgrade, like, because serious vulnerabilities have been publicized and, like, everyone is screaming, they would still ask us, like, okay, but is this vulnerability actually exploitable in our systems? [16:40.610 --> 16:48.270] Like, I had, like, these questions, like, before many times in Cloudflare where we have a, like, privilege escalation published. [16:48.470 --> 16:52.670] But, for example, they say, we don't run, like, third-party code on our servers, right? [16:52.790 --> 16:55.410] Like, if we, like... [16:55.410 --> 17:06.370] Like, yeah, we really can't protect from an insider attack, but, like, we don't have a possibility of someone running their third-party code on the system and getting... [17:06.370 --> 17:09.230] So, is this security vulnerability actually exploitable? [17:09.230 --> 17:11.130] Like, can you prove that it is exploitable? [17:12.230 --> 17:15.710] But the problem is, this is the wrong question to ask, right? [17:16.170 --> 17:20.310] And, like, think of it from this perspective. [17:20.330 --> 17:29.990] So, you're running a Linux kernel with a non-vulnerability which you really don't know if it's exploitable or not, right? [17:29.990 --> 17:36.670] And let's consider the potential attacker's perspective, the person who broke into Equifax, right? [17:36.930 --> 17:39.410] So, who is this person, the attacker? [17:39.590 --> 17:43.790] The attacker is a person who is highly motivated to break into the system. [17:43.950 --> 17:46.210] This is their primary source of income. [17:46.870 --> 17:48.890] They know you're running a vulnerable... [17:48.890 --> 17:51.690] They potentially know you're running a vulnerable system, right? [17:51.690 --> 18:01.570] And they spend exclusively, almost 24-7, all their resources, time and effort just to break into your system and find a successful exploit, right? [18:02.970 --> 18:12.630] Yeah, so the attacker is very, very determined to get into your system and does nothing else just to do it, right? [18:14.330 --> 18:17.510] But we're asking this question not for the attacker. [18:17.790 --> 18:25.470] We're asking this question for a security engineer, a Linux kernel engineer, or someone who is reviewing these patches. [18:26.070 --> 18:27.590] They are different, right? [18:27.790 --> 18:30.570] So, they are highly motivated to go home on time. [18:32.310 --> 18:36.130] And most likely, they are not, like, reviewing just this one patch. [18:36.230 --> 18:42.230] They are reviewing, like, many patches, not from the Linux kernel, from all the production software that you are using. [18:42.590 --> 18:53.150] And they also have, like, other competing priorities, like security folks, build security architecture, create tools, like do consulting, many, many, many other things, right? [18:53.330 --> 18:55.510] So, it's kind of like a multitasking person. [18:57.010 --> 19:08.330] So, the disconnect here is that that person on the left is the person who will most likely has the answer, is this exploit applicable to you? [19:08.330 --> 19:11.550] But you cannot ask them because you don't know them, right? [19:11.730 --> 19:12.810] And they're bad. [19:13.130 --> 19:20.790] But you're asking a question of the wrong person who cannot really devote that much time to produce this answer. [19:21.670 --> 19:36.790] So, therefore, it's not really correct into assuming that, you know, like, if a security researcher would say, yeah, this vulnerability is not applicable, it's not really applicable. [19:36.790 --> 19:40.410] The safest course of action is to take and patch it, right? [19:40.610 --> 19:41.110] If it's known. [19:44.850 --> 19:54.150] Another anti-partner I see from SREs and production engineers, they say, like, okay, this is a kernel very, like, scary thing. [19:54.390 --> 19:56.750] It breaks all the servers if it goes wrong. [19:56.750 --> 20:02.350] So, let it soak for one month somewhere in Canary to ensure it's stable, right? [20:03.310 --> 20:06.570] But again, why it's an anti-partner? [20:07.110 --> 20:15.510] Because the more you soak, the more you delay the upgrade, the more changed delta you accumulate, right? [20:16.910 --> 20:20.210] Secondly, you have the security portion of it. [20:21.190 --> 20:29.490] And the more you delay the upgrade, the more you're running with production system with more than one CVE and patch. [20:29.690 --> 20:35.890] And, like, the thing is, this one month in this example was not an arbitrary number. [20:36.090 --> 20:41.250] People usually somewhat come up with two weeks or one month because they think it's, like, enough. [20:41.470 --> 20:43.130] Like, they perceive it as enough. [20:43.130 --> 20:48.090] But remember, CVEs are being patched in the kernel every week. [20:48.170 --> 20:51.310] So, your soak time cannot be larger than that, right? [20:51.470 --> 21:11.890] Because you're not only risking the security of your system, you're, again, risking your compliance because if your soak time is one month, there is no way you're going to deploy a known patch vulnerability to production in one month because you'll soak it in one month in Canary and then you'll have another one month to release it, [21:12.150 --> 21:12.350] right? [21:12.350 --> 21:13.950] Or two weeks or whatever. [21:17.190 --> 21:20.050] So, why people come up with the soak times? [21:22.030 --> 21:29.950] Like, in my experience and my opinion, it comes down to the fact that we don't know what we're looking for. [21:30.030 --> 21:35.490] When I say, like, ask people, why do you think one month is enough or two weeks is enough? [21:35.570 --> 21:38.830] Like, because they will just wait and see what happens. [21:38.830 --> 21:42.190] But they're not, they don't know what they're looking for. [21:42.190 --> 21:50.370] Instead, you should have metrics will tell you, like, is this kernel behave the same as the previous one and it's acceptable to you, right? [21:50.910 --> 21:53.670] We also don't know our workload. [21:54.490 --> 22:04.470] So, oftentimes, and not SREs, but other engineering teams from Cloudflare come to me and ask, like, will the new kernel break my software? [22:04.470 --> 22:05.970] I said, I don't know. [22:07.030 --> 22:08.750] What does your software need? [22:09.110 --> 22:23.810] And, like, when you follow the 5Y rule questions, you start understanding that, like, okay, some teams, they write, like, a key value store and, like, their performance is heavily dependent on the kernel page cache. [22:23.970 --> 22:27.330] So, like, kernel page cache performance is very important to them. [22:27.330 --> 22:36.640] Other teams are writing, like, like, we have Cloudflare workers, which is a third-party code execution system. [22:37.510 --> 22:40.170] So, like, scheduling and CPU performance is important. [22:40.430 --> 22:48.270] So, every workload underneath has this, like, one or two kernel features they require. [22:48.270 --> 22:53.770] Like, the other team we have, like, they heavily rely on Linux network namespaces. [22:53.890 --> 22:57.650] So, Linux network namespaces is the needed feature of them. [22:57.750 --> 23:01.470] And for all these features, you can actually write tests, right? [23:01.610 --> 23:13.810] And you can exercise them and, like, compile a test suite of kernel pre-production testing, which can have unit tests, integration tests, performance tests. [23:13.970 --> 23:16.130] But the old thing, I call it the acceptance test. [23:16.130 --> 23:22.650] So, now I'm telling all the teams which come to me and say, like, oh, will the kernel break the build? [23:22.810 --> 23:30.390] Or the most hilarious I heard, like, oh, can we have a say if you're allowed to release a new kernel or not? [23:30.510 --> 23:34.930] And, like, we have, like, thousands of engineering teams in Cloudflare. [23:35.050 --> 23:41.010] And if I allow every one of them to have a veto, I will never upgrade the kernel because some of them will say no. [23:41.010 --> 23:48.770] But I said, you can have a veto if you write a test for us in our test framework, which we provide, and it will fail. [23:48.870 --> 23:52.030] If your test fails, we'll not release the kernel until we debug it, right? [23:53.490 --> 23:57.490] But to write tests, they have to learn their workload and what do they need from the kernel. [24:00.660 --> 24:08.160] Yeah, final thing I saw, like, in the early days of Cloudflare and which we successfully removed now is... [24:08.590 --> 24:10.160] When you come to people, they just... [24:10.630 --> 24:12.660] It goes back to the beginning of my presentation. [24:12.700 --> 24:15.780] They just perceive kernel as too scary, too risky. [24:16.040 --> 24:22.000] And they start, like, okay, well, like, kernel is this, like, mega beast, right? [24:22.000 --> 24:26.600] So let's have a different approval process for release. [24:26.800 --> 24:31.920] Like, if regular software will require one approval, for kernel we'll have, like, three approvals from different teams. [24:34.040 --> 24:35.520] Which is actually nonsense. [24:35.820 --> 24:41.880] So what if I told you that the kernel deploy is safer than any other software, right? [24:42.880 --> 24:46.220] And, again, I give it in the Cloudflare example. [24:46.500 --> 24:48.900] So this is, like, Cloudflare network today. [24:48.900 --> 24:57.080] All these blue dots in the world are data centers and each data center can have, like, many, many, many service, right? [24:57.720 --> 25:01.860] So how does a regular software update look for us? [25:01.860 --> 25:04.100] So we have, of course, it's automated. [25:04.260 --> 25:07.100] We have configuration management system doing all of that. [25:07.240 --> 25:20.200] But when a new team, a team, like, a team who manages our web server, Nginx, for example, releases a new version, the way how it works is the configuration management sees there is a new version available. [25:20.260 --> 25:27.920] It upgrades the software package on each server and then does a server restart, service restart. [25:28.260 --> 25:35.260] So depending on if the service is critical or not, you can be a graceful restart or non-graceful restart. [25:35.420 --> 25:35.920] It doesn't matter. [25:35.960 --> 25:37.580] But it's still a service restart. [25:37.740 --> 25:39.120] So the new code takes over. [25:40.500 --> 25:59.880] So the gist of that process is if you don't put explicit safeguards to deploy new code, like, in a stage and slow down manner, a typical Nginx release, if it's bad, it can break the whole network. [25:59.880 --> 26:03.340] And unfortunately, Cloudflare learned it the hard way. [26:03.340 --> 26:15.320] So we had a couple of, like, global outages because a bad software deploy got spread across the network too quickly and almost took down the whole network. [26:17.580 --> 26:19.100] Linux kernel, for example. [26:19.200 --> 26:25.460] The biggest blessing and a curse of the Linux kernel upgrade, it requires a system reboot. [26:26.060 --> 26:33.080] Unless you do, like, live patching, but then you're crazy and, yeah. [26:33.720 --> 26:36.900] So, yeah, kernel upgrade requires a reboot, right? [26:37.480 --> 26:42.300] And reboot, again, for us, it's all automated, but it requires more steps. [26:42.540 --> 26:45.320] You need to drain the traffic from the server. [26:46.000 --> 26:52.400] You need to put it out of production, means silencing the monitoring and alerts because the server will be down for some time. [26:52.540 --> 26:54.120] Then to do actual reboots. [26:54.340 --> 26:56.960] Reboots are not fast, especially on the servers. [26:56.960 --> 27:04.660] Then when it's actually booted, our configuration manager steps in and reconfigures the server. [27:04.960 --> 27:08.440] And sometimes it takes, like, tens of minutes. [27:09.960 --> 27:15.560] Then when the server is configured, we run the acceptance test and then we put it back into production. [27:17.780 --> 27:27.320] And because nobody's crazy and, like, we're not crazy, so you don't really reboot all the servers at once, right? [27:27.440 --> 27:33.580] Like, if you have a kernel upgrade, I cannot see a process which says, like, let's pull down the whole network and do a reboot. [27:33.580 --> 27:36.420] So you will naturally have this process, right? [27:36.480 --> 27:38.060] You will reboot servers one by one. [27:38.400 --> 27:42.040] Or, like, in our case, we do it in batches to speed it up a little bit. [27:42.160 --> 27:48.740] But the thing is, like, the kernel deploys inherently slow-paced and gradual rollout. [27:49.180 --> 27:56.800] So you have these safeguards out of the box if you're managing your production sanely, right? [27:56.800 --> 28:01.060] So it gives you plenty of time. [28:01.220 --> 28:10.840] So, like, we did deploy bad kernel upgrades, but we have noticed after, like, two or three servers rebooted and it had almost no visible impact on our network. [28:13.900 --> 28:17.320] Did I convince you that Linux kernel upgrades are safer? [28:19.740 --> 28:20.300] Okay. [28:21.520 --> 28:23.740] So now let's a little bit... [28:23.740 --> 28:26.480] Like, we talked about the risks of not applying kernel upgrades. [28:26.480 --> 28:31.220] We talked about what problems you might have. [28:31.900 --> 28:33.980] Let's talk about the Linux kernel releases. [28:34.180 --> 28:41.140] One of the other problems I've encountered that people just perceive every kernel release as the same. [28:41.320 --> 28:42.580] It's a big, scary thing. [28:42.700 --> 28:47.760] But if you learn the kernel release process, you will see that not all releases are created equal. [28:47.900 --> 28:52.680] So some releases are safer to apply and some releases require more testing. [28:54.900 --> 28:57.760] So kernel version numbers look like this. [28:57.980 --> 29:01.920] So you have, like, a number dot, another number dot, and another dot. [29:02.400 --> 29:07.280] For example, like, 5.15.32, right? [29:07.400 --> 29:10.220] And it kind of looks like a semantic versioning system, right? [29:13.040 --> 29:15.400] This is the biggest mistake everyone makes. [29:15.540 --> 29:18.180] This is not a semantic versioning system in the kernel. [29:18.760 --> 29:22.860] So one thing you have to remember, kernel does not follow a semantic versioning system. [29:23.480 --> 29:25.660] Instead, it just has two components. [29:26.700 --> 29:32.040] The first two numbers, they usually call it a major or stable kernel release. [29:33.260 --> 29:36.420] And the final, and not major, minor. [29:36.760 --> 29:38.800] Like, two numbers designate one thing. [29:38.900 --> 29:40.580] It's a major kernel release, right? [29:41.240 --> 29:45.440] And the second number is our bugfix releases. [29:45.920 --> 29:52.060] There is no established terminology whether I should call it patch releases, but I call them bugfix releases. [29:52.260 --> 29:55.960] So these contain only bugs and security fixes. [29:56.680 --> 30:06.560] Another important thing to understand is these bugfix releases will never contain new features or subsystem rewrites. [30:06.680 --> 30:09.220] So they're usually quite, quite safe to apply. [30:09.360 --> 30:11.420] They only fix bugs and security vulnerabilities. [30:14.540 --> 30:18.980] So how does kernel release flow works in general? [30:19.280 --> 30:26.740] So the main line, the main bleeding edge kernel code is maintained by, still maintained by Linus. [30:26.740 --> 30:29.860] It's, let's call it the main branch. [30:30.440 --> 30:42.760] So the features are actually developed in various other branches from subsystem, managed by subsystem maintainers. [30:42.920 --> 30:46.160] So for example, there is a driver's subsystem. [30:46.320 --> 30:47.820] There is a memory management system. [30:48.060 --> 30:48.680] There is networking. [30:48.940 --> 30:50.760] And these are like other branches. [30:50.920 --> 30:52.620] And features are being developed there. [30:52.620 --> 30:57.580] And then the Linus like pulls from these branches these features like once in a while. [30:58.440 --> 30:58.780] Okay? [30:59.440 --> 31:09.020] And at some point, when Linus considers that his branches is stable, he cuts a stable release. [31:09.240 --> 31:13.400] And the way how they do it, it branches out from the main branch. [31:13.400 --> 31:20.580] So they create a dedicated branch, they call it a stable branch, for major Linux kernel release. [31:20.720 --> 31:23.500] So we have like 5.10, 5.11, 5.12. [31:23.980 --> 31:27.360] And this usually happens every 9 and 10 weeks. [31:29.220 --> 31:32.700] And then the stable branches leaves. [31:32.960 --> 31:40.940] And at some point, a tag is created on the stable branch, which designates a bugfix release. [31:41.660 --> 31:47.380] So 5.11.1 is a bugfix release on the 5.11 stable branch, right? [31:47.900 --> 31:51.180] But how are actually bugs getting there, right? [31:51.440 --> 31:55.060] They don't immediately get into the stable branch. [31:55.060 --> 32:04.480] So the process is if you find a bug in subsum system, you actually usually fix it on the subsystem maintainers tree. [32:05.040 --> 32:07.000] Then eventually that bug... [32:07.000 --> 32:08.820] But you mark it as a bugfix, right? [32:09.680 --> 32:15.480] And then eventually this bugfix gets pulled in by Linus into the main branch. [32:18.020 --> 32:23.780] And stable branches cherry-peak these commits onto themselves. [32:23.780 --> 32:25.080] And when... [32:25.080 --> 32:29.760] So this is where new features are never introduced into the stable branches. [32:29.960 --> 32:34.020] So stable branches do not completely merge the main line Linus branch. [32:34.160 --> 32:36.860] They only cherry-pick bugs and security vulnerabilities. [32:37.260 --> 32:47.800] And eventually when enough bugs have been cherry-picked, a new bugfix release is being cut, and it usually happens like every week with another tag. [32:50.750 --> 32:51.310] Right? [32:51.570 --> 32:52.710] So a new merger... [32:52.710 --> 32:56.870] Major or stable kernel version is released every nine to ten weeks. [32:57.030 --> 32:59.090] And like even there the process is quite rigid. [32:59.290 --> 33:05.930] So they only allow like two weeks for feature development and seven weeks for bugfixing and stabilizing the kernel. [33:07.310 --> 33:09.050] They call it a merge window. [33:10.690 --> 33:17.110] The another thing is because it's not a semantic versioning, the leftmost version means nothing. [33:18.110 --> 33:29.690] So 4.19 upgrade to 4.20 might have more breaking features, like breaking changes or features, than an upgrade to 4.20 to .10. [33:29.710 --> 33:40.110] This is an error like many SRE and production teams make because they say, oh, we used to upgrade from 4.19 to 4.20, but now we are going to 5.0. [33:40.150 --> 33:42.070] This is probably a super major version. [33:42.230 --> 33:43.370] Like we needed to exercise. [33:43.630 --> 33:45.170] And no, it's the same. [33:45.430 --> 33:45.790] Right? [33:47.090 --> 33:54.510] It's just like this leftmost number is incremented when Linus decides to. [33:57.310 --> 34:02.210] And yeah, and bugfix and patch releases are released around once a week. [34:02.710 --> 34:04.830] It's the rightmost version number. [34:04.830 --> 34:06.710] They just cherry-pick bugs. [34:07.210 --> 34:12.650] And while they're propagated through the whole tree, they get some initial testing from the Linus kernel community. [34:12.650 --> 34:14.670] So they're pretty much in good shape. [34:14.810 --> 34:19.410] And therefore, you don't have new features and regressions are quite rare. [34:21.090 --> 34:31.510] Yeah, but it may contain critical security patches and you're almost want to apply them because of the sheer commits and security patches going to bugfix releases. [34:31.510 --> 34:34.430] There is most likely something which is affecting your system. [34:36.510 --> 34:40.330] Yeah, let's talk about a little bit major version and stable releases. [34:40.650 --> 34:47.810] So a stable release, once it's branched out, it's being supported around two or three months. [34:48.170 --> 34:53.910] So supported means these bugfixes and security patches are being backported and cherry-picked into this branch. [34:57.290 --> 35:00.110] But after two or three months, it's considered end-of-life. [35:00.110 --> 35:05.770] So at this point, you will likely might need to evaluate a new major version with new features. [35:06.170 --> 35:22.690] If it's too costly, and like, for example, in Cloudflare, we still want some stability and don't want to evaluate a new major kernel release every two to three months, there are long-term stable releases, which is usually the latest stable release of the year. [35:22.690 --> 35:30.990] And there, the Linux kernel community backports bugfixes and security vulnerabilities for at least two years. [35:31.130 --> 35:42.370] So it provides you more room not to evaluate major kernel releases too often, but still be on top of every bugfix and security patch out there. [35:44.930 --> 36:04.870] Yeah, and this... I encourage you to read this usually overlooked page about the Linux kernel releases on the officialkernel.org site, and it has an explicit paragraph saying that the major... leftmost major number means nothing, and don't bother... don't bother, [36:04.870 --> 36:06.410] like, worrying about it too much. [36:08.550 --> 36:16.150] Okay, so based on what we learned today, we can get some safe and easy production kernel upgrade tips, right? [36:16.810 --> 36:18.710] So the first thing is... [36:18.710 --> 36:26.850] what we can take out from this presentation is don't create a dedicated deploy process for the Linux kernel, right? [36:28.270 --> 36:36.750] Because, as we just learned, kernel upgrades are usually less risky than any other software, on the contrary, which everyone thinks. [36:38.690 --> 36:53.910] And simple stage rollout is usually enough, and kernel upgrades are naturally slow-paced because they require a reboot, and you will have plenty of time noticing a bad deploy and pulling the plug on the deploy process. [36:56.290 --> 37:05.690] Secondly, you have to work with your organization, with your SRE teams, to avoid justifying bug-filled kernel upgrades. [37:07.410 --> 37:11.930] So bug fix releases should be deployed with no questions asked. [37:12.410 --> 37:23.250] And because of the volume of the commits there, and fixes and security patches, there is most likely there is something which is affecting your workflow. [37:23.250 --> 37:29.230] So it doesn't make sense to actually analyze if it's there, and if it's worth supplying. [37:30.310 --> 37:45.210] Because bug fix releases do not contain new features, regressions are quite uncommon, and because of the compliance risks of security features, you should work out, build out the process, which will minimize the required soak times. [37:45.430 --> 37:49.910] Instead, try to move to a metric-driven approach, instead of validating a new kernel. [37:56.010 --> 37:56.570] If... [37:56.570 --> 38:03.810] Again, if validating a new major kernel release is too much trouble for you, consider staying on the long-term branch. [38:06.090 --> 38:08.090] This is what actually Cloudflare does. [38:08.730 --> 38:15.590] So it gives you at least two years of bug fixes security patches, but we don't stay there for two years. [38:15.590 --> 38:22.890] Actually, after a year, the next stable long-term release is already available, and we immediately start evaluating. [38:23.250 --> 38:25.010] So we still have one year. [38:26.210 --> 38:26.770] If... [38:26.770 --> 38:29.310] Like, we usually transition much faster than that, but we... [38:29.310 --> 38:47.470] In the end, we still have one year of time to smoothly transition from the old kernel to the new kernel, still getting bug fixes and performance and security patches, but we also get the newer features earlier than we would wait for another year. [38:47.650 --> 38:51.790] And again, we accumulate less change-delta following this process. [38:51.970 --> 38:56.710] And as we learned before, change-delta is bad for your risk. [38:59.270 --> 39:02.170] Yeah, and implement and improve... [39:02.170 --> 39:09.510] If you didn't already implement, if you did, improve your pre-production kernel testing for major version validation. [39:10.690 --> 39:13.930] And this will help you understand your workload, actually. [39:14.690 --> 39:20.190] You can write tests which exercise various kernel subsistence required by your workload. [39:20.950 --> 39:25.290] And not only these tests will put you in a better place. [39:25.430 --> 39:28.810] If thing goes wrong, if something doesn't work, it will... [39:28.810 --> 39:32.110] They will help you to actually communicate with kernel community. [39:32.930 --> 39:36.250] Like, Linux kernel is huge, and nobody understands everything. [39:36.810 --> 39:38.250] But if you encounter problem... [39:38.810 --> 39:44.190] Like, on the contrary, there are some opinions out there that kernel developers are really... [39:44.890 --> 39:50.710] kernel upstream community is really not friendly, and they say bad things on the main links. [39:50.850 --> 39:55.270] But only because if you come to them with a problem, but they cannot reproduce it. [39:55.510 --> 40:07.010] Once you have a test which is easily reproducible, and you can deliver that test to the upstream community, describing your problem, see how it fails, you will get help. [40:07.250 --> 40:10.090] Like, I've did it many, many times. [40:10.190 --> 40:17.850] If I have a reproducer, if I have a test, like, I immediately get help in any subsystem, even if I am not an expert there. [40:19.430 --> 40:25.230] Yeah, and you should make metrics or data-driven decision, not time-based decisions. [40:25.670 --> 40:35.710] So, try to build your process in a way to decide if the kernel is group based on data and metrics and tests versus just let's run it for one month somewhere and see what happens. [40:37.330 --> 40:46.230] And finally, like, on continuing metrics, metrics monitoring and deploy automation can help with human risks perception, right? [40:46.430 --> 40:58.790] So, apart from having this data-driven approach to decide if a kernel grade is good enough or not, it also will provide you quick early signals about potential regressions, although they're quite rare. [40:59.470 --> 41:00.090] Yeah. [41:00.990 --> 41:03.910] The other point I wanted to make about automation. [41:04.170 --> 41:05.430] One other... [41:05.430 --> 41:08.070] Like, definitely consider kernel deploy automation. [41:08.430 --> 41:26.210] So, one of the early problems we had in Cloudflare because of this kernel risky perception we discussed in the beginning of this presentation, if you come to an SRE, various people perceive kernel grades with different levels of risk, right? [41:26.370 --> 41:42.870] And sometimes, if, like, a more junior employee is tasked to do a kernel grade, they are just afraid, more afraid that things will go wrong and they try all their best to actually avoid it somehow, to come up with an excuse not to do it. [41:43.050 --> 41:47.530] When you have kernel automation, the human factor is not a problem anymore, right? [41:47.710 --> 42:04.430] So, you can roll out your kernel of grades and you don't have to deal, you don't have to put other people in this weird position where they are afraid to do something wrong because it's done for them automatically and there is no decision, human making decision involved in the process. [42:04.610 --> 42:05.450] All this data-driven. [42:07.670 --> 42:12.010] I think that's mostly what I wanted to tell you today. [42:12.010 --> 42:19.510] So, in this presentation, we learned that Linux kernel upgrades are actually not more risky than any other software. [42:19.930 --> 42:24.850] You definitely need to patch early and patch often, especially for the Linux kernel. [42:26.490 --> 42:30.710] Always apply bug fix kernel releases with no question asked. [42:31.330 --> 42:35.390] And when it comes to the Linux kernel, try to understand your workload. [42:35.590 --> 42:44.130] Try to understand your workload requirements to the Linux kernel so you can actually design tests, metrics and monitoring to actually validate these requirements. [42:45.150 --> 42:48.690] And it will help you to stay patched and secure. [42:49.770 --> 42:50.610] Thank you. [42:53.070 --> 42:54.050] Thank you. [42:54.350 --> 42:56.690] We have two questions from the Matrix chat. [42:57.870 --> 43:08.950] The first one is from Band-Aid asking, Fedora core kernel upgrades require two reboots, whereas Ubuntu upgrades can be done in place followed by a reboot. [43:09.170 --> 43:10.170] Any idea why? [43:13.790 --> 43:14.990] No, unfortunately. [43:16.130 --> 43:19.130] So, in Cloudflare, we don't use a distribution kernel. [43:19.130 --> 43:22.410] We build our own kernel directly from kernel.org site. [43:22.410 --> 43:26.650] We call it upstream kernel, and that requires one reboot. [43:27.170 --> 43:33.490] I played with Fedora a long time ago, and at that point, it required only one reboot. [43:33.670 --> 43:36.890] So, I'm actually not sure why now it requires two reboots. [43:37.030 --> 43:38.750] So, sorry, I cannot answer that question. [43:39.710 --> 43:40.270] Okay. [43:40.510 --> 43:42.310] The next one is from Arc6. [43:42.850 --> 43:46.230] You mentioned that live kernel upgrades are crazy. [43:46.230 --> 43:53.410] Can you explain your thoughts on not using Ksplice or Kpatch, particularly for high-severity CVs? [43:54.730 --> 44:09.490] So, my opinion on Kpa, like, various live-patching techniques, if you understand the internals of how kernel works, you will see that they have a very limited applicability. [44:09.730 --> 44:11.030] So, you can only... [44:11.030 --> 44:18.070] So, the way how live-patching works, you have, like, let's say, a vulnerable piece of code, an algorithm, right? [44:18.290 --> 44:27.110] And you kind of, like, put a new one in place and, like, redirect all the code execution from the vulnerable part to the new part. [44:27.190 --> 44:31.670] But that only works if the data structures themselves do not change. [44:31.830 --> 44:38.150] So, Linux kernel internal API is not stable, unlike, like, Windows and other operating systems. [44:38.410 --> 44:44.810] So, like, the structures can be modified at any point of time, even in bug fix releases. [44:45.350 --> 44:51.950] And if you have a CVE which requires you to modify the structure, you cannot apply live-patching anymore. [44:52.070 --> 44:53.310] So, basically... [44:53.310 --> 44:58.590] And that's why I don't like live-patching, so you have to actually know when it works or not. [44:58.970 --> 45:02.910] But, like, many organizations think, oh, we have live-patching, so we can... [45:02.910 --> 45:05.050] we're safe from, like, high-critical CVEs. [45:05.050 --> 45:09.330] No, you can patch with them only, like, a small subset of critical CVEs. [45:10.140 --> 45:21.010] This morning, to stable branches, I just checked, if you... did you hear about the red bleed vulnerability in Intel CPUs? [45:21.170 --> 45:34.350] So, like, patches have just dropped to bug fix releases on all stable branches, supported stable branches in Linux kernel, and you cannot live-patch it because it requires recompiling your code to actually have the mitigation. [45:34.350 --> 45:51.170] So, in my opinion, it's better to build a robust kernel deployment process where you can continuously reboot your machines and deploy kernel, whether it actually wastes a lot of effort making live-patching work for you and only covering a small subset of cases. [45:53.450 --> 45:57.190] We have one more in the chat, but let's turn to the audience here. [45:57.430 --> 45:58.330] Anybody have any questions? [46:01.880 --> 46:02.620] Yes, please. [46:02.620 --> 46:03.260] Yeah, [46:13.150 --> 46:13.610] so... [46:14.090 --> 46:16.630] So, the question was, because there was... [46:17.110 --> 46:23.590] So, the graph shows how many CVEs are usually patched per release, but how many CVEs are introduced. [46:24.210 --> 46:30.770] This goes back to the question of the kernel releases themselves, right? [46:31.330 --> 46:34.370] So, bug fix releases... [46:35.450 --> 46:41.910] Like, CVEs are usually introduced by new code, by new features, some optimizations or whatnot. [46:42.390 --> 46:45.230] So, bug fix kernel releases do not introduce this. [46:45.410 --> 46:47.310] They only patch what is available. [46:47.630 --> 46:59.090] So, unless the bug fix itself introduces a new vulnerability, which happens very rarely, usually there are no new CVEs introduced, only, like, they're getting fixed. [46:59.090 --> 47:04.170] And even if, like, all the things, bad things happen, right? [47:04.270 --> 47:05.610] So, sometimes you get this. [47:05.690 --> 47:11.870] You get a regression, you get a new CVE introduced by a bug fix, but the overall trend is always up, right? [47:11.990 --> 47:20.050] So, you always, by continuously applying bug fix releases, you will end up, usually end up in a better place than you were before. [47:24.100 --> 47:25.220] Any more questions? [47:28.060 --> 47:28.680] Yes, please? [47:35.470 --> 47:38.910] I would recommend rolling your own kernel... [47:39.230 --> 47:44.190] Sorry, the question was, when would you recommend rolling your own kernel versus using distribution kernel? [47:44.450 --> 47:50.670] I would recommend rolling your own kernel when you really know your workload, right? [47:52.130 --> 47:52.850] So... [47:52.850 --> 47:57.510] And, like, in Cloudflare, for example, why we run quite... [47:57.510 --> 48:01.070] Like, distributions kernel are usually older than the upstream ones. [48:01.210 --> 48:07.950] And, like, nature of our products and services, we try to utilize as many new kernel features as possible. [48:08.290 --> 48:15.830] And, like, very advanced things in Linux network namespaces where heavy users often kernel BPF. [48:16.070 --> 48:18.930] So, like, distributions kernel are usually a little slow for us. [48:19.010 --> 48:25.290] And sometimes, for example, you need a particular behavior from a particular subsystem, and, therefore, running your... [48:25.290 --> 48:28.290] You can get the improvements faster and, therefore... [48:28.290 --> 48:30.190] And sometimes you want to make some tweaks, right? [48:30.290 --> 48:44.030] Because the kernel itself, like, is a good generic piece of software which should be good enough for small IoT devices as well as high-performance systems and, like, servers in the cloud. [48:44.030 --> 48:45.970] But sometimes they make trade-offs. [48:46.210 --> 48:54.770] And, like, you would get the benefit of, like, removing the trade-off they made for IoT device to make your, like, server workload perform better. [48:55.470 --> 48:58.630] But that requires you to knowing what you need from the kernel. [49:01.520 --> 49:02.280] Yes, please. [49:09.960 --> 49:13.680] The question was, is cloud for Linux distribution based on another distribution? [49:13.940 --> 49:28.160] So, our production distribution is based on Debian, except for the fact that we don't run Debian kernel, we compile our own upstream kernel, as well as we don't install operating system on disk. [49:28.180 --> 49:29.840] So, our systems are stateless. [49:30.020 --> 49:33.750] So, we kind of run the whole operating system as a live CD from RAM. [49:33.750 --> 49:38.290] And this gives us, like, more flexibility and ease of deployment. [49:38.290 --> 49:45.120] So, when we need to update the operating system or the kernel, we don't have to care about the state. [49:45.120 --> 49:48.620] We just, like, reboot the server, and, like, the new operating system is there. [50:01.110 --> 50:07.790] I haven't checked it for a while now, but from the top of my head, somewhere around 5 gig. [50:10.710 --> 50:13.870] When I say stateless, we also don't... [50:15.030 --> 50:21.010] For example, if we need to store something which is big, we bind mount the directories. [50:21.110 --> 50:24.790] Let's say, like, for packages, we can store them on disk, right? [50:24.930 --> 50:27.030] Or some configuration or data. [50:27.250 --> 50:29.390] It's like, the main code is in the rootFS. [50:32.310 --> 50:34.150] Okay, that is all our time. [50:34.350 --> 50:35.790] Thank you so much, Ignat. [50:35.990 --> 50:37.450] Give him another round of applause, please. [50:42.410 --> 50:43.210] Thank you. [50:44.150 --> 50:50.710] The next talk here is COVID making from cyber pantries to cyber glasses. [50:51.270 --> 50:52.190] Stick around if you want. [50:52.370 --> 50:54.130] But otherwise, enjoy the rest of your day. [50:54.270 --> 50:57.290] And don't forget, hackers got talent tonight. [50:57.650 --> 50:59.050] Hope to see some of you there. [50:59.050 --> 50:59.150] And we'll be right there.