[00:11.380 --> 00:12.280] Hello, everyone. [00:12.520 --> 00:15.600] Welcome to the next talk in Track 3 today. [00:16.580 --> 00:17.880] We're glad you're all here. [00:18.100 --> 00:20.080] Hope you're enjoying the conference so far. [00:20.180 --> 00:20.940] I hope you enjoyed the keynote. [00:21.320 --> 00:22.660] A couple points of business. [00:22.920 --> 00:25.820] Please silence your cell phones when you're in the track room itself. [00:25.980 --> 00:29.120] The audio is pretty sensitive and it can pick it up from quite a distance away. [00:32.260 --> 00:34.100] And I forgot my second point. [00:34.100 --> 00:41.920] So, this talk is by nwf and it is on CHERI, a modern capability architectural framework. [00:42.480 --> 00:44.600] So, we will be doing the talk. [00:44.620 --> 00:46.060] He says about 40 minutes. [00:46.060 --> 00:49.200] We should have about 10 minutes of QA at the end. [00:49.340 --> 00:50.440] So, we get the QA. [00:50.700 --> 00:53.000] I will be over here on the mic. [00:53.000 --> 00:54.940] So, let me know what your question is. [00:55.080 --> 00:58.440] I'll repeat the question back and then we can... he'll be able to answer questions. [00:58.540 --> 01:02.000] We also have some chat questions coming in through our Matrix chat forum. [01:02.000 --> 01:04.320] So, as I said, enjoy the talk. [01:06.980 --> 01:07.800] Hello, everyone. [01:08.140 --> 01:13.620] I'm nwf and today I'm going to talk to you about CHERI, a modern computing architecture centered around capabilities. [01:15.300 --> 01:22.600] Before we dive in too much, I should point out I work on CHERI for Microsoft, but I'm not speaking for my employer. [01:23.240 --> 01:24.680] Opinions in this talk are mine. [01:25.700 --> 01:32.280] CHERI is a science experiment and so please don't take this as, you know, a commitment or promise of future products. [01:32.960 --> 01:38.080] As said, questions during the Matrix and there will be some time for Q&A at the end. [01:40.060 --> 01:45.840] So, a more inflammatory title for this talk might have been, Modern Computing Architecture, Unsafe at Any Speed. [01:47.120 --> 01:49.520] And why might someone claim that? [01:49.520 --> 01:55.640] So, a less contentious phrasing of that is that software security really isn't great. [01:56.740 --> 02:04.840] As computers continue to infiltrate every aspect of existence, we're seeing a steep upward trend in yearly CVEs. [02:07.640 --> 02:22.300] But embarrassingly, these tend not to be, like, new kinds of bugs from new kinds of domains, but rather, year after year, 70% of these things turn out to be from memory safety problems, by which I mean pointer injection, buffer overflows, use after free, [02:22.440 --> 02:23.120] and so on. [02:23.920 --> 02:30.620] These problems have been with us since the beginning, at least the beginning of UNIX, so about 50 years, if not a little bit longer. [02:31.520 --> 02:39.120] And, okay, they weren't widely known, but they came to broader awareness in 1996 with Aleph-1s smashing the stack for fun and profit. [02:39.960 --> 02:43.340] And so, even by that standard, it's been 25 years and counting. [02:46.700 --> 02:54.060] And it turns out that if you take 70% of an increasingly bad time, it turns out to still be an increasingly bad time. [02:54.760 --> 02:59.500] Microsoft Security Response Center is handling more and more memory safety issues every year. [03:02.730 --> 03:06.310] Of course, we're not the first people to identify a 50-year-old problem. [03:06.530 --> 03:12.150] Lots of people have tried lots of things, ranging from minor tinkering to vast sweeping overhauls of everything. [03:12.390 --> 03:13.670] Some examples are on this slide. [03:14.570 --> 03:19.410] Unfortunately, nothing really seems to have moved us closer to done for commodity computers. [03:20.390 --> 03:22.230] But for all that, don't get discouraged. [03:22.570 --> 03:26.870] I'm going to try to convince you that there's hope if we make a slightly different kind of change. [03:28.090 --> 03:29.450] And so, enter CHERRY. [03:29.450 --> 03:33.330] On the one hand, it is indeed a pretty radical new computer approach. [03:33.550 --> 03:35.650] We're going to change how pointers work. [03:36.210 --> 03:43.210] If you remember a time before, this is a foundational shift on the same scale as adding virtual memory to the computing architecture. [03:44.510 --> 03:48.250] On the other hand, I hope to convince you that it's not so radical after all. [03:48.850 --> 03:52.670] I hope to show you that CHERRY composes well with modern microarchitectures. [03:52.870 --> 03:58.090] And that maybe C and C++ and foreign function interfaces to those can be made safer. [03:59.830 --> 04:01.670] CHERRY has also taped out. [04:01.810 --> 04:04.610] ARM has an experimental Morello prototype SOC. [04:04.610 --> 04:09.270] This is a quad core two and a half gigahertz ARMv 8.2a with CHERRY extensions. [04:09.630 --> 04:11.970] So that's surprisingly real for an experiment. [04:15.410 --> 04:25.210] So, to understand the changes that CHERRY makes to a computing architecture, it will be helpful to have a small example of some of the kinds of unsafety that we're designing it to inhibit. [04:31.580 --> 04:34.060] I swear, these things are not user friendly. [04:34.340 --> 04:34.560] OK. [04:34.860 --> 04:37.320] So, here's a little C program. [04:38.360 --> 04:41.540] And all it does is a stack allocation calls a function. [04:41.880 --> 04:45.080] And, OK, it has two rather glaring problems in it. [04:46.440 --> 04:50.240] We can work out what the stack might look like when we make that function call. [04:50.980 --> 04:52.940] And there's nothing really surprising here. [04:53.080 --> 04:54.380] Reading upwards from the bottom, [04:57.900 --> 05:03.460] we have that the lowest addresses are 16 bytes for the buff allocation. [05:03.620 --> 05:06.320] Above that are 16 bytes for the pad allocation. [05:06.600 --> 05:08.780] And above that, main has saved its return address. [05:10.020 --> 05:16.520] So, here's one possible compilation of our program into RISC-V, which is a pretty boring standard RISC architecture. [05:17.160 --> 05:19.120] Don't worry if you can't read RISC-V assembler. [05:19.220 --> 05:20.680] I'll walk you through the highlights. [05:21.980 --> 05:28.900] Gazing into the assembler, we see that the stores that we're performing are relative to an address in the register A0. [05:29.900 --> 05:39.300] And if we look a little bit further down, we can see that the compiler has inserted code before the function call to copy the stack pointer into A0. [05:39.740 --> 05:43.100] And as we said, the stack had buff at its lowest address. [05:43.400 --> 05:44.920] So, that's all as expected. [05:46.240 --> 05:48.160] But what happens when this program runs? [05:48.160 --> 05:52.780] Well, several things go wrong in rapid escalating succession. [05:53.760 --> 06:01.500] The first thing that happens is that first store instruction writes outside of buff and clubbers something in pad. [06:02.860 --> 06:08.060] That's bad, but at least it's something we can kind of explain using the names of things visible in the language. [06:09.520 --> 06:12.160] The next thing that happens, though, is really mysterious. [06:12.160 --> 06:16.840] We write outside the language visible allocations into something that's just magic. [06:17.000 --> 06:20.460] The compiler and ABI have inserted this anonymous return address. [06:20.680 --> 06:25.300] So, this is already really far off into the weeds. [06:26.080 --> 06:35.400] But then, when main goes to actually return far after our bug, it's going to load a corrupted pointer and jump who knows where. [06:35.400 --> 06:37.980] It's just going to go and do whatever it's going to do. [06:40.980 --> 06:44.760] So, let's spend a moment being sort of philosophical about what just happened. [06:45.520 --> 06:50.980] I'd contend that each thing stems from the CPU not really knowing enough about what's going on. [06:53.120 --> 07:01.640] Nothing Foo had to hand, that is in its registers or on its stack, told it how big the buffer was or really even where it was. [07:01.700 --> 07:03.460] It was just, here's an address, go for it. [07:06.780 --> 07:14.400] When main... when Foo wrote out of bounds, the store silently corrupted a pointer, which is something semantic. [07:15.400 --> 07:16.480] It overwrote it with bytes. [07:17.680 --> 07:28.820] And then when main jumped to its popped return address, we didn't notice, the CPU did not notice that there was anything amiss because bytes are just bytes and pointers are just addresses and addresses are just bytes. [07:30.160 --> 07:40.700] So, the common cause here, in some sense, is that we compiled C pointers, these semantic objects, down to fixed width integer addresses. [07:44.450 --> 07:47.950] So, with that example in mind, what's Cherry going to do differently? [07:48.190 --> 07:50.990] And what are these capability things that I've alluded to? [07:54.480 --> 07:57.720] So, let's ponder what it would take to fix these problems. [07:58.340 --> 08:04.700] And by fix, I mean cause to fail stop, deterministically, and ideally close to the actual problem. [08:05.060 --> 08:06.820] What would we need to pull that off? [08:07.600 --> 08:12.860] We'd need some kind of new abstract data type that we could use instead of integers for pointers. [08:13.980 --> 08:20.120] Sometimes things like this go by the name of fat pointers, but we're actually aiming for something a little better than that phrase usually means. [08:20.320 --> 08:23.380] So, we might call them just better pointers for the moment. [08:24.740 --> 08:28.540] Of course, a better pointer still needs to carry an address around. [08:29.120 --> 08:30.900] So, we have to have that in this thing. [08:32.300 --> 08:34.540] But we also want to carry some bounds. [08:34.540 --> 08:36.740] So, this is two more addresses, right? [08:36.800 --> 08:38.380] The lower base and an upper limit. [08:38.880 --> 08:43.740] To say that, you know, we're describing an object that goes from here to here and currently pointing there. [08:45.000 --> 08:55.780] And as we saw with the return address, we need to distinguish somehow between valid better pointers and those that have been tampered with somehow, including by clobbering some of their bytes. [08:56.460 --> 08:58.140] So, this has to be special. [08:58.540 --> 09:04.040] So, we'll use a bit that we'll kind of set aside from the rest of our structure. [09:05.820 --> 09:08.180] And of course, since we've opened the floodgates, right? [09:08.280 --> 09:09.220] Everybody loves metadata. [09:09.460 --> 09:11.920] There's probably going to be some other metadata in here too. [09:16.550 --> 09:20.590] So, abstract data types are all well and good, but you know, come on, we're trying to build systems here. [09:20.750 --> 09:22.210] So, what does this actually look like? [09:23.670 --> 09:31.250] So, Cherry defines an architectural representation for these better pointers with mysterious valid bits on the side, which it calls capabilities. [09:32.090 --> 09:35.590] The pointer bits are twice the size of the integer address. [09:35.890 --> 09:41.850] So, if you're on a 64-bit machine, that means that there's 128 bits in memory for these capabilities. [09:43.010 --> 09:47.050] Actually, it's 129 because there is that one bit sort of floating off to the side. [09:49.730 --> 09:52.130] These things are understood by the CPU hardware. [09:52.370 --> 09:55.350] So, we extend the registers to hold capabilities. [09:55.350 --> 10:01.210] In some sense, we have 129-bit registers, although they mostly still act like 64-bit registers. [10:03.110 --> 10:09.950] Every load and store instruction that gets executed must be to an address that's in the bounds of a valid capability. [10:10.350 --> 10:16.770] If ever that isn't true, the CPU will trap, raising a capability fault, which is rather like a page fault. [10:20.710 --> 10:25.730] And there are, concretely, some permission bits and some other metadata bits in the capability structure as well. [10:25.830 --> 10:27.430] We'll get into that a little bit more later. [10:32.080 --> 10:35.550] So, let's talk about that valid bit that's been kind of floating off in space. [10:36.160 --> 10:43.170] For historical reasons, it also gets called the cherry tag, which is a horrifically overloaded word, but is mercifully short. [10:43.170 --> 10:45.100] I will probably continue to call it the tag. [10:46.640 --> 10:58.160] So, cherry systems associate one bit of tag for every 16-byte granule, that is, 128-bit, 16-byte granule of physical memory. [10:59.560 --> 11:04.560] And as I just said, we extend the registers to hold capabilities and their tags. [11:04.860 --> 11:07.980] So, here's a small system, right? [11:08.080 --> 11:16.800] Let's say that register six is holding a capability that's pointing to a location in memory that's holding a capability, and register one is holding some data. [11:18.660 --> 11:26.740] And in some kind of made-up assembler, here are some instructions that transfer capabilities and data between the CPU registers and memory. [11:26.740 --> 11:28.760] So, what happens when we run them? [11:29.940 --> 11:37.580] So, if we load a capability, the tag that was in memory comes along with it. [11:37.860 --> 11:44.120] And so, now register two, which was the target of our load, has a set tag and is holding a valid capability. [11:45.640 --> 11:53.320] If we store some data out to memory, we always clear the tag that's out in memory. [11:54.920 --> 11:58.720] So, now that location in RAM is holding a zero tag. [12:01.580 --> 12:10.340] When we try to load a capability-sized thing from a location that has a clear tag, again, the tag just comes along for the ride. [12:10.480 --> 12:15.620] And so, we do transfer the 128 bits of the capability, but the valid bit remains clear. [12:16.420 --> 12:22.240] If we then try to use that thing because the tag is clear, the processor will trap. [12:27.120 --> 12:35.060] So, one way to think about this, if you like, is that Cherry embodies a very simple one-bit dynamic type system or tag system. [12:35.580 --> 12:40.480] Every word, every 16 bytes is either a capability or an integer. [12:40.480 --> 12:45.760] And if you ever try to use an integer where a capability is required, the processor will trap. [12:50.120 --> 12:56.040] So, other than push capabilities around like we were just doing and load and store through them, what can we do with them? [12:57.120 --> 13:04.840] So, one thing we'd better be able to do is change the address within bounds without really impacting the rest of the machinery. [13:05.960 --> 13:14.280] So, the processor has instructions, a special case instruction for adding or offsetting to an address. [13:14.560 --> 13:16.660] And for anything else, there are getters and setters. [13:16.840 --> 13:19.880] So, you can pull the address out, do whatever you need, and shove it back in. [13:23.200 --> 13:29.180] If we're going to actually use these things to save us from ourselves, we need to be able to change the bounds on them. [13:29.560 --> 13:31.880] And so, there is indeed a set bounds instruction. [13:32.900 --> 13:41.060] This generates a valid capability only if the requested bounds are smaller than the existing bounds. [13:41.240 --> 13:45.040] So, you can raise the base and lower the limit, but you can't do it the other way around. [13:47.820 --> 13:53.260] And the last thing we can really depend on is that the architecture will do provenance tracking for us. [13:53.400 --> 14:10.320] So, if we manipulate or do something bad to the bytes of the capability, other than through these capability-manipulating instructions, the architecture will clear its valid bits and then prevent us from using it as a capability. [14:13.430 --> 14:20.250] So, with those operations in mind, and by way of reminder, here's what things looked like before on a non-Cherry compilation target. [14:23.820 --> 14:28.500] And now, if we compile to Cherry RISC-V, the program looks quite similar. [14:29.180 --> 14:42.680] The first thing to note is that our store byte instructions have become capability-authorized store byte instructions, and they cite the capability in register CA0, which is just A0 extended to hold the capability. [14:44.700 --> 14:57.140] And the second thing to note is that the call site, where main calls foo, has not simply copied the stack pointer across, but now builds a capability with narrower bounds to pass as the argument. [14:57.140 --> 15:03.660] The 16 in the assembler is an immediate form because we statically know that buff is 16 bytes long. [15:04.040 --> 15:09.140] There's also a form that takes its length from a register for doing dynamic bounding. [15:10.340 --> 15:12.460] So, what happens when we run this program? [15:13.080 --> 15:22.300] Well, it crashes on the first meaningful instruction in foo, because 16 up from CA0 is outside of the capability bounds. [15:25.930 --> 15:29.070] We can use capabilities for more than just stack allocations too. [15:29.450 --> 15:34.690] So, the malloc that we run on top of Cherry, for example, can return bounded capabilities to heap objects. [15:35.630 --> 15:38.670] Internally, malloc has the authority to access the entire heap. [15:39.770 --> 15:46.070] But when responding to a client request and deriving a capability, it can set the capability bounds. [15:46.070 --> 15:53.610] And then nothing that the client of the allocator does will let it use that capability to access beyond those initial bounds. [15:55.570 --> 15:59.670] The client is, of course, free to derive its own subsets of that. [15:59.830 --> 16:04.530] But those are, again, subsets of the bounds enforced by malloc. [16:08.460 --> 16:10.220] So, that's the core of Cherry. [16:10.220 --> 16:13.060] We add architectural capabilities to the machine. [16:13.780 --> 16:19.460] And we ensure that they come about only through legitimate operations, clearing the valid bit if not. [16:20.800 --> 16:25.620] We check that every dereference is permitted by a capability. [16:26.780 --> 16:31.780] And then we rewrite the compilers and runtimes and so on to use the capabilities for pointers. [16:36.130 --> 16:42.850] Just, I want to pause for a moment and note that Cherry is, unlike much of its competition, secret-free and deterministic. [16:44.230 --> 16:50.310] So, an adversary cannot forge a capability, even if they know every bit of the system state, right? [16:50.310 --> 16:56.210] If I tell them every bit that's in RAM, including the valid bits, they can't construct a capability. [16:56.550 --> 17:01.210] This is unlike ASLR or stack canaries or other mitigations that you're probably familiar with. [17:03.270 --> 17:12.790] So, because we can't re-inject the data as pointers, most of the things that are kind of stem from smashing the stack for fun and profit no longer work. [17:14.850 --> 17:19.330] If you attempt an out-of-bounds or invalid dereference, it will always trap. [17:19.590 --> 17:23.910] There's nothing you can do to take an invalid capability and turn it back into a valid one. [17:24.870 --> 17:28.730] And byte-level corruption or attempts to widen the bounds or so on are always cost. [17:31.860 --> 17:32.340] Okay. [17:32.700 --> 17:38.040] So, now that we understand the architectural nature of Cherry, let's see how to build software on top of it. [17:40.950 --> 17:45.390] We can take this idea of using capabilities for pointers to its logical conclusion. [17:45.790 --> 17:49.250] We're going to use capabilities for every pointer in a process. [17:50.010 --> 17:58.010] So, that means both the pointers that you see and think of in the language, as well as the ones below the language, the implicit ones. [17:58.530 --> 18:04.010] And so, this means the compiler, loader, and even the kernel have to be active participants in this implementation. [18:06.730 --> 18:09.830] And we do also have to slightly change the C semantics. [18:10.010 --> 18:12.770] If you're curious for more details, please do see our programming guide. [18:12.970 --> 18:16.230] But, by and large, most C just works. [18:19.270 --> 18:28.430] If we do use capabilities to represent every pointer in a process, what we get is a capability graph between different objects with the thread registers sort of forming the roots. [18:28.850 --> 18:35.150] We call this environment Cherry ABI because it is an application binary interface that uses capabilities. [18:38.210 --> 18:39.530] But, wait, hold on. [18:39.710 --> 18:43.350] User processes do more than just follow user space pointers. [18:43.690 --> 18:50.650] Sometimes they interact with the outside world by taking advantage of this complicated thing called the kernel, and they make system calls. [18:52.350 --> 19:01.430] So, there's a risk that the kernel could be tricked into violating our carefully orchestrated capability system, making it what's called a confused deputy. [19:01.890 --> 19:08.130] After all, the kernel has legitimate intended access to the entirety of the user space address space. [19:10.070 --> 19:24.070] So, here's a short example where user space has allocated a one kilobyte buffer and is asking the kernel to write into that buffer some larger number of bytes, perhaps because an attacker has control over the length of the request. [19:25.190 --> 19:27.130] This is completely implausible, I know. [19:27.250 --> 19:28.030] Hearts never bleed. [19:29.870 --> 19:38.630] In order to limit its own behavior, a Cherry ABI aware kernel changes the system call interface so that pointers are now passed as capabilities. [19:39.850 --> 19:46.390] This way, that overlong read request will fail gracefully when the kernel goes to copy data out. [19:47.070 --> 19:50.730] And in fact, in the implementation, we can take advantage of this fact. [19:50.930 --> 19:58.030] We can take advantage of the existing fail-safe copy out, which aborts on trap. [19:58.530 --> 20:04.210] And all we have to do is pass the capability that the user gave us to copy out. [20:04.510 --> 20:08.050] So, we don't have to insert bounds instructions, or bounds checking instructions. [20:08.250 --> 20:10.230] There's actually very little code to change. [20:11.350 --> 20:16.890] And the information flow, the capability flow, through the system will enforce the bounds for us. [20:19.350 --> 20:23.510] So that's, in a nutshell, how Cherry is different than current architectures. [20:23.630 --> 20:27.270] But I also promised it wasn't as disruptive as it might first have sounded. [20:28.730 --> 20:41.550] So perhaps the greatest indication so far that Cherry is practical is that ARM and its partners, including us at Microsoft, are doing an industrial-scale science experiment named Morello, which is a kind of Cherry I had to look it up to. [20:42.650 --> 20:46.650] This is an ARM v8.2 chip with Cherry features added. [20:47.210 --> 20:48.950] It's clocked at two and a half gigahertz. [20:49.110 --> 20:50.910] It has 16 gigs of RAM by default. [20:51.170 --> 20:52.070] It's really quite nice. [20:53.330 --> 21:00.790] ARM really wants me to tell you that it is emphatically not, but will hopefully influence, successors to ARM v8.9. [21:01.430 --> 21:05.130] Morello is a dead-end experimental architecture, but it's still really cool. [21:06.670 --> 21:15.050] So, all of that to say, for present and future systems programmers, it's looking increasingly likely that Cherry will be part of the world that we live in. [21:18.380 --> 21:22.900] One of the central design objectives and why Cherry has been able to... [21:22.900 --> 21:28.460] why we have been able to make Cherry into a real chip was that it couldn't meet a whole new everything. [21:29.600 --> 21:33.620] Importantly, it needed to be compatible with commodity memory and buses and so on. [21:33.720 --> 21:38.560] We called up the DRAM manufacturers and said, hey, could you make 129-bit DRAM for us? [21:38.740 --> 21:40.880] And they looked at us like we were from Mars. [21:41.700 --> 21:42.820] So how do we do this? [21:42.920 --> 21:47.800] How do we work with ordinary DRAM, but have these weird 129-bit data structures floating around? [21:50.560 --> 21:55.000] So, as I said before, we augment the CPU core to hold capabilities in registers. [21:55.380 --> 21:57.820] So that's roughly doubling the size of the register file. [21:59.560 --> 22:04.100] And in the cache hierarchy, we carry the tags around with data. [22:05.460 --> 22:08.860] But our last level cache will split cache lines. [22:09.340 --> 22:15.000] And the data bits of them will go out to DRAM as if they were ordinary data, because they are. [22:15.260 --> 22:20.220] And the tag bits will go separately to this new dedicated thing that we call a tag controller. [22:20.820 --> 22:26.560] And the tag controller is, in turn, is backed by a reserved tag table in memory. [22:27.460 --> 22:32.880] This tag table is not architecturally accessible as data to the CPU. [22:33.200 --> 22:37.680] You can imagine there's some hardware filter in the way that says the CPU doesn't get to see those bits. [22:41.500 --> 22:45.900] So, at a glance, Cherry has two primary architectural incarnations. [22:46.760 --> 22:53.100] There is the Morello SoC, and there's also RISC-V, which is mostly in QEMU and in FPGA. [22:54.100 --> 22:58.380] Both of those have executable and human-readable ISA specs. [23:00.300 --> 23:03.160] And, of course, there is an emulator for the Morello as well. [23:05.700 --> 23:10.700] Atop these, we do most of our work in a modified FreeBSD that we call CherryBSD. [23:11.440 --> 23:14.680] The kernel and C runtime components have been made Cherry aware. [23:15.040 --> 23:20.680] There's also early work on Linux, FreeRTOS, and some other things of this ilk. [23:22.660 --> 23:30.420] The whole software stack is built mostly in cross-compilation, using a capability-aware branch of LLVM, so modern Clang and LLD. [23:31.260 --> 23:36.200] And we have educated GDB for both cross-architecture and native debugging. [23:38.340 --> 23:46.480] Continuing up the stack, we have all of CherryBSD user space, Postgres, Apache, Nginx, WebKit, Qt, and KDE ported. [23:48.600 --> 23:51.080] And we can actually do some really interesting analysis. [23:51.540 --> 23:55.520] There's a whole lecture's worth of material about porting C and C++ programs to Cherry. [23:56.040 --> 24:00.300] But generally, the higher up in the stack you go, the less work-rated it is. [24:01.760 --> 24:10.080] So when we were manipulating things in the kernel, we had to change, you know, 0.2, and in libc, it was like less than 0.5% of lines. [24:10.720 --> 24:13.980] Jits are really complicated because they are intimately aware of the architecture. [24:14.200 --> 24:19.100] But as you move up into applications, it's very, very little code that has to be changed. [24:19.260 --> 24:27.340] In fact, many KDE applications required no modifications for Cherry at all once the Qt and the KDE libraries have been ported. [24:29.160 --> 24:31.040] And so everybody likes screenshots, right? [24:31.220 --> 24:37.100] So this is KDE and some of its applications running completely cherified on RISC-V, InPremio over BNC. [24:37.400 --> 24:39.040] This all works on Morello, too. [24:39.320 --> 24:42.300] Morello has a GPU and open GPU drivers. [24:42.840 --> 24:45.480] So it will real soon now be a viable workstation. [24:45.660 --> 24:48.020] We're just going through the throes of platform bring-up. [24:51.190 --> 24:54.530] And everyone's next question, I'm sure, is, okay, how much does it cost? [24:55.870 --> 24:59.130] For complicated reasons, I don't have numbers to give you about Morello. [24:59.310 --> 25:00.850] As I said, we're still doing some bring-up. [25:02.650 --> 25:20.790] But looking back a couple of years, as of 2019, on a slightly different CPU in FPGA, we saw between 0% and 10% cycle time overhead, which for this CPU was equivalent to wall clock, with many programs actually having essentially no performance difference. [25:22.390 --> 25:29.310] The biggest cost that we see is, indeed, because we have doubled the size of pointers, L2 cache misses increase for pointer-heavy workloads. [25:30.390 --> 25:35.610] Real soon now, we should get a much better understanding of how this works on a modern microarchitecture, thanks to Morello. [25:39.580 --> 25:40.060] Okay. [25:40.340 --> 25:43.940] So the fact that all of that works is, I think, pretty exciting. [25:44.060 --> 25:47.980] But it turns out there's a lot more that we can gain from Cherry in Cherry. [25:48.840 --> 25:49.400] I see. [25:49.500 --> 25:50.120] Sorry, there's a question. [25:54.800 --> 25:55.980] I will take that. [25:56.120 --> 25:59.920] The question is about running capability-style OSs on top of Cherry. [26:00.140 --> 26:01.280] Let's hold that to the end. [26:01.420 --> 26:02.260] It's a bit of a discussion. [26:05.300 --> 26:05.700] Right. [26:05.980 --> 26:12.140] So it turns out that there's much more that we can do with Cherry than just mitigate existing problems. [26:12.880 --> 26:21.160] We can actually use it to build compartmentalized software so that we can confine the impacts of arbitrarily bad behavior to just one compartment. [26:23.440 --> 26:30.680] And the key insight here is that without a transitive capability to a given resource, there's no way to access it, even if you know the address. [26:32.560 --> 26:38.340] And so the one really attractive thing to do is to sandbox things like codecs that face untrusted data. [26:38.340 --> 26:56.180] If the only thing that you have as a codec is access to your own code, your access to your input buffer, your output buffer, maybe some scratch space, and the ability to stop running, then there's not a whole lot that you can do even as a fully compromised codec. [26:56.340 --> 26:56.720] Right. [26:56.840 --> 27:00.440] Attacker-controlled input gives rise to attacker-controlled output, but that's it. [27:00.580 --> 27:01.080] Sort of ho-hum. [27:03.380 --> 27:03.740] Okay. [27:03.940 --> 27:06.540] But there is this little caveat, right, of like, how do you... [27:06.540 --> 27:08.040] Like, what is that execute only thing? [27:08.120 --> 27:09.460] How do you get back out of one of these? [27:09.540 --> 27:10.240] It's easy to get in. [27:10.360 --> 27:13.060] You just delete things from the register file, but how do you get back out? [27:17.840 --> 27:22.860] So one answer, which we have done, is to enrich Sherry with additional kinds of capabilities. [27:23.380 --> 27:27.480] I'm going to talk about two of them, sealed and sealing and unsealing capabilities. [27:29.380 --> 27:33.620] This is exploiting some of that other metadata in the capability form that I talked about. [27:34.900 --> 27:40.700] So a cherry capability can be combined with a sealing capability to produce a sealed capability. [27:41.720 --> 27:43.560] Sealed capabilities are immutable. [27:43.560 --> 27:46.300] If you try to change anything about them, you'll clear the tag. [27:46.780 --> 27:51.720] And they are inert, in that they don't authorize other operations, including loads or stores. [27:51.900 --> 28:04.180] So you can hold onto these things, but you can't use them until they get recombined with an unsealing capability, which gives you back the original pre-sealed thing, which now you can use. [28:07.200 --> 28:12.440] There are multiple kinds of seals, and the sealing and unsealing capabilities have to match. [28:12.600 --> 28:15.660] If you try to use the wrong one, you get an untagged or a trap. [28:18.860 --> 28:22.300] Building on this functionality, we can also do something really interesting. [28:22.300 --> 28:29.540] If we have two capabilities under the same seal, we can invoke them as a sealed pair. [28:29.760 --> 28:41.300] We hand both of them to an instruction, and that instruction checks that they have the same type, unseals both of them, and installs one of them into the register file, and the other one as the program counter. [28:41.520 --> 28:42.700] So this is a kind of jump. [28:43.400 --> 28:54.120] You can think of this as object-oriented method invocation, where the executable capability, the one that gets installed as the program counter, names the method, and the other one names an object. [28:54.420 --> 28:59.720] It's do this to that, but you have no access to either of those things beyond the ability to call them. [29:01.120 --> 29:05.540] And this is one way we can get out of these sandboxes in a very continuation passing style. [29:06.220 --> 29:13.860] If the data pointer or sealed capability is the outer context continuations data, and the method is the continuations code. [29:14.320 --> 29:19.040] So this is a way for us to not have access to the outer context, but be able to return to it. [29:23.040 --> 29:27.620] Another thing we can revisit with Cherry is the need for process isolation in the first place. [29:28.580 --> 29:33.240] Traditionally, processes live in different address spaces, and we use the MMU to isolate them. [29:34.040 --> 29:41.140] And if we want to do IPC, we have to context switch between them, probably by having the kernel do some data copies. [29:42.000 --> 29:43.740] I write and you read from a pipe. [29:44.380 --> 29:51.040] This incurs TLB switching costs, which are paid in time, power, and or silicon area. [29:53.440 --> 29:58.980] We can also establish shared pages with the MMU, but notice there's something funny here, right? [29:59.080 --> 30:01.580] Pointers to the shared region are fine. [30:02.580 --> 30:12.920] If we're very careful, we can have pointers within the shared region, but there's now a risk that we might have pointers that leave the shared region, right? [30:12.940 --> 30:19.520] Which is, I get to store something that means something to me, but it means nothing to you, or worse, means something vulnerable to you. [30:25.490 --> 30:32.610] So Cherry lets us tear down the MMU-based walls between processes so that we can run many processes in a single address space. [30:33.370 --> 30:35.890] Isolation is maintained thanks to the capability system. [30:36.010 --> 30:38.130] Again, you can't access what you can't point at. [30:39.310 --> 30:43.230] And in this model, we can do IPC via those sealed capabilities. [30:43.230 --> 30:53.570] And if we want to have copy semantics, we can have a trusted switcher in user space that we trust to do the copy before completing the call. [30:54.130 --> 30:55.510] So this is really exciting. [30:55.670 --> 31:01.690] This is kernel bypass IPC with user threads directly crossing the traditional process boundary. [31:03.310 --> 31:09.490] And moreover, we get really fast sharing in this model if we just pass a capability across that IPC layer. [31:10.550 --> 31:21.050] And note that there's no risk of misinterpretation of those capabilities because there's no misinterpretation risks because it's all within the same address space. [31:24.110 --> 31:35.110] I'd like to very quickly touch on the part of the Cherry project that I'm most directly involved with, which is investigating using Cherry to build temporal safety as well as spatial safety. [31:37.470 --> 31:41.170] So another way of phrasing temporal safety is, what about use after free? [31:41.410 --> 31:43.610] And that's a perfectly reasonable question. [31:44.090 --> 31:54.190] After all, having gone through the gyrations of making pointers into capabilities at runtime, it is possible to use a capability after freeing it still. [31:54.950 --> 32:03.030] So in this example, the allocator is likely to return the same capability that I just handed back in free, and I'm still allowed to write to it. [32:06.040 --> 32:16.280] So we're going to focus, or at least I have been mostly focused on heap temporal safety, because heap objects have more complicated life cycles than stack objects, and they tend to resist static approaches. [32:17.600 --> 32:21.640] And as part of those complicated life cycles, pointers to the heap tend to spread. [32:21.820 --> 32:28.280] They end up in other heap objects, in globals, on the stack, even into the kernel heap, for example, as part of asynchronous IO. [32:30.200 --> 32:38.020] Right, so this means that, again, that there's a risk that the application inadvertently retains a reference to a freed object, which then comes to overlap a new allocation. [32:38.860 --> 32:42.180] This is undefined behavior in C, but that doesn't mean it doesn't happen. [32:45.480 --> 32:55.300] So we can eliminate the risk of use after reallocation, right after the allocator has repurposed memory, if we first revoke dead references. [32:55.940 --> 33:01.460] So this does leave a little bit of a use after free window, but it just means that we've extended that object's lifetime a little bit. [33:05.500 --> 33:08.600] Note that revocation is the dual of garbage collection, right? [33:08.760 --> 33:17.780] So rather than extending the lifetime of objects until there are no references, we're going to say, you told me this object was dead, I'm now going to delete all of the references to it. [33:20.800 --> 33:33.160] So to pull this off, we're going to expand the usual view of heap memory, in which things are either free or allocated, and just sort of cycle back and forth between the two, by introducing a third state called quarantined. [33:34.500 --> 33:44.380] Address space becomes quarantined when the application calls free, and only actually becomes free that is ready to be reallocated after a global sweep through the application's memory. [33:45.140 --> 33:49.080] This sweep will remove capabilities pointing into any quarantine region. [33:49.420 --> 34:00.640] And since sweeping is global and involves testing every capability in the address space, we allow quarantine to accumulate for a while before we make revocation pass, making it effectively a batch operation. [34:03.580 --> 34:08.580] So this turns out to be quite feasible for Cherry, in some sense, because it is a capability architecture. [34:08.900 --> 34:13.780] We don't have to guess whether words are pointers to objects or just suspicious numbers. [34:14.400 --> 34:19.120] And since we know with certainty, we're justified in erasing capabilities, right? [34:19.120 --> 34:22.280] It would be really bad if we erased a suspicious number. [34:24.680 --> 34:31.160] Beyond merely being possible, it turns out that we can add just a little bit of architectural support to speed things up really quite significantly. [34:31.760 --> 34:38.580] We can have the CPU assist us in tracking which pages have capabilities on them, so we don't need to sweep the ones that are just holding data. [34:39.660 --> 34:46.900] And we can also avoid stopping the world by configuring the processor to trap on pages that we haven't yet looked at. [34:47.340 --> 34:55.920] So if the user program tries to read a capability that we haven't yet scanned, it will take a trap, we'll scan that page, and then allow it to access just that one more page. [35:00.140 --> 35:07.620] So we have an implementation of this from a couple of years ago, and of course, a bunch of work in progress, but we haven't done a rigorous study on the work in progress. [35:08.140 --> 35:20.920] But to give you some idea, on spec 2006, on the same CPU from the last set of benchmarks, the Geo mean overhead here, if we have a second core that we can offload onto, is 2.5% on top of the cherry costs. [35:21.100 --> 35:21.700] That's pretty good. [35:23.100 --> 35:29.460] The work in progress that I mentioned of using these load traps lowers overheads across the board by about 10%, it seems. [35:30.940 --> 35:35.680] And it significantly improves, by which I mean nearly eliminates, application pause times. [35:36.920 --> 35:43.140] And of course, in the background, we're doing additional software and architectural work to try to even further tamp down on these costs. [35:47.000 --> 35:53.000] So the last thing I'd like to touch on today is, is Cherry in competition with safe languages like, for example, Rust? [35:53.940 --> 35:57.600] If you know Betteridge's Law of Headlines, you already know that the answer is no. [36:00.520 --> 36:07.520] But depending on which side people think they're on, this supposed competition between the two begins the same way. [36:07.760 --> 36:10.820] Okay, yes, everything is on fire, but... [36:10.820 --> 36:15.180] And then it diverges, with some people saying, it's all C's fault. [36:15.440 --> 36:18.500] Safe languages solve all of these problems, so why do we need Cherry? [36:20.160 --> 36:22.700] And other people say, it's all the architecture's fault. [36:22.900 --> 36:26.360] Cherry fixes the architecture, so why do we need to invest in safe languages? [36:29.160 --> 36:30.900] But I think it's important, right? [36:31.020 --> 36:33.940] So why might people think that there's competition between the two? [36:34.500 --> 36:36.020] We should look in a little more detail. [36:36.240 --> 36:46.540] So if I run an unsafe language on an unsafe architecture, so C without Cherry, then spatial and temporal errors lead to arbitrary code execution. [36:46.880 --> 36:49.300] You know, 90% of the time, that's just what happens. [36:51.980 --> 36:54.880] So one answer is, okay, I'm going to make the architecture safe. [36:55.260 --> 36:57.660] And now spatial errors fail, stop. [36:58.020 --> 37:02.240] And if you believe the cornucopia implementation, right, then heap temporal errors do too. [37:03.940 --> 37:07.140] Or you could say, okay, no, I'm just going to go switch to a safe language, right? [37:07.260 --> 37:10.840] Java or C-sharp, TypeScript, ML, Haskell, Rust, Ada, there's a whole list of these. [37:10.840 --> 37:19.460] In these safe languages, array index errors throw exceptions, which are very nice little prepackaged, well-behaved things in the language. [37:20.020 --> 37:22.220] And other spatial errors are impossible. [37:22.660 --> 37:25.860] And all temporal errors are also impossible by construction. [37:25.860 --> 37:29.480] So obviously that last box has a lot going for it, right? [37:29.580 --> 37:33.340] We should try to rewrite everything into that box. [37:36.140 --> 37:43.560] Unfortunately, just in the open world, there's about 10 billion lines of C and about 3 billion lines of C++. [37:44.420 --> 37:52.680] That probably works out to between $130 and $1,300 billion to rewrite just open source. [37:56.240 --> 38:05.280] Moreover, even if we tried to do that, there is some code that is intrinsically unsafe because it sits below the language abstraction. [38:05.780 --> 38:09.360] So these are things like your memory manager, your garbage collector, your context switcher. [38:11.740 --> 38:22.600] And moreover, different safe languages, even runtimes of the same language, likely view each other as unsafe because the runtimes will maintain different invariants. [38:24.640 --> 38:28.100] Okay, so let's try to rewrite parts of the program instead. [38:31.420 --> 38:36.720] So when we think about rewriting part of a program, the model that we have is a two worlds model. [38:37.040 --> 38:43.960] There's the safe world with the new safe code, which communicates with the unsafe old world through some well-defined interface. [38:46.000 --> 38:48.460] But in the real world, things are quite different. [38:48.460 --> 38:58.040] The safe code is inside the unsafe world, and any memory safety bug in the unsafe code can violate any of the invariants that the safe language code depends on. [39:00.160 --> 39:08.820] So the sandboxing functionality that Cherry provides gives us a mechanism to confine memory safety errors to instances of unsafe code. [39:08.820 --> 39:24.300] We can catch Cherry's architectural traps within a sandbox and turn them into error reports for the safe language or exceptions in the safe language, which can then gracefully recover because the error cannot have corrupted the state of the safe world. [39:26.960 --> 39:35.580] I'd like to give a shout out to the Rust community, by the way, where for not entirely unrelated reasons, there's already a bit of a move towards a very compatible story. [39:36.220 --> 39:43.120] They have recently come to be fretting about the semantics of unsafe Rust because it turns out compilers like to make assumptions. [39:44.160 --> 39:50.780] And there's a recent proposal that basically says we should use something that is very much like Cherry in unsafe Rust. [39:50.780 --> 39:58.920] And so if you write code, if you write unsafe Rust using this strict provenance model, it should be less unsafe on Cherry. [40:01.520 --> 40:03.360] And so with that, I'll wrap up. [40:04.000 --> 40:09.380] So Cherry enriches CPUs to have tagged capabilities with architecturally enforced invariants. [40:09.720 --> 40:17.920] This addresses many root causes of longstanding security vulnerabilities and promising and offers promising new compartmentalization mechanisms. [40:19.240 --> 40:20.720] It looks to be quite real. [40:20.980 --> 40:25.040] There is an FPGA RISC-V, and there's this ARM Morello SOC. [40:25.260 --> 40:28.480] We have LLVM, FreeBSD, and most of KDE. [40:29.800 --> 40:33.120] If you want to know more, please do get in touch on CherryCPU.org. [40:33.200 --> 40:37.020] There is a whole bunch more reading material, probably about a thousand pages at this point. [40:38.100 --> 40:40.560] There's a Slack, an email list, and so on. [40:41.540 --> 40:43.420] And of course, you're welcome to play along at home. [40:43.680 --> 40:46.160] Almost everything here is open source. [40:46.820 --> 40:48.240] So there's a getting started guide. [40:48.420 --> 40:50.260] We have a one-stop shop cross-build system. [40:50.440 --> 40:54.760] And there's even a little bit of a textbook to help you get used to programming on Cherry. [40:55.800 --> 40:57.120] And with that, I'll take questions. [40:58.340 --> 40:58.800] All right. [40:58.860 --> 40:59.940] He can hear you, everyone. [41:00.120 --> 41:02.600] So if you have any questions, let me know, and I'll read them out to him. [41:04.680 --> 41:05.660] Any questions from there? [41:05.660 --> 41:08.940] Oh, there was one on Matrix that I can answer. [41:08.940 --> 41:09.280] Yeah. [41:09.420 --> 41:11.140] So the question was, have you tried... [41:11.140 --> 41:11.960] That's the Matrix one first, and then we'll do yours. [41:12.920 --> 41:13.380] Sorry. [41:14.260 --> 41:20.340] The question was, have you tried running a capability-style OS on top of Cherry, like SEL4 or things like Gnode? [41:20.820 --> 41:24.660] We have not run those specifically, to the best of my knowledge. [41:24.660 --> 41:40.620] However, there have been, over the years, some blue-sky, green-field experiments in writing operating systems on the assumption of, you know, what if we had had Cherry from the beginning? [41:41.780 --> 41:57.740] And so, this thesis by Lawrence Esswood, Cherios, Designing an Untrusted Single Address-Based Capability Operating System Using Capability Hardware and a Minimal Hypervisor, is a really excellent look at that kind of a direction. [41:58.240 --> 42:02.580] And if people want to try porting Gnode, we would be all ears. [42:04.960 --> 42:05.720] All right. [42:05.820 --> 42:06.560] Question from the room. [42:15.500 --> 42:16.090] Okay. [42:16.280 --> 42:23.500] The question is, why did you reduce the amount of metadata from 127 to 127 bits from 128 bits so you could use the tag? [42:24.900 --> 42:25.500] Ah. [42:25.820 --> 42:27.220] So, the reason that the... [42:27.590 --> 42:28.550] That's an excellent question. [42:29.800 --> 42:37.340] So, the reason that we keep the tags separate is so that memory continues to act like memory. [42:37.340 --> 42:38.130] Right? [42:38.220 --> 42:47.740] If we put the tag in with the other chunk of bytes, then a data store to memory could set that tag, right? [42:47.920 --> 42:52.500] Or could, you know, could do whatever it needed to do to magic up a capability out of nowhere. [42:53.280 --> 43:07.190] Because the valid bits are off on the side, the architecture can manipulate them separately and enforce this invariant that if the tag is set, then the corresponding bytes represent a well-formed, legitimately derived capability. [43:09.040 --> 43:10.260] Does that answer the question? [43:11.380 --> 43:11.690] All right. [43:11.760 --> 43:12.550] That answered the question. [43:12.900 --> 43:14.480] Any other questions from the room? [43:16.860 --> 43:17.650] All right. [43:17.760 --> 43:18.540] We have one more. [43:30.770 --> 43:32.730] Doesn't the tag explorer... [43:32.730 --> 43:33.690] Run through it again. [43:33.790 --> 43:34.970] Doesn't the tag explorer... [43:34.970 --> 43:36.230] I'm sorry. [43:40.220 --> 43:42.680] We're going to bring the questioner up to the mic. [43:44.180 --> 43:44.620] All right. [43:47.000 --> 43:57.760] Doesn't the tag controller being external and inaccessible have implications in, for example, hibernate and resume for processes, because the tags... [43:59.800 --> 44:00.440] Yes. [44:00.860 --> 44:02.440] So, what an excellent question. [44:02.720 --> 44:10.660] And so, more generally, paging is a really interesting question because, for example, discs also don't understand capabilities, right? [44:10.680 --> 44:11.680] They just understand bytes. [44:13.180 --> 44:25.940] And so, the way that paging and hibernate and resume would work is the kernel or the hypervisor retains a capability to the entire address space. [44:27.680 --> 44:37.440] And when it pages something back in, it uses that authority to reconstruct capabilities from the bytes on disc. [44:37.640 --> 44:40.540] So, we write out all of the bytes in a page. [44:40.580 --> 44:46.700] We then write out the tag bits that correspond to those bytes as a separate chunk of bytes on disk. [44:46.980 --> 44:54.160] And then we pull both of those back in and recombine them using the capability manipulation instructions of Cherry. [44:56.100 --> 45:00.320] So, there are... yes, there is some complexity there, but it can be made to work. [45:03.610 --> 45:04.330] All right. [45:04.410 --> 45:05.250] Any final questions? [45:07.050 --> 45:13.190] I see one on matrix, which is, would this mitigate speculative execution vulnerabilities that expose data via side channels? [45:13.190 --> 45:15.410] And, oh, that is an excellent question. [45:15.630 --> 45:16.190] We have a... [45:16.190 --> 45:22.010] We, the computer lab at Cambridge University, have a short report that looks into this. [45:22.190 --> 45:24.530] But, boy, it's just a... [45:24.530 --> 45:29.190] Well, speculative vulnerabilities in general are just a minefield, and so why would we be different? [45:30.390 --> 45:43.330] But, yes, because leaking the capability bits themselves doesn't buy an attacker anything, we do reduce the impact of a bunch of speculation vulnerabilities. [45:43.730 --> 45:48.190] I won't say we mitigate them, but right off the bat, we reduce their... [45:49.830 --> 45:52.390] Sorry, right off the bat, we reduce the impact of some of them. [45:52.390 --> 45:53.930] Some of them we do mitigate. [45:54.310 --> 46:09.130] So, for example, if you think of, like, the Spectre V1 chain load gadget, because you pick up bounds when you do the first load, even in speculation, you won't execute the second load. [46:09.290 --> 46:14.110] It would speculatively trap, but instead you'll abort and unwind. [46:16.210 --> 46:20.170] So, some vulnerabilities also just basically, right off the bat, get mitigated. [46:20.970 --> 46:28.050] But, we don't do anything, for example, about side channels that might leak your cryptographic keys, because those are just data. [46:28.730 --> 46:37.610] So, there is this really interesting landscape of, like, splitting how cherry and speculation work together. [46:42.360 --> 46:46.760] All right, I think that's the last question, unless there's one more from the audience. [46:48.700 --> 46:49.100] Excellent. [46:49.280 --> 46:51.460] Well, thank you so much, Nuf, for the talk on cherry. [46:51.680 --> 46:52.460] It was fascinating. [46:52.880 --> 46:53.440] Thank you, everyone. [46:53.940 --> 46:56.620] Will the resources be available to the audience anywhere? [46:57.360 --> 46:57.760] Yes. [46:57.900 --> 47:00.260] I will share these slides and all of the links in them. [47:01.140 --> 47:07.660] Yeah, if you share them in the Matrix chat channel, everyone can access it through the Matrix chat channel throughout the rest of the conference and afterwards. [47:08.760 --> 47:09.160] Perfect. [47:09.860 --> 47:10.340] All right. [47:10.440 --> 47:11.860] Thank you so much for your time. [47:12.040 --> 47:12.480] We really appreciate it. [47:12.480 --> 47:12.980] Thank you, everyone. [47:17.680 --> 47:18.000] Thank you, everyone. [47:18.000 --> 47:18.360] All right, everyone. [47:18.440 --> 47:24.280] The next talk will be cybersecurity certification, the good, the bad, and the ugly at 2 o'clock. [47:24.600 --> 47:25.740] We look forward to seeing you there. [47:25.740 --> 47:25.800] Also, let's see what he gets there. [47:25.800 --> 47:25.840] Thank you.