[00:02.630 --> 00:15.570] The reason I decided to give this talk was because at the last HOPE conference I was in a bar with some guys after a long day of talks and we were talking about the way information is encoded in various ways and various systems. [00:16.410 --> 00:21.090] And at the time I was working on a DNA project trying to figure out how information was encoded in DNA. [00:21.750 --> 00:24.910] And I thought that they would be really interested in this idea so we started talking about it. [00:25.250 --> 00:31.850] And there was clearly a lot of interest in it but all throughout the conversation they kept coming back to the What's DNA? [00:32.510 --> 00:42.930] And I sort of had to realize that I was talking to people who, you know, spent their lives setting up honey pots and filtering packets and stuff and maybe they never took a biology course. [00:44.390 --> 00:56.590] So the second reason I decided to give this talk is because I see a lot of similarities between what molecular biologists do in trying to figure out systems that are difficult to understand and what other people do trying to explore systems. [00:56.590 --> 01:00.230] And I think that there's a sort of growing interest in molecular biology. [01:01.150 --> 01:04.210] And I think you can see that just in this next slide. [01:04.490 --> 01:11.850] Here's a cover of a recent issue of Make Magazine in which there's an article talking about how to precipitate DNA in your kitchen. [01:12.590 --> 01:21.270] And this article describes how to isolate DNA, how to run a gel to characterize DNA, how to build a thermocycler to use PCR to amplify DNA. [01:21.630 --> 01:27.890] And what you see in the picture here is a glass with some ethanol on the bottom and some salt water on the top. [01:28.050 --> 01:31.290] And DNA is really soluble in salt water but it's not very soluble in ethanol. [01:31.610 --> 01:34.350] So right at the interface there's some precipitation going on. [01:34.450 --> 01:37.790] And you see these little tiny strands of DNA precipitating out of solution. [01:41.630 --> 01:44.070] So this leads to the question, what is biohacking? [01:44.750 --> 01:47.110] I think in general it's two things for me. [01:47.230 --> 01:52.010] And one of them was, you know, in honor of Steven Levy who spoke this afternoon. [01:52.790 --> 01:58.470] He talked about some of the characters in that book in the following way. [01:58.550 --> 02:03.930] They had grown up with a specific relationship to the of the world wherein things had meaning only if you found out how they worked. [02:04.650 --> 02:07.130] And how would you go about that if not by getting your hands on them? [02:08.110 --> 02:12.490] Now with biological systems, I tend to think they have meaning even if I don't really understand how they work. [02:12.630 --> 02:21.870] But if you consider that when you walk down the street, every leaf on every tree is executing operations that have been going on since longer than humans have existed on the planet. [02:21.870 --> 02:24.070] And we don't even understand most of those things. [02:24.490 --> 02:25.590] I think it's fascinating. [02:25.690 --> 02:30.450] And I think it's also fascinating to try to understand what those mechanisms are. [02:30.750 --> 02:34.010] And there are methods for getting your hands on the insides of living cells. [02:34.510 --> 02:45.150] And the second aspect of, I guess, biohacking would be given a little bit of information and given a little bit of knowledge, is there anything that I can do that's novel? [02:46.610 --> 02:58.190] I should say on that last slide, one of the main points I wanted to make here was that there was a time when the only people isolating DNA were PhDs and researchers. [02:58.590 --> 03:00.810] And now, you know, high school students are doing it. [03:01.270 --> 03:05.850] And I think these two groups of people are still isolating DNA, but they're doing it for two different reasons. [03:07.290 --> 03:11.890] Researchers are trying to find out more and more information about systems which they already know a lot about. [03:12.990 --> 03:17.290] Whereas a high school student might say, look, there's a lot of stuff I don't have to know to do something useful. [03:17.930 --> 03:19.250] And that's a form of abstraction. [03:19.650 --> 03:21.930] And abstraction is one way that engineers get things done. [03:22.370 --> 03:23.630] And this will come up later. [03:26.130 --> 03:30.890] So these two different aspects, one is sort of a top-down approach, trying to probe a black box and see what happens. [03:31.210 --> 03:35.630] And the other one is, if I have a few components, what can I do with this? [03:35.750 --> 03:39.070] You know, I know what this protein is used for, but is there anything else I can do with this protein? [03:41.030 --> 03:44.650] So we could ask, are biological systems hackable? [03:45.310 --> 03:46.390] I think they are. [03:46.990 --> 03:48.730] Consider this picture here. [03:48.950 --> 03:54.730] Where if you make a small change to the information content of the system, suddenly it has a very drastic effect. [03:54.990 --> 03:58.610] So this is a disease called brachydactylism. [03:59.290 --> 04:06.310] And a small change to the information content of DNA causes the second bone in your index finger to change length. [04:07.930 --> 04:11.950] So does that mean, then, that there's a gene which controls the length of every bone in your body? [04:12.990 --> 04:15.410] Is this information used anywhere else in your body? [04:15.530 --> 04:22.410] In other words, if you saw someone with a short index finger, would you expect their toes, I guess their mama toe, also to be extra short? [04:23.250 --> 04:29.470] It turns out that people with this disease do have shortened toes. [04:29.770 --> 04:34.470] But you might ask, why doesn't it affect the length of every, you know, the second bone of all of your fingers? [04:34.470 --> 04:35.630] It's a very similar bone. [04:36.470 --> 04:37.510] But it just doesn't. [04:38.670 --> 04:45.910] Another example of this is shown here, where a small change to the information content of the system causes an entire program to misfire. [04:46.290 --> 04:52.190] So on the left-hand side, I mean, on the left-hand side of this slide, is a regular fly head. [04:52.510 --> 04:55.950] And there's these two little structures right here called antennae. [04:56.830 --> 05:01.090] But on the right-hand side, those antennae have been replaced by legs. [05:01.410 --> 05:06.350] So these two pictures are pictures of flies with legs growing out of their head instead of antennae. [05:07.050 --> 05:09.510] And it takes a developmental program to grow a leg. [05:09.650 --> 05:11.290] It's not just one, you know, little change. [05:11.790 --> 05:16.410] So, yet, it's only one change in the DNA that causes this whole program to misfire. [05:16.410 --> 05:24.570] So, I think these are both ways in which you can see that the information content of living systems executes very much like programs do. [05:25.190 --> 05:27.030] And things can go wrong and things can be altered. [05:28.590 --> 05:31.550] Both of these changes I just showed were things which were found in nature. [05:32.050 --> 05:38.490] Nobody knows enough about biology to go around giving people short fingers or make, you know, flies grow legs out of their face. [05:39.930 --> 05:41.610] But that's not to say that that's not changing. [05:42.910 --> 05:45.770] There's programs trying to figure out, you know, like the salamanders. [05:45.890 --> 05:48.090] If you cut off one of their limbs, the limb grows back. [05:48.830 --> 05:52.150] In zebrafish, if you cut up some of their organs, the organs regenerate themselves. [05:52.790 --> 06:00.730] So these mechanisms exist and there are research programs going on to try to figure out how do we, you know, understand all this stuff. [06:02.270 --> 06:05.330] So in this talk, I'm basically going to talk a little bit about DNA. [06:06.390 --> 06:13.250] I'm going to give a little bit of a historical perspective to try to set the stage for why is thinking about hacking molecules relevant now? [06:13.730 --> 06:16.990] What's changed in the past few years to make this such a pressing question now? [06:17.430 --> 06:20.470] And then I'm going to talk a little bit in detail about two biological circuits. [06:20.470 --> 06:26.510] Just to sort of get a sense for what kind of things people do and maybe what they can do. [06:27.290 --> 06:30.750] So first of all, DNA is a molecule. [06:30.910 --> 06:31.590] It's in all your cells. [06:32.910 --> 06:36.730] Every living system is composed based on the DNA in their cells. [06:37.010 --> 06:39.450] All these organisms share DNA. [06:39.850 --> 06:42.290] In other words, they all have genes. [06:42.650 --> 06:45.290] The genes often will work between organisms. [06:45.450 --> 06:46.370] So these are yeast cells. [06:46.730 --> 06:47.630] This is a human being. [06:47.810 --> 06:48.370] There's a frog. [06:48.590 --> 06:51.290] This organism here is like a big nose with legs. [06:51.470 --> 06:53.470] It's got amazing olfactory capabilities. [06:53.470 --> 06:57.150] And the molecular mechanisms are often preserved between these organisms. [06:57.350 --> 07:02.930] So if you study the way cells divide in yeast, you can often understand mechanisms of cancer. [07:03.550 --> 07:06.810] So the molecular mechanisms are often preserved between cells. [07:07.010 --> 07:13.890] But the DNA content... these organisms exist because of the way that the information content in their DNA is executed. [07:15.270 --> 07:19.750] The way information is encoded in DNA is in the form of four... what are called bases. [07:20.750 --> 07:22.470] And these are structures of those bases. [07:22.690 --> 07:23.590] A, G, C, and T. [07:23.830 --> 07:25.010] Or A, C, G, and T. [07:25.250 --> 07:27.690] This stands for adenine, cytosine, guanine, and thymine. [07:29.010 --> 07:32.030] And these bases are linked together in chains to form a pattern. [07:34.090 --> 07:35.970] So they're attached to sugar molecules. [07:36.050 --> 07:39.310] And then there's a phosphate backbone, which you can use to connect these things together. [07:39.310 --> 07:43.850] And once you can connect molecules like this together, you have a way of encoding information. [07:43.850 --> 07:48.030] Because these things can stretch on for hundreds, thousands, and millions of units long. [07:51.090 --> 07:53.010] So it forms a way to make patterns. [07:53.950 --> 07:57.490] In cells, most of the time DNA doesn't exist as a single strand like that. [07:57.610 --> 07:59.150] Most of the time it exists as a double strand. [07:59.710 --> 08:01.250] You've probably heard of the double helix. [08:01.370 --> 08:04.430] This is with the two strands that are bound together. [08:04.430 --> 08:07.050] And they form this sort of twisty structure. [08:07.050 --> 08:18.210] But on that last slide where I showed one strand, because DNA exists in two strands, there's a set of pairing rules which determines how these two strands interact. [08:20.050 --> 08:22.510] So C's are always paired with G's. [08:22.630 --> 08:24.330] And A's are always paired with T's. [08:24.790 --> 08:29.030] The idea there is if you know the sequence of one strand, then you know the sequence of the other strand. [08:29.850 --> 08:35.750] The other thing to consider is that the strands are held together by a series of very weak interactions called hydrogen bonds. [08:35.750 --> 08:38.710] And these interactions are easily broken and can easily come back together. [08:39.510 --> 08:44.510] Whereas the chain that makes up the strand of DNA is formed by very strong bonds. [08:44.770 --> 08:48.770] So it's hard to break those chains. [08:49.630 --> 08:54.930] A fundamentally important property about DNA is its ability to come apart and then come back together. [08:56.390 --> 09:00.590] Because these hydrogen bonds are weak bonds, it's really no different than boiling water. [09:00.590 --> 09:03.030] If you add heat to water, the water turns into steam. [09:03.250 --> 09:05.170] If you add heat to DNA, it comes apart. [09:07.030 --> 09:15.430] But because there are these pairing rules, so DNA can melt and it can come back together, this provides a mechanism of addressability. [09:15.990 --> 09:22.190] So the pattern of bases on one strand will only come back together with its opposing pattern on the opposite strand. [09:23.470 --> 09:25.750] So this is the basis of a lot of things in biology. [09:25.990 --> 09:29.110] It's how organisms replicate themselves. [09:29.950 --> 09:38.410] It's how on CSI you can find some DNA at a crime scene and then go out in the population and match that DNA to some unknown population. [09:38.730 --> 09:45.950] So one of these strands, if put into a complex mixture of all of humanity, can find just the one strand that matches itself. [09:46.710 --> 09:49.730] This is also the basis of using DNA as a computer. [09:51.350 --> 09:59.510] If you could encode a problem in little snippets of DNA, you can generate the opposite strand as an infinite number of solutions. [10:00.130 --> 10:06.110] And if those little pieces can find the opposite strand in a library of infinite solutions, a certain structural form. [10:06.190 --> 10:07.430] And you can isolate that structure. [10:07.570 --> 10:09.230] And this is how people use DNA as a computer. [10:10.250 --> 10:12.610] So it's a really powerful thing, but it's not very practical. [10:12.610 --> 10:16.950] Because it takes, you know, with current methods, it takes about a week to solve a simple problem. [10:18.950 --> 10:22.310] If you write DNA out, like on paper, this is what it looks like. [10:23.850 --> 10:26.730] It's just a set, you know, there's two strands, a top strand and a bottom strand. [10:28.010 --> 10:29.510] If you...so it's not that readable. [10:29.690 --> 10:33.150] But if you are familiar with DNA at all, you can start seeing certain structures in here. [10:33.310 --> 10:36.370] Like this ATG here, that's often the beginning of a protein. [10:37.410 --> 10:41.070] There's another sequence in here, GAA TTC. [10:41.070 --> 10:44.810] This is a recognition sequence for an enzyme which can cut DNA into two pieces. [10:46.150 --> 10:51.870] So there's a lot of enzymes in biology which will recognize various signature sequences in DNA. [10:52.010 --> 10:52.910] So the DNA can be cut. [10:53.130 --> 10:56.150] And since it can be cut, there's also enzymes for putting it back together. [10:56.450 --> 10:59.790] And this is the way that you can sort of move genes between organisms. [10:59.930 --> 11:00.970] You can rearrange genes. [11:01.530 --> 11:04.750] You can make a lot of DNA structures using these enzymes. [11:07.070 --> 11:09.110] So I said that DNA encodes information. [11:09.330 --> 11:14.250] And the way that that information is expressed in the world is through something called the central dogma. [11:14.730 --> 11:23.010] So given a piece of DNA, if you want to express that DNA, the first step is to express it in the form of a single-stranded molecule called RNA. [11:23.830 --> 11:25.310] That's one level of gene expression. [11:25.350 --> 11:27.610] And it's tightly controlled by various things. [11:27.610 --> 11:30.850] The next thing is to then translate that RNA into protein. [11:31.930 --> 11:34.270] Proteins, you know, they consist of 20 amino acids. [11:34.530 --> 11:36.510] The amino acid code is the same in all organisms. [11:38.110 --> 11:40.290] And proteins are the way that we get things done in the world. [11:40.950 --> 11:44.350] Basically, all these 20 amino acids have slightly different chemical properties. [11:44.990 --> 11:51.850] So it's the configuration of amino acids in a protein which gives you access to chemistry so that you can do stuff. [11:52.590 --> 11:55.750] All the DNA in an organism is something we call the genome. [11:56.570 --> 12:03.910] If you consider all the DNA that makes up a single organism, just represented as one long strand here. [12:04.670 --> 12:08.610] So here's sort of a summary of the genomes of a few different organisms. [12:09.870 --> 12:12.750] E. coli is probably the most studied organism on the face of the earth. [12:12.950 --> 12:14.450] It's got about 4,300 genes. [12:14.910 --> 12:19.270] And its DNA content is about 4.5 million bases of DNA. [12:20.510 --> 12:24.370] Yeast and humans are both complicated structures that have nucleus, nuclei. [12:25.070 --> 12:26.810] Yeast has about 6,000 genes. [12:27.050 --> 12:28.550] It's about 12 megabases. [12:29.410 --> 12:31.610] Humans are about 3 gigabases. [12:32.750 --> 12:36.710] But the number of genes is not really known because humans are complicated. [12:37.030 --> 12:41.390] And this number has ranged from 20,000 upwards of 150,000. [12:42.150 --> 12:44.950] And the best guess now is somewhere around 23,000. [12:45.070 --> 12:45.170] Yeah? [12:45.570 --> 12:46.870] How about nematodes? [12:47.050 --> 12:47.730] Do you know offhand? [12:47.730 --> 12:49.610] Yeah, nematodes have about 20,000 genes. [12:50.570 --> 12:53.130] So are we really no different than worms you find in your salad? [12:53.750 --> 12:55.490] I mean, this is the kind of thing you'd read in the paper. [12:55.590 --> 12:58.170] Like, oh, we're... how can we be the same as worms? [12:59.610 --> 13:02.490] But there's more to it than that because it's not just genes. [13:02.830 --> 13:04.930] It's controlling how those genes are expressed. [13:05.830 --> 13:10.190] So consider that in 12 megabases you've got 6,000 genes. [13:11.090 --> 13:13.150] In yeast, about half of that. [13:13.270 --> 13:16.810] So about half of the content of this DNA codes for those genes. [13:17.110 --> 13:20.250] So about 50% of the genome is... codes for proteins. [13:20.530 --> 13:24.670] And about 50% is regulatory information or just spaces between genes. [13:25.530 --> 13:30.010] In humans, it's only about 3% of the genome which codes for genes. [13:30.010 --> 13:34.210] So 97% of the DNA is there and we don't know what it's for. [13:34.350 --> 13:35.010] We don't know what it does. [13:35.430 --> 13:36.950] Why is there such a different density? [13:37.170 --> 13:38.350] And I've shown a little plot here. [13:38.450 --> 13:42.290] This is the yeast genome at 12 megabases relative to the human genome. [13:42.770 --> 13:43.630] Just in length. [13:44.890 --> 13:53.370] Consider that DNA as a single molecule, if you were to stretch out all the DNA in one of your cells, it would be 6 feet long. [13:54.170 --> 13:58.650] And somehow all of your cells are able to wrap that DNA into a tiny little nucleus inside the cell. [14:00.670 --> 14:01.730] That's just fascinating. [14:04.150 --> 14:08.630] But given that... given that the information content... [14:09.930 --> 14:13.470] I mean, that organisms are determined by the information content in their genome. [14:14.290 --> 14:18.130] A few years ago, people thought, well, why don't we just sequence genomes? [14:18.390 --> 14:25.450] If it's... you know, at the time that genome sequencing was proposed, people were studying stuff one gene at a time. [14:26.050 --> 14:26.170] Yeah? [14:30.420 --> 14:32.860] The genome doesn't specifically encode it. [14:33.000 --> 14:34.820] Is it junk DNA anymore or is that...? [14:34.820 --> 14:35.620] Well, yeah. [14:35.740 --> 14:38.200] So the 97% of the genome that we don't know what it does... [14:38.700 --> 14:40.560] The term junk DNA is thrown around. [14:40.980 --> 14:44.900] But any serious scientist doesn't really, you know... [14:45.680 --> 14:46.600] It does something. [14:46.920 --> 14:47.420] Oh, yeah. [14:47.580 --> 14:49.540] It certainly does something, because you can't do without it. [14:50.760 --> 14:53.160] So to call it junk, it's just like, well, you don't know what it does. [14:53.240 --> 14:54.400] I guess that means it's junk. [14:54.540 --> 14:56.220] I mean, it's just... it's not a very useful statement. [14:56.560 --> 14:58.120] How do you know you can't do without it? [14:58.680 --> 14:59.180] What's that? [14:59.320 --> 15:03.440] How do you know you can't do without it? [15:10.060 --> 15:15.720] So, consider that at the time that people realized that we have to get the information out of DNA. [15:19.220 --> 15:21.520] It... studying a single gene made a PhD thesis. [15:22.320 --> 15:23.320] And took years. [15:23.540 --> 15:27.040] So if you wanted to start on a project, like let's say you were studying some aspect of cancer. [15:28.180 --> 15:34.340] It would take you a long time, like maybe two or three years, just to try to find a gene that's involved in the thing you're studying. [15:34.960 --> 15:38.100] And then you could do some functional studies on that particular gene and then do something. [15:38.260 --> 15:43.180] So you might spend five or ten years... five or ten years of your life just trying to find a single gene. [15:44.560 --> 15:48.100] Because no... you know, cells and living systems are just black boxes. [15:48.220 --> 15:49.460] We don't know what gene controls what. [15:51.180 --> 15:53.280] So, yet sequencing technology existed. [15:53.560 --> 15:55.000] So why not sequence genomes? [15:55.180 --> 15:58.800] Well, the reason is because at the time this was being proposed, I was sequencing DNA. [15:59.040 --> 16:04.980] And it took me about three days of work with radioactivity, trying to label molecules, to get just a hundred bases of sequence. [16:05.520 --> 16:10.120] That meant if I wanted to sequence the three billion base human genome, it would take 250,000 years. [16:12.740 --> 16:16.260] So, one solution is, well, hire 250,000 people and it'll take a year. [16:18.480 --> 16:26.520] But if you consider, like, the space program, some people get really pissed off about the space program because they say it's going to take enormous resources to do something like, you know, go to Mars. [16:27.140 --> 16:29.920] And if we were to use that money in another way, we could accomplish a lot. [16:30.860 --> 16:32.680] But this is an investment in infrastructure. [16:34.680 --> 16:41.280] And so, if you set a goal and then provide funding resources, people will attack that goal and the technology will change. [16:41.280 --> 16:50.540] And it has changed and such that today, using just off-the-shelf technology, we can collect a billion, I mean, a human genome, three billion bases in about three weeks. [16:50.960 --> 16:53.120] And that is changing on a daily basis. [16:54.340 --> 17:01.460] Which means that, you know, in the not-too-distant future, when you go to your doctor's office, you might, you know, you might know the sequence of your own genome. [17:02.500 --> 17:05.040] Which can be really important for various things. [17:05.040 --> 17:10.840] Another way to look at this problem is if information content is the limiting, is a limiting aspect in biology. [17:12.340 --> 17:15.600] Around the time that genomes are starting to be sequenced was around here. [17:16.200 --> 17:17.520] And this is where we are now. [17:17.640 --> 17:20.500] So this is a graph of the information content in GenBank. [17:21.860 --> 17:30.060] So when you have an event like this, where you go from a long time with almost no information, and then suddenly a huge shift like this, things change. [17:30.240 --> 17:33.680] It causes a paradigm shift in the way that you think about the system that you're studying. [17:34.320 --> 17:36.740] And you can, this even looks a little bit like Internet usage. [17:36.920 --> 17:38.880] Where, you know, back here, almost no one was on the Internet. [17:39.300 --> 17:41.000] And now, almost everybody is. [17:42.880 --> 17:43.940] So, genomes again. [17:44.600 --> 17:48.800] Just because we have the sequence of a genome, doesn't mean that we know what any of it codes for. [17:49.060 --> 17:49.200] Right? [17:50.240 --> 17:52.180] We can try to predict where genes are. [17:54.060 --> 17:55.960] But we still don't know what those genes do. [17:57.280 --> 18:12.260] In addition, one of the fundamental components which makes like an ape or a worm different than a human, is not just the genes, because we share so many genes in common with yeast and worms and apes, but the control elements that turn those genes on or off. [18:13.120 --> 18:16.560] This is a ripe area of biology which is waiting to be explored. [18:17.640 --> 18:18.980] It's called the CIS regulatory code. [18:19.120 --> 18:21.480] It's the control mechanisms for genes. [18:22.220 --> 18:27.060] So having all this information, primarily what it does is it causes people to think about the system in a new way. [18:27.740 --> 18:34.300] If you're studying a black box and you don't have a good map of it, you know, you're going to sort of push on it in one place and see what happens in another place. [18:34.380 --> 18:40.260] But if you have an entire map, it's like suddenly having a map of the world and you get to now annotate it. [18:40.420 --> 18:41.100] Did you have a question? [18:41.300 --> 18:51.120] Is the information content of a genome both necessary and sufficient? [18:52.580 --> 18:54.140] It depends on the organism. [18:57.180 --> 19:01.240] For instance, so let me just talk about this for a second. [19:01.600 --> 19:02.760] Thinking about the whole system. [19:03.080 --> 19:07.860] In 1996, yeast was the first eukaryotic organism sequenced. [19:08.080 --> 19:11.300] Now you've got 12 megabases of sequence. [19:12.280 --> 19:14.800] You can predict that there's about 6,000 genes. [19:16.160 --> 19:21.860] For the first time, people are realizing, wait a minute, I need to know what a database is because I need to keep track of all these genes. [19:22.600 --> 19:24.920] There are methods for knocking genes out of yeast. [19:25.840 --> 19:33.380] So if you've got 6,000 targets and a lot of graduate students, you can say, well, let's start knocking out genes and see what happens. [19:33.660 --> 19:37.400] So what was done is that the yeast community made a knockout collection. [19:37.640 --> 19:43.720] They tried to make 6,000 different strains of yeast in which each yeast was missing a different gene. [19:44.760 --> 19:51.860] So of that, they were able to make about 4,000 strains because about 2,000 of the genes are absolutely essential. [19:52.060 --> 19:54.400] If you knock out the gene, the organism is dead, right? [19:55.680 --> 19:58.660] So necessary and sufficient, it sort of depends. [20:00.220 --> 20:05.920] Some single knockouts, you knock out gene X, the yeast are alive, you go to knock out a second gene, now the yeast are dead. [20:05.920 --> 20:12.520] I asked people to say, is there information content that is necessary for encoding the organism outside the genome? [20:13.780 --> 20:19.160] Is there information content required for organisms that's outside the genome? [20:20.100 --> 20:20.540] It's passed down. [20:20.880 --> 20:21.720] What's that? [20:21.900 --> 20:22.740] It's passed down. [20:22.880 --> 20:24.580] Yeah, reproductively, not actually. [20:25.440 --> 20:27.980] Well, there's metagenomic elements. [20:30.100 --> 20:31.040] So, like, E. coli has a genome. [20:32.100 --> 20:33.220] Yeast have genomes. [20:33.360 --> 20:34.900] But they also have plasmids. [20:35.080 --> 20:38.380] And these plasmids are like little circles of DNA that can be shifted between strains. [20:41.820 --> 20:46.740] Well, mitochondrial DNA is also, yeah, so mitochondrial DNA is sort of an extra-genomic element. [20:48.520 --> 20:52.320] Yeah, so, you know, what's an organism anyway? [20:52.320 --> 20:53.720] I mean, look at mad cow disease. [20:53.900 --> 20:55.260] You have a protein, which is infectious. [20:55.520 --> 20:56.520] There's no nucleic acid. [20:56.760 --> 20:57.840] It's a structure. [20:58.300 --> 20:59.980] It's a self-propagating structure. [21:02.360 --> 21:06.420] So, one thing that changed, along with all this, okay, yeast. [21:06.540 --> 21:08.480] You've got 6,000 genes, 6,000 targets. [21:08.900 --> 21:10.940] Well, make copies of all the genes. [21:11.100 --> 21:16.380] So, do 6,000 PCR reactions and make 6,000 pieces of DNA, and then build a robot. [21:16.620 --> 21:21.920] I mean, most molecular biologists are not robot builders, but it's actually not hard to build these robots. [21:21.920 --> 21:23.800] They use off-the-shelf technology. [21:24.040 --> 21:25.300] So, here's an XYZ robot. [21:25.440 --> 21:26.400] This is a linear stage. [21:26.620 --> 21:27.740] This is a linear stage. [21:28.020 --> 21:28.460] This little thing. [21:28.540 --> 21:29.220] Here's a linear stage. [21:29.520 --> 21:33.480] So, here's a homemade robot that can execute motion in three dimensions. [21:35.100 --> 21:41.480] So, people started realizing, well, if you can have an XYZ robot and you've got lots of genetic elements, there are a lot of things you can do. [21:42.280 --> 21:49.060] One thing that we do is we take and we create slides in which we spot down a little bit of every yeast gene on a slide. [21:49.060 --> 21:54.960] So, here's a slide that's got 6,000 spots on it representing the entire yeast genome on a hunk of glass. [21:55.760 --> 22:00.000] So, for the first time ever, remember I said that people used to study genes one at a time, right? [22:00.520 --> 22:04.700] Well, why do an experiment to measure just one gene if you can measure 6,000 genes? [22:05.140 --> 22:11.120] If you do measure 6,000 genes, this is sort of what the data looks like if you scan that slide after doing an experiment. [22:11.120 --> 22:13.160] Each one of these dots is a different gene. [22:13.900 --> 22:16.880] The color of the dot represents the level of expression of that gene. [22:18.580 --> 22:21.440] But you don't even have to know what the genes do to measure them. [22:22.000 --> 22:26.580] But let's say you're comparing, you know, healthy tissue to cancerous tissue and you see some spots fluctuating. [22:26.900 --> 22:28.660] Those are your things to go study, right? [22:29.400 --> 22:31.100] And now you've found a bunch of genes. [22:31.420 --> 22:35.000] The difficulty here is that nobody knows what to do. [22:35.000 --> 22:38.940] This is a 6,000-point vector representing the state of a system. [22:40.560 --> 22:43.080] Again, most molecular biologists are not very good at math. [22:43.260 --> 22:44.840] That's why people go into biology, right? [22:49.840 --> 22:54.640] So, how do you, you know, and this is every experiment you do is going to generate a 6,000-point vector with yeast. [22:54.840 --> 22:56.980] You do it in humans, you're going to have 40,000-point vectors. [22:57.860 --> 22:59.760] How do you compare a 40,000-point vector? [22:59.920 --> 23:01.660] How do you compare 100 different experiments? [23:01.760 --> 23:03.500] So, the data generation is huge. [23:05.460 --> 23:08.040] You know, physicists and mathematicians are getting involved now. [23:08.240 --> 23:13.920] But the main point I want to make is that now you've gone from a black box that you used to be able to poke and measure one output. [23:14.240 --> 23:17.980] But now you can poke it and measure the entire state of the transcriptional system. [23:18.260 --> 23:22.420] You can measure the... you can get a snapshot of the activity of all the genes at once. [23:23.180 --> 23:24.460] And this is really powerful. [23:25.600 --> 23:28.060] And so this is... you've probably heard the term genomics. [23:28.300 --> 23:29.580] This is what genomics is. [23:29.740 --> 23:33.300] It's instead of thinking of things one at a time, you think of them everything at a time. [23:33.300 --> 23:36.640] If you want to... like... you want to study spinal cord development? [23:37.420 --> 23:38.200] Study snakes. [23:38.420 --> 23:39.840] They're all spinal cord, right? [23:40.680 --> 23:42.480] But that means you have to sequence the genome. [23:42.780 --> 23:44.340] But we can sequence the genome now. [23:45.980 --> 23:47.600] So, the whole genome of the organism. [23:47.740 --> 23:48.860] You want to measure all the genes? [23:49.060 --> 23:50.100] Do an array experiment. [23:50.180 --> 23:51.160] You can measure all the genes. [23:51.860 --> 23:56.060] There's... there's... but, you know, what I showed you was measuring the RNA level of a gene. [23:56.660 --> 23:58.700] That RNA has to be translated into protein. [23:58.900 --> 24:00.400] And that is also a regulated step. [24:01.200 --> 24:04.680] So, there are methods for studying all the proteins of an organism at once. [24:06.560 --> 24:08.700] All these proteins interact with each other. [24:08.940 --> 24:10.140] Proteins interact with each other. [24:10.280 --> 24:11.040] They interact with DNA. [24:11.380 --> 24:12.840] There's all kinds of interactions in the cell. [24:14.000 --> 24:20.920] So, if you want to think about trying to catalog those things, you can do high-throughput screens to try to measure all the interactions at any given time. [24:22.320 --> 24:24.260] And... but again, it's just one point in time. [24:24.700 --> 24:26.640] You know, if you heat the cells, this happens. [24:26.800 --> 24:29.200] If you cool the cells down, this happens. [24:29.300 --> 24:30.800] If you expose the cells to aspirin, that happens. [24:30.940 --> 24:34.100] If you, you know, do something else, put them in UV light, whatever. [24:36.000 --> 24:40.300] So, with all this information, now, people are thinking in a different way. [24:40.920 --> 24:43.140] Everybody wants to have a network map of their organism. [24:43.960 --> 24:45.240] So, this is a map of E. coli. [24:45.880 --> 24:47.680] And it's protein-protein interactions in E. coli. [24:48.200 --> 24:50.460] And you can start seeing structures from these maps. [24:50.780 --> 24:53.980] I mean, people have been studying networks for a long time. [24:54.160 --> 24:55.960] But we don't know really what they mean in biology. [24:57.720 --> 25:00.280] Here's another map of yeast-protein-protein interactions. [25:00.620 --> 25:03.280] More proteins, more interactions. [25:05.740 --> 25:06.720] What are the lines? [25:07.160 --> 25:09.960] Yeah, so the dots are individual proteins. [25:10.200 --> 25:14.420] And the lines between the dots reflect that somebody has measured an interaction between those proteins. [25:15.100 --> 25:26.800] So, one question is, I mean, if you're going to measure 6,000 things here and 6,000 things there and do, you know, a year of experiments to measure all the protein-protein interactions, you have to have a way to summarize the data, to view the data, to do something with it. [25:27.300 --> 25:31.880] Most people I know, they want to be able to draw these pictures, but they don't know how to do any sort of network math. [25:31.880 --> 25:33.580] They don't know how to compute on these things. [25:33.860 --> 25:36.860] They don't necessarily know, does this thing have any predictive power? [25:37.020 --> 25:37.700] What does it tell me? [25:37.700 --> 25:40.020] How can I compute on this structure? [25:42.740 --> 25:52.240] Which is, you know, I think that's a good thing to do, but even if you can't do that, you could probably look at this structure right here and realize that the protein at the center is probably a really important protein. [25:56.360 --> 26:03.460] But along with these sort of network maps, people also started thinking of their data in terms of these kinds of structures. [26:03.660 --> 26:06.900] So here's a little map summarizing sea urchin development. [26:07.660 --> 26:10.240] It turns out that this map is also true for humans. [26:10.440 --> 26:15.000] That sea urchins and humans, early on, have very similar developmental events. [26:15.000 --> 26:20.240] And you can sort of represent those events in terms of figures like this. [26:20.500 --> 26:24.640] And when I first saw this, the first thing I thought of was, it looks like a circuit diagram. [26:25.400 --> 26:28.240] There's things which are activators, there's things which are repressors. [26:29.080 --> 26:36.740] And if you can encode, you know, your biological system with something like this, now I kind of want to know, how do I tweak this? [26:36.920 --> 26:38.520] How do I understand the different components? [26:38.520 --> 26:39.900] It's obviously composed of modules. [26:41.080 --> 26:42.720] It's got all of these different parts. [26:43.860 --> 26:47.960] Shouldn't we be able to sort of look at this big catalog of parts and make our own network diagram of something? [26:48.360 --> 26:48.420] Yeah? [26:48.640 --> 26:52.230] Can that network a table? [26:52.610 --> 26:54.950] Can it be displayed as an algorithmic table? [26:57.070 --> 26:58.090] Yeah, I think it can. [27:00.890 --> 27:01.330] Yeah. [27:02.950 --> 27:06.090] So if we revisit now, what does biohacking mean? [27:09.010 --> 27:11.090] So where does all that just leave us? [27:11.210 --> 27:13.650] It leaves us, now we've got, we're swimming with information. [27:13.950 --> 27:15.870] We're surrounded by genes and parts. [27:16.210 --> 27:21.070] We're surrounded by entire systems level snapshots of activity. [27:22.130 --> 27:24.170] So now people think about things in a different way. [27:31.590 --> 27:35.410] So if we want to now probe systems, we can start asking questions along this level. [27:36.050 --> 27:40.290] You know, you want to understand the system, so you can sort of change existing systems. [27:40.290 --> 27:42.550] Now you can see the system, you can change it. [27:42.650 --> 27:43.570] You can knock out genes. [27:44.450 --> 27:48.210] You can study, you know, remember there's genes and there's promoters which control the genes. [27:48.510 --> 27:52.150] So you can try to understand what the activity of the promoters is. [27:52.390 --> 27:54.670] You can put genes in places where they're not supposed to be. [27:54.890 --> 27:58.190] You know, maybe there's a gene in your toenail that's not supposed to be expressed in your eyeball. [27:58.350 --> 28:00.870] So what happens if you express a toenail gene in your eyeball? [28:03.050 --> 28:10.090] If you have genes, you can not only sort of mis-express them and knock them out, But you can try changing them to see what happens if you tweak them. [28:11.410 --> 28:22.130] You can try to create new proteins, which is almost impossible because trying to start with just general sequence space and evolve something new, it's basically impossible. [28:22.310 --> 28:28.530] It's impossible because it takes too much DNA to try to search through any kind of sequence space. [28:28.730 --> 28:35.630] But it's also impossible because when you actually try to do it, you end up evolving things which bind to glass and plastic and you don't get anywhere. [28:36.190 --> 28:39.890] But that doesn't mean that there aren't other ways to go about making new proteins. [28:41.330 --> 28:45.670] And these days, we also talk a lot about, can we just create whole new systems? [28:46.050 --> 28:49.950] Given that we've got so many views of these systems, can we just create new ones? [28:52.010 --> 29:00.410] So, before I talk about some individual circuits, I wanted to just remind people or maybe introduce people to a few basic components. [29:01.350 --> 29:03.110] Here's a figure of a gene. [29:03.990 --> 29:05.910] It would be expressed in this direction. [29:06.330 --> 29:08.330] And remember that genes have promoters. [29:08.750 --> 29:14.570] And whether this gene gets expressed or not depends on the sequences in this promoter. [29:14.730 --> 29:22.230] And it depends on whether there's proteins around which will bind this promoter and turn the gene on, or other proteins which will bind this promoter and keep the gene off. [29:24.270 --> 29:27.010] Proteins which bind and keep the gene off are called repressors. [29:27.850 --> 29:30.570] And repressors can be inactivated. [29:31.090 --> 29:39.290] For instance, if you're a bacteria and you have a gene for lactose, you don't want to express that gene unless there's actually lactose around in your environment. [29:39.770 --> 29:44.450] So the lactose gene will be repressed until there's lactose around. [29:44.830 --> 29:51.050] And what lactose will do, it interacts with the protein, and when it interacts with the protein, the repressor comes off and now the gene is expressed. [29:51.850 --> 29:55.690] So there's activators and there's repressors, and this is at least in bacteria. [29:56.190 --> 29:58.790] In mammalian systems, things are just way more complicated. [29:59.830 --> 30:14.550] But nonetheless, given a bunch of genes, given a bunch of transcription factors, the general system is that signals come in through the outside world, and they transmit those signals through transcription factors to activate gene expression. [30:14.770 --> 30:20.390] So if you want to try to understand this system, there are certain tools that are available. [30:21.310 --> 30:31.170] One sort of tool is something called a reporter gene, where you can hook things up to a gene such that when that gene is expressed, you can see it being expressed. [30:32.370 --> 30:37.170] Another thing to do is to simply not hook something to the gene, but to just replace the gene with the reporter gene. [30:37.390 --> 30:42.050] That way you can measure the activity of this promoter by looking at the reporter. [30:42.470 --> 30:46.010] And some examples of reporter genes are shown here. [30:46.010 --> 30:53.590] There's a bunch of different kinds of reporter genes, but one of the most popular ones is a protein that's found in jellyfish. [30:54.290 --> 30:56.070] It's called green fluorescent protein. [30:56.330 --> 30:57.230] It was isolated from jellyfish. [30:57.910 --> 31:00.890] You know, there's a long history to GFP. [31:02.470 --> 31:08.110] But it turns out to be a really powerful tool because you can shine light on an organism and tell if it's expressing GFP or not. [31:08.910 --> 31:11.570] There's another, you know, you've all seen fireflies maybe. [31:11.810 --> 31:14.830] There's an enzyme which allows fireflies to make light. [31:15.010 --> 31:17.590] And so you can take that enzyme and use that as a reporter gene. [31:18.510 --> 31:21.110] There's a protein which metabolizes sugar. [31:21.490 --> 31:22.830] You can use that as a gene. [31:23.090 --> 31:28.310] This is a hundred different yeast colonies, all with a particular promoter hooked up to this reporter gene. [31:28.310 --> 31:34.490] And the reason why there are different levels of blue there is because each colony has a different mutation in the promoter. [31:35.050 --> 31:38.310] So the mutation in the promoter is affecting the expression of this gene. [31:38.450 --> 31:39.350] And you can see that here. [31:41.030 --> 31:45.130] Some genes, some reporter genes, here's one that codes for diphtheria toxin. [31:45.410 --> 31:48.250] So if this reporter gene is expressed, it actually kills the cells. [31:48.990 --> 31:51.770] Now why would you want a reporter gene which kills cells? [31:53.290 --> 31:55.470] Consider that, you know, diabetes. [31:55.790 --> 31:59.850] In diabetes, in your pancreas, there are certain cells called beta cells which produce insulin. [32:01.570 --> 32:07.230] For various reasons, like pregnant women go through fluctuations in the mass of their beta cells. [32:07.830 --> 32:11.110] Sometimes the beta cells just go missing and then you've got diabetes. [32:11.770 --> 32:18.650] So if you want a system in which you can study the regeneration of beta cells, you could target this gene into beta cells. [32:20.230 --> 32:27.430] In this case, there's someone doing this where they take diphtheria toxin and they put it under the control of a controllable reporter. [32:27.690 --> 32:31.010] So, I mean, a controllable promoter. [32:31.370 --> 32:34.070] In this case, the promoter can be turned on and off with tetracycline. [32:34.490 --> 32:44.230] So you have a strain of mice which are expressing, which have this gene in all of the beta cells, and you can activate it by simply giving the mice a dose of tetracycline. [32:44.230 --> 32:47.190] So when you do that, the mice suddenly become diabetic. [32:47.490 --> 32:52.470] And that allows you to study what happens when the cells regenerate themselves, if they regenerate themselves. [32:52.990 --> 32:58.110] So you can get this gene to express in a particular tissue, and you can use it to study certain phenomena. [33:00.410 --> 33:00.930] Yeah? [33:01.210 --> 33:03.070] Is the luciferase always on? [33:03.910 --> 33:05.070] It's not always on. [33:05.070 --> 33:09.890] So whether it's on or not depends on if this promoter is being activated or not. [33:10.070 --> 33:10.750] No, I mean, yeah. [33:10.910 --> 33:13.310] Assuming the promoter is activated, is it just always glowing? [33:13.570 --> 33:16.030] It would glow if there's the right metabolite around. [33:16.170 --> 33:18.070] So it uses some metabolite to generate light. [33:18.650 --> 33:19.210] And, yeah. [33:21.070 --> 33:28.790] So here's some examples of organisms which have been transfected or changed or modified by putting reporter genes into them. [33:28.790 --> 33:33.250] These are two pictures of mice that have GFP inside the mouse. [33:33.530 --> 33:34.330] So they're green. [33:34.570 --> 33:35.150] The mice are green. [33:35.990 --> 33:37.190] Somebody made a green rabbit. [33:39.350 --> 33:41.490] These are fish that you can find in fish stores. [33:41.770 --> 33:44.430] So in this case, these fish are expressing GFP. [33:45.290 --> 33:47.590] But there's a lot of fluorescent proteins in the ocean. [33:47.770 --> 33:49.310] And so people have isolated either... [33:49.310 --> 33:54.950] They've either mutated GFP to be a different color, or they've looked at other sea species and found other colors of proteins. [33:54.950 --> 33:57.430] One of them is YFP, yellow fluorescent protein. [33:57.570 --> 33:59.810] One of them is RFP, red fluorescent protein. [34:00.090 --> 34:03.930] And these are just zebrafish that have been transfected with these reporter genes. [34:06.150 --> 34:08.290] Here's an example of a pig. [34:09.250 --> 34:11.390] So the pig on the right is a regular pig. [34:11.510 --> 34:14.510] The pig on the left is a pig that's been transfected with YFP. [34:16.310 --> 34:18.590] Now, I should point out that people don't just do this for kicks. [34:18.970 --> 34:22.970] I mean, the rabbit was actually sort of an art project. [34:22.970 --> 34:24.210] But the other things... [34:26.710 --> 34:29.370] The reason for doing this is so you can study certain things. [34:29.550 --> 34:32.180] For instance, here's a picture of a mouse with lymphoma. [34:32.890 --> 34:38.230] And you can actively, in real time, without killing the mouse, study the progress of lymphoma. [34:39.810 --> 34:47.420] Another phenomena you can study is when people have blood transfusions, you know, or like a bone marrow transplant. [34:48.900 --> 34:54.120] Often, you end up finding cells, like let's say 20 years after someone has a bone marrow transplant. [34:54.400 --> 34:59.020] If you section their brain, you'll find nerve cells which came from the donor. [34:59.880 --> 35:04.080] Even though there were no nerve cells in the donor, bone marrow. [35:06.660 --> 35:09.480] Also, you know, once you... this is a Purkinje neuron. [35:09.660 --> 35:11.920] And once you're born, you don't make any more Purkinje neurons. [35:12.800 --> 35:26.540] So, yet, if you take a green mouse and have it give some kind of transfusion or bone marrow transplant to a sick mouse, and then look in that mouse later, you can find cells that were never transferred but yet are there. [35:27.140 --> 35:31.640] Meaning that either cells are fusing or there's some strange kind of event going on. [35:31.820 --> 35:34.800] But it's... the point is that you can study these things if you have markers. [35:36.680 --> 35:41.000] Here's an example of modifying herpes simplex virus by adding the luciferase gene. [35:42.060 --> 35:46.160] So, you can... what this allows you to do is track an active infection. [35:46.980 --> 35:50.960] Here's a mouse, which on day one was exposed to virus in its eyes. [35:51.740 --> 35:55.760] And then, over time, the virus spreads into the nasal cavity and everything. [35:55.760 --> 35:57.900] But then, by day nine, the infection's gone. [35:58.620 --> 36:01.260] You know, you can't study this without some kind of reporter gene. [36:01.860 --> 36:07.700] Or at least, the way they studied it before reporter genes was to kill the mouse and slice up its head and sort of, you know, look at stuff. [36:11.220 --> 36:25.700] So... so with all this genomics, with all these modifications we can do now, sort of this notion has been coming up in a lot of conversations I've had in the last two or three years with people, where people are just starting to think of biology as parts. [36:25.760 --> 36:26.160] Yeah? [36:26.360 --> 36:29.240] Can you attach a reporter gene to any particular gene? [36:29.480 --> 36:31.280] Have you had to insert a continuous sequence? [36:31.780 --> 36:35.260] You can... you can try to put a reporter gene on any particular gene. [36:37.680 --> 36:38.160] Um... [36:38.160 --> 36:44.860] You... what you don't know is if you attach a glob onto the end of some gene, if the original protein will be affected somehow. [36:45.480 --> 36:49.140] You know, sometimes the protein will cease to function if it's got a reporter gene attached to it. [36:50.420 --> 36:52.600] In that case, you can just try taking the promoter. [36:52.800 --> 36:55.700] You know, it's okay if a cell has two copies of a promoter. [36:56.020 --> 37:04.000] So you can leave the original gene in place with its promoter and its gene, but then copy the promoter and hook a reporter up to that and put it somewhere else. [37:04.060 --> 37:09.460] So that way you can sort of view the activity of the reporter with the reporter gene. [37:10.360 --> 37:10.800] Yeah? [37:10.920 --> 37:12.820] How reversible is this process right now? [37:12.920 --> 37:14.920] It's not... well, how reversible is it? [37:15.060 --> 37:25.440] That's a good question because, for instance, if you're going to engineer a virus to go attack cancer cells, you want to know that you can control that event and that something's not going to happen to it. [37:25.760 --> 37:30.040] So reversibility is sort of an active area of, you know, engineering. [37:33.020 --> 37:34.540] And also it depends, though. [37:34.760 --> 37:35.680] Like E. [37:35.800 --> 37:41.820] coli and yeast, which are transfected with things, depending on how you transfect them, they can sometimes kick out the DNA and not have it anymore. [37:42.960 --> 37:43.360] Yeah? [37:43.520 --> 37:45.800] What's the mechanism for reducing the reporter genes? [37:46.480 --> 37:47.200] It depends. [37:47.340 --> 37:48.460] It varies from organism to organism. [37:48.720 --> 37:57.400] So typically you would sort of try to clone the reporter gene onto a piece of DNA in the lab, but then getting it into the genome and into the organism, it just varies. [37:59.180 --> 38:01.820] Like the... well, yeah, it just varies. [38:03.620 --> 38:08.180] So here's a guy that... he's a developmental biologist, but now he's into systems biology. [38:09.380 --> 38:14.920] And, you know, he makes the comment that scientific fields like species arise by descent with modification. [38:17.100 --> 38:21.440] So things have been the same in molecular biology for the past couple of decades. [38:21.760 --> 38:25.260] But now, with all this information and all these genes, things are changing. [38:25.260 --> 38:30.240] to sort of this... to this notion, that people are surrounded by piles of genes and proteins. [38:30.740 --> 38:34.700] So the idea that biology is parts is kind of out there. [38:34.900 --> 38:37.740] And if you're surrounded by parts, don't you sort of want to do something with them? [38:38.020 --> 38:41.040] If you think you understand the system, can't you put them together in some new way? [38:43.940 --> 38:46.560] So what would be the simplest thing you could do with some parts? [38:48.600 --> 38:52.780] So these guys decided, well, let's see if we can do something very, very simple. [38:53.200 --> 38:57.680] Let's see if we can design a cell which exists in either of two states. [38:58.100 --> 38:59.520] And we'll call it a toggle switch. [39:02.120 --> 39:07.180] And as a construct, this is what it would look like, given the things sort of I've already told you. [39:08.340 --> 39:12.480] You have a reporter gene, so you've got something that you can track and see if it's on or not. [39:13.620 --> 39:14.400] And that's it. [39:14.480 --> 39:16.540] If the cell is expressing the reporter gene, it'll be green. [39:17.300 --> 39:19.440] If it's not expressing the reporter gene, it won't be green. [39:20.660 --> 39:29.400] And so how can you sort of make that cell either exist in a green state or an odd green state, using off-the-shelf components? [39:30.300 --> 39:31.580] Well, here's one way to do it. [39:32.860 --> 39:34.620] You can take two different promoters. [39:35.060 --> 39:39.540] Each of these promoters is a repressible promoter, meaning it can be turned off with a protein. [39:39.540 --> 39:46.240] And you make each promoter encode, I mean, control the repressor of the opposite promoter. [39:47.540 --> 39:50.100] So this is sort of a confusing way to look at it. [39:50.180 --> 39:53.920] But the other thing is, remember, repressors can be inactivated by inducers. [39:55.280 --> 39:56.420] Here's just a different view. [39:56.780 --> 39:57.280] Same thing. [39:57.680 --> 39:59.160] Two promoters facing apart. [39:59.400 --> 40:01.780] Each one is going to control the expression of a gene. [40:03.060 --> 40:07.420] And depending on the configuration of this whole thing, a reporter will be expressed. [40:07.420 --> 40:16.360] So if you can... if you turn... if this promoter is on, it's going to make this repressor by expressing transcripts, which gets translated into protein. [40:16.600 --> 40:19.920] So if this promoter's on, this promoter's going to be off. [40:21.640 --> 40:24.780] If this promoter's off, you're not going to make the reporter gene. [40:25.000 --> 40:26.180] The cells are not going to be green. [40:26.580 --> 40:27.540] That's one state. [40:28.720 --> 40:39.380] The other state is, if this promoter is on, then you're going to express this whole transcription unit, Meaning you are going to make a repressor and a reporter gene. [40:39.760 --> 40:45.600] That repressor is going to turn this promoter off, so you'll have green cells and an off promoter. [40:46.000 --> 40:47.080] So that's it. [40:47.320 --> 40:53.120] It's that simple that you've designed a cell which exists only in either of two states. [40:53.380 --> 40:54.880] Not partially between the states. [40:54.920 --> 40:56.900] It's either green or it's not green. [40:57.120 --> 41:01.100] So what characteristics does this have that people might think about? [41:01.400 --> 41:05.840] One is that it has memory. [41:06.540 --> 41:09.580] Because these repressors respond to things. [41:10.820 --> 41:18.160] So the cell can exist in one of two states and you can add the inducer of either repressor whenever you want. [41:18.680 --> 41:24.220] So consider, for instance, like if you design a repressor to recognize some environmental contaminant. [41:25.100 --> 41:27.560] And these cells are sitting out there in a stream bed. [41:28.180 --> 41:33.240] And they're supposed to recognize, I don't know, if any cadmium comes by. [41:34.720 --> 41:38.160] If the cells are exposed to cadmium, maybe they'll turn green. [41:38.380 --> 41:40.980] And they're going to stay green even after the cadmium goes away. [41:41.100 --> 41:43.880] Because once the switch has been flipped, it stays in that state. [41:45.920 --> 41:48.400] But then you can also reset it by adding the other inducer. [41:48.580 --> 41:51.460] So you can stably make the cell go into one state or the other. [41:51.740 --> 41:53.520] It'll stay in that state so it's got memory. [41:53.680 --> 41:54.780] And you can always reset it. [41:55.340 --> 41:57.280] So this is a really simple example. [41:57.420 --> 42:01.240] It's probably one of the simplest examples of taking components and doing some engineering. [42:01.440 --> 42:04.300] Some intentional engineering to make the cell do one of two things. [42:04.900 --> 42:06.680] And this is what the DNA would look like. [42:06.860 --> 42:09.600] So you can make this in the lab or you can order it from a company. [42:09.780 --> 42:11.320] There's companies which do DNA synthesis. [42:12.680 --> 42:18.760] But it's basically just putting a bunch of components together in a particular configuration. [42:18.760 --> 42:22.120] So this is an example of what you would call bioengineering. [42:22.680 --> 42:23.340] Did it work? [42:23.700 --> 42:24.400] Yeah, it works. [42:24.760 --> 42:25.080] It works. [42:25.160 --> 42:26.580] So they were able to make the green E. coli? [42:26.920 --> 42:28.220] Yeah, they can make the green E. coli. [42:28.760 --> 42:33.340] But it has, I mean, you could imagine, here's an example. [42:33.480 --> 42:40.620] A few years ago, there used to be this column in Nature where this guy would come up with wacky ideas. [42:40.760 --> 42:41.820] So this is like 20 years ago. [42:41.940 --> 42:46.220] This guy said, why don't we just take heroin addicts and infect their gums with E. coli that make heroin. [42:48.340 --> 42:49.980] Because then, or methadone. [42:50.640 --> 42:54.480] Because then you can just, you know, have all the methadone you want. [42:54.880 --> 42:56.480] You won't go breaking into people's houses. [42:56.740 --> 42:58.780] And if you ever want to get off it, just take some antibiotics. [42:59.280 --> 42:59.580] Right? [43:00.880 --> 43:02.440] But it was a joke back then. [43:02.560 --> 43:09.300] But this would be a system where if you need some drug, you can induce the production of that drug for some short time. [43:09.480 --> 43:11.360] And when you want to turn it off, you can induce its off. [43:11.460 --> 43:12.040] Its offness. [43:12.300 --> 43:14.780] So the reporter, you know, I said it could be green. [43:15.020 --> 43:16.300] But the reporter could be anything. [43:17.780 --> 43:19.940] So there's a lot of uses for this kind of thing. [43:21.920 --> 43:25.820] And, but again, that's probably the most basic thing that one could think of to do. [43:28.380 --> 43:29.120] And it's hard. [43:29.220 --> 43:31.340] Most molecular biologists don't think in terms of this stuff. [43:31.380 --> 43:35.000] Because usually biology is way too complicated to try to do something intentional like that. [43:35.660 --> 43:41.240] But that doesn't mean that given all these parts and given certain principles that you can't try to do something. [43:41.940 --> 43:45.360] And so it turns out that engineers have been dealing with these kinds of problems for a long time. [43:46.560 --> 43:51.140] First of all, one reason why it's hard to do these things in molecular biology is because there's no standardization. [43:51.620 --> 43:53.540] None of the parts fit together in the same way. [43:53.740 --> 43:59.120] If in my lab I'm working on something and someone else's lab they're working on something, we're not going to do it the same way. [43:59.380 --> 44:01.540] When our pieces aren't going to fit together in the same way. [44:01.780 --> 44:06.500] So there's been no effort at standardizing stuff, especially outside the Chinese community. [44:07.320 --> 44:10.720] But that doesn't mean that it's not possible to make certain standard fittings. [44:11.020 --> 44:13.640] The second thing is something called abstraction. [44:14.780 --> 44:19.980] One way to deal with complexity is to just hide it and make it not necessary for you to know it. [44:20.280 --> 44:25.580] For instance, if you buy a bunch of stereo components, you can hook them together without knowing anything about voltages. [44:25.980 --> 44:28.620] You don't have to have an oscilloscope or a voltmeter to hook up your stereo. [44:29.400 --> 44:30.820] And that's because of abstraction. [44:31.940 --> 44:38.040] And the third principle is something called decoupling, where a lot of biological elements are just way maybe over-engineered. [44:38.200 --> 44:40.340] They have so many functions that we don't know about. [44:40.440 --> 44:41.480] There's hidden stuff in there. [44:42.600 --> 44:45.340] And a lot of that stuff could be re-engineered to be simplified. [44:45.620 --> 44:52.900] If you want a protein which is one function, study it and recode it so it just has that one function and not ten other functions. [44:54.340 --> 44:56.640] So these are principles that one could apply to biology. [44:56.840 --> 44:59.980] And if you do apply these principles, you can get stuff done. [45:00.120 --> 45:03.360] So there's been a group at MIT who's trying to do these kinds of things. [45:03.920 --> 45:14.280] So here's something called the registry of standard biological parts, where they look into the genome and they say, look, we all use promoters, we all use activators, we all use repressors. [45:14.520 --> 45:16.140] Why don't we characterize these things? [45:16.380 --> 45:19.280] Let's measure their activities so we can report them in standard ways. [45:19.840 --> 45:24.020] Let's give them the same fittings on the end so that we can put them together in a standard way. [45:24.280 --> 45:24.720] Yeah? [45:25.000 --> 45:31.260] Is there a risk of the by-dream things like this in the genome of nature? [45:31.380 --> 45:31.920] Oh yeah, yeah. [45:32.140 --> 45:33.920] So the question is, is there a risk for doing this? [45:34.280 --> 45:34.880] There is. [45:36.060 --> 45:38.640] I'll try to address that in the end. [45:41.600 --> 45:43.820] But you can make catalogs of parts. [45:43.880 --> 45:46.240] And if you can make catalogs of parts, can you do something interesting? [45:47.640 --> 45:51.900] So let's just apply these engineering principles to a simple system. [45:53.520 --> 46:00.200] Here's where you want to engineer some bacteria, which usually don't smell very good, to do either of two things. [46:00.380 --> 46:13.040] If they're growing an exponential phase, so if you take some liquid broth and put bacteria in there, they'll grow in what's called like wide open, you know, pedal to the metal, full out mode, where they're consuming all the nutrients and they grow at a certain rate. [46:13.160 --> 46:16.540] But at some point they exhaust all the nutrients and they slow down and that's called saturation. [46:16.860 --> 46:20.380] So that's two different phases, exponential and saturated. [46:21.620 --> 46:26.380] And this group wanted to make it so that when they were growing in saturation phase, they smelled like mint. [46:26.840 --> 46:29.920] But when they exhausted all the nutrients and stopped growing, they smelled like bananas. [46:32.100 --> 46:37.180] It's not that hard to do because there's a gene in petunias which can take known metabolites in E. coli and convert those to something which smells like mint. [46:40.900 --> 46:46.960] There's a gene from yeast which can convert a known alcohol into an acetate, which smells like bananas. [46:47.740 --> 46:52.460] So if you wanted to make this construct, you would simply have to take E. coli and these two genes and make something called devices. [46:57.660 --> 46:59.740] So here's where abstraction gets involved. [47:00.000 --> 47:08.080] Instead of thinking of the complexity, if you could go to the standard registry of biological parts and pull off these devices from the shelf, then you could put them together. [47:08.400 --> 47:10.920] In this case, the chassis is a strain of E. coli which somebody has already made so that it doesn't stink very much. [47:14.760 --> 47:16.240] It doesn't really smell like anything. [47:17.140 --> 47:18.480] It's an off-the-shelf part. [47:19.920 --> 47:24.620] Each of these devices is composed of those two genes I showed with some regulatory stuff. [47:25.540 --> 47:28.060] So in the end, you've got two reporters. [47:28.320 --> 47:30.620] One is going to smell like mint, one is going to smell like banana. [47:30.940 --> 47:35.660] But remember, they have to be controlled by the cycle of the cell, like what state the cells are in. [47:36.120 --> 47:37.740] So you can use a promoter. [47:37.880 --> 47:43.580] In this case, the promoter here is a promoter that's on when the cells are in stationary phase. [47:44.660 --> 47:48.240] But it's not on when they're in rapid growth phase, okay? [47:49.020 --> 47:50.440] But that's really all you need. [47:50.480 --> 47:56.440] Because you can take that same promoter and go to the registry of parts and get something called an inverter. [47:57.340 --> 48:02.000] So what the inverter does is it says, it just inverts the signal coming from here. [48:02.160 --> 48:09.940] So if this is going to be on in stationary phase, then you can, then putting an inverter in front of that element, that control element, will invert its activity. [48:10.260 --> 48:13.420] So now you're going to control this gene and make it do the opposite of what that does. [48:14.000 --> 48:19.900] So I haven't even told you what an inverter is at the molecular level, but you don't even have to know what it is, because that's abstraction working for you. [48:20.600 --> 48:24.400] And of course, the inverter has parts, which cause it to have inverter activity. [48:25.820 --> 48:33.920] So there's a whole movement of, and it's founded by these guys, Tom Knight and Drew Endy, called the Internationally Genetically Engineered Machines Competition. [48:34.280 --> 48:38.560] It started just a few years ago, and every year it's been more than doubling in size. [48:39.160 --> 48:45.900] And the idea is for people to get together and think of ways to take parts and put them together in novel ways to do things. [48:46.460 --> 49:01.820] Here's just a couple of examples of a cell-cell signaling network, where cells with the same proteins are signaling each other, and depending on their location, depending on their growth location, they're sending out signals and responding in different ways. [49:03.640 --> 49:06.520] These are E. coli cells, which are cycling through three different colors of proteins. [49:07.400 --> 49:16.800] Here's a network for engineering three different cell types to play the Simon game, where the cells will only respond if they're given three cues in a particular sequence. [49:17.100 --> 49:19.040] If you mess up the sequence, they won't respond. [49:21.100 --> 49:23.380] These are E. coli, which have been engineered to respond to light. [49:23.580 --> 49:26.240] So in response to light, they pump out blue molecules. [49:28.380 --> 49:32.040] Here's an example of using yeast to grow a malaria drug. [49:33.120 --> 49:38.760] And this is an example of getting mice to see a color that they couldn't see before. [49:39.580 --> 49:45.400] It's not necessarily a circuit, but it's just a way of changing and adding a new activity to a system that I thought people would find interesting. [49:47.960 --> 49:51.860] So I think the last thing is decoupling. [49:52.020 --> 49:53.300] I'll just go through it really quickly. [49:53.460 --> 49:55.160] The idea here is to reduce complexity. [49:56.080 --> 49:57.420] This is a viral genome. [49:57.720 --> 49:59.620] Viral genomes are packed really, really tight. [50:00.020 --> 50:03.780] They're hard to study because there's so many elements which overlap. [50:04.060 --> 50:06.260] So here's two genes which overlap with each other. [50:07.060 --> 50:08.280] So they're hard to study. [50:08.340 --> 50:09.380] You don't know why they overlap. [50:09.520 --> 50:12.560] You don't know if there's some... if they have to overlap. [50:13.480 --> 50:24.020] So this guy, Drew Endy, took this T7 genome, which I've shown here, and he separated all the overlapping genes and basically made 600 simultaneous changes to the genome. [50:24.900 --> 50:27.760] And yet the phage are still alive after all that. [50:28.820 --> 50:29.260] Huh? [50:29.500 --> 50:30.560] It's a form of compression. [50:30.760 --> 50:36.200] It's a form of compression, but the notion of biological compression is difficult. [50:36.660 --> 50:43.840] And you know, this was the first example of someone taking a genome-level approach, making so many simultaneous changes, and the thing is still alive. [50:44.820 --> 50:48.120] So this is something that people are doing now, is trying to tweak genomes. [50:48.600 --> 50:52.200] And how many changes... so you're going to simplify the system, so it's easier to study. [50:52.520 --> 50:56.380] This could be really important if you want to try to design a virus that's going to go attack cancer cells. [50:58.620 --> 51:05.580] So that's... I'm just going to stop there because there's... well, last thing is... parts, we need more parts. [51:05.740 --> 51:08.440] If you want to make parts, find parts and use parts, we need more parts. [51:08.540 --> 51:15.080] So this guy went out and filtered seawater around the earth and ended up isolating like a million genes just from seawater. [51:16.800 --> 51:18.020] So now we've got lots of parts. [51:20.980 --> 51:29.860] It's very in principle, very important that... these engineering principles I talked about, there's no known engineering principles for dealing with machines that reproduce or evolve. [51:31.880 --> 51:48.500] This is essential to think about because, like you said, there are risks here, and dealing with them is going to require a lot of open standards, an open community, sharing information, much like the same ethics that exist in the hacker community, and so on. [51:49.000 --> 51:50.800] So that's it. [52:02.050 --> 52:02.570] Question? [52:16.350 --> 52:18.970] Yeah, so the question has to do with cost. [52:19.590 --> 52:28.730] It's a... it can be expensive to do biological research, but if you've ever thought about biotech startups, they often take place in people's garages. [52:28.910 --> 52:34.270] So there are ways to... to... there are ways to get things done that aren't that expensive. [52:34.590 --> 52:40.430] But there's also things like Make Magazine, which will tell you how to build a thermocycler, for like a hundred bucks. [52:40.850 --> 52:43.970] So you can do a lot of powerful things without a hundred dollars of technology. [52:44.910 --> 52:48.330] For that... for that hundred dollar piece of technology, you can build new genes, which... [52:49.650 --> 52:52.010] So it is possible to do things even without a lot of money. [52:57.660 --> 52:58.040] Yeah? [53:21.950 --> 53:24.710] Yeah, well the lunatics, that's why we need an open community. [53:25.830 --> 53:27.530] Because you have to be able to monitor each other. [53:27.630 --> 53:30.690] If somebody does something bad, you need a lot of knowledgeable people to respond. [53:31.290 --> 53:36.190] The last thing is that... that... you know, the defense department is thinking about this. [53:36.190 --> 53:42.310] So companies which do gene synthesis, if you decide not to do it yourself, you can order genes from a company, they're watching. [53:42.490 --> 53:45.570] If you order smallpox virus, you're going to have a knock on your door. [53:46.150 --> 53:47.470] So we have to be done now. [53:47.510 --> 53:49.030] I think... I think the next talk is coming up. [53:49.770 --> 53:52.370] Thank you. [53:52.790 --> 53:52.990] Yes? [53:53.010 --> 53:56.110] I have some other questions, if you are not running off elsewhere. [53:56.230 --> 53:56.770] I'm not running off. [53:58.070 --> 53:58.770] Fantastic talk. [53:58.930 --> 53:59.050] Yeah. [53:59.710 --> 53:59.750] Yes. [54:02.090 --> 54:02.610] Are you next? [54:02.610 --> 54:02.830] Are you up next? [54:03.010 --> 54:03.250] No.