[00:05.570 --> 00:07.090] This would actually be really difficult. [00:22.180 --> 00:23.600] I wish I had another machine. [00:27.200 --> 00:29.680] Yeah, the front panel is not coming up. [00:33.400 --> 00:33.880] Yeah. [00:38.900 --> 00:39.840] Make this easy. [00:41.200 --> 00:43.800] Kill time because I was up until 4 a.m. [00:44.040 --> 00:44.860] last night working on this. [00:45.180 --> 00:45.900] So, yeah. [00:55.640 --> 00:56.180] Finish. [00:56.180 --> 00:57.460] What the hell is that? [01:00.430 --> 01:01.970] Of course it starts out of minimized. [01:06.270 --> 01:06.810] Boo! [01:09.130 --> 01:10.350] That's a wrong one. [01:13.170 --> 01:14.070] There we are. [01:15.830 --> 01:16.730] There we are. [01:21.940 --> 01:22.840] Yeah, I noticed. [01:25.220 --> 01:26.200] Okay, well. [01:30.120 --> 01:30.620] Yeah. [01:31.400 --> 01:32.340] Okay, well. [01:32.840 --> 01:33.120] F*ck it. [01:33.220 --> 01:34.160] Maybe we can use a sidebar. [01:38.560 --> 01:39.060] Yeah. [01:39.940 --> 01:41.940] This is not this good. [01:42.180 --> 01:42.360] Okay. [01:49.100 --> 01:50.500] Well, sorry about that. [01:50.740 --> 01:52.080] My name is Matt Joyce. [01:52.400 --> 01:52.620] Hi. [01:53.380 --> 01:54.920] I'm sorry for the delay. [01:55.300 --> 01:59.100] As you can tell, I didn't actually expect this machine to completely foul up this bad. [01:59.100 --> 02:01.840] So, the art of do-foo. [02:02.620 --> 02:06.700] The question is, I get it from a lot of people, is what the hell is this talk about? [02:07.280 --> 02:11.020] So, obviously it's about data analysis, which that's not obvious at all. [02:11.020 --> 02:16.140] So, there's a lot of different components that go into data analysis and a lot of different fields that deal with it. [02:16.400 --> 02:19.540] Data mining is one that a lot of people like to talk about in security. [02:20.340 --> 02:26.660] Statistics, inference, there's a ton of different mathematical equations used to sort large volumes of data. [02:26.660 --> 02:29.760] And it has done a ton of different scientific fields that deal with it. [02:30.260 --> 02:43.240] So, I wanted to give a talk that kind of briefly touched on how to deal with large quantities of data and explain to people just how dangerous it can be that there's a lot of large quantities of data out there that we can all access. [02:44.040 --> 02:46.780] So, I also wanted to make sure that it was kind of entertaining. [02:47.000 --> 02:47.340] I'm sorry. [02:47.420 --> 02:49.000] Thus far, I've been incredibly boring. [02:49.200 --> 02:50.720] And that's my fault. [02:50.940 --> 02:51.960] You can all heckle me. [02:53.600 --> 02:54.560] You suck. [02:55.000 --> 02:55.600] Thank you. [02:56.240 --> 02:58.900] So, I've covered most of this. [03:00.060 --> 03:02.640] Does anyone, like, want to know what inference is? [03:04.140 --> 03:04.620] Yes. [03:04.820 --> 03:04.960] Okay. [03:04.960 --> 03:05.180] Yes. [03:05.580 --> 03:05.960] Okay. [03:06.200 --> 03:09.400] So, inference algorithms are basically learning algorithms. [03:09.640 --> 03:10.580] They infer things. [03:10.780 --> 03:12.860] They make assumptions and then they make predictions. [03:13.560 --> 03:14.040] Inference. [03:14.800 --> 03:18.620] So, 80% of this is showing up. [03:18.720 --> 03:19.700] This is obviously not true. [03:19.820 --> 03:20.700] I'm going to fail horribly. [03:21.020 --> 03:24.000] It's all breaking apart and I'm imploding inside. [03:26.120 --> 03:27.780] This isn't going to teach you statistics. [03:28.400 --> 03:29.820] Statistics is really f*cking hard. [03:30.120 --> 03:31.820] There's a lot of math that goes into it. [03:31.900 --> 03:35.800] There's a million different specialized equations for solving specific data sets. [03:35.800 --> 03:42.100] The idea here is coming up with a single set of data sets that can sort a large volume of data quickly and efficiently. [03:42.800 --> 03:44.380] That doesn't mean sort it right. [03:44.600 --> 03:46.060] Which brings us to our next slide. [03:47.940 --> 03:48.860] Enrico Fermi. [03:49.120 --> 03:50.300] Now, you can't really read that. [03:50.420 --> 03:52.260] But basically, Enrico Fermi is this guy. [03:52.700 --> 03:53.220] What? [03:54.160 --> 03:54.680] F5? [03:57.260 --> 04:02.750] Yeah, it's... Yeah, exactly. [04:08.350 --> 04:08.750] Yeah. [04:09.230 --> 04:14.610] So, I'm going to just not do that because my machine is about to fall apart. [04:15.170 --> 04:18.990] And I'll just cover it with my voice and hopefully this all makes sense in the end. [04:19.810 --> 04:21.890] When I get the applications up, it'll be okay. [04:23.570 --> 04:25.710] Enrico Fermi saw the first atomic bomb drop. [04:25.870 --> 04:27.630] The anniversary was three or four days ago. [04:27.770 --> 04:30.610] He's also known for a sort of type of math called Fermi math. [04:30.610 --> 04:32.870] Fermi math is usually taught to children. [04:33.170 --> 04:39.810] And the reason is because you can deal with very large sets of data, very large numbers, and you can solve things relatively quickly. [04:40.190 --> 04:46.990] So, the story that's famous about Enrico Fermi is while everyone else is watching the bomb explode, he's dropping pieces of paper. [04:46.990 --> 04:52.190] And he's watching the paper, waved in the wind, and he's actually calculating the blast yield from the bomb. [04:54.130 --> 05:01.090] So, small volumes of data, in some cases, with the right brain, the right causalities and correlations, you can figure out cool things. [05:01.850 --> 05:05.850] In this particular case, you can figure out the blast yield of a bomb by watching paper fall. [05:07.090 --> 05:10.970] So, data analysis is all about doing shit like that. [05:11.550 --> 05:19.550] Seeing something that looks seemingly unrelated to something else, but finding a correlation and defining a predictive causality. [05:19.790 --> 05:20.810] Doesn't have to be right. [05:21.210 --> 05:22.690] Doesn't have to be correct. [05:23.030 --> 05:23.830] We're not scientists. [05:24.070 --> 05:25.270] We're not writing PhDs. [05:25.450 --> 05:28.070] Most of us are in the security field or we're hackers. [05:28.250 --> 05:30.630] We're looking for a solution to a problem. [05:30.630 --> 05:34.090] That problem might be, how do I break into this website? [05:34.390 --> 05:48.850] Or, I'm doing a search on my own network and looking at all of our own CVS code, and I want to make sure that my developers aren't doing something stupid, like using usernames inside their application based off of, let's say, the name of the application. [05:52.130 --> 05:53.090] So, ethics. [05:54.050 --> 05:55.850] Some of you have them, some of you don't. [05:55.990 --> 05:58.450] When dealing with data analysis, it's very dangerous. [05:59.890 --> 06:03.670] If you look at data analysis the wrong way, you can basically confuse people. [06:04.110 --> 06:09.730] You can analyze things and do things like racial profiling and say that all Arabs are terrorists. [06:09.990 --> 06:10.790] That's not true. [06:11.450 --> 06:13.170] Most Arabs are very nice people. [06:13.290 --> 06:13.910] They're good people. [06:14.190 --> 06:17.990] But there are a large number of terrorists in the world who happen to be of Arabic descent. [06:18.450 --> 06:23.690] Therefore, racial profiling works on paper, but in reality, it's not fair to people individually. [06:24.290 --> 06:25.010] That's ethics. [06:26.690 --> 06:27.450] Be ethical. [06:27.690 --> 06:28.550] Don't be a douchebag. [06:28.670 --> 06:29.850] Unless, of course, you have no ethics. [06:33.340 --> 06:36.980] Or you're an American because we rock and we have manifest destiny and we can do whatever we want. [06:39.660 --> 06:40.340] The datum. [06:40.890 --> 06:41.650] What is data? [06:41.830 --> 06:42.450] Where is data? [06:43.020 --> 06:46.150] Basically, what I'm coming down to here is how can we catalog data? [06:46.220 --> 06:47.100] How can we capture data? [06:47.340 --> 06:48.580] Data exists everywhere. [06:49.120 --> 06:58.920] Knowing the number of people in this room, knowing the number of people with certain types of colors of eyes, knowing genetic information, knowing the temperature of the room, knowing how many lights are in the room, how many sockets, how many doors, how many windows, how many... everything. [07:01.480 --> 07:02.240] Data is everywhere. [07:02.400 --> 07:03.340] How do we catalog it? [07:03.480 --> 07:04.960] What data do we want to go after? [07:05.620 --> 07:07.460] Is it relevant to what we want to search for? [07:07.920 --> 07:14.210] So, we've got a little picture of Turing dressed up as the alligator hunter guy. [07:15.480 --> 07:16.620] So, trapping data. [07:17.020 --> 07:25.770] A lot of people for trapping data will use things like with the AMD project, you're basically coming in here and they're saying, hey, you've got this cool nifty badge. [07:25.770 --> 07:27.800] If you want to use it, give us some information about you. [07:28.420 --> 07:29.580] They're trapping your data. [07:30.150 --> 07:30.710] Collecting it. [07:31.020 --> 07:34.080] There's other ways of trapping data because data exists in the wild. [07:34.740 --> 07:38.720] Things like, and we'll get to them, CVS repositories are online. [07:38.720 --> 07:40.650] There's code online. [07:41.000 --> 07:43.300] There's websites with lots of information. [07:43.840 --> 07:48.340] I don't think it's Eventbrite, but there's something like Eventbrite for... [07:51.100 --> 08:00.220] If you ever go biking or running or you're in a marathon, these guys take pictures and they match it up to your number on your chest based off of an OCR algorithm. [08:00.450 --> 08:05.060] So, if you know someone's name and they're into this sort of thing, you can pull pictures of them online. [08:07.450 --> 08:11.620] So, there's a lot of data out there and it's pretty easy to find. [08:11.740 --> 08:16.080] If you're looking for something in particular, you'd be amazed at what you can find out there. [08:17.150 --> 08:20.000] Spidering mailing lists is one of the things I was looking into. [08:20.660 --> 08:21.830] It's readily available. [08:21.950 --> 08:23.270] It's basically pretty obvious. [08:23.710 --> 08:25.780] And it helps us understand communities. [08:26.330 --> 08:28.160] There's a lot of different mailing lists out there. [08:28.270 --> 08:28.840] They're all open. [08:29.520 --> 08:32.500] LKML, NYlog, NYC Resistor, all this stuff. [08:32.500 --> 08:47.500] And I wanted to do some experiments to show you that even in a cultural setting or a societal setting, you can see how analyzing data can be efficient or inefficient, as the case may be, in defining correlations. [08:48.220 --> 08:56.520] Is there a correlation between the size of a mailing list, the size of the emails, the number of people on the list, the content of the list? [08:56.520 --> 09:00.950] And does that correlate to the success of the project or the failure of the project? [09:01.600 --> 09:04.210] As it turns out, most of the case, that's not true. [09:04.960 --> 09:07.170] There's really almost no correlation in that stuff. [09:18.470 --> 09:18.950] Yep. [09:19.690 --> 09:23.470] So extrapolating data is really not that hard. [09:23.930 --> 09:26.010] You find a large data set on the Internet. [09:26.290 --> 09:28.690] For instance, the schedule of talks on the website. [09:29.490 --> 09:33.770] I'll actually show you a little bit of a script that can do this. [09:33.910 --> 09:36.750] But you look at the talks on the schedule of talks. [09:36.970 --> 09:47.790] The way it's broken down in HTML allows you to very easily extrapolate all that data and dump it into either a database or a flat file that you can easily sort and enter into applications for data mining. [09:48.510 --> 09:49.610] Extrapolating data is easy. [09:49.610 --> 09:51.450] Perl is great for it. [09:54.940 --> 09:56.420] This is my troll post. [09:56.820 --> 09:58.520] There's a lot of people who hate languages. [09:58.800 --> 10:00.040] Don't hate languages. [10:00.500 --> 10:02.300] Some languages are actually pretty useful. [10:02.580 --> 10:03.540] Some aren't. [10:03.920 --> 10:06.460] Perl is very useful at parsing large volumes of data. [10:06.740 --> 10:10.540] And there's a shit ton of modules written for it for dealing with Bayesian algorithms. [10:10.740 --> 10:16.120] Largely because of things like spam assassin and neural network programmers being lazy when it comes to parsing data. [10:17.360 --> 10:19.880] There are other languages as well that are really good. [10:20.060 --> 10:21.060] C is really good. [10:21.640 --> 10:22.700] That's actually bullshit. [10:22.880 --> 10:23.800] Lisp is also really good. [10:23.940 --> 10:25.520] I just wanted to piss off the Lisp programmers. [10:29.780 --> 10:31.580] Again, this isn't about math. [10:31.800 --> 10:33.140] There's a lot of math in this. [10:33.240 --> 10:34.200] It's about getting it done. [10:34.440 --> 10:37.380] You don't need to know statistics to use this stuff. [10:37.720 --> 10:43.900] The simple fact is a guy who's into computer security probably knows jack shit about statistics because he didn't go to school for it. [10:44.060 --> 10:45.540] He went to school for other things. [10:45.680 --> 10:46.500] He's not an economist. [10:46.780 --> 10:48.220] He's not a mathematician. [10:48.500 --> 10:49.060] Well, he might be. [10:49.060 --> 10:51.180] I know a couple of mathematicians who are into the field. [10:51.500 --> 10:56.220] But in most cases, if you're actually dealing with auditors, auditors are not brilliant people. [10:56.420 --> 11:01.240] They're people who are technically gifted and they're good at finding bugs. [11:01.940 --> 11:05.720] That doesn't necessarily mean that they're scientifically brilliant. [11:06.520 --> 11:15.780] Finding a bug in a data structure that isn't really logic but is, say, a correlation. [11:16.140 --> 11:17.000] Great word. [11:18.180 --> 11:25.100] Is something anyone can do without having super huge amounts of knowledge regarding statistics. [11:32.620 --> 11:33.220] Probability. [11:33.340 --> 11:37.660] There's two basic types of probability that we're going to discuss because we're doing a linear analysis. [11:37.660 --> 11:40.180] is forward probability and reverse probability. [11:40.500 --> 11:42.440] In one situation, you know the outcome. [11:42.580 --> 11:45.880] In the other situation, you don't know the outcome. [11:46.100 --> 11:49.060] So, in forward probability, you don't know the outcome. [11:49.180 --> 11:53.960] In reverse probability, you know the outcome but you don't know some of the variables going into the data set. [11:54.800 --> 11:59.520] With reverse probability and forward probability, usually this feeds into Bayesian algorithms. [11:59.520 --> 12:01.140] How do you want to solve it? [12:01.260 --> 12:02.100] What are you trying to predict? [12:02.300 --> 12:04.880] Are you trying to predict members of the data set or outcomes? [12:08.240 --> 12:11.220] Regression is very simply explained. [12:12.340 --> 12:21.580] Gauss is a famous guy because he claimed to have invented basically the first linear regression algorithms. [12:22.040 --> 12:28.630] And the reason that he's credited with it is because he helped a guy track a comet back in the 1700s. [12:30.090 --> 12:33.850] The late 1700s, 1795, 1800 area. [12:34.090 --> 12:36.510] A man lost a comet he was tracking behind the sun. [12:36.870 --> 12:40.430] Back then you didn't have a lot of ways to track celestial bodies. [12:40.690 --> 12:43.290] So they were trying to guess where it would reappear again. [12:44.130 --> 12:49.390] They had this man's record of where he saw it and tracked it in the sky, which wasn't exactly correct. [12:50.150 --> 12:56.530] But he was able to graph it and basically with regression you'll have a graph with data points on it. [12:56.530 --> 13:05.210] And you draw a line through it and try and define a function that defines how that data set can be modeled. [13:06.350 --> 13:14.090] With some things, especially astral bodies, it's really easy to get a very clean line and follow it through to where it's going to appear. [13:14.090 --> 13:14.750] It worked. [13:15.730 --> 13:20.310] So least squares is what we usually use for linear regression. [13:20.950 --> 13:22.350] It's pretty straightforward. [13:22.830 --> 13:28.950] Most of you who have done calculus or worked in a lab in college have probably done this, whether or not you know it. [13:29.210 --> 13:34.770] It was probably brought up in your lab class whenever you needed to do a linear regression. [13:34.770 --> 13:37.090] Even if you didn't know you were doing linear regression. [13:40.370 --> 13:40.850] Correlation. [13:46.270 --> 13:56.850] So, correlation is basically once you have two sets of data that you've linearly regressed, you can try and see if those waveforms, those functions, match up in any way. [13:57.010 --> 13:59.830] If they match up in any way, they may correlate. [14:00.050 --> 14:01.650] They may not, but they may correlate. [14:01.650 --> 14:09.810] And if they do, you might be able to find something out that's either a third factor or you might just be finding out what you wanted to know. [14:11.290 --> 14:13.850] So, that brings us to the next slide. [14:17.010 --> 14:20.490] Correlation is something that's usually used improperly. [14:20.910 --> 14:26.530] It's one of these things that's great about statistics is you can lie really, really well using statistics. [14:27.150 --> 14:30.770] Throwing numbers at people tends to shut people up and dump them into idiot mode. [14:30.770 --> 14:34.410] There are fallacies about correlations that are pretty straightforward. [14:36.410 --> 14:41.570] Sometimes, just because a waveform matches up with another waveform doesn't mean that one's the cause of the other. [14:41.690 --> 14:43.770] It just means that these things match up. [14:43.850 --> 14:47.510] There could be underlying causes, data sets that you're not actually taking into account. [14:48.170 --> 14:49.670] Or it could just be coincidence. [14:49.990 --> 14:51.170] It's unlikely, but it could be. [14:51.910 --> 15:04.170] Other things that run into a problem with correlations, when you graph stuff, sometimes you'll see it, it's very obvious when you graph it, that sometimes the way the waveform pops up is a function of how the spread happens. [15:04.530 --> 15:08.850] Sometimes the spreads can be completely different, but still bring up the same waveform. [15:09.430 --> 15:18.490] Which basically means your prediction is mathematically correct, but probability-wise, it's completely wrong. [15:18.490 --> 15:29.910] So, just going through the motions of solving a correlative function isn't going to necessarily bring you any closer to finding any real correlation. [15:33.350 --> 15:40.290] Naive Bayesian algorithms are usually used for correlation in data sets because they're very fast, but they are not usually very correct. [15:40.970 --> 15:46.650] They're good for classifying documents, which is largely what you're doing when you're doing data analysis on a network. [15:48.610 --> 15:50.410] Naive Bayes makes certain assumptions. [15:50.770 --> 15:52.690] Bayesian algorithms in general make assumptions. [15:52.970 --> 15:55.870] There are schools of statistics that are not fans of assumptions. [15:56.110 --> 16:01.370] If you go into your PhD thesis with a Naive Bayesian algorithm in there, odds are you're going to get tore up. [16:02.390 --> 16:06.630] Because if you're on a PhD level, they want your numbers to be exact and accurate. [16:07.270 --> 16:10.830] Bayesian algorithms don't produce exact and accurate results. [16:11.130 --> 16:16.670] Just like when Fermi was figuring out the blast yield of the bomb, his accuracy was based off certain assumptions. [16:16.670 --> 16:18.070] He was making guesstimates. [16:18.290 --> 16:18.830] Lots of them. [16:19.450 --> 16:22.490] You don't functionally need to always be correct. [16:22.670 --> 16:23.970] You don't functionally need this. [16:24.090 --> 16:25.690] And that's why Bayesian algorithms are great. [16:25.950 --> 16:33.290] When we were trying to discover a way to figure out how to stop spam from showing up in our inbox, some guy thought, hey, we don't need to be right all the time. [16:33.290 --> 16:35.610] We just need to be right 99% of the time. [16:35.610 --> 16:37.410] Bayesian algorithms work for that. [16:40.590 --> 16:41.930] But most of you know all this. [16:44.150 --> 16:44.630] Weka. [16:45.670 --> 16:48.370] Weka is... well, actually, I think I skipped a couple here. [16:48.950 --> 16:51.910] Yeah, okay, so here's a picture of not Thomas Bayes. [16:51.970 --> 16:53.490] He was a religious nut, an accomplished mathematician. [16:54.030 --> 16:58.690] He gave us a Bayesian theorem from which many other algorithms are derived, including Naive Bayes. [16:59.950 --> 17:01.150] That's Bayesian theorem. [17:01.510 --> 17:02.810] It's pretty straightforward. [17:02.810 --> 17:07.130] It's not super hard, but it is a shit ton of iterative math. [17:10.350 --> 17:10.790] Weka. [17:11.030 --> 17:12.010] Weka is really cool. [17:12.450 --> 17:16.310] Weka is like many other very expensive tools, except it happens to be free. [17:19.750 --> 17:20.190] Yes. [17:20.570 --> 17:22.710] So, the good side is it's open source. [17:22.930 --> 17:24.270] It means you're not going to pay a lot for it. [17:24.430 --> 17:25.610] It can be modified. [17:25.730 --> 17:26.850] There's a lot of plugins for it. [17:26.850 --> 17:29.170] It works with a lot of different data acquisition tools. [17:30.290 --> 17:34.890] The negative sides are... it's slow. [17:35.130 --> 17:36.770] And it stores everything in main memory. [17:36.810 --> 17:39.610] Which means if you've got a very large data set, it's going to crash. [17:39.850 --> 17:40.870] And it's going to crash hard. [17:41.570 --> 17:44.150] It does, however, have the option to work off of JDBC. [17:44.350 --> 17:47.310] And it can work off of a SQL server, which is pretty straightforward and useful. [17:47.890 --> 17:51.470] But the reason we use Weka is not because it is your end-all. [17:51.770 --> 17:55.050] The idea is to use Weka to figure out how you're going to model your data. [17:55.050 --> 18:02.990] You throw a small part of your data set in, figure out where the correlations lie, or how you want to go through it and look for correlations. [18:03.490 --> 18:05.450] Then you do it in another set of code. [18:05.930 --> 18:07.570] Probably see if it's a large data set. [18:07.670 --> 18:09.110] If not, you can probably model it in Weka. [18:09.790 --> 18:10.470] It works. [18:11.450 --> 18:11.850] Okay. [18:12.430 --> 18:14.370] This is going to get really dicey now. [18:21.900 --> 18:23.960] If I can see my mouse... [18:44.300 --> 18:44.700] Okay. [18:50.570 --> 18:50.970] Wow. [18:51.790 --> 18:53.110] I should have brought a spare machine. [18:56.170 --> 18:56.730] All right. [19:02.870 --> 19:03.690] You know what? [19:05.370 --> 19:07.130] I don't think I'm going to bore you with this. [19:07.590 --> 19:12.330] Weka is really cool, but I'm only going to show you a small part of it because I can't see the screen. [19:12.550 --> 19:14.310] Then I'm going to have discussion with something else. [19:17.610 --> 19:18.050] Open. [19:19.050 --> 19:19.750] Go up. [19:26.940 --> 19:28.780] No, that's not the right directory. [19:29.080 --> 19:29.760] Is that the... [19:32.240 --> 19:33.800] Yeah, I have no idea where I am. [19:35.880 --> 19:36.480] Yeah. [19:36.840 --> 19:37.360] Okay. [19:37.620 --> 19:38.960] So that's not the right directory. [19:41.220 --> 19:42.420] Do you... [19:42.420 --> 19:42.780] HOPE? [19:45.020 --> 19:46.220] Is that... [19:48.480 --> 19:49.080] Right. [19:51.600 --> 19:52.940] Should be in demo one. [19:56.200 --> 19:57.980] Is this because it's Java? [20:01.620 --> 20:02.220] Yes. [20:04.420 --> 20:05.020] Okay. [20:05.580 --> 20:07.000] What is it looking for here? [20:07.600 --> 20:08.200] EXP. [20:08.720 --> 20:09.180] All right. [20:10.540 --> 20:12.040] That's not going to help us. [20:15.280 --> 20:15.660] Man. [20:17.140 --> 20:18.220] Yeah, I can't see this. [20:18.440 --> 20:18.580] Okay. [20:18.800 --> 20:19.400] So Weka... [20:19.400 --> 20:20.800] I'm just going to explain it to you. [20:21.120 --> 20:22.700] Weka allows you to do visualizations. [20:22.700 --> 20:29.160] If you open up an ARFF file, which is basically kind of like really weak XML, you basically dump it in. [20:29.280 --> 20:31.220] You define your data types like a very simple database. [20:32.040 --> 20:35.240] So say you're doing a Perl parse off the schedule of talks. [20:35.640 --> 20:41.340] You can grab the title, the descriptive analysis, and the author. [20:41.760 --> 20:43.560] You've got three different sets of data there. [20:43.780 --> 20:49.300] You're going to describe basically other stuff if you want it is location, the size of the room. [20:49.360 --> 20:57.180] We know that Ingritia is the smallest room, which means obviously no one thinks highly of me, which is probably correct since my laptop doesn't work and I'm looking like a fool. [20:57.540 --> 21:00.240] But you throw all this data into Weka. [21:00.240 --> 21:04.340] You define the data types and then you can run certain sorts and searches. [21:04.880 --> 21:08.800] And you can run linear regression by itself in about 15 different algorithms. [21:09.080 --> 21:15.560] You can run correlation off of about 15 other different algorithms. [21:15.760 --> 21:18.040] And it'll give you a correlative coefficient. [21:18.320 --> 21:23.400] And if the correlative coefficient is near zero, it means that there's probably no correlation here. [21:23.620 --> 21:27.780] So functionally, if you're looking for data, you extrapolate a shit ton of data on your target. [21:27.780 --> 21:44.500] Great targets of past exploration have consisted of, and this asshole over here forgot to give it to me, was he's actually sitting on psychological profiles of hackers in the New York City area over the years. [21:44.500 --> 21:48.480] That would make a great data set. [21:48.580 --> 21:50.560] Another great data set I've used is paste bins. [21:51.200 --> 21:57.820] Paste bins contain code snippets and configuration files from the paste bins on Freenode. [21:58.020 --> 21:59.880] Does anyone not know what a paste bin is? [22:01.240 --> 22:08.700] Okay, so if you're on IRC and you're looking for help, someone gives you a website to paste your configuration or your source code into. [22:08.700 --> 22:10.860] And then you give them the link and they can see it. [22:11.360 --> 22:16.240] A lot of these are configured with either linear serial accounts or a recent post page. [22:16.700 --> 22:21.140] So you could develop a spider to sit on there and pick that shit up and you can play with it. [22:21.420 --> 22:31.100] Do simple stuff like searching and sorting keywords like password or MySQL Connect or just off of extrapolation you can find a shit ton of really cool stuff. [22:31.100 --> 22:36.540] But the really cool stuff is when you start looking for failures that exist in code. [22:36.720 --> 22:40.260] You can look for buffer overflows in printf. [22:40.400 --> 22:43.640] You can look for system calls with variables in them. [22:43.900 --> 22:50.260] And then you can start figuring out which developers, because they usually sign their names, are the bad ones. [22:50.340 --> 22:55.580] You can also see across the board what mistakes are likely to happen with certain types of code. [22:56.400 --> 23:03.000] This is useful data for a person who is a systems administrator and wants to make sure that his developers are not screwing up constantly. [23:03.280 --> 23:05.400] But doesn't want to get involved in looking at their code. [23:05.940 --> 23:09.840] You can just throw out at a meeting, hey, listen, you're a PHP developer. [23:10.020 --> 23:13.520] We just want to make sure that you're not using a shit ton of persistent connections constantly. [23:14.880 --> 23:19.660] Or, you know, you need to make sure that anything with a MySQL connected is not set up so that anyone can read it. [23:20.520 --> 23:21.980] You know, simple stuff like that. [23:23.840 --> 23:29.080] Schedule of talks, getting back to that, is when you extrapolate data based off of people and keywords. [23:29.080 --> 23:35.700] You can do lexical analysis thanks to an application by Princeton University called WordNet. [23:36.160 --> 23:42.720] WordNet allows you to look at a sentence, break apart the words, and pull the noun values out of the sentence. [23:43.080 --> 23:47.600] Using that, you can actually basically set up keywords. [23:47.600 --> 23:50.320] I'm looking for usage of the word hack. [23:50.680 --> 23:53.360] So it'll give you hack, hacking, all the other stuff you need. [23:53.760 --> 23:55.540] I'm looking for usage of data mining. [23:55.680 --> 23:58.660] I'm looking for usage of culture jamming. [23:58.740 --> 24:01.320] I'm looking for usage of whatever's happening. [24:01.460 --> 24:03.540] And you could basically look at the... [24:03.540 --> 24:06.400] I actually wish I could show you, but I got the Perl code to do this. [24:06.960 --> 24:12.100] You can extrapolate, pull this data out, look at those numbers, and dump them into the Weka. [24:13.560 --> 24:18.580] So you can look at the numbers in such a way where you're saying, okay, we're all in this room. [24:18.720 --> 24:19.680] We're all going to talk. [24:19.860 --> 24:21.040] Certain talks are in bigger rooms. [24:21.160 --> 24:23.000] Certain talks are obvious of greater value. [24:23.240 --> 24:27.260] So what this year is of the greatest importance to all of us? [24:28.320 --> 24:31.560] It's pretty easy to figure that out just based off the weight of numbers. [24:33.000 --> 24:35.160] But you might also look for correlations. [24:35.860 --> 24:40.180] What people throughout the years are seeing this and going, hey, this is great. [24:40.820 --> 24:48.260] And added to that with the attendee metadata project, you can actually see who's going in and out of rooms based off of the location tracking on here. [24:48.860 --> 24:55.020] So we can know whether or not a lot of people come to a talk and leave because it sucks. [24:55.020 --> 24:56.020] Like my talk, I'm sorry. [25:00.100 --> 25:01.300] Moving on from there. [25:02.060 --> 25:02.820] Extrapolating data. [25:03.180 --> 25:07.340] So what it basically comes down to is you first want to define your data set. [25:07.440 --> 25:07.860] You're looking. [25:08.040 --> 25:08.360] You're hunting. [25:08.400 --> 25:12.020] You're saying, I want to know more about a specific area. [25:12.800 --> 25:16.460] Let's say I work as a system administrator. [25:16.460 --> 25:23.060] So what I'm concerned about is whether or not our sensitive data sets are available to someone or predictable. [25:23.640 --> 25:26.280] So I want to define, here's my sensitive data. [25:26.400 --> 25:27.940] Here's the stuff that's connected to it. [25:28.840 --> 25:31.900] I pull all of that into a room based off keywords. [25:31.980 --> 25:35.360] And I throw that into Weka and start analyzing. [25:35.580 --> 25:41.400] I look for correlations between the full set of code base and path data for all of our web apps. [25:43.120 --> 25:47.000] Do that and we'll end up with possible correlations on usernames. [25:47.480 --> 25:50.840] Possible correlations in usage of functions by specific developers. [25:51.040 --> 25:51.860] Things of that nature. [25:51.860 --> 25:59.300] And you'll be able to, in theory, shut down uses of predictable data. [25:59.520 --> 26:05.320] Where, say, a developer likes to use this God username, God. [26:05.960 --> 26:09.200] So one of your developers likes to use God a lot, hides it in his code. [26:10.140 --> 26:12.040] You don't want that popping up a lot. [26:12.520 --> 26:18.660] So developer X correlates with rapid recurring God in username. [26:19.720 --> 26:29.400] You see that pop up because instead of doing the searching of all the code yourself, you're running it against Weka or an application that's otherwise doing linear regression. [26:29.660 --> 26:34.440] And that you're extrapolating it, doing linear regression, then doing correlative analysis. [26:34.940 --> 26:47.760] And if the correlative coefficient is above zero by a pretty decent margin, say, closer to 2 to 90, you might want to be sitting there going, hey, this might be something I want to take a look at. [26:47.920 --> 26:49.880] You look at the graph without even looking at the data. [26:49.940 --> 26:52.020] Just look at the graph and see if the graphs match up. [26:52.020 --> 26:56.040] If they do, then you look at the data and say, could there be a correlation here? [26:56.140 --> 26:57.600] Could there be a causality that makes sense? [26:58.560 --> 27:03.400] If not, then obviously you didn't find anything. [27:03.560 --> 27:05.720] There's a lot of data analysis that goes bad. [27:05.920 --> 27:11.900] You won't find anything as we didn't get to see with the mailing list. [27:12.040 --> 27:16.800] Mailing list, in theory, you would think that high traffic means that the project's doing well. [27:17.920 --> 27:19.320] Usually that's not the case. [27:19.680 --> 27:23.340] A lot of projects have a lot of high traffic and don't actually do jack shit. [27:25.160 --> 27:26.260] But some do. [27:27.700 --> 27:35.280] The other thing that you would think is that you can push things like looking at... [27:35.280 --> 27:36.360] How do I explain this? [27:36.500 --> 27:37.960] Okay, so you have a mailing list. [27:38.280 --> 27:40.400] Sometimes you won't see certain things immediately. [27:40.700 --> 27:43.400] And sometimes you'll be able to use things that you didn't otherwise think you could. [27:43.620 --> 27:47.920] For instance, when you're looking at a mailing list, archive, you'll pull up... [27:47.920 --> 27:51.100] A lot of them will have gzip digest, some won't. [27:51.560 --> 27:56.920] Can you pull the values of gzip digest and use them against the values of the non-gzipped? [27:57.680 --> 28:04.640] Usually yes, because you can actually figure out, predictably, what the difference in size will be. [28:04.640 --> 28:17.520] We don't care about exact numbers, so you can actually generate, using predictability, naive beige, what the actual value difference should be on a scale. [28:17.700 --> 28:23.200] Ranging from starting at 5k on up to like 96k on up to 70k digest. [28:25.800 --> 28:27.100] I am dying up here. [28:29.180 --> 28:32.020] I'm going to take a question or two, see if anyone's lost. [28:32.200 --> 28:32.320] What? [28:39.030 --> 28:39.950] Say the question again. [28:40.210 --> 28:40.810] How much [28:45.640 --> 28:46.160] does... [28:49.700 --> 28:50.540] Define representation. [28:50.920 --> 28:53.220] Well, so, I mean, you said you can look for like... [28:54.180 --> 28:54.600] Yeah. [28:54.800 --> 28:55.860] It seems like... [28:55.860 --> 28:56.180] Oh. [28:57.320 --> 28:57.880] It... [28:57.880 --> 28:59.980] In some cases, it screws it up thoroughly. [29:00.160 --> 29:01.400] In some cases, it doesn't. [29:01.740 --> 29:03.240] It really depends on what... [29:03.240 --> 29:05.320] This is what correlative analysis is about. [29:05.760 --> 29:12.720] You need to look at the correlation and extrapolate all other data and isolate only the ones that matter. [29:12.920 --> 29:14.840] That doesn't mean all but two. [29:14.980 --> 29:16.320] That means only the ones that matter. [29:17.060 --> 29:23.280] Sometimes, the underlying stuff that does matter is extrapolated as well. [29:23.920 --> 29:25.780] For instance, there was a... [29:26.620 --> 29:37.080] There was a weapons testing program done by the military where they set up a camera that was supposed to look at an image and define whether or not it was an enemy or not an enemy. [29:37.640 --> 29:39.520] And it was working in the lab every day. [29:39.800 --> 29:40.500] Works fine. [29:40.660 --> 29:43.340] They assume, you know, the picture's an enemy, not an enemy. [29:43.520 --> 29:44.100] It's pretty straightforward. [29:44.500 --> 29:46.940] But they bring it outside and all of a sudden it's not working. [29:47.540 --> 29:48.320] Why is that? [29:48.860 --> 29:51.540] Pictures taken of it, enemy day, we're nice day. [29:52.140 --> 29:53.720] Allied day, not a nice day. [29:54.480 --> 29:57.020] It started picking up the weather patterns instead of the enemy, not enemy. [29:57.380 --> 30:01.220] So, there's a correlation there that isn't actually correct. [30:02.040 --> 30:04.000] So, I guess, is that... doesn't that answer your question? [30:04.280 --> 30:05.000] Okay, yeah. [30:15.490 --> 30:15.910] But... [30:15.910 --> 30:23.390] Well, you can try and do predictive analysis and predict what the values will be and match it up against what actually happens. [30:23.390 --> 30:29.330] You're basically using your own analysis as test data for prediction analysis and see if it's actually matching up. [30:30.810 --> 30:34.090] Sometimes, you might be completely wrong, but completely right. [30:34.590 --> 30:42.130] Looking at Newton's laws, for instance, you'll see that he was pretty much right in all cases, but his laws are still completely wrong. [30:43.370 --> 30:49.810] So, we don't care if we correlate wrong so long as our prediction is correct functionally. [30:49.810 --> 30:56.510] So, whether or not the representation is wrong or right doesn't matter so long as the prediction works. [30:56.790 --> 31:01.370] If you happen to get it right without knowing why, functionally, why not use it? [31:01.430 --> 31:08.210] You might run into trouble down the line, but, you know, in this case, for our usage, it usually isn't that important. [31:08.290 --> 31:09.450] We're not building nuclear weapons. [31:13.170 --> 31:13.970] I can't hear you. [31:14.070 --> 31:15.290] Can I just say something about that? [31:15.450 --> 31:15.670] Shoot. [31:15.670 --> 31:33.860] Can you flip on an outfit that you can use if you have very large financial data before you're grinding out if you know what the promising dimensions are? [31:35.660 --> 31:37.160] Yeah, you guys will wait at yourself. [31:39.320 --> 31:40.780] But, anything else? [31:45.180 --> 31:47.140] Nobody can predict the stock market, actually. [31:47.340 --> 31:49.500] Most of it's based off of immediate prediction. [31:49.980 --> 31:55.020] Actual prediction is almost impossible because you have so many different systems working against each other. [31:56.940 --> 31:59.400] You do have to predict... you can predict major shifts. [31:59.760 --> 32:03.240] Large companies will do major things based off events that are predictable. [32:04.820 --> 32:06.820] But, with... it depends on what you're talking about. [32:06.940 --> 32:07.920] There's two different types of trading. [32:08.060 --> 32:09.780] There's day trading and then there's long-term trading. [32:10.760 --> 32:16.240] And then there's other problems like guys who use the small trading systems to force through larger trades. [32:18.780 --> 32:22.180] In specific scenarios, you can predict major shifts. [32:22.640 --> 32:28.480] In day trading, you can predict day trading responses because day traders all generally respond in the same way to certain actions. [32:28.920 --> 32:34.020] They'll sell off all of their stocks if they see a dip below a certain amount. [32:34.180 --> 32:35.920] A lot of these guys follow scripts of sorts. [32:36.140 --> 32:38.920] They work out their own methods and they work as quick as possible. [32:39.380 --> 32:43.080] So, in a day trading situation, you put a machine with a simple predictive algorithm in there. [32:43.400 --> 32:46.860] As long as it's hooked up and moving quickly, it'll move faster than the guy. [32:47.020 --> 32:49.940] And it'll probably be one of the first guys into the day trading and get in and out real quick. [32:50.220 --> 32:53.760] Which means, predictively, it doesn't have to predict anything major. [32:53.940 --> 32:57.660] It just has to get in and out quick on a regular basis and it will make money. [32:58.120 --> 32:59.620] A lot of guys have systems like that. [32:59.620 --> 33:04.240] They're not super intelligent, but they can predict movement and then act on it. [33:05.840 --> 33:12.400] With larger long term shifts predicting large movements of money, you can do that based off certain events. [33:12.580 --> 33:18.040] You see the price of oil going up, you know that alternative fuel sources are suddenly going to do well. [33:19.100 --> 33:21.020] But, you also have the human element. [33:21.200 --> 33:22.880] Why is the price of oil going up? [33:23.580 --> 33:28.440] Sometimes, it's simply because somebody decided to buy a shit ton of oil to affect the market on purpose. [33:28.440 --> 33:37.920] Like the instance where a guy several years ago, I guess it was a year or two now, pushed the price of oil up to $100 a barrel just so he could say he was the guy who did it. [33:39.480 --> 33:43.720] There are problems with predicting data sets that large. [33:43.900 --> 33:48.560] Especially with exchanges where there's so many different elements coming in and out. [33:49.560 --> 33:55.260] That you can functionally, like we're talking about here, but doing actual prediction on it's very difficult. [33:57.680 --> 34:00.480] I mean, that would be my approach to it. [34:01.420 --> 34:04.300] Does anyone else have any comments or suggestions on that? [34:04.860 --> 34:09.620] Have you kind of brainstormed about anomaly detection? [34:13.440 --> 34:15.880] Anomaly detection for intrusion detection systems. [34:16.080 --> 34:21.120] Anomaly detection is used a lot in financial systems or spending monitoring. [34:21.280 --> 34:21.900] Like auditing? [34:22.140 --> 34:22.760] Yeah, auditing. [34:23.800 --> 34:25.060] Are there anomalies? [34:25.380 --> 34:26.700] Actually, this is really useful. [34:27.020 --> 34:29.520] Anomalous detection is really useful for a couple different things. [34:30.480 --> 34:36.700] Say you have a guy who's a programmer and you're watching the code he submits. [34:36.800 --> 34:39.860] He submits 30 SPN entries a day. [34:40.580 --> 34:44.180] Not that much, but say 3,000 lines of code a day. [34:44.440 --> 34:46.840] He's doing a fair amount every day in and out. [34:46.920 --> 34:49.420] And all of a sudden his numbers drop down to 1,000 lines a day. [34:51.000 --> 34:51.920] That's an anomaly. [34:52.160 --> 34:52.880] What does that mean? [34:53.640 --> 34:57.400] Could it mean he's looking for new work or he's not being utilized properly? [34:57.980 --> 35:05.880] And then for intrusion detection systems, you'll run into a guy who's, say, logging in, logging in, logging in at a certain rate. [35:05.880 --> 35:10.560] And then all of a sudden he logs in 16 times, including five in the middle of the night, six nights in a row. [35:11.080 --> 35:13.180] Is this guy breaking in? [35:15.060 --> 35:16.520] Predictive analysis is great for that. [35:16.620 --> 35:23.560] You can predict what somebody's going to do and when they don't do what you're expecting, they will basically fall under the guise of, take a look at this guy. [35:23.560 --> 35:32.020] The problem with this is, and it's a big problem, is with a large amount of prediction, there's a small percent error. [35:32.320 --> 35:37.820] That small percent error when projected on a large set of data will end up being a very large error. [35:38.240 --> 35:48.080] For instance, there was HIV testing just recently where a guy basically, I think it was about 50% failure rate on an HIV test product in New York City. [35:48.080 --> 35:50.980] So about 50% of people who were told they had HIV didn't. [35:51.620 --> 35:52.680] That's a huge problem. [35:53.400 --> 35:56.820] The FDA's acceptable levels are 2%. [35:57.440 --> 36:05.280] So 2% of, say, 3 million people getting tested is going to be, what, 20,000 people? [36:06.300 --> 36:08.300] I think, if my math is right. [36:08.960 --> 36:12.420] That's a lot of people who have just been told that they're going to die, and they're not. [36:13.100 --> 36:14.580] So, in some cases... [36:14.580 --> 36:15.120] Well, they're all going to die. [36:15.320 --> 36:16.520] Yeah, well, they're going to die eventually. [36:17.340 --> 36:27.280] But anomalous detection can be very dangerous, especially with situations like the TSA, where they're going to say, this person is anomaly, he could be a dangerous flyer, we need to, you know, harass him. [36:28.440 --> 36:34.180] And that's why we generally have a problem deploying anomalous detection systems in security situations. [36:34.700 --> 36:49.620] If you decide to go on your once-in-a-lifetime vacation to Honolulu, and you have your card with you, and you try and use it, and all of a sudden they're screaming, that's an anomaly, you're going to be pissed because your credit card's not working, and you look like an idiot at the local restaurant. [36:50.300 --> 36:53.960] Anomalous detections can work as long as they're reviewed properly. [36:55.440 --> 36:59.100] And that's basically a big problem with anomalous detection systems. [36:59.320 --> 36:59.720] Anything else? [36:59.720 --> 37:06.100] Yeah, there has to be a human in there, but what you're really doing is you're just knocking down this massive flow of data, a smaller subset. [37:06.280 --> 37:06.940] Yep, but... [37:06.940 --> 37:07.840] You can't come through. [37:08.040 --> 37:18.740] The other great thing about anomalous detection systems is, say, we're doing something like we have ECHELON out there, which is analyzing a large volume of data, especially now with FISA, yadda yadda yadda. [37:18.880 --> 37:30.660] You could start, say, throwing off that volume of data by throwing random shit out there that says, I'm an anomaly, I'm an anomaly, I'm another anomaly, I'm still an anomaly, I'm going to continue to be an anomaly, all the live long day. [37:30.920 --> 37:34.640] You get enough people doing that, you can start convincing it of things that aren't true. [37:35.740 --> 37:51.680] A couple of people have written books about things like this, where in a situation where they're looking for anomalous activity in a city, the way people move on buses, if you start f*cking with the way the tracking system functions, it's not going to be able to predict any more how that works. [37:51.840 --> 37:58.180] The problem with an anomaly detection system is once one of the variables is compromised, the entire detection system is thrown out of whack. [37:59.740 --> 38:07.120] So, if I compromise the value of my location in a city and feed it over to someone else, suddenly they're an anomaly. [38:07.520 --> 38:11.260] If I keep doing that 16 other times, everyone's an anomaly. [38:11.800 --> 38:13.160] Then the whole system breaks. [38:16.540 --> 38:18.120] Is anomalous detection useful? [38:18.260 --> 38:18.460] Yeah. [38:19.440 --> 38:20.360] It's very useful. [38:22.060 --> 38:22.700] Anyone else? [38:24.740 --> 38:25.140] Okay. [38:27.640 --> 38:31.300] So, I don't know if I've even gotten this anywhere near where I want it to be. [38:31.460 --> 38:32.120] Here's the problem. [38:32.380 --> 38:36.360] I did a lot of security auditing for a while, and we usually followed scripts. [38:36.940 --> 38:38.100] And scripts are kind of like this. [38:38.240 --> 38:42.940] You come in, you see an application, you run it through a bunch of fuzzers, brute force detectors. [38:43.100 --> 38:43.980] They look for problems. [38:44.220 --> 38:45.960] If they find a problem, they report the problem. [38:45.960 --> 38:48.000] You have to manually detect whether or not the problem exists. [38:48.240 --> 38:49.580] They don't do logical analysis. [38:49.780 --> 38:50.740] They don't do code review. [38:50.940 --> 38:52.240] They don't do a lot of things. [38:52.480 --> 38:55.400] A lot of our banking systems work off this theory. [38:55.640 --> 38:57.300] You go in, you check it. [38:57.700 --> 38:58.660] Is it working? [38:59.000 --> 38:59.480] Okay. [38:59.920 --> 39:02.940] Could I break it by throwing a thousand different problems at it? [39:03.340 --> 39:03.700] No. [39:04.380 --> 39:06.700] We can do the same thing with large volumes of data. [39:06.940 --> 39:10.680] We can throw a predictive analysis. [39:10.680 --> 39:16.100] We basically go extrapolation, linear regression, correlation, predictive analysis. [39:16.480 --> 39:18.100] And that's just the intro. [39:18.320 --> 39:23.380] It's not important to do severe stuff unless you've got the money to do it and you think it's worthwhile. [39:24.040 --> 39:28.780] But for everyone who's out there got a large volume of data, you can do a small number of steps. [39:28.960 --> 39:30.640] There's a shit ton of tools for doing it. [39:30.700 --> 39:32.220] There's a ton of libraries for doing it. [39:32.380 --> 39:37.740] There's no reason that one or two system administrators couldn't throw this out on a network and do data analysis. [39:40.220 --> 39:48.040] There aren't very many tools out there that are geared towards the security industry for that sort of data analysis. [39:48.160 --> 39:49.780] A lot of books are starting to come out about it. [39:49.860 --> 39:54.560] A lot of people are starting to figure out that, yeah, we have a lot of this data floating around. [39:54.680 --> 39:55.600] There's a lot of reports. [39:55.780 --> 39:56.900] There's a lot of... [39:58.060 --> 40:04.920] For instance, back in the day we were outside of the Citigroup Center and Citigroup decided to throw out four or five boxes of paper. [40:05.360 --> 40:12.580] And those four or five boxes of paper consisted of financial records for most of City Bank's customers. [40:12.800 --> 40:13.040] Wait, wait. [40:13.180 --> 40:14.400] Was this on the first Friday of the month? [40:14.680 --> 40:15.000] Yes. [40:15.340 --> 40:15.660] You were there. [40:15.860 --> 40:16.000] Wait. [40:16.160 --> 40:18.980] So 2600 was going on in the bottom atrium? [40:19.180 --> 40:19.260] Yes. [40:19.780 --> 40:21.440] And Citigroup says they threw out their f*cking financial records? [40:21.520 --> 40:21.740] Yes. [40:22.420 --> 40:28.240] So Citigroup actually threw out their financial records during a 2600 meeting in front of us, unshredded. [40:30.420 --> 40:32.020] Several thousand sheets of paper. [40:33.540 --> 40:39.460] Aside from the immediate threat posed by this, you have their financial data, their name, their address, and yada, yada, yada. [40:39.980 --> 40:47.240] You can do a shit ton of data mining on all of those people and figure out things about Citigroup that they may not want you to know, like who they're marketing to. [40:47.980 --> 40:57.460] Another cool thing is, I saw this one on the Internet, was a guy basically sat down and said, why do people go to check cashing places in the ghetto? [40:58.300 --> 41:00.560] And as it turns out, the reason is very simple. [41:00.720 --> 41:02.480] There are no banks in bad neighborhoods. [41:03.000 --> 41:07.020] If you look at the locations of banks, they avoid bad neighborhoods for a reason. [41:07.660 --> 41:09.100] People who are poor don't have money. [41:09.300 --> 41:11.260] They don't spend money in banks. [41:11.520 --> 41:13.800] They go to the check cashing place because it's cheaper. [41:14.420 --> 41:17.740] Banks can be expensive for someone who doesn't have good financial sense. [41:18.840 --> 41:24.220] So, doing analysis of data is very, very, very, very important for us. [41:24.320 --> 41:26.280] And I'm trying to stress that with this talk. [41:27.400 --> 41:29.980] It's trying to give you the tools that you need to do it. [41:30.140 --> 41:36.220] And if anyone doesn't know how to go out and figure out how to do this, please ask questions. [41:41.360 --> 41:45.740] I'm going to throw this stuff up and I'll make it available at... [41:45.740 --> 41:46.920] Let me just think. [41:47.700 --> 41:55.840] I'll put it up at Sendai, S-E-N-D-A-I dash systems, S-Y-S-T-E-M-S dot com. [41:56.380 --> 42:05.720] And I will post a link to all the application, the presentation, the code, and a link to Wicca and a tutorial for getting started with using it. [42:06.340 --> 42:08.580] I'm sorry I couldn't get this work on my laptop. [42:09.040 --> 42:15.580] But hopefully you guys are going to go out and actually take a look at some of these tools. [42:15.580 --> 42:25.280] Grab them, throw them in on your environment, and start dealing with large volumes of data and saying, hey, I'm seeing shit that is bad. [42:26.720 --> 42:27.240] Okay? [42:28.400 --> 42:28.920] Anyone? [42:29.360 --> 42:30.160] More questions? [42:32.560 --> 42:33.080] Shoot. [42:33.340 --> 42:37.080] Yeah, what are some more examples of things that might be fun to move forward? [42:37.720 --> 42:39.160] Well, let's think. [42:39.700 --> 42:41.020] I've done a couple of ones. [42:42.420 --> 42:46.660] I really got to tell you that the code Pastebin one was probably a treasure trove. [42:47.040 --> 42:50.220] You'd go in there, you'd find configuration files, you'd find zero-day exploits. [42:50.860 --> 42:54.740] You could actually do a search for shellcode inside of Pastebin data. [42:55.080 --> 42:58.700] And you'd pull up pretty much every exploit ever dumped in there. [42:58.700 --> 43:02.700] And some of them would be unreleased, because guys would be testing them out with their friends. [43:04.700 --> 43:10.540] So, depending on where you expect data to be that's considered of high value, you probably want to choose your target first. [43:11.020 --> 43:17.800] And then think about all the different effects, like a butterfly theory thing, that they have on the world. [43:19.240 --> 43:20.980] You see this a lot with reverse engineering. [43:21.820 --> 43:26.300] Physicists, for instance, are trying to figure out how the universe began. [43:27.400 --> 43:33.360] And by the time we figured out when it began, it was because we figured out background radiation and how it emanated. [43:33.480 --> 43:41.420] And we were able to look back at it and say, oh, well, the background radiation of the universe tells us that it began, I think it was, what, 3 trillion years ago? [43:41.500 --> 43:42.000] 13 trillion? [43:42.160 --> 43:43.100] No, 4,000 years ago. [43:43.260 --> 43:43.960] Yeah, 4,000. [43:46.840 --> 43:48.260] Raptor Jesus loves you. [43:50.720 --> 43:57.640] So, if you analyze the things coming out and the things going into a black box design, you can figure out what's going on inside the black box design. [43:57.880 --> 44:06.840] If you are looking at a bank, for instance, or you're looking at a hospital, and you want to figure out, well, how many drugs are going into the place? [44:06.920 --> 44:09.800] Where are the drugs stored so I can break in there with my friends and steal them? [44:11.440 --> 44:12.360] You can do that. [44:12.580 --> 44:15.460] You can say, okay, the trash is being thrown out every other day. [44:15.460 --> 44:17.820] We know that they're moving this many needles. [44:18.000 --> 44:20.800] We know that the needles go up and down and fluctuate during the months. [44:20.860 --> 44:25.740] We know usually when they're delivered, they'll be more free with throwing them out, especially if it's like a small clinic. [44:26.920 --> 44:33.020] And you can pretty much, just by sorting through their trash, go, oh, hey, now's a good time to rob them. [44:33.500 --> 44:34.940] Same is true of anyone. [44:35.640 --> 44:40.780] If you hang out outside a bank, you can say, okay, when do they come for pickups with their little armored truck? [44:41.140 --> 44:44.020] Oh, okay, they come every fourth Wednesday of the month. [44:44.020 --> 44:47.100] Okay, then I'll hit them on the day before, Tuesday. [44:49.220 --> 44:51.140] So, this is all unethical stuff. [44:51.280 --> 44:58.280] It's also good for ethical things, like doing this to yourself and saying, okay, well, obviously we need to randomize things a bit. [44:58.400 --> 45:01.140] We need to not do this in a predictable fashion. [45:01.580 --> 45:08.180] We need to tell the bank company that they need to deliver their money at random intervals so that we don't get hit real hard. [45:10.980 --> 45:12.160] Great stories. [45:13.200 --> 45:13.800] Okay. [45:14.840 --> 45:15.620] Here's one. [45:16.000 --> 45:20.260] How we discovered that Cuban Missile Crisis, that the Cubans had missiles. [45:20.640 --> 45:27.700] It wasn't really data analysis, but they're flying overhead and they see what look like soccer fields. [45:29.920 --> 45:32.240] And that doesn't exist in Cuba. [45:32.420 --> 45:34.880] They're a baseball playing country, so why are there soccer fields? [45:35.940 --> 45:41.260] So, the way we figured out that there were nuclear weapons in Cuba was because we saw the Russians building soccer fields. [45:41.420 --> 45:42.340] Not because we saw missiles. [45:42.460 --> 45:44.740] Not because we knew anything about the movement over there. [45:44.860 --> 45:46.440] We just knew that there were Russians playing soccer. [45:48.680 --> 45:53.760] So, doing analysis of extraneous data can be incredibly useful. [45:55.740 --> 45:57.100] Does anyone else have any questions? [45:57.840 --> 45:58.320] Huh? [45:59.080 --> 45:59.380] Shoot. [45:59.740 --> 46:00.340] How big [46:03.760 --> 46:04.100] is it? [46:06.040 --> 46:10.500] They didn't actually list numbers and I was pushing it pretty hard, but I couldn't get it up to anything where it would crash. [46:10.660 --> 46:15.480] My guess is since it's being stored in main memory, the crash occurs when your system runs out of memory. [46:16.560 --> 46:22.240] So, if you have a very small machine, it will crash a lot quicker than if you have a very large machine. [46:23.300 --> 46:28.020] So, you analyze your data set and it's, you know, $3,000 loan. [46:31.820 --> 46:33.120] That could be one sale. [46:43.140 --> 46:46.940] Agent 1000 is actually marginally big. [46:47.100 --> 46:48.480] I'm pretty sure it would handle it fine though. [46:49.600 --> 46:51.080] Sales records are very small numbers. [46:51.260 --> 46:52.680] It would take a while to churn through that. [46:53.260 --> 46:57.040] If you start doing transforms on it, it will do sets of transforms. [46:57.040 --> 46:58.800] And Wicked can be kind of slow. [46:59.320 --> 47:03.160] It's not designed to be the optimized approach. [47:03.340 --> 47:07.700] It's designed to be the one that gives you a very easy, smooth interface with good results. [47:09.120 --> 47:13.460] So, if you're planning on doing it constantly, you may want to look into a more specialized solution. [47:13.940 --> 47:17.360] But if you're just doing it for the first time to see what you need, it will be fine. [47:21.820 --> 47:22.740] Any other questions? [47:23.100 --> 47:23.500] Concerns? [47:23.700 --> 47:26.760] Anyone want to tell me I'm horrible and they hate me? [47:27.000 --> 47:27.700] Painfully white? [47:28.000 --> 47:29.300] Did you at least get a tan while you were here? [47:30.920 --> 47:32.740] I'm getting one under these lights actually. [47:32.860 --> 47:33.840] I can't see any of you. [47:34.160 --> 47:35.960] So, I'm actually extremely paranoid. [47:36.160 --> 47:38.440] I think that there's people back there making faces at me. [47:39.080 --> 47:40.900] No, I'm right here making faces at you. [47:43.160 --> 47:45.100] Yeah, so, I nosedive on this one. [47:45.180 --> 47:45.720] I'm sorry guys. [47:48.080 --> 47:49.500] Other, any other questions? [47:50.380 --> 47:51.320] I'm wrapping it up. [47:52.140 --> 47:53.340] He's telling me to wrap it up. [47:53.480 --> 47:54.140] I'm out of time. [47:54.860 --> 47:55.280] Good night.