[00:44.050 --> 00:45.150] Hello, everyone. [00:45.410 --> 00:46.550] Welcome to our next talk. [00:46.870 --> 00:49.050] How's the conference going so far for the last day? [00:50.510 --> 00:54.270] Did any of you make it to the Hotel Pennsylvania presentation last night? [00:55.810 --> 00:57.490] Yeah, that was quite interesting. [00:58.250 --> 01:02.450] All right, well, a couple of brief announcements, a couple of brief business before we start the next talk. [01:02.850 --> 01:04.530] Closing ceremony is at 6 p.m. tonight in 416. [01:06.150 --> 01:07.990] That's upstairs at the end of the building. [01:08.230 --> 01:11.350] It will be simulcast in this room instead of a little theater. [01:11.350 --> 01:18.530] So if you don't make it into the larger theater, into 416, come down here and you'll be able to watch the closing ceremonies from here. [01:19.390 --> 01:21.910] We do need volunteers to help us do the loadout. [01:21.910 --> 01:24.850] If you're interested, Monday morning at 9 a.m. here at DiAngelo at the loading dock, whatever help we can get. [01:28.770 --> 01:33.830] Stop by the info desk or reach out to one of the other crew staff members and let them know that you would like to help. [01:35.610 --> 01:37.630] There are some great talks in the coffee house today. [01:37.950 --> 01:46.310] So if you have the time to stop in and listen in on some of the speakers who didn't get into the normal tracks and have some interesting things to say, go and support them. [01:46.430 --> 01:47.430] It would be great for them. [01:48.070 --> 01:48.930] Keep hydrated. [01:48.930 --> 01:52.270] For those of you who are staying in the dorms, check out us at 8 p.m. tonight if you had the dorm until today and 2 p.m. tomorrow if you were staying the night tonight. [01:58.230 --> 02:10.630] And lastly, if you have any feedback for us, any comments, any criticisms, any applause, sensitive feedback at hope.net, F-E-E-D-B-A-C-K at hope.net. [02:11.150 --> 02:18.310] And with that, let's jump into our next talk, Hackers Can Help, Open Technical Problems in Investigative Journalism by Brandon Roberts. [02:18.850 --> 02:20.930] Brandon will be joining us through a virtual Zoom. [02:21.150 --> 02:31.150] If you can send comments to the Matrix chat or hold them to the end, we have a microphone over here and we will, when we get to the Q&A session, you can step up to the mic and we will ask your questions. [02:31.450 --> 02:32.390] Over to you, Brandon. [02:33.550 --> 02:34.190] Hello, everyone. [02:34.430 --> 02:35.110] Hello, HOPE. [02:35.370 --> 02:36.750] I wish I could be there. [02:36.890 --> 02:37.690] I love HOPE. [02:38.670 --> 02:39.450] Hotel Penn. [02:40.090 --> 02:42.070] Rest in peace, you dirty mofo. [02:42.070 --> 02:42.990] Okay. [02:43.530 --> 02:47.690] Yes, this is called Hackers Can Help, Open Problems in Investigative Journalism. [02:48.130 --> 02:49.530] I'm a data journalist. [02:49.950 --> 02:50.990] My name is Brandon Roberts. [02:52.550 --> 02:56.830] I specialize in data-heavy investigative projects. [02:57.330 --> 02:59.490] I collaborate with a lot of newsrooms. [02:59.510 --> 03:07.030] I'm independent, so I don't have an employer, but I do work with a lot of large and small news organizations. [03:07.830 --> 03:18.710] Recent work I've done with ProPublica, Newsday, Associated Press, and then also some local organizations like Oregon Public Broadcasting and Seattle Times. [03:19.110 --> 03:20.270] I'm based in Western Washington. [03:21.790 --> 03:24.330] Yeah, and basically, I'm self-taught. [03:24.650 --> 03:27.350] I didn't go to school for any of this stuff. [03:27.910 --> 03:36.450] For many years, I had, you know, I had more computer skills than I had journalism skills, and I really wanted to apply them to journalism. [03:36.450 --> 03:41.490] But it took me a really long time to figure out what kinds of things are actually useful to journalists. [03:41.770 --> 03:48.730] And the way I did that was trial and error, you know, just smashing my head against my laptop for, you know, almost a decade. [03:49.610 --> 03:54.310] This talk is going to not prevent you from doing that, hopefully. [03:54.670 --> 04:00.610] I'm going to go through problems that journalists face, particularly data journalists. [04:01.710 --> 04:07.570] And, you know, point out things that, you know, if you have technical skills, you'd be very useful. [04:07.970 --> 04:12.170] So, and good ways to get involved in your local journalism community as well. [04:13.990 --> 04:14.510] Okay. [04:14.510 --> 04:15.030] Okay. [04:15.030 --> 04:28.530] So, when I was, when I was started, and other people that I talked to, a lot of, like, CS type things, other conferences, non-journalism conferences, and people are doing something related to news. [04:28.930 --> 04:30.930] Most of the time, it's one of these things. [04:30.930 --> 04:32.910] They're analyzing news. [04:33.070 --> 04:34.790] They're building bias detectors. [04:35.630 --> 04:40.190] They're working with, like, you know, non-governmental data. [04:40.490 --> 04:45.710] So, like, you know, like the rumor mill, Twitter data, stuff like that, user-generated content. [04:46.650 --> 04:54.310] And they're doing kind of, like, simplistic, simplistic analyses that just are not useful. [04:55.550 --> 04:59.090] Yeah, like, journalists, we work with primary sources. [05:00.650 --> 05:03.370] We, you know, we do records requests. [05:03.370 --> 05:04.290] We get data. [05:04.430 --> 05:07.550] And we know what kind of data we are going to get. [05:07.690 --> 05:11.010] Because we typically know what we're going to, like, look for. [05:11.290 --> 05:17.550] If we're doing police investigations, we're going to ask for police misconduct data or use of force. [05:17.730 --> 05:19.010] We know what we're looking for. [05:19.050 --> 05:20.390] And we're asking for that. [05:20.610 --> 05:22.510] So, techniques that are just very general. [05:22.750 --> 05:26.790] And we'll tell you, like, this data set is generally about these things. [05:27.010 --> 05:28.290] Usually not very useful. [05:29.690 --> 05:39.010] Most journalists are looking for, you know, that needle in the haystack or just, you know, looking for a really specific kind of thing. [05:39.010 --> 05:44.070] So, a lot of these technologies are not useful. [05:44.370 --> 05:46.030] I have this example here. [05:46.530 --> 05:47.890] This little bias detector. [05:47.970 --> 05:48.890] I see these all the time. [05:49.070 --> 05:50.010] I just think they're really funny. [05:53.010 --> 05:57.090] Because it's like, this is, like, supposed to be very neutral. [05:57.090 --> 06:00.470] And then, you know, it's annotated better than not reading news at all. [06:00.590 --> 06:02.570] Which is, like, a very, very subjective thing. [06:05.630 --> 06:07.830] So, this is the kind of stuff I did when I got started. [06:09.090 --> 06:11.250] This is not... I don't do any of these things anymore. [06:13.370 --> 06:13.910] Okay. [06:14.430 --> 06:16.170] So, why this talk? [06:16.870 --> 06:19.830] I did a GitHub search for news bias detector. [06:19.990 --> 06:21.150] I got 23 repos. [06:21.350 --> 06:25.250] You know, a bunch of random code snippets and search hits basically related to this. [06:27.930 --> 06:30.090] The thing about those bias detectors, too. [06:30.230 --> 06:34.510] Anything related to climate instantly just goes straight over to the left-wing extremism. [06:34.870 --> 06:37.890] If you talk about budget, they're going to put that as, you know, right-wing. [06:38.130 --> 06:39.530] So, it's more like... I don't know. [06:39.750 --> 06:40.950] They're just... this is not very useful. [06:41.650 --> 06:42.450] Russian troll. [06:42.710 --> 06:44.130] This was a big thing a few years ago. [06:44.270 --> 06:46.770] People wanted to do a Russian troll detector. [06:48.150 --> 06:49.950] 538, the newsroom. [06:50.150 --> 06:51.350] They have their own. [06:52.610 --> 06:54.430] This is just absurd. [06:55.490 --> 07:02.150] What's the difference between a Russian troll and just someone who's, like, really passionate about, you know, whatever political thing. [07:02.430 --> 07:07.410] It's just... it's very hard to figure that out, particularly in an automated way. [07:07.890 --> 07:11.870] And then, yeah, this, like... it's just... it's not useful. [07:12.890 --> 07:14.990] Here's something that is useful, though. [07:15.450 --> 07:16.450] Records requests. [07:16.930 --> 07:17.510] FOIA. [07:17.830 --> 07:21.930] Tools that, you know, help journalists do records requests and manage FOIA. [07:22.910 --> 07:25.070] Barely any repositories around this. [07:25.910 --> 07:26.390] And... [07:27.470 --> 07:31.390] Every newsroom, every journalist, technical or not, has to do records requests. [07:31.790 --> 07:33.790] That's, like, you know, it's a huge part of the game. [07:34.430 --> 07:37.430] And the fact that there's just so little on GitHub about that. [07:37.430 --> 07:42.550] It just kind of shows the lack of, like, literacy amongst technical people. [07:42.990 --> 07:44.590] For journalists... [07:47.130 --> 07:50.510] Civically, a records request within, like, the journalism community. [07:50.510 --> 07:56.170] It's, like, almost a running joke that, like, every journalist has their own FOIA tool. [07:56.170 --> 08:00.550] The New York Times reportedly had, like, four internal FOIA tools. [08:00.970 --> 08:09.870] So, not only are we, like, reinventing the wheel as an industry, individual newsrooms are also reinventing the wheel internally. [08:12.170 --> 08:12.810] Okay. [08:13.130 --> 08:17.470] And these are the things that data journalists actually do. [08:17.470 --> 08:19.370] We scrape data. [08:19.610 --> 08:20.890] We request data. [08:21.190 --> 08:25.670] We gather data from government agencies and corporations when possible. [08:25.950 --> 08:29.330] It's not always possible, but, you know, that's where scraping comes in. [08:30.790 --> 08:37.130] And then once we have our data, we need to extract, you know, meaningful machine-readable information from it. [08:37.130 --> 08:40.190] So, that's either extracting data from a PDF. [08:40.510 --> 08:43.770] It's doing something like OCR or handwriting analysis. [08:44.970 --> 08:49.230] Basically turning human-readable data into machine-readable data. [08:49.990 --> 08:53.230] And then on the right here, we have building data sets. [08:53.370 --> 09:00.930] It's kind of like the final thing, the final step that you do when you're gathering data and all this. [09:01.490 --> 09:05.110] These are what I call the tools of the trade of data journalism. [09:05.110 --> 09:06.870] These are the things that we do. [09:07.050 --> 09:08.650] These are the things that we need help with. [09:10.430 --> 09:15.510] So, the rest of the talk, I'm going to go into these in detail and talk about the problems. [09:16.450 --> 09:17.830] Talk about our needs. [09:18.210 --> 09:22.990] And, you know, if you have ideas about this kind of stuff, we need you. [09:25.370 --> 09:25.770] Okay. [09:25.910 --> 09:26.030] Yeah. [09:26.130 --> 09:27.530] And like I said, we need your help. [09:27.930 --> 09:32.370] So, data journalists, people who are like the top of the game. [09:32.370 --> 09:36.230] You know, these are the most technical people in the industry. [09:36.870 --> 09:40.970] Not just like a random reporter from a small newsroom who might not be technical. [09:41.210 --> 09:43.970] I'm talking about data journalists who are like they're supposed to be pros. [09:45.310 --> 09:51.870] Amongst us, you know, we all have our own like cobbled together scraper scripts. [09:52.310 --> 09:55.490] They're trash, trash code, you know, just trying to get them done quickly. [09:57.190 --> 10:02.130] We all have our own like we manage our own Selenium stacks because we need browsers, automated browsers. [10:02.470 --> 10:06.590] So, like you can't just you need a headless browser. [10:06.750 --> 10:07.730] So, you need to have a whole stack. [10:07.890 --> 10:08.630] So, we all have that. [10:08.790 --> 10:09.650] We all do that ourselves. [10:11.110 --> 10:15.170] We all have like countless one-off throwaway extractor scripts. [10:15.330 --> 10:19.010] So, I need to get specific information out of this one PDF. [10:19.270 --> 10:20.250] I'm going to write a script for that. [10:20.450 --> 10:23.970] You know, just thousands and thousands of code snippets just thrown away. [10:24.150 --> 10:24.630] We all did it. [10:25.270 --> 10:32.130] And then we all participate in like large amounts of manual tedious data entry or transcription and stuff like that. [10:32.130 --> 10:33.890] And we're all reinventing the wheel. [10:34.430 --> 10:38.690] Very little code sharing because the code we're writing is so specific to the problem. [10:40.010 --> 10:40.450] Yeah. [10:40.690 --> 10:42.970] So, there's just so much room for improvement here. [10:44.090 --> 10:48.530] Reinventing the wheel is going to be a common theme that you're going to see here. [10:49.830 --> 10:56.750] The reason why I think our code is like this, like why it's sloppy, why we don't share is because we're on deadlines. [10:56.990 --> 10:58.230] We're trying to get stuff done fast. [10:58.230 --> 11:05.610] And that's just not, you know, it's not a great environment for like writing like a great library that everyone's going to use and share. [11:08.130 --> 11:08.650] Okay. [11:10.010 --> 11:10.690] All right. [11:10.990 --> 11:16.030] So, data journalism as a scientific-ish process. [11:17.150 --> 11:20.230] I just want to kind of sketch out how we do our work. [11:22.070 --> 11:26.450] You know, like breaking news and other forms of news, it's a lot less structured. [11:26.450 --> 11:28.530] You know, it's kind of like a reaction to an event. [11:28.690 --> 11:31.210] You know, you dig in, you get context-raised stories, whatever. [11:32.090 --> 11:34.070] Data journalism works a little differently. [11:34.330 --> 11:35.190] It's a little slower. [11:36.290 --> 11:38.550] You know, so I'll break it up into these three phases. [11:39.690 --> 11:40.070] Hypothesis. [11:40.290 --> 11:42.170] You know, what is our idea? [11:42.290 --> 11:43.090] What are we investigating? [11:43.510 --> 11:46.410] Can start with a tip, a rumor, just a hunch. [11:46.410 --> 11:52.370] You know, I think, oh, I think that certain people might be evading campaign finance rules in a specific way. [11:52.370 --> 11:54.130] That's an idea that I can test. [11:55.170 --> 12:02.430] Once I have this idea, you know, I'll gather data to, you know, see if I was right or wrong or see what's actually going on. [12:02.430 --> 12:11.350] If I think that the police in my area are like, you know, beating people up, I would gather use of force data. [12:11.550 --> 12:16.990] If I think police are arresting people unfairly, maybe I'll scrape the municipal court database. [12:18.130 --> 12:19.950] So that's where the web scraping comes in. [12:20.550 --> 12:22.750] Or, you know, request information from them. [12:23.650 --> 12:28.930] Then I would have to extract it, get into something that I can use an algorithm to analyze. [12:30.510 --> 12:34.470] And then once I have the data, you know, I can actually do an analysis on it. [12:34.570 --> 12:35.850] Then I would check my hypothesis. [12:36.510 --> 12:37.950] Are people breaking the law? [12:38.130 --> 12:40.730] Are the police, you know, beating up people? [12:40.910 --> 12:42.070] Like, you know, then I can test. [12:42.070 --> 12:46.190] So we gather data, we use tools to process that data. [12:46.490 --> 12:51.430] And then we check, check our, you know, check our suspicions, see what's up. [12:51.690 --> 12:53.570] And then, you know, that repeats itself. [12:55.230 --> 13:00.570] Just because you're using data, it doesn't mean that your method is science. [13:00.810 --> 13:02.710] It doesn't mean that your method is flawless. [13:03.690 --> 13:05.110] All data has limits. [13:05.450 --> 13:08.290] It's very important to know the limits of the data that you're working with. [13:09.990 --> 13:15.990] And you yourself can be a very strong and powerful source of bias. [13:16.450 --> 13:22.430] It's very easy to, you know, when you're going through files, extracting stuff. [13:22.630 --> 13:26.810] You know, you can easily, you know, put your finger on the scale a little bit, you know. [13:26.950 --> 13:33.170] So it's really important to be thorough and to be honest and to do things in like a repeatable way. [13:33.170 --> 13:44.230] And that's why having tools that are shareable is important because then other people can do the same, the same experiment, you know, the same investigation and get the same results. [13:48.890 --> 13:50.230] Just some definitions. [13:50.930 --> 13:54.150] A CSV, a comma separated values file. [13:55.450 --> 13:58.050] We love these as data journalists. [13:58.670 --> 14:05.590] Basically, the goal of most investigations is take some data, somehow get it into a CSV. [14:05.890 --> 14:09.830] Because once you have it in a CSV, then you can do all kinds of stuff with it. [14:09.830 --> 14:12.050] Every language can read it easily. [14:12.690 --> 14:18.410] There's all kinds of awesome tools that can help you do, you know, all kinds of analysis on it, just as a CSV. [14:19.490 --> 14:22.230] We have this great tool called CSV Kit. [14:23.290 --> 14:28.450] It allows you to do joins on CSVs, do all kinds of stacking and filtering and all kinds of stuff. [14:28.570 --> 14:28.950] It's amazing. [14:28.950 --> 14:38.490] So a lot of the stuff I'm going to show you is ways to turn human readable data into an CSV file. [14:45.300 --> 14:47.700] Oh, skipped a slide. [14:47.820 --> 14:47.920] Okay. [14:48.340 --> 14:49.020] First problem. [14:49.500 --> 14:50.120] Website. [14:50.500 --> 14:54.200] That is bots and browsing the web. [14:56.920 --> 15:00.480] Usually, when you're scraping, it's downloading HTML, right? [15:00.480 --> 15:02.080] But not always. [15:04.000 --> 15:07.180] Sometimes it's, you know, interacting with the... [15:07.180 --> 15:08.760] You get downloading some PDFs. [15:08.920 --> 15:10.140] It doesn't really matter. [15:10.440 --> 15:15.100] The point is that a script is interacting with a web browser, interacting with the website. [15:17.800 --> 15:20.780] Webscraping used to be, like, a part of my work. [15:23.120 --> 15:27.920] Recently, like, governments, especially in the local area, have, like, becoming... [15:27.920 --> 15:29.800] They're becoming more technical savvy. [15:29.800 --> 15:32.380] So it's easier to get, like, an Excel spreadsheet from them. [15:32.580 --> 15:34.880] But five, ten years ago, that was not the case. [15:34.900 --> 15:35.900] I scraped everything. [15:36.160 --> 15:42.000] So one of my largest scrapes was I scraped the entire Seattle Municipal Court database. [15:42.000 --> 15:46.540] So I had, like, nearly every misdemeanor going back to the 1970s. [15:48.360 --> 15:52.080] It's really important for us to be able to use a real browser. [15:52.380 --> 15:55.760] Because most of the sites that we're trying to scrape use a lot of JavaScript. [15:56.920 --> 16:02.360] Like, really old sites using, like, old ASP and .NET stuff. [16:02.660 --> 16:04.580] Just really, you have to have a browser. [16:05.520 --> 16:09.020] So that means that we're doing a lot of Selenium. [16:10.220 --> 16:10.740] Yeah. [16:14.060 --> 16:14.580] Okay. [16:14.940 --> 16:16.320] So reason why a web screen. [16:17.700 --> 16:20.040] There's a lot of ways to do one thing. [16:20.700 --> 16:26.380] These are the search buttons from all of the Canadian search pages. [16:26.900 --> 16:33.120] So for each province, we have, like, a way to search lobbyists to see, you know, lobbying for what? [16:33.120 --> 16:36.240] What companies are paying for you. [16:36.520 --> 16:38.060] All the information is online. [16:38.680 --> 16:39.980] Each process is a database. [16:40.720 --> 16:42.040] And these are the search buttons from each. [16:44.300 --> 16:44.860] Code. [16:44.920 --> 16:46.120] You can see the code from each, too. [16:46.420 --> 16:48.960] The code is not always the same. [16:49.320 --> 16:54.560] This search button, it's, you know, it's an italic thing. [16:55.900 --> 16:59.960] And it's in a button, which is good, but anchor tag with an image. [17:00.540 --> 17:06.080] So your scraper needs to be able to figure out how to submit, you know, how to do a search. [17:06.600 --> 17:08.880] And in a generic way, it's hard. [17:09.300 --> 17:17.240] So right now, there's not really any tools out there that can be like, hey, find a scraper at a site and hit submit. [17:18.660 --> 17:23.120] Hit the search button, you know, even if we just want a search button, that doesn't exist. [17:24.620 --> 17:25.740] And that would be useful. [17:28.770 --> 17:29.270] Okay. [17:29.810 --> 17:33.910] So yeah, scraping is hard because you have to handle unexpected things. [17:34.150 --> 17:40.550] This site here, under certain conditions, will give you a search result. [17:40.790 --> 17:46.010] But under other conditions, it'll give you an old school JavaScript alert box. [17:46.090 --> 17:48.570] And you have to click OK, you got to make that go away. [17:48.690 --> 17:53.010] So if your scraper didn't know that that was going to happen, your scraper would break. [17:53.010 --> 17:58.810] So having a robust web scraper tool is, it's really hard to build. [18:00.170 --> 18:06.030] Just because you have to, like, know all the different things that are going to happen or have thought of, oh, alert. [18:06.150 --> 18:07.070] What if an alert pops up? [18:07.150 --> 18:08.470] OK, let's just put some code in there. [18:08.930 --> 18:09.590] Hit a submit. [18:09.810 --> 18:11.170] And then if there's an alert, click that. [18:11.770 --> 18:13.050] All these things need to... [18:13.830 --> 18:19.890] So if there was some kind of, like, you know, scraper library that did this, it would be super useful. [18:24.000 --> 18:24.640] OK. [18:25.120 --> 18:27.280] And here's the Seattle Municipal Court site. [18:28.800 --> 18:33.380] The load times on this are an issue because they vary. [18:33.920 --> 18:36.180] You got to wait for that little scroll thing to go away. [18:36.180 --> 18:41.840] And, like, knowing how to code for that is hard. [18:43.480 --> 18:50.040] If your scraper, like, hit submit and then it looks for the page immediately, there's a chance the document won't be loaded yet. [18:50.220 --> 18:52.800] Maybe that little scroller didn't even load yet. [18:52.960 --> 18:58.900] So it's like being able to figure out page transitions, that's super hard because every page is different. [18:59.820 --> 19:03.600] This is using, like, some old-school JavaScript stuff. [19:05.220 --> 19:10.620] Sometimes this site just really lags and takes, like, five minutes to load a page. [19:10.780 --> 19:12.200] Sometimes it's super fast. [19:13.240 --> 19:15.120] You need to be able to handle all those scenarios. [19:17.280 --> 19:17.800] Yeah. [19:18.940 --> 19:26.320] You know, similarly, when I was looking through the province lobbyist searches, one of them's down. [19:26.320 --> 19:29.980] So the scraper needs to know what to do when the site's down. [19:30.660 --> 19:32.620] Or, you know, your script can just explode. [19:32.620 --> 19:36.820] But if you're only scraping something once a day and this happens, like, maybe you want to wait. [19:37.000 --> 19:37.540] I don't know. [19:38.860 --> 19:44.740] It would be useful if there was some kind of, like, you know, library that helped with these kind of things. [19:48.320 --> 19:48.760] Okay. [19:48.980 --> 19:51.040] So the existing technologies. [19:51.740 --> 19:55.200] Web scraping, this is, you know, there's something called Scrapey. [19:57.300 --> 19:58.040] Scrapey's cool. [19:58.880 --> 20:00.240] But you've got to write a lot of code. [20:00.420 --> 20:01.380] There's a lot of boilerplate. [20:01.800 --> 20:04.200] The more code you write, the more code you have to maintain. [20:05.880 --> 20:08.100] And scrapers are super brittle often. [20:09.160 --> 20:10.280] They break a lot. [20:11.340 --> 20:14.740] And, you know, the more code there is, the more annoying it's going to be to maintain. [20:15.480 --> 20:17.400] So writing less code is always best. [20:18.820 --> 20:20.320] CSS selectors, XPath. [20:20.540 --> 20:23.960] That's, like, how most, like, web scraping tutorials would tell you to do things. [20:24.660 --> 20:26.360] Like I said before, it's super brittle. [20:27.360 --> 20:32.440] Susceptible to, like, small changes in the page layout or, like, the design or whatever. [20:32.760 --> 20:34.640] The function of the page might not change. [20:34.800 --> 20:36.640] But, you know, maybe they change the IDs. [20:36.840 --> 20:38.760] Maybe the IDs are automatically generated. [20:38.860 --> 20:39.760] The classes and IDs. [20:39.980 --> 20:41.080] That's how Facebook works. [20:41.080 --> 20:47.020] If your scraper relies on the classes being the same, it's only going to work for a small period of time. [20:48.080 --> 20:51.020] So we need things that can be robust. [20:52.200 --> 20:53.560] Selenium, talked about this. [20:53.740 --> 20:57.340] But, you know, working with Selenium is super tedious, verbose. [20:57.700 --> 21:00.860] You've got to catch every possible issue. [21:01.200 --> 21:02.380] It's just super hard. [21:03.600 --> 21:09.820] And the reason I think it's hard is because we're really lacking domain-specific languages around web scraping. [21:10.840 --> 21:14.960] We have one called hexed for extracting data from HTML. [21:15.820 --> 21:16.520] And this is kind of cool. [21:16.700 --> 21:17.640] This is a hexed template. [21:18.120 --> 21:27.620] So this is saying, for all anchor tags, extract the href as a link and the text of the tag as title. [21:27.820 --> 21:31.980] And then if we have some input HTML, we would get this JSON back. [21:32.440 --> 21:34.140] So that's, like, a really cool idea. [21:34.400 --> 21:37.180] A domain-specific language for extracting data from HTML. [21:37.400 --> 21:38.480] I use this a lot. [21:38.480 --> 21:41.520] It would be awesome if there was something similar for web scraping. [21:44.060 --> 21:44.780] And there's... [21:44.780 --> 21:47.620] I mean, there's so many things that we could do with web scraping. [21:48.600 --> 21:50.660] It's just a huge, vast area. [21:52.680 --> 21:53.180] All right. [21:53.280 --> 21:55.640] The next problem is PDF data extraction. [21:55.940 --> 21:56.720] PDFs are everywhere. [21:56.940 --> 22:00.620] If you request data from a clerk, you're going to get a PDF, like, nine times out of ten. [22:00.620 --> 22:07.080] Again, what a PDF is, is like an unstructured view into some structured data. [22:07.980 --> 22:12.540] They're often not OCRed or they are poorly OCRed. [22:13.960 --> 22:20.880] Clerks in a lot of places, especially like police departments, they love doing a print to PDF for an Excel spreadsheet. [22:21.200 --> 22:23.260] They say that they do this, you can't change it. [22:23.460 --> 22:25.780] But, like, if I was gonna change it, I could do that anyways. [22:26.520 --> 22:27.240] But, yeah. [22:27.280 --> 22:29.960] So if you're a journalist, you're gonna be working with a lot of PDFs. [22:30.720 --> 22:32.820] We do have some tools to extract this stuff. [22:33.580 --> 22:34.420] They're not perfect. [22:36.540 --> 22:37.220] Yeah. [22:37.480 --> 22:39.040] So here's, like, just an example PDF. [22:39.580 --> 22:44.280] This comes from this project called PDF Flummer, written by a journalist I know. [22:47.060 --> 22:49.420] PDF is the portable document file. [22:50.440 --> 22:59.400] What a PDF is, is it is really setting around displaying information on all the same way. [22:59.400 --> 23:04.840] So you can guarantee that if you have this PDF, it's gonna look the same on all computers. [23:05.060 --> 23:09.920] And that was, like, kind of a revolutionary idea, you know, a long time ago. [23:10.320 --> 23:12.860] Because, you know, you're, you're sending text files around. [23:13.180 --> 23:17.140] It depends on the font that the person has, depends on the font size that the person has. [23:17.580 --> 23:24.940] This was, like, the best technology for ensuring that a document looks the same across, you know, a bunch of environments. [23:27.180 --> 23:31.000] Yeah, so this is, like, a really good spreadsheet. [23:31.260 --> 23:32.200] Because this one is OCR. [23:32.380 --> 23:32.940] You can't see that. [23:33.040 --> 23:34.660] But it has, like, these little lines. [23:34.820 --> 23:35.560] It's really cool. [23:35.760 --> 23:36.620] Nothing's cut off. [23:37.240 --> 23:39.580] This is, like, an ideal spreadsheet. [23:41.180 --> 23:42.720] Here's what a spreadsheet looks like. [23:42.800 --> 23:44.200] That same spreadsheet looks like. [23:44.800 --> 23:48.140] I've taken the raw data from the PDF and turned it into a spreadsheet. [23:48.880 --> 23:51.440] And you can see this one. [23:51.520 --> 23:52.920] This says notice date here. [23:54.640 --> 23:57.040] We have N-O-T-I-C, notice date. [23:57.320 --> 24:08.080] So each thing on a PDF, each letter, each line, it has a location, where it is, what it looks like, all that. [24:08.480 --> 24:13.300] So it's, like, there's no such thing as, like, a word. [24:13.460 --> 24:14.160] There's no sentence. [24:14.360 --> 24:14.740] There's nothing. [24:14.740 --> 24:17.820] There's just individual characters scattered all over on a page. [24:17.960 --> 24:18.900] That's what a PDF is. [24:19.900 --> 24:22.340] And that's why it's so hard to extract data from it. [24:24.620 --> 24:24.980] Yeah. [24:25.120 --> 24:26.220] So here's another example. [24:26.340 --> 24:28.460] This is, like, a more realistic example of a PDF. [24:29.520 --> 24:30.800] This is from Vancouver. [24:31.420 --> 24:32.520] Vancouver Police Department. [24:32.680 --> 24:34.000] This is their use of force records. [24:34.180 --> 24:37.120] So if they use force against someone, there's a log of it. [24:37.240 --> 24:38.000] This is that log. [24:38.100 --> 24:39.160] This is just one page of it. [24:39.360 --> 24:44.260] This is, like, you know, obviously an Excel spreadsheet of some kind, and they printed it to a PDF for me. [24:45.200 --> 24:46.540] You know, there's some problems here. [24:46.660 --> 24:47.260] It's cut off. [24:48.640 --> 24:52.100] You know, the names are kind of split up all weird. [24:52.700 --> 24:56.500] This one here, just, you know, it's not even on the same line anymore. [24:56.500 --> 24:57.300] That's a problem. [24:58.920 --> 25:05.660] But this is still considered, like, a really high-quality PDF in terms of, like, extraction needs. [25:06.740 --> 25:08.940] We have a cool tool called Tabula. [25:09.580 --> 25:10.940] It can extract... [25:10.940 --> 25:13.880] It can turn a PDF into a CSV file. [25:14.240 --> 25:15.140] It's awesome. [25:15.300 --> 25:16.840] But it either works or it doesn't. [25:19.100 --> 25:24.340] Using this, like, all of us data journalists, like, this is a rite of passage, pretty much. [25:24.560 --> 25:26.640] Writing something that works with Tabula. [25:27.680 --> 25:36.760] Yeah, if you look at the contributors of this project, most of the people are, like, really highly skilled data journalists. [25:38.040 --> 25:40.320] Where it doesn't work is, like, right here. [25:40.320 --> 25:45.400] So, this PDF, this is another use of force. [25:46.420 --> 25:47.040] Oh, sorry. [25:47.100 --> 25:49.460] This is a complaints log from police near me. [25:50.360 --> 25:52.200] You can highlight the text here. [25:52.300 --> 25:55.240] But if you do that, highlight it, and you paste it, you get this. [25:55.400 --> 25:59.100] And it's, like, not in, you know, not in, like, a useful order. [25:59.280 --> 26:07.180] So, just because the document is OCRed, it doesn't mean that it's, like, going to be easier to extract it, necessarily. [26:08.860 --> 26:14.740] We need a way to, like, be like, okay, I need all of these blocks to call on, you know. [26:14.900 --> 26:19.220] And something, like, some kind of a structure to do this, it just doesn't exist. [26:19.640 --> 26:25.920] So, what we do as journalists, we have a one-off script that can handle the exact PDF. [26:33.150 --> 26:33.670] All right. [26:33.810 --> 26:39.410] So, more reasons why you help extract data from PDFs. [26:39.410 --> 26:42.630] Here's a historical climate data record. [26:43.310 --> 26:44.470] This is, like, some data. [26:45.370 --> 26:46.650] This is what it looks like. [26:46.750 --> 26:51.810] It's got this brown background that has these weird symbols that's going to make it hard to extract. [26:52.050 --> 26:54.490] But climate science information. [26:55.370 --> 27:01.010] So, by writing tools that can do this, you would not just be helping journalists about climate science. [27:01.010 --> 27:05.290] And just, here's another example of a similar format. [27:07.030 --> 27:07.710] Okay. [27:07.810 --> 27:11.210] The next problem we run into is OCR optical character recognition. [27:11.310 --> 27:16.690] And that is turning printed letters into computerized text. [27:17.050 --> 27:20.590] This is, like, a fun little portable OCR tool. [27:22.790 --> 27:23.590] And, yeah. [27:24.510 --> 27:27.050] So, this kind of shows here. [27:27.170 --> 27:29.310] But it's often not perfect. [27:29.610 --> 27:30.630] We can see that, yeah. [27:31.190 --> 27:32.510] Oh, there's a lot of problems. [27:32.510 --> 27:33.670] A lot of typos. [27:33.830 --> 27:35.090] But it's pretty close. [27:35.290 --> 27:35.770] You know? [27:35.890 --> 27:36.690] We got a little... [27:36.690 --> 27:39.530] We got a squiggly brace instead of a parentheses. [27:41.250 --> 27:42.570] Why is OCR important? [27:42.790 --> 27:45.170] Have you ever heard of the Panama Papers? [27:48.150 --> 27:53.850] This was, like, one of the first, like, huge OCR heavy and huge just data journalism projects. [27:54.590 --> 27:58.450] There was really good talk about this and a hope in the recent past. [28:00.270 --> 28:07.810] Long story short, an employee from the Panamanian law firm, Mossack Fonseca, leaked data. [28:08.070 --> 28:10.750] This is, like, the chat log that they had with the journalists. [28:12.190 --> 28:18.510] This law firm managed the money of some very powerful organizations and very, very wealthy people. [28:18.770 --> 28:21.090] So, this is, like, super rich. [28:22.110 --> 28:27.250] Really important is being able to get people places from the data. [28:29.630 --> 28:29.960] Yeah. [28:30.450 --> 28:34.350] So, being able to link and trace people through this data. [28:35.470 --> 28:36.030] Yeah. [28:36.490 --> 28:42.530] So, it was leaked to Deutsche Zeitung, the South German newspaper, or SC. [28:43.290 --> 28:44.790] It came in over time. [28:45.050 --> 28:49.150] So, eventually, it was 2.6 terabytes of data, 11.5 million files. [28:49.830 --> 28:52.770] And then here's kind of the breakdown of all the types of files. [28:52.930 --> 28:57.790] So, there's emails, database entries, PDFs, a lot of PDFs, and images. [28:58.710 --> 29:03.690] So, this is all places where OCR would come in helpful, come in handy. [29:04.410 --> 29:06.730] I talked to one of the journalists who worked on this project. [29:06.950 --> 29:20.450] They used a tool called Nuix, which is something that, like, police and, like, homicide investigators and, you know, those kinds of higher-level law enforcement agencies use to go through, like, evidence logs and stuff. [29:20.450 --> 29:21.850] It allows people to search. [29:22.090 --> 29:24.570] It does, like, some rough entity extraction. [29:24.570 --> 29:32.650] So, it'll, like, identify people's names and, like, places and the names of things, proper nouns, you know, that kind of stuff. [29:32.990 --> 29:40.710] It'll allow you to search them and be like, okay, give me all the documents talking about the queen or something like that. [29:41.330 --> 29:42.570] You can do that with Nuix. [29:42.890 --> 29:46.250] It has a poor OCR engine. [29:46.630 --> 30:04.090] So, if there was any kind of, I don't know, problem with the scan of the document, you know, some dirt on the Xerox copier or whatever, maybe the person's name, you know, got messed up a little bit and now they're not going to be in the search results because no one's going to read 11.5 million files to find it. [30:06.410 --> 30:07.250] I talked to him. [30:07.350 --> 30:12.230] I said, I was like, hey, what if someone developed a better OCR engine for Nuix? [30:12.370 --> 30:14.810] Would you rerun all the data? [30:15.030 --> 30:17.070] And he said, he's like, yeah, we would definitely redo it. [30:17.170 --> 30:20.090] We'd probably get, like, you know, a lot more leads out of it. [30:21.470 --> 30:26.190] So, like, being able to do better OCR would have real-world impact. [30:27.710 --> 30:31.590] So, this here is a file that I requested from a police department. [30:31.590 --> 30:36.290] This is a complaint log against police officers, and it's in handwriting. [30:36.630 --> 30:39.670] This is how they keep track of complaints against police officers. [30:42.430 --> 30:46.810] I'm involved in a project where I'm, like, investigating police misconduct in my area. [30:47.070 --> 30:48.430] And, you know, here we go. [30:48.550 --> 30:52.950] Favoritism, striptease dance, enticing female inmates to expose their breasts. [30:53.310 --> 30:54.750] Just a handwritten record. [30:54.850 --> 30:59.330] If I hadn't just stumbled across this, probably no one would ever know about this. [30:59.850 --> 31:03.610] And there's really no good tools to do this kind of stuff. [31:04.890 --> 31:06.430] There's, you know, there's research papers. [31:06.650 --> 31:09.850] There's machine learning techniques that supposedly can do really good on handwriting. [31:09.850 --> 31:12.550] But, you know, there's no easy-to-use tools. [31:14.090 --> 31:16.190] Something like this would be just so helpful. [31:18.670 --> 31:21.690] Okay, the next problem here is form data extraction. [31:22.070 --> 31:23.730] This is kind of similar to PDF extraction. [31:24.210 --> 31:30.310] Here we have Romney for president placing an ad on First Coast News. [31:30.530 --> 31:31.510] First for you, I don't know about this. [31:31.610 --> 31:33.090] This is a TV station. [31:33.090 --> 31:42.450] So, $150, we're going to run some airtime ads between July 30th to August 7th, 2012. [31:43.150 --> 31:52.170] So, being extracted data from this, it's not like in the PDF spreadsheet example where everything's in nice rows. [31:52.590 --> 31:56.710] You know, this data, the fields for this data, they're scattered all over the page. [31:57.610 --> 32:03.630] So, being able to turn a form into like a row in a spreadsheet would be very, very useful. [32:04.190 --> 32:16.430] And what this is, is the government does require political action committees and other political actors to file, you know, documents saying that they've purchased ads. [32:16.850 --> 32:22.090] The agency, sorry, the broadcasters file these, but there's no standard form. [32:22.990 --> 32:29.230] So, here's two different forms from different media organizations. [32:30.610 --> 32:32.970] And, you know, they have similar information. [32:33.110 --> 32:35.370] They look pretty similar, but they are slightly different. [32:37.990 --> 32:39.410] Here's another example. [32:39.730 --> 32:41.510] You know, this one is completely different. [32:41.530 --> 32:45.270] And this one's, you know, more similar to the others, but again, different. [32:48.930 --> 32:51.610] Yeah, so, this problem is huge. [32:51.610 --> 32:52.490] You see this everywhere. [32:52.870 --> 32:55.310] A lot of campaign finance data has stuff like this. [32:57.310 --> 32:59.530] Often a PDF is sometimes just images. [32:59.990 --> 33:02.590] I describe them as like visually structured. [33:02.930 --> 33:09.190] So, like, you know, there's lines, there's kind of loose structure, and it's organized like in a block kind of visually. [33:12.050 --> 33:16.750] ProPublica did a huge project, a crowdsource project to free the files, as they say. [33:18.250 --> 33:23.790] They built a web tool where you could go in, you could pull up one of these, and then you could annotate one. [33:24.010 --> 33:26.230] You can like, oh, this is a product field. [33:26.450 --> 33:29.130] This is the amount right here, the rate. [33:29.130 --> 33:30.290] You could do that. [33:30.670 --> 33:48.350] So, what they were trying to do is have people all these, and then the hope was that they would build a machine learning tool or, you know, some algorithm that would be able to look at this specific form, like this kind of form, and pull out all the fields they need. [33:49.850 --> 33:54.590] It worked pretty good, but, you know, it's nowhere near perfect. [33:54.830 --> 34:04.190] And this is a problem that a lot of people active in journalism, machine learning research, this is like another good data sets that we use, models and stuff. [34:08.450 --> 34:21.170] So, together, yeah, all the techniques that I listed here, extracting data from PDFs, scraping, one investigation will typically use all of them. [34:22.490 --> 34:24.850] So, this is a panel of my papers. [34:25.430 --> 34:30.650] There's a lot of really good information about how the investigation, if you're interested, you can search for it. [34:32.430 --> 34:37.610] Yeah, they had to turn PDFs and this raw data into something readable. [34:38.050 --> 34:43.270] Because, like I said before, it was human readable, but it needs to be machine readable. [34:43.270 --> 34:51.450] So, they are talking about how the computer hardware that they use, using the technology at the time, use more than 30. [34:53.230 --> 34:58.850] And it took them two months, the first data ready for reporters to even start to analyze. [34:59.990 --> 35:01.850] There's a really good talk about this. [35:02.150 --> 35:07.410] One of the reporters was saying, like, describing how they had to constantly up their servers. [35:09.950 --> 35:17.670] This was at a point where they were doing it in the cloud, but when they found the data, they had to keep it in a room and make sure that nobody else had access. [35:17.930 --> 35:25.670] They had to control access to it, so they couldn't just put it on Amazon, because they didn't want the government to seize it, or someone to seize it, because it was stolen property. [35:27.250 --> 35:32.610] So, yeah, scaling up this one server, you know, using the technology at the time, took a long time. [35:32.790 --> 35:37.890] So, tools to make this easier, to make this faster, would just be so helpful. [35:38.250 --> 35:43.130] And would really just unlock a lot of investigations that aren't really possible right now. [35:46.570 --> 35:47.130] Okay. [35:47.470 --> 35:49.810] And why do all this? [35:50.050 --> 35:54.670] The point of doing all these things, scraping, extracting, is to build data sets. [35:54.970 --> 35:58.450] Data sets are like the bread and butter of data journalism. [35:59.790 --> 36:06.830] And we need to build our own data sets, because most of the questions that we're interested in, there's no data set for that. [36:07.490 --> 36:15.970] If you want to know, like, I don't know, like a lot of climate stuff, hey, pipelines across the country, how many abandoned pipelines are there in the US? [36:16.490 --> 36:22.390] The answer is, there's a lot, because, you know, we've been mining in this country for a long time, and drilling. [36:23.230 --> 36:28.130] And there's no one nationwide data set describing that. [36:28.810 --> 36:30.610] Same thing with police misconduct. [36:31.830 --> 36:38.330] Newsrooms across the country are building their own data sets, trying to answer the question of, you know, what happens to bad cops? [36:39.690 --> 36:41.710] What causes police to be fired? [36:43.130 --> 36:45.470] How many use of force events are there? [36:45.910 --> 36:47.850] Does my state have more than other states? [36:48.210 --> 36:52.870] These are not questions that you can answer right now, because there's no data set for it. [36:54.030 --> 37:00.110] So if you want to help journalists, the single most helpful thing you could do would be to build a data set. [37:01.930 --> 37:03.110] But that's not easy. [37:03.350 --> 37:03.850] You know, that's hard. [37:03.950 --> 37:05.090] That's a long haul. [37:05.290 --> 37:06.830] That requires dedication. [37:08.370 --> 37:10.210] But it can have huge impact. [37:11.770 --> 37:13.130] I have this example here. [37:13.290 --> 37:15.950] This is the California reporting project starting in 2018. [37:15.950 --> 37:24.110] 2015, they gathered basically like anytime an officer got in trouble, they got that record. [37:24.410 --> 37:25.410] They requested that record. [37:25.550 --> 37:30.610] And they tried to request it from as many police departments across the entire state as they could. [37:30.770 --> 37:35.690] And that was like a team of 50 newsrooms across the whole state. [37:35.690 --> 37:41.570] And they were using very low, low technology methods to do that. [37:42.870 --> 37:52.830] So if you had, you know, some technology, some scraping, some requesting ways to automate records requests and manage those and pull in the data, you could do really great work. [37:53.830 --> 37:55.990] And you could do it on any subject you're interested. [37:56.270 --> 37:58.490] You know, I just talked about climate because that's a big one. [37:58.490 --> 38:03.070] But, you know, anything related to criminal justice would be huge. [38:03.230 --> 38:04.050] Campaign finance. [38:04.230 --> 38:07.590] There's campaign finance data like all over. [38:07.870 --> 38:13.170] But bringing it all together in, you know, a specific way, always helpful. [38:14.170 --> 38:22.110] Advertising information, you know, Google, Facebook, these companies are not transparent about advertising. [38:22.950 --> 38:26.230] You know, you can help build a data set about that. [38:27.090 --> 38:31.790] And, you know, these things would just be so important. [38:32.010 --> 38:32.610] So, so useful. [38:34.670 --> 38:35.150] Okay. [38:35.370 --> 38:39.090] So I'm going back to like the science process of data journalism. [38:39.770 --> 38:41.910] This is the data set building process. [38:41.910 --> 38:42.930] You're going to prepare. [38:43.250 --> 38:44.290] So you're going to do some research. [38:44.290 --> 38:48.310] You're going to figure out who or how to get the data you need. [38:48.670 --> 38:50.410] Who do you have to ask? [38:51.610 --> 38:52.950] Where do you find it? [38:53.490 --> 38:54.270] Things like this. [38:55.090 --> 39:01.590] You want to do this in like a repeatable and reproducible way. [39:02.690 --> 39:06.190] So that if you need to gather more data later on, you can do it in the same way. [39:07.550 --> 39:09.230] So, yeah, you just really want to be systematic. [39:09.630 --> 39:14.490] And, you know, what you put in in the beginning in the preparation phase, just pay dividends later on. [39:15.030 --> 39:18.350] The gathering phase, you know, that's where these tools come in. [39:18.530 --> 39:24.990] You know, extracting forms, getting stuff out of PEFs, gathering, you know, scraping tools. [39:24.990 --> 39:27.570] This is a very tool-based part of the process. [39:28.650 --> 39:31.630] And then once you have it, you also need to process the data. [39:31.810 --> 39:32.990] That's where more tools come in. [39:33.670 --> 39:36.550] These two parts of the process, very tool heavy. [39:36.790 --> 39:40.050] And this is the area where journalists need help. [39:41.490 --> 39:47.110] Right now, the machine learning tools that we have, they're oriented towards the needs of big tech. [39:47.110 --> 39:48.930] So, you know, recommenders. [39:49.070 --> 39:51.950] Yeah, if you need a recommender, there's so much technology available for you. [39:52.110 --> 39:57.150] If you need to analyze handwriting from a handwritten police report, you're out of luck. [39:57.330 --> 39:59.750] Because Google and Facebook aren't going to make money off that. [39:59.830 --> 40:01.930] So there's just not as much research into that area. [40:03.150 --> 40:09.490] So, yeah, once you've gathered and processed the data, analyzing it, that's your next step. [40:09.650 --> 40:12.090] You need to turn your data into a story. [40:12.610 --> 40:14.330] You know, computers understand data. [40:15.190 --> 40:19.810] Humans understand narratives and, you know, cause and effect. [40:20.070 --> 40:21.910] That's what humans understand. [40:23.190 --> 40:25.850] And that's not always easy for technical people. [40:28.050 --> 40:31.870] So once you have the data, like I said, the data is not the story. [40:32.070 --> 40:33.930] You can share the data, but that's not a story. [40:34.670 --> 40:36.030] You need to know your data. [40:36.710 --> 40:37.590] Read, read, read. [40:37.890 --> 40:40.290] So, you know, you've done all this work. [40:40.470 --> 40:41.690] You've built this amazing data set. [40:41.810 --> 40:43.070] You need to understand it now. [40:43.070 --> 40:46.230] You have a good idea of where it came from and how it was built. [40:46.430 --> 40:50.050] So you might know the pitfalls of it and what it's not, what it doesn't contain. [40:50.330 --> 40:55.870] But you really do need to read it to understand what it is, what it is you've collected, gathered. [40:58.490 --> 41:07.770] To go beyond the data phase and start publishing, I recommend that you partner with a local news organization. [41:09.250 --> 41:10.870] That's going to be your best bet. [41:10.870 --> 41:19.890] They don't have the resources or the know-how to build data sets to do these advanced scrapes, to build these tools, to extract stuff out of PDFs. [41:20.170 --> 41:24.010] Chances are they don't know how, but they do know how to build stories. [41:24.410 --> 41:27.730] They do know the context of the thing that you have, most likely. [41:27.730 --> 41:35.190] They do know, like, what kinds of things are important politically in your area that, you know, would have ramifications for your data. [41:35.410 --> 41:40.990] So it's potentially a very mutually beneficial partnership. [41:41.830 --> 41:44.330] You can also reach out to academics. [41:45.530 --> 41:51.650] Academics are kind of in a similar boat, you know, can always use help with technical type things. [41:51.850 --> 42:00.070] And obviously, a combination of the two, academics and reporters, that's even better. [42:00.550 --> 42:06.510] I helped with a story, ProPublica and the Palm Beach Post did something called Black Snow. [42:07.190 --> 42:09.310] It was about sugarcane burning. [42:09.870 --> 42:14.470] So in Florida, they allow sugarcane farmers to burn their crops. [42:14.910 --> 42:21.770] And they do it in an area, only on certain days when the wind's not blowing towards, like, where rich people live, basically. [42:22.450 --> 42:26.850] And the whole industry was just saying, well, it's fine. [42:26.910 --> 42:28.110] It's not toxic. [42:28.110 --> 42:32.010] Even though there's black smoke and people get hospitalized, they're like, no, no, no, that's unrelated. [42:33.190 --> 42:34.370] We have it tested. [42:34.550 --> 42:34.910] We're fine. [42:35.070 --> 42:36.290] There's no problems here. [42:37.770 --> 42:39.450] Local people didn't believe it. [42:39.930 --> 42:41.130] Scientists didn't believe that. [42:41.310 --> 42:47.930] So they actually partnered with academics, you know, pollution, like, air scientists or whatever. [42:48.670 --> 42:52.990] And they put out sensors and, you know, they did the science that the EPA wasn't doing. [42:54.130 --> 42:54.730] Yeah. [42:54.890 --> 43:02.030] So that's a case where partnerships, news and academia could be very, very fruitful. [43:06.090 --> 43:13.670] Something I like to tell people is, like, if you're interested in news and reporting and investigative stuff, there's a place for you in the industry. [43:13.890 --> 43:14.710] There really is. [43:16.010 --> 43:18.550] It can take a while to get involved. [43:18.830 --> 43:21.270] But the journalism scene is changing. [43:22.710 --> 43:28.290] Collaboration and cooperation amongst people is becoming increasingly embraced. [43:29.770 --> 43:30.450] Yeah. [43:30.650 --> 43:31.810] And local, local, local. [43:32.090 --> 43:36.010] Like, if you're going to reach out to someone, find your local newsroom. [43:36.010 --> 43:43.350] It might be small, you know, a lot of local newsrooms are only like two people, three people. [43:43.610 --> 43:49.110] I live in the capital of my state and the local newspaper here is three reporters. [43:49.630 --> 43:50.010] That's it. [43:50.270 --> 43:53.650] So local news is shrinking and it needs your help. [43:55.170 --> 44:02.210] And like I said, collaboration, you know, it just... you do such good work when you collaborate. [44:02.470 --> 44:05.430] So the more people that you can talk to, the better. [44:06.910 --> 44:13.190] If you are building a data set about something, find someone who covers what you are, like, interested in. [44:13.570 --> 44:18.710] You know, find... if you're looking into criminal justice type stuff, find a police beat reporter. [44:19.230 --> 44:23.310] Figure out who covers the thing that you want to work on and talk to them. [44:23.550 --> 44:24.850] Because that's going to be your best bet. [44:27.170 --> 44:28.710] Budget might be a problem though. [44:29.050 --> 44:31.230] News, you know, they only have a few people working. [44:31.430 --> 44:32.750] They probably don't have a huge budget. [44:33.010 --> 44:35.170] So they might not be able to pay you very much. [44:36.190 --> 44:38.890] That's just one of the sad realities of the industry right now. [44:39.850 --> 44:41.370] There's just not a lot of money in it. [44:41.650 --> 44:44.050] Especially smaller, smaller organizations. [44:44.370 --> 44:46.010] Like a lot of them don't have a freelance budget. [44:47.930 --> 44:54.670] So what I do find is, like, if you work with them, they will try to do their best to compensate you. [44:55.650 --> 44:58.670] But, you know, you can't do it for the money for the most part. [44:59.790 --> 45:06.230] Cons, potential cons of, you know, working with other newsrooms is they might ask for exclusivity. [45:06.510 --> 45:11.890] You know, they might want to say, hey, can we not share this with anyone until we publish our first report? [45:12.350 --> 45:17.850] That's going to be something that you're going to have to, like, really consider and say, is that... [45:17.850 --> 45:19.330] Am I okay with that? [45:21.950 --> 45:25.430] And that's, you know, it's always a trade-off because people want to release data now. [45:26.210 --> 45:38.110] But by holding on to it and working it, as long as you are actually going to share it, you could increase the impact that you could have by, you know, kind of carefully crafting your story and doing a thorough investigation. [45:38.590 --> 45:44.230] Rather than just gathering some data and just spitting it out there, which you could put a lot of work into it. [45:44.350 --> 45:51.130] And that can have very little effect because, you know, signal to noise ratio, there's a lot of noise out there. [45:51.370 --> 45:52.330] It's hard to find the signal. [45:52.490 --> 45:54.570] So your data could get lost. [45:55.490 --> 45:57.190] The other thing is bias. [45:57.650 --> 46:03.510] Really, really try... you have to try to not be biased, not appear biased. [46:04.350 --> 46:10.770] If you are scraping police sites and somehow getting, let's just say, use of force data, and you've built this huge database. [46:11.010 --> 46:19.490] If your Twitter feed is, f*ck cops, f*ck cops, the whole time, no one's going to want to work with you because they're not going to necessarily know that they can trust you. [46:20.770 --> 46:24.910] I don't believe that impartiality is, like, you know, a real thing. [46:25.110 --> 46:27.570] Like, everyone brings their own bias to whatever they work on. [46:27.930 --> 46:35.390] But trying to, you know, trying to, like, really try to be as neutral as possible goes a long way. [46:35.390 --> 46:38.990] And particularly stuff about that affects you personally. [46:39.710 --> 46:47.370] When I was working at a newspaper in Texas, I would have people all the time come up to you and be like, Hey, I got this story. [46:47.630 --> 46:48.850] I gathered all these documents. [46:49.030 --> 46:51.370] They're going to build this bridge in my neighborhood. [46:51.370 --> 46:54.130] And it's going to allow poor people to come over here. [46:54.250 --> 46:56.330] And here's all the reasons why it's f*cked up. [46:56.450 --> 46:59.210] But none of the reasons were the real reason that you knew they didn't want it. [46:59.210 --> 47:02.810] So it's things that affect people very personally. [47:03.050 --> 47:06.430] It's just hard not to be, you know, biased in your own favor. [47:08.410 --> 47:11.790] So just be mindful of these pitfalls. [47:12.690 --> 47:13.310] Okay. [47:13.510 --> 47:14.350] And then resources. [47:14.370 --> 47:21.290] If you want to get involved in data journalism, these are resources that you can reach out to, look into. [47:21.770 --> 47:23.490] NICAR and IRE. [47:24.150 --> 47:28.230] IRE is the Internet is Investigative Reporters and Editors. [47:28.450 --> 47:30.310] This is our trade group, basically. [47:30.650 --> 47:34.270] And we put on an awesome conference called NICAR. [47:35.770 --> 47:38.390] National Institute for Computer Assisted Reporting. [47:38.930 --> 47:44.310] Best of all, doing journalism with computers, which is not something I run this. [47:44.710 --> 47:46.930] But there's an amazing conference. [47:47.850 --> 47:49.510] And there's a mailing list on there. [47:49.630 --> 47:50.690] You can reach out on there. [47:50.810 --> 47:52.390] You can only be able to work with on there. [47:53.090 --> 47:55.230] Open News is an organization. [47:55.670 --> 47:57.370] They help facilitate collaboration. [47:57.730 --> 48:01.730] And they have, like, a technology side too. [48:02.030 --> 48:07.890] So, you know, if you just have a tool that you built and you think journalists would be interested, they have open news. [48:08.010 --> 48:08.850] They have a blog. [48:09.130 --> 48:10.430] They have a large network. [48:11.170 --> 48:12.570] Big local news. [48:12.750 --> 48:16.110] They also facilitate collaborations between news. [48:16.430 --> 48:20.010] They are particularly focused on trying to help smaller news. [48:21.290 --> 48:23.790] They are centered around data. [48:24.290 --> 48:29.470] So, if data set that you want to get out there, big local news will host it for you. [48:29.950 --> 48:32.290] They will put it on their site, even if it's huge. [48:32.870 --> 48:34.730] They have all kinds of ways to archive it. [48:35.130 --> 48:38.890] And even, like, they have a network of small news. [48:38.890 --> 48:44.090] So, you might be able to hook up with a drone through them. [48:44.710 --> 48:48.650] The IWL Society, a really cool organization. [48:48.650 --> 48:57.890] They have trainings and stuff around kind of, like, helping non-representable criminalism. [48:58.690 --> 49:04.050] And then City Bureau, a really cool example of what can be done. [49:05.290 --> 49:20.750] They hired, like, just average people, like, go and sit in on council meetings and meetings at the school, just training people to get in their local, their local basically. [49:20.750 --> 49:23.530] And they have, like, a subgroup called City Scrapers. [49:23.790 --> 49:27.790] You know, people scrape data from their local governments and gather it. [49:28.670 --> 49:38.790] So, yeah, these are resources that you should look if you're interested in data journalism, want to get involved, and want to try to find, you know, something local that you can get involved with. [49:39.030 --> 49:39.770] These are all good resources. [49:42.590 --> 49:48.670] So, yeah, so you can always contact me as well on Twitter. [49:49.090 --> 49:50.370] You can go to my site. [49:51.310 --> 50:00.370] Chat, you know, if you're trying to get involved or you have a tough, or, you know, if you think you have an awesome tool that journalists will need, chat to me. [50:02.310 --> 50:03.590] I'd love to chat with you all. [50:05.890 --> 50:07.670] And that's what I got for slides. [50:10.650 --> 50:14.010] So, if any questions, I would take them. [50:15.370 --> 50:16.030] All right. [50:16.150 --> 50:17.450] We have time for just a couple. [50:17.630 --> 50:19.650] We've got, like, two minutes for Q&A. [50:19.770 --> 50:22.030] So, we have one live Q&A question here. [50:23.130 --> 50:29.310] How much time do you spend scraping social media sites looking for data about individuals? [50:31.610 --> 50:33.650] Data on individuals, I don't do at all. [50:33.650 --> 50:36.950] Data on what YouTube recommends me. [50:37.930 --> 50:39.710] Actors are advertising on Facebook. [50:40.230 --> 50:44.050] I don't do that very much now, but that sometimes. [50:44.670 --> 50:47.870] Especially when I'm doing research into, like, and stuff. [50:47.970 --> 50:50.470] I want to see what they are paying people to see. [50:51.590 --> 50:52.330] Like that. [50:52.710 --> 50:53.750] And a related question. [50:53.870 --> 50:57.450] How much effort do you spend looking at the Wayback Machine? [50:57.850 --> 51:00.250] I work at the Internet Archive, so I'm particularly... [51:00.250 --> 51:02.230] Oh, I love the Wayback Machine. [51:02.930 --> 51:05.070] I use it, like, every day, probably. [51:06.810 --> 51:07.210] Yeah. [51:08.330 --> 51:08.850] All right. [51:08.930 --> 51:12.510] One more question from a speaker here in the audience, and I think we're going to be at time. [51:12.670 --> 51:12.930] Okay. [51:13.090 --> 51:18.550] So, I'm an editor with IEEE Spectrum, so it's great to see you looking at these tools instead of us. [51:18.810 --> 51:23.350] One thing I would also make an appeal for is maybe a RegEx helper. [51:23.350 --> 51:25.410] So, these are data sets, as you know, they're... [51:25.410 --> 51:26.830] We can't use machine learning on them. [51:26.930 --> 51:28.230] They're not really that big enough. [51:28.630 --> 51:33.370] But something, if I'm searching for, say, let's say, scheme programmers, job listings for scheme programmers, just... [51:33.990 --> 51:40.970] And then, I don't want to pick up ads that are pension scheme, but I don't want an ad that is, I want a scheme programmer, and I'm offering you a pension scheme. [51:40.970 --> 51:49.830] So, tools that would help people build RegEx, and also, journalists, I think, trust RegEx more because they feel more deterministic than ML. [51:49.950 --> 51:52.630] So, I would just add that to your pile of wish lists. [51:53.590 --> 51:53.810] Right. [51:54.010 --> 51:54.130] Yeah. [51:54.190 --> 51:55.030] I think that's a great one. [51:55.670 --> 51:57.430] RegEx, you know, super... [51:58.210 --> 52:05.490] Use it way more than machine learning, because it's understandable, you can explain it to people, and you know how it's going to work every time. [52:06.310 --> 52:06.790] All right. [52:06.850 --> 52:10.790] Well, thank you so much for your talk, Brandon, and thank you for the audience for participating in the talk. [52:15.110 --> 52:17.570] There were a number of questions in the Matrix chat stream. [52:17.970 --> 52:18.930] Go ahead and keep asking them. [52:19.130 --> 52:24.590] Brandon can access that and he can answer your questions there if you want to ask questions since the session is now over. [52:25.210 --> 52:27.070] Come back in 10 minutes. [52:27.210 --> 52:31.690] The next session will be, School districts should not be in the business of intelligence collection. [52:31.950 --> 52:33.110] Should be a good talk. [52:33.370 --> 52:40.030] And remember, the closing ceremonies will be simulcast in this room, not in Little Theatre at 6 o'clock. [52:40.170 --> 52:42.290] So, come for the closing ceremonies.