The State of Human Rights in the Pandemic | Stats + Stories Episode 151 by Stats Stories

price.jpg
gargiulo.jpg

Megan Price is the Executive Director of the Human Rights Data Analysis Group, Price designs strategies and methods for statistical analysis of human rights data for projects in a variety of locations including Guatemala, Colombia, and Syria. Her work in Guatemala includes serving as the lead statistician on a project in which she analyzed documents from the National Police Archive; she has also contributed analyses submitted as evidence in two court cases in Guatemala. Her work in Syria includes serving as the lead statistician and author on three reports, commissioned by the Office of the United Nations High Commissioner of Human Rights (OHCHR), on documented deaths in that country. @StatMegan

Maria Gargiulo is a statistician at the Human Rights Data Analysis Group. She has conducted field research on intimate partner violence in Nicaragua and was a Civic Digital Fellow at the United States Census Bureau. She holds a B.S. in statistics and data science and Spanish literature from Yale University. She is also an avid tea drinker. You can find her on Twitter @thegargiulian.


Episode Description

Almost every day we seem to get new data about the COVID crisis. Whether it’s infection rates, death rates, testing rates, false-negative rates, there’s a lot of information to cull through. Making sense of COVID data is the focus of this episode of Stats and Stories with Megan Price and Maria Gargiulo.

+Timestamps

2:55 What’s the reaction been?

11:10 How important is the information in supporting these decisions.

14:30 What stories are we missing?

18:14 Schools and Covid.

23:30 How to Make Sense of all of the COVID data.


+Full Transcript

Rosemary Pennington: Almost every day we seem to get new data about the COVID crisis. Whether it’s infection rates, death rates, testing rates, false-negative rates, there’s a lot of information to cull through. Making sense of COVID data is the focus of this episode of Stats and Stories where we explore the statistics behind the stories and the stories behind the statistics. I’m Rosemary Pennington. Stats and Stories is a production of Miami University’s Departments of Statistics and Media, Journalism and Film, as well as the American Statistical Association. Joining me are regular panelists John Bailer, Chair of Miami’s Statistics Department and Richard Campbell, former Chair of Media, Journalism and Film. Our guests today are Maria Gargiulo and Megan Price of the Human Rights Data Analysis Group, or HRDAG. Price is the Executive Director where she’s worked on projects related to human rights issues in Guatemala, Colombia, and Syria. Gargiulo is a statistician with HRDAG and was also a data science fellow at the US Census Bureau. They’re here today to talk about some of the group’s work on the COVID crisis. Maria and Megan, thank you so much for being here.

Megan Price: Thank you for having us.

Maria Gargiulo: Yeah, thank you.

Pennington: Megan, I’m going to start with a question for you. So, HRDAG describes itself as quote -a non-profit, non-partisan organization that applies rigorous science to the analysis of human rights violations around the world- end quote. You’ve been publishing a bit about COVID including some pieces in Significance Magazine, how do you situate the work on COVID within the human rights framework that your group, you know, is sitting in?

Price: Yeah, that’s a great question, thank you. Well, everything that we do stems from the Universal Declaration of Human Rights. That’s the starting point for all of our thinking about our work and we’re also just humans. And so, when this crisis started, of course, understanding it and trying to just get some handle on how to even go about making decisions about how to live our lives was at the forefront of all of our minds. And through our work, we’ve had so much experience as what we think of as science communicators, thinking about how to explain really complicated, emotionally-fraught ideas to folks who may not have much or any grounding in statistics or data analysis or science work. And so, we really felt like that was not only a role that we could step into but also something that could help us as a team to focus on something that felt urgent and useful.

John Bailer: So, what’s been some of the reactions that you’ve had to these columns? I mean, you’ve been writing a number of these explanatory pieces to try to convey and communicate some of these issues that are emerging with the pandemic. Do you have any feedback?

Price: We have, and I have to say this is a little bit biased because it was one of my friends, but my favorite reaction has so far been to a column we wrote in a literary magazine called Granto, which is perhaps not a common outlet for statisticians, about essentially what role does stats play in interpreting screening tests and how do you know what your personal screening test means? And one of my particularly math-phobic friends reached out and said I actually understood that, thank you. And that is just the most gratifying feedback we can get.

Richard Campbell: So, can you talk a little bit about the undercounting of COVID infections and what some of the obstacles are in getting good data in your work?

Price: Sure, I think I might start that- I’ll start with your second question which is getting good data in our non-COVID work. Our non-COVID work is focused on human rights violations as our name implies, and specifically on types of violence. And there are a whole variety of reasons why that might not be fully documented. And some of them are pretty benign, some of them are just the violence wasn’t witnessed or the individuals who are doing the best they can to document and describe that violence just didn’t have the resources that week, didn’t have enough people the ground and then other times they’re pretty intentional, a lot of violence is hidden and very intentionally kept from the public eye, and so I would say that a variety of those same things are happening in our attempts to understand COVID-related deaths. There are certainly a lot of incentives to not categorize something as a COVID death or to choose different metrics in terms of positive rates of tests or numbers of tests or who gets tested, and those incentives are not always going to lead to the most complete and the best data collection, unfortunately. But then again there are also just lots of perfectly benign reasons in New York at the peak of the outbreak there, everyone was just overwhelmed, and the idea of writing everything down, you know, certainly came far lower on the list of priorities than helping everyone you could help. And so that’s, I think where statisticians can come in and say look, you don’t have to write everything down, we can use the tools in our toolkit to fill in those gaps.

Bailer: So, you write in one of the essays that your group wrote that science starts with theories and stories about how the world works. Now, does the idea of trying to- you know this is a really hard story to tell- that people, you know, they may have last thought about theory as something they heard about in the scientific method when they were at school and didn’t really think a lot about since then. What are some of the challenges and some of the potential solutions when trying to communicate these more complex stories? Whether they are SAR models and some of the nuance of finding them to an audience that may not think a lot about theories and background?

Price: Maria, can I put you on the spot? Do you want to take that one?

Gargiulo: Yeah, sure. So, when I think about theories personally, the thing I really like to try and figure out is how do I test if I think a theory holds in this situation. And I think in communicating science, giving people things to look for is really helpful so I think a lot about- I think the piece you mentioned the Director of Research, Patrick Ball wrote and he kind of provides a list of like things you might look out for, so for example when we’re testing a theory, a rigorous theory is really careful about the types of assumptions it makes. So, in order to come to our conclusion that we made about the way the world works, what are the things we assumed? And once someone kind of delineates those really clearly it’s a lot easier to say oh I think those assumptions are reasonable. I can kind of hold on to the threat here, that makes sense, or I don’t think that’s true. And if that’s not true you might have a way to start thinking about oh, if that’s not true, what other things might not hold? So, I guess trying to communicate the ideas that let people test the theory for themselves, even if that’s an informal way, I think that’s really important for things like this.

Campbell: So, this morning, speaking of stories, there’s a story on the front page of the Dayton Daily News about area residents could be part of a virus study. And through this podcast and talking to scientists and statisticians I’m just confounded by the fact that we haven’t done more random studies of COVID. And I’m wondering both at the regional level and at the national level; and that Ohio just now is going to do a random study of 1200 randomly selected participants. What’s the problem here? I mean we’ve talked to statisticians who have said this should have been going on much earlier and we’d have a much better idea of who’d infected and who’s not. And I’d like both of you to talk about this.

Gargiulo: I can start. I think for me, and part of this is I don’t actually understand, to the full extent, resource constraints right now, but I think a lot of this is resource constraints. It’s a lot easier, I think, to say oh we have these 20 people in the hospital right now, we can test them, we can talk to them, we can do these things. Rather than okay, you know thinking about what does a representative sample look like and finding that representative sample within the community. Do we actually want it to be fully representative in that normal sense? Do we want to oversample certain groups who want to sample other groups? So, I just think it’s harder. It’s- you know, convenient samples are nice because they’re convenient. Random samples are hard because they need to be really carefully constructed and under constrained resources, it’s not clear to me how feasible that is or how hard or easy it is.

Price: Yeah I’m mostly going to second everything Maria just said. I mean I think much like kind of prioritizing that happened around New York around do we just try to get everyone we can to the hospital? Or do we keep perfect records? I mean one of the things that I think is hardest about this moment in time is that just everything needs massive resources and figuring out how to allocate those and how to balance the really urgent today priorities, while also like recognizing that we need to make some long term- we need to make some decisions with a long term vision that you know our future selves will be grateful for, and I’m certainly grateful that it’s not my job to make those kinds of decisions. And I think also coming from- I have a public health background where, you know, there are lots of situations where you can’t do a randomized control trial for ethical and logistical reasons and I think there’s a certain amount of that at play here, too, and I think that because of the way the United States is set up- you know, something that public health has done for years and years is to identify these natural experiments that happen because different regions make different decisions and take different actions, and so, personally, I think that it’s as important and as valuable to identify those comparisons that are more readily available as it is. I mean I certainly- let’s also do randomized trials and let’s get those organized, but I think that both of those things happening at once is the way to go.

Bailer: So, this part of the conversation makes me think a lot about the value of information. You know, so, in some way what we’re saying is that we’re taking these samples of convenience we’re looking at individuals who are probably symptomatic and that are showing- that are of gravest concern, but they’re telling us about, you know, are people that are symptomatic, are the disease, do they have the disease as opposed to knowing what’s going on in the population? And so I think it’s a hard question, you know what’s- you talk about decisions and what’s the value of the information that you gain from knowing more about what’s going on in the population than knowing about what’s going on in some small symptomatic subset of the population, and I agree completely about the, you know, that resources have to be allocated in a way that – there’s a triage component to this, to solve this problem in a sensible order, but if we’re – how important is it to have the information that’s unbiased and kind of meaningful for supporting these decisions? That’s-

Price: I mean, yes.

[Laughter]

Price: And you know, but again I think that that’s where, you know, as statisticians I mean we should always recognize when our data are incomplete and biased but we also shouldn’t just sort of throw up our hands and say well, then we can’t use that data. There- we should recognize when a particular class of methods is appropriate to either adjust for those things or to account for them in some way. And I think also you know kind of coming back to natural experiments you know we do have a couple of really-I hesitate to use the word interesting in this setting, but really interesting things that have happened, specifically on cruise ships, which is a closed population and where they were able to collect data about every single person and so that again like the population on a cruise ship isn’t going to represent general populations anywhere but it gives us a chance to say okay if we test every single person, what’s the difference that we’re seeing between symptomatic and asymptomatic and I know here in San Francisco they did a very similar thing just at a microlevel they picked like a four-block radius in one of the neighborhoods in San Francisco and said we’re just going to test everybody in this four-block radius. And so, I think there’s also opportunities to do that kind of hyper-localized thing to start to learn more information.

Campbell: You know what that- what did that yield? That four-block study that was interesting?

Price: Oh man, that yield- so this was a UCFF study and in partner with another organization that I’m not going to be able to come up with but what they found was the kind of racial disparity that we’re now seeing at large, especially in the latest New York Times data. So, in this four-block radius this four-block neighborhood; it was in the mission. And I can’t remember now but I want to say like maybe five percent of the Hispanic residents were positive. Not necessarily symptomatic, not necessarily [inaudible] but they gave everyone a diagnostic test and they were positive. They literally could not find a single Caucasian member of that neighborhood that tested positive.

Pennington: Wow. That’s incredible. You’re listening to Stats and Stories and today we are talking with Maria Gargiulo and Megan Price of the Human Rights Data Analysis Group. We see a lot of coverage in news media of infection rates, of death rates, of hospitalizations. Given the work that you have been doing on HRDAG on this issue are there stories in the data that are under-reported that you think people should be paying more attention to?

Price: That’s a great question. Um. Hmm. To be honest I can’t really think of one because the one that has been pressing on my mind the most has been the racial and ethnic disparities and I think that we are starting to see more attention being paid to that so I’m grateful to see that coming to light. You know I think as with anything else that’s really scary, we’re seeing a lot of stories about how bad things can be, but I’m also really hesitant to say hey we should tell more stories about people who are recovered and are fine because we need people to take action to protect their community. So no, actually on balance I kind of think that most of the stories are out there. I don’t know, Maria, what do you think?

Gargiulo: So, a story I would like to hear more about in a non -U.S. context is what the intersection of say COVID at conflict or COVID at displacement is going to be. So, I’m thinking for example COVID arrives at a refugee camp, you know what happens? And that is terrifying because I think the only conclusion that I come to in my head is the results are going to be grim, but what does that like- what happens? Do people leave the camp? Do people stay in the camp and get sick? So that’s a space I’m watching to just see what happens and also how does humanitarian aid react to that? I have no idea. So we don’t- you know so that’s not so relevant in the U.S. context, but you know as we consider COVID as a global pandemic I think that’s something I will be watching and really hoping goes better than I’m expecting it to go.

Pennington: Do you know of any work that’s looking at infection rates along class lines? Because I would imagine that there could be particular breakdowns along with class in some places. And it’s not something that I can remember having seen like you’ve pointed out Megan, I think the reporting on race has just sort of started emerging in a lot of the coverage but I can’t remember seeing much about class. I’ve seen it about the geographic breakdown like rural versus urban, but then this issue of are poor communities being impacted more or less or anything like that, so I just wanted to ask that question.

Gargiulo: Yeah, not that I’m aware of and in fact, earlier in the pandemic, which I mean is such a weird way to describe things because as much as we’re all in this time dilation, you know it honestly hasn’t been that long, but earlier I did see some comparisons of occupation, of risk and infection rate by occupation which is a bit of a proxy for that and I, haven’t seen much follow up on that. so, I think that that is another thing that deserves more attention.

Bailer: And it seems like some of the things related to- some of the exposures related to occupation may also play out in terms of living conditions. So if you’re- the concern I guess, in the U.S. it’s something like 40% of the fatalities are in nursing homes, you know and as you look in other environments it tends to be where people are living in more group housed environments and if you live in a high-density area as well as go out and work it seems like that just kind of explodes it. So that runs a little bit counter to my earlier comment about who we’re studying and how. And in some ways, if we’re looking at the people who are going to be most dramatically impacted then you might want to be targeting what we’re doing. I thought that I saw that there was some recent work that’s starting to come out related to the COVID impact in Central and South America, and I won’t swear to it; I’d have to dig that up too, so I’m not sure.

Price: Yeah, there has been and so I guess that’s sort of the coda, to my comment to- you know, what stories are getting told is highly correlated with what media source you’re consuming and so, yeah. Because we have a lot of projects and partners and collaborations in Central and South America, I have a lot of sources who have information on that part of the world and so yes there is a fair amount of coverage coming about how the infection rates are unfolding there. But yeah I’m not seeing that in perhaps more conventional mainstream US media.

Campbell: One of the things that relates to the sort of class problem that Rosemary brought up is there’s a lot of discussions now should we send our kids back to school and part of is it is that wealthier school districts are in better shape to do this than poorer school districts and I guess my question is if you have children or if you don’t have children, I mean what should we do? What’s the best advice? Or is it all sort of just a regional or local problem?

[Laughter]

Price: So, I have two kids. My daughters are 14 months and 3 and a half years old and they’re at daycare right now, and I kind of am both like really happy about that and really scared about that. and also, my husband is a public-school teacher so schools and kids and what to do is like all we think about right now. And you know, it’s interesting I think that operationally it has to be regional because it’s going to be so contingent upon just what the situation is on the ground but on the other hand, you know a top-down national you know like threshold guidelines; you can only even consider opening up the schools if your case count per capita is X. You know to safely have in-person learning you need Y dollars per student. We’re going to provide these grants that are going to cover you know PPE and sanitation services. I mean that kind of thing can be in a bigger framing, but yeah I mean just to kind of answer the question as a statistician, I have no idea.

Bailer: Well, you’re telling us something because you’re both working from home now. So, there’s clearly a policy decision that you’re making at a very local level about kind of what can we do to prevent potential infection within our community, within our workforce. Maria, did you want to add to that too?

Gargiulo: I mean really just to reiterate what Megan said and I really have spent no time thinking about this but I think like you know one I have no idea like statistically speaking and two though I think like the whole idea of like either all schools opening or no schools opening, that’s not it for me. Like I think these decisions really, they need to be made in the communities because if something goes wrong it’s those same communities that are going to be affected. So, it’s not just about are the kids in school but if the kids are in school and something goes wrong what are the potential repercussions? And I don’t think while we might have really great- it would be great to have some national guidelines to help school districts out at the end of the day the national government isn’t suffering if something bad happens, the community is suffering; they need to make that decision.

Bailer: But those communities that need to make decisions, just getting back to what you’ve been producing, and some of the things that you’ve been writing about are that they need good data; they need good information. And in some ways you know, you- if you’re a- so now, Maria I’m going to make you the superintendent of our local school district.

Gargiulo: Excellent.

Bailer: Congratulations and condolences, by the way, because you’re the one that has to make a decision about how many kids can come back to school. How should they be spaced in their classrooms, how should- you know, all of these things? And by the way you’ve got ten parents on the line waiting to talk to you about why they need their kids back in school. I mean, so how does science help, you know, how does science and the study of some of the data that’s associated with this pandemic- how can that be communicated to help these local decision-makers that you’ve appropriately mentioned to make the calls that they need to make?

Gargiulo: Yeah, so I think that if I were the superintendent in charge of this I’d want to talk to different people. So, I’d want to talk to these parents on the phone, I’d want to talk to my teachers. Do they feel like, you know, part of this is not necessarily about the science, the ground troops it’s also like do you feel safe going to work? How do the kids feel about going to school? I’d love to talk to some of them and figure out you know if you had the opportunity to go back to school would you feel safe doing that? or would you just sit in class being really anxious all the time you know thinking today is the day I’m going to get sick or I’m going to get one of my classmates sick or my teacher? So, I’d want to start with conversations there and then I’d start asking questions like how much money do I actually have for personal protective equipment? do I have backup plans for when and if things go wrong? What do those look like? What are the effects of starting a school year in person and then sending kids home? This is I think a different kind of data collection that isn’t necessarily like you know biological data about the virus. You know like do students have the internet at home, right? These are other types of data collections we need to do. So virus biology and you know everything we know about the spread about the epidemic I think helps us make decisions about okay we can only have you know 15 students in the classroom so maybe 50% full, we’ll call that, but also then there are these other types of data collection that need to happen that really has nothing to do with the spread of the virus and everything to do with you know the upside kind of social dynamics of what’s happening to all kinds of angles, so I really want to get more data sources involved even though it would complicate things.

Bailer: Well, you know, if this stat thing doesn’t work out I think there might be a superintendent gig in your future.

[Laughter]

Gargiulo: My retirement job.

Bailer: That’s a well thought out response.

Pennington: I’m going to swoop in with a final question and steal it from John, you know, people I think are overwhelmed with data related to this you know because it’s coming out every day. Given the work that you’ve been doing what advice would you have for our listeners about you know how to wade through the data and how to make sense of it in their own lives?

Price: You want to go first Maria, or do you want me to?

Gargiulo: No, you go first.

Price: So, you know what I personally have been doing has been to have really strict news and data consumption diet and to really stay focused hyper-locally. And it’s hard because my phone at any moment wants to tell me about these headlines about how there’s a spike in cases in the state of California, but the state of California is really big and in my city, there is an increase in case but it’s not quite as scary and so working really hard to contextualize those big stories with the hyperlocal data and I do think that that’s actually something that most cities and counties that I have looked at have been doing a really good job of being transparent and saying look this is what we know and this is how we know it but that said, I am a statistician and so I find data very comforting. And I think that if that is not the place you’re coming from then even that can still feel really overwhelming because these hyperlocal dashboards do still contain a lot of information and they get updated every day and so you know in that case what I would really recommend is to identify one or two sources who you absolutely trust who are filtering and contextualizing that information for you and that may be a news source that maybe a friend that may be an expert on twitter, it can be hard to vet those sources and to really know that you’re getting really reliable information that way but I think that if you personally don’t have kind of the comfort to deal with that raw data that’s coming at you that would be my recommendation.

Gargiulo: Yeah, I’ll just kind of second everything that Megan just said I, in particular, don’t look at the data every day. call me crazy but I do read a lot of epidemiologists on twitter and you know it’s really nice I get really good synthesis and for me they also sometimes kind of write about kind of the news studies that are coming out and I could sit down and read those studies and you know I might understand bits and pieces of them but for me, it’s nice to have these data contextualized with like what are the advances we’re making, where are we making progress, where are we really struggling right now? And getting that from someone who is an expert not only in that field but there are lots of folks on Twitter being really thoughtful about science communication, that’s where I’ve been doing a lot of my learning and I think that’s just helped me kind of you know to find the signal in the noise and get out at least what I want to understand which is mainly like what does the general trajectory look like? And Megan is right with these hyperlocal news sources. Like that’s really helpful to me especially because I have not been really leaving my house so really the most relevant thing for me is that hyperlocal geography but then also understanding like here’s the trajectory we’re going on in terms of scientific methodsso balancing both like research like with what’s actually happening is what I look for and I just try and read experts on that.

Pennington: Well, Megan and Maria, thank you so much for being here today.

Megan and Maria: Thank you guys so much.

Pennington: That’s all the time we have for this episode of Stats and Stories. Stats and Stories is a partnership between Miami University’s Departments of Statistics and Media, Journalism and Film, and the American Statistical Association. You can follow us on Twitter, Apple Podcasts, or other places where you can find podcasts. If you’d like to share your thoughts on the program send your emails to statsandstories@miamioh.edu or check us out at statsandstories.net and be sure to listen for future editions of Stats and Stories, where we explore the statistics behind the stories and the stories behind the statistics.


Risk Assessment Biases | Stats + Stories Episode 147 by Stats Stories

tarakshah.jpg

Tarak Shah is a data scientist at HRDAG, where he cleans and processes data and fits models in order to understand evidence of human rights abuses.

Prior to his position at HRDAG, he was the Assistant Director of Prospect Analysis at University of California, Berkeley, in the University Development and Alumni Relations, where he developed tools and analytics to support major gift fundraising.


Episode Description

Protestors have taken to streets across the U-S this summer in order to fight back against what they see as an unjust criminal justice system – one that treats People of Color in prejudicial and violent ways. The concern over racial bias in policing has long been a concern of activists, but there’s an increasing focus on other ways racial bias might influence decisions made in America’s courts and police stations. The statistics related to race and the criminal justice system is a focus of this episode of Stats and Stories.

+Timestamps

What spurred this research? (1:33)

What is a risk assessment model? (2:12)

What ore these tools suppose to do? (4:00)

What is fairness? (5:18)

What did you learn? (10:12)

What is the takeaway for the layperson? (15:20)

What’re some parallels to this work? (19:35)

How do you make this interesting? (22:07)

What’s the flow of your work, for reproducibility? (25:00)


+Full Transcript

Rosemary Pennington: Protesters have taken to streets across the U.S. this summer in order to fight back in what they see as an unjust criminal justice system. One that treats people of color in prejudicial and violent ways. The concern over racial bias in policing has long been something activists were thinking about, but there’s an increasing focus on the other ways racial bias might influence decisions made in America’s courts and police stations. The statistics related to race in the criminal justice system is the focus of this episode of Stats and Stories where we explore the statistics behind the stories and the stories behind the statistics. I’m Rosemary Pennington. Stats and Stories is a production of Miami University’s Department of Statistics and Medial, Journalism and Film, and the American Statistical Association. Joining me are regular panelists John Bailer, Chair of Miami’s Statistics Department and Richard Campbell, former Chair of Media, Journalism and Film. Our guest today is Tarak Shah. Shah is a data scientist at the Human Rights Data Analysis Group, or HRDAG where he cleans, processes, and builds models from data in order to understand the evidence of human rights abuses. He was the co-author of a report released last fall that examined whether a particular risk assessment model reinforces racial inequalities in the criminal justice system. Tarak, thank you so much for being here today.

Tarak Shah: Thank you for having me.

Pennington: Could you explain what spurred this particular bit of research into this risk assessment model and what your report found?

Shah: Sure. So, there’s been interest in these pretrial risk assessment models in particular for a little while now. Partly because of how much public opposition has grown to the money bail system. And because of that these other risk assessment tools have been proposed as more objective or more neutral alternatives to decisions by judges which may be considered biased. And this kind of fits into that atmosphere.

John Bailer: So, can you talk- just to take a step back, just to help fill in the gaps for people that are new to this? I mean- this idea of what happens in a pretrial process. And then as- sort of to jump off of that, what does a risk assessment tool do in the context of this pretrial process?

Shah: Yeah. Excellent question. So, in general when a person is arrested they- a court must decide and depending on what state you live in they’ll have either one to two days to make this decision. Whether you can go home while you await the beginning of your trial, or whether they need to take some kind of action, whether that’s detention or some kind of supervisory condition in order to ensure that you will appear for your court date and or that you will not be a danger to your community during the time that the trial hasn’t happened yet. So, those decisions are being made, as I mentioned before, the judges have historically relied on bail to make sure that people appear for their court dates, there’s increasing recognition that that disproportionately harms people who are poor and so there’s been interest in alternatives, but the basic kind of decision that either a judge or some kind of decision-making system is required to make has to do with usually one or both of those two elements that I mentioned. So, either whether a person is going to be a danger to their community or whether they are going to flee the jurisdiction and escape accountability. So that’s kind of the decision in front of us, historically made by judges. More and more judges are getting information from these risk assessment tools as just additional information to make that decision.

Pennington: So, are these tools like a technology that they’re relying on to help them understand what possible behaviors of particular defendants might be?

Shah: Yeah exactly, and they so these are tools that will take data in about characteristics of the arrested person. So, things like their age and sex and other demographic information as well as things like their arrest history or other kind of encounters with the court system. And I mentioned those kinds of two high-level principles like danger to the community and risk of flight. In practice, like those are kind of fuzzy concepts that need to be made concrete when we’re talking about actual measurements, and so the way those things get measured is in terms of danger to the community we- those who developed these risk assessment tools look at rearrests. So, was a person who was going to face trial- were they rearrested before their trial concluded? Sometimes that’s narrowed down somewhat. So maybe in a given jurisdiction, they’ll only look at felony rearrests or violent rearrests and there’s all sorts of logic that goes into what counts in each of these categories. Similarly, with flight risks- that’s also a little bit fuzzy in some ways, and so what we can measure is failure to appear for a court date.

Richard Campbell: I was interested in you talked about fuzzy right there, and you talk about a definition of fairness, which seems like- that’s not something I hear statisticians talking about very much, having done a hundred and fifty shows- No offense, John. But how do you- what is the definition of fairness in risk assessment modeling?

Shah: Yeah, that is an excellent question and in fact, one that there are multiple definitions fairness in this context and I will give a couple of examples. So one is maybe just going back to what these models look like there are very often logistic regressions or some other kind of predictive model which will classify people into yes, they will re-offend or no they will not or something like that. And so, within that context, there’s kind of basic notions that I think anybody without any kind of statistics background might be able to pick up. Things like demographic So, do black individuals who appear before the court- are there similar decisions made about them versus white individuals? So, in practice there’s- we tend to rely on somewhat more complicated definitions of fairness. The idea being that- well, let me give you an example, so in addition to demographic, there’s things like equal false-positive rates. So, the argument here is that the biggest cost of one of these risk assessment decisions is when somebody has to be incarcerated or otherwise supervised as a result of that score. And so false positive here is somebody who is determined to be high-risk by this tool, but who in fact would not have gone on to re-offend or miss their court date if they were left to go home, so that’s one example. Another example is just kind of equal calibration across race groups. So that means like if you- so your logistic progression puts out the number 0.47, so like 47% likelihood that you’re going to re-offend or something. So white people who get that score people versus black people who get that store. SO, everybody got a 0.47 among the white group did about 47% of them reoffending versus similar numbers for the black group. There is=- so we have equal calibration, we also have false-positive rates. We also have a similar notion of equal false-negative rates- that is people who did go on to re-offend or miss their court date; how often were they actually labeled high risk or low risk and are those rates equal across race groups or other protected characteristics. The kind of challenge, well one challenge is just what I mentioned is that there are multiple different definitions and there’s not an official correct definition and in addition the examples I just happen to give there is like an important result in fairness which is that they are under most realistic circumstances, they are mutually incompatible. That is, you can’t meet all three of them at the same time. So, it’s- which I think makes sense from a non-statistical perspective. People have different notions of what fairness means; I think. But it does make it challenging to talk about fairness in these contexts and I just want to kind of add, so everything I’ve been kind of talking about is kind of fairness within the system defined by the model. Like where the outcomes that are measured in the model, and what were the data inputs that went into the model and they’re kind of taking those data as a given. A separate level of analysis here for fairness is whether or not there is bias in the data itself. Whether these measures are fair measures of the thing that we’re interested in measuring. And so that kind of goes back to what I said before where I said we have these notions of danger to the community or flight risk, but in practice when we’re talking about creating a progression model we need these measures and what we have is rearrests or failure to appear for court and so often there’s problems with both of those measures. A lot of people fail to appear not because they fled the jurisdiction but because they forgot or the court date got changed and they moved so the postcard that they received, they never got it. And similarly, with re-arrest, the assumption there if you’re using that data is that an arrest is an unbiased measure of criminality or dangerousness and there’s a lot of evidence that that’s not the case.

Bailer: So, what did you learn? Let’s get back again to kind of the punchline to the work that you’ve done. You’ve helped us frame what a risk assessment model is and how that’s being used in the context of establishing or evaluating fairness.

Shah: So, in this particular research, we were looking at a tool used in New York City to determine eligibility for a supervised release program. This is kind of an alternative to being detained while you await your trial. So, the idea was that individuals that get a low-risk score would become eligible for this supervised release program whereas those who received a high score would not and the alternative is that you are detained. As I mentioned we looked at some of these different fairness measures that I mentioned such as false-positive rates and accuracy across race groups, and the particular model met some of those but not others so it had much higher false-positive rates for black and Hispanic people than it did for white people and also in terms of demographic parody it was much more likely to give black individuals a higher risk than white individuals. But as I mentioned- well I don’t have anything else to say about that. but one kind of deeper thing that came out of that was a couple of things that we noticed about the data and the process used to build the model itself. So, in terms of the data, we had that question about whether in this case felony re-arrest was the outcome variable that the developers were modeling in their regression, whether felony re-arrest is a fair measure of dangerousness. And that was an important question for us because when we looked into it a little bit the training data for this model had all been collected during the height of New York’s Stop and Frisk program. And sometime after that data was collected New York courts themselves had determined that this was an unconstitutional program because it was disproportionately applied against black and Hispanic people. And something like 87% of people who were stopped under Stop and Frisk were black or Hispanic. So unfortunately for us I guess the arrest data that we got that was the training data for the model did not contain information about whether each arrest was a result of the Stop and Frisk or anything else, however, we’re able to kind of look into- like make some inference about how many of these arrest outcomes could have been affected by the Stop and Frisk program, so we looked back at what the most common arrest resulting from a Stop and Frisk stop were and they were either drug-related or weapons-related, so then we went back to our data and found that just under 40% of the arrests in our outcome data was either drug-related or weapons possession related. So, we can’t say that all of those were Stop and Frisk, but we can guess that a good number of them were because this was during that program. We also found- so one of the things that we had to do when we were writing this report was attempted to recreate the scoring model that’s used. So we kind of read the paper and it goes through logistic regression and so it kind of goes through the variables and so forth and so we have the same training data as the original developers did so we were trying to replicate their model at the beginning. And we ran into some challenges. We had the same train and test split as the developers did and we got the same coefficients for each of those different pieces of the data and that all made sense but then when we looked at the final model that New York was using the point scores that they gave for each characteristic did not match up with the coefficients that we found in regression and in fact appeared to be kind of picked and chosen from the different splits of data. So there were like some coefficients that were found from fitting the model to the training data, some coefficients that came again from the text data and some that came from fitting the model to the entire data set and it wasn’t- so that, I don’t think we would have known if we had not tried to replicate the model, to begin with. Once we saw that we contacted the developers and tried to get more information about what was going on and the best that we can find out is that there was like a decision process by a committee that was like oh this looks good, this doesn’t look good and they kind of worked their way to scores. And I point that out because that’s a little bit separate from the types of fairness that we were talking about but I think these tools often get packaged as like objective or neutral thing and here we see an illustration that that packaging is really hiding a lot of political decisions that are going on.

Pennington: You’re listening to Stats and Stories and today we are talking to Tarak Shah of the Human Rights Data Analysis Group. So, I’m going to go back to that sort of thing you were saying there at the end. So why should someone who is not a statistician, someone who is not engaged in activist morale criminal justice, you know what is the takeaway for just the layperson, as far as it relates to this particular report? Why should my mother or my best friend or my colleagues here care about what you found in this report?

Shah: Great, so I- that’s a good question, I would maybe think about that in a couple of ways. So one thing that I would think about going back to what I was just saying where often data-informed tools are presented as kind of a more objective alternative to some other procedure and I think it’s important for all of us regardless of where we’re working to be somewhat critical of that because often that’s just a way of kind of sneaking in whatever kind of biases we already had in through this kind of packaging. And I think that applies not just to incarceration decisions but often any kind of automated decision-making systems. I also think just- maybe a lot of us are concerned about policing and incarceration right now so as somebody who is concerned about that especially with pretrial incarcerations and these are people who are presumed to be innocent by the legal system, I think worrying about how those people are treated and how decisions about their liberty is made is an important thing on its own. And in particular, I mentioned that the risk assessment tools are often positioned as neutral or objective alternatives to judge’s decisions which are known to be biased and there is a lot of evidence that they are and so just to understand that we don’t get to wash our hands of the bias just by putting it through these systems. And maybe, hopefully knowing that would lead people to think a little bit more- so I don’t know if this is helpful or not but there’s kind of two ways to look at this, like are the decisions fair in terms of equal treatment under the law for different race groups. Another kind of level of wondering about this is whether- there’s kind of a larger issue here that’s at play and sometimes I feel like these risk assessment tools is to kind of push through an idea that like there’s a fairer way to incarcerate people who are presumed to be innocent and so I hope that p[ that the biases that we see everywhere else in our society also sneaked their way into the model, the data that we used to make kind of databased tools will force people to think about that larger picture and about what it actually means to fairly incarcerate innocent people.

Pennington: And it seems like it’s in line with a lot of research around technology that has shown that we have believed these technologies might, as you suggested, give us the get out of jail free card, to use that unfortunate phrase, right? When it comes to this issue of bias, oh we’ll allow the AI or the technology to handle everything because it’s not biased, and what we’re increasingly coming to realize in whether search engines or video games where you have certain avatars is that the bias is built-in because it’s not been challenged outside of technology. It would seem like your report is kind of suggesting that when it comes to this very important issue of incarceration of people before they are ever actually in court, right before they go to trial, that that bias has also been dealt into that technology. So, it feels like it’s within the framework of our larger understanding of the way AI has maybe not been as critically engaged with other technologies because we see them as these arbitrators of truth in a way that perhaps they’re not.

Shah: Yeah I think that’s right, and we see, I mean the data that we have are generated by existing social processes and the existing social processes that we have are not at all free from racial bias. And so that works its way into the data and like you say into all these different AI systems. And of course, in this case, making very high stakes decisions about people’s freedom. Yeah so I think that’s definitely an issue

Bailer: So, what if I- I’ going to hire you now to build a risk assessment tool for me. So, as you think about this and you know there are cases where we see this- in the banking industry, there are certain variables that are just not acceptable for use in prediction as inputs the models. Is there some sort of parallels to this? I mean if you were thinking about this as saying I would like to try to build as fair a risk assessment tool as possible, what would be some of the steps that you would think about and that you would need to consider in doing so?

Shah: So there are like- there have been people who’ve kind of pointed out or asked that or kind of define fairness as the absence of certain predictors in kind of the way that I think I’m less familiar with the banking industry but I think that they just dot have race as a variable and that’s how they kind of deal with that. I assume this is true in banking, it’s definitely true in criminal justice data that there are lots of things that correlate along with race and so you can’t really remove race from your predictors. You can remove that one column, but you can’t remove it from the zip code or from previous arrest history and stuff like that. So, I would start with that. like it’s not enough to sort of close your eyes and hope that you don’t see race because it’s there and so and then you know I think it would be important to kind of have like we talked about these technical definitions about fairness within the model and the predictions it makes and are they treating people equally across races. As I mentioned there are multiple different definitions and some of them might be mutually incompatible so it’s important if you are going to go down this route to pick one that makes sense for your application. And so I kind of- and I don’t have the absolute answer to this but I kind of mentioned that one way to think about this is how the cost of incorrect predictions is distributed across races and whether that cost and that burden is shared equally or not. And so in the case of- again I don’t want to come down on the side of one version or the other but in the case of these pretrial risk assessment tools that seems to suggest something like false-positive rates as like a more important thing to look at maybe because the- when you’re considered high-risk, that’s when you’re really paying the high cost of the predictions.

Richard Campbell: Some of your findings are really important. Even the phrase risk assessment modeling- I’m imagining- how do you get a journalist interested in that, and what the implications are of that, because I think these kinds of findings are really important for [people to understand, especially now. And have you seen any good coverage about your findings and how do you interest a journalist in this? I mean the way that I would go about it, of course, is to find somebody that was really treated unfairly in the system and tell that story, but then how do you make sure you explain what you found in the data understandable to a journalist who could then communicate that to the public?

Shah: So one thing before I go into my answer, this is like an interesting question because the fields that I’m talking about in evaluating fairness in terms of these technical definitions and so forth I don’t know if this is where I started but definitely my introduction and a lot of peoples introduction to it was actually through journalism it was a story in ProPublica about the compass risk assessment tool. That kind of set off a lot of this study. So, in some sense- like in that case the journalists were ahead of me. I was learning from them. But I think you pointed out the language itself risk assessment either sounds boring or technical but it’s also kind of a useful entry point because I think it’s a term that frames the decisions being made or the scope of what decisions can be made and so when we talk about pretrial decision making if we frame this decision in terms of risk assessment then the person in front of you all they are is a possible risk. And that’s- there’s been some push back against that in various places but I don’t know- so a different way to help think about what I mean when I say risk assessment frames that the decision in a very particular way. So, another alternative might be something like a needs assessment. Like I mentioned before one of the reasons a lot of people miss their court dates is not because they got on a private jet and went to the Bahamas, it’s because they don’t have a home address[ so they’re not able to receive updates about when their court date is. They couldn’t get childcare, or time off work, or they forgot. And so all these things are preventable in all of these other ways but if you only see the decision as a risk you’re only thinking of it in terms of well, regardless of what the reason the person didn’t show up is and that’s the only thing I can worry about. So, kind of opening up that decision to divert- not allowing it to necessarily be framed in terms of risk from the get-go is a useful entry point I think.

Bailer: You know that’s almost impossible, well the question might be short, but the answer might be long. I don’t know the risk of that response. You know I guess one of the things that I have hears, and I’m going to sneak in here anyway and just make Rosemary mad at me because she can’t hit me because we’re doing this all remotely.

Pennington: Right, I would never anyway.

Bailer: No that’s true. So, there’s a theme that you talked about here and that’s the transparency of research that came out. You know the importance of being able to reproduce what was done. And it sounds like you had to do some forensic analysis to figure out what occurred in this previous work. So, what that suggests is that you value this in what you do in your current work. So I was wondering in deference to my friend Rosemary here if you could give us a quick summary of what is the flow that helps ensure in your work that you make it reproducible and that you have accountability baked into it that others could follow.

Shah: Thank you for asking that question and it’s important to lots of people right now, like lots of scientists are worried about reproducibility and I think it’s particularly important to who work in human rights because it’s possible that, well you want to get it right first of all, and also you want it to stand up to scrutiny because you may have results that are not- that doesn’t make people that happy and so you need to be prepared for all kinds of scrutiny and attacks and one way to do that is to be very confident in the work that you’ve done. And in terms of like I found that like having specific ways of working and specific structures for how I manage a project are one of the best ways that I can kind of guarantee those sorts of results and in particular the way we do things at HRDAG is we use this system called principled data processing where we kind of very explicitly set up our pipeline where we work on individual tasks in the pipeline, for instance, importing data from an Excel file or something is one task that may be standardizing code values within the columns is its own task and reproducing a regressions model that we found in a paper is its own task and kind of having- being able to do those things distinctly so you’re not worried about is the entire thing reproducible or correct but like is this thing is this link in the chain strong enough and then moving on to the next piece. That’s helped us a lot and then we kind of manage everything technically with these files so I had to learn a little bit of computer engineering as part of this job but I think starting with that idea of breaking things up into tasks that can be tested individually so that you have a little bit more confidence once the project starts to get bigger and more unwieldy and you’re not worried about the core values when you’re working on the model because solve already kind of tested those in an earlier step.

Bailer: Thank you.

Pennington: Well, that is all the time that we have for this episode. Tarak thank you so much for being here today.

Shah: Thank you.

Pennington: Stats and Stories is a partnership between Miami University’s Departments of Statistics and Media, Journalism and Film, and the American Statistical Association. You can follow us on Twitter, Apple Podcasts, or other places where you can find podcasts. If you’d like to share your thoughts on the program send your emails to statsandstories@miamioh.edu or check us out at statsandstories.net and be sure to listen for future editions of Stats and Stories, where we explore the statistics behind the stories and the stories behind the statistics.


Pets During Quarantine | Stats + Stories Episode 145 by Stats Stories

Allen McConnell is University Distinguished Professor and Chair of the Department of Psychology at Miami University. His research examines how relationships with family and pets affect health and well-being, how people decode others’ nonverbal displays, and how self-nature representations influence pro-environmental action with this work supported over the years by National Institutes of Health (NICHD and NIMH) and National Science Foundation grants.

Read More

Big Data Policing | Stats + Stories Episode 143 by Stats Stories

Sarah Brayne is an Assistant Professor of Sociology at The University of Texas at Austin. In her research, Brayne uses qualitative and quantitative methods to examine the social consequences of data-intensive surveillance practices. Her forthcoming book, Predict and Surveil: Data, Discretion, and the Future of Policing, draws on years of ethnographic research of the Los Angeles Police Department to understand how law enforcement uses predictive analytics and new surveillance technologies. Prior to joining the faulty at UT-Austin, Brayne was a Postdoctoral Researcher at Microsoft Research. She received her Ph.D. in Sociology and Social Policy from Princeton University. Brayne has volunteer-taught college-credit sociology classes in prisons since 2012. In 2017, she founded the Texas Prison Education Initiative.

Read More

The Science of Sex | Stats + Stories Episode 126 by Stats Stories

Debby Herbenick is a sex educator, sex advice columnist, author, research scientist, children's book author, blogger, television personality, professor, and human sexuality expert in the media. Dr. Herbenick is a professor at the Indiana University School of Public Health and was lead investigator of the National Survey of Sexual Health and Behavior.

Read More

From the Royal Statistical Conference | A Stats + Stories Special Episode Pt. 2 by Stats Stories

This episode features a number of interviews from the recent Royal Statistical Society International Conference from last month. Today's guests include, Amy-Jayne McKnight discussing the challenges associated with genomic data. We consider how data is captured, processed, and then analysed and AJ outlines the challenges presented by sources of variability. Then, Lancaster University’s Harry Spearing about the application of extreme value theory to the ranking of Olympic swimmers, as well as ranking in sports more generally.

Read More

From the Royal Statistical Conference | A Stats + Stories Special Episode by Stats Stories

This episode features a number of interviews from the recent Royal Statistical Society International Conference from last month. Today's guests include, Iain Flint of G’s Growers talking about the IceCAM project, which helps to minimize food waste by adapting the growing programs of iceberg lettuces according to weather predictions. We also have James Tucker, head of the Quality Centre and Methodology Advisory Service at the Office for National Statistics talking about respondent confidentiality, and data privacy and protection. As well as, Kevin Johanson from the Expert Group on Sámi Statistics based in Norway, on how the group is working on developing statistics on the Sámi people and how these statistics can lead to better policymaking.

Read More

Who’s Behind All These Gig Economy Jobs? | Stats + Stories Episode 110 by Stats Stories

Siddharth Suri is a computational social scientist whose research interests lie at the intersection of computer science, behavioral economics, and crowdsourcing. His current work centers around the crowd workers who power many modern apps, websites, and artificial intelligence (AI) systems. This work culminated in a book he coauthored with Mary L. Gray titled Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass (May 2019).

Read More

Is Wearable Tech Worth It? | Stats and Stories at JSM by Stats Stories

Dr. Jane Paik Kim is Clinical Assistant Professor in the Department of Psychiatry and Behavioral Sciences at Stanford University School of Medicine. Her professional aim is to improve public mental health through the application and development of statistical methods in mental health research. Her research interests are statistical methods for digital health interventions delivered through mobile or wearable devices, and psychiatric ethics research. Her statistical interest areas are in the robustness of regression-based inference for both clinical trials and observational studies, as well as methods development for survival data arising from non-standard biased sampling schemes.

Read More

Making Forensic Science Scientific | Stats + Stories Episode 91 by Stats Stories

Dr. Alicia Carriquiry is a Distinguished Professor of Liberal Arts and Sciences and a Professor of Statistics at Iowa State University. She serves as Director and lead investigator for the Center for Statistics and Applications in Forensic Evidence. The NIST Center of Excellence’s mission is to increase the scientific rigor of forensic science through improved statistical applications. Dr. Carriquiry provides scientific oversight and research expertise to the center. She participates in the Organization of Scientific Area Committees subcommittee on Materials and Trace Evidence and serves as a technical advisor for the Association of Firearms and Tool Mark Examiners. Dr. Carriquiry was recently named to the National Academy of Medicine and elected as a fellow to the American Associations for the Advancement of Science.

Read More

Understanding Conflict Resolution | Stats + Stories Episode 89 by Stats Stories

Dr. Sara Cobb has a Ph.D. in Communication (UMASS Amherst) and is the Drucie French Cumbie Chair at the School for Conflict Analysis and Resolution (S-CAR) at George Mason University, where she was, from 2001-2009, the dean/director. In her current role as faculty she teaches and conducts research on the relationship between narrative and conflict. She is also the Director of the Center for the Study of Narrative and Conflict Resolution at S-CAR, which provides a hub for scholarship on narrative approaches to conflict analysis and resolution. She is co-editor of the journal Narrative and Conflict: Explorations in Theory and Practice.

Read More

The U.N. and Statistics | Stats + Stories Episode 86 by Stats Stories

Stefan Schweinfest was appointed Director of the Statistics Division (UNSD/DESA) in July 2014. Under his leadership, the Division compiles and disseminates global statistical information, develops standards and norms for statistical activities including the integration of geospatial, statistical and other information, and supports countries' efforts to strengthen their national statistical and geospatial systems.

Read More

Using Data to Protect Human Rights | Stats + Stories Episode 74 by Stats Stories

Megan Price is the executive director of the Human Rights Data Analysis Group (HRDAG), and designs strategies and methods for statistical analysis of human rights data for projects in a variety of locations including Guatemala, Colombia, and Syria. She has contributed analyses submitted as evidence in two court cases in Guatemala and has served as the lead statistician and author on three UN reports documenting deaths in Syria. Megan is a member of the Technical Advisory Board for the Office of the Prosecutor at the International Criminal Court, on the Board of Directors for Tor, and a Research Fellow at the Carnegie Mellon University Center for Human Rights Science. She is the Human Rights Editor for the Statistical Journal of the International Association for Official Statistics (IAOS) and on the editorial board of Significance Magazine. Before she was executive director at HRDAG, Megan was the director of research there.

Read More

Holding Up A Mirror To Society - A Tale Of Official Statistics In Greece | Stats + Stories Episode 60 by Stats Stories

Andreas V. Georgiou is an economist with specializations in Monetary Theory and Stabilization Policy and in International Trade and Finance. After working for the International Monetary Fund, he returned to Greece in 2010 to head the newly established Hellenic Statistical Authority (ELSTAT)-the successor of the National Statistical Service of Greece following the onset of the economic crisis in Greece. He was President of the Hellenic Statistical Authority for 5 years. He worked to re-organize and rebuild the institution, on a new basis of fully conforming to international and European statistical standards and practices, leading to the establishment of the credibility of Greek statistics. He is an elected Member of the International Statistical Institute . He has been a Visiting Associate Professor in Finance, Banking and Investment, at the Economics University Bratislava, Slovak Republic, and a Visiting Lecturer at Amherst College. He is currently a Visiting Scholar at Amherst College.

Read More

Chins And Ears Are Not Information Rich - Awkwardness And Social Relationships | Stats + Short Stories Episode 57 by Stats Stories

Ty Tashiro (@tytashiro) is an author and relationship expert. He wrote Awkward: The Science of Why We're Socially Awkward and Why That's Awesome and The Science of Happily Ever After . His work has been featured at the New York Times, Time.com, TheAtlantic.com, NPR, Sirius XM Stars radio, and VICE. He received his Ph.D. in Psychology from the University of Minnesota, has been an award-winning professor at the University of Maryland and University of Colorado, and has addressed TED@NYC, Harvard Business School, MIT's Media Lab, and the American Psychological Association.

Read More

Are Communities Helped By Terraforming Food Deserts? | Stats + Stories Episode 56 by Stats Stories

Bonnie Ghosh-Dastidar is a senior statistician and Director of the Statistics Advisory Center at RAND Corporation. Her current research interests center on leveraging natural experiment designs to estimate the effects of neighborhood-level 'interventions' or changes on residents' health behaviors (e.g. diet, physical activity) and outcomes (e.g. obesity). Her statistical expertise is in the areas of study design, survey methods, non-response, and analysis of longitudinal and multilevel data.

Read More