Fairness in decision-making with AI: a practical guide & hands-on tutorial using Aequitas
Recent work has raised concerns on the risk of unintended bias in AI systems being used nowadays that can affect individuals unfairly based on race, gender or religion, among other possible characteristics. While a lot of bias metrics and fairness definitions have been proposed in recent years, there is no consensus on which metric/definition should be used and there are very few available resources to operationalize them. Therefore, despite recent awareness, auditing for bias and fairness when developing and deploying AI systems is not yet a standard practice. In this tutorial, we present Aequitas(http://github.com/dssg/aequitas), an open source bias and fairness audit toolkit that is an intuitive and easy to use addition to the machine learning workflow, enabling users to seamlessly test models for several bias and fairness metrics in relation to multiple population sub-groups. Aequitas facilitates informed and equitable decisions around developing and deploying algorithmic decision making systems for both data scientists, machine learning researchers and policymakers.
In this tutorial we will cover the following how tos: How to think about fairness and equity when building and evaluating AI systems; How to define fairness goals and manage efficiency and effectiveness tradeoffs; How to select fairness metrics; How to build and select ML models that achieve those fairness goals; How to validate that the AI system is fair;How to monitor a deployed AI system for fairness and adapt if necessary;
This session took place in track PyData and was classified suitable for some domain / basic python by the speaker.
Transcript (auto)
Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.
Speaker 1 [00:04]
In this tutorial, we are going to talk about bias and fairness, how to think about these issues when you are developing your models, when you are deploying your models, and this is really aimed to be a practical kind of presentation, because I often tend to feel like the general presentations and the most common presentations about bias and fairness in AI are very fluff. They really don't discuss really practical bits and insights. And this is really about this really like more practical discussions. And when I submitted this tutorial back in the proposal back in May, I had this part of end zone tutorial and after giving this talk a few times I really realized based on feedback from the audience is like the tool is very straightforward and easy to use what is really hard is to understand and reasoning about the results that the tool gives you so there is a section kind of in the middle of the tutorial where we go actually through a jupyter notebook with some code and examples but i will expect this to be really straightforward for you and you can try out that the tool um afterward and you can once again just message me and i will help you with that uh but um this is really about discuss discussing what is fairness, when should we care about the bias, how algorithmic decision-making works, and some kind of case studies where I actually show real results of fairness audits. So, just the typical about me. When I submitted the tutorial, I was in Chicago doing my postdoc at the Center for Data Science and Public Policy. Also, together with the Data Science for Social Good Summer Fellowship, I was a technical mentor last year edition in 2018. But since then, I moved to Lisbon, and I'm now in the industry, but I'm still working on these issues. I'm creating a new research group. The typical name is either FAT or FATE. I always like it to be faith, so it's fairness, accountability, transparency and ethics in AI and just to give you a little bit context about the work that I developed in Chicago where the in which the Akitas toolkit is part of this work so I'm one of the main contributors of the toolkit and the Center for Data Science and Public Policy And there is a science for social good. They really aim to tackle real-world problems with data and machine learning in the intersection of public policy, social good kind of projects and problems from criminal justice, workforce development, education, healthcare. these are some of the partners that they are the ones that provide us data you can think about this center as kind of a mini consulting group within the university but for helping non-profit and government agencies and you can see here the White House but just a disclaimer this is the White House before 2016 and And we have interesting projects that received really immediate attention, like the police projects. I don't know if you recall, but there was some wave of adverse incidents between the police and the African-American community in the U.S. And what these projects aim to solve was to predict the police officers that were at risk because of stress levels, because of working too many hours, of having an adverse incident, so we can adopt a proactive approach of redirecting these officers to psychological and counseling therapy and this kind of a group workshops. And this is just one example. About the industry. So just a small review. Fidzai is a startup company that basically works for fintech and online payments to detect fraud and anti-money laundering. So in the private sector, when I was deciding to move to industry, I was considering either healthcare or fintech because there are really areas similar with this kind of predictive and risk assessment for decision making. And here we really aim to catch fraudsters while not hurting or not allowing people to make online payments if they are not really fraudsters. and that's one of the issues that we want to tackle and to make sure that we have fair decision-making. So, as I said before, we are really, in the first part, discussing algorithmic decision-making, sources of bias, the fairness definitions, common misconceptions in this area. In the part two of the tutorial, we'll go and deep dive a little bit about the toolkit and actual results and some additional considerations. So the idea is that at the end of this tutorial, when you are in your company or in your organization, and you are developing models for decision-making, you can really reason about what you should care, what are the consequences of the models that you are developing, who gets impacted, how to get data and collect these results to give the decision-makers or the product managers or the leadership on your organizations to make more informed decisions. Actually, I want to ask you, how many of you are data scientists, data engineers? Okay. So, it's really not about you to make the decisions of what is fair or not fair, what should get deployed or not. It's really about us data scientists to really understand what's going on, collect data about what's going on, and give the decision-makers a systematic menu of options. Because in this area, often people don't really agree. And what we really want to avoid is to have this kind of subjective decision-making where if you are working with a given product manager or leadership, you make a given decision. If you are working with other middle management, you have a different kind of decision. It's really about creating this systematic kind of analysis about what's going on and then someone will make the decisions in our projects in chicago and with our partners it was actually the the organization that would take the model that was really defining what they care about what would be what should be the impact if they care about hurting specific groups or not not as a data scientist to really define that so um from you uh and most of you are data scientists um how many of you have worked in decision making and not um text or image classification but rather like classification for scoring and really uh okay good so this initial part might be trivial for you but because of requirements uh or like entry level and let's go a little bit dive on this so some examples of algorithmic decision making so fraud prevention what you are really like deciding and predicting is are you getting a charge back from this transaction was this credit card stolen or not you're making a call on credit lending is the classic example of decision making insurance policy surprising nowadays is also based on decision-making of the risk of your behaviors. Screening job candidates, and we can discuss some examples on this. Oh, by the way, there will be some question Q&A breakpoints throughout the presentation, so not just at the end. Preventive healthcare, it's really interesting because now we are shifting from reactive, so people get a disease and you want to get treatments, to kind of proactive and preventive healthcare. You are actually predicting the risk of someone developing a disease and based on these risk scores you really try to avoid that most these people forgetting that and getting a better outcome. Also in the criminal justice and in the police there has been especially in US the adoption of this kind of algorithmic decision-making. So just to give you a little bit more of a context, this is an example of a decision-making project I worked on at the University of Chicago, and our partner was Johnson County, Kansas, and this is kind of project where you really, really face risks of being unethical, and it's really, really important to base on how you scope the projects, how you define interventions and who is going to take interventions then the same machine learning model can be used for good or for evil and here is an example of this is actually a timeline of a real person that was living in Johnson County Kansas and this was the first project where all the public data from public services in the county and got together and linked and we were able to really trace all the interactions all the touch points of the individuals with emergency medical services 911 bookings from the police all interactions with the police officer get some kind of log we had access to them mental health consultation and services public health consultation and the goal here is to tackle the problem of recidivism in people with mental health issues so in the u.s. access to mental health to people that they that don't really have health insurance is really tough so if someone gets mental health breakdown and they punch someone on the streets. They are running naked on the streets. They, even in the case, they call 911 and they say they want to hurt themselves. Police goes along and goes to the site. And often, what they do is they arrest the people. And the people get a jail booking. And, as is kind of obvious, jail is not the right place for people with facing a mental health breakdown. And Johnson County in Kansas is really progressive on that and they developed this co-responder program where a mental health provider, an expert, would go along with the police for calls with specific codes like substance abuse, domestic battery suicide calls and so on and they tried to stabilize the person and, if possible, to divert them to the hospital and not to the jail. But this is reactive. And once again, I'm repeating myself, it is proactive versus reactive, but you'll see what I'm getting into with this decision-making and predictive kind of setup. With this project, our goal was to base on people. So our cohort was people with mental health, slightly indications of mental health issues since ever in our data that got a jail booking in the last three years to predict the risk of them getting a jail booking in the next year. So we can rank people and we can offer access to mental health services. So mental health providers would call people, would go to people's houses and realize people's lives are going, if they need some help, and if they are receptive to the intervention, they would start enjoying some kind of preventive program. And why there are unethical or ethical risks in this project? Because if we get this prediction to the police, they could actually arrest people before, like kind of a minority report kind of setting. But thankfully, the good news is that the intervention is a mental health intervention, not a police intervention. So the only partner, the only organization that got access to the full integrated data and the prediction was the mental health service, and not the police, not the courtrooms, just the mental health providers. And this is just to get a little bit on the context on when we move on to bias and fairness, the implications of deploying specific models versus others. So, in kind of a nutshell, algorithmic decision-making can be seen as statistical risk assessment. So, you want to predict the likelihood of an outcome. And this outcome can be this patient developing a given disease, this credit card being stolen, the likelihood of the job candidate being successful in the job. And usually this is solved using training a binary classifier on historical cases. And contrary to image classification, multi-class classification, where we usually take the class that gets the highest score, here in the binary classification, we just get one score. And in order to really have control on what you are deploying, we don't use the typical predict on 0.5 thresholding. We use the actual predict proba or the scores. If it's a probabilistic model, the probability score. If it's like a tree-based model, just we rank the examples by the instance by score. So here is an example, a fake example. If we have a sample with 20 instances and we already train our model and we get the scores and we have the true label. So you can think about this as a validation set or test set. And here we have a total of six positive examples and 14 negative examples. So our prevalence, our base rate, our prior, or I'll usually call the ratio between positive examples and the total examples is 30%. And let's assume that the model is the same, the scores are the same, the ranks are the same. We are just now defining a way of testing the model. And the way we define how we want to test the model can either be based on a threshold or based on a top K capacity. If you think about job candidates, you receive thousands of resumes and you want to rank the top 100 candidates that can proceed on the next steps of the interviewing process. In these situations, these type of predictive problems, you have a capacity. And you define K, and you always take the top K examples. If you think about event-based decision-making, like fraud, or if you have a specific event, like interaction with the police, or you go and see the doctor, then we need to define a threshold on our validation set, and we deploy this threshold on our test set in production. So, if the score gets above a given threshold, then we consider this as positive. In this example, we have defined our threshold as 0.98, or our k is 4, and we got right about all the predictions. So, we have four 1s in the label column, we have zero false positives, we have two false negatives, though. there were two positive examples instances that didn't get classified as positive and we have 14 true negatives so I don't know if you are familiar with the ROC curve if you go to the Wikipedia page of the ROC there is this table there that strangely is not present in any other web page on article on on Wikipedia but this is really really useful because it really defines the confusion matrix and all these supervised error-based metrics that we typically calculate so it's in order to really understand the results that the Akitas toolkit gives you, you really, if you don't know these rates, all these rates, it's good if you have on the side this table and can help you really understand what you get. So, as we were there, we defined a threshold on top k equal 4 and now we can compute these metrics. And the false positive rate, we don't have any false positives, it's zero. The recall, we caught four out of the six positive instances, so it's two-thirds. The false negative rates is like we have two, we missed two out of six positives, so it's one-third. And our precision, because based on the predicted positive set, we got all correct so our precision is one congratulations but the thing is maybe you have capacity to interview more than four candidates so maybe you can just go a little bit down on the list to consider the predicted positives and if we now have a threshold of 0.92 or K equal then the story is a little bit different, but once again, it's exactly the same model and the same predictions, but we get very different results. So, now we have six true positives, four false positives, and no false negative. So, our recall is one. As we can see, we didn't miss any positive example. Our false positive rate, though, is not 0 anymore, because we have 1, 2, 3, 4 false positives. So we get 4 out of 14 negatives. So it's 0.29. Our false negative rate, we are not missing anyone. So it's 0. And the precision is 6 out of 10, because out of the 10 positives, we miss 1, 2, 3, 4. So we just got six true positives. And about the ROC space, it's really important to understand how this works. So for the same model in example one, we are operating in this point. So we got zero false positive rate and we got two thirds on recall. For the second example, when you go to 10, And we already got the max recall, and our false positive rate is 0.29. If we keep going down the list until we consider everything as positive, we will keep along this axis, because we already caught all the positive examples, so recall cannot be higher than 1, and the false positive rate will go until the max, because you are considering all the false positives as positives now. And it's also really important to understand the separation line here. So in this line, we get recall equal to false positive rate. And why this is important? Because if we define recall in a probabilistic way, it's given that we have a set of examples that the label is positive, what is the probability of classifying a given instance of the positive examples as positive. And the false positive rate is the same thing but now conditioned on the negative examples. So, based on not being a positive label, what is the probability of misclassifying the non-positive as positives? And as you can see, when these two probabilities are equal, basically we are getting probabilities independent of the label. That's why we say this is random, because it doesn't really matter the label, the true outcome of the examples, the probability of predicting as positive is the same regardless of the label. And this is also very important because often people, and this is a misconception, and it's not like because you didn't figure out or it's just because based on practice, often in text classification or in image classification we have a balance data set for training and the base rate the prevalence is 0.5 so we often say it's like flipping the coin to get a random prediction but when we don't have balanced classes as it happens in decision-making because what is common decision-making is to have fewer positive examples what really means to be random is if your recall is equal to false positive rate for any threshold that you define. So it can be a random for specific threshold on 5, on 10, on 50. You get always on this line here if the classifier is really random on the positive predictions. Thankfully this example it's far better than random. Any question so far? No? Good. But not all instances are the same. So now assume that you have an attribute like skin color and this analysis is independent of you using this attribute for the model or not. And we will get there if you should use these attributes on your models. This is now just doing the exactly same analysis, but now you have two subgroups, white versus non-white. And you can count, do the same exercise for non-white. How many level positives? We have three, so one, two, three. How many label negatives, we have 6, which is 1, 2, 3, 4, 5, sorry, how many label negatives, the zeros, for, I mean non-white, sorry, 1, 2, 3, 4, 5, 6, and the group size is 9, and the prevalence of this group is one-third. For the whites, we do the same thing. We also have three level positives, so one, two, three. Level negatives, we have eight, one, two, and then four, six, eight. Group size is 11, and the prevalence is 0.27. And this is really important to keep in mind that we actually have more examples from a given group than the other, and the prevalence is higher on the smaller group though. If we use the same threshold for making positive decisions, now we can calculate the same metrics again, and now you see that the recall is the same, the precision is the same, the false negative rate is the same, also because the false negative rate is one minus recall, and the false positive rate though is different is 0.25 on white the larger group and with lower prevalence and is 0.33 on non-white so if we think about the probabilities so given that you are from the group non-white on skin color and the label is zero the probability of make of having a false positive is higher for non-white than white. So now, any question about this subgroup decomposition? No, straightforward. So what about this question? Is this previous difference on false positive rate between white and non-white a problem? Anyone can chance an answer? Take the risk. Okay. Um, actually, yes. that's also a very good point but let's assume this is an example but this is a very good point and we'll get there as well um so it really depends so as you were saying wrongly accused it really depends on the intervention and the context if you're providing help and benefits if you don't need the benefit and you get the benefit that's really there is a cost opportunity because often we have limited budget and we are spending money on people who actually don't need help for a specific healthcare or education. But if you think about the people that get interventions, it's really not really bad to get an intervention if you don't need it. But if intervention is punitive like being accused in criminal justice or something like that, then the story is very different and we get there so decision making is about predictions and often there will i can say this like unless you have a really limited capacity and you are operating really on the top there will always be errors and bias is about in this context this is not about statistical bias when you are learning and machine learning if you're background in machine learning And bias in this context is used to define if there are disparate errors across different specific subgroups. And decision-making has been around for thousands of years. And this is often interesting also in the generic and common talks about fairness. it seems like, oh, we are smart and we have all this knowledge and power and we know how to use these tools and we are building deep learning models and we are changing the world, we are so powerful. We need to be really cautious about the impact of what we are doing. But bias in decision-making has been around for thousands of years. Like if you read Plato and the Republic and you think about ethics, all the policy decisions have been around and every policy decision will get some people that will get more benefit or some people will get more help so it's really about the scale we are now in a in a time of auto machine learning auto ml metal learning kubernetes deployment everything is getting automated and because we often just care about the overall performance metric that our client defined or that we define that we want to solve an issue and we define our global metric as precision or recall or f1 or whatever by the way aoc doesn't make any sense in decision making because it's not the same to operating on top for top 10. So AUC gives you the kind of a proxy of the overall performance along all these thresholds. But in practice incision making, you operationalize in a specific threshold. So it's really about precision recall and not so much about AUC. So you often maximize your performance for this global metric. And you have some heuristics or something or some rankers. And let's see, even if we have some kind of a fancy hyperparameter optimization in place, we just pick the models that optimize our global performance. So the best gets selected. And often we don't understand that two models with same performance, global performance, might have very different differences in disparate errors across these groups. So when we are scaling and automating the decision-making process, we are not really understanding the impact of the models in the system. And often, if we have automatic decision-making as well, if we don't have any human reviewing and getting the prediction, we are creating feedback loops in which we deny bail to someone and people stay in jail and we don't actually get the true label if the people be released do they recidivate if we allow this transaction to go would it get like charged back was it really fraud we are automatically declined the problem of rejecting for instance so we are creating all these feedback loops and all this automation and so on and we really don't understand the impact and i really talk with people from software engineers, data science, to product managers, and when I really come to, do you really know what is the impact in production of your systems? People often divert and change subject or talk about other things, because maybe I'm getting too boring. But the thing is really about understanding the impact. So, there are many sources of bias, and the first are very obvious, like the world and the people is really biased. We are really biased, that's the way it is, and values are very different in Germany, in US, in Chile, in China, in Japan, in Fiji. Values really are very tied with culture and societies in general. So, in different regions of the globe, bias will be different. And about the data. When we are developing these models, we need historical data. And just to recap, there are two types of bias in the data, the sample and the label. Let's just focus first on the sample. often we are developing our machine learning models in a sample of the population that is different from the target population that will get the decisions so imagine that there are new types of candidates, new types of jobs that the companies create and maybe 10 years ago there was no product manager or maybe there was but there was no machine learning engineer or there was no data engineer and now we have so but we have all these models um the models created on this is all this historical data and um this is really one source of bias that and often this is also very dependent on time so because the behavior is shifting and we have all these feedback loops if there is already like a machine learning model in production the the sample is really is really a problem here and people in social sciences and people in the surveys, analyses, they have dealt with a sample bias in a long time. They try to extrapolate, given on a sample, the political polls for all the country for uses. But in our case, machine learning, we really need to use all the data, as much data as possible to really capture the signal. So this is one of the sources of bias that we have. Another one is the label. And based on my experience, this is the most troublesome source of bias besides world and people. But I think it's too early for us to try to change the world and people for now. But the labels are really a problem because we often take for granted that the label is really true. It's, I got this data set and the label is one. So are you really sure that the label should be one? is there really no noise in your labels and if there is noise is it really random or is it specific for some subgroups there is all this area of counterfactual fairness and I really advise you to search and read some papers about it but it's really interesting to think about causal kind of methods to try to figure out what if the label was different would the model really learn a different decision boundary or not. And also, when we think about healthcare job candidates, and we were saying like, let's predict the likelihood of the candidate being successful. How do we measure success? It's really, once again, often there is a heuristic. Oh, if he got promoted in the first two years of the job, and are you sure that being promoted doesn't introduce any noise. It's not conditioned on the managers in the relationship. It's not conditioned on gender, ethnicity, background, education, political views. Are you really sure about that? So this is really, really an issue, and we should all think about it. An example in the U.S. is the police investigations. often really have some kind of political drive on that. So label is really one of the issues. And you can even think of, for instance, in fraud prevention, we often get the analyst to make the last call. So it's really the analyst figuring out if it's fraud or not. And it's really, really hard to make a call for that. And then we came here. We come here. I'll get some more. If you guys want some more, you can get. Every step of the machine learning pipeline, in every step we are making a decision. Which, how we get data, how we aggregate data, how we store and link data. You recall the example of the project of Johnson County in Kansas where we are linking data from different places. In US, there is now actually national ID card, so we use social security number, first name, last name. But often, data from interaction with the police and 911, you get errors and typos and so on. So record linkage is a huge source of bias because you are introducing bias in two dimensions. So you are interested in bias on the rows because you are propagating the population, the subgroup size because you have duplicates and the real size of the group is not the one that you get because instead of being one person, you are counting as two. And also on the features, on the attributes level, there is also bias because before there was some behavior that you were capturing that was divided between two people. but the behavior should be aggregated in just one person. So we are actually introducing bias as well on the feature level. The models you choose, hyperparameters, the metric that you define to test, just that can also introduce a use of bias because you are often selecting X number of models based on that metric. The way you deploy, how the intervention is going to be used, The way you teach people to use the model, the way you maintain, are you often really retraining or not? Do you really care if the data, if there is any kind of concept drift? The way you communicate to the people how to use interventions or not? Everything is introducing a specific bias. And at the end, there is also the action or intervention bias. Because the intervention might be mental health outreach, as I was saying, or specific nutrition consultation, or some training, a workout, whatever. And not all the groups will be prone to accept the intervention the same way, and not all the groups in the intervention will be as effective. If we think about disease and you have just one intervention, it might have different effects on older people and younger people. It might have a different effect on male versus female. So if you just define one set of intervention or even a couple of interventions, the way the heterogeneity effect might exist. And also there is a human that can override the model recommendation, and that is also a source of bias. So, as I was saying in the beginning, often we got into these discussions of what it means to be fair. And as you know, we are really responsible for data-driven decision-making, data-driven analysis we should not get caught on this trap so there is really no universally accepted definition of what it means to be fair and it really means on the intervention and I'm going to give an example actually I have two examples but I'm going to skip the second one and is rather long so let's think about people got the jail booking and they go to pretrial and we have a predictive risk assessment of them being of high risk of recidivating, of getting a jail booking next year. And often that's the proxy problem that we try to fit in order to decide the bail amount that someone should get or even don't get a bail at all and just keep in jail. And let's consider this example, just release or not giving a bail to someone. so different people might consider it fair if the model makes mistakes about denying bail to an equal number of white and non-white individuals, so absolute numbers regardless of how many people got arrested, regardless of how many people there is in the city in the states, in the country in the world some people might say we should make errors in equal numbers to the two groups this is one definition and here c is a constant so what really matters is that regardless of the group you get the same contest yes this is just one definition another definition would be that the chances that a given white or non-white person will be wrongly denied bailey is equal regardless of the race this is not looking at the labels. This is just looking to the attributes values. So, we have two groups. The probability of making false positives is the same for two groups. And this can be defined based on the population size in the city, for instance. And this is one definition. Another one is the one minus precision. So, this is often thought as kind of a score calibration. So, it's among the jail population so the people that got positive predictions so you remember the threshold the ones that get above the threshold among all that population the probability of having been wrongly denied so the probability of having false positives condition on being denied bail is independent of the of the skin color so this is the equal false discovery rate and once again this is equal precision. Because false discovery rate is 1 minus precision. It's from the predicted positives, how many false positives you get. And, personally, the ones I find more interesting is reasoning about this one. So, among the decisions that you have of keeping people in jail, how many you were wrong for the two groups. To this one, that is conditioned on the label. So, given that you actually are innocent, and you should not be in jail, and you are from a specific group, what is the probability of innocent people from the white group or innocent people from the non-white group to get a false positive and keep arrested? And, as I said, there are very different definitions. There are more, like demographic parity and I will go over that later on, but just you to really be mindful that we are in the same set of context of punitive examples. We already constrain ourselves to just punitive interventions, and we have at least four different, and even if you don't agree with even one, let's say like two different metrics that are often competing against each other, you often don't, you cannot guarantee both at the same time sometimes. Any question? No. So, I also have the same example but now for the assistive, and it's exactly the same reasoning, but the metric is on the other side, and I will really go… this is more for archive and you get on the slides, so it's the same thing but now thinking about false negatives, thinking about the probability of missing people based on the group size, probability of missing people based on the negative predictions, so people below the threshold, what is the probability of being from one group or another and the label being positive so you're not getting an intervention you were considered not a good candidate for the job and the false negative rate is condition on you actually being a good candidate for the job but you didn't get selected what is the probability once again same kind of reasoning many different metrics often the way we are measuring fairness is based on parity. So, we compare one of these metrics that I just mentioned between two groups. So, we compare the false positive rate of group A and group B and we see the ratio. And this is really an interesting metric or way of calculating because it allows you to communicate like simpler. So, if you you tell to a non-technical person that you are making three times more mistakes for innocent people that are non-white than white is easier to communicate than your false positive rate is 0.85 versus 0.77. This is really important. No, just communicate in a clear way. This is how parity measures are useful. And fairness is normative. So it's about taking these disparities and really define boundaries. And this is of an arbitrary boundaries. The machine learning community is adopting these kind of metrics from the criminal justice cases on fair lending, equal opportunity act, where there were some historical cases where the jury decided that there was discrimination if the benefit was lower than 80% of the majority group. And this is really arbitrary. They define 80%. If it's lower than 80%, it's fine. I don't know if it's fine or not, but this is just for you to get a notion that when you are creating these kind of systems for analysis to define and you can go and see. And often it's very hard to get really equal numbers, also because of statistical significance, but it's important to keep in mind that these thresholds are rather arbitrary. And here is a summary that we created to help our non-technical partners to reasoning about all these metrics. So the way to navigate this is asking the question, do you want to be fair based on disparate representation or disparate errors of your system? So you care about false positives or false negatives, or you don't care, you just care about who gets the intervention and the action, regardless of the labels. And if you don't care about labels and errors, then there are many possibilities. Two examples here are just equal numbers. So you select the same numbers for a group and for another, for the intervention. Intervention meaning getting accepted in the university or passing to the next round of interviews in the job market. Or demographic parity. Demographic parity here means that based on some definition of a population, and the population can be, once again, the city, the country, the state, if there are 60% from group A and 40% of group B, If you pick for intervention, 60% of interventions are for group A and 40% are for group B. And this is demographic parity. If you care about errors, then it becomes a really tricky question. Do you really trust the labels? If you don't trust the labels, just try counterfactual fairness and see if you get interesting results. and often you get exactly the same vector input vector and the label is different and those are kind of obvious candidates for being suspicious but often you can think about what if there is an attribute that is not correlated in a specific group to the outcome and the other group is correlated and let's just skip that attribute, like h, and see if the input vectors are the same besides these attributes. And then, is the label really different? Is just the attribute that is changing the label? And then you can figure out those kind of counterfactual fairness approaches. But let's assume you trust the labels. And then is really the question, are your interventions punitive or assistive? If they are punitive, this is just a list of metrics I told before. If they are assistive, these are the other metrics. And this is really about narrowing down the scope. So I really strongly advise to not look to these metrics if you are providing punitive intervention, because this will just distract you from what is really making a difference on your models and the impact is really like the false positives. The false negatives doesn't really, people don't get hurt because they got released. But there are some tricky situations, which is the intervention, the people affected by intervention might have two distinct groups. So if you think about child welfare cases, I've run some audits on models like this like using social workers to inspect homes of people that are at high risk of having some kind of adverse incident with children like abuse or mistreatment of children and then the risk of being a false positive is really high for the families because you actually you are a good parent and social services are coming to our house and trying to inspect your home and how you behave with your children. So it's really a punitive intervention for the parents. But if you are a child and you are suffering some kind of abuse, then you really care about false negatives. So you are really missing children that are at risk and they are not being considered by the law. So often, it really depends on the context and the application, but if you have these different groups that get affected from the same decision, then you should care about the two sides of the of the metrics, but if not, I advise to just focus on one leaf and not the other. Any questions? Anything about the society? Yeah, it really depends on the context. If we can think about, first, healthcare. If the treatment has some kind of side effects, then you have both sides. But if the treatment is just a nutrition consultation or active lifestyle, then it's good for everyone. And there is really no risk on being punitive, on really redirecting people to good practices. but this is really about the menu of options so at the end, as I was saying, decision making has been around for many many years so at the end it's a policy decision and the people, the companies, the organizations that are carrying out interventions, they should have a clear picture of what are the target population they really care about and once again, it's not us that we should make the decision. But if we think about the fraud prevention cases, for instance, you might think at first that oh, I just care about false positives because my transaction was declined and my credit card was not stolen. But if you think about protecting people that get the card stolen and you should block it as soon as possible, then you really care about false negatives. So, yeah, it's really about the menu and now, for instance, banks care about which one of the banks should care about to protect customers that get credit cards stolen or to not create friction on transactions and not deny transactions that are okay, yeah. No, to explain it is not to make the decision. Yes, but how do you learn in practice? How do you make it transparent to them that they think it's not? That's actually, we didn't arrange this question, but this is a really good question. The question about how do we inform people and how we get this more transparent is about using these toolkits like Akitas and giving training like this one and creating reports where you really inform people about what's going on and how that can affect different people. Helping people define who is affected by the intervention, what are the specific groups you care about. And it's really about making these kind of processes and practices that will help people make more information. So this is really the goal of the tutorial is that you go out and you follow up on these and you may really inform people about what's going on. Just a quick overview about, this is more on the research side. So it's like the main research problems in this area is detection. How can we measure? And this is about definitions and so on and so forth. Like, I don't know if you're aware, like this very famous paper of equality of opportunity that is about parity and recall. And there are many papers on definitions of sufficiency and so on. Another issue is tracing the sources of bias. So really go back and see, was this decision about imputation, about record linkage that really introduced the bias? How do we really isolate what are really sources of the bias? And the third problem is mitigation. How can we fix and try to mitigate the risks? And regarding the research panorama, there are lots of papers on detection and mitigation. It's really common. Machine learning, computer science people, they really love optimization problems. So they define fairness as an optimization problem. They add some regularization parameter and voila. What happens is that in real life, in practice, you often cannot play around with thresholds. You cannot play around with capacity. You have fixed budgets. You cannot increase the positive cases. It really depends. And another really, really issue is, in these examples, because it's easy to communicate, we are talking about binary groups. So, white versus non-white, male versus female. But in practice, I've dealt with the problems in the US where the ethnicity column has a cardinality of 13 or 15. How can we accumulate all these specific groups together? And the research work is really theoretical on that. And we started 10 minutes later, but I think I should go faster. Famous misconceptions is that, shall I use race or gender in my models? If the model doesn't see it, how can it be fair bias based on that? Well, this is the most common misconception in this area. Because often other features of some of those features, those attributes. Like in the US, socioeconomic background, education, the age of the first jail booking is highly correlated with your ethnicity. really really high correlated so it's really easy to infer one based on the others same thing with browsing history or instagram likes it's somehow easier to detect if you are men or not based on your interests so you really will not fix it by just remove it and in same and at the same time, there might be the other way around. That attribute might have some features that you actually are not able to measure, but you can use that attribute as a correlation of another signals that are really important. So if this attribute is correlated with the label, you should include that attribute in your models because it will help you make overall better decisions. But then it's really about you then defining how to operationalize and how to compare different models. Another one is like, if you just follow demographic parity, you get fair outcomes. This is not true because this is not conditioned on label. So actually what you are doing is like you're pushing to the top examples that are negative. The label is zero and you are pushing to the top. And then basically you are crippling your overall performance. This is not about forgetting the overall performance. No, overall performance is still the main goal. You actually want to help people or you want to get good interventions. But then the secondary goal would be I really care about the biases and the fairness of my models. This is like a corollary. There's lots of papers about it. It's really easy to get to these just using definitions of false positive rate to negative rate. You get to know that if the prevalence of the two groups are different, then it's impossible to have FPR and recall and precision parity at the same time and this is like a corollary you can research on this so what do we know as a mitigation plan nowadays you have two options if you don't agree with the results or you need to improve the results, you go back you try different approaches in your pipeline you use one of these methods of fairness based optimization you try to use it for a specific problem and you see if the results improve or you get budget and you change the way you deploy the models and you define different thresholds and maybe if you get more capacity to instead of interviewing other candidates for the job if you can interview 200 maybe you get equal error rates across different groups and this is really like at the end you see even to solve these issues money might might be the solution so what is missing tools is missing case studies really understanding if these kind of patterns about fairness and bias really hold on different geographies in different types of interventions in different types of policy areas from healthcare to criminal justice, are we really getting common kind of results across these groups? Is record linkage really one source of the bomb? Is the label noise? So we really, nowadays, we are not really sure of what is the panorama. Regulation will come. You might be sure of that. GDPR is just a start, and it's a start tackling privacy in the day ownership. but accountability, explainability, fairness, there is regulation that is coming and that's good. And we also need review boards. So external entities like SEA for food or for drugs, we need some kind of external boards for algorithms and not algorithms, for models and how to deploy them and how to assess impact continuously, not just the specific sample, but also through time, we need to conduct regular audits of the models. So, audit is the first step towards mitigation. This is really the old school cliche that if you cannot measure it, you cannot improve it. This is really, really a cliche, but that's true and people didn't choose to measure it. So, that's why we developed ECITAS. When it was released, it was the first toolkit that would allow you to, for decision-making in policy to really get audits for specific groups on binary classification and now there are other tools out there most of them really focus on the mitigation side not much on really understanding and the understanding the metrics and and and tailor what metrics you should care about we are also targeting a non-technical people like public policy people like management as well not just machine learning people or data science people um and today there are the toolkits out there um the way we envision akitas is to be used when you are evaluating selected models if you are a data scientist and if you're a policy maker or a product manager or a client that is getting a model that was outsourced to a consulting company you should audit the model before you accept it to deploy and also uh you should conduct regular audits of your models to maintain the models. As people often try to get concept drift, try to get and see if there is model degradation, we should really start getting this as a standard practice of actually seeing if there is a degradation on the fairness goals. How can you use it? It's a web audit tool for non-technical people. You can install it locally and it's a local web a site where people can upload a CSV with predictions, and they can click and get the results of the bias report. There is the Python library, which is just used on your notebooks or wherever you just import and you audit a specific model. And there is also the CLI, which is good for batch. Like, you want to audit like a thousand models that you train, and you just, if you have the predictions and the attributes on the SQL table, you You just write down the SQL query and it interfaces with your database and you fetch the predictions and it also uploads the results for all these 1,000 or 2,000 models. What do you need to audit the model? You need a prediction and you need to define a K or a threshold to create this binarization. You need the attributes that you want to audit for. And once again, these attributes might or not have been used on the modeling side. And you need labels if you are interested in the disparate errors. So, what are the weaknesses of Akitas? The most obvious are that we assume random label noise. So, we assume the labels are the true outcome of the past events. And we really treat them equally regardless of the group. And also, it's sample specific. But this is not a weakness of Akitas itself. is a weakness of this paradigm of auditing and bias and fairness because we need to create the confusion matrix on a given sample. The size of the groups depends on the sample. Maybe not all the patients or all the people interacting with criminal justice systems or your whatever product you have. In a given week or in a given month or in a given year, the population and the subgroups might be different. the sizes might be different, the prevalence and the base rates might be different and if we are really measuring specific samples the audit for bias and fairness is bias itself to the sample but for now this is really an unsolved issue few ideas is creating some kind of a cumulative kind of auditing through time or just continuous sliding windows, but just be mindful of the results you get is just based on the sample you tested. So what you need to define before accepting or rejecting a model to go to deployment. Assuming that your organization has evolved to really don't accept a model to production if it doesn't meet fairness criteria, you need to define what groups you care about. In the legislation in the US, there is a definition of what is a protected group. Usually it's based on race, ethnicity, nationality, religion, veteran status, sexual orientation, age, but you might define a different type of attributes you care about for your situation and you need to define what metrics concerns you care about you need to define this so we started 10 minutes later so I'd say more 10 to 15 starting 1130 it's there is a coffee break now yeah at 11.20 it finishes the tutorial and it starts the copy break but 10 minutes after that it starts the next one ok so ok ok ok so the before the paradigm is before you pick if you just care about global performance you pick maybe the yellow model. Nobody actually does this kind of plots, but this is for you to really get an idea. But as you audit for bias inference, you unfold a new dimension. And maybe instead of picking the yellow model, you might pick a model that compromised the global performance a little bit, but is as some kind of a lower bias. So this is really the paradigm for model selection. This is an example of Compass. I believe most of you have heard of this. So, there is a Compass notebook and let's go there. And you can find the case studies on the slides. I think we'll not have time for that, unfortunately. So, if you go to Akitas and after you install it, there is this demo notebook. And if you want to find the data is on the examples and you have the raw data with the scores and you have the data that we use for the notebook so basically this is actually the data that was used to do a compass article in ProPublica and you have the entity ID you have the score you have the label value you have race sex and category. Akitas needs categorical variables. If the input is non-categorical, it will create automatic bins, unless you define what bins make sense. And I think the bins should be predefined by you and your specific application, because just following the distribution might not be meaningful. And basically this is an example of using Akitas on a Jupyter Notebook. It's based on four major classes group, bias, fairness and plotting. Group is about just calculating the confusion matrix for specific subgroups as the really beginning example I showed. Bias is about the parity measures and you need to define what is the reference group you care about. You might define the reference group based on specific predefined groups, like men as historical favorite, or white people, or you can use the majority group, so the group that is the larger group will be the reference, or you can use the mean metric group, so the group that for specific metric gets as lower value, and then you'll get the reference. So this disparity will be bounded between one, which is for the minimum, two plus infinity, the disparity is always compared with the minimum value. And fairness is based on thresholds you care about and we'll tell you like just fair not fair and the plotting is just we are not actually data visualization engineers if there is any here we are open to get contributors so the charts are not amazing so you can go throughout this but the group is just the first main method is just we get all these predictions, let's just create a crosstab where we group by the attribute for each attribute and we calculate false positive, false negative, true negatives and true positives and then you get something like this, now you are aggregating by the attributes and you get how many were predicted positive and blah blah blah, the number of label positives, label negatives, the group size and the total entities of the whole dataset then you have the matrix recall, true-negative rate, false-ambition rate, false-discovery rate. Do you remember that table of metrics of ROC? So, they are here. And then you can plot it right away. This is helper. And here we are showing the number, the size of the group. And the coloring is really based on the size of the group. and this is just the specific metric value. This is for false negative rates, for instance. You can plot all metrics at the same time. You can basically, here you define just the number of columns you want. Then here you are calculating the bias, and you get the disparity based on predefined groups, where you pass a dictionary with the predefined groups you care for. each one of the attributes will be the denominators. And then you just pass the crosstab that you got, the original data frame for statistical significance calculation. So basically what we are doing here is based on the metrics, for instance, if the metric cares about false positives, we get all the false positives for a specific group, false positives for another group, like just binary, and we calculate the t-tests for this and this is as simple as getting the disparities and then you can plot the disparities here in 3Maps is the size is the size of the group and the color is the the darker is the larger disparity the lighter is lower disparity so the in this case if you see here the the groups are always the same size but based on the metric they might have higher or lower higher or lower disparities so yeah I think I need to wrap up I'm sorry for this we started later because people were coming in but it seems we need to go any question and reach out to me on LinkedIn or so and And you can ask questions based on the slides. There is more things there. But, well. Yeah, here you can see this is actually the data from Compass. and you can see like one of the things that people were complaining was the false positive rate of african-american being twice as a false positive rate of caucasian so this was actually part of the article and the tool detect this for the women versus men thing it yeah so one of the misconceptions in this area is like um it seems like it's ai that is creating bias? No. AI is actually showing the bias of the historical case. So this is a misconception. So depending on the metrics, one metric that they figured out was like just the selection parity. So it was selecting much more males than females regardless of the label. So if you really care about the label, people that actually deserve to be selected, the predictions were basically kind of independent on that. But also based on just the base rates of historical data they found that the base rate for men were much higher on success than than women so yeah just it is about using these kind of tools to really assess what's going on on your test data and do it regularly actually there is a bit on that on the on the end of the presentation is about you can override and put in place unfair i'm wrapping up unfair unfair interventions unfair decisions um in order to override the historical case so you might decide to actually be unfair and start selecting some kind of affirmative action, that is an example, but for the historical unfavored. So you are changing the way you are collecting data to favor things. Sorry, guys. I'm really sorry.