No, you can't 'eval' your way to fairness

Algorithmic fairness cannot be achieved solely through quantitative metrics or evaluation libraries because fairness is a lived experience rather than a calculable state. The core problem is that technical evaluations often rely on a model-free approach, using tools like LLM-as-a-judge to evaluate other models, which creates a circular logic that fails to address systemic bias. Furthermore, relying on a single metric often leads to misalignment; for example, a machine learning engineer may optimize for equalized odds (balancing false positives and false negatives), while a product lead may prioritize demographic parity (equal selection rates across groups).

To address these failures, system designers must distinguish between harms of allocation—who receives a resource or penalty—and harms of representation—how groups are depicted or omitted in data. High-risk examples include the COMPAS algorithm, which demonstrated higher recidivism flags for black defendants, compounding existing systemic injustices. Because sensitive attributes in datasets are imperfect maps of lived experience, the approach shifts from purely mathematical optimization to human-centric design. This involves red-teaming assumptions, utilizing feminist datasets for fine-tuning, and applying principles from Design Justice, such as centering the voices of those directly impacted by the system.

Key takeaways emphasize that choosing a metric is an implicit choice of a value system. Effective intervention requires moving upstream in the product cycle to make system decisions visible and interpretable to users, inviting their consent rather than applying "silent" adjustments under the hood. Ultimately, technical tools are secondary to qualitative engagement, requiring designers to act as facilitators who incorporate the lived expertise of marginalized communities to mitigate intersectional harm.

This description was generated by Open-Source AI using the transcript of the session and the original submission contents.

This session took place in track Ethics & Privacy.

Submission

The proposal as submitted by the speaker before the conference.

Cold open Fairness is fundamentally not tractable to classic optimisation techniques.

The exposition Fairness is not a state of the world, it's an experience of it. No technology is fair in a vacuum. Fairness can only be understood when a technical system collides with humans in the world. It is felt as much as it is calculated. We can look at statistical results in aggregate to understand patterns, but these do not tell the story of the individual.

Further, attempting to optimise numerical fairness metrics is fundamentally coercive and technocratic: putting our thumb on the scale globally, injecting "positive bias" into single dimensions, framing fairness as a data problem rather than a problem of human dignity. It's a "one metric to rule them all" approach that fails to acknowledge differences in preference, culture, experience. To build systems that support human agency we must first abandon our idea of a single moral machine which consistently outputs correct answers from inputs and algorithms. Any system treating people as fungible or undifferentiated is structurally unfair.

What might consent-based fairness look like instead? Asking "Do you want extra help?", making sure individual preferences and self-reported disadvantage can add a layer of human respect into the equation. But we rarely see even this. Instead we see universalist design that decides what's good for people without consulting them - the same pattern that Design Justice critiques as erasing those who experience intersectional disadvantage.

What does this have to do with evals? We're seeing a wave of off-the-shelf libraries measuring bad behaviours in LLM outputs, often simplifications of older fairness metrics. And yes, they can catch obvious failure modes like slurs in outputs. But this is one failure mode among many. Installing a library and calling the job done is fairness washing. The harder, more fruitful approach is to explore the space of failure modes, consider what an ideal world would look like, and design measures, mitigations, and feedback loops accordingly. It also means grappling with the fact that we cannot avoid doing harm. What we can do is harm reduction, humility, and striving toward something better while acknowledging the impossibility of the task.

Third act This talk won't offer easy answers. Attend if you want to grapple with the gnarly problems of building systems for humans. We'll borrow ideas from Design Justice and the disability rights movements: nothing about us without us. Let's ask and answer better questions. You'll leave with sharper mental models and tools for the next tricky conversation at work.

Outline (30 minutes): The problem (10 min):

  • Fairness as experience, not state.
  • Why optimisation fails.
  • The individual vs the aggregate.
  • Why treating people as fungible is structurally unfair.

The critique (10 min):

  • Off-the-shelf fairness evals as fairness washing.
  • The temptation to install a library and call it done.
  • What these tools can and cannot catch without further analysis.

The alternative (10 min)

  • Borrowing from Design Justice and disability rights.
  • Exploring failure modes rather than optimising metrics.
  • Harm reduction over false perfection.
  • Transparency, explanation, empowerment.

What you'll take home You'll leave with sharper mental models for thinking about fairness in technical systems, frameworks borrowed from Design Justice and disability rights movements, and tools for the next tricky conversation at work about what fairness actually means. There are no easy answers here, but there are better questions.

Transcript (auto)

Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.

Speaker 1 [00:00]

Welcome to Laura Sommers. Laura is a very technical designer working at Pydentic as Lead Design Engineer. Take it away Laura, thank you. Thanks so much.

Speaker 2 [00:16]

Thanks so much, Florian. Thank you, everyone, for coming. And thank you for coming to my spicy titled talk. I hope I do not disappoint. Good. That went forward. These are all photos of me as a young person, just so that we try and ground this in humans and history. So, yeah, I have been working in tech for 20 years. I work mostly in startups. I care a lot about the responsibility and ownership that we as technical designers and system builders should take on. But also I will say the more I have worked in the field, the more I have grappled with these questions and the more I have thought about it, the more I realize how big and hard it is and the more I realize that there's no magic pill, there's no silver bullet, there's only difficult questions and grappling in the gray. And that's what I want to encourage all of you to do. So just a quick summary of what we're going to talk about. I'm going to talk about why I am so dogmatic that fairness is not a metric that we can optimize for in this experience of the world. I'm going to talk a bit about fairness libraries, like the ones that came up around 2017, 2018, and then again, like what's been happening in the eval space over the past couple of years. And then I'm going to offer some alternative activities, things to do, ways to think that I hope will help sharpen your mental models and give you something practical able to take away from this talk. So to start off, oh, by the way, I live in Berlin, and there's the most amazing graffiti. This is outside my dance studio, and it's, like, incredible. So let's kick off, like, warm up a little bit with an activity. So I just want you to put your hands up. Who would prefer to see a film in 70 mil versus IMAX? So, like, old-school analog versus, you know, megapixels. So, 70 mil. Ooh, okay, we've got a couple of old-school nerds. Yep, good. And then IMAX. Okay, that's the clear winner. So, the decision is, I'm booking tickets for all of us to go see Project Hail Mary tomorrow, and we're all going to IMAX. Cool. How does that feel? Oh, Paloma's happy. No unfairness. That was the easy mode. Okay, next one. Are you a night owl or an early bird? Hands up, night owls. yeah okay cool and hands up early birds oh the clear winner okay so based on that next year the conference is going to start an hour earlier okay that landed good interestingly there's some there's actually some really interesting research about early birds versus night owls and it seems to have something to do with what time of year you're born and how much light you're exposed to like in your early like days of infancy. So anyway, side note. Okay. Last one. Last activity. I know. Hands up. I just, I'm not asking you to reveal anything too personal, but I'm just wanting to get a rough idea of the ages in the room. So give me a hands up if you're in your twenties ish. Cool. All right. People in their thirties. Okay, that feels like the majority, awesome. And 40s and more. Yep, okay, there's some of us. Good, I don't feel too alone. Great. Okay, so based on that, I'm a system designer, and I have decided that everyone who's over 45 is going to get the easy mode of the app that we're building. And that's because a lot of research shows that people who are over 45 have a hard time adopting new technologies. So they're not going to see things like advanced filtering or keyboard shortcuts or API docs. Fair? Not fair. Fair? Not fair. Awesome. Good. Okay. Talk is done. Okay. So why does that feel so bad? Something was decided about you on your behalf without your knowledge and consent. And if you take away one thing from this talk, it's maybe this idea that we should never, if we can avoid it, ever do this. We should always try and make system decisions visible and interpretable to users, and if possible, also invite their preference or consent. I want to talk a little bit about fairness, even in a non-human setting. I have two cats there, Rios on the left and Soji, and this is literally a photo of me giving them treats, and when I give them treats, like these little hard nuggets, they look at me, and if I say give Rios two and Soji one, she will literally look at me and be like... like, and there's this like deep and abiding sense of injustice. And I bring this up to offer the idea that there might be something in the world that just, you know, before intellect, before language, before what we think of as human consciousness, which is a sense of justice or a sense of injustice. And it's obviously related to allocation, which we're going to come back to in a minute. So yeah, my pitch is that fairness is felt before it is calculated. And that's how I want you to, to like think of it going forwards. Calculation is the end point of a lot of thinking and a lot of interpretation, but the gut feeling, that's the important bit. Why did Andy get the job and not me? Why did Florian get on the bus first? You know, like these are the things, I know Florian's so rude, dude. Come on. But these, these are, these are the ways that we experience fairness on a day-to-day basis, like in individual moments and individual experiences. So to offer a little case study. So we have a system, and it's trying to make some assessments about what people might need a helping hand. So who needs an intervention? Who needs, like, positive bias injected on their behalf? So they're looking at two candidates. They're both young American women. They're both black women. And they both have been marked as candidates for intervention. But what the system does not know, and often will not know, is something like this, that one of them lives in the White House, the other lives in a group house. Now, when we think about socioeconomic status, when we think about relative privilege, these are often the things that make a huge difference to the lived experience of bias. And it is not to say that being a young black woman in America means you have no experience of bias if you have significantly more privilege, but you would have a very different idea about who needs an intervention and I certainly wouldn't want to be the system offering one of the Obama sisters an intervention when I could be offering it to someone who is arguably much more in need of it. And oh yeah, I think the broader point of course is that systems looking at sensitive attributes are only ever in an imperfect map. They're not capturing the full texture of someone's life and this is a thing that we have to grapple with as data scientists, as machine learners, as people working in a field of quantitative data that we just never know the full detail and guess what? There's a wonderful solution for that. You can ask. Awesome. So yeah, fairness, not a state of the world experience of it. Next section, I want to talk about metrics and evals. So just to like ground our understanding and make sure we're all on the same page, there's two kinds of harms that we generally talk about in the fairness space. One is harms of allocation. So that's who gets what. And that allocation can be a positive thing like access to a loan or getting into university. And it can be a negative thing like staying in jail for longer or getting a bigger fine. And then there's harms of representation. And that's basically who shows up in the data and whether an underrepresented group is even less represented or represented in a negative way by the system. So just to offer a couple of examples to ground these ideas. and of course, like anything, there's a big spectrum of types of harms that we could be talking about and some of them will be very low risk and some of them will be extremely serious. So a low risk harm of allocation is something like, you know, you go to a doctor's surgery and you know, you've got 10 children and two of them don't get candies because the system didn't allocate it to them. It's annoying, but it's not the end of the world. Maybe you have a tantrum, but it doesn't derail your life. A high-risk harm of allocation, a classical example is this COMPAS algorithm. Has everyone heard of this? Just quick hands up. Couple people, okay, I'll explain it. So this was a system, ProPublica did a really good breakdown of the algorithm if you're interested, but essentially it was an algorithm used to try and detect recidivism, which is a way of saying like, is someone likely to re-offend once they've been released from prison or from jail? And so what was happening was it was more than twice as likely to flag black defendants as potential future offenders, meaning they were more likely to stay in jail for longer, which amplifies these existing harms that we know happen, which is that black offenders tend to be put in jail more often and for longer in the first place. So it's sort of compounding these existing harms that we know about. harms of representation. I'm on Snapchat and I want the bunny ears filter, but for some reason it's not so good at reading the edges of my face or maybe my hair has a crazy shape. So my ears are up here and it's annoying, but you know, it's not the end of the world. And a high risk example of a harm of representation is like, you know, a young person looking for mathematicians and sees only old white dudes. And it's like, well, I can't see myself there. I have no, and you know, Whether consciously or unconsciously, that starts to winnow down their idea of who they can be in the future. And these are just a couple of examples. I'm assuming because you're in this talk that you're kind of already a warm audience and you know about these kinds of problems and I'm not gonna spend a lot of time on them, but there are a bunch of great websites that you can go to to learn if you want to learn more about the kinds of problems machine learning systems can propagate on the world. Awful AI is a good one and there's tons more, so if you want more links, ask me later. Okay, so let's look at tools. So I did a little analysis of the classical machine learning fairness libraries and then libraries for evals. And I was looking specifically for evals that had pre-canned metrics. So a lot of them basically have very few or no available metrics and they just let you write your own custom metrics and the job that they're doing is giving you essentially the eval harness and not actually the eval itself. But what I was interested in was how much are these fairness topics showing up in the eval's library space. And I suppose it won't surprise you to know that it's not very much at all. So just to give you an idea, the Fairlearn AI fairness 360, Equitas responsible AI toolboxes are all kind of the classical, like let's call them old school AI libraries, fairness libraries, and then these are some of the ones that show up in the eval space. Excuse me. So let me show you. When I was trying to map them, You can see on the left, those are all the kinds of harms that the classical libraries try to look for. In the middle is where they overlap. And really, most of these evals libraries are only discovering things like stereotypes or toxic language. They're really not doing much at all in the way of doing analytic assessment of the space of overall outputs. So they're not really attempting to determine if the system is biased as a whole. Probably because they know they're definitely biased as a whole. And then there's, of course, other evals, concerns that they're trying to interrogate that are perhaps not exact fairness topics. And evals, I'm not here to tell you evals are bad or unhelpful. But I do think that we have to be really pragmatic about how useful they can be to us in this domain. So I would argue that if you are installing an off-the-shelf library, like an off-the-shelf evals library, and then turning on an LLM as a judge and saying, cool, I've done fairness, it's like you've asked the unreliable narrator to judge its own reliability. Like you're really kind of doing this Aurora Borossian snake-eating-its-tail thing. So I think we really have to be wary of the model-free approach to fairness. Like we should have a model, we should have a mental idea of what it is we're trying to determine, what harms we care about, what harms we're trying to prevent, and we should be, like, making sure that we're injecting those opinions at every stage of our analysis and intervention. So let's look at what that can look like. I have a lot of Venn diagrams for some reason in this talk, but let me try and, like, map for you what I think, like, the jobs to be done are in this space, because I hope this will be helpful to you so like one big one is understanding the world like understanding the space of intersectional harm intersectional bias that has happened historically and that is an entire job that is a thesis that is a you know like that's that's a company there's a lot of work to do there and then a sort of related but smaller job to do is understanding foundation models and LLMs and I mean as we saw in the keynote yesterday like that the architectures are manifold there are so many different systems being built and deployed in the world. And I would suggest that there's probably differences in the kinds of harms they can propagate. And I would say the research and the understanding of fairness cross transformer architecture is basically non-existent right now, but hopefully at some point we will catch up. Then there's the question of your own system. So if you're in academia or if you're in a company or if you're working on a product or and an analytic system, understanding the specific domain of your system, understanding the potential harms that are more likely to happen in your system. For example, if you're working in medicine, you probably care about the harms and the potential bias that are relevant to your field. You might be saying, hey, we know that computer recognition that does scanning skin tones does really poorly on darker skin tones, so we're going to be very careful not to take its word for it if we're scanning for, like, moles that might be cancer if the person's skin tone is over a certain darkness. You don't probably care about the bunny ears scenario, right? Like, the point that I'm trying to make is that we have to really be pragmatic about winnowing down our attention and making sure that we don't get too focused on, you know, like, topics that are interesting to us but maybe don't matter in the broad space of, like, what are the potential risks? What are the potential harms that can happen. And then we want to intervene. And the point I'm trying to make here is simply that there is some overlap, but I would argue that there's a huge amount of activities we can do in the intervention space that are not purely modeling the world or running these fairness metrics over them. So let me keep going because I have quite a lot to cover. Here's a very large diagram that was my attempt to map all the kinds of activities we as humans could be doing to try and improve ourselves, to improve our systems, to improve our networks, to improve our communities. Let me just try and pick a couple. Stay curious. Talk to people. Red team your assumptions. I think there's a huge amount of overlap between security process and fairness process. If your team is interested in security but not fairness, it's an easy way to move them along the spectrum to say, hey, well, we can apply this security process, which is red teaming, but you know, we're going to investigate these kinds of problems. There's, yeah, there's a huge amount to do. You can, you can be an activist. You can work on building communities. You can make reading groups. You can learn for yourself. Like there's of course impact and analysis work to do as well. But the thing I want to flag for you is like, really don't feel like if you can't write a metric about it, the job is done. There's so much else you can do. And to try and land where my thinking has got to, when I've been asking myself these questions about why have I got such an ick when I see people just focusing on a single fairness metric but not really having a bigger conversation about what their system is doing, I would offer this meme. I think me and my little Python script. It's just no match for all of law, philosophy, religion, academia, culture, and art on this topic of justice and fairness. We have thousands of years of history of trying to ask this question, like what is right? What is fair? How do we treat people? So yeah, me and my little Python script, maybe like we need to have a sense of our place in the history. So more about intervention. I want to talk a bit about some of these topics from design justice. I realize I've only got a few more minutes, but I really love this book, and it does a lot of heavy lifting around the topic of who has power, how it affects oppressed groups, and it invites you to go on the history journey of understanding different historical movements in design and how they've grappled with the topics of fairness. So there's universalist design, there's individual design and fairness. There's a bunch of different movements it takes you through, and it's a really useful way for learning more about what there is to do and some of the challenges there are with different approaches. I also really love designjustice.org, which Sasha Constanza-Chalk, who's the author of this, also flags that everything in their book and everything that they're flagging is a whole bunch of communities. I can't even say all the names of all the people who are involved in this kind of work, but please, I invite you to go read and learn. But there's some principles from design justice that I think really give us useful frameworks as people working in AI and ML. So I'm just going to pick a couple I really like. One is principle two. We center the voices of those directly impacted by the outcomes of the design process. I think often we think of the users of the system or even the people who the system makes decisions about as almost the least important person in the system. And I think that's completely wrong. how did we need to invite them in as a primary stakeholder, as an important participant in the process. We see the role of designer, but I'm going to sub in here data scientist as a facilitator rather than an expert. I really love this as well because it means we make sure that we invite learning what we don't already know. Another principle I like is that we believe everyone is an expert based on their own lived experience and that we all have unique and brilliant contributions to the design process, and I mean, you know, again, like systems design, we're not talking about UI per se, we're talking about what goes into the sausage, and, you know, there's more there, but the reason I really like this is it helps you reframe how to think about talking to and working with people that are not the people in your team. A few other resources I'm just going to wave at. And also, I have a QR code at the end, and all of these are linked, so you don't have to take pictures. And if you're interested to read any of these, I encourage you to go learn. But there's a bunch of toolkits. There's a bunch of additional books that express some of these kinds of challenges that I've been talking about with a fairness base. And yeah, I basically want to pitch you, like, yes, metrics can be good. Evals can even be good. but they're only as good as the hypothesis behind them as your understanding of the world and how that's reflected in the metric that you have chosen. And I want to give you another example before I stop for questions. Yes, five more minutes, good. So imagine you have a CV screener. We're building a system that's screening CVs. We all know this is an easy way to propagate harms. And we have two colleagues who are having a discussion. They're like, okay, we want to make it fair. Good, good, we've had the discussion. We want to do it. Great. So the ML engineer goes off to tune the model and the product lead goes off to watch the dashboard. And the thing that's difficult is no one has really had that difficult possibly confronting conversation. What do you mean by fair? So six months later, the engineer comes back and says, hey, I tuned this model to achieve equalized odds. That means I have the same success rate for false positives and false negatives. But the product lead sees that there are two demographics with very different advancement rates and says that's not fair. So their understanding of what fairness was was not aligned in the first place. So let me give you a confusion matrix and hopefully this will land the idea. So one, the engineer was optimizing for equalized odds. So they were looking for predictive positive, actual positive, predictive negative, actual negative across all demographic groups. And that's what they were optimizing for. Whereas the product lead wanted demographic parity. So they wanted to see this selection rate, the success rate of both groups looking more equal. And this is a really good example of how even two very simple, very commonly used fairness metrics can be completely misunderstood. And I can't go forward. Cool. Yes, this is the fun thing about this talk is that I, let me try and grow my slides again. Okay, cool. Great way to waste some time, Laura. All right, let me try and go forward. There's a bug there in the slide. So yes. Always debug your slides. Yeah, I wrote these slides with Claude and it was quite a ride. that's another talk okay so yeah choosing your metric is choosing your value system and that's what i'm trying to land all right um and i want to offer just a little bit more data um so i did some analysis of all of the speakers in this talk in this conference and i was looking for this question of who even has access to sensitive data to these protected classes we care about like race age, socioeconomic status, gender presentation, any of these others, and I did a little bit of analysis to attempt to work out who has no or some or direct access, and this is, you know, it's a very lazy data analysis, so, you know, if I'm wrong or if you're in one of these companies and you're like, that's not true, I do have access, you know, I was just trying to make some high-level aggregate assumptions based on what I could see on the internet, but the thing that I wanted to flag is that most of us, my gut feeling has been for some time, don't even have access to these things. So we should be looking for other tools and other ways of learning and interrogating our systems and making them better. So I would offer you the further upstream in the product system that you can intervene, the better. And I would invite you as much as possible to design with people, not for people. I'm going to just flip to the end because we're out of time. The last thing that I really want to land for you is this might be an uncomfortable thing to say out loud. Do you identify as a member of a disadvantaged group? If so, we'd like to adjust the system so it works in your favor. Would you like us to do that? And that may be weird and uncomfortable to say out loud, but the point I would like to suggest is that it's still less bad than doing that silently under the hood, which is what we did in our activities in the beginning of this. like doing it silently under the hood is a consent-free, visibility-free way of working and it's less respectful of people's agency and preferences and humanity. So in conclusion, fairness is not a metric. Now go make some good trouble.

Speaker 1 [24:18]

Thank you so much, Laura, for your interesting talk and thank you for already asking some questions. You can also ask some more and upvote questions at talks.pycon.de. Laura, the first question to you is, do you have any good practices or approaches how to get real life the experts in and how to convince the organization to make fairness one central pillar?

Speaker 2 [24:43]

Ooh, okay, kicking off with some big ones. With the experts, I'm going to assume, well, sorry, if the person's in this room, do you mean experts as in fairness experts or experts as in, like, the subject matter expert? Yeah, you mentioned earlier that the people with their life experience are the experts. Yeah, right, right, so the people who you're intervening on. Yeah, it's a great question. My, look, I mentioned before using security is like kind of the skinny edge of the wedge. And I think that you can also use usability or like UX performance testing as a kind of skinny edge of the wedge and be like, hey, we're just gonna do a little UX research and we're gonna find out what's happening with people who are using our system. And if, hey, presto, bingo, you're learning that maybe some harms are happening that you hadn't planned for, you can say, well, is this a reputational risk? Could this affect our bottom line? If you have stakeholder bosses who are not so sympathetic to this topic, I would try and tie it back to bottom line. It's not my personal preference because I think we should do good work from a principled reasons standpoint. But if you have a hard time getting it to land otherwise, then I would try and try and find a way to tie it to either reputational risk or like possible ways that it could cost the money, cost the company money. If it's contravening GDPR or some other data privacy, that's also a nice way to say, hey, there's law about this, we can't fight it. But yeah, I would basically be saying, if you have a system that operates across a bunch of people and there are no feedback loops baked in at any point, you're not finding out if people are having a good time or not. That's just an easy low-hanging fruit way and it's basically saying, hey, we should find out if people are even happy with this thing. Yes, they're using it, but for how long?

Speaker 1 [26:32]

Thank you, Laura. The next question is, I would say that most LLMs are not fair by design, meaning the training data. What would you recommend regarding choosing a model and adopting it in an application?

Speaker 2 [26:45]

That's a great, great question, and you're absolutely right. We have a world where there's known intersectional harms, intersectional bias, and those harms have been trained into the models because it's in the training data. I have seen some people doing some things like designing feminist data sets and doing a fine-tune round, but it does really depend on whether you have the space to do something like fine-tuning or to host your own open source or if you're required to be working with a paid API and using one of the foundation models provided. I would say if you're in the world where you have to use the API and you don't really get to choose your model, focus more on the upstream question of defining your problem statement. Look at the Rubik's Cube twisted a few more times and see if there's another way you can try and design your system to avoid the worst harms and then otherwise assume the harms are baked in and there's not much you can do about it. And if you can, fine-tune a model or pick an open-source model. Look for feminist data sets. Look for data sets that try and inject the positive bias that fight the harms that you're worried about and see if you can try and improve your system in that way. But yeah, I think we have to be pragmatic that we can't fix everything all at once. So I try to recommend picking one thing to do and try and fight for that rather than trying to fix everything in the world. And of course, we're individuals. These are system problems. So find friends, make community, work together.

Speaker 1 [28:14]

Thank you, Laura. The next question is, what options do we have to assess fairness with topics that impact virtually everyone, like climate change? Not everyone is impacted equally, but due to its indirect impact, people who are or will be impacted the most do not have the resources to get into the topic in detail.

Speaker 2 [28:32]

detail oh my god such a good question um oh i don't know i don't know i don't really have a good answer to this i mean if you have people using your system you have the excuse to talk to them but when you're talking about someone in a pacific island who's not even online like you're not gonna you're not an ethnographer you're not an anthropologist you're not gonna go and like have a you know like sit in with them and talk to them so um I'm just going to admit it I just I don't have a good answer for this I'm sorry but I'd love to talk about it more afterwards

Speaker 1 [29:07]

Thank you. And then potentially the last question. What can I do, knowing my little Python script has a minimum space within other weights, to make it not to reproduce harm? How to recognize it within the fake neutrality of technical choices?

Speaker 2 [29:24]

How do, oh, man, these are such good questions. I mean, look, I think maybe if I just go back to this, if we just remember, go here, please, yeah. If we just remember, like, me and my little Python script is, you know, it's a drop in the bucket. It's a normative statement, and it's saying, I care about our system not harming people. I care about trying to reduce the worst of the harms. But I also know it's not the full solution.

Laura Summers

Laura is a very technical designer™️, working at Pydantic as Lead Design Engineer. Her side projects include Sweet Summer Child Score (summerchild.dev) and Ethics Litmus Tests (ethical-litmus.site). Laura is passionate about feminism, digital rights and designing for privacy. She speaks, writes and runs workshops at the intersection of design and technology.

Social card for talk: No, you can't 'eval' your way to fairness