The Sound of Silence: Online Misogyny and How we Model it
At Opt Out, we’re building a set of tools designed to help female-identifying people engage with healthy debate online safely. Victims of online misogyny experience a sense of fear and attack on their integrity. Amnesty International found that of the women who experienced abuse or harassment online, 41% of responding women felt that their physical safety was threatened and 1 in 2 women experienced lower self-esteem or loss of self-confidence as well as stress, anxiety or panic attacks. By pushing female-identifying people out of online spaces, because of fear of victimisation or retaliation, online misogyny punishes female-identifying people further by affecting their economic potential and political representation. Many cannot rely on the internet for a living and in addition, the Chartered Institute of Marketers, UK reported that 83% of women had self-censored indicating that voicing a political opinion could be under threat. In a world that is more entrenched on the internet than ever, these human rights violations are contributing to an oppression similar to those of decades long gone.
Opt Out's founding idea is a browser extension that removes misogynistic comments from an individual’s social media feed. But what is online misogyny? How do you capture the nuances whilst ensuring that the richness of respectful online interaction is maintained? In this talk I will discuss:
- How we generated our dataset - annotators, snorkel-metal
- Our modeling techniques and architectures
- Network - the influence of user profiles on classification
- Our investigation into the syntactic structure of online misogyny
The General Data Protection Regulation (GDPR) has changed our lives online on social media platforms. We have the right to be forgotten, to see what is being collected about us and to opt out if we wish. The current abuse that female-identifying people suffer is not avoidable. We see Opt Out as an extension of the GDPR that also protects the human rights of these people online. These human rights are the right to security of person, non-discrimination, and the right to freedom of expression and opinion. Under the United Nations Guiding Principles on Business and Human Rights, social media platforms have a specific responsibility to respect these human rights. While steps have been made to protect these people online, not enough has been done. Let’s not let hate win, lets Opt Out.
This session took place in track PyData and was classified suitable for none domain / none python by the speaker.
Transcript (auto)
Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.
Speaker 1 [00:04]
Good afternoon. My name is Teresa Ingram and I am the co-founder of OptOut, an open source set of tools designed to help female identifying people engage with healthy online debate. But before I talk to you about OptOut, I'm going to talk to you about why we need to exist. So most people will recognize this. Fridays for Futures is a global climate action movement fronted by 16 year old Greta Thunberg from Sweden. um now Fridays for Futures was not always the global phenomena that it is today late August last year Greta was out protesting alone um but it wasn't long before others started to join her and one of those was 10 year old Lily Platt from the Netherlands Lily started to organize her own weekly protests and it wasn't long before the two were protesting together So that's supposed to be a picture of the two of them protesting. So as Greta's popularity blossomed and bloomed, the hate also boomed. And when Lily jumped into Greta's defense, she received and was inundated with anti-Semitic slurs, um racist threats um links to porn on her instagram and one troll attack that was so severe her whole family had to change their mobile phones and these two are not the only ones we have havana chapman edmondson on the right hand side she has eight years old and has received racist threats death threats and even um it was found out by her family that somebody had tried to contact her who was on the registered sex offenders list and then we have jamie margolin who founded the activist group zero hour now the other girls have normally their parents to handle their social media um stuff but jamie does it herself and jamie is uh gay latina and jewish and she's inundated with hate um to quote jamie from a tweet that she posted last year like she said i was terrified they were asking me things like where is my synagogue and what was creepy was they got more and more deep and found out more and more personal information about me and so i decided to just do like a really brief search trying to figure out um what was going on with these girls what was happening and so i don't think the images are going to work sadly never mind um but really these young women are being targeted online so digital media has opened the door for new forms of oppression and violence against women and when those women also happen to be human rights defenders politically active vocal or challenging the status quo or commenting on areas that are seen to be male this can be even more difficult and so the UN recognizes that there's a problem but even for us at opt out this is not good enough we say that all female identifying people are disproportionately suffering abuse online and when they happen to be challenging the status quo or if they happen to have an intersectional identity, be it trans, black, body weight, you name it, the abuse is yet more intense. Online misogyny is real and it's silencing the voices that society so desperately needs to hear. So what are the social media platforms doing about it? Well, to give you an example from the British MP Jess Phillips, she reported a user who was threatening her with rape threats and this was twitter's response so as you can see not very useful the user was not banned they have been subsequently banned um but initially they weren't um so let's have this larger conversation then a failure to act means that some of our most vulnerable online are receiving rape and death threats and as i've already explained with the example with lily and her family, the boundary, the binary identification between online and offline is not as clear-cut as it seems. So, if big tech are not willing to do anything about this, if they're going to act as neutral bodies, it's up to little tech to try to do something about it. So, what are we doing? How are we going to do this? First step, define and conquer. so what is misogyny to opt out misogyny is not sexism somebody can like female identifying people like women but believe in the power structures that are in place to keep women in their places so women belong in the kitchen things like this um we see um misogyny as the law enforcement of the patriarchy. It's also important for me to tell you why we've chosen misogyny over sexism and also misogyny over online gender-based violence, which is the kind of more broader term which describes all forms of online violence that's gendered. And that's because Opt Out currently is only working on comments and rhetoric, and it's not working on stopping all forms of online violence such as doxing which is the releasing of private details or revenge porn for example. Okay so I'm just going to tell you about some of the ideas that we've been seeing being articulated in our data sets. So people tend to either describe men or women as masculine feminine and there's no in-betweens and they have intrinsic characteristics because of their masculinity or femininity um women are identified with their bodies and by their body parts um dead naming of trans women that's where somebody is referred to by a name they no longer identify with um claiming that women can't do sport that's a real favorite of mine and refocusing the conversation to do with the male angle agenda or perspective and i was playing around with the data sets and i ran some random forest scripts just to try and find some important features of the of the data and a really important characteristic of a misogynistic tweet is this you can't make it up so i sat down with our social scientists on the team and we went from these ideas to these groups. So we've got seven of them for now. So insults, self-explanatory, something that is said just to devalue a person. Sexual harassment, harassment that is sexualized in some form, be it innuendo, a risque joke, something like this. A threat of violence, anything that has a violent physical dimension. Gender essentialism. So this is your men are from mars women are from venus um trans misogyny um i'm going to explain a little bit more about that with the example just later on um objectification treating somebody by their body parts and not as an individual um you know like referring to them or discussing their worth just based on how they look things like this and derailing which is something is probably a bit easier to explain with the example later on but it's essentially just trying to change the theme of the conversation to something that the perpetrator feels more authoritative to talk about. So mansplaining is a great example of derailing. Okay, so here are some examples. So transmisogyny, often what will happen with transmisogyny is that a trans woman is referred to by very male characteristics, you know, the nice knobbly knees there. And these are some of the hashtags that are used often. objectification and derailing so i've already mentioned about the mansplaining another example of derailing is something called a mob attack um which is where a user's social media feed is flooded with messages be them banal or be them totally sexist racist threatening whatever and this cannot just last for hours or a day the some of the worst cases i've seen have been months And what this does is it renders the user unable to really use their social media how they should be able to. So it's also important to just mention quickly that currently the model that we're using is not a multi-classifier. These are kind of our academic groups, just so we can understand the face of the beast that is online misogyny. So that's online misogyny for opt-out. So our solution. At opt-out, we're hoping to put a stop to the silencing nature of online misogyny. The General Data Protection Regulation, GDPR, has changed our lives on social media platforms. We have the right to be forgotten, to see what is being collected about us, and to opt-out if we wish. But the abuse that female identifying people are receiving online is not optional. We see opt-out as an extension of the GDPR that also upholds the human rights of these individuals, allowing them to engage with respectful online debate once more. So we have an open source set of tools designed to help any female identifying person get back to the online spaces that they've been chased out from. And our main tool is our browser extension that works essentially like an ad blocker but filters out online misogyny instead of adverts but we're also much more than that we're holding workshops that are bringing together female identifying people to not only offer support but to act in a form of protest that enough is enough and by doing this we not only provide a valuable resource for these people but we also are able to identify needed technical infrastructure ensuring that our tech is as community driven as it possibly can be we also have a awesome website which I'll get to in a second and we are trying to develop our own antidote to Silicon Valley KPIs so other KPIs that measure participation, measure things like number of users or number of shares, what we're trying to do is to develop KPIs that measure things like diversity and inclusivity and health of online discussion and finally we're just being as loud and proud as we possibly can be So, this is our website. Our website is inspired by Egyptian-based NGO HarassMap. HarassMap is a website where you can go and submit anonymously details of a physical incident harassment that you've received and it's mapped, it's put onto a virtual map. Opt-out version, you can submit anonymous details of an online incident that you have received and that data is then stored, studied, and feeds the models that our tools depend on. We're also showing transparently how the data is being used and usage statistics and our KPIs on the website, so we can try to show female identifying people how their submission is helping their online sisters and help to fuel the movement. Finally, in the long-distance future, we hope to have a virtual harass map, which will enable female identifying people to navigate the murky waters of online life as best they can. And finally, our browser extension. So, our browser extension just works with a simple binary classifier sentiment analyzer sat on a server somewhere. What happens is as the page loads, the tweets are sent to the back end. The model evaluates them, returns a score. And if it deems it misogynistic, they're removed. And if not, they stay. And if there's an image attached to the comment, then the image is also removed. So simple. Not too difficult. Well, the browser extension has also got a lot more coming. That's what it is currently. But there's a lot more to be added to it. So there is going to be automatic reporting of misogynistic tweets. Anything that is deemed misogynistic will be automatically sent to the moderators to be checked over. And also, the browser extension and opt-out in general is consent, not censorship focused. What's going to happen is that the browser extension will have its own local instance of the model that you can supply feedback to because we're not we're not going to get it right each time and we're not going to get it right for everybody um and so currently what we have is everything apart from the local instance model um and the current model that we have built um still has a lot of work to do but luckily we have a few tricks in our tools in our toolbox to help us with that. So, the ticket to any successful supervised learning problem is a really good labeled data set. But where do you obtain a misogynistic or a feminist data set? We tried many different places and And we had to make our own. So first we started by searching for key terms on Twitter, camel toe being a really important key term to use if you want to find misogynistic comments. But then we found that our search was as good as we were mean or as good as we were creative. So we decided to then search for politicians' names and any other outspoken women online, for example, Zoe Quinn, the games developer, who is a victim of Gamergate. And I suggest that you go read about that if you don't know what that is. And so we ended up with too many comments, too much data, and it was very diluted. So then we reached out to our academic friends, and Zirik Razim very kindly gave us his data set, which was developed using these search terms you can see here um which are to do with my kitchen rules which is an australian cooking tv show um and also um the final data set that we were able to obtain is the automatic misogyny identification data set from elizabetta facini um and this is quite a new data set for us so i'm not entirely sure how she collected it and developed it but um opt in to opt out and we'll let you know and we know so we have the data great but it's not labeled and after a disastrous round of um our own attempts to annotate we decided to go snorkeling on the great advice of one of our team members andrada um i mean our agreeability score was so low that i was wondering whether we should really continue the project at all so this was really a savior and what has who's heard of So, Snorkel is a framework for multitask week supervision. Essentially, what you do is write some labeling functions, which are then used to augment your data set. So, you manually label some tweets, some comments, and then you write some rules. And these rules can take these forms. regex being the most popular. For example, these are some that we wrote. So, rape glitch is an online dialect that's used to attack women, but I'll go more into that later. And then you evaluate. You see how well your rules accurately predict the comments that you've labeled, and then you also see how much conflict there is between the different labeling functions. And you end up with a table like this, and the two really important columns are the empirical accuracy and coverage. You want your empirical accuracy to be above at least 0.5, because that's the accuracy of your labeling function to get it right, and then coverage to be as high as possible. And from this, Snorkel builds a generative model that becomes your labeling model. And what this does is it produces labels for unlabeled data. And you have your golden ticket. Great. But this wasn't the end for us. We reached out to somebody who wrote a brilliant tutorial on Snorkel called Abraham. and what's come from this is that we've begun a wonderful partnership with Sculpt and Sculpt is essentially a web app version of Snorkel and what this does is it allows our it not only gets rid of the labeling bottleneck but it allows our domain experts our social scientists to distill their knowledge directly into opt-out. I'm no longer the middle person between the data scientists and the social scientists. They can sit there, understand what they're doing, understand the labeling, their label functions, and the impact of the different ones and which ones fit and which ones don't. The best thing is, though, that it allows us to mess up. We can change our mind about misogyny. We can change our different categories. And it doesn't matter. Instead of taking weeks, maybe even months, to relabel a data set, it takes days. And so we have now a data set of about 20,000 labeled comments. And we've just got some initial results. results. These are the 20 top curse words found on Internet, in Internet language. And as you can see, there's already a gender difference in the language between non-misogynistic and misogynistic. And it's a very similar sort of thing here. So, here, all I did was find the used spaCy dependency trees to try and find the subject and verb of a sentence and And filtered it based on whether the subject was male or female. And these are the verbs attached to it. And all I'm really trying to get at is that already, just very basic labeling schema and stuff. We're already seeing gendered language. Differences in the gendered language. So, now we have a data set. We can start playing around with some cool modeling stuff. Actually, what's working best for us currently is our logistic regression model that we built in Sculpt. But that's probably because we haven't got around to really tuning our deep learning models. We also have a hunch that state retaining deep learning architecture won't really cut it for tweets because they're too short. And so this is what we have so far. But really, this is really, really basic and not going to stay like this for long. So I really urge you to just check out this repo. This is where all our data science is happening because where we're going next is looking at graph convolutional networks, trying to develop or trying to understand the influence of different users on our ability to classify whether something is misogynistic or not. And as I mentioned about the rape glitch before, we're also going to look at treating misogyny as a dialect. and after being inspired by a talk that I heard here at PyCon also how to model meaning and I need to find the person who gave that talk and talk to them a bit more about it okay so where we're at currently so we currently have a MVP browser extension an MVP website and we're being active we're getting out there talking to people being out there in the community but we always are in need of more people our website is in need of a serious makeover and this is where we're going. We've just completed our first round of funding applications and for the rest of 2019 we're just going to be looking at improving what we've got and beginning the implementation of the local model and then in 2020 we see ourselves going multi-platform, multilingual and eventually building the virtual harass map but with anything like this it's important not only to talk about what we're doing and why but who we are and we're a range of people from data nerds to social scientists but what's a commonality between us all is that we won't let hate win. Our vision is we want to champion women back into the online world they've been chased out from, support them and their voices while still protecting them and holding the perpetrators accountable. Online misogyny is real and it's silencing the voices that society so desperately needs to hear. Let's opt out.
Speaker 2 [22:48]
Well, thank you very much. That was a fascinating talk as well. Are there any questions from the audience that we can take or? Sure. Can I ask everyone to put their hands up quite tall, so I know where to go afterwards.
Speaker 3 [23:10]
Thank you for an excellent talk and excellent initiative. I personally would like to thank you for setting up opt out and looking out for other people in online. So I kind of have two doubts. I wouldn't say it as doubt maybe one of them is suggestion. So first doubt is that when you say or the browser extensions that opt out produced, does it only filter female identifying online misogyny? Or is it gender unbiased? The way I see online abuses, internet is a harsh place and everyone will be a victim eventually. that's what i feel so is it going to be only redacting the female identifying online misogyny or is it gender unbiased and the second is kind of a suggestion since you guys are basically redacting the online misogyny when the extension is activated in browser extension you can get the usernames and report it to the government or even to the tool like twitter or facebook so that they can ban the people because at least the country where I am from we have online rules and stuff so if you do come in some abusive stuff you can get even arrested so is it possible that you can create a tool where you can automatically report them to the respective organizations or governments to take action on them
Speaker 1 [24:38]
uh so the first question as i understood it was are we just focusing on female identifying misogyny um i think we're just focusing on misogyny it doesn't really i think it doesn't really matter who it's coming from misogyny is is is gender neutral almost i mean it's yeah i mean all it is is law enforcement of the patriarchy so even if it's another woman saying oh you should be in the kitchen that's misogyny um and secondly um as far as i understood again um is there going to be some way to automatically report people hold them accountable yes that's um so that was supposed to be there in the mvp already the problem is in every other country apart from germany that automatic reporting function is really easy to do but in germany for twitter it's a legal document so that's taken a little bit longer but that's because we can't we don't just want to hide the problem we want to hold people accountable for it um yeah
Speaker 2 [25:39]
Thank you very much, and thank you. If anyone else has a question could you also put your hands up at the moment so I can see where to go afterwards?
Speaker 1 [25:51]
Does it only work with...
Speaker 2 [25:52]
work with English or
Speaker 1 [25:53]
English or have you been
Speaker 2 [25:53]
Have you included other languages as well?
Speaker 1 [25:55]
as well we will we certainly will and so andrada who gave this brilliant talk just before on um does hate sound the same in all languages is a key member of opt out so this is you know we're definitely going to do this sooner rather than later because because misogyny exists across but as soon as we transfer from different languages also about transferring the cultural understanding as well so it's it's not even just as simple as retraining some models so it will take time but We're growing. We can do it.
Speaker 2 [26:27]
Thank you. And are there any more questions? I think there's one hand over here.
Speaker 1 [26:38]
I just want to ask how big is the
Speaker 3 [26:41]
How big is your data?
Speaker 1 [26:41]
your data set and how long have you been collected it
Speaker 4 [26:45]
collected it
Speaker 1 [26:46]
Um, so, um, our working, the one that we're most happy with currently is about 20,000. Um, but we've got too many comments. We've got too many. Yeah. It's something like it's over 26 million comments. So yeah, too many.
Speaker 2 [27:05]
And just before I go to the top, I actually had one question myself. I know that a lot of the kind of misogynistic or any kind of hate-filled language is sometimes kind of reclaimed by the communities to which that kind of hate is directed. And so I wondered, especially with things like the browser extension, do you have a way of managing what might be like a potentially false positive, as something that's falsely identified as being hate speech when in fact it's kind of protest.
Speaker 1 [27:35]
protest against it it's like a reclamation of yeah yeah um that's a really good question i think my idea is to have um something that is not just looking at the content of the tweet but also the conversation so who is this comment coming from um and hopefully that will be um make it easier to classify between when something is intended as negative and when something is more of a protest but yeah no that's a really good point and we're not sure yeah Amen.
Speaker 2 [28:05]
further on the road map can I okay further questions here
Speaker 1 [28:14]
You mentioned that the browser plugin would be automatically reporting tweets. Is that in line with the terms of service of Twitter? Can you automate reports? I haven't seen the counter. I haven't seen that we can't. So like I said, though, an opt-out is really about protesting and activism as well. So if we ruffle some Twitter feathers, I'm fine with that. Hi, I have a question because I moderated social media and I know that but so for public fears some things can be
Speaker 3 [28:56]
things can be allowed.
Speaker 1 [28:57]
out. Did you consider that? Because to protect female public figures could be difficult compared
Speaker 3 [28:57]
Yeah.
Speaker 1 [29:05]
to protect just females, like regular women. I think we'd have to look more, I wasn't aware. So no, we haven't. So yeah, we'll bear that in mind. Yeah.
Speaker 2 [29:21]
Uh, okay.
Speaker 4 [29:27]
And thanks for a talk. Very inspiring. I have two questions. Is it a business? You mentioned your raised money. So if it is, what's your model? And the second one is, so you have a definition of misogyny, which I personally agree with. But, you know, how do you deal with going somewhere and people telling you, actually, it's not misogyny. It's just how we do things here. It's part of our culture. You know, we're traditional. We're conservative. And you're a liberal. And, you know, you bring your hippie ideas. You know, how do you deal with this?
Speaker 1 [30:08]
Can I answer the second question first? Just tell them not to download it. It's fine. They don't need to use it. And the first question, so we have ideas for turning this into a business. We can sell an API subscription-based for, say, like Tinder, for example. When we eventually do things like removing dick pics and stuff like this, we can sell our services to dating apps also another idea is selling an API to you know you have like MSN chats with customer service representatives online they get a lot of abuse as well so developing an API to help protect those people
Speaker 2 [30:55]
Thanks, and just one final question.
Speaker 5 [30:58]
Thank you. My question goes away from online to offline to events like this one. And I wonder if you reflect about events like PyData or the general community of data science that the vast majority, well, I read as male. And still, in this talk, I perceive the room as much emptier than in most other talks. so I think I don't want to go as far as asking whether is misogyny also represented here it certainly is but if you try to develop like let me call them tools on a battlefield how do you or do you also work on encouraging those who are affected by misogyny to like get on the boat and develop with you is the question clear Thank you.
Speaker 1 [31:56]
As far as I understood it, how are we doing anything to help get rid of the misogyny within computing or tech on some level? Did I not understand?
Speaker 5 [32:13]
Yes, also. No, I just thought it may be very hard in this community to actually just find people who are affected to join you developing.
Speaker 1 [32:24]
you're
Speaker 5 [32:25]
But you could probably get a bunch of males to work with you on the tool, but that will not be the same.
Speaker 1 [32:32]
So, scarcity is in our favor, actually, because there are very few topics like this. And, you know, there are women in this community. So, you know, so they all come and they all come help. But also, can I answer the question that I thought you asked as well? We're really trying to make it as beginner friendly and trying to allow people to do things that are outside of their comfort zones and career change as well as much as possible. so yeah we're really trying to encourage as many first contributors women changing careers all this stuff into our open source stuff as much possible
Speaker 2 [33:12]
Well, that's brilliant. Thank you very much to everyone who's come along, and can we have a big thank you to both of our speakers today. Thank you very much.