Modern NLP for Proactive Harmful Content Moderation

The rise of large language models (LLMs) has revolutionized natural language processing (NLP), creating opportunities to address complex societal challenges, including the pervasive issue of harmful online content. Despite global regulations and platform-specific policies, abusive speech and toxic content continue to plague digital spaces, highlighting the need for smarter, scalable, and multilingual solutions.

This talk explores how modern NLP technologies can play a transformative role in content moderation, moving beyond traditional detection methods to proactive measures that promote healthier online interactions. We will cover key topics, including:

  • Understanding the Landscape: Definitions and nuances of harmful content categories, including hate speech, misinformation, and harassment. We will bring practices not only from CS field, but from communication with social scientists and NGOs.
  • Hate Speech Detection: Can LLMs detect hate speech? How the models can be adapted to new languages?
  • Text Detoxification: Diving into nuances of toxicity of 9 languages (from our recent shared task) and sharing best practice on LLMs prompting for texts detoxification.
  • Counter-Speech Generation: Our recent research results on how make LLMs generate not a very general "Please, it is not ok to talk like this report" but indeed address the targeted group.
  • Ethical Considerations: Who, in the end, responsible for the content moderation? How the community can help to bring best practices? How the measure the "effectiveness" of LLMs for content moderation?

This session took place in track Natural Language Processing & Audio (incl. Generative AI NLP) and was classified suitable for intermediate domain / novice python by the speaker.

Transcript (auto)

Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.

Speaker 1 [00:07]

Yeah, so we're going to discuss about how NLP, like modern models, can help for proactive measurements for harmful conduct moderation. And, like, yeah, so the side question will be, can NLM solve it all? So I would like to start with a citation. So understanding the problem is actually half of the solution. And also today at, I think, Alex's talk, he also was mentioning that there is a call for interdisciplinary discussions about like what actually AI is going to help society or in other industries and also like yeah there is a question indeed like how can NLP or AI help for example to moderate content and I actually would have loved to say that we actually already understood the problem but actually it's not so much true and I will tell you like in this talk like what we actually as researchers know so far and what are the challenges and yeah maybe like what also can be some new ground for thoughts and yeah so let's go and so yeah a couple of words about myself indeed my name is Iryna Dementyeva and I'm a postdoc at Techno Western Munich and I'm doing a lot of research in applying NLP for different social good aspects like how we can fight fake news how we can fight toxic speech hate speech also to support my native language Ukrainian and any other underrepresented languages and also like yeah so this year we're gonna have as a first workshop at NLP for positive impact at ACL, like one of the biggest conference in NLP in Vienna during summer. If you want to come, yeah please come. Yeah, so online moderation is actually how it's happening right now. It's first of all exists like there is indeed still hate speech and any other harmful content in their social networks and social platforms are indeed trying somehow to tackle them. However usually what happens there are some detectors, actually not super advanced once and then messages or users are just blocked. However, there's discussions that actually people sometimes don't even understand like why they're blocked or message is deleted and that can even cause even bigger hate, even bigger anger that can then be applied not only in online realm but actually can go unfortunately in even physical realm. So like our idea as researchers like taking into account like what the research is happening around this topic was so maybe we can have like some automatic more proactive moderation so and the key here is indeed automatic because like what is happening right now that yeah it's just blocked because there is no automatic measurements how to do something differently and that can be indeed some warnings explanations so there is a whole huge call for explainable ai so we can somehow explain why something is toxic or hateful or why it's not okay to behave in such way and maybe we can supposedly propose for users to detoxify their message and maybe somehow rewrite it and like just make a suggestion or maybe then write some counter-measurement to what the user wrote and explaining like yeah why did something was okay and like then indeed if the user ignores already like everything was happening then indeed like call for human moderator and then call for authorities or any higher measurements yeah so also in this slice you're going to see further very uh sometimes examples of very very hateful and rude phrases of course it's not to insult you it's just like yeah research purposes we have to read this as well uh yeah but actually what we're talking about like what is hate speech what is toxic speech or abusive speech or like whatever it's like what is happening with all this terminology and uh yeah i mean there are there is social science and actually they indeed thought for some time about this and there is a list of what what should be called what actually i really encourage you maybe to have some time and read also all this terminology what is important for us in this presentation is that we have abusive speech going to be like an umbrella term for all these different kinds of expressions that are not okay. Hate speech, if it's targeted against some group or persons, like there is a target. And toxic speech, it can be just like a general toxic speech with some obscene lexicon, but it's not usually targets like some very specific person or group, like with some insult. And also there is a list of text abuse speech. Yes, of course, there is many, many other things, and also in Germany, there are special NGOs and companies that are taking into account many different aspects of what can really happen in the digital realm. Of course, there is many of these happening with images, with undesired visuals, what you can receive, and there is also, of course, a call to the Texas and somehow to tackle this. Yeah, but when we looked into how we can understand all these terms and actually what is happening in NLP research, but also not even there, but also what's happening with social platforms and also with kind of regulations. We know all about European DCA here, and if there's even any alignment between all of these definitions or regulations or like really what is happening. So we try to have a look on this. Just for statistics, I will tell you that, yeah, we looked into the countries, platforms, and papers from NLP fields, around each 20 per each category. And you're going to see this list of many, many, many, many questions and a lot of text. The idea was to see, indeed, if in the countries, hate speech is defined on the low level. If there is any understanding, like if there is online hate speech, like a distinction between hate speech and online hate speech. What has happened? What are the countermeasurements against this? If there are the same definitions as the platforms. If NLP papers understand that the user is doing some research for language of the platform, if the definition aligns with the original definitions for the platform and from the country. To make a story short, looking at this very long text, some interesting statistics per each category. The main thing to see here is that a lot of numbers are far from 100%, first of all. Indeed, the majority of countries define hate speech. They understand that it exists. However, far from the majority realizes the difference between online hate speech and just hate speech. Also have special case of USA, you know, like what's happening with this country. And yeah, also there is a call, but of course not also every country, but there is anyway a call to somehow have counter measurements against hate speech. Like in whatever form it means, like there is a call for this. In platforms, also not every platform explicitly defines in the guidelines what is hate speech and what is not okay, how to behave in the community as a platform. There are even platforms without any definition about hate speech at all. And yes, there are also not every platform verifies the users, it can be also anonymous in some platforms, it can be really a dangerous platform where you can express yourself whatever you want. And the big unfortunate thing was, like, yeah, if we take into account research papers, we will see that even if a paper scrapped their data, like, it's research that scrapped data from open source, they do not check definitions of the platforms from where they scrapped the data. So they just, like, yeah, define them by themselves or by from another video science papers, but not checking, like, what was really happening in this social platform. And even the last papers then report, like, what annotators then annotate the hate speech, who was their background, who they were at all, and so on. And so it's really like a huge lack of details and a huge lack of alignments between definitions within all these categories. So yeah, from here, it's really called, if you're doing some research, please check definitions, please check with social scientists, like in the field where you're doing your research or further technology, like yeah, please do cross-check with the reality. Yeah, but then anyway, like what is still, We had many data sets, we have many models, open source, like yeah, still what we have. We have some data, especially we have a lot of data for English, if we want to do some classification, and also there is like a very nice page on hate speech data set catalog. If you have a look, there are actually many data sets from many languages, also with very mixed definitions of hate speech, abuse speech, and toxic speech, but yeah, anyway, like it's list is very nice, and you can really, yeah, have a look and download many of them, it's open sourced. But then if you have the data, you can, on one hand, just go and prompt LLMs, or on the other hand, you can just fine-tune steel encoders and use the data set for supervised fine-tuning. In terms of LLMs, you can go to different also size of LLMs. It can be, for example, 25 structure to LLM. It can be mid-sized. We can go to LAMA, to MISRAL, or we can even go to child GPT. So you can see really like resources allocated for this is really very different. But what are the results? So there are several papers for rugged tested LLMs for high-speed classification, and we can see that if we just prompt LLMs, the results are far from being good. Even sometimes it's really like a little bit better than just random, but if you fine-tune the model, just some encoder-based model, it's really already way better results and it's way less resources. Also about strategy for prompting, there are also like you can, of course, you can do it in different ways. You can just do zero-shot. You can just few-shot. can be a case that, for example, you can indeed use definition from the law or from the platform policy and then insert the definition into the prompt. You can also adjust maybe some key symbols for LLM, I don't know, with dashes or whatever, so maybe it also can help. However, even with all these tricks, accuracy is, like, FN score is not super high, it's not higher than 80%, which is at the same level like its encoder-based model. So the takeaway here is the destruction fine-tuned LLMs can be promising, however to fine-tune some transform basic coder is really still in the game, especially if you want to have a solution for many languages. It's better to collect a couple of hundreds per language and still fine-tune the encoders and go to LLM. So another technology can be, for example, to detoxify the text. What it can be? Yeah, so here are some examples. So indeed, if you have a user that wants maybe to post some toxic message in the social network, We can have some warning and say, like, yeah, if you're really sure if you want to post this, maybe you should rephrase this in another way. Still, also, even LLMs can be quite toxic even without the safety measures trying to be taken against. You can still go and unlock toxic persona in LLMs, and it can be potentially like a chatbot in production. It really can misbehave with the user. So then the solution can be maybe we can change the style of the text from toxic to non-toxic. and saving this text and also the style of the original author as much as possible. Of course, we cannot handle all types of toxicity or abuse with this method. We can maybe tackle what I call a very light case of toxicity. If you think about how we can rephrase racism or another very severe hate speech, we actually cannot, because it cannot be done like this. So we can talk about some obscene presence lexicon in the sentence. So for now, actually, we have already a corpus for non-languages, like we have these parallel pairs, and then you can fine-tune a model on this corpus, and again, we have already nine languages, and yeah, this multilingual, quite multilingual setup. Yeah, so we, for example, can fine-tune mt5.xl, or mt0, or mbarst, or any other modern sequence model. But on the other hand, so we had like a shared task, and you can see like here's a list of different LLMs people tried, from LAMA to ChatGPT and so on. And then the question is if these LLMs can really handle this task. Okay, so there is a big table with the final results for all the languages and what we can see. So comparing to how humans perform tasks to human references for resource-rich languages, actually LLMs really can be already quite good. But yeah, because indeed these languages are present in the models, and to some extent it really can be as a human performance. However, if we look into maybe not so super present languages, the performance is really far from good. Also about Chinese, like you can see here, and you can think about, yeah, but actually Chinese has really a lot of data in the internet, like what is happening. So also with Amharic, it's also quite a popular language in Africa. These languages are not like Latin script languages, it's really something different. And for models, it's still super difficult to perform the organization of these languages. So that's why performance. And also we can see that even if for some languages LLMs can be good, there is no single model that performs pretty well for all the languages simultaneously. So multilingual performance is really still a challenge. But also we tried LLMs for explanations, because we wanted to see actually what can be toxic different languages and indeed annotation with humans can be quite expensive so that's why at least we try to generate explanations with LLMs and then cross-check with humans like what was generated. So there was a prompt so we indeed went to this property very weak gtp4 and asked to explain like yeah can you please define what was toxic span as a sentence can you please describe the tones the language and so on and what we found out is actually LLMs can be really very good in explanation. So toxicity detection is still not super perfect, however it can define at least a scale, like yeah, something is less toxic, something is more toxic. And yeah, if we want to describe like more descriptive features like tone, language type, this really can be very nice. And our annotators, native speakers are really super surprised, it can be really very precise. And yeah, so about toxic spans, I just wanted to share with you some interesting findings. Yeah, So if you are a native speaker of any other language, please be prepared. Yeah. So we also tried to see the most popular tokens within toxic words. And interesting thing that indeed there are universal things. Yeah, I've seen lexicon like it can be like some things can be really pretty universal across many cultures and many languages. However, there are really cultural specifics. For example, in Ukraine, Russia and in Chinese, there is still a stigma about the LGBT community in Hindi and Amharic. It's really a culture to compare like foreign cells with animals. In Germany, there is, you can see a huge discussion about refugees and there is a play of words about this topic. So yeah, we really can see the scenes and this is a call that we really need to define cultural specific corpus anyway, even like LLMs can be quite big, but we anyway need culture specific solutions. Yeah, and we have even more languages. So if you want to see our data participate in the shared task, please go ahead and follow the links. Also, these datasets are open sourced and also you're very welcome to check everything online. And the last part is counter speech. So for example, we have not just toxic speech, but hate speech. So again, like some insults against some targeted group. And then there's a question like what we can generate against this hate speech, how we can argument, what we can do. And of course, if you can go to any LLM right now and ask to do this you will get some response but then the question like how effective can be this response and here you can see the example of not so nice counter speech it can be offensive for example of course we don't want to have like offensive response to the message we want something more polite and more calm uh something can be really irrelevant or maybe too generic however if we want to have successful arguments then we want to really targeted, aware counter-measurement against the hate speech. And yeah, so in this case indeed we have to take into account the context, at least take into account the target group which the hate speech is targeting, and even if you can take it out even something bigger, who is the author of the hate speech and who is the receiver of the hate speech, yeah, really take the context as much as possible, it would be great, At least we did some experiments with their targets. In encoder-based methods, you really can have additional vector which you will incorporate into the decoder, and then it generates more target-aware response. And you can also see this table with numbers, yeah, many numbers, but the short story is that we wanted to increase the relevance and to decrease toxicity, and in this case, yeah, if we incorporate target-awareness, we indeed decrease the relevance, sorry, increase the the relevance of our generic counter-speech. And yeah, with LLMs, the scores actually were not super great. However, if then we looked into their examples, and here we can see example of hate speech and then different counter-speech generated by LLAMA3, and we can see, yeah, either it can be super irrelevant and super generic, but if we're asking like, yeah, please really be a super target aware for this hate speech, then yeah, we can have more relevant example. Yeah, so final takeaways. Yeah, if you really plan to do some social-related NLP or a project, yeah, please have a look into definitions and also try to understand the problem as much as possible with law, social science, with ethics and so on. LLMs are not the most efficient solution for everything, but they indeed can be really good explanations. And also if you would like somehow to have a cooperation between human and LLM, it really can be a solution. But yeah, it's still worth the benchmark, whereas baseline especially still transform encoder models. Cross-lingual and multilingual models are still a challenge, and yeah, if you want to have a multicultural model, you really need to take account of different languages and different cultures. Yeah, so the next steps, it would be really great to see indeed some more methods how we can incorporate this cultural knowledge into LLMs more efficiently because now we have, yeah, it was a big choice because it can be fixed. However, maybe we can additionally continue the models and they can be more cultural specifics. And again, finally, all those ideas were from research perspective, and it would be really great to see automatic right of content moderation in the wild in real platforms. Yeah, thank you very much. You will be very welcome to reach out to me and all my socials, and yeah, thank you for your time.

Speaker 2 [18:27]

Thank you, Darina, very much for this wonderful talk. And there aren't any questions yet, so I will start with my own. So given that I'm a researcher myself, I was often enough... I was facing the problem that you described earlier regarding annotators, right? That there's absolutely... It's not just in that domain, it's in many domains. So there's absolutely zero information in the paper about who the annotators were, How many there were the inter annotator agreement seems to be like one existence in most papers? So but very often that's not laziness, but a lack of funding So my first before you like because I was thinking about this question at the beginning and I wanted to ask you how to tackle That problem, but now you gave a solution which is yeah, actually we could ask LLMs, but Then you in some cases, but then like again, so you then give it to people to cross-check But isn't there still a problem because like I might give it to you and you say no I don't consider it a speed you give it to somebody else they say yes, so like you will still have 10,000 opinions But if something is either one or the other

Speaker 1 [19:28]

Yes, it can be still a problem, of course, yes, and in an ideal world, if we have, like, this internal funding for everything, of course, it would be nice to hire every person from every background, from every country to annotate hate speech. Unfortunately, we don't have those cases, but at least, yeah, right now is to find, at least in terms of genders, for example, or at least in terms of ethnicity, like, within your lab. So try at least to do something, not just only wired mail and usage, so at least time diversity will be great.

Speaker 2 [20:01]

Any more questions? Do you know how to use Slido? Okay, that's fine. We have a lot of time. So I will just walk around and hand around the mic. That's okay. We'll make an exception. It was a long day.

Speaker 3 [20:15]

Thank you. Hi. In that table that you showed where the human annotators were a benchmark in the first row, I was wondering why there are so big differences between the countries or between the languages. Yeah, exactly. Like the human references, they vary a lot across language. Do you know why? Like the performance, basically? Yes, like in this...

Speaker 1 [20:36]

Yes, in this specific case, their evaluation also was automatic, meaning that even if the text was written somehow by humans, how toxic or non-toxic was this text was identified by the model, by classifier. Then how similar is text to the original saving the content, it was also cosine similarity between LabC and Bay. Again, it was a model. And of course, again, their equality of scores between languages in those multilingual models is not equal. Like, it doesn't exist because they're pre-trained data sets of different sizes, so we don't know, indeed, if in the space these clusters of language are really equally distributed. We don't know this. And anyway, we know that it's far from being perfect, so that's why it's like the scale for each language is different, so that's why we're providing this code for human references. At least we can see that human text performed like this by these models, and then we can and compare all other models with its references.

Speaker 3 [21:37]

Okay, thanks.

Speaker 2 [21:39]

I think there was somebody else who was raising their hand, because I also got one on Slido, so I'm just thinking whom to... So you submitted this? Ah, okay, cool, then let's read it out. How do you tackle emojis? Nowadays it feels like a common language for young people.

Speaker 1 [21:56]

Yes, actually, it's a good question. Also, I didn't mention in this slide since the beginning, so we also actually have statistics, or at least we tried to catch statistics about from which years were the posts scrapped from the social networks. And it's actually a really super important question, especially after COVID events, especially after war events, Russian invasion of Ukraine, especially what's happening in Palestine and Israel. So again, like in many countries, the mood is very much different and the topics are different. and which emojis I use and also even how emojis I use is also changing. And we didn't specifically look so much into this. However, yeah, like even worse can change the meaning and become hateful after some events. So it's really like, yeah, sounds to take into account, of course. We didn't tackle this, but yeah, it's a field for future research.

Speaker 4 [22:54]

Hi, um, do you know if they already like plans? I don't know to exclude basically hate speech from the training data for those big LLMs or Particularly like in Europe probably yes but do you know if it's already done like in the future in the past and how it's gonna happen because this is kind of like the

Speaker 1 [23:12]

Like the,

Speaker 4 [23:13]

The reason why LLMs produce hate speech because they already seen it

Speaker 1 [23:18]

Yes, actually it's a good question and I think there were try-outs to make it So people just yeah indeed like filter out Some by some classifier again like it's not my normal filtering is by some classifier and they filter out something if it was good or bad Like we don't know it and so it was not something very spectacular At least we know that it wasn't like some safety progress like it's all everything But actually it's also question like if we really should filter it Because if the user will then write some hate speech, and maybe the model still should identify like, yes, hate speech, and I need to write something maybe like, again, counter-speech, like some counter-response to this hate speech. Yeah, so I also can point out here, there was indeed research from Meta to fight hallucinations and to prevent toxic token generation during the inference time. So there was this research from NoLanguageLab behind team, I really can recommend to check it out. So and I think also another problem here indeed like how to identify this toxicity even for tokens or for instances because yeah if we have some classifiers they can give you some F1 score or 95 on some data set if it really covers all existing in the world, maybe not. So still there can be something in the corpora. I think yeah it's a good research question, it still can be another research problem.

Speaker 2 [24:38]

Thank you, and another question. I really like this one Haven't thought of it myself, but now I'm like that's a good one Is there any research of LLM performance of content moderation of politically funded speech in online settings so-called? troll farms Will not mention any country specifically

Speaker 1 [24:58]

Actually, I can't recall any specifically, I don't know, paper, tables that will give you these results for troll detection. I think there are definitely data sets. Again, not recent ones. I think maybe of the year 2017, 2018, about how we can identify the trolls in the internet. And yeah, I mean, so definitely there is research right now how we can detect AI-generated texts, like for sure. but I really doubt there is any open data set on AI-generated social media posts.

Speaker 2 [25:30]

It's also usually not based on what they write, but like other metadata, whether it's from a troll farm or not. Yeah, absolutely. Are there more questions? Because we still have four minutes. Yeah, well, it's been a long day. Yes. Cool, then I don't know then let's finish the session early Darina. Thank you very much

Speaker 1 [25:53]

very much as well so

Speaker 2 [25:53]

There's a round of applause.

Daryna Dementieva

About — in the speaker's own words

Hello, I’m Dr. Daryna Dementieva. Driven by both personal experiences and a deep passion, I am a dedicated advocate and researcher focused on leveraging AI and NLP for Positive Social Impact. Currently (as a technical person) I am exploring collaborations with NGOs and social scientists to bridge the gap between cutting-edge AI technology and societal needs. My goal is to share insights on responsible AI and Data Science, inspiring and enabling projects in these fields to transition from concept to impactful reality.

Social card for talk: Modern NLP for Proactive Harmful Content Moderation