Grounding Toxicity in Real-World Events across Languages

Research output: Working paper / PreprintPreprintAcademic

9 Downloads (Pure)


Social media conversations frequently suffer from toxicity, creating significant issues for users, moderators, and entire communities. Events in the real world, like elections or conflicts, can initiate and escalate toxic behavior online. Our study investigates how real-world events influence the origin and spread of toxicity in online discussions across various languages and regions. We gathered Reddit data comprising 4.5 million comments from 31 thousand posts in six different languages (Dutch, English, German, Arabic, Turkish and Spanish). We target fifteen major social and political world events that occurred between 2020 and 2023. We observe significant variations in toxicity, negative sentiment, and emotion expressions across different events and language communities, showing that toxicity is a complex phenomenon in which many different factors interact and still need to be investigated. We will release the data for further research along with our code.
Original languageUndefined/Unknown
Publication statusPublished - 22 May 2024

Bibliographical note

Paper accepted for at The 29th International Conference on Natural Language & Information Systems (NLDB 2024)


  • cs.CL

Cite this