#869 · Primary category: Education & Research
hate-speech-and-offensive-language
Repository for the paper "Automated Hate Speech Detection and the Problem of Offensive Language", ICWSM 2017
Project last updated:06/12/23
GitHub Stars
848
Forks
329
Contributors
2
License
MIT
Why we included this project
This repository holds the labeled Twitter dataset behind the Davidson et al. 2017 ICWSM paper, which separated hate speech from merely offensive language, a distinction that still informs how moderation systems are built and evaluated. Researchers and ML engineers working on abusive-language detection get the data as CSV or pickle files, plus the lexicon and a classifier script that reproduces the paper's three-way classification. The bundled notebook shows how the labels were derived, which helps you understand the dataset's assumptions before training on it. Note that the code targets Python 2.7 and is no longer maintained, so treat this as a data and methodology reference rather than a drop-in production pipeline. The authors' follow-up paper on racial bias in this and similar datasets is also worth reading before you rely on the labels.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
prompts.chat
f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
JavaGuide
Java Interview & Backend General Interview Guide, covering computer fundamentals, databases, distributed systems, high concurrency, system design, and AI application development.
system-prompts-and-models-of-ai-tools
A curated collection of system prompts, internal tools, and AI models from popular AI assistants and coding agents.
30-seconds-of-code
Coding articles to level up your development skills
generative-ai-for-beginners
21 Lessons, Get Started Building with Generative AI