The World's Best Toxicity Dataset

Saving the internet is fun. Combing through thousands of online comments to build a toxicity dataset isn't. That's why we're creating the world's largest dataset of social media toxicity — so you can skip the slog and get to work.

Dataset Preview

Built by an Elite Workforce

Surge AI is a data labeling platform and workforce. Our data labeling team pored over tens of thousands of social media comments to build this toxicity dataset. Each comment was then evaluated by multiple members of our team to determine its severity level.

Other Datasets

InstructGPT-style Dataset
Build state-of-the-art large language models, in the style of InstructGPT and ChatGPT.
RLHF Dataset for Reinforcement Learning with Human Feedback
Build state-of-the-art AI by training your large language models on human feedback.
Japanese Hate Speech, Insults, and Toxicity Dataset
A dataset of online comments in Japanese that contain hate speech, insults, and toxicity.
Dataset of Search Queries and Intents
This dataset contains search queries, as well as the user's intent when performing the search query.
Google Search Quality Dataset
This Google Search Quality dataset contains search queries, intents, result URLs, and a human-evaluated rating.
Search Evaluation Dataset
This search evaluation dataset contains search queries, the intent behind each search query, result URLs, and a human-evaluated search quality rating.
Get notified

We're Launching More!

Thanks!
Oops! Something went wrong while submitting the form.

Love language?
So do we.

We're a team of engineers and researchers from Google, Facebook, Harvard, and MIT. We're building the modern data labeling infrastructure needed to power the next wave of AI.

Our data labeling platform and data labeling teams help AI companies around the world solve their core machine learning and language problems — from detecting hate speech and categorizing user reviews, to training powerful language models.

Our team comes from

Meet the world's largest
RLHF platform