Human data you can verify.

Every dataset traces back to real, identity-verified people - not an anonymous workforce and not a benchmark score you have to take on faith. Built for training, evaluation, and alignment work where who answered matters as much as what they said.

Social reasoning
This dataset that aims to provide signal to how humans navigate social situations, how they reason about them and how they understand each other
Get it on Huggingface
Humaine evaluation
This dataset contains human evaluations of AI model interactions across diverse demographic groups and conversation contexts.
Get it on Huggingface
Autoresearch HITL
This dataset from the study "When does autoresearch need a human?" - a case study running Karpathy's autoresearch on a DPO task.
Get it on Huggingface
Human-AI Team Interaction
This dataset captures interactions from a study on presenting large language models as teammates versus tools.
Download
Longitudinal preference tracking
This dataset follows the same verified participants over time, built to measure how people's preferences and judgments shift rather than treating every response as a one-off.
Request a sample
Representative population sampling
This dataset that aims to reflect a real cross-section of the population - not just the people most available to annotate - for tasks where whose opinion counts is the whole point.
Request a sample
Human feedback, live
This dataset is continuously refreshed with new human judgments as they come in, rather than a fixed snapshot collected once and reused.
Request a sample
Evaluation and safety
This dataset that aims to surface how humans judge harm, refusal, and appropriateness - real people testing the edges of what a model should and shouldn't do.
Request a sample
Domain expert reasoning
This dataset pairs complex questions in law, medicine, and finance with answers from credentialed experts, verified before they're included.
Request a sample
Task compliance
This dataset captures how well real people judge whether a response actually did what was asked - built to catch the gap between sounding right and being right.
Request a sample
Bias and fairness
This dataset captures how people from different demographic backgrounds interpret the same ambiguous scenario - built to surface bias a single annotator pool would miss.
Request a sample
Multilingual reasoning
This dataset that aims to provide native-language reasoning and preference data across multiple languages, collected directly rather than translated from English.
Request a sample
Real-world agent tasks
This dataset captures real people completing everyday tasks with an AI agent in the loop, labeled for where the agent succeeded and where it didn't.
Request a sample

Proven people. Proven data.

Preview a live sample of any dataset before you commit - see the people, the data, and the quality bar for yourself.