See what we can build.

These samples show the range and quality of data we can deliver, built by real people. Every engagement is scoped to your specification. Get in touch to see full examples.

ALIGNMENT & PREFERENCE

Tracking judgments over time

Preferences shift over time. This sample tracks the same verified participants across multiple rounds to show exactly how - proof we can capture drift, not just a single snapshot.
Request sample
POPULATION COVERAGE

A real cross-section of people

This sample reflects a genuine cross section of the population, not just the people most available to annotate - proof of representative judgment for tasks where whose opinion counts is the whole point.
Request sample
LIVE FEEDBACK

Continuously refreshed judgments

New judgments arrive continuously in this sample, drawn from a live feedback pipeline - evidence we can deliver an ongoing signal, not a single static batch.
Request sample
EVALUATION & SAFETY

Harm and refusal judgments

This sample shows real people judging harm, refusal and appropriateness in model outputs - proof of the safety evaluation judgment we can build around your team's standards.
Request sample
DOMAIN EXPERTISE

Specialist-level reasoning, verified

This sample shows how we source expert answers to complex questions in law, medicine, and finance - evidence of the specialist bench we can draw on, not crowd consensus standing in for expertise.
Request dataset
INSTRUCTION FOLLOWING

Task compliance

This sample judges whether a response actually did what was asked - the gap between sounding right and being right, made visible.
Request dataset
EVALUATION & SAFETY

Bias and fairness

People from different demographic backgrounds judged the same ambiguous scenario in this sample - proof we can surface bias a single annotator pool would miss.
Request sample
MULTILINGUAL & CROSS-CULTURAL

Native language reasoning

This sample was collected natively in each language, not translated from English - evidence of reasoning and preference data with no translation layer in between.
Request sample
AGENT & REAL-WORLD TASKS

Where the agent succeeded - and didn't

Real people completed everyday tasks with an AI agent in the loop for this sample, labeled exactly where it succeeded and where it didn't - proof of the real-world signal we can capture.
Request sample

Published and open to everyone.

Research we've run and published in the open - download, cite, and build on it directly.

OPEN ACCESS

Social reasoning

This dataset that aims to provide signal to how humans navigate social situations, how they reason about them and how they understand each other
Get it on Hugging Face
OPEN ACCESS

Humaine Evaluation

This dataset contains human evaluations of AI model interactions across diverse demographic groups and conversation contexts.
Get it on Hugging Face
OPEN ACCESS

Autoresearch HITL

This dataset from the study "When does autoresearch need a human?" - a case study running Karpathy's autoresearch on a DPO task.
Get it on Hugging Face
OPEN ACCESS

Human-AI Team Interaction

This dataset captures interactions from a study on presenting large language models as teammates versus tools.
Download it here

Proven people. Proven data.

Real people complete everyday tasks with an AI agent in the loop. We label exactly where it succeeded and where it didn't.