When AI evaluators disagree: Insights from Prolific and AI Circle’s NYC panel
An evening on expert judgment, the limits of agreement, and what it takes to evaluate an AI agent.
Who should get the final say on whether an AI model is doing a good job?
That question opened the technical panel at Prolific’s community meetup with AI Circle on a beautiful evening in Midtown Manhattan. AI Circle founder Albert Chun moderated the discussion with experts Matt Lincoln (Sr. Software Engineer at Prolific), Yihsuan Hsu (Applied Scientist and AI Research at Amazon) and Pier Paolo Ippolito (Forward Deployed Product Manager at Google), among sushi, pizza, drinks and a great crowd of technical AI leaders.
Their answers took the conversation from expert verification and data labeling to enterprise agents and the future of human evaluation. They also exposed a recurring difficulty: people can assess the same output, disagree about its quality, and each have a reasonable explanation.
For teams building evaluations, those differences deserve attention.
Who is the first person researchers should turn to judge a model?
Yihsuan made a case for the end user. He argued that the people relying on a model to complete a task should have the strongest say in whether it works or not. When teams can’t get feedback from those users, they need to consider how well their evaluators understand the people they represent. “Humans should never be taken out of these equations. We’ll hear human in the loop from here on out,” said Yihsuan.
Matt went one step further, distinguishing between kinds of judgment. Users may be best placed to assess cultural fit or interaction preferences, while technical correctness and safety can require specialist knowledge. He also questioned the assumption that a disagreement must immediately produce a winning answer. It might reveal that the rubric, or the question behind the evaluation, needs to change.
Pier added that selecting evaluators determines who gets a voice in shaping the model. People who experience a system’s consequences need a way to inform its development. Later, he illustrated the point with a restaurant. To paraphrase: A diner, a chef, and a safety inspector can visit the same place and assess it differently. Their responsibilities and experience lead them to notice different things. An evaluation needs to be clear about which of those perspectives it is asking for. “When you think about evaluation, you are deciding who is in or out of the room - who you give voice, who you don’t,” said Pier.
Knowing the subject is only the beginning
Asked what it means to verify an expert, Matt described several sources of evidence, including professional credentials, certifications, and tests of relevant knowledge. But verification has limits. “Domain expertise is necessary for this work, but it's not sufficient for the work,” he commented.
Someone may recognize a good answer without being able to explain their judgment in a way that others can consistently apply. Writing rubrics with the right criteria to assess an output is a skill that needs attention in its own right.
Yihsuan mentioned that he encountered this in his previous work at Amex. Experienced customer service staff understood the conversations they were labeling, but the team still needed to identify who could explain that knowledge and help colleagues improve. An expert committee worked with researchers and compliance colleagues to align the guidance with their different requirements.
Albert later posed a related scaling problem: how do you take knowledge held by one specialist and use it to support a much larger data collection effort? Matt’s answer began with the rubric. Help the specialist break a complex judgment into distinct criteria, using clear yes-or-no checks where appropriate. That gives other evaluators something they can learn and apply.
Consistency alone, though, does not establish quality. Yihsuan warned that unusually high agreement can reflect a dataset dominated by one category or questions that are too easy to teach the team much. Difficult cases near a decision boundary may reveal more about where a model needs work.
Make evaluation clear without making it automatic
A discussion about user experience brought out one of the evening’s most useful distinctions.
On the evaluator’s experience, Matt remarked that teams can recruit qualified people and still get disappointing feedback because the task is confusing. Instructions may be insufficient, excessively detailed, or awkward to use alongside the work. When that happens, he argued, researchers need to read the responses and inspect the study design before deciding if the evaluators lack expertise.
Pier approached the issue from the perspective of someone using an enterprise agent. He worried that a smooth, agreeable experience could encourage people to accept an output too readily. In an investment-related workflow, he described actively seeking evidence that would disprove the agent’s conclusions. Matt clarified that they were discussing different users: the person working with the agent and the person judging its performance. Pier’s concern nevertheless extended to evaluation. A task that becomes too routine can encourage people to stop thinking carefully.
One solution proposed by Matt was a prompt within a task that reminds evaluators when they have missed a criterion. Those reminders can support completeness while leaving the judgment itself to the evaluator. The challenge is practical. People need to keep track of a demanding rubric, remain attentive, and finish the work without burning out.
Human judgment remains part of automated evaluation
On LLM-as-a-judge, Yihsuan saw human input and automation as closely connected. People help establish the standard and calibrate the judge. Automation extends evaluation to volumes of work that would be impractical to review entirely by hand.
Pier then raised concerns about culturally nuanced assessments. An automated judge that smooths differences into a conventional answer can lose information worth preserving. He also pointed out that model judges can differ according to their training, instructions, and context, just as people bring different perspectives to a task.
That leaves teams with a question automation can’t settle for them: whose judgment should the system reflect? The discussion also acknowledged that researchers sometimes have to make a choice despite unresolved disagreement. During the audience Q&A, Yihsuan noted the need to examine evaluators’ reasoning, document the decision, and learn from it. There is no universal method that removes that responsibility.
Watch what the agent does along the way
Pier used a driving test example to explain his approach to agent evaluation. A theory test establishes whether someone knows the rules. A practical test checks whether they use the brakes, lights, and other controls appropriately as events unfold.
Agents need that second kind of assessment. A plausible final answer can conceal problems in the actions and decisions that produced it. Evaluators need to examine the trajectory: what happened, in what order, and with which tools.
Earlier in the panel, Matt described a related problem around collecting evidence from desktop tasks that can take hours. His team at Prolific was working on capturing screen recordings, keyboard and mouse activity, and spoken explanations of the work.
He described a browser-accessible virtual desktop designed by Prolific to make participation easier and avoid requiring people to record activity on their own computers. For complex workflows, the quality of the evidence depends partly on how manageable the collection process is for the person doing the task.
Ask what the benchmark helps you improve
An audience member asked the panel how you can tell if a new benchmark provides useful information beyond another performance score.
Yihsuan focused on finding weaknesses. Start with the capability being measured, identify where the model falls short, and construct tests that expose those failures. A harder evaluation may produce worse-looking numbers while providing better guidance for improving the product.
Pier was skeptical of treating any single score as a complete account of capability. He described benchmarks as directional evidence and cautioned against optimizing for metrics that become ends in themselves. Yihsuan added that benchmarks can have a short useful life. As models improve, teams need to expand coverage and complexity to keep measuring the things that matter.
The discussion gave teams a practical way to assess an evaluation, considering what its results would help them change.
Where human evaluation goes next
Before the audience Q&A, Albert asked the panelists to look three to five years ahead and give their predictions and hot takes on where human evals are going.
Pier anticipated more AI-generated content being created for other systems to consume. He raised questions about distinguishing human and synthetic data in that environment and “dead internet theory.” Yixuan expected new rules and infrastructure as the assumption of a human behind online activity becomes less dependable.
Matt anticipated evaluation moving toward more abstract questions about judgment and intention. He saw cultural and philosophical differences as enduring challenges, with no clear method for resolving them.
Those forecasts remain open questions. But they illustrate a question and concern that many people had voiced throughout the evening. As tasks get more complex, evaluators will need to explain their reasoning in greater depth. Teams will also need to stay willing to revise their criteria when that reasoning reveals something the original evaluation missed.
If you’re interested in attending Prolific’s next in-person AI event, follow us on Luma. Or to scope out an eval for your AI model or any other new studies, reach out to our team here.






