Articles

What’s New: Run multi-modal evaluations end-to-end, smart participant filtering, and more

Anisha Osbourne
|August 10, 2026

From building multimodal evaluations in Prolific to a much slicker way to manage your filter sets, this month's updates make it easier to run precise, high-quality studies. Scroll down to explore what's new.

Video is here: run every multimodal eval on Prolific

Video is now live in AI Task Builder, which means you can handles the full range of multimodal evaluations end to end with Prolific. Recruit verified participants and run your text, image, audio or video study in a single workflow.
 
  • Every media type, one task. Present any combination of text, image, audio and video in a single item. Video renders as a native player (play, pause, seek), right alongside your evaluation questions.
  • Side-by-side comparisons for preference evals. Use the configurable two-column layout to put Video / Image A next to Video / Image B for pairwise evals, or RLHF-style preference judgements. Ask questions against each item as participants compare, then collapse into summary questions about the pair (which is better overall, how confident and why). Available in the API and CLI first, coming soon to the UI!
  • Streamlined flow. Split a task into distinct pages so you can build a clear flow for participants, complete with an outro once the task is done. Drop images into your task guidance too: screenshots, scoring rubrics, anything that helps participants understand exactly what to do. Higher participant comprehension means higher quality data.
  • Real responses at API speed. Describe your evaluation and our API handles the rest for you: targeted recruitment from 300,000+ verified participants, study execution and delivery, with a median time to first response of just 4 minutes. Build it into your pipeline via the API or CLI, so no need to tool switch. 
Get started
 
All live now, you can build multimodal studies directly in the UI, or via the API and CLI to drop them straight into your pipeline.
 
Not sure where to start? Talk to our team and we'll set your first multimodal study up with you.

More granular participant recruitment

When you need to target a niche audience, you often have to combine multiple conditions at once, for example if you need participants who are employed full-time or part-time, and within a certain age range.

Now, you can build more complex queries more easily, with up to 5 OR branches in your filter set. This gives you the flexibility to define exactly the audience you need without compromising on precision or efficiency.

Explore in platform or via the API.

A smoother filter set experience

Reusing and managing your saved audiences is now much easier too:

  • Search for your filter set and view the participant count: Quickly find the set you're after, and see the eligible participant count for each one, so you immediately know whether it'll fill your study.
  • Created by and created date columns: See who created a filter set and when, sorted newest first by default, useful if you reuse filter sets often or share across teams.
  • Rename filter sets: Rename or delete a filter set directly from its detail page.
  • Descriptions: Add an optional description to any filter set, so your team knows exactly what it's for and when to reuse it, which cuts down on guesswork or duplicate sets.

Take a look in the UI and the API.

Reach the right participants, faster

Our Network holds strong at 295,000 90-day active participants. During summer, the smartest way to keep studies filling quickly is to target the right way.

Take a look at our quick tips to optimise your participant recruitment:

Filter by language, not just country

If you need English speakers, narrowing to a single country like the US or UK limits you to a slice of the people who'd qualify. Instead, use the language fluency filter and select ‘English’ to reach 267,000+ active, fluent English speakers across more countries to get the same fluency from a much wider pool. The same trick works well beyond English: language filters reach fluent participants wherever they are.

Audiences ready for your next study

Some of our deepest pools are wide open right now, so studies targeting them fill fast:

  • South African participants: a sizeable, active pool with plenty of availability
  • Verified Spanish and Dutch speakers: thousands of active language Experts, ideal for localised or language-specific studies.

One verified Expert Network

Our Expert Network now spans 90,000+ verified Experts across language, healthcare and coding, each individually verified for their skills or credentials. This month brings fresh depth in emerging languages like Thai and Punjabi, alongside established pools such as German (1,470+) and Mandarin (1,130+).

Not sure who's on Prolific?

Use our Audience Finder to check your audience, or find specialists in-app under Participants > Add screeners > Domain Experts.

To find them under the API, take a look here.

Watch: The latest episodes in the Frontier Series

In our Frontier Series, we interview frontline AI researchers who build the systems, run the studies and explore the questions that don't yet have clean answers.

Catch up on the latest episodes recorded at this year’s ICLR:

Episode 4: LLM Training as Lossy Compression

Prolific's Nora Petrova and John Burden sit down with Henry Conklin (Princeton/Cohere) and Tom Hosking (Cohere), to unpack their paper on treating LLM training as a form of lossy compression. They dig into the two-phase pattern they observed across model sizes during training, and why smaller models may fundamentally struggle to compress enough to generalise well.

Watch Episode 4

Episode 5: Inside the Evaluation Problem

The conversation continues with Henry and Tom exploring what their compression framing means for evaluation: why compression quality might be a more meaningful signal than benchmark scores, and what that implies for how we train, evaluate and select models going forward.

Watch Episode 5

Meet the team in person 

26 August: Fine-tuning the agent stack, London

Join Prolific and AI Circle to hear from technical leaders on how teams are fine-tuning agentic systems and evaluating them under real-world conditions. We’ll examine how to design stronger agent harnesses, build evaluations that capture long-horizon behaviour and close the gap between prototypes and dependable systems. Register here.

10 September: A New Era for Data Quality in Online Research, Online

With MTurk closing to new customers, a lot of researchers are rethinking how they approach online recruitment and data collection altogether. Join us for a fireside conversation grounded in real evidence. We'll cover an independent, peer-reviewed PLOS ONE study comparing data quality across major research platforms, co-authored by Dr. Andrew Gordon. Register here

That's all for our August update! See you next month.

The Prolific team