Articles

The new era of data quality in online research

Dr Andrew Gordon
|September 29, 2026

On September 30, Amazon Mechanical Turk is shutting down for good, which makes it time to say a final goodbye to one of the first platforms that made large-scale online research possible.

When MTurk announced that it was closing its doors to new customers on July 30, I wasn’t surprised - I’ve been following the sharp decline of data quality on the platform since 2018. Its closure is symptomatic of a wider shift in the research landscape, where increasing fraud and the ever-present issue of inattentive, dishonest, or outright lazy respondents have redefined what data quality means today.  

We find ourselves in a new era of online research. But how did we get here? And how does this new era change the way panels and researchers approach vetting, trust, and data quality? 

The three eras of online research

Since online sampling began over twenty years ago, the way we look at and assess data quality has evolved dramatically. There are many ways that we could break up this period into defined ‘eras’, but I tend to err on the side of simplicity with just three: scale, curation, and verification. 

1. The era of scale (2005-2014)

We begin with MTurk in 2005. MTurk started as a marketplace for microtasks and wasn’t actually built for researchers at all. Most tasks on the platform were commercial, like data cleaning, image tagging, and content moderation, posted by businesses rather than academics. But academics quickly realized it solved one of the biggest constraints in human research: access to large and diverse participant samples. 

The model MTurk used was an open door. Anyone could join, everything was anonymous, and the focus was squarely on volume. This was new territory to academics, so a field of research quickly sprang up seeking to validate whether such an approach was reliable. The results showed that it was. Some major early validation studies, like Paolacci et al. (2010) and Casler et al. (2013), found comparable effect sizes between MTurk, university, and social-media samples, while Behrend et al. (2013) provided one of the most highly-cited guidance papers on the approach. This work gave the field license to move online. 

It’s worth noting that industry research was already well on the way to this shift by this point. In the late 1990s, driven by collapsing telephone response rates, commercial research had moved towards opt-in online panels. This same period was the maturation phase in industry where panels began to consolidate, and the first programmatic sample exchanges were built. Academia was not a pioneer in online data collection; they were just catching up to the industry trend. 

2. The era of curation (2014-2020)

Around 2014, the philosophy moved away from the fully open, free-for-all model towards managed pools with more built-in ethical standards. During this period, more academically-inclined, curated platforms (such as Prolific and CloudResearch) began to appear to offset what was seen as a major weakness in the MTurk approach: the lack of vetting. While Prolific approached the problem with a bespoke respondent pool, Cloud Research (at the time TurkPrime) opted to build a vetting layer on top of the already created MTurk pool. 

As these new panels emerged, the research focus shifted from in-person vs. online sampling to how online samples compare to the established MTurk benchmark. For a while, MTurk still held its own. Peer et al. (2017) found its quality nudging ahead of Prolific in its early days, and Kees et al. (2017) rated it best on quality metrics, at around 75 cents per quality response when compared to industry and in-person samples. Nonetheless, it was clear that the newer panels were already holding their own. 

Things changed dramatically in 2018 with the "bot panic", the observation that data quality on MTurk had dramatically deteriorated. The name of the panic is somewhat of a misnomer, as it wasn’t just bots impacting the platform; it was also more generic fraud such as click farms and VPN abuse. The openness that had been MTurk’s strength (open access) became its biggest weakness, as it lacked the controls to prevent fraudulent activity. By 2020, it appeared the writing may be on the wall for MTurk, with highly cited studies such as Chmielewski & Kucker (2020) appearing to have great influence. 

As the academic platform wars began to rage, industry panels were taking a different path. While the academic panels prioritized curation and direct participant relationships, industry panels like Dynata and Cint optimized for scale and volume through aggregation and routing between surveys to fill quotas rapidly. This model of scale at all costs was driven by specific demand from industry researchers, and the industry panels optimized accordingly. There was now a very clear split in terms of the panels researchers of different stripes would choose to use. 

3. The era of verification (2020-Present)

Now we reach the era we currently live in: the verification era. 

When the COVID pandemic struck, demand for online research exploded almost overnight. Labs moved their research online, new platforms emerged to meet the demand, and as demand increased, so did issues related to fraud and data quality. 

Given this increased use of online panels, it was only natural that researchers sought to put them under the microscope. We therefore saw more research turning towards comparing these providers in detail. 

By the time my colleagues and I published a comparison of MTurk, CloudResearch, Prolific, Qualtrics, and Dynata in 2022, the findings from those influential 2017 papers had essentially inverted, with Prolific at the top and MTurk near the bottom. And this wasn’t just a one-off finding. A methodologically similar study by Douglas et al. (2023) showed the same thing, perhaps in even greater detail. They found that only 26% of MTurk participants qualified as high quality, while it was 62% on Cloud Research and 68% on Prolific. Another study that got a lot of traction at this time, Webb & Tangney (2024), was even more damning, finding that out of a sample of 529 participants on Prolific, only 2.6% could be verified as genuine human responses. The writing was on the wall for MTurk. 

Conversations around data quality and AI have also dominated the last few years. The AI threat has reignited a panic similar to the bot scare of 2018, this time centered on LLM-assistance and LLM-powered bots. In either case, the question we're asking has changed from "is this data good?" to "are these responses even human?" 

It is in these last years that we saw the final nails in the MTurk coffin from two published papers. Stagnaro et al. (2026) found that MTurk was lowest in terms of response validity compared to eight other providers. And in a shameless self-plug, Gordon et al (2026) found that not only did MTurk have the lowest overall data quality score out of ten providers, it also had the highest cost per quality response and the highest bot pollution. 

This brings us to today, with MTurk permanently closing its doors as the model of unvetted, large-scale crowd sampling comes to an end. The original Mechanical Turk, which gave the platform its namesake, was a chess-playing robot secretly operated by a human. Ironically, we now face a landscape where humans can potentially be powered by machines. In many ways, this inversion is what led to MTurk’s demise.   

In this new era of data quality, rigorous verification has begun to underpin everything to ensure research findings can be trusted. But who is responsible for this? And how do we build the tools to ensure researchers can get high-quality data from real people?    

Vetting and trust: Whose job is it?

To answer this question, first we need to define what we mean by “vetting.” 

There are two kinds of vetting. One is the verification that someone is a real, unique person before they even take a study. In my opinion, this should be the provider’s responsibility - after all, providing real people for online studies is the entire premise of their business (well, most of them). 

There are many measures that providers can (and likely should) put in place to ensure participants are real and who they say they are, including ID verification, blocking VPNs, and ongoing fraud detection. It is, however, worth being honest about the potential drawbacks of heavily vetted samples, as it can have an impact on the representativeness of the final pool. We must remember that opt-in samples are already not analogs of the general population, and by requiring IDs we may be reducing that representativeness even further - but that’s a discussion for another article. 

The second kind of vetting is study-specific. A provider can confirm that a participant is a real, unique person, but it cannot confirm that they are the right person for your study or that they will give you data you can actually use. That is the researcher's job. 

Providers have limited (or even no) visibility at the point a study begins, so the study itself has to do the checking. That means building in demographic checks and screening questions that test whether the participant actually holds the qualifications your sample requires. It also means designing the study so that a genuine, qualified participant can succeed. Participants need the tools to answer, and they need to understand what is being asked of them. If they fail on either count, the resulting bad data is a design problem, not a fraud problem. 

Data quality: Whose job is it?

Vetting gets the right person into your study. Data quality is about what they do once they are there. A perfectly genuine, qualified participant can still fail attention checks or rush through your study, so passing vetting guarantees nothing about the data itself. 

As with vetting, I believe that the answer to who ensures data quality is once again both the provider and the researcher. One of the clearest indicators of this I have seen recently was from RepData’s ESOMAR paper, where they showed that across roughly 13,000 participants, pre-survey screening removed about 16% of low-quality respondents, and in-survey detection removed another 17%, with only about 4% overlap. 

What does this mean? It means that both the provider and the researcher can catch low data quality, and they often aren’t catching the same people.  

The provider's job is to build a pool that behaves well even when nobody is watching, because once a participant enters a study, the provider usually cannot see what they do. Some of this works through rewards. Ethical pay and fair treatment give participants a reason to stay in the pool and protect their standing in it. Some of it works through penalties and monitoring, such as bans following study rejections, regular pool audits, and bot detection. 

Providers can also support researchers directly, whether through guidance on which quality checks to use or through purpose-built tools like SENTRY or Prolific’s Authenticity Checker. Done well, this produces a pool in which most participants are motivated to give good data. Done poorly, you get MTurk. 

The researcher holds the other half of the responsibility, as it’s the half a provider can’t reach. However good the pool is, the provider hands over participants, not data. So what those participants produce depends on the type of study they encounter, meaning researchers need to design with detection in mind, building in attention checks, comprehension checks, and open-ended questions (among others) that make low-quality responses visible. 

It also means designing studies that participants can succeed in - bad instructions, unclear requirements, and overlong or boring studies will degrade the responses of even the best participant. A good participant in a bad study gives you bad data just as surely as a bad participant in a good study does.  

Designing trusted platforms that work for researchers and participants

At the center of all of this lies an interesting tension: what works well for a researcher may not work well for a participant. When discussing vetting and quality, we mustn’t forget how this affects the participants, since the participant experience directly impacts data quality. 

Put it this way, you could technically maximize scrutiny in your study by having a camera on the participant at all times, an attention check on every page, and a comprehension check after every statement. As a researcher, this would be great because you end up with a set of data you have absolute trust in. But it would ruin the experience so thoroughly that nobody would take the study, and the people who did would give worse data out of sheer frustration. There's a trade-off between having enough checks to trust the data and enough goodwill left over to get good data in the first place.

To put it simply, you won’t get good data if the platform only serves one side of the researcher-participant relationship. A platform built purely for researcher control, with no thought for the participant on the other end, will eventually degrade the very data it's trying to protect. It is no coincidence that the platforms that invest the most in ethical treatment of their participants also end up at the top of most quality benchmarks. 

A new baseline for data quality

The MTurk model was never going to survive the era of verification we find ourselves living in today. It's tempting to treat it as an isolated story about one platform's decline, but behind it there’s a more important trend we need to pay attention to. 

The bar for data quality in online research has moved, permanently, and it's still moving. Verification, trust, and participant experience are the core pillars that will ensure findings are based on genuine responses from real people. And neither panels nor researchers can shoulder this responsibility alone. Good data requires effort from both sides, with multiple measures working in tandem to close as many quality gaps as possible. 

I recently discussed this topic in more detail with Rebecca Richards, Prolific’s Senior Project Manager. You can watch the full conversation here.