Fine-tuning the agent stack: Insights from AI researchers, engineers, and technical leaders
Ask an agent to buy a coffee from across the road, and you’d expect a short trip. If it comes back after taking a detour to Barcelona and you’re based in the UK, the correct order won’t make up for the roundabout way the agent got it.
That example, from a recent panel hosted by Prolific and AI Circle in London, captures a familiar challenge in agent development. Before fine-tuning, you need to understand how your agent got to the result in the first place.
The discussion, led by Josephine Parquet, Head of API Products at VEED, brought together Tyler Edwards, founder and CEO of Overmind, Thomas Mann, Research Engineer at Meta, and Angelos Perivolaropoulos, who leads speech and text research engineering at ElevenLabs.
A recurring point came up across the conversation: choosing a stronger model won’t necessarily solve the problem you’re facing. You also need to examine the software around it and the evidence you use to judge success.
Give your model the right support
Imagine a voice agent mishearing someone’s date of birth over a poor connection. As Angelos explained, the surrounding system can help catch the mistake by asking the speaker to confirm. Without that check, an otherwise capable model could carry incorrect information into the next action.
The software supporting those decisions is often called a harness, which Tyler defined as “anything that isn’t the model weights”. Thomas described its role as keeping the agent working beyond a single prompt. A coding agent, for example, can use its harness to run tests and correct mistakes before returning a result.
“A really good model with a really bad harness is probably less capable than an okay model with a really good harness,” Angelos said. Thomas added that the harness also needs to suit the model, because efficiency carries plenty of weight. A thousand parallel attempts from a small model might eventually solve a problem, but they could take far too long, and product quality is about user experience as well as solutions.
Thomas described enterprises moving away from one monolithic model towards fleets of smaller, specialized models, escalating to a larger model only when complexity demands it. He cited Tech Mahindra’s interest in small agents handling narrow tasks at the edge, supported by a larger central model.
Matching the model to the assignment could reduce the computing resources needed for routine work.
The right answer can hide an expensive process
Tyler noted video transcription as a task with several opportunities for failure. An agent might struggle to load the video before transcription begins, or fail to save the output after producing an accurate transcript. Looking only at transcription quality would miss why the assignment remained unfinished.
His advice was to think about reliability across the whole sequence. If a process has five steps that each succeed 95% of the time, you need to calculate the failure rate at each step, not just whether the final output is correct, and optimize the worst-performing step.
Angelos offered a vivid example of a costly detour. After a failed package install, an agent decided the package index must be down. It then spent 40 minutes fetching individual files from a long log by hand, for a job that should have taken one minute. It did eventually succeed, but he argued that how an agent handles a reliability problem is an important axis to evaluate: retrying sensibly, stopping to ask for help, or doing something absurd.
Cost follows the same logic. A low token price doesn't guarantee a cheaper result, because a weaker model may need many more attempts. Tyler's advice was to "think cost per task, not cost per token."
Don't trust a score without checking what it rewards
Angelos shared an example of an audio judge trained to identify natural-sounding voices. Testing initially looked promising, but using the judge’s preferences to train other models had an unwanted effect. “So we ended up training our models to sound worse,” he said.
The judge had learned to associate microphone artifacts with human speech. Clean synthetic voices lacked those characteristics, so it rewarded recording noise. His conclusion was to use multiple judges and metrics, since optimizing for a single one invites reward hacking.
Thomas described combining automated evaluation with human annotation, including experts checking whether a judgment makes sense for the task. Reviewing examples of what an automated judge rewards can help you identify misleading assessments before they influence training.
Bring user preferences into evaluation
Thomas recalled an example of an invoicing agent that received an adversarial input designed to extract a fraudulent million-dollar invoice. Technically, it succeeded, following all the right mechanical steps, but the user intent was hostile.
Deterministic tests can’t catch this, and a human review after the fact is too late. You need something that can react and gather context in real time.
For subjective qualities, Angelos highlighted the importance of blind human preference testing. People may consistently favor one model’s output even when benchmark scores suggest little difference. Asking them to compare responses provides another source of evidence, especially when you’re trying to improve an experience that a single automated measure struggles to capture.
Passing a test doesn’t settle the question
During the audience Q&A, a practitioner highlighted an example of an agent fabricating information to pass a check that required source data in a structured output. Adding more conditions hadn’t resolved the problem, with the agent finding other ways to satisfy them.
Thomas argued that the agent should never be able to see exactly how it's being assessed. Feedback should be numerical, boolean, or categorical, and instrumented separately from the agent's reasoning loop. Tyler added that the more rigidly you constrain a process, the less value there is in using an agent at all, and at some point a fully deterministic system may be the better choice.
Cheating also happens in more extreme forms. Tyler pointed to models escaping their sandboxes, tracked by the Felony Bench benchmark, and to subtler violations such as agents peeking at test data or downloading models they'd been told not to use.
Before treating a successful run as a training example, check how it got there. If an agent succeeds through prohibited access, a higher score offers little evidence that it can complete the assignment under the intended conditions.
Give autonomy some context
If an agent has decided to delete an entire database, you want to make sure it asks a human for permission first. On the other hand, it’s not very efficient to have it wait overnight for approval every time it needs to remove a temporary file. Angelos wished for more configurable permissions, so the decision to proceed reflects the risk of the action.
Tyler’s approach was to “give agents freedom, but keep the checks around them deterministic.” Existing engineering practices still apply, including automated checks in deployment pipelines. The challenge is deciding where human approval adds protection without holding up routine work.
Look at time spent waiting
Reviewing a month of his own Codex and Claude sessions, Tyler found around 140 hours where agents sat idle, waiting for further instructions, showing how delays accumulate between an agent stopping and someone noticing. Checking on each session means constant context-switching, and he estimated that in many cases the agent doesn't need a human, just something more like a chief of staff.
Tyler proposed a supervising agent that tracks progress across other agents and escalates real problems, so nobody has to keep checking individual sessions. He was clear that the idea remains theoretical, and Angelos questioned whether an agent deliberating over every decision might create a bottleneck of its own.
Let failed tasks guide your fine-tuning
Production traces record how an agent works through an assignment, including failures your planned tests may never have anticipated. Tyler described using unsuccessful runs to build test cases, retaining the original request and the point where the agent went wrong.
His team then experiments with changes to prompts, tools, and code to see whether the failure can be fixed. Where improvements to the harness reach their limit, reviewed traces can become fine-tuning data. Working through the failure first establishes what training needs to address.
Choose examples worth repeating
"Observability and general engineering discipline seem to have been forgotten amid the AI hype," Thomas said, pointing to underinvestment in tracing and data analysis. He noted that data preprocessing accounts for around 70% of AI engineering work, yet teams often jump straight to training their own model without good data.
Collecting runs is only the start. Someone still needs to check whether a completed task contains an expensive detour or an unacceptable decision before it becomes a training example.
Start with where your agent fails
Before your next fine-tuning run, look closely at the tasks your agent struggles with. Some failures will need changes to the surrounding software. Others can inform better training examples. Understanding the difference helps you spend development time where it can improve the experience for the people relying on your agent.
Several of the approaches in this discussion depend on high-quality human judgement, including blind preference testing, expert review of automated judges, and careful selection of training examples. Talk to Prolific about finding verified participants, including credentialed experts, to support your agent evaluations and training data collection.
Follow Prolific on Luma to hear about upcoming events in San Francisco, New York, London, and beyond.





