
I spend the majority of my time running evaluation experiments on agents to figure out a few things: How are they performing under various conditions? Where are they failing against what’s expected? Where are they succeeding against what’s unexpected? How can we objectively measure performance and improvement?
To do this effectively, I first have to answer a foundational question: how do you build an experiment that lets an agent improve at an existing or new task? The goal is to be able to say with confidence when a change to a prompt/harness/environment actually improves the agent’s performance. Easier said than done, of course, but the actual granular work required in doing this does look similar across applications, generally following the steps of:
- Dataset curation and labelling
- Task and environment shaping
- Reward and verification defining
- Running and adjusting based on results
Which at a high level doesn’t sound too complicated, but, when breaking it down, the act of creating a representative, useful, and effective evaluation environment takes a lot of careful consideration and effort. So I wanted to discuss the importance of crafting evaluation environments (more commonly understood as benchmarks) for agents, how they propagate further to optimizations like language model post training, and document my learnings from building eval environments for productionized agents at scale.
Firstly, what am I referring to when I talk about an environment? You’ve probably heard of the hype around building “RL environments” popularized by the likes of Prime Intellect (and corporatized by the likes of Mercor/Scale AI/Snorkel). These environments (we can drop the RL from this as that’s more a reference to the consumption of the environment) are developed to accurately mock a scenario that an agent/model would be operating in and score its performance on a related task. While this seems benign, the nuance of all the different variables and integrations needed to make an accurate environment is where the work lies.
At a high level, evaluation environments are made up of a few core components:
| Component | Description |
|---|---|
| The Harness | Tooling and prompting given to the agent that allows it to take actions within the environment |
| The State | The world / moment in time that the agent is operating within or reacting to |
| The Verifier | The scoring rubric and reward (i.e. success measurement) |
| The Config | The different settings that can be varied across runs |
These all manifest as what’s colloquially referred to as “tasks,” essentially all these components packaged up together in a way that targets the recreation and measurement of a given behavior. Conceptually this can be broad, so it’s useful to inspect what an existing published task looks like, using what is probably the most well known of these environments/benchmarks in the space currently, Legal Agent Benchmark from Harvey.
Note: we’re mixing terminology a lot here when discussing “environments” and “benchmarks” since, in reality, the makeup of these share many primitives. I.e. A benchmark can be a group of tasks and environments that provide coverage for a domain, such as different areas of legal practice. It is not uncommon to see more mature suites of effective tasks used as or referred to as benchmarks.
In the LAB
Starting from the top, Harvey’s goal is to, obviously, evaluate an agent’s ability and performance in the legal field. They break this down into various domains such as taxes, real estate, white collar defense investigations (???), etc.
Looking into one such domain, taxes, the folders become more specific as they contain what we’re referring to specifically as the tasks, including such riveting scenarios as drafting a response to an IRS information document request, extracting tax attributes from audited financial statements, and comparing current year filing positions against prior year returns. One can begin to see how the tasks for the specific domains in the legal benchmark are positioned theoretically as various scenarios or units of work that an attorney would complete.

Breaking the task down further into its underlying components via the example of drafting a tax due diligence report, we’re given a task.json file and a /documents directory. The state of this task is held within the directory, which includes the variety of documents required to perform the drafting of the due diligence report, paired with the instructions:
Review the attached data room documents for this acquisition and prepare a buy-side tax due diligence report for the investment committee. Output:
tax-due-diligence-report.docx.
When running the benchmark as a whole, these tasks get built alongside a global harness, which Harvey implements as a relatively simple wrapper with available tools for bash, read, write, edit, glob, grep, and finish (which allows the agent to exit out). They do include a few skills for working with Word, PowerPoint, and Excel file types but for the most part maintain this simplicity with a minimal system prompt.
Note: Harvey’s built in LAB harness is notably simplistic in their public release of the benchmark and tasks, as they likely keep this here to provide a minimal working example for a runner that can consume and complete their task/environment format. As we’ll get into shortly, these environments are used for both agent development as a product and model post-training, of which the specifics and IP would not be included in an open source release.

Each task is scored via LLM-as-a-Judge against a list of criteria, where each criterion names the deliverable(s) it applies to and defines what counts as a PASS or FAIL. Our example includes 75 criteria covering such dimensions as “Flags intercompany management fee as additional TX apportionment risk” and “Employment tax section addresses Section 530 safe harbor applicability.” Interestingly, the final score of the task is a strict all or nothing aggregate of the criteria, where passing 75/75 nets a 100% and anything below a 0.
The config in this case covers any implementation and supporting variables used to carry out the intended harness action set and complete the task, including (but not limited to) other harnesses, the language model chosen, varying system prompts, turn budget, tools/skills, temperature, etc. etc. etc. These are the parameters you configure and tune during product development while simultaneously using the benchmark to measure how well your current agent is performing, which in this example case would be measuring the act of drafting a tax due diligence report. When combined with the rest of the LAB tasks, we’d get a holistic view of our implementation’s “capability for supporting legal work” overall.
Benchmark Breakdown
With our understanding grounded in a live example, we can break down what goes into each step of creating one of these tasks. As mentioned, environments and benchmarks are notoriously difficult deliverables to do well. Other benchmarks, like Cognition’s FrontierCode, have noted more than 40 human hours spent per task created and training data brokers like Mercor pay upwards of $200+ per hour of domain specialists’ time to assist in data curation for building these tasks and environments.
So what does the process for coming up with a task or benchmark for an agent look like currently? I feel it’s useful to work backwards from my definition of a benchmark in A Field Guide to Agent Evaluation:
“Benchmarks intend to accurately measure performance for a specific capability in an agent-agnostic manner.”
Thus, we need to define the capability or behavior that we want to measure the performance of. In the case of Harvey this was the aforementioned capability of supporting legal work, but other examples exist including agentic coding (CursorBench), computer use (OSWorld), general knowledge work (GDPval), etc. When considering the common denominator here, each company behind the benchmarks mentioned is directly developing (or helping to develop!) agents for these specific capabilities (Cursor, OpenAI/Meta/Anthropic, and OpenAI respectively). This is then where you consider the capabilities of the agent you’re developing and its core responsibility you need to measure.
Personally, I actively contribute to LangChain’s internal benchmark for their agent Engine, which has evolved since the initial creation of the linked blog, but started with the core idea of trying to measure how well Engine could identify and cluster agent issues given a set of traces. From a first principles lens, that breaks down into the data the agent operates over (e.g. the traces) and what it’s expected to transform them into (e.g. clustered issues). Which naturally leads to the second step of data labelling.
The data curation and labelling step is probably the most tedious and strenuous step, hence why it’s often outsourced (and the companies doing this well are raking in a shit ton). This step requires developing two sections in tandem, both the relevant dataset and, importantly, the desired outcome. In other words, you’re building a golden dataset that includes labels by which the success or quality criteria have been objectively pre-defined. Taking our previous examples, we can see some approaches:
| Benchmark | Capability | Dataset | Measurement |
|---|---|---|---|
| Legal Agent Benchmark (Harvey) | Legal support | Curated client matters (data room docs) with partner-style instructions | Expert-written pass/fail rubric, LLM-judged, all-pass |
| CursorBench (Cursor) | Agentic coding | Codebase snapshots and requests pulled from real Cursor sessions | Agentic graders on implementation correctness |
| GDPval (OpenAI) | General knowledge work | Tasks and reference files written by industry professionals across 44 occupations | Pairwise LLMaaJ against a human deliverable (win rate) |
| IssueBench (LangChain) | Agent issue triage | Trace batches with labeled failures and expected clusters | Classification metrics against identification & grouping |
In short, you need to clearly be able to articulate the input to the agent, and how to judge the output against your best measurement of success. This often includes a mixture of heuristic measurements (e.g. F1, exact match, held out unit test suites passing, etc) and fuzzy measurements (e.g. LLM-as-a-Judge aligned via inter-rater agreement measurement). As this is the core of what defines the success of a given task, time and attention to detail are often spent here. We see the biggest inclusion of domain-specific SMEs brought in to instill their knowledge into this step, both labelling the data directly and providing their judgement as a reference for automated grading. As I’ve been stressing, it’s imperative to get this step as accurate as possible, since everything downstream assumes the score can be trusted, whether that’s reading it at face value or using it to train and hill climb against.

With the dataset, capability, and success criteria defined, the final component needed is an environment to execute the agent within. Often, this requires mocking the exact world that the agent would operate within for the given task. This can be quite simple, like Harvey’s example runner that loads the task’s documents into a sandboxed workspace and gives the agent basic file operation tools, but often specific task environments are more complex. If you were making a task that required measuring an agent’s ability to run data analysis using BigQuery, the environment required would have to provide something like: a simulated BigQuery instance, seeded rows with realistic skew and nulls, partitioned and clustered tables, so query cost behaves like production, an inspectable INFORMATION_SCHEMA, half-stale column descriptions, because who updates those, soft-deleted rows and test accounts that a correct answer excludes, real error strings for resource limits and permission boundaries, result sets too large to hold in context, etc. etc. etc. And that only really covers making a realistically mocked BigQuery instance.
Going back to my personal angle, the IssueBench team had to develop a system that could sandbox and mock the entire LangSmith backend, as our agent operates directly within the platform itself. Thankfully, we also did some OSS work at the time with benchmark runners like Harbor to make our lives, and the lives of those using LangSmith to run similar evals, easier. From what I’ve seen, similar environments are used for other in-product assistants when undergoing similar benchmarking and evaluations.

Note: LLMs these days are particularly astute at recognizing patterns, making the fidelity of a mocked system paramount. While creating and testing the environments to run these tasks, I’ve often seen something as small as mocked metadata cause an agent to notice it’s being evaluated or recognize that it’s operating in a non-production environment. It’s up to the developer to ensure these small details don’t slip as they can affect the validity/production-parity of a run.
The environment also ties into the harness, as the toolset required to interact with the data in the environment is often bundled with the LLM and its scaffolding. The implementation of the harness, however, is the most variable as that’s often what’s being explicitly tested under the constraints of the environment and various configurations. It’s also largely what’s being optimized by the team consuming the benchmark if the goal is to improve an agent product. On that point, only one half of the puzzle is actually creating the benchmark and evaluation, the other half is how it gets applied in practice.
Running up that Hill
To level set, the reason that effective benchmark evaluations like these are so useful is that they not only give us an objective measurement point, but also act as a target for optimization for a desired behavior. This is generally where we start to see references to “Hill Climbing” appear, referring largely to the loop of running an agent on a benchmark, observing its behavior and score, then iteratively tweaking the agent to improve its performance on the task (or set of tasks).

There exist many levers in the harness that can be adjusted to push performance (or other related efforts such as lowering cost while maintaining performance), including but not limited to: the system prompt, provided tools, the model of choice, the architecture (e.g. multi-agent, RLM, using or not using subagents, etc), and all the components that make up the program of your agent system. But the usage of these environments does not stop at the harness layer, and we know this to be true as major labs are spending lots of cash to purchase these datasets and implementations for model training.
In a similar manner to how the program of an agent can be tweaked to improve performance, the weights of an LLM can be updated as well. This would reframe the benchmark as a training environment to run methods like reinforcement learning during the post-training phase of an LLM, with the verifier score acting as the reward. In this case, the harness is actually not as important as you’re isolating the measurement of the model under the given environment, held constant. One can go as far as to say it’s even beneficial to use multiple harnesses to better generalize the LLM’s ability to operate within and succeed at a wide variety of tasks and implementations.
Note: This can also be a dual optimization problem which we see as models are trained with various environments under the same harness. For example, Claude models being trained with the Claude Code harness, and GPT models with Codex. This lets the labs own both the intelligence and implementation layers to push further performance and usefulness.
Outside of these two major use cases, having an objective measurement and target allows one to do all forms of experimentation while staying grounded in real, trustworthy metrics. Other use cases could include training a smaller parameter language model on a task to swap out via RL, generating rollout data for SFT, eval driven agent development, and any improvement/regression checks throughout AI application or model development.
Keeping Pace with the Frontier
Building effective benchmark shaped evaluations and tasks allows us to move past evals as a regression check, and towards evals as optimization targets that can drive effective, measurable improvement while developing AI related applications. The importance of this (and doing it well) is offset largely by the difficulty in getting it right, requiring many hours (and dollars) for creating effective datasets, realistic environments, and trustworthy grading rubrics. But when done well, it gives the user an indispensable tool for optimizing their agent or model.

Two additional factors that add to this importance when considering broader trends: 1) it’s profitable when done well- we see major data vendors rising in both mind share and capital as they curate and sell domain/behavior specific environments to labs and businesses, and 2) frontier labs are investing heavily in building both the tools needed and systems in place to automate this work. It’s incredibly high leverage to be able to create and understand this style of evaluation, and I believe forms and systems related to making this process easier will become increasingly valuable.