We started looking seriously at factory evals because we had a fairly ordinary question: when a new model or agent harness comes out, should we switch?
At the scale of one developer and one agent, people answer this by trying it for an afternoon. It feels smarter, the first patch looks good, and somebody changes the default. Once agents are opening around twenty pull requests a day and writing most of the code, that stops being a tooling preference and starts looking more like a production migration, but most teams still evaluate it by feel.
The obvious alternative is to replay work the factory has already shipped. Take fifty old tickets, restore the code from before each fix, run them through the proposed configuration, and compare the result with what actually happened: did it fix the bug, catch the same edge cases, need a human, take more attempts, or cost more?
This sounds straightforward until you try it, because the answer already exists and the agent needs the internet.
The offline version rewards the wrong agent
The simplest setup is a VM containing the old repository, with git history cut before the fix and the network disabled so the agent cannot find the merged pull request.
We tried versions of this, and the problem is that a real engineering task almost never consists of a repository alone. The product needs packages from a registry, authentication needs an identity provider, the task lives in Linear, the discussion lives in Slack, and the agent may need GitHub, internal services, or documentation before it can even reproduce the bug.
When those systems disappear, a capable agent does what capable agents do: it works around the environment. It mocks the auth service, replaces an integration with a stub, skips the end-to-end test that cannot run, and eventually gets a green test suite. If the grader only checks the visible tests, this can score as a success even though we would never accept the same work in the real factory.
The perverse part is that the eval can prefer the worse agent. A cautious agent that says it cannot verify the fix fails, while an agent that removes the inconvenient check passes.
Turning off the internet prevents one kind of cheating and creates another.
The live internet is worse
Giving the replay normal network access makes the task realistic, but now fifty tickets across three candidate configurations can create 150 real pull requests, send messages to real Slack channels, update old Linear tickets, trigger email, or charge a card.
It also breaks the benchmark in less spectacular ways. Dependencies change, APIs move, rate limits vary, and runs interfere with each other through shared state. More importantly, GitHub and Linear contain the merged fix, the review comments, and often a neat explanation of the root cause, so the agent can solve the task by finding the answer we were trying to hide.
Telling the agent not to look is not a solution. We are testing a system designed to search for useful context, so the internet it searches has to already be the internet from the moment the work was done.
The conclusion we reached was slightly strange but technically simple: an eval needs that internet frozen in time.
What an eval internet looks like
Each replay starts from a snapshot at time T containing the repository, database, services, tools, and agent configuration. The internet is part of the snapshot. It is frozen there.
GitHub still has the ticket and the unfixed code. The merged pull request, the review, and the root-cause writeup are not in it yet. Slack and Linear still have the discussion as it stood. The registry and the docs still answer the way they answered then. The agent can search, install, and read, and what it finds is the world at T.
When the agent writes, the write lands in that same frozen world. Creating pull request 42 titled fix auth, reading pull request 42, and commenting on it are three operations on one object. A canned 200 OK cannot do that. Agents read back what they wrote, wait for state changes, retry calls, and choose the next step from the result, so the copy of the service has to keep state, implement the parts of the API the agent uses, and record every mutation.
That is what the fakes are for. They are the services as of T, small enough to run inside the eval and faithful enough that an unplanned path still behaves. Registries and public documentation are a read-only cache from the same moment. The model API stays live, because the model is under test and the snapshot has nothing to say about which model will be asked. We built DoubleAgent around those stateful copies, because a folder of mocked responses fell apart as soon as a run took a path we had not predicted.
The event log also gives us something we did not have when we were only grading patches: a record of what the agent did to the world.
A correct diff is not enough
We recover hidden tests from the real merged pull request and run them against the replay, outside the agent’s workspace. That tells us whether the behavior is equivalent without requiring the agent to produce the same diff as the human or original agent.
Then we inspect the service event logs. Did it open the expected pull request? Did it send the right notification? Did it make twelve failed attempts first? Did it write to billing when the task had nothing to do with billing? Did it update the ticket before it had verified the fix?
This matters because a factory can produce the right code through behavior we would never allow in production, while a transcript can look thoughtful right up to the moment the agent performs the wrong external action. Grading the final diff misses both.
For each replay, we care about:
- Whether the hidden tests pass and the fix is behaviorally equivalent.
- Whether the agent catches the known bugs and edge cases.
- How many attempts it makes and whether a human has to step in.
- Which expected and forbidden side effects appear in the event log.
- What the accepted outcome costs across model usage, compute, builds, and tests.
I would not collapse those into one magic score. A model that costs twice as much but catches a class of security bugs may be the right choice for one line and the wrong choice for another, while a cheaper harness that creates more review work has mostly moved cost out of the model invoice and onto the engineering team.
We also run each case more than once. One successful run tells us the configuration can solve the task; it says very little about whether it usually will, so pass rate, attempts, intervention, side effects, and cost all need a distribution rather than one lucky result.
The failed tickets are the benchmark
There is a temptation to fill the eval set with clean, successful tasks because they are easy to reconstruct and they make the chart look sensible.
The cases we learn most from are the ugly ones: tickets where a human rescued the agent, fixes that passed review and later broke production, loops that created several bad pull requests, and incidents where the agent optimized the test instead of fixing the product. Those cases describe the actual limit of the factory, which is exactly what a model or harness change can move.
The set also cannot be static. New work should enter, stale work should leave, and public artifacts become more likely to appear in training data over time. The fakes need contract tests against the real APIs too; otherwise the benchmark slowly turns into an accurate evaluation of a version of GitHub that does not exist.
This changes how we choose models
When we compare a model or harness, I want a fairly boring answer: on tickets like ours, did it merge more often, catch the same bugs, need fewer interventions, avoid unexpected writes, and what did a successful run cost?
That will not always pick a winner. Paying twice as much may be sensible for a security review and absurd for a dependency update, but at least the argument is about results from the same work rather than impressions from three unrelated runs.
This is the part I think most factory benchmarks skip over. Replaying the code is manageable; replaying the services around it, without exposing the old answer or changing anything real, is where the eval becomes difficult.
Until that middle ground exists, a benchmark may tell you which agent performed best in the benchmark environment. I would be careful about assuming it tells you which one should run the factory.
