Google's harness engineering lesson: test how coding agents behave
TL;DR
Google's September 2026 account of harness engineering argues that end-to-end benchmarks are useful report cards but weak debugging instruments: a benchmark can grade the destination and still hide the route. The more actionable layer is a behavioral evaluation suite that makes the harness observable through small checks for tool calls, file changes, validation steps, and safe task completion. Google says dogfooding comes before a grand evaluation framework, then recommends a repeatable loop in which small failures become assertions, with flexible checks, batch evaluation, and regression protection. Our view: the important shift is not abandoning benchmark scores; it is giving teams a way to explain and improve them.
Google's developers are not presenting another leaderboard result. In its September 9, 2026 Google Developers Blog post, principal engineer Taylor Mullen and staff software engineer Christian Gunderman describe a practical problem in building AI coding agents: a composite end-to-end score can move without telling the engineering team what actually changed. Their proposed answer is harness engineering, with behavioral evaluations acting as the feedback mechanism between a model change and the final score.
The distinction matters because an agent is not only a function that maps a prompt to a patch. It is a system that chooses tools, interprets ambiguity, edits files, validates work, and decides when it is finished. A final answer can look acceptable while the path taken is fragile. Conversely, a successful task may follow a route that a rigid test never anticipated. Google's approach is useful precisely because it treats those two problems separately.
A benchmark can grade the destination and still hide the route
The post starts with familiar end-to-end benchmarks such as Terminal-Bench and DeepSWE. Give an agent a large codebase, impose a time limit, and count how many tests pass or fail. That produces an understandable scorecard. It is also expensive to investigate when the score changes.
A lower score might reflect overconfidence on an ambiguous prompt. The agent might have failed to verify the test suite before submitting, or hallucinated a command-line flag. An end-to-end result does not typically identify which behavior broke. Teams are left with a number and a costly debugging exercise.
Google's case for behavioral evaluations is therefore not “benchmarks are useless.” It is that benchmarks answer a different question. The large evaluation checks the final destination. Smaller behavioral checks provide guideposts along the route. This is an important caveat in the source, and it keeps the proposal from becoming another false binary between macro and micro testing.
For teams building production systems, the practical implication is straightforward: do not ask a single score to serve as both product report card and engineering diagnostic. Keep the scorecard, but add tests that can tell you why it moved.

Behavioral tests make the harness observable
Google describes behavioral evaluations as integration tests for improving harness operation. Instead of demanding final string equality or measuring whether an agent completed an entire multi-file refactor, the test asserts on a discrete, observable action.
The examples are deliberately ordinary. Given an underspecified prompt, does the agent ask a clarifying question rather than guess? When it modifies a build file, does it run the local validator before declaring completion? When it writes documentation, does it include canonical repository links? These are not cosmetic preferences. They are operating behaviors that can make an agent safer to iterate.
The source includes a concrete Python example written for the Antigravity SDK. The test asks an agent about the weather in Mountain View, California, then collects the names of the tools called. Its assertion checks that the built-in web-search tool was used. The test is not comparing the prose of the answer with a golden sentence. It is checking whether the agent consulted live information instead of answering from memory.
The strongest idea in Google's article is this: Output matching is often brittle because many answers can be correct. An intermediate assertion can be more stable when the desired behavior is clear. It also gives the engineer a useful failure message: the agent answered without consulting live search.
This is the kind of engineering detail that deserves more attention than a new benchmark headline. A harness that records what happened can support a tighter loop between prompt design, tool schemas, model selection, and observed behavior. For more of our practical perspective on dependable automation, explore Netics Labs.

Dogfooding comes before a grand evaluation framework
Google's sequencing advice is just as important as its test design. When bootstrapping an agent from scratch, the team should first follow developer instinct and dogfood the system. The source names a threshold of basic usefulness: the agent should be capable of working on its own codebase, handling boilerplate, writing its own Markdown renderer, and executing routine developer tasks.
Only then, Google argues, does a formal evaluation suite become the right next phase. At that point, the purpose is not primarily to celebrate a two-percent improvement. It is to provide confidence that a prompt tweak, a tool-schema change, or a model upgrade did not make the agent holistically worse.
This is a welcome antidote to evaluation theater. A sophisticated harness cannot compensate for an agent that has not yet demonstrated basic utility to its own builders. Dogfooding produces the failure modes worth encoding. The evaluation suite then turns those failures into repeatable checks rather than abstract guesses about what users might encounter.
The sequence also makes the work cheaper. Start with a real mistake, not a sprawling taxonomy of hypothetical failures. Once an agent is useful enough to expose meaningful weaknesses, formalize the behaviors that matter.
Small failures become a repeatable three-step loop
Google suggests starting with a three-step behavioral testing loop.
First, pick one failure mode. If the agent forgot to run unit tests before marking a task complete, isolate that missed action and make it the target. The point is to convert an incident into a testable behavior rather than immediately redesigning the entire system.
Second, choose an assertion that matches the task's complexity. For a simple task with one optimal solution, a strict single-turn assertion can check whether the agent reached a specific milestone, such as calling the test runner. For a complex task, a rigid tool sequence can punish a correct alternative. In that case, Google recommends a fuzzier outcome-based check, including an LLM-as-a-judge approach, to assess whether the chosen steps solved the problem successfully and safely.
That flexibility is not a concession to weak testing. It is recognition that agent trajectories are variable. A deterministic assertion is valuable when the behavior is deterministic. It becomes counterproductive when it mistakes one implementation path for the only safe outcome.
Third, automate batch evaluations to monitor stability. Google cautions against blocking pull requests on single evaluation runs that may be noisy because AI models are nondeterministic. A larger batch provides aggregate pass rates over time. That directional signal lets a team tune prompts and upgrade models without halting development for variance it already expects.

Regression protection is the real product of the suite
The post gives a local command—pytest evals/behavioral/ -v—and says the behavioral suite can run in under five seconds. The value is not the command by itself. It is the possibility of a fast, repeatable safety net around a system whose behavior changes as often as its prompts and models do.
Google also describes a loop in which an LLM tweaks its own system prompt until a failing behavioral test passes, while the rest of the suite acts like a CI/CD guardrail. That arrangement still needs judgment: making one test pass is not enough if existing behaviors regress. The surrounding suite is what keeps local optimization from becoming system-wide damage.
Our independent read is that this is where “harness engineering” earns its name. The model is only one component. The harness defines how the component is observed, constrained, and improved. The best score is not necessarily the immediate objective; trustworthy iteration is.
End-to-end suites remain necessary because they verify whether the complete agent can reach the final destination. Behavioral evaluations make that destination reachable by exposing the small decisions that determine whether an agent is reliable along the way. Teams that use both gain a clearer signal when changing prompts, adding features, or deploying a new model.
For builders, the next test may be sitting in a recent failure log. Write that one first, make the assertion fit the task, and measure the behavior in batches before treating a score movement as a verdict.
Sources
- Google Developers Blog, “The Anatomy of Harness Engineering: How to Evaluate, Iterate, and Guard AI Coding Agents,” September 9, 2026: https://developers.googleblog.com/the-anatomy-of-harness-engineering-how-to-evaluate-iterate-and-guard-ai-coding-agents/
Source: The Anatomy of Harness Engineering: How to Evaluate, Iterate, and Guard AI Coding Agents — developers.googleblog.com, September 9, 2026