The OpenAI Agents API Makes the Harness a Managed Product
OpenAI's Agents API turns the Codex harness into a managed service with hosted sandboxes. Netics on why the interesting boundary is the sandbox, not the model.
TL;DR
- OpenAI launched the Agents API in public beta on September 10, 2026: the same harness and infrastructure that powers Codex, exposed as an API for building and running cloud agents.
- OpenAI handles the agent loop — model calls, tool use, context management, subagent coordination; you control the agent's capabilities and choose where code runs.
- Two sandbox paths: bring your own sandbox or connect a provider (Blaxel AI, Cloudflare Dev, Daytona, DigitalOcean, E2B, Modal, Oracle Cloud, Runloop AI, Vercel), or use OpenAI-hosted sandboxes with CPU, GPU, and memory options, including VPC deployments.
- Customer-reported numbers on launch: eval score 0.71 to 0.85, 4x lower orchestration latency, 60% lower cost per case, 86% fewer failed agent responses.
- Netics' take: the interesting boundary moved from the model to the sandbox — once OpenAI runs the loop, your remaining decision is where the agent's code and files actually execute.

The harness becomes the product
The Agents API announcement is short and strategically blunt: OpenAI learned what it takes to make long-running agents work by scaling Codex and ChatGPT for Work to millions of users, and it is now packaging that harness and infrastructure as an API. The components it names are the parts that usually break in homegrown agent stacks: a harness that manages context, uses tools efficiently, and coordinates subagents, plus infrastructure that keeps agents running reliably for days, with environments where they can work with files, run code, and save intermediate results.
That framing matters because it names the actual cost center of agent development. Most teams building agents do not struggle to call a model; they struggle with everything around the call — context that grows without bound, tool calls that fail because state was lost, subagents that cannot report back, sessions that die overnight. OpenAI's pitch is to absorb exactly that layer. The customer keeps what makes the agent unique: capabilities, tools, and the environment it runs in. The API's subagent support, in particular, is where the harness earns its keep: coordinating workers, merging their results, and keeping each one's context isolated is precisely the orchestration work that homegrown stacks reimplement badly and repeatedly.

The sandbox is now the decision
The part of the announcement that deserves more attention than it will get is the sandbox model. OpenAI runs the agent loop; the developer controls the agent's capabilities and chooses where it runs code and works with files. There are two paths. Bring your own sandbox, or connect a sandbox provider — the announcement lists Blaxel AI, Cloudflare Dev, Daytona, DigitalOcean, E2B, Modal, Oracle Cloud, Runloop AI, and Vercel as first-class integrations. Or use OpenAI-hosted sandboxes, where the agent can run code, work with files, and produce artifacts while OpenAI provisions and manages the environment, at standard container rates with model usage billed separately.
This split — harness managed by OpenAI, sandbox chosen by the customer — is the real architecture statement in the product. It means the execution environment is an interface, not an implementation detail. A team can prototype on hosted sandboxes and later move the same agent definitions to a sandbox inside its own VPC without rewriting the orchestration layer. Hypha's launch quote makes the operational point concretely: separating the harness from the sandbox reduced failed agent responses by 86%. That number is not about model quality; it is about the harness no longer being entangled with the environment's failure modes.
Existing analysis on this blog has already made the durable-state argument for agent sandboxes: long-running tasks only work if the runtime keeps state between steps. The Agents API extends that argument to procurement. When you previously chose a sandbox, you were choosing your own runtime; now the choice is which execution boundary you hand to a managed service. The governance question shifts from "can we run this loop?" to "where does this agent's code execute, and who can see the files?"

What the launch numbers do and do not prove
The announcement carries five customer quotes, and they read like a portfolio of exactly what the harness should fix. Ciridae's CTO reports an evaluation score moving from 0.71 to 0.85 with subagent support, and a 4x latency reduction from the new APIs versus a previous setup where orchestrating subagents was cumbersome. SafetyKit reports a 60% reduction in cost per case after migrating its case review workflow, with lower latency and better token efficiency. Hypha reports the 86% reduction in failed responses. Dwelly notes the API handled bursty workloads — fanning out work across hundreds of agents asynchronously without idle infrastructure between peaks.
Those are real, attributed, and directionally coherent. They are also all vendor-selected. The honest reading is that the harness layer was genuinely costly for these customers — orchestration, context, and session reliability — and the managed version removed real friction. The numbers that would tell the second half of the story are not in the announcement: what happens to cost when the same workload runs at steady state for a quarter, how the hosted sandboxes' container rates compare to running your own fleet, and how much of the "no additional fees" pricing survives real production usage. Pricing on tokens and tools is a clean launch story; the follow-up question is what a heavy agent workload actually accumulates.

Where the boundary sits for European teams
For a French or European engineering team, the Agents API is useful precisely because it makes one decision very explicit: the execution boundary. European deployments that touch personal data will not run agent workloads on hosted sandboxes in a region they do not control. The bring-your-own-sandbox path — including deployments within the customer's VPC — is the compliance-relevant option, and the announcement's explicit support for VPC deployments matters more than any model capability in the same paragraph.
The practical checklist for evaluating the Agents API is short. First, define where the sandbox lives for each workload class, because that determines the data-processing story. Second, measure the full cost of a real workload, not the launch-pricing symbols. Third, decide whether the managed loop removes enough operational burden to justify the coupling — because once your agent's orchestration runs on OpenAI's infrastructure, your failure domain includes their availability. The architecture is honest about this trade: OpenAI runs the loop, and you choose the sandbox. The choice is the product.

For teams weighing managed agent runtimes against self-managed stacks, the same control-plane reasoning that applies to agent sandbox durability applies here one level up: once the harness is managed, the sandbox is the last infrastructure you own outright. Plan that boundary first, then let the API remove the rest. The Netics homepage is where we start these evaluations.
Sources
- Introducing the Agents API — OpenAI, September 10, 2026.
- Introducing the Agents API and hosted sandboxes — OpenAI Developer Community, September 10, 2026.
Source: "Introducing the Agents API" — openai.com, September 10, 2026.