The OpenAI Agents API Makes the Harness a Managed Product

OpenAI's Agents API turns the Codex harness into a managed service with hosted sandboxes. Netics on why the interesting boundary is the sandbox, not the model.

Netics feature card for the OpenAI Agents API article with the official OpenAI identity
Netics editorial feature card using the official OpenAI identity

TL;DR

  • OpenAI launched the Agents API in public beta on September 10, 2026: the same harness and infrastructure that powers Codex, exposed as an API for building and running cloud agents.
  • OpenAI handles the agent loop — model calls, tool use, context management, subagent coordination; you control the agent's capabilities and choose where code runs.
  • Two sandbox paths: bring your own sandbox or connect a provider (Blaxel AI, Cloudflare Dev, Daytona, DigitalOcean, E2B, Modal, Oracle Cloud, Runloop AI, Vercel), or use OpenAI-hosted sandboxes with CPU, GPU, and memory options, including VPC deployments.
  • Customer-reported numbers on launch: eval score 0.71 to 0.85, 4x lower orchestration latency, 60% lower cost per case, 86% fewer failed agent responses.
  • Netics' take: the interesting boundary moved from the model to the sandbox — once OpenAI runs the loop, your remaining decision is where the agent's code and files actually execute.
Official OpenAI announcement image for the Agents API launch
Official OpenAI announcement image for "Introducing the Agents API" (openai.com, September 10, 2026).

The harness becomes the product

The Agents API announcement is short and strategically blunt: OpenAI learned what it takes to make long-running agents work by scaling Codex and ChatGPT for Work to millions of users, and it is now packaging that harness and infrastructure as an API. The components it names are the parts that usually break in homegrown agent stacks: a harness that manages context, uses tools efficiently, and coordinates subagents, plus infrastructure that keeps agents running reliably for days, with environments where they can work with files, run code, and save intermediate results.

That framing matters because it names the actual cost center of agent development. Most teams building agents do not struggle to call a model; they struggle with everything around the call — context that grows without bound, tool calls that fail because state was lost, subagents that cannot report back, sessions that die overnight. OpenAI's pitch is to absorb exactly that layer. The customer keeps what makes the agent unique: capabilities, tools, and the environment it runs in. The API's subagent support, in particular, is where the harness earns its keep: coordinating workers, merging their results, and keeping each one's context isolated is precisely the orchestration work that homegrown stacks reimplement badly and repeatedly.

Official OpenAI architecture diagram of the Agents API: the application sends tasks and tool calls to the harness, which coordinates the sandbox
Official OpenAI diagram from the Agents API announcement: tasks and tool calls between your application and the managed harness, with the sandbox as the execution boundary (openai.com, September 10, 2026).

The sandbox is now the decision

The part of the announcement that deserves more attention than it will get is the sandbox model. OpenAI runs the agent loop; the developer controls the agent's capabilities and chooses where it runs code and works with files. There are two paths. Bring your own sandbox, or connect a sandbox provider — the announcement lists Blaxel AI, Cloudflare Dev, Daytona, DigitalOcean, E2B, Modal, Oracle Cloud, Runloop AI, and Vercel as first-class integrations. Or use OpenAI-hosted sandboxes, where the agent can run code, work with files, and produce artifacts while OpenAI provisions and manages the environment, at standard container rates with model usage billed separately.

This split — harness managed by OpenAI, sandbox chosen by the customer — is the real architecture statement in the product. It means the execution environment is an interface, not an implementation detail. A team can prototype on hosted sandboxes and later move the same agent definitions to a sandbox inside its own VPC without rewriting the orchestration layer. Hypha's launch quote makes the operational point concretely: separating the harness from the sandbox reduced failed agent responses by 86%. That number is not about model quality; it is about the harness no longer being entangled with the environment's failure modes.

Existing analysis on this blog has already made the durable-state argument for agent sandboxes: long-running tasks only work if the runtime keeps state between steps. The Agents API extends that argument to procurement. When you previously chose a sandbox, you were choosing your own runtime; now the choice is which execution boundary you hand to a managed service. The governance question shifts from "can we run this loop?" to "where does this agent's code execute, and who can see the files?"

Official OpenAI screenshot of the Agents API sandbox integrations row with Daytona, Runloop, and Vercel
Official OpenAI screenshot from the Agents API announcement showing sandbox provider integrations including Daytona, Runloop, and Vercel (openai.com, September 10, 2026).

What the launch numbers do and do not prove

The announcement carries five customer quotes, and they read like a portfolio of exactly what the harness should fix. Ciridae's CTO reports an evaluation score moving from 0.71 to 0.85 with subagent support, and a 4x latency reduction from the new APIs versus a previous setup where orchestrating subagents was cumbersome. SafetyKit reports a 60% reduction in cost per case after migrating its case review workflow, with lower latency and better token efficiency. Hypha reports the 86% reduction in failed responses. Dwelly notes the API handled bursty workloads — fanning out work across hundreds of agents asynchronously without idle infrastructure between peaks.

Those are real, attributed, and directionally coherent. They are also all vendor-selected. The honest reading is that the harness layer was genuinely costly for these customers — orchestration, context, and session reliability — and the managed version removed real friction. The numbers that would tell the second half of the story are not in the announcement: what happens to cost when the same workload runs at steady state for a quarter, how the hosted sandboxes' container rates compare to running your own fleet, and how much of the "no additional fees" pricing survives real production usage. Pricing on tokens and tools is a clean launch story; the follow-up question is what a heavy agent workload actually accumulates.

Original Netics diagram: self-managed agent loop versus the OpenAI Agents API — what stays with the customer and what OpenAI now runs
Original Netics diagram: what changes when the harness becomes a managed product — context, orchestration, and session reliability move to OpenAI; capabilities, tools, and execution environment stay with the customer; source: OpenAI Agents API announcement.

Where the boundary sits for European teams

For a French or European engineering team, the Agents API is useful precisely because it makes one decision very explicit: the execution boundary. European deployments that touch personal data will not run agent workloads on hosted sandboxes in a region they do not control. The bring-your-own-sandbox path — including deployments within the customer's VPC — is the compliance-relevant option, and the announcement's explicit support for VPC deployments matters more than any model capability in the same paragraph.

The practical checklist for evaluating the Agents API is short. First, define where the sandbox lives for each workload class, because that determines the data-processing story. Second, measure the full cost of a real workload, not the launch-pricing symbols. Third, decide whether the managed loop removes enough operational burden to justify the coupling — because once your agent's orchestration runs on OpenAI's infrastructure, your failure domain includes their availability. The architecture is honest about this trade: OpenAI runs the loop, and you choose the sandbox. The choice is the product.

Original Netics diagram: measured results from Agents API launch customers — evaluation score, latency, cost per case, and failure reduction
Original Netics diagram: the launch-reported metrics — evaluation score 0.71 to 0.85, 4x lower latency, 60% lower cost per case, 86% fewer failed responses; source: OpenAI Agents API announcement.

For teams weighing managed agent runtimes against self-managed stacks, the same control-plane reasoning that applies to agent sandbox durability applies here one level up: once the harness is managed, the sandbox is the last infrastructure you own outright. Plan that boundary first, then let the API remove the rest. The Netics homepage is where we start these evaluations.

Sources

Source: "Introducing the Agents API" — openai.com, September 10, 2026.