An AI model can propose a plan. An agent needs somewhere to put its tools, files, context, mistakes, and the task it promised to finish three hours ago. OpenAI’s newest developer product is aimed squarely at that less glamorous machinery.

OpenAI launched the Agents API in public beta, offering the execution harness used by Codex as a programmable service. Developers can run tasks that last for days, use built-in tools, preserve context, create subagents, and choose hosted, self-managed, or partner sandboxes. OpenAI says the API carries no additional platform fee beyond the models and tools consumed.

A model gets a workshop

The distinction between a model API and an agent API is operational. A model request is usually bounded: send input, receive output. An agent may inspect a repository, run a command, discover a failure, retry, delegate a narrow subtask, and return later with an artifact. That requires durable state and a controlled place where actions can occur.

Packaging those pieces as a service could make sophisticated workflows much easier to build. It also concentrates responsibility. The agent may arrive with a workshop, but the developer still decides which doors it can open, what it is allowed to change, how much it can spend, and when a human must approve the next move.

The sandbox choices are therefore as important as the headline. A hosted environment reduces setup. A self-managed environment offers more control. Partner options may fit existing infrastructure. None automatically answers the uncomfortable question: what happens when a long-running process confidently pursues the wrong objective overnight?

The signal

Agent products are moving from impressive loops in a demo to infrastructure with scheduling, persistence, tools, and boundaries. That is progress, but it changes the evaluation target. Developers must test not only whether the model gives good answers, but whether the whole system stops, recovers, records its actions, and respects permissions.

The agent now has an office, a toolbox, and permission to stay late. The next feature request should be a very good lock on the supply closet.

The product is the boundary, not the demo

Turning agents into a service gives developers reusable machinery for tools, state, tracing, and long-running work. It also centralizes risk. An agent that can call external systems needs explicit limits on what it may read, change, purchase, send, or delete. A model deciding the next step is only one component; authorization, observability, rate limits, and human review determine how expensive a bad step can become.

This makes evaluation more concrete. The important question is not whether an agent can finish a polished benchmark task. It is whether the surrounding service exposes enough evidence to reconstruct what happened when the task fails halfway through and leaves three systems disagreeing about reality.

TINA’s view: managed agents need boring controls

TINA’s view: packaging agent infrastructure as a cloud service can reduce duplicated engineering, but convenience must not blur responsibility. Developers still own the permissions they grant and the consequences of an action. The strongest counterargument is that demanding enterprise-grade controls before experimentation would keep useful automation inaccessible to smaller teams. The answer is safe defaults and staged capability, not an empty permissions page.

This judgment improves if the platform makes narrow scopes, approval gates, complete traces, budget caps, and reliable cancellation standard rather than premium decoration. It worsens if success is measured mostly by how autonomously a demo runs. Watch what happens after an agent is interrupted. A system that cannot stop cleanly is not autonomous; it is merely unattended.

Cost is part of that boundary. Long-running agents can consume compute and paid tools after the useful work has stopped. A visible budget and a hard ceiling are safety controls, not merely finance features. Autonomy without an expense limit is a surprisingly traditional way to learn about governance.