Cloud Agent APIs Compared in 2026
How to launch agents from your own code, and which APIs let you run them for your users
On this page
Cloud and background agent APIs let your code start an agent on a remote machine and collect the result, whether that's a pull request, a file, or an answer. You can start one from your own product, a CI job, or a Slack bot.
The main differences are who pays for the agent, whose account it runs as, and whether the terms let you offer it to your own users. This guide sorts the options by that, then covers how results come back, what we saw when we ran the same task on two of them, and what to test before production.
Everything here is as of October 9, 2026. Almost all of these APIs are in beta or preview.
The Short Version
Who Pays, and Whose Account Is It?
There are two ways to run agents for your users. Your product can pay for the agent and call the API with your own key, or each user can connect their own account and pay for their own usage. The APIs below support one or the other, and their terms decide which.
- Some terms limit use to your own business. Devin's terms say "for your internal business purposes only". LangChain's and Factory's standard terms say much the same, so check your agreement before building a product on them. DigitalOcean's preview terms allow only non-production workloads for now.
- Some run as each user. GitHub's agent API only accepts user tokens, like a personal access token or a GitHub App user token, not an app's installation token. That works when each user has a Copilot plan and authorizes your app. Cursor's service accounts can only act as members of your own Cursor team.
- Subscriptions can't be shared. Unless you have a separate agreement, Anthropic lets you run unmodified Claude Code in your product only if each user signs in with their own Claude account, API key, or cloud provider credentials and pays for their own usage. You can't route their requests through your subscription or pay for their usage. Codex cloud and Jules run under one person's account.
If your product pays and your customers just use it, the options here are Claude Managed Agents, OpenAI's Agents API (or its AWS version), Gemini, and Omnara. They differ on how they keep each user's credentials separate.
You can also skip the hosted APIs and run an agent SDK inside your own application, like the Claude Agent SDK or the OpenAI Agents SDK. That gives you full control of the agent, and leaves running, scaling, and storing sessions to you. I covered that path in How to Deploy an AI Agent to Production.
Built to Run for Your Users
Claude Managed Agents
Anthropic's Claude Managed Agents runs Claude in a sandbox behind an API, in beta. You create an agent, an environment, and a session, then stream events back or get signed webhooks. Each end user can have their own credential store, and the sandbox can run on your own infrastructure through a self-hosted worker. It runs Claude models only and isn't eligible for Zero Data Retention.
OpenAI's Agents API
OpenAI's Agents API runs the Codex harness for you, in beta. One call creates the agent, the sandbox, and the first task together. Credentials can be tagged per end user, results come back as a stream or signed webhooks, and the sandbox can be OpenAI's, a partner's, or yours. It runs OpenAI models only, stores data in the US only, and doesn't offer Zero Data Retention.
Amazon Bedrock Managed Agents, powered by OpenAI, runs the same agent loop and models inside AWS, with AWS IAM for authentication. It fits apps that already use AWS IAM.
Gemini
Google's Managed Agents in the Gemini API is in preview and doesn't charge for compute yet. It stores credentials per project and refreshes OAuth tokens for you, but it has nothing per end user, so your app maps each user to the credentials their agent may use.
Omnara
Omnara is an open-source platform with an API, a TypeScript SDK, a CLI, and signed webhooks, and it works with any model. Agents can run on Omnara's machines, on sandboxes in your own account with providers like Modal and Daytona, on your own servers, or on a fully self-hosted install, where you run the platform yourself. Access is organized by organizations, projects, and roles. Omnara Cloud has no platform fee, though models and machines still cost money. You can keep customer credentials in your backend and call them through custom tools, with your app deciding which user's access each agent gets.
Built for Your Own Team
These are the coding-agent APIs. They're good at turning an issue into a pull request for your own engineers.
- Cursor Cloud Agents API. Formerly the Background Agents API. You send a prompt and a repo and follow progress over a stream you can reconnect to, on Cursor's machines or your own workers. The current version is in beta. Webhooks are only on the older version for now.
- GitHub Copilot's cloud agent. Assign an issue or call the tasks API, and it opens a pull request from a GitHub Actions runner. The API is in preview, you can pick the model, and results come back by polling.
- Devin. Sessions are created under your organization, with optional workers on your own machines. Results come back by polling and pull requests.
- Factory. Its sessions API is only enabled for selected organizations.
- LangChain Managed Deep Agents. Not a coding-specific product, and it supports private conversations for each end user. Its standard terms say internal business use, so confirm your agreement covers a customer-facing product.
- OpenHands. Has a cloud API for coding tasks, and its open-source version can be self-hosted.
Tied to One Person's Account
- Codex cloud has no documented public REST API. An experimental
codex cloud execcommand can submit tasks from a script. For an API, use OpenAI's Agents API. - Claude Code on the web can be triggered by an experimental routines endpoint that returns a session link. It runs on your Claude subscription.
- Jules has an API that's still in alpha, with keys tied to one account.
How Results Come Back
If an API only supports polling, you'll be running a scheduler to check on every agent. If it has webhooks, keep a way to poll anyway, since Claude Managed Agents drops a webhook after three failed deliveries. Webhooks can also arrive twice or out of order, so deduplicate on the event ID and fetch the agent's current state instead of relying on the order.
Same Task, Two APIs
We gave Claude Managed Agents and OpenAI's Agents API the same task: write a small Python module and its tests, run pytest, and report. Both finished and their tests passed.
Claude finished faster and cost less in this run, and OpenAI needed fewer calls to get started. Claude's four calls include creating the agent and environment, which you can reuse for later tasks. At standard rates, GPT-6 Astra's input and output tokens cost five times Claude Sonnet 5.5's, OpenAI's run also sent about three times as many input tokens, and its container bills a five-minute minimum. We used different models and ran each once, so this is one sample.
What to Test Before Production
- A run that never finishes. Set a deadline for each run, and check what cancelling it stops and whether compute keeps billing.
- A launch that times out. It may have worked. Check whether the agent exists before starting another. Cursor accepts your own agent ID so a retry returns a conflict instead of a duplicate, but you can't combine that with per-session environment variables.
- Credentials inside the sandbox. Proxies that inject secrets can fail or strip them, so test with the exact secrets setup you'll use in production.
- Commits and tokens. Check what the agent's GitHub token can push and whose name ends up on the commit.
- A follow-up after the sandbox is gone. Ramp's background agent snapshots the workspace when it finishes and restores it if a user replies later. Check what your API keeps between messages.
FAQ
What is a cloud agent API?
It's an API that starts an AI agent, usually a coding agent, on a machine in the cloud. You send a task and often a repo, and get back a stream of progress, a result, or a pull request. Cursor calls its version Cloud Agents, GitHub calls it the Copilot cloud agent, and Anthropic and OpenAI call theirs managed agents.
Can I use Claude Code programmatically for my own users?
Only if you run Claude Code unmodified and each user signs in with their own Claude account, API key, or cloud provider credentials and pays for their own usage, unless you have a separate agreement with Anthropic. You can't share your own subscription or pay for their usage. If your product pays, use the Claude Agent SDK with your API key or Claude Managed Agents, both under Anthropic's commercial terms.
Does Codex cloud have an API?
There's no documented public REST API. The experimental codex cloud exec command can submit tasks from a script, under your ChatGPT account. For an API, use OpenAI's Agents API, which runs the same Codex harness.
What happened to Cursor's Background Agents API?
Cursor renamed Background Agents to Cloud Agents. The current Cloud Agents API is in beta and streams progress, and webhooks are still only on the older version.
Can I run cloud agents on my own infrastructure?
Partly, on most platforms. Claude Managed Agents and OpenAI's Agents API can run the sandbox on your machines, while they keep running the agent and storing the transcript. Cursor and Devin can run work on your machines too. Omnara can run the whole platform on your infrastructure.