Self-Hosted Managed Agents in 2026
What Claude, OpenAI, AWS, and open-source platforms let you run on your own infrastructure
On this page
A "self-hosted" managed agent can mean very different things. Sometimes only the commands run on your machines, while the agent loop, the model calls, and the transcript of everything the agent saw and did stay in the vendor's cloud. Sometimes you run the whole platform.
Either can be the right fit, depending on why you want to self-host. This guide sorts the options by what runs where, as of October 9, 2026, and includes a test where we self-hosted a full platform ourselves.
The Short Version
What You Can Self-Host
In the first two options, a local sandbox doesn't keep your code on your network. When the agent reads a file, the contents come back as a tool result, and that result goes to the model and into the vendor's transcript. A self-hosted sandbox controls where commands run and what the agent can reach. It doesn't control where the content returned to the agent ends up.
Your Machines Run the Tools
Claude Managed Agents with self-hosted sandboxes
You run a worker on your own Linux host that picks up work from a queue, runs each tool call locally, and posts the result back over an outbound connection. You operate the workers and their network access. On Anthropic's own platform, the worker authenticates with an environment key you create in the Claude Console. On Claude Platform on AWS it uses AWS authentication, though Anthropic still runs the service.
Anthropic still runs the agent loop and stores the transcript, and memory stores stay hosted by Anthropic with a synced copy on your machine. Tool inputs and outputs pass through Anthropic, and the data retention table lists self-hosted sandboxes as not eligible for Zero Data Retention or HIPAA. A few features don't carry over. Anthropic won't mount files or GitHub repos into your sandbox, and the credential type that injects secrets into sandbox requests isn't supported yet.
OpenAI's Agents API with your own environment
You run codex exec-server inside your environment, and it connects out to OpenAI to receive commands and return results. Each session needs its own executor, authenticated with a restricted environment key from the OpenAI dashboard.
OpenAI runs the harness and the models and keeps the session state. Its docs say "choosing a self-hosted sandbox does not make the Agents API ZDR-eligible," and data residency is US only. Files from self-hosted environments don't show up in the artifacts API, and vault credentials for sandbox code only work in OpenAI-hosted environments, so you'd bring your own proxy for secrets.
Cursor and Devin
Cursor's self-hosted machines keep the agent loop and models in Cursor's cloud while your worker edits files and runs commands. Cursor's docs list what the worker sends back during a run: file contents, terminal output, diffs, and screenshots. Team pools need an Enterprise plan.
Devin Outposts works the same way, with Devin's agent loop in its cloud and command execution on your machines. Devin doesn't install firewall rules on outposts, so network limits are yours to enforce.
The Agent Runs on Your Machine
Claude Code self-hosted runners
Claude Code in your terminal or IDE already runs on your machine. Self-hosted environments are for Claude Code's cloud sessions, the ones started from the web, the desktop app, or your phone. A runner on your machine claims the session, clones the repo, and starts Claude Code there, so the checkout, build files, and secrets stay on your hardware unless the agent reads them into the conversation.
Anthropic still runs the queue and the interface, and the conversation, including prompts, responses, and tool results, goes to Anthropic, which stores the transcript. Model calls go to Anthropic unless you point the runner at Amazon Bedrock or Google Cloud. It's in beta on Team and Enterprise plans, and it isn't available to organizations with Zero Data Retention or the HIPAA configuration.
GitHub Copilot on your own runners
Copilot's cloud agent can run on self-hosted GitHub Actions runners, ideally short-lived ones managed by Actions Runner Controller. The agent runs inside the Actions job on your runner, calls GitHub's model API, and its session log stays on GitHub. You have to turn off Copilot's built-in firewall for self-hosted runners, so network limits become your job here too.
Your Cloud Provider's Managed Service
If the requirement is your cloud provider's permissions, networking, and data residency, rather than servers you operate, look at your provider's own agent service. Check which resources run in your account and which run in the provider's.
- Amazon Bedrock Managed Agents, powered by OpenAI, runs OpenAI's agent loop and models inside Amazon Bedrock, with AWS IAM for access. Commands run on your own hosts or in AgentCore. It's in preview.
- AgentCore Instances run your agent on EC2 machines that AWS provisions and operates inside your own account, for long-running sessions.
- Microsoft Foundry hosted agents run in Microsoft-managed, VM-isolated sandboxes that you can connect to your own Azure network. Its standard setup keeps conversations, files, and search indexes in your own Cosmos DB, Storage, and AI Search.
- Google's Agent Runtime runs agent code you bring in a Google-managed project that reaches your network through Private Service Connect, with regional data residency, customer-managed keys, and HIPAA support. Google's Managed Agents API, in preview, supplies an agent and sandbox instead.
These are still managed services, and some features, like AgentCore's memory, are separate regional services with their own data rules.
You Run the Whole Platform
Here you host the agent loop, its records, and the credentials. The models, tools, and sandbox services you connect can still receive data, and some platforms report license or usage data back to the vendor, so the boundary depends on what you plug in.
- Omnara is open source under Apache 2.0 and lets you host the agent loop, the event history, and the API on your own Postgres, Redis, object storage, and a shared file system. Agents can use machines you connect or sandboxes from providers like Modal and Daytona, with any model, including one you host.
- LangSmith offers a full self-hosted platform and a smaller standalone Agent Server, and both need an Enterprise license. Without an air-gapped license, it still sends billing data back to LangChain.
- OpenHands is MIT licensed. Its Kubernetes deployment is a shared, single-tenant app without built-in login or user-level access control, which suits an individual or a trusted team. OpenHands Enterprise adds authentication and tenant isolation under a source-available license.
- Coder Agents runs the agent loop and stores chat history in your own Coder deployment, with Coder workspaces as sandboxes. The community license limits how many agents run at once, and it reports agent runtime to Coder for billing.
Visual builders like n8n and Dify can also be self-hosted, and I covered them in Four Ways to Build an AI Agent. Dify's license limits multi-tenant use.
Our Omnara execution and storage test
We ran Omnara from its published images with one docker compose command, connected a model key, and connected a Linux container as the agent's machine through Omnara's daemon, which only makes outbound connections. Then we gave an agent a task: report the machine's hostname and OS, and write a file with the current time.
It finished in 9 seconds. The file was on our machine, and every event in the run, including the commands the agent ran, was stored in our own Postgres. During the run, the only traffic that left our infrastructure was the call to the model.
This was a first check of where execution and storage happen. Test recovery, concurrent work, and isolation separately before production.
What Full Self-Hosting Costs
Running the platform yourself means running everything the vendor was running:
- Databases and storage, with backups that restore together. LangSmith warns that restoring Postgres, ClickHouse, and blob storage to different points leaves broken references. Omnara needs its memory file system backed up alongside Postgres, and on several hosts it has to support shared file locks.
- Upgrades across every service and every machine's daemon, tested on your own setup.
- Recovery testing. Kill a worker during a tool call, and check what resumes and what runs twice. I wrote more about this in How to Deploy an AI Agent to Production.
- On-call. Your team handles stalled sessions and service failures.
Compare the cost per finished task, including model calls, machines, storage, licenses, and the time your team spends running it.
How to Choose
- If agents need to reach internal systems and your security team is fine with prompts and tool results going to the model vendor, a self-hosted sandbox or worker is enough. Use the one from the vendor you already use. The worker only connects out, but it still runs whatever the hosted agent sends, so give it a narrow machine identity and limit which internal services it can reach.
- If you need Zero Data Retention, Claude Managed Agents, OpenAI's Agents API, and Claude Code's cloud sessions don't support it, even when the sandbox or runner is yours. Check HIPAA eligibility separately for each service and feature.
- If the requirement is your cloud provider's controls, look at Bedrock Managed Agents, AgentCore, Foundry, or Google's Agent Runtime, and check which parts run in your account.
- If transcripts have to stay on your infrastructure, choose a platform that lets you host them, and a model you host if prompts can't leave either.
- Check where credentials live. Some hosted agents can use credentials held by a tool service or proxy you run, so secrets stay with you even when the agent loop doesn't.
FAQ
Can I self-host Claude Managed Agents?
Partly. You can run the sandbox on your own machines with a self-hosted worker, but Anthropic still runs the agent loop and stores the transcript and memory. Self-hosted sessions aren't eligible for Zero Data Retention or HIPAA.
What is a Claude Code self-hosted runner?
It's a process on your machine that claims Claude Code cloud sessions, the ones started from the web, desktop, or mobile apps, and runs them on your hardware. Anthropic still handles dispatch and stores the transcript. Claude Code in your terminal or IDE already runs locally, and for agents inside your own app, the Claude Agent SDK runs a process you operate.
Can I self-host OpenAI's Agents API?
Only the sandbox. You run codex exec-server in your environment, and OpenAI runs the harness, the models, and the session state. Self-hosting the sandbox doesn't make it eligible for Zero Data Retention.
What are the open-source managed agent platforms?
Consider Omnara for a shared agent service with any model and your own machines, and OpenHands or Coder Agents for coding workflows. Check the license of the edition you need, since LangSmith's self-hosted version needs an Enterprise license and OpenHands' multi-user version isn't open source.
Does a self-hosted sandbox keep my code on my network?
Not entirely. File contents and tool output returned to a hosted agent leave your network, go to the model, and on most platforms land in the vendor's stored transcript. Check each service's storage, retention, and training policies.