How to Deploy an AI Agent to Production in 2026
What changes when an agent serves real users, and what we saw when we crashed one
On this page
In production, an agent serves other people, runs for hours, waits on approvals, spends money, and gets killed by deploys while nobody is watching.
This guide covers what has to change to run AI agents in production, the defaults that catch people, and the ways to deploy, with a setup I'd start from. We also crashed an agent on purpose to see what survives.
The Short Version
What Happened When We Crashed an Agent
We ran the Claude Agent SDK in Docker with Claude Sonnet 5.5. The agent had two tools: one that sends an email to a customer and logs it, and one that waits on a slow report job. We told it to send one email, wait for the report, and finish. Then we killed the container at different points, started a new one, and told the agent to continue.
In the mid-send case, the resumed agent recognized that the outcome was uncertain and asked a person. Without the transcript, our job started over and sent the email again.
Saving the transcript took more than one folder. Claude Code keeps sessions in ~/.claude, but its config lives in ~/.claude.json, outside that folder. Mount both, or set CLAUDE_CONFIG_DIR to a volume so Claude Code keeps everything under it.
The hard case is a crash between doing something and recording it. In our run, the agent stopped and asked, but the job still needed a person. Give tools that change things an idempotency key that the receiving service checks, or a way to check whether the action already happened, so the agent can continue on its own.
This was one run of each case, so other models and versions may behave differently.
Sessions That Survive Restarts
The agent's state has to live outside the process running it. There are a few ways to do that:
- A durable execution engine, like Temporal or the Workflow SDK, records each step so a run resumes after a crash.
- A checkpointer in your own database. LangGraph's in-memory checkpointer loses everything on restart, so production setups use the Postgres checkpointer. It doesn't clean up old checkpoints, so plan for that.
- The harness's own session store. The Claude Agent SDK can mirror transcripts to S3, Redis, or Postgres, but the mirror is best-effort and doesn't cover files on the machine.
Resuming usually runs some code again. LangGraph restarts the interrupted node from the beginning, so anything before the pause has to be safe to repeat.
Sandboxes and Isolation
Run untrusted code in a sandbox, usually one per session. You can keep the agent loop in your backend and have it run code through sandbox tools, which is how Vercel's Open Agents template works. The Claude Agent SDK also supports running the whole loop inside a container.
Permission prompts don't contain what the agent runs, so put a container or microVM around anything it executes, limit where it can connect to, and keep credentials out of it. Stripe's coding agents run on dev machines that are cut off from production and the internet, so they run without approval prompts at all. On GitHub-hosted runners, GitHub's agent runs behind a firewall and can only push to its own branch.
Cold starts matter once people are waiting. Ramp keeps a pool of warm sandboxes and starts one as soon as a user begins typing.
Many Users: Identity and Credentials
When an agent acts for a user, it needs that user's access, not yours. Claude and OpenAI can keep each user's credentials in a vault and inject them outside the sandbox, so the agent only sees a placeholder. That injection doesn't work with self-hosted sandboxes yet. In other setups, tokens reach your tool code or the runtime directly, so check where credentials become readable before running untrusted code. AgentCore Identity and some Foundry tools can run the OAuth sign-in and store tokens for you. Your app still has to connect each sign-in to the right user.
None of this checks who owns what. AgentCore, for example, doesn't enforce which user a session belongs to, and leaves that mapping to your backend. Whatever platform you use:
- Take the tenant from the authenticated request, not from anything the client sends.
- Check ownership when someone starts, reads, streams, resumes, or approves a run.
- Test that one tenant can't reach another tenant's sessions, files, memory, or credentials.
- Set concurrency and spending limits per tenant.
Memory needs the same care. Check which memory stores each user can attach to a session, and mount shared reference material read-only. A prompt injection can write to memory the agent can change, and later sessions read it as trusted.
Approvals That Can Wait for Days
Use a pause that outlives the process, like LangGraph interrupts or the Claude Agent SDK's defer decision from a PreToolUse hook. Then someone can approve hours later and a different process can resume the run. defer only works when Claude makes one tool call in that turn, and Claude Code deletes saved sessions after 30 days by default.
Save these alongside each pending approval:
- The run state at the pause, in durable storage
- The exact pending call, with its tool name and arguments, so the reviewer approves what will run
- Who decided, what they decided, and when
- An approval record that only one worker can claim, plus a stable action ID the tool uses to reject a duplicate
- The version of your code, since a run paused on Monday may resume into Friday's deploy
Don't rely on notifications to find pending approvals. Claude Managed Agents drops a webhook after three failed deliveries, so keep a list you can poll. When a reviewer says no, record it and enforce it in the tool, then tell the model so it can pick another approach.
Cost Limits
Set explicit limits on model turns, tool retries, and elapsed time for every agent. Framework defaults range from a handful of turns to none, they change between versions, and a framework "step" isn't always a model call. Test your limits on the version you deploy. Stripe lets its coding agents fix CI failures once and then stop, because more rounds rarely help.
Provider spend caps lag behind usage. OpenAI says its spend limits aren't instantaneous, Gemini's billing data can lag behind its project caps, and calls already running can push past any cap. Claude Managed Agents has a per-session dollar cap that stops a session between model calls. Stop accepting new work before a budget runs out.
One startup was charged $82,314 in 48 hours after a Gemini key was stolen. Keep keys on the server, restrict them to the APIs they need, and alert when the cost per finished task goes up.
Tracing
Trace every model call and tool call, with the user, the agent version, and the cost attached. The OpenTelemetry conventions for AI are still in development, and the defaults differ. The Claude Agent SDK's telemetry leaves out conversation content unless you opt in, though its local transcripts still keep it. The AI SDK records inputs and outputs by default once telemetry is on, and the OpenAI Agents SDK sends traces to OpenAI by default. Set what gets recorded and where it goes on purpose, especially if users' data passes through.
Ways to Deploy
The first decision is whether you're deploying an agent you already built, or adopting a platform that brings its own agent loop.
Self-host a framework agent
Run your agent on a framework with a durable layer under it:
- LangGraph with the Postgres checkpointer behind your own API and worker. LangChain's own server needs an Enterprise plan to self-host.
- OpenAI Agents SDK inside Temporal, where model calls become retried, recorded steps.
- Vercel AI SDK with
WorkflowAgenton the Postgres backend outside Vercel. - Mastra with durable agents for runs that have to survive restarts. They're in beta, so pin the version and test recovery before upgrading.
Some of these include authentication or sandbox integrations, but you still configure them and check the access rules and spending limits for your app.
Self-host a harness
Run a harness like the Claude Agent SDK or Codex on your own machines. Anthropic's hosting guide is clear that this isn't like hosting a normal API. Each active session runs its own CLI process, and its transcript stays in local files unless you add a session store. You can give each session its own container or run several in one, depending on what those processes can reach. Everything from the crash test applies here.
Host your agent on a platform
LangSmith Deployment, Google's Agent Runtime, and AWS AgentCore can host an agent you built with a framework. They run it and scale it, while the agent's logic stays yours.
Use a managed agent platform
Claude Managed Agents, OpenAI's Agents API, and Omnara bring their own agent loop, session log, sandboxes, and approvals behind an API. You configure the agent instead of writing the loop. Omnara's core platform is open source, so you can use it hosted or self-host it, which also means owning upgrades, backups, and recovery. I compared these platforms in more detail in Claude Managed Agents alternatives.
Even with a platform, check who maps your users to their credentials, what happens when a webhook is missed, and where the transcripts are stored.
A Setup to Start From
This is the architecture I'd start with. Ramp and Stripe describe similar setups.
- An API that authenticates users, starts runs, and streams progress, with a way to reconnect to a stream.
- A durable run layer, either a workflow engine or a job queue in your database, with step and spend limits per run.
- A session log outside the worker, holding events, checkpoints, and pending approvals.
- A sandbox per session for anything the agent executes, with limited network access.
- Credentials kept out of the sandbox, injected at a proxy or in your tool code.
- Tracing from the first day, with the user and version on every run.
For a small team, deploy an authenticated API and one worker, with Postgres as both the session log and the job queue. Save each job before you return its run ID, and store checkpoints outside the worker. Add a hosted sandbox provider if the agent runs code. Before launch, test a worker crash, a stream reconnect, a duplicate tool call, and recovery after a deploy. Keep a small set of real tasks with checks you wrote yourself, and run them before changing the model, prompt, tools, or harness.
FAQ
How do I deploy a LangGraph agent without LangSmith?
You can deploy the open-source LangGraph library on its own. Run it behind your own API and worker with the Postgres checkpointer, and check that the caller owns a thread before resuming it. Paused approvals resume from the saved checkpoint. Add a job that deletes old checkpoints. LangChain's ready-made server, LangSmith Deployment, needs an Enterprise plan to self-host.
How do I run the Claude Agent SDK in production?
Keep each session's transcript and config on a volume or in a session store, put the sessions in containers with limited network access, and keep credentials out of them. Set your own turn, retry, and spend limits. Use the defer decision if approvals can take longer than the process stays up.
How do AI agents wait for human approval?
The framework saves the paused run, including the exact tool call waiting for approval, and resumes it when someone decides. The run can resume in a different process days later, as long as the saved state, the pending call, and your code version are all still there.
How do I stop an agent from running up a huge bill?
Set limits on turns, retries, and time for every run, and a budget per user if users can start agents. Use provider spend caps as a backstop, since they lag. Keep API keys on the server and restricted to the APIs they need.
What are my options for a self-hosted AI agent platform?
If you want a platform with its own agent loop, evaluate Omnara. If you have a LangGraph agent and want LangChain's server for it, LangSmith's self-hosted deployment needs an Enterprise license. For each, check the production setup, the upgrade process, backups, and which APIs it supports before choosing.