AI Agent Run Budgets: When to Stop and Escalate
Set separate turn, tool, and time limits for a production support agent, then return evidence and an explicit incomplete status when its budget runs out.
Short answer: give a production support diagnostic agent separate limits for model turns, tool calls, and elapsed time. Define what counts as progress and what happens when any limit is reached. Stop gathering evidence, record why the run ended, and hand the collected facts and unanswered questions to the support engineer. Reaching a budget is an incomplete outcome, even if the agent produces a polished summary.
Use the limits already available in your runtime. Choose their values from representative workflow evaluations and the time an operator can wait. There is no universal number of turns or seconds that makes an agent production-ready.
| If you only read one section | Start here |
|---|---|
| You need to define the controls | The run-budget worksheet |
| You are choosing configuration or custom work | Native controls versus custom integration |
| You need a release decision | Test the stopping path before increasing the allowance |
Why a turn limit is only one boundary
The OpenAI Agents SDK defines a turn as one model invocation, including the tool calls that may occur in it. Its runner normally raises MaxTurnsExceeded when the configured limit is exceeded. A turn counter therefore measures something different from the number of tool executions or the time spent waiting for one. (OpenAI Runner reference)
LangGraph similarly reports GRAPH_RECURSION_LIMIT when a graph reaches its step limit before a stop condition. Its documentation notes that a cycle can cause this, but a legitimate complex graph can also reach the limit. Investigate the trace before increasing it. (LangGraph recursion-limit guidance)
For a support agent that checks an account's integration status, one successful tool call may reveal the answer. Another case may need several independent lookups. Counting calls alone does not tell you whether the case is resolved. Write the success condition first: for example, return verified connection status, relevant error evidence, and an actionable next investigation step, or explicitly hand off what could not be checked.
A run-budget worksheet for support diagnostics
Use this table as a design worksheet. The proposed controls are recommendations for this workflow, not vendor defaults or measured benchmarks.
| Boundary | What to decide before release | What happens at the boundary |
|---|---|---|
| Model turns | The maximum reasoning iterations justified by evaluated cases | Stop the loop and record a turn-limit outcome |
| Tool calls | Total calls and any tighter limit for a frequently repeated lookup | Block further calls under that limit; do not silently reset it |
| Tool deadline | How long each external lookup may wait | Record the missing evidence and follow the defined failure path |
| Overall task deadline | How long the support workflow may remain active, including its waits and retries | Close the diagnostic task as incomplete and notify its owner |
| No progress | Which repeated results or missing prerequisites mean another iteration has no useful next step | Ask for the specific missing input or hand off |
| Final handoff | Which evidence, unknowns, and stop reason an operator needs | Return a structured partial result without claiming resolution |
A useful no-progress rule might detect the same normalized lookup against the same data version returning the same result repeatedly. That is a workflow-specific hypothesis to test. Repetition can be legitimate when polling for an expected state change; that path needs its own bounded waiting rule.
Keep one budget across the support task
Consider a read-only agent investigating why a SaaS customer's integration stopped syncing. It can read the connection status, inspect a permitted error window, and retrieve the matching troubleshooting guide. It cannot change the integration configuration in this workflow.
At task creation, record a stable case identifier, a deadline, permitted diagnostic steps, and the terminal outcomes. Before dispatching each lookup, check the remaining time and applicable call allowance. Make the lookup timeout fit within the remaining task time, leaving room to record the outcome. These checks belong in application or workflow logic; asking the model to remember its allowance is insufficient as an enforcement mechanism.
Document whether a tool-call counter includes retries inside its adapter. Record each attempt so the diagnostic owner can distinguish repeated dispatches from one slow lookup.
Preserve that task budget when a worker retries or an operator resumes the same investigation. A fresh invocation should not automatically grant the case a fresh overall deadline. LangChain's documented middleware distinguishes per-invocation call limits from limits across a thread; thread limits require checkpointed state. Confirm that the chosen scope matches your support task rather than assuming a conversation and a case are identical. (LangChain prebuilt middleware)
Keep attempt and total duration separate. Temporal documents Start-To-Close as a limit on a single Activity Task Execution and Schedule-To-Close as the overall Activity Execution duration across its attempts. Those are Activity boundaries; your full diagnostic workflow may contain several Activities and still need its own overall deadline. (Temporal Activity timeouts)
Native controls versus custom integration
Start with the platform you operate today. OpenAI's SDK offers a max_turns error handler that can return a controlled final output. LangChain supplies model-call and tool-call limit middleware. Its tool limiter's default continuation behavior blocks exceeded calls while allowing the agent to continue, so verify the configured exit behavior instead of assuming every limit ends the run. (OpenAI running agents, LangChain limit middleware)
For work already orchestrated in AWS Step Functions, task timeouts and heartbeat intervals are native controls to evaluate. Heartbeats do not extend a task's maximum allowed duration. A heartbeat indicates worker liveness; it does not establish that the diagnostic question has been answered. (AWS Task state, AWS SendTaskHeartbeat)
Those controls may be sufficient when they cover the required scopes and the existing application already records incomplete outcomes. A native or no-code support workflow can also be sufficient if it exposes the needed limits and passes the failure tests below. Add custom integration only for a demonstrated gap, such as a case-wide allowance spanning several runs, a domain-specific no-progress rule, or a handoff record the platform cannot produce.
Do not assume a timeout or stopped runner cancels every downstream operation. Verify cancellation with the actual tool adapter and worker. Define how late results are handled so they cannot silently turn a closed, incomplete case into a successful one. If the agent later gains write tools, review retry-safe tool writes before treating a timeout as proof that no change occurred.
Test the stopping path before increasing the allowance
Compare candidate limits on the same representative cases. Record task completion, premature stops, repeated lookups, elapsed time, and operator corrections. A larger limit is useful only when it enables legitimate work that the smaller limit cuts off. Fix repeated failures or missing inputs at their source.
| Evaluation case | Required observable result |
|---|---|
| Ordinary diagnostic case | Verified evidence and the defined completion condition |
| Lookup that never responds | Bounded wait, explicit missing source, and an incomplete outcome |
| Repeated unchanged results | Tested no-progress handling instead of an indefinite loop |
| Valid case requiring more steps | A trace showing useful new evidence at each additional step |
| Retry or resume near the deadline | The original task allowance remains enforceable |
| Tool result arriving after closure | A defined late-result policy preserves the recorded outcome |
Reserve a bounded path for the handoff. If another model call is used to summarize evidence, it needs its own allowance. Provide a deterministic fallback record if that call cannot finish: case ID, stop reason, completed checks, source references, missing evidence, and the next owner. Keep the machine-readable outcome authoritative so fluent prose cannot label a budget-exhausted run as resolved.
Use the AI agent evaluation guide to connect these cases to a release decision and the production readiness checklist for the wider operating bar. If the task must survive long waits or restarts, background job execution addresses that separate lifecycle decision.
Where Zenovae fits
For a live or pre-launch agent whose traces show repeated loops, premature stops, or unclear handoffs, the Agent Reliability Audit is the relevant review path. Bring the task definition, runtime configuration, and representative failed runs. If your team can already evaluate and tune those controls, the audit versus observability comparison can help decide whether external review is useful.
The worksheet above is editorial guidance. It does not establish safe limits for your workload, certify a deployment, or promise a particular completion rate. Re-run the stopping cases when the model, tools, orchestration, or task scope changes.
Sources
Need Help with Your AI Project?
At Zenovae, we build production-ready AI systems that scale. From OpenClaw setup to custom integrations, Mission Control workflows, and full-stack delivery, we can help you ship faster and avoid costly mistakes.
Let's Talk