Back to Blog
    AI Agent Reliability
    October 11, 20267 min

    AI Agent Run Budgets: When to Stop and Escalate

    Set separate turn, tool, and time limits for a production support agent, then return evidence and an explicit incomplete status when its budget runs out.

    AI Agent ReliabilityRun BudgetsStop ConditionsProduction AISaaS Support

    Short answer: give a production support diagnostic agent separate limits for model turns, tool calls, and elapsed time. Define what counts as progress and what happens when any limit is reached. Stop gathering evidence, record why the run ended, and hand the collected facts and unanswered questions to the support engineer. Reaching a budget is an incomplete outcome, even if the agent produces a polished summary.

    Use the limits already available in your runtime. Choose their values from representative workflow evaluations and the time an operator can wait. There is no universal number of turns or seconds that makes an agent production-ready.

    If you only read one sectionStart here
    You need to define the controlsThe run-budget worksheet
    You are choosing configuration or custom workNative controls versus custom integration
    You need a release decisionTest the stopping path before increasing the allowance

    Why a turn limit is only one boundary

    The OpenAI Agents SDK defines a turn as one model invocation, including the tool calls that may occur in it. Its runner normally raises MaxTurnsExceeded when the configured limit is exceeded. A turn counter therefore measures something different from the number of tool executions or the time spent waiting for one. (OpenAI Runner reference)

    LangGraph similarly reports GRAPH_RECURSION_LIMIT when a graph reaches its step limit before a stop condition. Its documentation notes that a cycle can cause this, but a legitimate complex graph can also reach the limit. Investigate the trace before increasing it. (LangGraph recursion-limit guidance)

    For a support agent that checks an account's integration status, one successful tool call may reveal the answer. Another case may need several independent lookups. Counting calls alone does not tell you whether the case is resolved. Write the success condition first: for example, return verified connection status, relevant error evidence, and an actionable next investigation step, or explicitly hand off what could not be checked.

    A run-budget worksheet for support diagnostics

    Use this table as a design worksheet. The proposed controls are recommendations for this workflow, not vendor defaults or measured benchmarks.

    BoundaryWhat to decide before releaseWhat happens at the boundary
    Model turnsThe maximum reasoning iterations justified by evaluated casesStop the loop and record a turn-limit outcome
    Tool callsTotal calls and any tighter limit for a frequently repeated lookupBlock further calls under that limit; do not silently reset it
    Tool deadlineHow long each external lookup may waitRecord the missing evidence and follow the defined failure path
    Overall task deadlineHow long the support workflow may remain active, including its waits and retriesClose the diagnostic task as incomplete and notify its owner
    No progressWhich repeated results or missing prerequisites mean another iteration has no useful next stepAsk for the specific missing input or hand off
    Final handoffWhich evidence, unknowns, and stop reason an operator needsReturn a structured partial result without claiming resolution

    A useful no-progress rule might detect the same normalized lookup against the same data version returning the same result repeatedly. That is a workflow-specific hypothesis to test. Repetition can be legitimate when polling for an expected state change; that path needs its own bounded waiting rule.

    Keep one budget across the support task

    Consider a read-only agent investigating why a SaaS customer's integration stopped syncing. It can read the connection status, inspect a permitted error window, and retrieve the matching troubleshooting guide. It cannot change the integration configuration in this workflow.

    At task creation, record a stable case identifier, a deadline, permitted diagnostic steps, and the terminal outcomes. Before dispatching each lookup, check the remaining time and applicable call allowance. Make the lookup timeout fit within the remaining task time, leaving room to record the outcome. These checks belong in application or workflow logic; asking the model to remember its allowance is insufficient as an enforcement mechanism.

    Document whether a tool-call counter includes retries inside its adapter. Record each attempt so the diagnostic owner can distinguish repeated dispatches from one slow lookup.

    Preserve that task budget when a worker retries or an operator resumes the same investigation. A fresh invocation should not automatically grant the case a fresh overall deadline. LangChain's documented middleware distinguishes per-invocation call limits from limits across a thread; thread limits require checkpointed state. Confirm that the chosen scope matches your support task rather than assuming a conversation and a case are identical. (LangChain prebuilt middleware)

    Keep attempt and total duration separate. Temporal documents Start-To-Close as a limit on a single Activity Task Execution and Schedule-To-Close as the overall Activity Execution duration across its attempts. Those are Activity boundaries; your full diagnostic workflow may contain several Activities and still need its own overall deadline. (Temporal Activity timeouts)

    Native controls versus custom integration

    Start with the platform you operate today. OpenAI's SDK offers a max_turns error handler that can return a controlled final output. LangChain supplies model-call and tool-call limit middleware. Its tool limiter's default continuation behavior blocks exceeded calls while allowing the agent to continue, so verify the configured exit behavior instead of assuming every limit ends the run. (OpenAI running agents, LangChain limit middleware)

    For work already orchestrated in AWS Step Functions, task timeouts and heartbeat intervals are native controls to evaluate. Heartbeats do not extend a task's maximum allowed duration. A heartbeat indicates worker liveness; it does not establish that the diagnostic question has been answered. (AWS Task state, AWS SendTaskHeartbeat)

    Those controls may be sufficient when they cover the required scopes and the existing application already records incomplete outcomes. A native or no-code support workflow can also be sufficient if it exposes the needed limits and passes the failure tests below. Add custom integration only for a demonstrated gap, such as a case-wide allowance spanning several runs, a domain-specific no-progress rule, or a handoff record the platform cannot produce.

    Do not assume a timeout or stopped runner cancels every downstream operation. Verify cancellation with the actual tool adapter and worker. Define how late results are handled so they cannot silently turn a closed, incomplete case into a successful one. If the agent later gains write tools, review retry-safe tool writes before treating a timeout as proof that no change occurred.

    Test the stopping path before increasing the allowance

    Compare candidate limits on the same representative cases. Record task completion, premature stops, repeated lookups, elapsed time, and operator corrections. A larger limit is useful only when it enables legitimate work that the smaller limit cuts off. Fix repeated failures or missing inputs at their source.

    Evaluation caseRequired observable result
    Ordinary diagnostic caseVerified evidence and the defined completion condition
    Lookup that never respondsBounded wait, explicit missing source, and an incomplete outcome
    Repeated unchanged resultsTested no-progress handling instead of an indefinite loop
    Valid case requiring more stepsA trace showing useful new evidence at each additional step
    Retry or resume near the deadlineThe original task allowance remains enforceable
    Tool result arriving after closureA defined late-result policy preserves the recorded outcome

    Reserve a bounded path for the handoff. If another model call is used to summarize evidence, it needs its own allowance. Provide a deterministic fallback record if that call cannot finish: case ID, stop reason, completed checks, source references, missing evidence, and the next owner. Keep the machine-readable outcome authoritative so fluent prose cannot label a budget-exhausted run as resolved.

    Use the AI agent evaluation guide to connect these cases to a release decision and the production readiness checklist for the wider operating bar. If the task must survive long waits or restarts, background job execution addresses that separate lifecycle decision.

    Where Zenovae fits

    For a live or pre-launch agent whose traces show repeated loops, premature stops, or unclear handoffs, the Agent Reliability Audit is the relevant review path. Bring the task definition, runtime configuration, and representative failed runs. If your team can already evaluate and tune those controls, the audit versus observability comparison can help decide whether external review is useful.

    The worksheet above is editorial guidance. It does not establish safe limits for your workload, certify a deployment, or promise a particular completion rate. Re-run the stopping cases when the model, tools, orchestration, or task scope changes.

    Sources

    Need Help with Your AI Project?

    At Zenovae, we build production-ready AI systems that scale. From OpenClaw setup to custom integrations, Mission Control workflows, and full-stack delivery, we can help you ship faster and avoid costly mistakes.

    Let's Talk