Back to Blog
    AI Agent Architecture
    September 30, 20268 min

    When Should an AI Agent Run as a Background Job?

    Keep short agent calls synchronous. Move multi-step work that can wait, resume, or pause for review into a durable job with a clear status contract.

    AI Agent ArchitectureDurable ExecutionBackground JobsProduction AIIncident Response

    Short answer: keep an agent inside the request when the work is bounded, finishes within the caller's deadline, and the caller needs the result immediately. Return a job ID when the work may outlive the connection, wait for a person or external event, or needs to resume after a worker restart. Use a durable workflow runtime when the job has multiple steps whose progress must survive beyond one model response.

    There is no universal number of seconds that separates the two. The decision depends on the request deadline, failure recovery, and whether the workflow needs its own status and lifecycle. A background job is an execution boundary, not a reason to make every agent asynchronous.

    Quick take: a single long model response may fit a provider's native background mode. A multi-step run that calls other systems, pauses for approval, or must resume after a deployment needs durable application state. Return a status resource when the client should stop waiting; do not imply that an accepted job has already completed.

    Start from the caller's contract

    Imagine an incident-response assistant that receives an alert, gathers service metadata and relevant runbook passages, drafts a timeline and likely causes, then waits for an on-call engineer to review the proposed ticket update. The model's summary is only one step. The run also has tool calls, a review pause, a destination system, and a final status to communicate.

    If the alert and required context are already available and a short answer can be returned within the caller's request budget, a synchronous response keeps the interface simple. If a tool is slow, the run must wait for approval, or the caller could disconnect before the result is ready, create a job resource and let the caller check it later.

    HTTP's 202 Accepted status means the request was accepted for processing, not that processing finished. RFC 9110 says the response should describe the current status and point to a status monitor. In practice, return an opaque job ID and an authorized status URL, then report explicit states such as queued, running, waiting_for_review, succeeded, failed, or cancelled. Treat those state names as an application contract, not a universal standard.

    Choose the smallest execution model that fits

    Execution modelUse it whenWhat the caller receivesExample for incident response
    Synchronous requestWork is bounded, no person or outside event must be awaited, and the response is needed immediatelyThe completed result or a clear error in the same requestSummarize an already-loaded incident excerpt
    Provider background responseOne model response may outlast the caller's connection, but the application does not need a durable multi-system workflowA provider response ID and a way to check its statusGenerate a long incident analysis, then retrieve the completed response
    Durable workflow jobWork spans tools or services, can pause for a person, or must retain step progress after process failureAn application job ID, status resource, and a defined owner for completionGather evidence, draft a ticket update, wait for review, then post the approved update

    OpenAI's Responses API documents a native background option: create a response with background enabled and poll its response resource while it is queued or in progress. That is a reasonable first option when the unit of work is one model response. The same documentation says that store=false does not prevent temporary storage needed for background execution and polling, so review that behavior against the workflow's data requirements before using it. This is a product-specific data-handling note, not a broader privacy or compliance conclusion. (OpenAI background mode)

    For a multi-step job, a workflow engine can own the orchestration while the agent remains one activity in the workflow. AWS describes Step Functions Standard Workflows as a durable option for long-running work and documents different duration and execution-history behavior across Standard and Express types. Microsoft Durable Task documents checkpointing and recovery for long-running workflows, including agent workflows that wait for people. Temporal documents resuming workflow progress after failures and human waits. These are examples of capabilities to compare; they do not establish that one runtime fits every stack or workflow. (AWS workflow types, Microsoft Durable Task for AI agents, Temporal durable AI)

    Keep job progress separate from agent memory

    A job's durable state answers, “Where is this specific incident run, and what is the next authorized step?” It can hold a run ID, workflow version, completed step references, approval status, and the last safe checkpoint. It is not a long-term memory store, a source of current service ownership, or a replacement for the incident system of record.

    That boundary is similar to the distinction between workflow progress and cross-run agent memory: load current ownership and policy from their authoritative systems, and persist only the state needed to resume this job. A durable runtime also does not make every external write repeat-safe. Keep each ticket, paging, or notification action safe under its documented retry behavior; the separate tool-retry guide covers ambiguous outcomes and duplicate writes.

    When native background handling is enough—and when custom orchestration is justified

    Start with the native option already available in your stack. A provider's background response mode may be enough for one model call if its status, cancellation, storage, and access behavior fit. An existing worker queue or workflow service may already provide the job ID, status endpoint, durable storage, and operator controls your application needs. If it does, use that path rather than building a second orchestration layer.

    Use a durable workflow runtime when the run crosses service boundaries, can wait for hours or days, needs a human signal to continue, or must resume from named steps after a worker restarts. Add custom orchestration only for a documented gap—for example, if the existing runtime cannot express the required approval, authorization, cancellation, or recovery behavior. The production agent build is the relevant hub when that gap means a bounded agent workflow needs to be built and operated; AI integrations covers connecting the agent to the systems that own incident data and actions.

    Check each platform's actual limits and configuration. For example, AWS documents execution-history retention for Step Functions Standard Workflows as 90 days after completion; a workflow that needs longer business history must retain its own record. (AWS execution details)

    Release checklist for an asynchronous agent job

    • Give each run a stable job ID and record the workflow version that started it.
    • Return an authorized status URL and define every nonterminal and terminal state.
    • Persist step progress and outputs needed for resume; keep authoritative incident data in its owning system.
    • Define which person or service owns a waiting_for_review job and how that job resumes or expires.
    • Set cancellation, retry, and timeout behavior per step. Make side effects repeat-safe or reconcile their outcome before retrying.
    • Test a worker restart before and after an external action, plus a caller that disconnects and checks status later.
    • Review payload storage, retention, and access controls for both the model provider and workflow runtime.
    • Measure real queue age, completion time, failures, and human wait time before changing the execution model.

    Use the production-ready agent checklist for the broader release bar. If the agent is already live and the team needs an independent reliability score, review the Agent Reliability Audit after this execution-boundary decision is clear.

    Sources

    Need Help with Your AI Project?

    At Zenovae, we build production-ready AI systems that scale. From OpenClaw setup to custom integrations, Mission Control workflows, and full-stack delivery, we can help you ship faster and avoid costly mistakes.

    Let's Talk