Back to Blog
    AI Agent Architecture
    October 10, 20268 min

    Single-Agent or Multi-Agent for Production Incident Triage?

    Start incident triage with one tool-using agent; split into specialists only when independent evidence and evaluations justify the extra coordination.

    AI Agent ArchitectureMulti-Agent SystemsIncident ResponseProduction AIAgent Orchestration

    Short answer: start with one tool-using agent for incident triage. Split work across specialist agents only when the evidence-gathering tasks are genuinely independent, their outputs can be checked and combined, and a replay evaluation shows a useful improvement. Keep the agent read-only during initial triage and leave incident severity, mitigation, and communications under the incident team's control.

    For example, one agent can read an alert, fetch the relevant runbook, inspect a recent deployment, and return timestamped evidence plus missing information. A second agent is justified when it can independently investigate a separate source, such as deployment history or service telemetry, while a manager combines both results. More agents do not automatically mean more coverage.

    First separate tools, steps, and agents

    A workflow is not multi-agent just because it has several tools or several steps. A single agent can call a metrics query, read a runbook, and check a deployment record in sequence. A multi-agent workflow has multiple model-driven roles, each with its own instructions or context, and an orchestration rule that passes work or results between them.

    This distinction matters in an outage. Another agent adds another model decision and an output that must be checked. Microsoft's architecture guidance recommends choosing the lowest complexity that reliably meets the requirement, and names coordination overhead, latency, and cost as tradeoffs of more complex agent architectures. Anthropic's engineering guidance likewise recommends starting with the simplest solution and increasing complexity when needed. Those are design principles, not measured claims about your incident workflow. (Microsoft: AI agent orchestration patterns, Anthropic: Building effective agents)

    Choose the smallest topology that passes your incident cases

    Incident-triage conditionStart withWhy
    The agent gathers a bounded set of facts from a few known toolsOne agent with read-only toolsOne owner can collect evidence and produce a single structured summary.
    The steps depend on each other, such as finding the service before querying its deployment historyOne agent or a fixed, sequential workflowThe next lookup depends on the result of the previous one; parallel agents would duplicate or guess context.
    Two investigations use separate sources and can run without shared stateA manager plus bounded specialist tasksIndependent work can run concurrently, then return evidence to one place for review.
    The right domain is unknown until evidence arrivesA narrow routing or handoff patternDelegate only after a defined signal shows that a specialist is needed.
    A proposed action can change service state or page peopleA human approval step and a separate action boundaryTriage output should not silently become mitigation authority.

    The table is a starting hypothesis. Use examples from your own incidents to decide whether a split is worth retaining.

    A concrete SRE triage workflow

    Start with one read-only investigator and a typed result. For each incident, have it return:

    • incident ID and the alert facts it received;
    • evidence items with source, timestamp, and link or query reference;
    • hypotheses separated from observed facts;
    • missing or conflicting evidence;
    • the next recommended investigation step;
    • whether a human decision or escalation is needed.

    For instance, the agent could compare an error spike with the service runbook and the deployment record for the affected time window. It should report “deployment occurred at 14:05 UTC” only when the deployment source supports that fact; “the deployment caused the spike” remains a hypothesis unless validated. If the runbook or telemetry source is unavailable, the result should say what it could not check.

    If repeated evaluations show that independent investigations are being missed or the first agent is constrained by an unsuitable context, split only those investigations. One specialist can retrieve recent changes; another can query a defined telemetry window. Return bounded results with provenance to a manager that owns the final summary. Do not let specialist outputs page, roll back, or edit incident records directly during this test.

    This preserves the human incident structure. Google's SRE guidance describes a clear incident command line, assigned roles, and a working record of debugging and mitigation. A useful production boundary is therefore: agents may gather and organize evidence; a human incident commander or delegated operator decides whether to declare, escalate, communicate, or mitigate. This is a recommendation for the workflow, not a claim that Google's process prescribes AI roles. (Google SRE: Incident response)

    When native orchestration is enough—and when custom code is justified

    Review the orchestration features already available in your runtime before building another coordinator. The OpenAI Agents SDK documents a manager pattern where specialist agents act as tools and the manager retains control, as well as handoffs where a specialist takes over. Its guidance also describes code-controlled routing and parallel execution for independent tasks. Microsoft Agent Framework documents sequential, concurrent, handoff, group-chat, and manager-led workflow patterns. These are native options to evaluate against your requirements; their presence does not establish that a particular deployment meets your access, state, or operational needs. (OpenAI Agents SDK: Orchestration, Microsoft Agent Framework: Workflow orchestrations)

    Use a built-in pattern if it can enforce the needed task boundaries and provide a traceable result. Add custom orchestration only for a documented gap, such as a missing authorization check, required checkpoint, cancellation rule, evidence schema, or audit record. If custom coordination is becoming a core production workflow, the production agent build hub explains the bounded build path. For a live system with unresolved failures, the Agent Reliability Audit is the relevant review path.

    Prove the split before release

    Compare a single-agent baseline with the proposed specialist design on the same set of past incidents or carefully constructed fixtures. Keep the source data and expected human decisions fixed. Evaluate the final workflow, not just whether each specialist can produce a plausible paragraph.

    Test caseEvidence to recordFailure that should stop rollout
    Normal incident with a known service and recent deploymentCorrect source references, timestamps, and a useful next stepUnsupported causal claim or an omitted relevant source
    Conflicting alerts or telemetryWhich sources conflict and what remains uncertainThe summary silently chooses one account
    Missing or denied sourceExplicit limitation and a safe fallbackAgent invents a lookup result or retries outside its access
    Specialist timeout or malformed outputBounded retry or controlled partial resultManager treats missing output as confirmation
    Duplicate alerts for one incidentStable incident identity and one combined evidence recordDuplicate or inconsistent summaries create a second incident
    High-severity or ambiguous caseEscalation path and human ownershipAgent declares, pages, or mitigates without the required approval

    Track evidence coverage, unsupported assertions, missed escalations, completion time, model/tool usage, and operator corrections. Pick the multi-agent design only if it improves the outcomes your team cares about enough to justify the extra coordination. Do not infer production quality from a small demo or from a vendor's general architecture example.

    OpenAI's SDK provides deterministic utilities for checking orchestration behavior such as tools, handoffs, retries, and guardrails; its documentation says to use real provider adapters or integration environments for behavior owned by external models or networks. Its tracing documentation describes traces and spans for inspecting agent runs. Use those capabilities where they fit, and review trace data handling for your environment before exporting incident content. (OpenAI Agents SDK: Testing, OpenAI Agents SDK: Tracing)

    Release checklist

    • A single-agent baseline has been tested on representative incident cases.
    • Each specialist owns a distinct, bounded evidence task.
    • Every result separates observations, source references, hypotheses, and unknowns.
    • Agent tools are read-only unless a separate approval and action path is explicitly tested.
    • Missing, stale, conflicting, and denied evidence has a defined response.
    • A human remains responsible for incident command, mitigation, and communications.
    • The chosen topology improves measured workflow outcomes over the simpler baseline.

    For the wider release bar, use the production-ready AI agent checklist and the AI agent evaluation guide. If the work later needs to span long waits, restarts, or approvals, see when to run an agent as a background job; execution durability is a separate decision from the number of agents.

    Sources

    Need Help with Your AI Project?

    At Zenovae, we build production-ready AI systems that scale. From OpenClaw setup to custom integrations, Mission Control workflows, and full-stack delivery, we can help you ship faster and avoid costly mistakes.

    Let's Talk