Back to Blog
    AI Agent Architecture
    September 29, 20269 min

    AI Agent Memory Boundaries: What to Persist Between Runs

    Decide what a production agent should keep in session history, workflow state, long-term memory, and source systems.

    AI Agent MemoryProduction AIAgent StateWorkflow ArchitectureSaaS Support

    Short answer: do not treat an agent's memory as one bucket. For a support agent, keep the current conversation in session history, the progress of an unfinished task in durable workflow state, and only a small set of useful, scoped facts in long-term memory. Ticket status, account entitlements, and approved policy should still come from the system that owns them.

    The server should decide what gets saved and what context is loaded for each run. A model-generated summary can be useful context, but it is not the authoritative customer record.

    Quick take: Use built-in session history when the agent only needs continuity within one conversation. Add cross-session memory only for specific facts that improve later work, and give each fact an owner, scope, source, review date, and removal path. Keep current business data and workflow progress in their existing systems.

    If you only read one sectionRead this
    You are deciding what to storeThe four-boundary model
    You use an agent framework alreadyWhen native memory is enough
    You are preparing a releaseThe memory release checklist

    The four-boundary model

    For a multi-session SaaS support workflow, ask who owns each piece of information and how long it should live. These are different data lifetimes, even when all four may appear in an agent prompt.

    InformationPut it inLoad it whenBoundary to preserve
    Messages and corrections in the active caseSession or thread historyContinuing that case conversationScope the session to the authenticated actor and case
    Progress through an unfinished support taskApplication database or workflow runtimeResuming that specific taskStore typed status and ownership, not a free-form transcript
    Reusable team or user preferenceCurated, scoped memory storeA later task can benefit from itKeep source, scope, last-confirmed time, expiry, and a way to clear it
    Ticket status, customer plan, policy, or tool permissionsThe ticketing system, CRM, policy service, or application config that owns itWhen making a decision that depends on its current valueRetrieve current data; do not promote an old summary into a source of truth

    The split is supported by current agent-framework designs. The OpenAI Agents SDK describes sessions as conversation history for a specific session. LangGraph separates thread-level short-term state from application-defined data stored across threads. Those are useful implementation choices, but the architecture decision is broader: conversation continuity, task progress, and authoritative business data have different owners and update rules.

    Anthropic's context-engineering guidance also treats context as something to curate for each model call. It describes compaction and structured note-taking as ways to carry selected information across longer tasks, rather than sending every past message and tool result every time. Treat this as engineering guidance, not a guarantee that a summary will preserve every fact correctly.

    A support-ticket example

    Suppose an agent triages a SaaS support case, gathers product details, drafts a response, and may pause for an engineer.

    1. The application receives the case ID and authenticated workspace identity. It fetches the current ticket, plan details, and approved support policy from their owning systems.
    2. The current case conversation uses its own session or thread key. A correction such as “that log is from the staging account” belongs with this case.
    3. If the task may pause and continue later, the application stores structured progress separately: task ID, case ID, current stage, and who may resume it.
    4. A reusable preference such as “this team wants API examples in TypeScript” can become a candidate memory. The application applies its own rules for whether to save it, where it applies, and when it should be reviewed.
    5. Before a later case, the application retrieves only memories allowed for that workspace and relevant to the new task. It still checks current ticket and account facts live.

    The practical rule is simple: memory can help the agent find or interpret context; it should not silently replace the system that owns the fact. If the remembered preference conflicts with a current instruction or record, prefer the current authorized source and surface the conflict for review.

    Keep cross-session memory small and inspectable

    Long-term memory is useful when a fact is likely to improve future work and is expensive or awkward to ask for repeatedly. Examples include a confirmed output preference or a stable team convention. It is usually a poor place for current ticket status, temporary incident details, account entitlements, or copied policy text that already has an owner elsewhere.

    For each saved item, define:

    • Scope: the user, team, workspace, project, or case it belongs to.
    • Source: who said it or which system supplied it.
    • Freshness: when it was observed and when it should be confirmed again.
    • Purpose: which future task is allowed to use it.
    • Lifecycle: who can correct, expire, or remove it.

    Do not let the model choose a tenant or user namespace from untrusted conversation text. Resolve identity and scope in the application. Test that two workspaces with similar names cannot retrieve each other's notes, and that a changed or deleted preference stops influencing later runs.

    These are design recommendations, not a universal retention schedule or a compliance determination. NIST's Generative AI Profile identifies data privacy as a risk area and suggests practices for governing data collection and retention; it is voluntary risk-management guidance, not a product-specific memory rule or a law.

    When native memory is enough—and when a custom integration is justified

    Start with the session or memory features already in the runtime or platform. Custom work is justified only when a documented need is missing from those features or from the application's existing data layer.

    Existing optionNative capability documented by its providerUse it whenAdd a custom layer only if
    OpenAI Agents SDK sessionsStores conversation history by session and supports several backends, including SQLite, Redis, SQLAlchemy, and OpenAI-hosted conversationsThe requirement is continuity inside a defined session and its storage behavior fits your deploymentYou need cross-session facts with application-specific ownership, provenance, or lifecycle rules
    LangGraph persistence and memoryA checkpointer persists graph state for a thread; a Store holds application-defined data across threadsThread state and cross-thread memory already match the workflow's state modelYou need integration with a source of truth, identity policy, or data lifecycle the configured store does not provide
    Amazon Bedrock AgentCore MemoryConfigurable event retention, long-term memory strategies, namespaces, and IAM conditions for namespace accessIts supported retention and access controls fit the workflow and account boundariesA required data owner, deletion path, or authorization rule cannot be represented by the native setup

    If the existing support platform already preserves the same conversation thread, that may be enough: no custom memory layer is needed while each case stays isolated and current customer data is fetched live. For an agent runtime, the built-in session feature is usually sufficient when the requirement is conversation continuity; do not build a second memory product. LangGraph or AgentCore may also be sufficient for cross-session memory when their documented scope, storage, and access controls meet the application's requirements. Verify the exact backend and authorization behavior in your deployed configuration before relying on it.

    Build an app-owned adapter or memory service only for a specific gap, such as reconciling the memory lifecycle with an existing system of record or enforcing a tenant rule that the selected native configuration cannot express. A custom layer adds another store to secure, operate, migrate, and delete from; it is not automatically safer because the team controls the code.

    The memory release checklist

    Before enabling cross-session memory, answer these questions and add them to the release review:

    • Can an operator explain who owns each stored field?
    • Is the session key separate from the durable task ID and workspace identity?
    • Does the application, rather than the model, choose the memory scope?
    • Can a user or operator inspect, correct, expire, and remove a saved preference?
    • Does a stale-memory test prove that live ticket and policy data take precedence?
    • Do isolation tests cover two workspaces, concurrent cases, and similar user identifiers?
    • Does the native provider's retention, export, and deletion behavior fit the workflow's written requirements?

    If those answers are unclear, keep cross-session memory off. You can still ship a useful agent with live retrieval and scoped session history.

    Where Zenovae fits

    Zenovae's production agent build is scoped around one bounded workflow, with tools, approvals, an evaluation gate, and a runbook. If the hard part is connecting the agent to the systems that own current customer and task data, review AI integration services. The production-ready AI agent guide gives a broader release checklist, and an agent reliability audit is relevant when a live system needs an independent score before expansion.

    The useful first decision is not “Which memory database should we buy?” It is “Which facts need to persist, who owns them, and what should happen when they become stale?” Once that is written down, you can see whether a built-in session store is enough or whether a real cross-system gap remains.

    Sources

    Need Help with Your AI Project?

    At Zenovae, we build production-ready AI systems that scale. From OpenClaw setup to custom integrations, Mission Control workflows, and full-stack delivery, we can help you ship faster and avoid costly mistakes.

    Let's Talk