AI Agent Memory Boundaries: What to Persist Between Runs
Decide what a production agent should keep in session history, workflow state, long-term memory, and source systems.
Short answer: do not treat an agent's memory as one bucket. For a support agent, keep the current conversation in session history, the progress of an unfinished task in durable workflow state, and only a small set of useful, scoped facts in long-term memory. Ticket status, account entitlements, and approved policy should still come from the system that owns them.
The server should decide what gets saved and what context is loaded for each run. A model-generated summary can be useful context, but it is not the authoritative customer record.
Quick take: Use built-in session history when the agent only needs continuity within one conversation. Add cross-session memory only for specific facts that improve later work, and give each fact an owner, scope, source, review date, and removal path. Keep current business data and workflow progress in their existing systems.
| If you only read one section | Read this |
|---|---|
| You are deciding what to store | The four-boundary model |
| You use an agent framework already | When native memory is enough |
| You are preparing a release | The memory release checklist |
The four-boundary model
For a multi-session SaaS support workflow, ask who owns each piece of information and how long it should live. These are different data lifetimes, even when all four may appear in an agent prompt.
| Information | Put it in | Load it when | Boundary to preserve |
|---|---|---|---|
| Messages and corrections in the active case | Session or thread history | Continuing that case conversation | Scope the session to the authenticated actor and case |
| Progress through an unfinished support task | Application database or workflow runtime | Resuming that specific task | Store typed status and ownership, not a free-form transcript |
| Reusable team or user preference | Curated, scoped memory store | A later task can benefit from it | Keep source, scope, last-confirmed time, expiry, and a way to clear it |
| Ticket status, customer plan, policy, or tool permissions | The ticketing system, CRM, policy service, or application config that owns it | When making a decision that depends on its current value | Retrieve current data; do not promote an old summary into a source of truth |
The split is supported by current agent-framework designs. The OpenAI Agents SDK describes sessions as conversation history for a specific session. LangGraph separates thread-level short-term state from application-defined data stored across threads. Those are useful implementation choices, but the architecture decision is broader: conversation continuity, task progress, and authoritative business data have different owners and update rules.
Anthropic's context-engineering guidance also treats context as something to curate for each model call. It describes compaction and structured note-taking as ways to carry selected information across longer tasks, rather than sending every past message and tool result every time. Treat this as engineering guidance, not a guarantee that a summary will preserve every fact correctly.
A support-ticket example
Suppose an agent triages a SaaS support case, gathers product details, drafts a response, and may pause for an engineer.
- The application receives the case ID and authenticated workspace identity. It fetches the current ticket, plan details, and approved support policy from their owning systems.
- The current case conversation uses its own session or thread key. A correction such as “that log is from the staging account” belongs with this case.
- If the task may pause and continue later, the application stores structured progress separately: task ID, case ID, current stage, and who may resume it.
- A reusable preference such as “this team wants API examples in TypeScript” can become a candidate memory. The application applies its own rules for whether to save it, where it applies, and when it should be reviewed.
- Before a later case, the application retrieves only memories allowed for that workspace and relevant to the new task. It still checks current ticket and account facts live.
The practical rule is simple: memory can help the agent find or interpret context; it should not silently replace the system that owns the fact. If the remembered preference conflicts with a current instruction or record, prefer the current authorized source and surface the conflict for review.
Keep cross-session memory small and inspectable
Long-term memory is useful when a fact is likely to improve future work and is expensive or awkward to ask for repeatedly. Examples include a confirmed output preference or a stable team convention. It is usually a poor place for current ticket status, temporary incident details, account entitlements, or copied policy text that already has an owner elsewhere.
For each saved item, define:
- Scope: the user, team, workspace, project, or case it belongs to.
- Source: who said it or which system supplied it.
- Freshness: when it was observed and when it should be confirmed again.
- Purpose: which future task is allowed to use it.
- Lifecycle: who can correct, expire, or remove it.
Do not let the model choose a tenant or user namespace from untrusted conversation text. Resolve identity and scope in the application. Test that two workspaces with similar names cannot retrieve each other's notes, and that a changed or deleted preference stops influencing later runs.
These are design recommendations, not a universal retention schedule or a compliance determination. NIST's Generative AI Profile identifies data privacy as a risk area and suggests practices for governing data collection and retention; it is voluntary risk-management guidance, not a product-specific memory rule or a law.
When native memory is enough—and when a custom integration is justified
Start with the session or memory features already in the runtime or platform. Custom work is justified only when a documented need is missing from those features or from the application's existing data layer.
| Existing option | Native capability documented by its provider | Use it when | Add a custom layer only if |
|---|---|---|---|
| OpenAI Agents SDK sessions | Stores conversation history by session and supports several backends, including SQLite, Redis, SQLAlchemy, and OpenAI-hosted conversations | The requirement is continuity inside a defined session and its storage behavior fits your deployment | You need cross-session facts with application-specific ownership, provenance, or lifecycle rules |
| LangGraph persistence and memory | A checkpointer persists graph state for a thread; a Store holds application-defined data across threads | Thread state and cross-thread memory already match the workflow's state model | You need integration with a source of truth, identity policy, or data lifecycle the configured store does not provide |
| Amazon Bedrock AgentCore Memory | Configurable event retention, long-term memory strategies, namespaces, and IAM conditions for namespace access | Its supported retention and access controls fit the workflow and account boundaries | A required data owner, deletion path, or authorization rule cannot be represented by the native setup |
If the existing support platform already preserves the same conversation thread, that may be enough: no custom memory layer is needed while each case stays isolated and current customer data is fetched live. For an agent runtime, the built-in session feature is usually sufficient when the requirement is conversation continuity; do not build a second memory product. LangGraph or AgentCore may also be sufficient for cross-session memory when their documented scope, storage, and access controls meet the application's requirements. Verify the exact backend and authorization behavior in your deployed configuration before relying on it.
Build an app-owned adapter or memory service only for a specific gap, such as reconciling the memory lifecycle with an existing system of record or enforcing a tenant rule that the selected native configuration cannot express. A custom layer adds another store to secure, operate, migrate, and delete from; it is not automatically safer because the team controls the code.
The memory release checklist
Before enabling cross-session memory, answer these questions and add them to the release review:
- Can an operator explain who owns each stored field?
- Is the session key separate from the durable task ID and workspace identity?
- Does the application, rather than the model, choose the memory scope?
- Can a user or operator inspect, correct, expire, and remove a saved preference?
- Does a stale-memory test prove that live ticket and policy data take precedence?
- Do isolation tests cover two workspaces, concurrent cases, and similar user identifiers?
- Does the native provider's retention, export, and deletion behavior fit the workflow's written requirements?
If those answers are unclear, keep cross-session memory off. You can still ship a useful agent with live retrieval and scoped session history.
Where Zenovae fits
Zenovae's production agent build is scoped around one bounded workflow, with tools, approvals, an evaluation gate, and a runbook. If the hard part is connecting the agent to the systems that own current customer and task data, review AI integration services. The production-ready AI agent guide gives a broader release checklist, and an agent reliability audit is relevant when a live system needs an independent score before expansion.
The useful first decision is not “Which memory database should we buy?” It is “Which facts need to persist, who owns them, and what should happen when they become stale?” Once that is written down, you can see whether a built-in session store is enough or whether a real cross-system gap remains.
Sources
- OpenAI Agents SDK: Sessions
- LangGraph: Memory
- LangGraph: Persistence
- Anthropic: Effective context engineering for AI agents
- Amazon Bedrock AgentCore: Organize long-term memory with namespaces
- Amazon Bedrock AgentCore: Create a memory store and set event retention
- NIST: Artificial Intelligence Risk Management Framework—Generative Artificial Intelligence Profile
Need Help with Your AI Project?
At Zenovae, we build production-ready AI systems that scale. From OpenClaw setup to custom integrations, Mission Control workflows, and full-stack delivery, we can help you ship faster and avoid costly mistakes.
Let's Talk