User Memories
Give your agents long-term memory across conversations.
Overview#
AG-UI Streams remember a conversation. User Memory remembers a person. This page explains what a memory is, how recall selects them, and what has to be true of your deployment before the memory surfaces exist at all.
If you are looking for the persistence architecture beneath a single conversation, read AG-UI Streams & Framework Threads instead.
Start with your coding agent#
Copy this prompt into your coding agent to inspect your existing CopilotKit app and configure long-term memory for your users. Prefer to work through the setup yourself? Follow the manual steps below.
Open your coding agent in your project's folder, or in an empty folder for a new app.This runs in a coding agent on your computer.
What is a memory?#
A memory is a short, durable statement about a user or a project, stored outside any single thread. "Prefers concise status updates" is a memory. The forty messages that revealed the preference are a thread.
The distinction matters because the two have different lifetimes. A thread is finished when the conversation is. A memory is meant to outlive it, and to be recalled into a conversation that has not happened yet.
Memories are stored as text plus a vector embedding, so recall is semantic rather than a keyword match. Asking for "how does this user like updates" can surface "prefers concise status updates" without sharing a word with it.
Set up User Memories#
When you are done, your app will be able to save a memory for a signed-in user and recall it in a later conversation.
Connect CopilotKit Intelligence#
Complete the Intelligence quickstart. That page signs you in with the CLI, selects the project, and configures server-side user identity. Your existing app and agent stay in place.
Confirm memory access#
Check your deployment's memory entitlement and embedding configuration using Activating memory below. Cloud-hosted deployments require no embedder configuration on your side. Self-hosted deployments need both a memory-enabled license and a configured embedder.
Save and recall a memory#
Choose the appropriate reading and writing interface for your app. Save a memory for the signed-in user, then recall it with a related query in a new conversation. Confirm the saved memory is returned.
Key concepts#
The three kinds#
Every memory is one of three kinds. The kind is supplied by whoever saves it and is never inferred.
| Kind | What it holds |
|---|---|
topical | A durable fact about the subject matter, independent of when it was learned. |
episodic | Something that happened, anchored to an occasion. |
operational | A preference or working instruction about how to behave. |
The kind is not cosmetic. Deduplication and supersession are same-kind
operations, so a topical fact and an operational preference with identical
text are two memories, not one.
User and project scope#
A memory is scoped to a single user or shared across a project.
useris the platform default. Omitscopeand you get it. This is the only scope delivered over realtime today.projectis shared, and must be requested explicitly.
Scope is enforced per request against the caller's grant, so a client holding read-only access to project scope cannot write to it.
Saving a near-duplicate absorbs it#
If you save content that closely matches a live memory in the same tenant, scope, and kind, the platform does not create a second row. It absorbs the new content into the existing memory, unions the source threads, and refreshes recency.
The save still succeeds, and the result tells you which happened, so a UI or an agent can say "absorbed into an existing memory" rather than implying something new was written.
Updating is a full replacement, not a patch#
Updating a memory supersedes it: the old memory is retired and a new one is
created with a new id. The change set you supply is the complete definition of
the new memory.
Omitting sourceThreadIds on an update does not preserve the previous value, it
resets the new memory's source threads to empty. Re-send content, kind, and
any sourceThreadIds you want to keep.
Removing is a retirement, not a delete#
Removing a memory retires it rather than erasing it. Retired memories are excluded from recall and from the default list, and can be surfaced again by asking for invalidated rows explicitly.
Activating memory#
Memory is not a feature flag. There is no MEMORY_ENABLED environment variable
and no memory.enabled Helm value. Access is granted by entitlement, and the
embedder is configured separately at startup.
Two independent things must be true before a caller can use memory:
- The deployment or organization is entitled to memory.
- app-api has valid embedder configuration.
Entitlement resolution fails closed. An unresolved entitlement, an inactive one, or a dependency outage all deny memory rather than allowing it.
Self-hosted#
The signed deployment license is the authority. Memory is available when that
license carries the memory feature, which today ships in the enterprise plan.
If your license does not include it, no amount of configuration will mount the
surfaces, and the correct next step is to talk to us about the license rather
than to keep editing values files.
The embedder is configured through these variables, which the chart templates for you:
| Variable | Purpose |
|---|---|
MEMORY_EMBEDDINGS_URL | Base URL of an OpenAI-compatible embeddings endpoint. Required. |
MEMORY_EMBEDDING_MODEL | Embedding model id sent to that endpoint. Required. |
MEMORY_EMBEDDINGS_API_KEY | Optional bearer token. Leave unset for the bundled in-cluster embedder. |
MEMORY_EMBEDDINGS_DIMENSIONS | Optional. Must be 1024, because storage is a 1024-dimension half-precision vector. |
Two consequences worth knowing before you deploy:
- app-api refuses to start if a required variable is missing or invalid. This is deliberate. Entitlements can change while the process is running, so it must never come up with memory routes and no embedder behind them.
- pgvector must be installed on the database server. The migration that creates the memory table also creates the extension, and the migration job fails without it.
By default the chart deploys an in-cluster text-embeddings-inference workload and points app-api at it. To use a hosted provider instead, set an external embeddings URL and model, at which point the in-cluster workload is not rendered:
embeddings:
external:
url: https://api.openai.com
model: text-embedding-3-small
dimensions: 1024
apiKey:
existingSecret: openai-embeddings
secretKey: api-keyChanging the embedding model or provider changes the vector space. Existing memories were embedded in the old space and are not comparable in the new one, so recall quality degrades until they are re-embedded. Treat a model change as a migration, not a config tweak.
Cloud-hosted#
On cloud-hosted Intelligence, memory is resolved per organization from that organization's effective entitlement, combining its plan features with any enterprise override. Nothing needs configuring on your side.
Limiting access per request#
Entitlement decides whether memory exists for a deployment. memory.access
decides what one request may do with it. It is a single policy on the
Intelligence runtime, and it governs both memory surfaces: the tools an
agent gets during a run, and the browser routes behind useMemories().
Omitting the memory option does not switch memory off — it removes the limit.
On its own it leaves both surfaces off, but either can still be turned on without
a policy, and then it runs with read-write on both scopes for every
authenticated request:
- With
enableEnterpriseLearningon, agent runs get all three memory tools, subject only to your organization's entitlement. - The deprecated
exposeMemoryRoutes: trueflag mounts the browser routes with no policy behind them.
This is the opposite of Channels, where a run gets no memory unless it asks.
Configure memory.access whenever either surface is on.
const runtime = new CopilotRuntime({
intelligence,
identifyUser: async (request) => resolveUser(request),
memory: {
access: ({ user, consumer }) => {
// Tenants on the free plan get recall, but record nothing new.
if (isFreePlan(user)) return { user: "read", project: "none" };
// The browser inspector is read-only; the agent may write.
return consumer === "client"
? { user: "read", project: "read" }
: { user: "read-write", project: "read-write" };
},
},
});The callback runs once per request, after identifyUser has resolved, and
receives:
| Field | Meaning |
|---|---|
request | The incoming Request, if the decision needs a header, cookie, or path. |
user | The application user identifyUser returned. |
consumer | "agent" for an agent run, "client" for the browser memory routes. |
Return a grant naming both scopes, or null. Each scope takes "none",
"read", or "read-write". The policy may be async.
What each outcome does#
null and a grant with every scope at "none" are one outcome, not two — both
mean "this request gets no memory". What that costs depends on what the caller
asked for:
| The policy returns | Agent run | Browser /memories/* |
|---|---|---|
A grant with any scope above "none" | Memory tools attach, filtered to the granted scopes | The request proceeds |
null, or every scope "none" | The run proceeds with no memory tools | 403 |
Nothing (undefined), or an access level outside the three | 500 | 500 |
| A thrown error | 500 | 500 |
The asymmetry is deliberate. A conversation asked for something other than memories, so switching memory off must not cost the user their assistant. The browser routes exist only to serve memories, so "you may not have them" answers the question actually asked — and an empty list would falsely imply none exist.
A missing return is treated as a broken policy rather than a restrictive one,
so a branch that forgets to return fails loudly instead of quietly switching
memory off.
A grant that allows one scope and not the other is honoured exactly: a
{ user: "read", project: "none" } grant registers recall_memory alone,
scoped to user memory, and leaves save_memory and forget_memory unregistered
so the agent never sees a tool it may not call.
memory.access is a memory policy, not an authorization gate on the request.
To reject a request outright, use beforeRequestMiddleware, which can refuse
before any of this runs.
It also mounts the browser routes#
Configuring memory turns on /memories, /memories/recall,
/memories/subscribe, and /memories/:id. Without it those routes 404 as if
they did not exist, so an un-opted-in deployment reveals nothing about memory
even when Intelligence is configured.
The older exposeMemoryRoutes: true flag did the mounting without the policy
and is deprecated. memory.access replaces it: one option that both exposes
the surfaces and bounds what they may do.
Channels grant memory per run instead#
A Channel run carries its own grant on runAgent, using the same three access
levels, rather than consulting memory.access. See
Identity and Memory for that path, including
how to choose the memory subject in a group conversation.
Reading and writing memories#
Angular#
injectMemories() exposes the same surface as signals. See
Threads, memory, attachments, and headless UI
for a worked component, and the
Angular public API
for the exported surface.
REST#
The memory routes are available to any backend or script holding a runtime credential. Authentication is the runtime tuple: a project API key as a bearer token, plus the app user's id in a header.
| Route | Purpose |
|---|---|
GET /api/memories | List the caller's memories. Pass includeInvalidated=true to include retired ones. |
POST /api/memories | Save a memory. |
POST /api/memories/recall | Semantic recall against a query. |
PATCH /api/memories/:id | Supersede a memory with a full replacement. |
DELETE /api/memories/:id | Retire a memory. |
POST /api/memories/subscribe | Mint a realtime subscription for memory metadata. |
A save takes the content, the kind, an optional scope, and optional provenance:
curl -X POST https://your-deployment/api/memories \
-H "Authorization: Bearer cpk-<project>_<short>_<long>" \
-H "X-Cpki-User-Id: <app-user-id>" \
-H "Content-Type: application/json" \
-d '{
"content": "Prefers concise status updates.",
"kind": "operational",
"sourceThreadIds": ["<thread-id>"]
}'content is capped at 8192 characters and sourceThreadIds at 100 entries.
Unknown fields are rejected rather than ignored, so a typo in a key is an error
instead of a silent no-op.
Recall takes a query and an optional result limit, which defaults to 5 and is capped at 20:
curl -X POST https://your-deployment/api/memories/recall \
-H "Authorization: Bearer cpk-<project>_<short>_<long>" \
-H "X-Cpki-User-Id: <app-user-id>" \
-H "Content-Type: application/json" \
-d '{ "query": "how does this user like updates", "limit": 5 }'Every route returns memories in the same shape: id, kind, scope,
content, and sourceThreadIds. Recall results additionally carry a score,
and the list route carries invalidatedAt when retired rows are requested.
Agent tools over MCP#
Three tools let an agent manage its own memory: save_memory, recall_memory,
and forget_memory.
They register on the MCP server only when that surface is mounted, which is a prerequisite for MCP generally rather than a memory setting. Registration is still checked per request against the caller organization's memory entitlement, so mounting MCP does not grant memory.
Registration is then narrowed again by the run's grant, so an agent is only ever
offered tools it may actually call: a readable scope registers recall_memory,
a writable one adds save_memory and forget_memory, and a grant of nothing
registers none of them. See
Limiting access per request for where that grant
comes from.
How recall works#
Recall is hybrid. A query is embedded and compared against stored vectors, and
the result is fused with other signals into a single score, which is what the
score field on a recall result reports. Retired and superseded memories are
excluded.
This is why the embedding space matters so much. Comparability is a property of the space, not of the text, so memories written under one model cannot be meaningfully ranked against a query embedded under another.
Next steps#
- Threads: AG-UI Streams & Framework Threads, the persistence model for a single conversation
- Platform: CopilotKit Intelligence, where memory sits among the other pillars
- Self-hosting: Self-hosting Intelligence, chart values, dependencies, and deployment modes