AI · 8 min read ·
AI Agent Memory: Types, Architecture and How to Build It
Models don't remember anything between calls. Memory is something you design — here's how to decide what an agent keeps, finds and forgets.

A language model is stateless: every call starts from zero. When an assistant "remembers" your company name or last week's decision, that memory was stored somewhere by the application and put back into the prompt. AI agent memory is that design — what to store, how to find it again and when to let it go.
The types of AI agent memory
| Type | What it holds | Example |
|---|---|---|
| Working (short-term) | The current conversation and task state | The order the customer is asking about right now |
| Episodic | Past events and interactions | "Last month this customer had a delayed delivery" |
| Semantic | Facts and preferences | "Prefers WhatsApp; ships to Casablanca" |
| Procedural | How to do things | "Refunds above €100 need manager approval" |
Working memory lives in the context window. The other three live outside the model and are retrieved when relevant.
A practical memory architecture
- Capture — after each interaction, extract candidate memories: stated facts, preferences, outcomes and corrections. Most of a conversation is not worth keeping.
- Store — save each memory with who it belongs to, when it was created, its source and a usefulness score.
- Retrieve — before the agent acts, fetch the top few memories relevant to the task, combining semantic similarity with recency and usefulness.
- Inject — add them to the prompt in a clearly labelled block, separate from instructions.
- Update and forget — reinforce memories that proved useful, let stale ones decay, and delete on request.
Scoring: what should the agent remember first?
Retrieval by similarity alone returns memories that sound relevant but are outdated. A better score blends three signals:
- Relevance — how closely the memory matches the current task.
- Recency — newer memories usually win, especially for preferences.
- Usefulness — memories that helped before get a boost; ones that led to corrections lose weight.
Our agent-memory-store is a dependency-free reference for this loop: it remembers facts per user, scores them by usefulness and recency, retrieves the top-k and lets stale memories decay. It is deliberately small — the logic is what matters, and you can swap its storage for Postgres or a vector database later.
Why forgetting matters
Memory that only grows gets worse. Old preferences contradict new ones, retrieval gets noisier and personal data accumulates beyond what you are allowed to keep. Build forgetting in from the start: decay scores over time, merge duplicates, and honour deletion requests completely — including in backups and embeddings. Under GDPR and UK GDPR, a customer's memory record is personal data like any other.
Memory inside multi-step pipelines
Not all memory is long-term. In multi-stage workflows, each step needs the right slice of what came before — the extracted facts, not the whole transcript. Our llm-prompt-chain framework passes each step's output and a shared context into the next prompt template (for example extract → summarise → translate → classify). Keeping that hand-off explicit makes pipelines cheaper, faster and much easier to debug than one giant prompt.
Common mistakes
- Storing whole transcripts instead of distilled facts.
- Mixing retrieved memories into system instructions, where they can act as commands.
- No owner or tenant on each memory, so data leaks between customers.
- No way to inspect what the agent "knows" about someone.
Designing an agent that should get better with every conversation? Our AI agent development team builds memory, retrieval and privacy controls as one system.
Frequently asked questions
Do LLMs have memory?
No. Models are stateless between calls. Memory is stored by the application and added back into the prompt when it is relevant.
Is a vector database required for agent memory?
Not at the start. A relational table with good scoring works for many agents; vector search helps when you have many unstructured memories to match by meaning.
How do you stop memory from leaking between users?
Store an owner and tenant on every memory, filter retrieval by them at the database level, and keep retrieved memories separate from system instructions.


