Skip to content
Huggehub Global Digital StudioParis --:--London --:--New York --:--Now accepting new projects →AI / Commerce / Technology / GrowthEurope • United Kingdom • United States

AI · 8 min read ·

AI Agent Memory: Types, Architecture and How to Build It

Models don't remember anything between calls. Memory is something you design — here's how to decide what an agent keeps, finds and forgets.

By Hatim El Badaoui

A black sphere linked to floating glass frames by thin chrome lines

A language model is stateless: every call starts from zero. When an assistant "remembers" your company name or last week's decision, that memory was stored somewhere by the application and put back into the prompt. AI agent memory is that design — what to store, how to find it again and when to let it go.

The types of AI agent memory

TypeWhat it holdsExample
Working (short-term)The current conversation and task stateThe order the customer is asking about right now
EpisodicPast events and interactions"Last month this customer had a delayed delivery"
SemanticFacts and preferences"Prefers WhatsApp; ships to Casablanca"
ProceduralHow to do things"Refunds above €100 need manager approval"

Working memory lives in the context window. The other three live outside the model and are retrieved when relevant.

A practical memory architecture

  1. Capture — after each interaction, extract candidate memories: stated facts, preferences, outcomes and corrections. Most of a conversation is not worth keeping.
  2. Store — save each memory with who it belongs to, when it was created, its source and a usefulness score.
  3. Retrieve — before the agent acts, fetch the top few memories relevant to the task, combining semantic similarity with recency and usefulness.
  4. Inject — add them to the prompt in a clearly labelled block, separate from instructions.
  5. Update and forget — reinforce memories that proved useful, let stale ones decay, and delete on request.

Scoring: what should the agent remember first?

Retrieval by similarity alone returns memories that sound relevant but are outdated. A better score blends three signals:

  • Relevance — how closely the memory matches the current task.
  • Recency — newer memories usually win, especially for preferences.
  • Usefulness — memories that helped before get a boost; ones that led to corrections lose weight.

Our agent-memory-store is a dependency-free reference for this loop: it remembers facts per user, scores them by usefulness and recency, retrieves the top-k and lets stale memories decay. It is deliberately small — the logic is what matters, and you can swap its storage for Postgres or a vector database later.

Why forgetting matters

Memory that only grows gets worse. Old preferences contradict new ones, retrieval gets noisier and personal data accumulates beyond what you are allowed to keep. Build forgetting in from the start: decay scores over time, merge duplicates, and honour deletion requests completely — including in backups and embeddings. Under GDPR and UK GDPR, a customer's memory record is personal data like any other.

Memory inside multi-step pipelines

Not all memory is long-term. In multi-stage workflows, each step needs the right slice of what came before — the extracted facts, not the whole transcript. Our llm-prompt-chain framework passes each step's output and a shared context into the next prompt template (for example extract → summarise → translate → classify). Keeping that hand-off explicit makes pipelines cheaper, faster and much easier to debug than one giant prompt.

Common mistakes

  • Storing whole transcripts instead of distilled facts.
  • Mixing retrieved memories into system instructions, where they can act as commands.
  • No owner or tenant on each memory, so data leaks between customers.
  • No way to inspect what the agent "knows" about someone.

Designing an agent that should get better with every conversation? Our AI agent development team builds memory, retrieval and privacy controls as one system.

Frequently asked questions

Do LLMs have memory?

No. Models are stateless between calls. Memory is stored by the application and added back into the prompt when it is relevant.

Is a vector database required for agent memory?

Not at the start. A relational table with good scoring works for many agents; vector search helps when you have many unstructured memories to match by meaning.

How do you stop memory from leaking between users?

Store an owner and tenant on every memory, filter retrieval by them at the database level, and keep retrieved memories separate from system instructions.

Let's build what's next

Let's build what's next.

Have a project, product or ambitious idea? Tell us where you want to go.