Our AI Agent Forgets Everything Between Sessions. What Should We Put Underneath It?
Our AI Agent Forgets Everything Between Sessions. What Should We Put Underneath It? The short answer: an agent is stateless by default, so it forgets unless something outside it stores and returns context. What goes underneath is a memory layer: a store that saves what matters from each session and hands the right pieces back at the start of the next one. You have five real options: a plain database, a vector store, an embedded memory library like Mem0, Zep, Letta, or Cognee, a hosted memory se
Our AI Agent Forgets Everything Between Sessions. What Should We Put Underneath It?
The short answer: an agent is stateless by default, so it forgets unless something outside it stores and returns context. What goes underneath is a memory layer: a store that saves what matters from each session and hands the right pieces back at the start of the next one. You have five real options: a plain database, a vector store, an embedded memory library like Mem0, Zep, Letta, or Cognee, a hosted memory service over MCP, or structured files. The honest fork is who operates the memory: you, or a vendor.
Why does a shipped agent forget between sessions?
Because the model itself keeps nothing. Every session starts with a fresh context window: the system prompt, the user prompt, and whatever your code loads. The conversation from last Tuesday is not in there. Anything the agent learned — the customer's plan, the decisions it made, the constraints it was given — vanished the moment the previous session ended, unless something wrote it down first.
This is why teams end up re-explaining the same project to their own agent every Monday. The agent is not broken. It was never given a place to remember.
What does "putting something underneath it" actually mean?
It means giving the agent three things that live outside the model:
- A place to write. After a session, the important bits get stored: facts, decisions, preferences, what is done, what is next.
- A way to retrieve. At the start of the next session, the agent pulls back only what is relevant, not the entire history.
- A policy for both. What counts as worth storing, how long it lives, and who is allowed to read it.
Storage alone is not memory. A database full of raw transcripts does nothing if the agent cannot find the right record when it needs it. Retrieval is the part that does the work.
What are the five real options?
A plain database (Postgres, SQLite). You store records keyed by user or project. Dead simple, you own everything, costs almost nothing. The catch: you write the retrieval logic yourself, and exact-match queries over conversation text are bad at finding "what did they mean" records.
A vector database (pgvector, Pinecone, Qdrant). You embed memories and search by meaning. Recall quality jumps: the agent finds what it meant, not just what it typed. The catch: chunking, dedup, and pruning are your problem, and "store everything" turns into a swamp fast without a write policy.
An embedded memory library (Mem0, Zep, Letta, Cognee). These are built for exactly this job. Mem0 extracts and stores memories from conversations; Zep adds temporal reasoning over what happened when; Letta manages agentic memory states; Cognee builds knowledge graphs. Real strengths, real open code. The catch: you run it. The database, the embeddings, the uptime, the scaling bill — that is your infrastructure now.
A hosted memory service over MCP. Your agent connects with an API key as a Bearer header to an endpoint like https://api.vilix.ai/mcp and gets full memory with zero memory infrastructure to run. Vilix AI works this way: it stores the full conversation exchanges, derives useful memories from them, retrieves by semantic plus keyword search, isolates data per user, and resolves conflicts with last-write-wins, all reachable from every MCP-connected tool. The honest tradeoff: it is cloud-hosted, so if your policy is self-host-only, this is not your pick.
Structured files (markdown notes, JSON). A decisions log, a facts file. Cheap and transparent. The catch: it rots silently, it does not scale past one agent, and retrieval is "read the whole file into context" until the file gets too big to read.
Which one should a small team actually pick?
Run the decision through three questions:
- Who operates the memory? This is the real fork. An embedded library or a raw database means you own the servers, the backups, and the 3am pager. A hosted service means you pay someone else to care. Pick based on what your team actually wants to run, not on which README looked nicer.
- Who are the users? If one agent serves many customers or clients, the memory layer must isolate per user or per client. Per-user data isolation is a must-have, not a nice-to-have, and bolting it on later is a migration.
- What does retrieval need to find? Exact identifiers (order IDs, plan names, policy names) want keyword search. Meaning ("the customer sounded frustrated last time") wants semantic search. Most shipped agents need both, and few DIY stacks ship both.
How do you wire memory into a shipped agent?
The pattern that holds up, regardless of which store you pick:
- Decide what a memory is. Facts, decisions, preferences, state ("deploy to staging done, prod pending"). Not raw transcripts, not every turn.
- Write after the session, not during. A save step at session end keeps the hot loop fast and the store clean.
- Load at session start. Retrieve the relevant records for this user or project before the agent answers anything.
- Key everything by user. Per-customer isolation from day one.
- Set a write policy. What gets stored, what gets updated, what gets deleted, who decides. Stores rot from lax writes, not from the wrong database.
- Test recall like a feature. Ask the agent questions from last week's sessions. If it cannot answer, the memory layer is decoration.
FAQ
Can the agent just keep a longer context window instead of a memory layer?
No. A longer window still resets when the session ends. Context windows are per-session; memory must live outside the model to survive between sessions.
Do we need a vector database?
Only if recall by meaning matters to your agent. If users ask about exact identifiers, keyword search is enough and simpler. Most shipped agents eventually need both, because real conversations mix exact names with vague references.
What if one agent serves many clients?
The memory layer must isolate per client or per user from day one: separate namespaces, separate keys, no cross-contamination. This is a data-isolation requirement, and adding it after launch is a migration, not a feature flag.
How do conflicting memories get resolved?
The common rule is last-write-wins: the most recently saved version overrides the older one, so correcting something once corrects it everywhere. Vilix AI works this way — say "we are not doing that decision anymore" once, and that becomes the truth going forward.
What exactly does Vilix AI store?
The full user and assistant exchanges your agent sends via save_turn, useful memories derived from those exchanges, plus projects, tasks, and rules. You can list, update, and delete any of it from any connected tool or the dashboard, export it in a portable format anytime, or wipe the account instantly. It is cloud-hosted: you manage nothing, and if your policy requires self-hosting, pick an embedded library instead.