Free forever, no credit card.Get Started for Free →
← All posts
October 5, 2026 · 6 min read

Our AI Agent Forgets Everything Between Sessions. What Should We Put Underneath It?

Our AI Agent Forgets Everything Between Sessions. What Should We Put Underneath It? The short answer: an agent is stateless by default, so it forgets unless something outside it stores and returns context. What goes underneath is a memory layer: a store that saves what matters from each session and hands the right pieces back at the start of the next one. You have five real options: a plain database, a vector store, an embedded memory library like Mem0, Zep, Letta, or Cognee, a hosted memory se

Our AI Agent Forgets Everything Between Sessions. What Should We Put Underneath It?

The short answer: an agent is stateless by default, so it forgets unless something outside it stores and returns context. What goes underneath is a memory layer: a store that saves what matters from each session and hands the right pieces back at the start of the next one. You have five real options: a plain database, a vector store, an embedded memory library like Mem0, Zep, Letta, or Cognee, a hosted memory service over MCP, or structured files. The honest fork is who operates the memory: you, or a vendor.

Why does a shipped agent forget between sessions?

Because the model itself keeps nothing. Every session starts with a fresh context window: the system prompt, the user prompt, and whatever your code loads. The conversation from last Tuesday is not in there. Anything the agent learned — the customer's plan, the decisions it made, the constraints it was given — vanished the moment the previous session ended, unless something wrote it down first.

This is why teams end up re-explaining the same project to their own agent every Monday. The agent is not broken. It was never given a place to remember.

What does "putting something underneath it" actually mean?

It means giving the agent three things that live outside the model:

  1. A place to write. After a session, the important bits get stored: facts, decisions, preferences, what is done, what is next.
  2. A way to retrieve. At the start of the next session, the agent pulls back only what is relevant, not the entire history.
  3. A policy for both. What counts as worth storing, how long it lives, and who is allowed to read it.

Storage alone is not memory. A database full of raw transcripts does nothing if the agent cannot find the right record when it needs it. Retrieval is the part that does the work.

What are the five real options?

A plain database (Postgres, SQLite). You store records keyed by user or project. Dead simple, you own everything, costs almost nothing. The catch: you write the retrieval logic yourself, and exact-match queries over conversation text are bad at finding "what did they mean" records.

A vector database (pgvector, Pinecone, Qdrant). You embed memories and search by meaning. Recall quality jumps: the agent finds what it meant, not just what it typed. The catch: chunking, dedup, and pruning are your problem, and "store everything" turns into a swamp fast without a write policy.

An embedded memory library (Mem0, Zep, Letta, Cognee). These are built for exactly this job. Mem0 extracts and stores memories from conversations; Zep adds temporal reasoning over what happened when; Letta manages agentic memory states; Cognee builds knowledge graphs. Real strengths, real open code. The catch: you run it. The database, the embeddings, the uptime, the scaling bill — that is your infrastructure now.

A hosted memory service over MCP. Your agent connects with an API key as a Bearer header to an endpoint like https://api.vilix.ai/mcp and gets full memory with zero memory infrastructure to run. Vilix AI works this way: it stores the full conversation exchanges, derives useful memories from them, retrieves by semantic plus keyword search, isolates data per user, and resolves conflicts with last-write-wins, all reachable from every MCP-connected tool. The honest tradeoff: it is cloud-hosted, so if your policy is self-host-only, this is not your pick.

Structured files (markdown notes, JSON). A decisions log, a facts file. Cheap and transparent. The catch: it rots silently, it does not scale past one agent, and retrieval is "read the whole file into context" until the file gets too big to read.

Which one should a small team actually pick?

Run the decision through three questions:

  1. Who operates the memory? This is the real fork. An embedded library or a raw database means you own the servers, the backups, and the 3am pager. A hosted service means you pay someone else to care. Pick based on what your team actually wants to run, not on which README looked nicer.
  2. Who are the users? If one agent serves many customers or clients, the memory layer must isolate per user or per client. Per-user data isolation is a must-have, not a nice-to-have, and bolting it on later is a migration.
  3. What does retrieval need to find? Exact identifiers (order IDs, plan names, policy names) want keyword search. Meaning ("the customer sounded frustrated last time") wants semantic search. Most shipped agents need both, and few DIY stacks ship both.

How do you wire memory into a shipped agent?

The pattern that holds up, regardless of which store you pick:

  1. Decide what a memory is. Facts, decisions, preferences, state ("deploy to staging done, prod pending"). Not raw transcripts, not every turn.
  2. Write after the session, not during. A save step at session end keeps the hot loop fast and the store clean.
  3. Load at session start. Retrieve the relevant records for this user or project before the agent answers anything.
  4. Key everything by user. Per-customer isolation from day one.
  5. Set a write policy. What gets stored, what gets updated, what gets deleted, who decides. Stores rot from lax writes, not from the wrong database.
  6. Test recall like a feature. Ask the agent questions from last week's sessions. If it cannot answer, the memory layer is decoration.

FAQ

Can the agent just keep a longer context window instead of a memory layer?

No. A longer window still resets when the session ends. Context windows are per-session; memory must live outside the model to survive between sessions.

Do we need a vector database?

Only if recall by meaning matters to your agent. If users ask about exact identifiers, keyword search is enough and simpler. Most shipped agents eventually need both, because real conversations mix exact names with vague references.

What if one agent serves many clients?

The memory layer must isolate per client or per user from day one: separate namespaces, separate keys, no cross-contamination. This is a data-isolation requirement, and adding it after launch is a migration, not a feature flag.

How do conflicting memories get resolved?

The common rule is last-write-wins: the most recently saved version overrides the older one, so correcting something once corrects it everywhere. Vilix AI works this way — say "we are not doing that decision anymore" once, and that becomes the truth going forward.

What exactly does Vilix AI store?

The full user and assistant exchanges your agent sends via save_turn, useful memories derived from those exchanges, plus projects, tasks, and rules. You can list, update, and delete any of it from any connected tool or the dashboard, export it in a portable format anytime, or wipe the account instantly. It is cloud-hosted: you manage nothing, and if your policy requires self-hosting, pick an embedded library instead.

Get Started for Free

Persistent memory across ChatGPT, Claude, and the AI tools you already use in Vilix AI.

Get Started for Free

Free forever, no credit card.

Keep reading
Your Scheduled Agent Has No Past. Give It One: Seeding Agent Memory From Existing Conversations

Your Scheduled Agent Has No Past. Give It One: Seeding Agent Memory From Existing Conversations You have spent two years telling ChatGPT about your business. Your Claude chats hold the naming conventions, the deploy targets, the API versions, and the hundred little corrections you made along the way. Then you deploy a scheduled agent in n8n or a cron script, connect a memory layer, and watch it wake up knowing absolutely nothing. That empty start is not a bug. Memory systems only store what fl

Your Relevance AI Agent Has a Memory Feature. Your Scheduled Runs Still Start Blind.

You set a Relevance AI agent on a recurring schedule. Every morning at 7 it wakes up, pulls the new leads, scores them, and fires off the follow-ups. It works beautifully for a week. Then one morning it re-scores a lead it already contacted on Tuesday, sends a second follow-up to a prospect who said no, and completely misses the one who said "call me next week" because nobody told the agent that last week ended. The agent did not malfunction. It did not hallucinate. It just started blank, the s

One Memory Across Every AI Tool: Tabula, Eling, openIME, and Vilix AI, Honestly Compared

One Memory Across Every AI Tool: Tabula, Eling, openIME, and Vilix AI, Honestly Compared Quick answer: If you keep re-explaining yourself every time you switch AI tools, you need a memory layer outside any one of them. Four real options do this today: Tabula, Eling, openIME, and Vilix AI. Same goal, one memory for every AI, but they differ in what gets stored, who hosts it, and how much you manage yourself. Tabula is strongest for dashboard-level control. Eling is strongest for the simplest hos