Your Pydantic AI Agent Forgets Between Runs. Here Are the Memory Options That Work.
Your Pydantic AI Agent Forgets Between Runs. Here Are the Memory Options That Work. You ship an agent with Pydantic AI. It runs on a schedule: every morning it reads the support inbox, drafts replies, and flags the tricky ones. Day one looks great. Day four, it flags a ticket it already resolved on day two, drafts a reply that contradicts the answer it gave yesterday, and asks the user a question it already asked last week. Nothing crashed. The agent just woke up blind, the way it does every ru
Your Pydantic AI Agent Forgets Between Runs. Here Are the Memory Options That Work.
You ship an agent with Pydantic AI. It runs on a schedule: every morning it reads the support inbox, drafts replies, and flags the tricky ones. Day one looks great. Day four, it flags a ticket it already resolved on day two, drafts a reply that contradicts the answer it gave yesterday, and asks the user a question it already asked last week. Nothing crashed. The agent just woke up blind, the way it does every run.
This is the question every Pydantic AI builder asks eventually: does it remember anything between runs? The honest answer has two layers, and the gap between them is exactly where your memory architecture lives.
The short answer: it remembers the conversation, not the past
Pydantic AI gives you continuity inside a conversation, not memory across days. Each call to agent.run() is a separate run with a fresh run ID. If you pass the previous run's message history back in, the new run sees the old messages and carries the same conversation ID. Correlate pause, resume, and multi-turn work with that ID and your traces line up.
That is genuinely useful, and a much smaller promise than it sounds like. Message history is the memory of the run it is in. When the process ends, the schedule fires a fresh run tomorrow, or a different user talks to the same agent, that history is gone unless you stored it and fed it back yourself. Pydantic AI does not persist anything between runs on its own. Its own storage docs put it plainly: conversation history is an agent's memory of the run it is in, and longer-lived memory is a separate layer you attach.
What the ecosystem ships today
The Pydantic AI Harness SDK added a Memory capability: a persistent, namespaced notebook with on-demand search, backed by in-memory, file, or Postgres stores, plus conversation search over stored history. That is real progress. It means the framework now has an official answer to "where do memories live."
Notice what it does not answer. A file store on your laptop is not a store your scheduled agent on a server can reach. A Postgres store is memory you host, back up, and keep available. Multi-user isolation, retention rules, what gets remembered and what gets forgotten, and the same memory following the agent when it also runs somewhere else: all of that is still your code, your infrastructure, your on-call rotation.
Option 1: build the memory layer yourself
The DIY path is a Postgres table (or a vector store), an embedding step, a retrieval function, and a write path that saves the useful parts of each run. It works, it is fully under your control, and every part is honest engineering.
Count the real cost before you start. You are building storage, retrieval ranking, per-user namespacing so one customer's data never leaks into another's, a forgetting path (stale preferences have to die or they poison future runs), and a backup story. For one agent serving one user, the file store is fine and you should just do that. For anything with users, schedules, or revenue attached, the DIY tax compounds: every new agent you ship re-solves the same problems.
Option 2: use an embedded memory library
Embedded memory libraries give you an API for storing and searching memories inside your own process or service, and many handle the embedding plumbing for you. A solid choice when the agent is one component of a bigger system you already operate, because you were going to run infrastructure anyway.
The tradeoff is scope. An embedded library remembers what you put into it, in the process that holds it. When your scheduled agent wakes up on a fresh machine tomorrow, when the same user talks to a second agent built on a different stack, or when you move from prototype to production, the memory does not follow on its own. You still own the database, the uptime, and the migration path.
Option 3: give the agent hosted memory over MCP
Pydantic AI agents can act as MCP clients. Its docs cover connecting to remote MCP servers over Streamable HTTP and registering them as toolsets on the agent, with per-user authentication for servers that need credentials. That is the opening most builders miss: your agent does not need a memory database at all. It needs a memory it can call.
A hosted memory layer over MCP means the agent's memory lives in one account, reachable from any tool that speaks the protocol. The agent saves what happened in a run; the next run, tomorrow or on another machine, loads the relevant context before it starts. The same memory follows the user when they switch tools: context saved by your Pydantic AI agent is visible to the coding assistant, the voice agent, and the scheduled cron, because they all read the same store.
How to choose
The decision comes down to who operates the memory, not who writes the agent:
- One prototype, one user, no schedules. Use the framework's built-in conversation history and the Harness Memory capability with a file store. You do not need anything heavier.
- One agent, real users, and you already run infrastructure. An embedded memory library plus your own store keeps everything in your stack. Budget real time for the forgetting path and the isolation rules.
- Scheduled runs, multiple agents, or the same user across tools. Use hosted memory over MCP. Vilix AI is built for this: cloud-hosted with zero infrastructure for you to manage, one memory that follows the user across every tool and device over MCP, full conversation history stored (not just extracted facts), and retrieval that combines semantic search with keyword matching. It is free forever, with a 7-day Pro trial that needs no credit card, and your data stays portable: export everything or delete individual memories, or wipe the account instantly, whenever you want.
The builders who get burned are the ones in the middle: a scheduled agent with a file store on a laptop, or a DIY Postgres table with no retention policy. Memory that works on day one and rots by day thirty is worse than no memory at all, because the agent sounds confident while acting on stale facts.
The bottom line
Pydantic AI handles the agent. It does not handle the past. Message history covers the conversation; everything older, everything scheduled, everything shared across tools is a layer you choose deliberately. Pick the layer that matches who operates it: your laptop for prototypes, your infrastructure for single systems you already run, hosted memory over MCP for everything that has to survive between runs. Your agent forgets everything between runs by default. Give it one memory, and make sure that memory is still there tomorrow morning.