Full Pro free for 7 days, no credit card. Start free →
← All posts
September 29, 2026 · 6 min read

Which MCP Memory Server Should You Give Claude Code? 7 Options, Honestly Compared

Which MCP Memory Server Should You Give Claude Code? 7 Options, Honestly Compared Every Claude Code session starts blank. You brief the agent, it does the work, the session ends, and next time you start over. CLAUDE.md slows the bleeding but only holds what you remembered to write down, in one repo, on one machine. A memory server fixes this at the protocol level: register it once, its tools appear in every session, and the agent saves decisions, project state, and facts, then pulls them back

Which MCP Memory Server Should You Give Claude Code? 7 Options, Honestly Compared

Every Claude Code session starts blank. You brief the agent, it does the work, the session ends, and next time you start over. CLAUDE.md slows the bleeding but only holds what you remembered to write down, in one repo, on one machine.

A memory server fixes this at the protocol level: register it once, its tools appear in every session, and the agent saves decisions, project state, and facts, then pulls them back semantically. The right choice depends on where your agents run and who operates the infrastructure. Here are the seven options worth evaluating in September 2026.

Start here: the reference server

@modelcontextprotocol/server-memory is the official knowledge-graph memory server: the agent stores entities and relations, and recalls them in later sessions. It is the smallest possible answer and the baseline everything else is measured against.

Pick it when you want to try the pattern with minimum commitment. Move on when you need multi-tool sharing, a dashboard, tunable retrieval, or memory that leaves one machine.

If you self-host: four local options

You run the infrastructure, your data stays on your hardware, and you pay in setup and maintenance time.

memdb ships the most configurable MCP surface of the bunch: 12 tools across four groups, with memory search that tunes injection profile, result count, relativity threshold, and dedup strategy. The register command comes straight from its docs:

claude mcp add memdb http://127.0.0.1:8001/mcp

It also offers a Claude Code plugin for automatic context injection. Pick it when retrieval control matters more than setup simplicity; the price is a local service you operate and knobs you tune.

persistent-memory-stack runs as a single Docker Compose service over streamable HTTP, shared by Claude Code, Codex CLI, and Claude Desktop. Its best design idea: only the recall tool loads at task start, with write, graph, and admin tools deferred until needed, so sessions skip the context cost of a full tool catalog. Choose it for one local memory service across tools with per-session context economy. The cost: a Compose stack with tokens and backups on your plate.

marm-memory is a local-first layer that fuses three memory types: session history, codebase indices, and concept graphs, all in SQLite. It talks to Claude Code, Codex, Grok, Gemini, VS Code, and Cursor, so it covers multi-client setups without any cloud. Two steps to wire it up:

claude mcp add --transport http marm-memory http://localhost:8001/mcp

Choose it for code-aware memory (codebase index plus concept graph, not just notes) when data must never leave your machines.

waggle-mcp is the lightest lift in the self-hosted group: a local graph-memory server installed with pipx, no API key, no Docker, no cloud account:

pipx install waggle-mcp
claude mcp add --transport stdio waggle -- waggle-mcp serve --transport stdio

SQLite-backed with automatic memory hooks in Claude Code, so reads and writes happen without narration. Choose it when you want graph-shaped memory with minimum setup. The multi-machine story is yours to build.

If you run zero infrastructure: two options

shared-agent-memory-mcp keeps long-term memory in Notion and exposes search, recall, get, add, update, and delete over MCP, with a CLAUDE.md convention telling the agent to read before working and write only durable knowledge. It also works from Cline, OpenCode, GitHub Copilot, and Hermes. Choose it when your team lives in Notion and you want hand-editable memory. The constraint: Notion's rate limits and structure become your memory's limits.

Vilix AI is the one fully cloud-hosted option: a memory layer you connect to over MCP instead of installing anything. You register Claude Code once, then connect Codex, Cursor, OpenClaw, Hermes, or any other MCP-compatible tool to the same account. The agent saves full conversations (not just extracted facts), decisions, and project state; any connected tool later retrieves what is relevant semantically, with keyword search alongside it for exact strings.

No local server means nothing to babysit and nothing tied to one machine. Memory follows you across devices and works for scheduled or background agents that never touch your hardware. You manage it from the dashboard or any connected AI: list, update, delete individual memories or wipe the whole account instantly, and export everything in a portable format whenever you want. When two tools disagree, last write wins, stated up front, so one correction becomes the truth everywhere.

The free plan is enough to start, and the 7-day Pro trial needs no credit card. Learn more at https://vilix.ai/?utm_source=vilix-blog&utm_medium=article&utm_campaign=which-mcp-memory-server-should-you-give-claude-code.

Choose it when your agents run in more than one place, or you would rather work than operate infrastructure. The honest tradeoff: cloud-only, no self-host option. If data cannot leave your hardware, the local group above is your answer.

The seven at a glance

Option Runs on Memory model Multi-tool Best when
Reference server Your machine Knowledge graph Claude only You want the baseline
memdb Your machine Scoped cubes, tunable search Any MCP client Retrieval control matters most
persistent-memory-stack Your machine (Docker) Task-start recall, deferred tools Yes You want context economy across tools
marm-memory Your machine History + codebase index + concept graph Yes, many clients Code-aware memory, zero cloud
waggle-mcp Your machine (pipx) Graph memory Any MCP client Minimum setup
shared-agent-memory-mcp Notion Searchable notes Yes Your team lives in Notion
Vilix AI Cloud Full conversations + facts + state Yes Agents run in more than one place

FAQ

Can one memory server really serve Claude Code and my other tools? Yes, with the multi-tool options above. The reference server is Claude-centric. Register the same server in each client and they share the memory.

Do agents save to memory automatically, or do I have to tell them? Some ship hooks or plugins that read and write automatically (task-start recall, memory hooks, auto-injection plugins); others rely on a CLAUDE.md convention. Check each README before assuming.

What does running memory in Claude Code actually cost? Local options cost your time: setup, updates, backups, disk, and RAM. The hosted route costs a subscription and zero operating time. Weigh the engineering hours, not just the sticker price.

How do I shrink CLAUDE.md once I have a memory server? Move facts, decisions, and corrections into the memory server; keep CLAUDE.md for standing rules the agent must follow every session.

Is it safe to keep codebase context in a cloud memory? A policy question, not a technical one. If data cannot leave your hardware, pick a local option. If it can, look for export-anytime portability and instant delete so you are never locked in.

The decision, in one paragraph

Smallest experiment: the reference server. Data stays on your hardware: waggle-mcp for minimum fuss, memdb for retrieval control, marm-memory for code-aware memory, persistent-memory-stack for sharing across tools. Agents across machines or schedules: the cloud option. Verify features before committing; this field ships weekly.

Try Vilix Pro free for 7 days

Persistent memory across ChatGPT, Claude, and the AI tools you already use in Vilix AI.

Get started free
Keep reading
Vilix AI vs Letta: Which Memory Approach Fits Your AI Agents?

Vilix AI vs Letta: Which Memory Approach Fits Your AI Agents? The short answer: Vilix AI and Letta both give AI agents persistent memory, but they live in different places. Letta is an open-source agent harness where the agent curates its own memory as versioned files you run locally with your own model keys. Vilix AI is a cloud-hosted shared memory layer you attach over MCP to tools you didn't build (Claude, Codex, n8n agents, headless runners), one memory across all of them. The honest tradeo

Human-in-the-Loop Agents Forget What People Tell Them. Memory Fixes That

Human-in-the-Loop Agents Forget What People Tell Them. Memory Fixes That Human-in-the-loop is the responsible way to run a scheduled AI agent. Any step that spends money, messages a customer, or deletes something gets gated behind a person. The agent proposes; the human approves. Nobody argues with that architecture. But there is a gap hiding inside it. The approval happens, the run finishes, and by the next run the agent has no record that a human ever weighed in. Approvals are single-use. Th

Your Scheduled Agent Has Two Memories. Only One Survives the Night.

Target query: short term vs long term memory ai agents Slug: scheduled-agent-two-memories-only-one-survives dev.to title: Short-Term vs Long-Term Memory in AI Agents: What Actually Survives Between Runs vilix.ai-blog title: Your Scheduled Agent Has Two Memories. Only One Survives the Night. Your Scheduled Agent Has Two Memories. Only One Survives the Night. Your n8n workflow fires at 6 AM. The AI agent inside it wakes up, triages the overnight support tickets, drafts the morning digest, and g