Make's AI Agents Have a Memory Problem Nobody Talks About
Make's AI Agents Have a Memory Problem Nobody Talks About Nobody talks about it because the demos never show week six. In the demo, the Make AI agent reads a fresh inbox, applies the instructions, and produces a tidy result. It looks complete. Six weeks later the cracks show: the same disqualified leads get researched again, the same false-positive alerts get escalated again, the content angles that flopped last month get repurposed again. The agent is not broken. It is doing exactly what it wa
Make's AI Agents Have a Memory Problem Nobody Talks About
Nobody talks about it because the demos never show week six. In the demo, the Make AI agent reads a fresh inbox, applies the instructions, and produces a tidy result. It looks complete. Six weeks later the cracks show: the same disqualified leads get researched again, the same false-positive alerts get escalated again, the content angles that flopped last month get repurposed again. The agent is not broken. It is doing exactly what it was set up to do. The setup is what is missing a piece.
The piece is memory, and the reason the gap stays invisible is that Make's interface seems to promise it already. The Knowledge tab is labeled as the agent's long-term memory. It is not. Understanding the difference is the whole game.
A library is not a diary
Make documents its AI agent modules clearly: choose a provider — OpenAI, Anthropic, Gemini, or Make's built-in AI Provider — define the instructions, and optionally attach Knowledge: files the agent can draw on, stored in a RAG vector database, with the relevant chunks pulled in per request. The suggested contents are FAQs, brand guidelines, company policies.
That is a library. Libraries are wonderful and static. Nothing the agent does during a run is ever written back into those files. The library does not know that Acme Corp was disqualified as a lead in February, or that the "urgent" alerts from the staging environment are always false positives, or that last quarter's content angles underperformed and should be retired. A diary records what happened. A scheduled agent needs the diary, and the Knowledge tab will never become one no matter how many files get uploaded.
The three things a scheduled agent must remember
Once the library stops being asked to do the diary's job, the requirements get clear. A scheduled agent needs three kinds of memory, and each one fails differently when it is missing.
First, durable facts about the world it operates in: customer tiers, SLAs, the vendors on the do-not-contact list, the products that were discontinued. Without these, the agent re-derives stable truths from scratch every run and occasionally gets them wrong.
Second, run history: what the last few runs actually did. The lead-scoring scenario should know it already researched and rejected a company last week, with the reason attached. The monitoring scenario should know the last three "critical" alerts from the same host were noise. Without run history, every execution is groundhog day with a compute bill.
Third, corrections: the fixes a human applied after the fact. When a priority got overridden on a ticket, or a draft got edited before sending, that correction is the most valuable training signal the system will ever receive. Without a place to store it, the agent needs the same correction forever.
What actually fills the gap
Four options cover the realistic space, and the honest version of each matters more than the polished one.
The built-in Knowledge, used as intended. For static reference material — the brand voice, the refund policy, the FAQ — it is the right tool and the cheapest one. Its limits are architectural: nothing writes back to it during runs, it is scoped to one agent, and it answers reference questions while the real problem is run-to-run continuity.
A data store wired as a scratchpad. A Make Data Store, Airtable, or Google Sheet: look up the relevant rows at the start of the scenario, map them into the prompt, write the outcome back at the end. Pragmatic for a single scenario. The costs are the ones DIY always charges: token usage grows with history, lookups are exact-key rather than semantic, and the schema, the cleanup, and the deduplication are yours to own. A second scenario that needs the same memory means rebuilding the plumbing.
A vector database. Embed memories, retrieve by semantic similarity, inject only what is relevant. This is the serious answer when memory is a product feature with engineering behind it — real infrastructure, an embedding pipeline, chunking and retrieval tuning, and per-query costs to show for it.
A hosted memory layer. Rent instead of build. The scenario's memory logic becomes API calls or an MCP connection instead of a database you operate. This is the slot Vilix AI is designed for: cloud-hosted with zero infrastructure to run, the same memory reachable over MCP from Make, n8n, Zapier, Claude Code, and everything else, so a fact written in one tool is readable in another. It stores full conversation history rather than just extracted facts, which means the agent can consult what actually happened instead of a summary of it. The free plan never expires, the Pro trial is seven days with no card, and export or deletion is available any time. The tradeoff is the category's own: the memory lives on managed servers, not inside your stack, so air-gapped or on-premise requirements rule it out.
The shape of the fix in a real scenario
Take the lead-qualification scenario. Today it pulls new signups, researches each company, and scores them — re-researching companies it rejected weeks ago because each run starts blank. The fix is two steps around the existing agent module: a lookup that pulls the known state for each company into the prompt, and a write-back that records the verdict and the reason. Disqualified stays disqualified, with the reason attached, until something actually changes. The scenario's logic is untouched. Only its amnesia is gone.
The question is which kind, not whether
Every scheduled agent that runs more than a handful of times develops the same need. The demos hide it, the Knowledge label obscures it, and the compute bill for re-learning the same lessons quietly grows. Once the library-versus-diary distinction is clear, the decision is just matching the tool to the shape of the problem — and making sure the diary exists at all.