Where Your Scheduled Agent's Memory Actually Lives (and Which Layer Breaks First)
Where Your Scheduled Agent's Memory Actually Lives (and Which Layer Breaks First) Picture the Monday 9 AM triage run. Your agent reads the overnight support tickets, escalates the urgent ones, drafts replies for the rest. At 9:20 it logs a note: the Acme Corp billing thread needs a human because the refund exceeds the auto-approve limit. Good run. Tuesday 9 AM, same agent, same workflow. It reads the Acme thread again — and approves the refund automatically, because this morning's run has no id
Where Your Scheduled Agent's Memory Actually Lives (and Which Layer Breaks First)
Picture the Monday 9 AM triage run. Your agent reads the overnight support tickets, escalates the urgent ones, drafts replies for the rest. At 9:20 it logs a note: the Acme Corp billing thread needs a human because the refund exceeds the auto-approve limit. Good run. Tuesday 9 AM, same agent, same workflow. It reads the Acme thread again — and approves the refund automatically, because this morning's run has no idea what yesterday's run decided.
The policy was in the prompt. The decision was in yesterday's run. The agent had one and not the other.
Place one: the conversation window
The most basic form of agent memory is the conversation itself. Every message, every tool result, stacked up and re-sent to the model on each call so it can see the whole exchange. Inside one session this works beautifully. It feels like remembering because the model is literally re-reading the transcript every time.
The catch is the word "inside." When the run ends, the window closes and the transcript is gone from the model's view. Scheduled agents die a little death every run — n8n workflow completes, cron job exits, Zapier run finishes — and the next run wakes up blank.
Place two: the summary file
The usual answer to a dying window is a summary: compress each run into a few paragraphs, keep the last few runs verbatim, and hand the model the summary plus the recent turns at the start of the next run. It is cheap, it is bounded, and it loses exactly the details that matter most.
Consider what survives compression. "Monday's run hit 429 rate-limit errors from the Shopify API; switching to 2 calls per second with the 9:40 retry fixed it" becomes "there were API rate-limit issues on Monday." The lesson — the retry cadence that actually worked — is exactly what summarization discards.
Place three: the fact database
The third place is where facts live independent of any conversation: client preferences, routing rules, decisions that should still hold next month. The minimal honest version is a plain key-value store — one documented pattern uses a simple JSON file on disk, saving pairs like a client's timezone or answer-style preference, then loading the whole set into the prompt at session start. Bigger setups use a vector database: embed each fact, retrieve the relevant ones per query, the same mechanism as RAG.
This is the place most memory tutorials actually teach you to build, and it answers the wrong half of the question. A fact database knows the auto-approve refund limit. It cannot tell you whether yesterday's run already approved that refund. Rules without a record of events produce the Tuesday-morning failure from the intro: the agent knew the policy perfectly and still made the wrong call.
Place four: the state file
The fourth place is the least glamorous and the most load-bearing for scheduled work: the agent's checklist. Which rows were processed, where the last run stopped, how many retries have fired, what is still pending. Frameworks are converging on a name for this — working memory, a persistent scratchpad of task state that survives across sessions and is separate from both conversation history and the fact database.
Without it, every run starts from the top of the queue. The triage agent re-reads all 200 tickets instead of the 14 new ones. The sync job re-uploads the whole catalog instead of the delta. And because re-processing is silent — no error, just wasted tokens and duplicated side effects — the cost shows up on the invoice, not in the logs.
The loop nobody ships
What separates a memory system from four disconnected files is the loop. Two steps wrapped around the actual work:
Read first. Before the agent touches anything, pull the stores: inject the facts, load the checklist, attach the summary of the last run. One documented DIY approach adds exactly this as a lookup step at the start of every run — query a log of past interactions by customer identifier, pull a two-line summary of the last touch, feed it into the prompt. A few hundred milliseconds, a fraction of a cent, and the agent stops treating every customer like a stranger.
Write last. After the work finishes, commit what changed: update the summary, save new facts, advance the checklist. A run that ends without writing is a run the next one cannot build on.
Expire on purpose. Old memories go stale. An eighteen-month-old preference might be actively wrong today. If nothing ever deletes, the fact database becomes a museum of outdated instructions, and every run pays token costs to read exhibits.
Which layer breaks first in practice
In order of how often operators get bitten: the write-back gets skipped (the agent "has memory" in the architecture diagram and amnesia in production); the fact database bloats (memory that is never read is an expensive decoration on the token bill); the summary replaces the transcript (keep full conversation history — summarize for the index, not for the archive); the checklist lives in the agent's head (the first long run that exceeds the window loses its place).
Skip the plumbing
Building all four places yourself means a transcript store, a summarizer, a fact database with retrieval, a state file, the read-write loop, and expiry rules — hosted and maintained. That is a genuine engineering project, and for most automation operators it is not the project they signed up for.
Vilix AI exists so you do not have to build it. Cloud-hosted, zero infrastructure. The same memory follows your agent across every tool over MCP: the n8n workflow, the cron script, Claude Code, the phone app, all reading and writing one shared store. Full conversation history, not just extracted facts or lossy summaries, so the raw material a future run might need is never thrown away.
The free plan is free forever, the Pro trial is 7 days with no credit card, and your data is portable. Export everything or delete it anytime, in a portable format.
FAQ
Why does my scheduled agent forget everything between runs? Because the language model stores nothing between calls, and scheduled runs are separate calls. Unless a store outside the model saved the last run's state — and the next run reads it — every run is a fresh start.
Can I just make the prompt longer instead of adding memory? Only inside one run. A longer prompt does not survive the run ending. Prompt length is capacity; memory is persistence. Scheduled agents need persistence.
What is the smallest memory setup that actually works for a scheduled agent? Three things: a checklist of what is done and pending, a short summary of the last run, and a write step at the end of every run. Add the fact database once corrections start repeating.
Put each kind of memory in a place that survives the run — window, summary, facts, checklist — connect them with a read-first, write-last loop, and your agents stop being new hires every morning.