Your Scheduled Agent Is Confidently Wrong About Last Week. False Memories Are Why.
Your scheduled agent ran fine on Monday. By Friday it was confidently acting on a Tuesday that never happened. The setup is familiar: a weekly report agent wakes up every Friday, reads its stored context, and drafts the client summary. Last Friday it wrote that the client had approved the new pricing on Tuesday's call. There was no Tuesday call. The "memory" came from the agent's own notes, written the week before when it misread a tentative discussion as a decision. One wrong sentence, stored
Your scheduled agent ran fine on Monday. By Friday it was confidently acting on a Tuesday that never happened.
The setup is familiar: a weekly report agent wakes up every Friday, reads its stored context, and drafts the client summary. Last Friday it wrote that the client had approved the new pricing on Tuesday's call. There was no Tuesday call. The "memory" came from the agent's own notes, written the week before when it misread a tentative discussion as a decision. One wrong sentence, stored as fact, laundered into truth within a week.
This is the failure mode operators underestimate: false memory. Not a prompt injection. Not a poisoned document. The agent invented something, wrote it down, and now believes its own invention, and no human is ever in the room to say "wait, that didn't happen."
Why agents write fiction into memory
Long-term memory for agents is usually a write path with no fact-checker. A typical scheduled setup ends each run the same way: the agent summarizes what happened, extracts "key facts," and stores them for next time. Every step of that pipeline can confabulate.
The summarizer promotes maybes into definitelys ("the client seemed open" becomes "client approved"). The extractor merges separate events into one. Retrieval then returns the fiction with the same authority as a real record, because the store has no column for "how sure are we."
A chatbot's confabulations get corrected by the human in front of it. A scheduled agent's memories are only ever read by the agent itself, so a false memory that survives its first retrieval tends to survive forever: each retrieval cites it, acts on it, and writes a new memory confirming it.
The research stopped being theoretical in July 2026
Two studies this year showed how cheaply false memories get planted and how hard they are to spot:
- MemGhost: one crafted email got an inbox-checking agent to write a false note into persistent memory, hide the write from its reply, and let it sway later sessions. It worked in 87.5% of background runs against OpenClaw on GPT-5.4 and 71.4% against a Claude Code SDK agent on Sonnet 4.6.
- GhostWriter (Torres et al., arXiv 2607.06595): injection into agent long-term memory succeeded at near-universal rates around 98%, with about 60% of planted memories later activated.
Those are attacks. The everyday version needs no attacker: the agent hallucinates a detail, stores it as a lesson learned, and the next run treats it as institutional knowledge. Same mechanism, no adversary required.
What false memories cost an operator
When a memory is wrong, the damage is not a bad answer in one chat. It is a wrong action, repeated on a schedule. The outreach agent "remembers" a prospect already replied positively, so it skips the follow-up that would have closed the deal. The reporting agent cites a revenue figure it invented, and the number goes to a client in a PDF. The support agent "remembers" a customer's issue was resolved, and closes the ticket without checking; and two agents share a store, so one writes a fiction the other acts on.
The agent never says "I think I remember." It retrieves, and retrieval looks like knowledge.
A runbook for memories that might be lying
You cannot stop a model from ever confabulating. You can stop confabulations from becoming permanent infrastructure. The controls sit on the write path and the read path, not in the prompt.
Stamp every memory with provenance. Each record should carry its run, tool output, and source document. "Client approved pricing, per the agent's summary of run #482, no source document attached" is a claim you can check.
Separate observed from inferred. "Zendesk ticket #1042 marked resolved" is observed. "Customer is satisfied with the resolution" is inferred. Store them apart, so a contradicted inference can be replaced without touching the observation. Most false memories live in the inferred column, wearing the authority of the observed one.
Verify before persisting consequential facts. For writes that drive money, client messages, or access changes, check deterministically: does this "fact" appear in the tool output it claims to come from? A string match against the source transcript catches a surprising share of confabulations.
Audit on retrieval. Before a memory enters a consequential run, check it against newer records. A stored "decision" from March that contradicts a May email gets flagged, not trusted.
Run contradiction sweeps. Periodically compare memories against each other and fresh ground truth. Two records that cannot both be true are a finding, not a tie to break by recency.
What the memory layer underneath should give you
The runbook is yours to operate, but the layer underneath decides how expensive it is.
The most useful property is that memories are not bare extracted sentences. A store of isolated "facts" is unauditable: no source, no run, no transcript to check the claim against. Full conversation history with source attribution lets every memory trace back to the run and tool output that produced it.
Second, the store should not live in a JSON file on the agent's VM, where anything with file access can rewrite the past. Server-side storage behind your account removes that class of silent tampering.
Third, deletion has to be real. A false memory you cannot remove is a lie you cannot retract. You need to delete one record or wipe the account instantly, from any connected tool or a dashboard.
That is how Vilix AI is built. It is a cloud-hosted memory layer: zero infrastructure to operate, no local database on a VM. The same memory follows your agents across every tool over MCP. Plan in Claude, build in Codex, run the nightly job in n8n, Make, or OpenClaw, and the context, decisions, tasks, and skills come along. It stores full conversation history, not just extracted facts, so every memory carries its provenance. Data is isolated per user. List, update, or delete anything from any connected AI or the dashboard at app.vilix.ai, export everything in a portable format, or wipe the account anytime. Free plan forever, with a 7-day Pro trial that needs no credit card.
The layer will not fact-check your agent for you. But provenance tracking and contradiction audits are far cheaper on a store where every memory has a source you can open than on a pile of confident sentences in a file somewhere.
FAQ
Is this the same as memory poisoning? Related but different. Poisoning is someone else planting a lie. False memory is your own agent inventing one. The defenses overlap: provenance, write verification, and audits catch both.
My agent only reads our own systems. Does this still apply? Yes. Confabulation needs no attacker: a misread ticket or a guess stored as a lesson produces false memories from internal content alone.
Can I just tell the agent to only store verified facts? It helps, then loses. The agent trusts its stored text more than your system prompt, because the stored text reads as its own past. Enforce verification in code on the write path.