Your Scheduled Agent Fabricated Last Run's Results. Here Is Why It Happens
Your Scheduled Agent Fabricated Last Run's Results. Here Is Why It Happens A lead-scoring agent runs every night at midnight. This morning its output mentioned "the demo we did with Acme on Tuesday" as a reason to bump the lead's score. There was no demo with Acme on Tuesday. There is no record of one anywhere: no calendar invite, no call log, no email thread. The agent invented a demo, scored the lead on it, and moved on to the next row. If you run scheduled agents, you have probably seen a v
Your Scheduled Agent Fabricated Last Run's Results. Here Is Why It Happens
A lead-scoring agent runs every night at midnight. This morning its output mentioned "the demo we did with Acme on Tuesday" as a reason to bump the lead's score. There was no demo with Acme on Tuesday. There is no record of one anywhere: no calendar invite, no call log, no email thread. The agent invented a demo, scored the lead on it, and moved on to the next row.
If you run scheduled agents, you have probably seen a version of this. The agent references events that never happened, quotes figures nobody computed, and summarizes "progress since the last run" as if it had any idea what the last run did. It does not, and that is exactly the problem. This post is the operator's guide to the failure: why it happens, the three tells that catch it, and the verification loop that stops it.
The mechanism, stated plainly
Three ingredients combine in every scheduled run, and the result is fabrication:
First, the run starts with amnesia. Nothing from the previous execution is in the prompt unless you put it there. Most operators do not.
Second, the engine underneath is a plausibility machine. Language models generate the most likely continuation of the text they see. They have no internal "I don't know" state that engages reliably, and training has historically rewarded confident answers over honest gaps.
Third, the task demands a story. "Score these leads based on recent activity." "Summarize what changed since yesterday." These instructions assume the agent knows what happened. It does not. So it writes the story anyway, and the story is fluent enough to pass.
Amnesia plus a plausibility engine plus a demand for narrative equals invented history. Every time.
The three tells
You do not need a hallucination-detection framework to catch this. You need to read the output like an auditor for ten minutes. Three patterns give it away:
1. Numbers that look too clean. Real operational data is messy: 4.17%, 23 leads, one weird outlier. Fabricated data rounds itself off. When three consecutive reports all cite tidy figures with no decimals and no anomalies, that is not stability. That is the model generating what a report usually looks like.
2. Events nobody can find. Pick one specific event the agent referenced: a meeting, a deployment, a customer reply. Search for it in the actual systems. Real runs leave traces in tools; fabricated events exist only in the agent's output. If you cannot find the event in any tool the agent can call, the agent could not have found it either. It wrote it.
3. The story changes on re-run. Run the same agent twice against the same inputs and compare the "since last run" sections. Real history is stable. Invented history drifts: the names change, the numbers change, the causes change. If the past is different every time you ask, there is no past. There is only generation.
The verification loop
The fix is a discipline, not a product. Four practices, in order of leverage:
Fetch before you write. The agent's runbook should be: call the tools, read the results, then write. Never write first and hope the tools agree. In practice this means the prompt structure puts tool calls before any summarization step, and the output format requires each claim to reference the tool call it came from.
Keep an append-only run log. Every run writes what it did, what it observed, and what it decided, to a log the next run reads before doing anything. The critical rule: the log records observations, not conclusions. "API returned 14 rows" goes in the log. "Churn is improving" does not, unless the agent can point at the rows.
Separate what happened from what you think happened. This is the one operators skip, and it is the one that matters most. The run log is facts. The report is interpretation. When the two mix, interpretation hardens into fact by the next run, and the fabrication becomes self-reinforcing: this week's invention becomes next week's source material.
Audit on a schedule, not on suspicion. Spot-check three factual claims per report against tool outputs, weekly. Suspicion-based auditing catches the fabrications you already noticed. Scheduled auditing catches the ones you did not.
Where a shared memory layer fits
The run log works, but it is infrastructure you build and maintain per workflow: the schema, the writes, the reads, the cleanup when it grows. The alternative is a memory layer that does this job for every tool at once.
Vilix AI is a cloud-hosted memory layer, so there is nothing to install or maintain, and the same memory is shared across every connected tool over MCP: the n8n workflow, the Make scenario, and the cron script all read and write one store. That matters for this failure specifically, because the fix is "read what actually happened before you write," and a shared store is what makes "what actually happened" available to every run, in every tool. It keeps full conversation history rather than just extracted facts, so the paper trail survives intact, and the data stays portable: export everything or delete it anytime. The free plan is free forever, and the 7-day Pro trial needs no credit card.
One honest caveat: it is a hosted service. If your environment requires the memory to live on your own machines, you want a self-hosted database, and you should budget the maintenance time that comes with it.
FAQ
Should I run the agent more often so it forgets less? Frequency does not create memory. An agent that runs every hour with no memory fabricates twelve times a day instead of once. Fix the memory, not the schedule.
Can retrieval alone cause this, even with a real memory store? Yes. Retrieval that returns the wrong slice, a stale cache, or an empty result the agent smooths over is one of the main mechanisms. The memory store is necessary but not sufficient; the agent still has to cite what it retrieved and refuse when retrieval comes back empty.
Is this worse with cheaper or smaller models? Somewhat. Smaller models confabulate more readily and follow citation discipline less reliably. But the failure is architectural, not a model tier issue. A top-tier model with no memory and a continuity-demanding prompt will still invent the demo.
What if the fabricated details are harmless? They compound. A harmless invented detail in this week's report becomes next week's cited fact. The failure mode is never one wrong number; it is a memory of events that never happened, growing more detailed with every run. Catch it early or rebuild the history later.
A scheduled agent with no memory is not a reliable narrator of your operations. It is a fluent one. The difference between the two is structure: a run log it must read, sources it must cite, and a store of what actually happened that outlives any single run. Put that structure in place and the midnight run stops writing fiction about your business.