Free forever, no credit card.Get Started for Free →
← All posts
September 30, 2026 · 5 min read

Your Scheduled Agent Re-Embeds the Same Memory Every Run. That Is the Bill.

Your Scheduled Agent Re-Embeds the Same Memory Every Run. That Is the Bill. Month-end arrives and the vector database invoice is triple what it was when the agent launched. Nothing about the workload changed. Same schedule, same tasks, same data sources. The agent is not doing more work. It is storing the same work more times. This is the failure mode nobody prices into a DIY memory layer. It does not show up in week one, because in week one there is nothing to duplicate yet. It shows up in we

Your Scheduled Agent Re-Embeds the Same Memory Every Run. That Is the Bill.

Month-end arrives and the vector database invoice is triple what it was when the agent launched. Nothing about the workload changed. Same schedule, same tasks, same data sources. The agent is not doing more work. It is storing the same work more times.

This is the failure mode nobody prices into a DIY memory layer. It does not show up in week one, because in week one there is nothing to duplicate yet. It shows up in week four, when every run has spent a month quietly writing copies of copies.

The anatomy of one expensive run

Walk through what a typical scheduled agent does at 2am. It wakes up, retrieves the top-k chunks it needs from the vector store, blends that context into its working prompt, does the task, and then stores the outcome. The storage step is where the money leaks: many implementations take the retrieved chunks, now woven into the response, and write the whole thing back as fresh embeddings. The original fact gets re-encoded, re-indexed, and billed again.

Now multiply by the schedule. A nightly lead-scoring agent re-reads the same ideal-customer-profile notes every run and writes most of them back. An invoice-triage agent pulls the same vendor rules each night and stores them again with slightly different wording. Each run adds a thin new layer of near-duplicates on top of yesterday's. After thirty nights, the store holds thirty versions of the same baseline, and the bill reflects thirty versions, not one.

Managed vector stores charge on two meters at once: stored vectors and query volume. The re-embed loop feeds both. Storage grows because every run adds vectors. Query spend grows because every retrieval scans a larger, noisier index. A 2026 cost breakdown put Pinecone serverless reads near $0.096 per million and a self-hosted Qdrant node at roughly $50 to $80 a month on a small VM. Neither figure is scary on its own. Both become scary when the write pattern behind them is exponential in practice and linear nowhere.

Why cron jobs are the worst case for this

Interactive agents leak too, but usage caps the damage: no conversation, no writes. Scheduled agents have no such governor. The schedule is the usage. Every tick of the cron is another retrieve-blend-rewrite cycle, running whether or not the world produced a single new fact worth remembering.

That is what makes the diagnosis easy, at least. Chart your vector count over 30 days. A healthy memory layer grows in steps, jumping when the agent genuinely learns something. A leaking one climbs on a smooth slope that never flattens, because the slope is the schedule, not the learning.

The runbook: four gates for the write path

Treat embedding writes the way you would treat any write path with a cost attached: gate it.

Gate one: similarity check before every write. Compare the candidate embedding against recent entries and skip the write when something close enough already exists. The retrieved context that causes most duplication is, by definition, already in the index, so this gate catches the biggest offender with the smallest code change.

Gate two: expiring raw history. Keep verbatim turns for a day or two, then collapse them into one summary embedding and delete the originals. Raw conversation turns are the highest-volume, lowest-value writes in the system. Letting them live forever is paying permanent storage rent for temporary information.

Gate three: two tiers of memory. Episodic and semantic memory should not share one write path. A resolved clarifying question is episodic; it dies with the session. A durable fact about the business is semantic; it earns a permanent embedding. One code path for both is easier to write and exactly how the bill spirals.

Gate four: a weekly consolidation pass. Merge near-duplicates, drop everything nobody has retrieved in 30 days, and re-embed the survivors as tighter summaries. Operators who run TTLs, dedup checks, and consolidation together report storage falling by more than half with no loss in recall quality. More engineering than free writes, and the difference between a memory system and a landfill with a search index.

When the gates are someone else's problem

Four gates, a summarization pipeline, a weekly cron, retention policies: that is a second system to build and babysit next to the agent itself. The hosted alternative exists for teams that would rather buy the outcome than build the machinery.

Vilix AI is cloud-hosted agent memory with no infrastructure to run: no vector database to provision, no embedding pipeline to tune, no metered invoice tracking every write. A single account keeps the complete conversation history, beyond just extracted facts, and every tool you connect reads that same memory over MCP, so Claude, Codex, Cursor, OpenClaw, and Hermes all work from one shared state. Recall is semantic combined with keyword search, and it loads only the relevant context rather than the entire archive, which ends the re-read-everything pattern at the source. The free plan is free indefinitely, the Pro trial runs 7 days with no credit card required, and your data stays portable: export it all in an open format whenever you want, or delete single memories or the whole account instantly.

If a DIY vector store already backs your agents, the four gates above apply wherever the embeddings live. If you are picking the memory layer today, do the honest math first: the DIY price is the database sticker plus deduplication plus summarization plus the consolidation cron plus the hours spent watching all three behave.

FAQ

My agent's traffic is flat. Why does storage keep climbing? Scheduled runs write on a fixed cadence regardless of traffic. Each run retrieves context and stores most of it back, so the vector count grows with the number of runs, not the number of users.

Will a cheaper vector store solve it? It cuts the unit cost, not the write pattern. The same loop fills a cheap index the same way it filled the expensive one. Gate the writes first, then optimize the store.

What is the single highest-leverage fix? The pre-write similarity check. It is a small change that directly breaks the cycle responsible for most duplicate embeddings.

How does a hosted memory layer change the cost equation? It replaces a metered invoice that scales with writes and queries with a flat subscription. The deduplication, summarization, and retrieval discipline become the provider's engineering problem instead of yours.

Get Started for Free

Persistent memory across ChatGPT, Claude, and the AI tools you already use in Vilix AI.

Get Started for Free

Free forever, no credit card.

Keep reading
Your Dify Workflow Starts Every Run From Zero. Here Is the Memory Fix

Your Dify Workflow Starts Every Run From Zero. Here Is the Memory Fix Picture the nightly workflow. Every evening at 9pm it pulls the day's support tickets, drafts replies, and files a summary. Monday night it handles a tricky billing dispute and writes a careful resolution. Tuesday night the same customer writes back, and the workflow stares at the thread like it has never seen it before. Because it hasn't. Every run is the first run. Dify documents this behavior without apology. Workflow app

OpenAI Dots Explained: The Always-On Agent With Memory You Can't Control

OpenAI Dots Explained: The Always-On Agent With Memory You Can't Control OpenAI used its DevDay keynote on September 29, 2026 to make its biggest bet yet on agents that keep working when you are not looking. The product is called Dots: always-on, personal agents powered by the GPT-6 Astra model, each with its own cloud computer and browser, pursuing goals in the background with minimal oversight. If you run scheduled or recurring AI workflows, Dots matters to you even if you never touch one. T

OpenAI vs Vilix AI: An Honest Comparison for Scheduled-Agent Memory

OpenAI vs Vilix AI: An Honest Comparison for Scheduled-Agent Memory People keep asking which one wins: OpenAI or Vilix AI. It is the wrong question, and the wrong framing costs real time when you build scheduled agents. OpenAI builds models. Vilix AI is a memory layer that sits outside any model, reachable over MCP from whatever client or agent you run. One is the brain, the other is the notebook. Brains forget by design. Notebooks remember. This is the honest comparison: what each one actual