Supabase for Scheduled Agent Memory: Where It Wins, Where It Breaks
Supabase for Scheduled Agent Memory: Where It Wins, Where It Breaks Every scheduled agent has the same unanswered question: where do its memories live between runs? Not in the prompt; the prompt is rebuilt from scratch every run. Not in the workflow tool; n8n executions and Zapier runs are stateless by design. The memory has to live somewhere outside the agent, somewhere the next run can reach. Supabase keeps winning this argument among automation operators, and it is worth understanding why:
Supabase for Scheduled Agent Memory: Where It Wins, Where It Breaks
Every scheduled agent has the same unanswered question: where do its memories live between runs? Not in the prompt; the prompt is rebuilt from scratch every run. Not in the workflow tool; n8n executions and Zapier runs are stateless by design. The memory has to live somewhere outside the agent, somewhere the next run can reach.
Supabase keeps winning this argument among automation operators, and it is worth understanding why: where it genuinely beats the alternatives, and where it quietly breaks down at 3am.
Why Supabase keeps getting picked
Strip the marketing away and Supabase is managed Postgres with a friendly face: a free tier that covers most agent workloads, the pgvector extension for semantic search, a REST API plus client libraries for Python and JavaScript, and ready-made nodes in n8n with plain HTTP modules in Make and Zapier. For an operator, that means one service covers three jobs: durable key-value facts, semantic recall over past runs, and a queryable log of what the agent did.
Compare it with the alternatives. Redis is fast, but you manage eviction and persistence yourself, and it cannot do semantic search. A local SQLite file dies with the container. A vector-only database handles similarity but is awkward for exact lookups like "the client's current plan tier." Supabase's pitch is one Postgres that does all three, and for scheduled agents that is usually enough.
The three layers an agent actually needs
A useful split: keep three kinds of memory in three places.
Layer one: identity and rules. Small, hand-curated, loaded on every run. Who the agent serves, what it is allowed to do, standing client preferences. In Supabase this is a tiny table, or even one row per client, read in full at the start of every run. It never gets big because it gets pruned by hand.
Layer two: append-only history. A run log. What the agent did, what it decided, what failed. Nothing here goes into the prompt automatically; it gets queried when a question comes up, like "why did the Tuesday digest stop going out?" This is a plain table with a timestamp index. It grows forever and that is fine, because you almost never load all of it.
Layer three: queryable state. The facts the agent needs to count, check thresholds against, or recall by meaning: per-client preferences, past decisions, deduplicated learnings. This is where the upsert pattern and pgvector live.
Splitting it this way matters because each layer fails differently. Confuse them and you end up with one table that is too big to load, too vague to query, and too stale to trust.
The schema pattern that survives real schedules
For the queryable layer, the shape that holds up is boring on purpose:
create table agent_state (
id uuid primary key default gen_random_uuid(),
agent_name text not null,
entity_key text not null,
facts jsonb not null default '{}',
updated_at timestamptz default now(),
unique (agent_name, entity_key)
);
entity_key carries the multi-tenant design: client:acme, workflow:daily-digest, project:launch. One agent serving twelve clients keeps twelve separate fact objects. Skip this and one client's "always CC legal" becomes every client's policy by run twenty.
The run loop is two database calls. At the start of the run: select the facts row and merge it into the system prompt. At the end: extract the durable facts from the run and upsert. For semantic recall across history, add a chunks table with a vector(1536) column and a matching function that orders by <=> distance, then inject the top matches.
Where it breaks
Supabase wins on the happy path. The breakage shows up in four places, all of them specific to scheduled runs:
1. Stale facts with no expiry. Nothing in Postgres knows that "paused until Q3" expired. Without a review cadence, the agent confidently acts on dead information. The fix is a reviewed_at column plus a scheduled job that flags facts older than N days for re-verification, which is to say, a second scheduled agent whose job is babysitting the first one's memory.
2. Overlapping runs. Schedules slip, retries fire, and two runs of the same agent execute concurrently. Both read the same facts, both write back their own version, and last-write-wins silently discards one run's learnings. Schedules that cannot overlap, idempotent keys, and treating the end-of-run upsert as a merge rather than a replace keep this survivable.
3. Embedding cost on every save. Semantic recall means embedding every chunk on every save, which is an API call with a price and a latency on the critical path of the schedule. Batch the embedding step, or keep semantic recall for the history layer and exact keys for the hot facts.
4. The service key problem. Headless scheduled jobs authenticate with the service_role key, which bypasses row-level security entirely. One leaked key is full database access. Scope a dedicated key per agent with minimum grants, rotate on a schedule, and keep the key in the secret store, never in the workflow definition.
The decision framework
Choose Supabase when the data must live in infrastructure you control, when your memory shape is unusual enough that a generic schema fights you, or when you want SQL-level introspection into what the agent believes. Debugging an agent whose brain you can query with select is a genuine superpower.
Choose a hosted memory service when the four failure modes above read like a second job description. The value proposition is the plumbing disappearing: no extraction prompts to tune, no staleness job to babysit, no embedding pipeline on the schedule's critical path. Vilix AI takes this side of the tradeoff. It is cloud-hosted with zero infrastructure for you to manage, and the same memory is available over MCP to every tool you connect, from Claude and Codex to Cursor, OpenClaw, and Hermes, so an agent's context follows it across the whole stack. It keeps full conversation history rather than just extracted facts, which removes the extraction prompt as a failure point entirely. The free plan is free indefinitely, the 7-day Pro trial asks for no credit card, and your data stays portable: export everything in a portable format whenever you want, or delete individual memories or wipe the account instantly. The tradeoff is the honest one: the memory lives in a service you do not operate. For teams whose compliance allows that, it is usually the cheaper total bill than maintaining the DIY version.
Whichever home you pick, the test is the same. Kill the agent mid-week, start a fresh run, and ask it what it learned on Monday. If it answers from memory instead of guessing, the memory layer is real. If it invents an answer, you have a database, not a memory.