Free forever, no credit card.Get Started for Free →
← All posts
October 2, 2026 · 6 min read

Airflow XComs Pass Notes Between Tasks. They Are Not Your Agent's Memory.

Airflow XComs Pass Notes Between Tasks. They Are Not Your Agent's Memory. Every morning your DAG wakes an agent to watch for incidents. It reads the overnight logs, decides what matters, and posts a summary. Last week it escalated a flaky payment webhook and told you it would keep an eye on it. This morning it escalated the same webhook again, as if the previous escalation never happened, and buried the one genuinely new incident halfway down the summary. The agent did its job inside the run. B

Airflow XComs Pass Notes Between Tasks. They Are Not Your Agent's Memory.

Every morning your DAG wakes an agent to watch for incidents. It reads the overnight logs, decides what matters, and posts a summary. Last week it escalated a flaky payment webhook and told you it would keep an eye on it. This morning it escalated the same webhook again, as if the previous escalation never happened, and buried the one genuinely new incident halfway down the summary. The agent did its job inside the run. Between runs, it remembers nothing.

If you run agents on Airflow, the tempting fix is already sitting in the toolbox: XComs. Stash the agent's state with xcom_push, pull it back with xcom_pull on the next run, and the agent carries on like nothing happened. Plenty of operators try exactly this. It holds together for a week or two, then the cracks show, and they all trace back to one misunderstanding: XComs were built to hand notes between tasks, not to remember anything.

The clipboard and the shift log

A useful picture: XComs are the clipboard handed from station to station along an assembly line. Each station reads the note, does its work, clips a new note on top, and passes it down. That is precisely what they do well. A task pushes a small value, a downstream task pulls it, and the handoff is clean, typed, and visible in the UI.

Agent memory is the shift log. It is the record of what happened across every shift: what was decided, what failed, what changed, what the next shift should do differently. Nobody expects the clipboard to serve as the shift log, because the clipboard is wiped and re-clipped every run while the log accumulates. XComs behave like the clipboard in every respect that matters, starting with scope.

Scoped to the run, blind to the schedule

XCom values are stored in Airflow's metadata database keyed by DAG, task, and run. That keying is the whole story. A scheduled run that fires on Tuesday has no natural access to Monday's XComs. You can force the issue with xcom_pull(dag_id=..., include_prior_dates=True), reaching back into the previous run explicitly, and many operators build the push-at-the-end, pull-at-the-start pattern around it.

Notice what that pattern really is: hand-maintained plumbing that simulates memory. It works until the day it does not. A backfill replays old dates and your "prior run" logic pulls the wrong history. A manual trigger creates a run id the pattern never anticipated. A renamed DAG orphans the entire chain. Real memory makes continuity automatic. The XCom pattern makes amnesia automatic and continuity a side project you debug at 2 AM.

Lookup by key is not recall

Set the plumbing aside and a deeper gap remains. Memory for an agent means answering questions like "why did we stop alerting on the EU region," "which supplier did we already disqualify," or "what was the conclusion of last month's capacity review." Those are questions about meaning.

XComs only answer questions about keys. To pull a value, the agent must know the exact key it was stored under, in the exact task, in the exact prior run. There is no search across values, no retrieval by similarity, nothing that finds the decision you forgot you made. The agent can only remember what it already knows how to ask for, which defeats the purpose. Scheduled agents fail most expensively not when they forget a value, but when they confidently reconstruct a past they cannot actually consult. Key-value lookup gives them no defense against that.

Your scheduler's database is the wrong vault

There is also a physical problem. XCom payloads live in the metadata database, the same database the scheduler, the webserver, and the workers depend on. The standing guidance from the Airflow community is unambiguous: XComs are for small values. Identifiers, short strings, compact JSON. They are not for conversation transcripts, tool-call histories, or state that grows a little more every single run.

Stuff a growing memory into that database and the whole instance pays for it: slower scheduling, slower UI pages, and size ceilings that vary by backend and bite without warning. You can bolt on a custom XCom backend that parks the bytes in object storage, but then you are operating storage infrastructure to prop up a handoff mechanism. Memory that grows daily belongs in a store designed for growth, not in the scheduler's own pocket.

One tool's notebook in a multi-tool world

The final problem is the one operators feel last and regret most. XComs are Airflow's private notebook. Whatever the DAG's agent learns stays in the DAG. The agent you run in Claude Code, the workflow in n8n, the assistant on your phone: none of them can read it, so each of them re-learns the same facts, re-makes the same mistakes, and re-asks the same questions.

This used to be a tolerable limitation. It is becoming the central one. Agent setups are multi-tool by default now, and Airflow itself is moving that way: the new common.ai provider runs agent loops inside the worker and lets those agents call MCP servers as toolsets. Your Airflow agent already reaches outside Airflow for its tools. Keeping its memory locked inside Airflow while its tools live outside is backwards.

Give each mechanism its real job

None of this is an argument against XComs. They remain the right tool for small handoffs inside a run: the run id, the artifact path, the row count, the boolean that says extraction succeeded. Keep them, and keep them small.

The agent's memory belongs in a shared layer with three properties XComs will never have: it is queryable by meaning, it grows without punishing the scheduler, and every tool the agent runs in can reach it. Vilix AI is built as exactly that layer. It is cloud-hosted, so there is zero infrastructure to run and nothing to maintain on your side. Every AI tool connects over MCP to one account, so the same memory follows the agent from the Airflow worker to Claude, Codex, Cursor, OpenClaw, Hermes, or any MCP-compatible tool. It keeps full conversation history rather than just extracted facts, so past decisions stay reviewable in their original context. And retrieval works the way agents actually ask: semantic search paired with keyword matching, so "why did we stop alerting on the EU region" finds the answer even when nobody filed it under a tidy key.

It starts free and stays free on the free plan, with a 7-day Pro trial that does not require a credit card. Your data stays yours: export it all in a portable format at any time, delete single memories, or wipe the account instantly.

The practical change is small. XComs keep doing the handoffs they were made for. The agent gains two new habits: read the shared memory when the run starts, write back what happened before the run ends. Do that, and the morning incident summary stops re-escalating last week's webhook, because the agent that wrote the escalation is the same agent reading the log, run after run. The clipboard stays a clipboard. The remembering finally has somewhere to live.

Get Started for Free

Persistent memory across ChatGPT, Claude, and the AI tools you already use in Vilix AI.

Get Started for Free

Free forever, no credit card.

Keep reading
Your AI Agent Forgets Everything Between Sessions. Here Are the 4 Fixes That Actually Work

Quick answer: AI agents forget everything between sessions because language models are stateless. Every run starts with an empty context window, and when the session ends that window is destroyed. Nothing carries over unless you deliberately stored it somewhere else. The fix is a persistent memory layer the agent reads when it starts and writes to before it stops. Four honest ways to do that: your provider's built-in memory, instruction files, a self-hosted memory layer, or a hosted shared memor

Context Lock-In: Why the Better AI Tool Feels Worse

Context Lock-In: Why the Better AI Tool Feels Worse You hear the buzz about a new coding CLI. Everyone says it is sharper than the one you use. You install it, point it at a task your current tool handles in seconds, and watch it fumble. It asks questions your old tool stopped asking months ago. It suggests patterns you abandoned back in March. It misses the conventions your whole codebase runs on. It feels like a junior developer who joined the team this morning. So you conclude the obvious

Your Lindy Agent Remembers the Chat, Not the Run

Target query: do Lindy agents remember between tasks dev.to title: Do Lindy Agents Remember Between Tasks? What Actually Persists vilix.ai title: Your Lindy Agent Remembers the Chat, Not the Run Your Lindy Agent Remembers the Chat, Not the Run Picture the Monday standup digest. Your Lindy agent has been running it for a month, and the setup promised it would get more helpful over time. This Monday it does the digest perfectly: same format, same sources, same tone. Then it flags a "new" compet