If you run agent memory in production, more than one process writes to it. Chat extraction writes. A background flush writes. Workflow runs write notes. A scheduler fires skills that read it. Each of those paths was written on a different day by someone looking at a different symptom, and each has its own ideas about limits, fallbacks and when memory is worth fetching at all.
Between August 29 and September 6 I shipped seven memory commits in Vodou. Two of them are features. Five fix a guard that sat in the wrong place, and those five are what this post is about.
A memory map that explains a memory, and a chat that begins with the agent
The two features are small. The memory map in the console can now produce a plain-language summary of a single memory (the endpoint is MCP-servers/Vodou-Console/src/api/brain-summary.ts). And a conversation whose first message comes from the agent, not the person, keeps that message when it is reloaded. The rehydration code in src/conversation-hydrate.ts had been written for conversations that start with the user. src/__tests__/hydrate-assistant-first.test.ts now covers the other case.
What matters more is what the week did to the path memory travels. Before it, only interactive chat reliably got memory. After it, every lane does.
Every lane asks memory now
ABORT … refusing to append on every cycle, and 9,874 more bytes in nine minutes
The first bug was a daily log file with a size ceiling. The flush process checked the size before appending, and every cycle it logged that it was refusing. The file kept growing anyway. I measured it live: 347,567 bytes, then 357,441 bytes nine minutes later, with the abort firing the whole time.
The file had three writers. The flush obeyed the ceiling. The chat extractor and the note a workflow run writes at the end never looked at the size. So the guard worked perfectly and meant nothing. The log line was accurate about the one writer it came from and misleading about the file.
The second half was worse. The ceiling was 100,000 bytes, documented as covering "heavy workdays; typical days are 10-30KB." I re-measured. Every daily file for the preceding ten days was over it: 600, 402, 398, 380, 353, 256 KB. The day I measured had 1,176 bullets, 669 of them from IDE capture. Deduplicating the whole file recovered 4 to 9 percent. That was real content. The ceiling had been wrong for weeks, and nobody noticed because the one writer that enforced it was the one nobody was watching.
The fix moved the rule into a single predicate that all three producers call. It took one bug to learn this, and I'll say it once: a guard in one writer is not a rule.
key_pool_ms=4049 on a twelve-word query
The same day, ordinary searches were logging four seconds in one stage while the rest of the search took about 300 ms.
One search, two stages
That stage reads from a RAM cache of every chunk vector and question-key vector. The cache rebuilt from scratch every 15 minutes, and also whenever a reconcile shrank a pool. It rebuilt on the request path, under the write lock, so the first search after expiry paid for it and every concurrent search queued behind it. The design note said "~0.5s pool load" for 92k keys. Measured on August 29: 57k chunks plus 140k keys, 3.2 to 3.6 seconds single-threaded, 289 MB. Every daemon restart that day started the cache cold.
The fix did not change what the cache holds. The request path still uses the full 15 minute staleness threshold. A separate lane on its own connection ticks every minute with a threshold two minutes shorter, so when the cache is about to expire, the rebuild happens there instead of inside someone's search. A tick with nothing to do costs two count/max queries.
One boolean, two subsystems, and 140 scheduled turns with zero memory
In MCP-servers/Vodou-Console/src/llm.ts, a flag called skipPrefetchForWorkbench skipped a prefetch step for scheduled skill runs. That was a real 1.5 to 2 second latency win, and it stays. But the same flag also gated getMemoryContext, and nobody ever argued for that. After a July change resolved scope for skill fires, every scheduled skill turn reached none of the 58k memory chunks. On receipts it was 140 of 140, logged quietly as {"lane":"hook_memory","chars":0,"ms":0}. Interactive chat was fine, so every manual test passed.
Memory now runs unconditionally for every lane. The split does not name any lane, on purpose, so a lane added next year cannot inherit the starvation.
A session limit that resets at 6:30 a.m. wrote permanent junk
On the night of September 3 the model subscription hit its session window. Every extraction came back with a CLI error saying the limit resets at 6:30 in the morning. Thirteen conversations stuck in the extraction ledger, and the flush treated that as a real failure and fell back to a heuristic extractor.
The heuristic copies any transcript line containing "will " or "should " verbatim. Five lines landed in the daily log. One was a whole QA-nightly JSON blob, because it contained "should be non-empty". Two were prompt-template lines. The extraction health check went red and paged every ten minutes for a day.
The flush already knew this class of failure. When the process cap is exceeded, it skips and keeps the buffer, because that condition clears on its own. A usage limit is the same condition on a longer clock, and it took the fallback anyway. All three fallback sites now skip on it.
The first-run fix belongs to the same family. The release archive shipped two embedding models and no cross-encoder, while reranking defaulted on for Apple Silicon. So a new user's first memory search silently started a roughly 1.1 GB download inside a query that looked hung. Bundling it would have quadrupled a 340 MB archive. Turning reranking off would have made retrieval worse for everyone to fix a first-run problem. The engine now resolves the reranker against what is on disk and says so when it substitutes. It has to say so, because relevance floors are calibrated per model: one model scores identical pairs at 0.875 where another scores 0.994.
The invariant: every constraint on a shared resource is checked by every writer, through one function
All five bugs have the same shape. A decision about a shared thing (a file's size, a cache's freshness, whether a turn gets memory, whether a failure is permanent) was made by one caller instead of at the resource. Here is a version you can check:
For every shared memory resource, list every code path that writes it or skips it. Every one of them calls the same predicate for each limit, and that predicate is the only place the limit is defined.
It is either true of your codebase or it isn't. A corollary covers the fallback bug: a failure class that clears on its own is classified before any fallback that writes durable state.
Check your own memory writers in five minutes
This check needs nothing from Vodou. First, find every writer of your memory files. Substitute your own path fragment:
grep -rnE "appendFile|createWriteStream|open\([^)]*['\"]a['\"]|O_APPEND" src/ \
| grep -iE "memory|daily|log|notes"
Passing: every hit sits in, or calls into, one module that owns the limit. Failing: two or more hits with their own size check, or none. Then confirm the limit actually holds while it claims to:
f=path/to/todays-memory-file
a=$(stat -f%z "$f" 2>/dev/null || stat -c%s "$f"); sleep 540
b=$(stat -f%z "$f" 2>/dev/null || stat -c%s "$f"); echo "$a -> $b"
If the size grows while your logs say the limit is refusing, you have my first bug.
Second, if you log per-turn context injection, group it by lane:
SELECT lane,
COUNT(*) AS turns,
SUM(CASE WHEN memory_chars = 0 THEN 1 ELSE 0 END) AS empty_turns
FROM turn_receipts
WHERE created_at > datetime('now', '-7 days')
GROUP BY lane
ORDER BY 1.0 * empty_turns / turns DESC;
Passing: empty rates are similar across lanes. Failing: one lane at or near 100 percent while chat looks healthy. If you don't log this per lane, that is the finding. You can't see this bug without it.
Third, grep your extraction fallback for rate-limit strings (429, rate limit, quota, session limit). If the only branch that handles them leads to a heuristic writer, a bad night becomes permanent memory.
Where the published memory advice stops
Cohorte's piece on governed writes frames memory as "a write-permission problem wearing a storage costume," and I agree. But it asks who may write. My bugs were about writers that were allowed to write and simply never checked the rules. LangChain's guide to agent memory recommends a background process that extracts and generalizes. It doesn't say what that process should do when the model behind it is unavailable until morning, and that gap is how junk ended up stored as fact. OptMem goes append-only with no daemon, which removes my cache problem entirely and keeps my growth problem: with 600 KB days of real content, a log nobody can afford to read comes back. The crystals post argues that memory should be pushed to the agent before it acts. My scheduled skills are the case for it. A lane that never asks never gets anything, and nothing complains.
The rebuild still costs 3.2 seconds, just somewhere else
The cache fix moved the cost. It did not remove it. A full rebuild is still single-threaded, still about 289 MB, and still grows with the corpus. It now happens in a background lane instead of inside a search, which covers steady state. But the numbers in the original design note were off by a factor of six within a few weeks, and I have no reason to think today's numbers will hold any better.
Source: Our memory ceiling logged ABORT while the file kept growing by Chad Priest, from Building Vodou in Public.
















