$ cat posts/kv-cache-elision.md

Paging a 27B model's memory to disk and back

How we kept Qwen3.8-27B going for 457,000 tokens without compaction by evicting and restoring its attention cache on MLX and llama.cpp, what that did and didn't prove, and where MDBX fits in.

12 min read#local-models#kv-cache#qwen#mlx#llama.cpp#mdbx

Anyone who has left a coding agent running on a long job knows the moment. The context window fills up, the harness quietly summarises everything so far into a few paragraphs, and the agent carries on with a slightly vaguer idea of what it was doing. It’s a sensible trick and it mostly works, but it’s lossy by design. The exact error from forty tool calls ago, the odd requirement buried in the third file, the reason you ruled out an approach an hour back: all of it gets squashed into whatever the summariser thought was worth keeping.

I wanted to see whether a local model could keep going without ever doing that. A bigger window only moves the wall, so the idea was to treat the model’s memory the way an operating system treats RAM, with a working set kept close and everything else paged out to disk and brought back when it’s needed. The model was Qwen3.8-27B, running on my M4 Max through MLX and on rented RTX 5090s through a llama.cpp fork, with Pi as the coding agent on top. We built it at Don Works over about a week in the middle of September. The repo is private for now, so every number below comes from its docs and result files.

Why Qwen3.8 made this worth trying

Qwen3.8-27B is a hybrid. Of its 64 layers, only 16 are full attention. The other 48 are Gated DeltaNet layers, which keep a fixed-size recurrent state rather than a cache that grows with every token.

That split is the reason the whole idea holds up. The attention layers are where memory grows with context length, so they’re the part you’d want to page out. The DeltaNet layers have already folded every token they’ve seen into their state, and that state never needs paging at all. If you evict a chunk of attention cache, the model hasn’t lost sight of it completely, because the recurrent layers still carry a compressed trace of the whole stream and the attention layers can have the exact version back whenever it’s restored.

It also means the surgery has to be careful. Anything that touches the cache in bulk, like stock llama.cpp’s sequence range removal, will happily wipe the recurrent state along with the attention entries, and at that point you’re back to running the whole prompt through the model again.

Paging the cache

The mechanism is simple enough to describe. The attention cache is split into 128 token pages, and each page is either hot (live in attention), cold (only on disk) or pinned (never evicted). When the hot set goes over a high watermark, pages are evicted until it’s back under a low one. Before a page leaves, its K and V tensors are written to an immutable canonical store on disk, and when it’s wanted again those exact tensors are loaded straight back into the running cache.

Paged attention cacheA row of 128 token pages. Pinned and hot pages stay in the attention layers, cold pages live only in the canonical store on disk, and a restore copies K and V back without a forward pass. The DeltaNet state spans the whole stream underneath.16 attention layers · live KV, bounded by watermarkspinned rootrecent tailevict: copy K/V out, drop pagerestore: 0 forward passescanonical K/V store on disk · immutable, bounded by free space48 Gated DeltaNet layers · fixed-size state over the whole stream, never touched
Each box is a 128 token page of history. Only pinned and hot pages take up attention. Cold pages exist only in the canonical store until the controller asks for them back.

What makes this different from truncation or RAG is that a restore doesn’t run anything through the model. Qwen uses partial rotary position embeddings, so when pages move about, the surviving keys have to be re-anchored to their new positions, but that’s arithmetic on tensors rather than a forward pass. We wrote the rule down as one invariant and checked it on every run:

model_forward_tokens == new_prefill_tokens
historical_forward_during_restore == 0

Put simply, every token the model has ever seen goes through it exactly once, and bringing old context back costs nothing in model compute.

Some things never leave. The system prompt, the tool schemas and the opening objective are hard pinned, capped at a quarter of the active budget so they can’t crowd everything else out. There’s also a small pinned ledger of tool call boundaries, so the model always knows which tools it called and in what order, even after their output has gone cold.

Deciding what to keep

Paging the cache is the mechanical half. The harder half is deciding which pages matter for the next turn, and I didn’t want the host code making that call with keyword matching or regexes, because that falls apart the moment the context has Japanese requirements in it, or a stack trace.

So a second, much smaller model does it. On the Mac that was Qwen3.5-2B, and on the CUDA box it was a 4B model on its own 5090. History is grouped into semantic units of roughly 1,024 tokens, eight of those make a cluster, and after the main model finishes a turn the controller writes a short summary of each new cluster in the background. Before the next turn it gets asked one question: given what’s about to happen, which blocks should come back, which should be pinned, and in what order should the rest be evicted? If a cluster looks relevant but too coarse, a second call picks within it.

The host treats everything the controller says as opaque IDs. It enforces the pins, the budget, the recent tail and a short lease on anything the controller restores, and if the controller returns rubbish, nothing gets evicted because of it. Because the host never reads the text, the same path worked across Japanese, Arabic, Spanish, Russian and code without any special handling. The 2B controller passed six out of six multilingual cluster probes and kept four rare identifiers intact in every run, with foreground decisions taking 0.8 to 2.3 seconds. The older Qwen3-1.7B was quicker but failed the repeat-run checks, so it didn’t make the cut.

By the end of the week the CUDA lane had moved that whole loop onto the server, running on its own worker alongside the target model’s generation. The plan it produces is only metadata, so if it isn’t ready in time the turn goes ahead on the host’s default retention and the plan gets picked up on the next one.

What we built it on

Pi is the coding agent, and the whole thing plugs into it as an extension that registers a kv-elision provider. It switches off Pi’s own compaction for that provider and leaves Pi’s transcript as the source of truth. The backend fingerprints the transcript, tokenizer, template, tools and thinking mode, so if anything earlier in the conversation changes it fails closed rather than trusting a cache that no longer matches.

On the Mac, the backend sits on mlx-lm and oMLX. MLX gives you direct access to the cache tensors, which made it the easiest place to prove the surgery was exact before scaling anything up.

On CUDA, the product we leaned on most was EVOKE, which runs on its own fork of llama.cpp. That fork exposes attention-only range operations (seq_rm_attention_only and seq_add_attention_only) along with kv_block_save and kv_block_load, which are exactly the primitives stock llama.cpp doesn’t have. Routing removal and position shifts through the attention-only calls is what stops the DeltaNet state getting cleared along with the pages. The one real bit of pain was that EVOKE’s Python shim predated two fields that upstream had since added to the llama.cpp parameter structs, so calling through the old struct shape overwrote memory and aborted the process. We ended up binding the exact pinned C ABI with ctypes before importing the engine.

We also got vLLM working as a third backend, which was more painful than either of the others. Stock vLLM 0.29 lays the attention and DeltaNet cache groups over the same physical pages, so writing into the middle of the attention cache corrupts recurrent state. Our patch packs them side by side instead, which works but roughly halves KV capacity on this model family. That lane also turned up a nasty one. At an idle streaming boundary, the worker’s count of computed tokens trails the scheduler’s by exactly one, and because vLLM keeps the rest of the output, a plan built on the stale count would have silently dropped a real token from the context every time the cache was rearranged.

Where MDBX fits

Every page, block, pin, lease and tag needs to live somewhere that survives the process and can be read without getting in the way of generation. That job goes to a small Go daemon called contextd, which talks over a Unix socket and stores everything in MDBX through Erigon’s mdbx-go bindings.

MDBX is a memory-mapped B+tree store from the LMDB family, with a single writer, lock-free readers and proper ACID transactions. That suits a control plane where one thing writes state as turns complete, while several others (the Pi status line, the controller, monitoring) want to read it constantly. contextd keeps eleven tables covering sessions, blocks, tags, the KV manifest, the page state index and a slot for DeltaNet snapshots, and for every page it holds identity, position, state, a locator and a checksum.

The one firm rule is that tensors never go into MDBX. The canonical K/V store is plain files on disk with a free-space reserve, and under disk pressure the host drops cold, unpinned blocks, least useful first, and never anything resident. MDBX only knows where things are and what state they’re in.

It turned out not to be a bottleneck either. A 128 page batch commit took 6.38 ms in fully durable mode, 0.48 ms with no-meta-sync and 0.28 ms with safe-no-sync, so we left it on durable. By the end of the long run below it was holding 621 blocks and 2,961 tag rows without anyone noticing it was there.

What the runs showed

The first thing to check was whether the surgery was actually exact. On the M4 Max with the 4-bit Qwen3.8-27B, we evicted a 128 token page from the middle of the cache, restored it and compared the logits against an untouched reference.

0.0max abs logit error after restore
20.6 mswarm restore of one page
1.35 swarm prefill of the same page
0 / 48DeltaNet layers modified

So bringing a page back cost about 1.5% of recomputing it, and the top token matched the reference. On CUDA, the equivalent gate (evict, compact, restore, carry on generating) produced 32 out of 32 identical greedy tokens.

The second test was long thinking. With attention capped at 8,192 tokens, the model generated 8,448 tokens of thinking, did it again on a follow-up turn and then answered a continuation correctly, with the forward-token accounting exact the whole way through.

The main one was a long coding run on two 5090s from Vast.ai, with the 27B target on one card and the 4B controller on the other. Pi had to read 24 large evidence files one at a time, with the requirements scattered across Japanese, Arabic and Spanish, then inspect the code, run the failing tests, fix the implementation, rerun the tests and finally answer a question about a fact from the very first file.

Tokens of history457,087
Tokens forwarded through the model457,087
Active attention at the end24,670
Peak active attention26,244
KV tokens evicted364,823
KV tokens restored38,726
Cold KV written to disk29.97 GB
Assistant turns / tool calls36 / 34
Pi compactions0
Tests passing at the end6 / 6
Recall of the fact from the first fileexact

Close to half a million tokens went through a model whose attention never held more than about 26,000 at once. Nothing was recomputed, nothing was compacted, and at the end it could still quote something from the start.

What it didn’t prove

It’s easy to read too much into that table, so it’s worth being clear about its limits.

It isn’t a quality comparison. There’s one synthetic coding workload, and the vanilla Pi control we ran used a smaller version of it (16 files rather than 24, on four 5090s through vLLM), where Pi compacted once and still passed. Both got the job done. Until the same task runs both ways on the same hardware, the claim is that the mechanism works, not that it makes the agent any better at its job.

It isn’t fast yet either. The run took 1,386 seconds, and the controller accounted for 1,016 of them across 250 separate calls, because the first version annotated every block synchronously before each turn. Reworking that into one blocking call per turn plus background cluster indexing should claw most of it back, but we haven’t rerun the comparison, so I don’t have a speedup number to give you.

There are some hard edges too. The DeltaNet state isn’t snapshotted yet, so a process restart or a Pi rewind means prefilling the whole transcript again, which is correct but slow. Images aren’t supported. The vLLM lane still pages mechanically (keep the head and the tail, evict the middle) because the semantic layer hasn’t been ported to it.

My favourite edge was that a bounded working set with greedy decoding has a fixed point. If a turn’s answer doesn’t change anything, the next turn sees exactly the same working set and produces exactly the same answer, and it’ll do that forever. The CUDA lane now spots a repeated answer and samples that turn instead of decoding greedily, but MLX and vLLM still don’t.

Why it was interesting

The thing I keep coming back to is how much the hybrid architecture did for us. On a pure transformer, evicting attention cache really does mean forgetting, and you’d be relying entirely on the controller to guess right. With 48 recurrent layers carrying the whole stream, eviction is closer to moving something from RAM into swap, where it’s still there in a rougher form and the exact copy is a 20 ms load away. I doubt anyone designed Gated DeltaNet with this in mind, but it suits it very well.

It was also a good reminder that deciding what to remember is a language problem, and that host code is better off not pretending otherwise. Handing it to a 2B model that costs next to nothing, while the host sticks to budgets and pins, got rid of a whole category of English-only heuristics that I’d otherwise have written and then regretted.

And it was satisfying to see the approach hold across three quite different runtimes. It needs two things from the runtime: a way to remove and re-add attention ranges without touching recurrent state, and a way to save and load KV blocks with stable layer views. MLX already had them, EVOKE’s llama.cpp fork had them, and vLLM needed patching to get there. Any runtime that exposes both could run the same thing.

Next up is the matched comparison against vanilla Pi, snapshotting the DeltaNet state so restarts are cheap, and porting the semantic layer to vLLM. I’ll write that up once there are numbers worth showing.