Your LLM Is Rereading the Same Docs Millions of Times and Paying Full Price
Your LLM Is Rereading the Same Docs Millions of Times and Paying Full Price
Every time a user asks a second question about your knowledge base, your inference stack throws away everything it just learned and starts over. A solo researcher just shipped a memory layer that makes that waste measurable — and claims to have fixed it. The numbers are real enough to make your cloud bill feel embarrassing.
What happened
Sietse Schelpe measured how stateless LLM serving actually wastes compute, then built Galahad, a memory layer for vLLM, SGLang, and llama.cpp that caches and restores the model’s key-value (KV) state byte-for-byte across requests. The core finding: across seven real-world datasets, 98.7% of prompt tokens were text the model had already read before — pure repeated work. The system has two components: Taliesin saves and reloads KV state for identical byte sequences; Blaise acts as a smart retriever, passing the model only the slice of a document that a given question needs — making it a more precise alternative to naive retrieval-augmented generation. On a 97,000-token recall test with Gemma 4 31B, the baseline (no Galahad, 12,000-token context window) answered 10 of 100 facts in 9.3 s at 2,754 J per question. Taliesin alone answered 98 of 100 at 3.0 s and 572 J. Blaise added on top hit 100 of 100 at 0.59–0.64 s and 200–213 J — beating a tuned RAGFlow pipeline that managed only 77 of 100. The restored KV state is bit-identical: all 262,144 output logits matched after restart and rehydration, and Galahad was verified across all 30 models tested under vLLM, failing closed on any state it cannot verify.
Cold read
This is a single-author preprint with a single benchmark dataset design (one corpus, one model family, one recall task format) — the 100-of-100 result is not evidence of general superiority over retrieval-augmented generation pipelines at production scale. The 28 kJ one-time corpus storage cost and the 13-question break-even point are reported for a specific hardware setup that is not disclosed in the abstract; your break-even on A100/H100 clusters at commercial electricity pricing could look very different. Byte-exact KV reuse only fires when prompts share identical bytes — any prompt templating variation, user personalization, or system prompt prefix change breaks the cache hit entirely, and real production traffic is rarely that clean. The comparison baseline (plain llama.cpp with a 12,000-token context and no cache) is a weak opponent; a well-tuned chunked RAG with prefix caching already deployed in vLLM would be a fairer fight. Factual consistency guarantees from bit-identical logits are a strong engineering claim, but the paper does not address what happens to cache validity when the underlying model weights are updated — a routine event in any live product.
What it means for you
- Signal maturity: 2/5 — single-author preprint, narrow benchmark, no production traffic validation
- Who gets hurt: RAG middleware vendors (LlamaIndex, LangChain wrappers, RAGFlow) whose differentiation is retrieval quality that Blaise’s 668-token precision starts to undercut
- What breaks if this is true: The business case for expensive embedding pipelines and vector databases weakens on static corpora — if KV state persistence is cheaper and more accurate, embeddings become the slower, lossier option for that use case
- Why it might not land: Prompt variance in real deployments destroys cache hit rates; the 98.7% reuse figure is a dataset average, not a production measurement, and byte-exact matching is a fragile dependency for any system with dynamic context injection
- Watch for: vLLM or SGLang merging a Galahad-adjacent KV persistence PR into their main branch — that is the signal that the inference infra community has validated the approach, not just the paper
Forecast as of 2026-10-02
By Q2 2027, at least one major hosted inference provider (Fireworks, Together, or Anyscale) will ship a named “persistent KV cache” or “stateful context” feature — but independent benchmarks will show cache hit rates below 60% on real multi-tenant traffic, well short of the paper’s 98.7% figure.
Source: Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost — Sietse Schelpe. https://arxiv.org/abs/2609.39358v1
