Your AI Agent’s Memory Is a Trash Fire. This Might Actually Fix It.
Your AI Agent’s Memory Is a Trash Fire. This Might Actually Fix It.
Every agent framework you’re using right now discards the wrong information at the worst possible time — permanently. A new paper argues the entire field has been solving memory backwards, and the numbers suggest they’re not wrong.
What happened

Most agentic workflow systems handle memory the same way: when a task finishes, they compress the trajectory into a summary, reflection, or skill artifact, then retrieve it later by similarity. The problem is this forces the system to decide what matters before it knows what the next question will be. Researchers from Salesforce AI Research built Just-in-Time Memory (JitMem) around a different principle: keep the raw trajectories, and only curate them at read time, once the current task is actually known. The curator synthesizes a compact, task-specific payload on demand rather than a one-size-fits-all artifact. Critically, because the payload is consumed on the same task that prompted its creation, the training signal is immediate — no long-horizon credit-assignment gymnastics required. Across three standard agentic AI benchmarks — ALFWorld, WebShop, and τ²-bench — JitMem beat the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Tellingly, even an untrained curator beat or matched existing learned write-time methods, suggesting the architectural shift alone does most of the work.
Cold read
These are benchmark numbers on ALFWorld and WebShop — controlled, narrow-domain environments that have been saturated by academic attention and carry real benchmark contamination risk; gains there don’t automatically travel to messy production environments. The τ²-bench gain of 3.9 points is the most credible signal precisely because it’s the smallest — and it’s the benchmark that probably most resembles a real support or operations workflow. Retaining raw trajectories instead of summaries sounds clean in theory, but the context window cost of pulling multiple full traces at read time could be prohibitive at scale; the abstract says nothing about latency, token cost, or what happens when trajectory libraries grow large. The curator is trained on “immediate task success,” which is a cleaner signal than delayed credit assignment — but it assumes you have a reliable success signal at all, which many real agentic deployments do not. The 16-point gains are dramatic enough to be suspicious; expect significant regression when the task distribution shifts outside training conditions.
What it means for you
- Signal maturity: 2/5 — Benchmark wins are real but production readiness is undemonstrated
- Who gets hurt: Vendors selling write-time memory layers — reflection stores, skill libraries, workflow distillation pipelines — as core moats
- What breaks if this is true: The case for expensive offline trajectory distillation collapses; the value shifts to real-time curation infrastructure and raw storage, commoditizing what several agent-framework startups currently charge for
- Why it might not land: Storage and inference costs scale badly; keeping raw trajectories for millions of agent runs is not free, and read-time synthesis adds latency on the critical path of every task
- Watch for: Any major agent-framework provider (LangChain, CrewAI, or a cloud provider’s native agent stack) shipping a “read-time curation” memory option; that’s the sign this architecture has cleared the cost-feasibility bar
Forecast as of 2026-09-24
By Q3 2027, at least one production agent framework will ship a configurable read-time curation memory module citing this paper or its descendants — but write-time memory will remain the default option in all major frameworks due to storage and latency constraints, meaning JitMem stays a research-first feature for at least 18 more months.
Source: Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents — Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih Yavuz, Shafiq Joty. https://arxiv.org/abs/2609.27334v1
