Skip to content
All discussions

Week 31 · September 2026

Does Your Agent's Memory Survive a Model Upgrade?

LinkedIn researchers tested four memory formats under model swaps. Fixed-schema KGs lost 0.0004 accuracy. Compressed notes swung 13 percentage points. No errors thrown.

Week 31 · September 2026·8 min read·Research & ideas
AI AgentsMemoryProduction AI

The Hook: Upgrading Should Not Mean Forgetting

Your AI agent has been running for six months. It remembers your preferences, your project history, your team's decisions. Then you upgrade the underlying model from version 4 to version 5. The agent still runs. It still responds. But it has quietly forgotten half of what it knew.

No error message. No crash. Just wrong answers where there used to be right ones.

This is the problem that Ankit Goyal and Jaideep Ray at LinkedIn set out to measure in their paper "Does Your Agent's Memory Survive a Model Upgrade?" published September 4, 2026. They tested four different memory storage formats under controlled model swaps, and the results reveal a landscape of silent failures that most teams are not testing for.

Why This Matters Now

Every major agent framework today offers some form of persistent memory. LangChain has memory modules. AutoGen stores conversation histories. Custom agents compress past interactions into summaries or knowledge graphs. The assumption is that this memory persists across upgrades.

But memory is not just storage. It is a contract between the writer (the model that created the memory) and the reader (the model that consumes it). When you swap the reader, that contract can break in ways that produce no errors but degrade accuracy significantly. The paper quantifies exactly how much.

The Four Memory Formats

The study compares four distinct approaches to agent memory, each representing a real-world pattern used in production systems:

LC-RAW
Complete conversation transcripts. Nothing compressed, nothing lost. The baseline.
RAG
Chunked text stored with vector embeddings. Retrieved by semantic similarity at query time.
NOTES
Model-compressed natural language summaries. Most common in agent frameworks. Most fragile.
KG-fixed
Fixed-schema knowledge graphs with structured records. Rigid format. Best portability.

Each format represents a different trade-off between storage cost, retrieval speed, and information density. What the paper adds is a fourth dimension: portability, the ability to survive a model change without accuracy loss.

The Experiment: A Rigorous Setup

The experimental design is notably careful, following pre-registered hypotheses with a signed Git tag to prevent p-hacking. Here is what they built:

48
Synthetic histories
160
Questions per history
2
Open-weight models
Exact
Match scoring

The two models are Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct-1M, both similarly sized but architecturally distinct. Questions cover direct facts, temporal changes, contradictions, multi-step relations, and aliases. Crucially, answers use randomized codes to eliminate guessing and make exact-match scoring reliable.

Why exact-match matters: Most memory evaluations use LLM-as-judge scoring, where one model evaluates another's answer. This introduces circular bias when testing model swaps. By using randomized answer codes and exact-match scoring, the paper eliminates model judge bias entirely. If the answer is "XK7-BLUE" and the agent says "XK7-BLUE," it is correct. No interpretation needed.

Finding 1: Fixed Schemas Transfer, Compressed Notes Do Not

The headline result is stark. When you swap the model that writes memory with a different model that reads it, fixed-schema knowledge graphs barely notice. Compressed notes can swing by over 13 percentage points.

Format Reader Own-Store Accuracy Inherited Accuracy Delta
KG-fixed Llama 0.8456 0.8445 -0.11 pp
KG-fixed Qwen 0.9878 0.9880 +0.02 pp
NOTES Llama 0.3762 0.4753 +9.91 pp
NOTES Qwen 0.4719 0.3391 -13.28 pp

KG-fixed shows a combined shift of +0.0004 with standard error 0.0020 across both migration directions. That is statistical noise. The schema enforces a contract that both models can read regardless of who wrote it.

NOTES, by contrast, show a 23-percentage-point spread depending on direction. When Llama reads Qwen's notes, accuracy jumps by 9.91 points. When Qwen reads Llama's notes, accuracy drops by 13.28 points. Same format, opposite outcomes. Averaging these two numbers gives a misleading -1.69 points that hides the real variance.

The averaging trap: If you test memory portability by averaging both migration directions, you will conclude NOTES lose only ~1.7 percentage points. In reality, one direction gains 10 points and the other loses 13. Averaging conceals the failure. The paper's insistence on reporting each direction separately is one of its most important methodological contributions.

Finding 2: Mixed Embeddings Are Worse Than Either Extreme

For RAG systems, model upgrades often mean new embedding models. The natural migration strategy is incremental: keep old embeddings, add new ones as new data arrives. The paper tested this with BAAI/bge-large-en v1.0 and v1.5 (both 1024-dimensional).

42%
Old index baseline
54%
Full re-embed
47%
50/50 mixed index
59%
Ideal routing bound

Full re-embedding gains 11.90 percentage points over the old index. The 50/50 mixed index recovers only 4.96 points of that gain, forfeiting nearly 60% of the improvement. No errors. No warnings. The system just retrieves worse chunks because embeddings from different model versions occupy slightly different regions of the vector space.

The ideal routing upper bound (59.60%, a hypothetical where each query is routed to whichever embedding version works better) shows that the information is there. The mixed index just cannot find it.

The practical trap: Incremental re-indexing is the default migration path in most vector databases. It is cheaper than full re-embedding and avoids downtime. But the paper shows it captures less than half the upgrade benefit. For production RAG systems, this means you should either commit to full re-embedding or accept that you are leaving significant accuracy on the table.

Finding 3: Same Symptom, Different Disease

One of the paper's most useful contributions is a diagnostic decomposition that isolates where information is lost. NOTES and RAG both lose accuracy compared to raw transcripts, but for entirely different reasons.

NOTES: Total Loss 0.584
Construction 80%
Reader 14%
Retrieval 6%
RAG: Total Loss 0.450
Retrieval 81%
Reader 18%
Construction (writing) Retrieval (finding) Reader (interpreting)

For NOTES, 80% of accuracy loss happens during construction, when the model compresses the conversation into a summary. Information is discarded at write time and cannot be recovered at read time. The summary itself is the bottleneck.

For RAG, 81% of accuracy loss happens during retrieval. The information exists in the chunks, but the wrong chunks are returned. The retrieval pipeline, not the storage format, is the bottleneck. Construction loss is only 1%.

Why this distinction matters: If your agent's memory is failing after a model upgrade, the fix depends on which format you use. For NOTES: re-derive the summaries from raw history using the new model. For RAG: re-embed the chunks. Same symptom (wrong answers), completely different intervention. Without this decomposition, teams guess.

Finding 4: Raw History Is Your Insurance Policy

Can you repair broken memory after a migration? The paper tests two repair strategies: store-only (work with what you have) and raw-retained (re-derive from original transcripts).

Raw-Retained Repair

  • NOTES (Qwen repairing Llama): 34/48 cases hit 90% target
  • RAG re-embedding: 48/48 at 90-99% targets
  • KG-fixed rebuild: 48/48 at 90%, 45-46/48 at 99%
  • Median cost for NOTES repair: $0.76
  • Median cost for RAG re-embed: $0.013

Store-Only Repair

  • NOTES repair success at 90% target: 0 out of 48
  • No raw history means no re-derivation possible
  • Compressed summaries cannot be uncompressed
  • Information lost at construction is gone permanently

The numbers are unambiguous. Without raw transcripts, NOTES repair succeeds exactly zero times out of 48 attempts. With raw transcripts, 34 out of 48 cases reach 90% recovery at a median cost of $0.76. The raw history is not just archival, it is operational infrastructure.

The $0.76 insurance policy: Keeping original conversation transcripts costs storage. Reprocessing them through a new model costs compute. But compared to the alternative (permanent accuracy loss with no recovery path), $0.76 per history repair is remarkably cheap. KG-fixed repairs are essentially free because the schema is model-independent.

Finding 5: Direction Matters

Perhaps the most counterintuitive result is that migration is not symmetric. Swapping Model A for Model B produces different outcomes than swapping Model B for Model A.

Llama reading Qwen's NOTES gains 9.91 percentage points. Qwen reading Llama's NOTES loses 13.28 points. These are not minor variations. They are directionally opposite effects that would cancel to a misleading average of -1.69 if you did not measure them separately.

This has direct implications for upgrade testing. If you validate a migration from GPT-4 to GPT-5, that validation does not tell you what happens when you roll back from GPT-5 to GPT-4. Each direction must be tested independently.

What This Means for Practitioners

Keep Raw Histories Use Fixed Schemas Test Both Directions Avoid Mixed Embeddings Diagnose the Stage

The paper distills into five operational rules for anyone building agent memory systems:

  • Keep raw conversation transcripts. Compressed memories are disposable derivatives, not the source of truth. Raw history enables recovery. Without it, you have no rollback path.
  • Prefer fixed-schema memory formats. Knowledge graphs with rigid schemas transfer across models with near-zero accuracy loss. The schema is the contract. Free-form notes are the most fragile format tested.
  • Test each migration direction separately. Averaging bidirectional results conceals real failures. Validate A-to-B and B-to-A independently.
  • Commit to full re-embedding or accept the cost. Mixed embedding indices (old + new) recover less than half the potential improvement. There is no cheap middle ground that preserves accuracy.
  • Diagnose before fixing. NOTES failures are construction problems (fix by re-deriving). RAG failures are retrieval problems (fix by re-embedding). Same symptom, different treatment.

My Take

This paper fills a gap that practitioners know exists but rarely measure. Everyone who has upgraded a model in a production agent system has encountered unexplained behavior changes. This study puts numbers on the problem and reveals a clear hierarchy: fixed schemas are safe, RAG is fixable, and NOTES are fragile.

The strongest contribution is the diagnostic decomposition. Knowing that NOTES lose information at construction (80%) while RAG loses it at retrieval (81%) transforms the debugging process from guesswork into targeted intervention. This is the kind of result that changes how you build systems, not just how you think about them.

The limitation is scope: two similarly-sized open-weight models, synthetic histories, and single-stage dense retrieval. Real production systems involve larger models, multi-stage retrieval, and subjective conversations. The directional asymmetry finding suggests that scaling up will reveal more surprises, not fewer.

For anyone building agent systems today, the practical message is clear: treat memory as infrastructure, not an afterthought. Model upgrades are inevitable. Memory migration plans should be part of the architecture from day one, not retrofitted after the first silent failure.

Discussion Questions

Open questions for the weekly discussion. No clean answers here.

1. The schema trade-off: Fixed-schema knowledge graphs win on portability but lose on expressiveness. Free-form notes capture nuance that rigid fields cannot. Is there a middle ground, perhaps a schema with a structured core and a free-text overflow field, that preserves portability without sacrificing expressiveness? Or does any free-text component reintroduce the fragility?
2. The embedding migration problem at scale: Full re-embedding is the right answer for accuracy, but at scale (millions of vectors), it can cost thousands of dollars and take hours. Is there a progressive re-embedding strategy that works, perhaps re-embedding the most-queried chunks first? Or does the paper's finding that 50/50 mixing underperforms mean any partial migration is fundamentally flawed?
3. The construction loss paradox: NOTES lose 80% of their information during construction, yet they are the most popular memory format in agent frameworks because they are cheap and fast. If the information loss is that severe, why do NOTES-based agents seem to work well enough in practice? Is it because users do not notice the errors, or because production workloads are simpler than this benchmark?
4. Larger model gap: The study uses 7-8B parameter models. Would the directional asymmetry shrink with larger models (70B+, frontier-class) that have more consistent internal representations? Or would it get worse because larger models develop more distinctive "writing styles" that smaller models cannot parse?
5. The memory versioning question: Should agent memory systems implement explicit version metadata, tagging each memory record with the model that created it? This would enable selective re-derivation during upgrades. But it adds complexity and assumes you know which records are affected. Is model-aware memory management practical, or is "re-derive everything from raw history" the only reliable strategy?

Memory is not a feature. It is infrastructure. The models will keep changing. The question is whether your memory system was built to survive that.

Read the Full Paper →

Weekly live discussion

Join the research breakdown on Zoom

Each article ships with a live session - deeper Q&A, practitioner takeaways, and how the ideas connect to production agent systems.

Reserve a spot