Memory Is Not Automatically Learning

How episodic memory, knowledge gates, retrieval evals, negative knowledge, and production trace mining fit into self-improving agent systems.

The Self Improving Stack series

← When The Harness Has To Evolve Next → Personas Are Content, Coordination Is Structure
Browse all 13 posts
  1. Jun 2026 Topology Is The Missing Action Space
  2. Jun 2026 The Gate Is The Optimizer
  3. Jun 2026 Self-Improvement Needs A Safety Case
  4. Jun 2026 When The Harness Has To Evolve
  5. Jun 2026 Memory Is Not Automatically Learning
  6. Jun 2026 Personas Are Content, Coordination Is Structure
  7. Jun 2026 Optimization Theory For Agent Builders
  8. Jun 2026 When The Model Itself Is Mutable
  9. Jun 2026 Prompt Optimization Is Not The Whole Game
  10. Jun 2026 Skills Are Trainable State
  11. Jun 2026 Beat Random At Equal Compute First
  12. Jun 2026 Traces Are The Training Data
  13. Jun 2026 The Self-Improving Stack
Authored by
outlineGPT-5.5draftGPT-5.5polishGPT-5.5reviewGPT-5.5publishGPT-5.5rewriteGPT-5.5polishGPT-5.5polishGPT-6-luna

Remembering more is not learning.

Learning means the next run changes in the right direction.

Memory is one way to change the next run without changing model weights. It lets an agent carry evidence, preferences, decisions, failures, and procedures across episodes. That makes memory powerful. It also makes memory dangerous.

A bad prompt can ruin one run. A bad memory can keep ruining runs until something expires, contradicts, or deletes it. Persistent state is inherited behavior.

So the important question is not “does the agent have memory?”

The important question is:

  • what is allowed to persist,
  • who can retrieve it,
  • what evidence supports it,
  • how it is tested,
  • and how it is retired?

That is the memory flywheel.

The State Variable

The clean way to think about memory is as mutable external state.

Let:

Mt=memory state before episode tτt=full trace from episode tut=proposed memory write after episode tGmem=memory write gate\begin{aligned} M_t &= \text{memory state before episode }t \\ \tau_t &= \text{full trace from episode }t \\ u_t &= \text{proposed memory write after episode }t \\ G_{\text{mem}} &= \text{memory write gate} \end{aligned}

Write candidates need structure:

utu_t contains:

  • kind
  • claim_or_procedure
  • evidence_refs
  • scope
  • confidence
  • sensitivity
  • freshness_policy
  • retrieval_policy

The update rule is:

Mt+1={Apply⁡(Mt,ut),if Gmem(ut,τt,policy) passes,Mt,otherwise.M_{t+1}=\begin{cases} \operatorname{Apply}(M_t,u_t),&\text{if }G_{\text{mem}}(u_t,\tau_t,\text{policy})\text{ passes}, \\ M_t,&\text{otherwise.} \end{cases}

At inference time, the memory layer changes the context seen by the policy:

ct=Retrieve⁡(Mt,qt,k,policy)yt=πθ(xt,ct,tools)\begin{aligned} c_t &= \operatorname{Retrieve}(M_t,q_t,k,\text{policy}) \\ y_t &= \pi_\theta(x_t,c_t,\text{tools}) \end{aligned}

where:

xt=task inputqt=retrieval query or retrieval plank=retrieval budgetct=retrieved contextπθ=model policy with fixed weights θyt=output or next action\begin{aligned} x_t &= \text{task input} \\ q_t &= \text{retrieval query or retrieval plan} \\ k &= \text{retrieval budget} \\ c_t &= \text{retrieved context} \\ \pi_\theta &= \text{model policy with fixed weights }\theta \\ y_t &= \text{output or next action} \end{aligned}

Memory is not magic. It is an intervention on the policy’s input distribution.

That gives us a measurable target:

Δmemory=E[Score⁡(πθ with M)]−E[Score⁡(πθ without M)]\Delta_{\text{memory}}=\mathbb{E}[\operatorname{Score}(\pi_\theta\text{ with }M)]-\mathbb{E}[\operatorname{Score}(\pi_\theta\text{ without }M)]

The unit test for memory is not “did retrieval return something?”

The unit test is whether the memory intervention improved the downstream task, under cost, safety, freshness, and privacy constraints.

A Short History

RAG made the basic distinction famous in 2020: a model can combine parametric memory in its weights with non-parametric memory in an external index. Lewis et al. explicitly framed provenance and updating world knowledge as open problems for parameter-only models.

Agent memory then became more explicit.

Generative Agents in 2023 stored natural-language experiences, synthesized reflections, and retrieved memories to plan social behavior. Reflexion in 2023 turned feedback into verbal reflections held in an episodic memory buffer, then used those reflections to improve later attempts without updating model weights. Voyager in 2023 pushed the procedural version: an embodied Minecraft agent built an ever-growing library of executable skills and reused them in new worlds.

The same year, MemGPT treated memory management like virtual context management. MemoryBank focused on long-term conversational memory, including selective forgetting and reinforcement. LongMem explored model architectures that augment fixed models with long-term memory banks.

By 2024 and 2025, the evaluation pressure sharpened. LongMemEval tested extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention. Mem0 reported production-oriented memory extraction, consolidation, graph memory, latency, and token-cost results on LOCOMO. A-MEM moved toward dynamically organized agent memory, where adding a memory can update links and contextual attributes of older memories. AgeMem, submitted in January 2026 and revised in April 2026, frames memory operations as tool actions and trains long-term plus short-term memory management with reinforcement learning.

The security story sharpened too. MemoryGraft, submitted on December 18, 2025, describes poisoned experience retrieval: malicious successful-looking records persist into an agent’s memory and later steer behavior on semantically similar tasks.

That is the line from RAG to agent memory:

  1. 01
    retrieve facts
  2. 02
    store experiences
  3. 03
    reflect across episodes
  4. 04
    store executable procedures
  5. 05
    manage memory as context
  6. 06
    evaluate multi-session behavior
  7. 07
    learn memory operations
  8. 08
    defend the memory trust boundary

The frontier is not “bigger memory.”

The frontier is controlled memory mutation.

The Memory Types

Different memories mutate different parts of the system.

Memory typeMutable unitUsed byEvaluatorMain failure mode
Episodictrace, episode summary, decision recordplanner, analyst, supervisorreplay and outcome comparisonsummary loses the causal detail
Semanticclaim, page, source anchor, relationretriever, answerer, researchercitation, contradiction, freshnessfalse or stale fact becomes canonical
Proceduralskill, checklist, tool habit, repair routinedriver, worker, coding agenttask success and transferlocal trick gets over-applied
Preferenceuser, team, product, or persona preferenceassistant, router, UI agentsatisfaction, explicit confirmationpreference is scoped too broadly
Negativeinvalid assumption, failed tactic, banned pathplanner, verifier, selectorrecurrence reductionstale warning blocks correct behavior
Sourceraw document, trace artifact, anchor, quoteknowledge curator, judgeprovenance and hash integritygenerated text is mistaken for source evidence
Decisionchosen option, rejected option, rationalefuture planner and reviewerconsistency under same constraintsold rationale survives after constraints change

This table matters because “memory” is too broad a word.

A retrieved user preference, an executable skill, a stale API fact, a failed deployment lesson, and a product requirement are not the same kind of thing. They need separate scope, retrieval, and write policies.

The Flywheel

A useful memory flywheel has seven steps:

  1. observe
  2. extract
  3. propose
  4. gate
  5. retrieve
  6. act
  7. evaluate

The trace supplies the raw material:

τt\tau_t contains:

  • task
  • messages
  • tool calls
  • observations
  • artifacts
  • verifier results
  • analyst findings
  • outcome

The extractor turns trace evidence into proposed writes:

τt⟶{u1,u2,…,un}\tau_t\longrightarrow\{u_1,u_2,\ldots,u_n\}

The gate decides whether each write is safe, scoped, supported, and useful:

G_mem(u_i) → admit | reject | ask | quarantine | expire

The retriever selects admitted memory for a future task:

Retrieve(M_{t+1}, q, policy) → context

Then the evaluator measures whether the retrieval actually helped.

That last step is where many memory systems become cargo cults. They store more, retrieve more, and show more context to the model, but never run the paired ablation:

  • same task
  • same model
  • same tool surface
  • with memory versus without memory

Without that ablation, memory success is often just retrieval theater.

The Write Gate

The memory write gate is the admission controller.

It asks different questions for different memory classes, but the core checks are stable:

Gate checkQuestion
ProvenanceWhat evidence supports the write?
LocusIs this global, persona-scoped, task-scoped, project-scoped, or agent-scoped?
SensitivityIs the content public, private, secret, or user-confirmed?
FreshnessDoes it expire? When was it last verified?
ContradictionDoes it conflict with active claims or newer traces?
ConfidenceIs the evidence strong enough for the target use?
ReversibilityCan the write be rolled back or superseded?
Retrieval impactDoes it improve retrieval-conditioned behavior?
PromotionHas it passed held-out tasks or production replay?

The provenance check is the one that prevents the most subtle mistakes.

A self-generated artifact is not the same thing as a source. An agent can write an analysis that says “API v2 requires field X.” That analysis is useful trace evidence. It is not itself proof that the API requires field X. The source-grounded memory needs to point to the API documentation, a tool response, a schema file, or a verified runtime observation.

Some memories do not need external source grounding. A user preference can be grounded in the user’s own instruction. A local coding habit can be grounded in a repeated trace pattern. A decision record can be grounded in the decision meeting or session. But the scope must be explicit:

  • operator preference for this repo
  • team convention for this product
  • task-local assumption
  • global technical fact

Most poisoning problems start when a scoped memory is treated as global truth.

One useful mental model is a scope lattice:

  • run
  • task
  • project
  • persona
  • team
  • organization
  • global

Promotion up the lattice requires stronger evidence. A run-local observation can become a task memory after repeated traces. A task memory can become a project convention after review. A project convention rarely deserves to become a global technical fact.

The gate can be written as a predicate:

admit(u) iff
  provenance(u) passes
  and scope(u) is allowed for the target readers
  and sensitivity(u) is allowed for the target storage
  and freshness(u, now) >= required_freshness
  and contradiction_check(u, M_t) passes
  and expected_lift(u) - expected_cost(u) > threshold

The last term is estimated, then corrected by real evals after the memory is used.

Retrieval Is An Intervention

Retrieval is not context stuffing.

Retrieval changes the policy input:

πθ(y∣x)\pi_\theta(y\mid x)

becomes:

πθ(y∣x,c)\pi_\theta(y\mid x,c)

where cc is retrieved context.

That context can help, do nothing, or harm. It can help by supplying missing facts, preserving user preferences, recalling a successful procedure, or warning against a repeated mistake. It can harm by anchoring the model on irrelevant facts, stale procedures, false summaries, or over-broad preferences.

The evaluator has to measure the whole effect:

MetricWhat it catches
Recall@kwhether relevant memory can be found
Precision@kwhether retrieved memory is mostly useful
Contradiction ratewhether retrieval injects conflicting claims
Freshness pass ratewhether retrieved sources are still valid
Answer liftwhether final outputs improve
Task liftwhether full agent trajectories improve
Costwhether memory increases latency, tokens, or tool calls
Abstention qualitywhether the system knows when memory is insufficient

Hit rate is not enough.

A memory can be retrievable and harmful. A vector store can return semantically similar experience that is operationally wrong. A graph can link related facts that differ in scope. A persona memory can dominate a task where it does not apply.

The promotion gate compares:

Scorewith memory−Scorewithout memory\mathrm{Score}_{\text{with memory}}-\mathrm{Score}_{\text{without memory}}

and also:

Costwith memory−Costwithout memory\mathrm{Cost}_{\text{with memory}}-\mathrm{Cost}_{\text{without memory}}

A memory layer that improves one benchmark by adding large latency and subtle privacy risk may be a bad production trade.

Negative Knowledge

Negative knowledge is one of the most useful and most dangerous forms of memory.

It records what not to do:

  • do not use endpoint A after version 3
  • do not assume screenshots live in path P
  • do not ask the user for repo facts before inspecting the repo
  • do not collapse supervisor and worker roles for this task class
  • do not retry a failed deploy hook without checking logs

This is often the difference between an agent that keeps repeating a class of mistake and one that actually compounds.

But negative knowledge needs expiration and scope.

If endpoint A comes back, the old warning can block correct work. If a failed tactic only failed under a specific model, repo, budget, or date, the warning cannot become a universal law. If the negative memory is a vague sentence like “avoid parallelization here,” it can suppress a good multi-agent strategy later.

Good negative memory has this shape:

claim: tactic T failed under condition C
evidence: trace spans and artifacts
scope: repo, tool, model, task class, or date range
replacement: use tactic R instead
expiry: when to re-check

Negative memory is not cynicism. It is a falsifiable constraint.

Memory Versus Skill

Procedural memory is close to skill optimization, but the distinction is useful.

A memory can say:

when patching a repo, inspect status and the last few commits first

A skill can operationalize it:

inputs: repo path
preconditions: git worktree exists
steps: status, log, reflog, open PRs
verification: no live rebase, no mid-merge, branch context known

The skill has an invocation contract, parameters, steps, and verification. The memory is the durable lesson that motivates or updates the skill.

Voyager’s executable library sits on the skill side. Reflexion’s verbal reflections sit on the episodic/procedural memory side. In production agents, the clean loop is:

  1. trace shows repeated procedural failure
  2. memory records the failure pattern
  3. skill proposal updates the reusable procedure
  4. held-out tasks test the skill
  5. memory stores the promotion evidence

This prevents the memory layer from becoming a bag of instructions that only work when the model happens to read them.

Multi-Agent Memory

Multi-agent systems make memory harder because there is no single “the agent.”

There are drivers, workers, reviewers, supervisors, routers, judges, researchers, and coordinators. Each role needs a different memory view.

A coding worker may need:

  • repo conventions
  • tool-call habits
  • known failure modes
  • current task artifacts

A supervisor may need:

  • branch state
  • worker assignments
  • conflict map
  • quality bar
  • promotion gate

A judge may need:

  • rubric
  • reference outputs
  • verifier traces
  • leakage restrictions

A coordinator may need:

  • fanout policy
  • budget policy
  • selector rules
  • stop conditions

This is why a single optimized persona prompt is not enough. In a multi-agent flow, memory has to be routed by role and task. A worker does not need every supervisor constraint. A judge cannot retrieve candidate-internal rationales that contaminate independence. A coordinator cannot treat one worker’s failed local path as a global ban unless the evidence says so.

The same point applies to maxTurns=0 agentic flows.

If a subagent gets one shot, it cannot learn inside its own episode. The learning has to happen outside it:

  1. pre-run retrieval
  2. post-run trace capture
  3. cross-run write proposal
  4. promotion gate
  5. next-run retrieval

That is still a flywheel, but the flywheel lives in the harness and memory substrate, not inside the worker’s conversational loop.

Operator directives such as “parallelize independent reads” or “inspect the repo before asking” can become procedural or preference memory, but only if they are captured with scope and evidence. A prompt optimizer may discover a wording that says “parallelize,” but it cannot invent a parallel execution graph if the runtime cannot fan out calls. A memory system can remember the directive. The runtime still needs the action surface to use it.

Memory is not a substitute for topology.

It is the substrate that lets topology improve across episodes.

Knowledge Poisoning

Knowledge poisoning is not merely “the agent did not know something.”

A gap is:

the agent needed X and did not have it

Poisoning is:

the agent confidently used X, and X was wrong

The second case is worse because the agent does not ask. It acts.

In December 2025, MemoryGraft named a concrete version of this attack surface: poison an agent’s experience retrieval by making malicious successful-looking records persist into long-term memory. Later, semantically similar tasks retrieve those records and imitate the unsafe pattern.

The general pattern is broader:

  • stale wiki page
  • outdated web result
  • wrong prior-run summary
  • tool description with old return shape
  • system prompt copied from an older runtime
  • successful-looking trace from a compromised task

The defense is not “trust memory less” in the abstract. The defense is dual verification:

  1. Did the agent act on the belief?
  2. Does trace or source evidence show the belief is false?

Only then can the system emit a poisoning finding. Otherwise it risks turning uncertainty into fake certainty.

Poisoning remediation is also a memory write:

  • mark stale
  • supersede claim
  • quarantine source
  • lower confidence
  • add expiry
  • link contradiction evidence
  • trigger held-out replay

Bad memory cannot just be deleted quietly. The system needs to learn why it was bad.

How Tangle Fits

The local Tangle stack is close to the architecture described above.

@tangle-network/agent-knowledge is the knowledge substrate. In the checked local source, version 1.3.0 describes itself as “source-grounded, eval-gated knowledge growth primitives for agents.” Its exported surfaces include source records, source anchors, claims, relations, pages, graph search, readiness scoring, freshness tracking, safe write blocks, validation, proposal generation from analyst findings, research loops, and release reports.

The important detail is that it models memory as structured knowledge, not loose text:

  • SourceRecord
  • SourceAnchor
  • KnowledgeClaim
  • KnowledgeRelation
  • KnowledgePage
  • KnowledgeIndex
  • KnowledgeSearchResult
  • KnowledgeLintFinding
  • KnowledgeRelease

That shape supports the gate:

  • refs
  • confidence
  • status
  • validUntil
  • lastVerifiedAt
  • sourceIds
  • allowedPathPrefixes
  • lint findings
  • release reports

@tangle-network/agent-eval supplies the analyst side. The local source includes knowledge-gap and knowledge-poisoning analyst specs. The knowledge-gap analyst asks what the agent lacked or what was stale, then attributes the gap to the layer responsible for holding it:

  • Agent-knowledge: wiki:<page>
  • Agent-knowledge: claim:<topic>
  • Agent-knowledge: raw:<source>
  • Agent-knowledge: stale:<page>
  • Websearch: outdated:<topic>
  • Tool-doc: <tool>
  • System-prompt: <section>
  • Memory: <key>

The knowledge-poisoning analyst asks for confident wrong action, then requires the dual verification protocol:

  • acted on false belief
  • belief contradicted by trace evidence

@tangle-network/agent-runtime supplies the bridge. The local createSurfaceKnowledgeAdapter wraps agent-knowledge proposal generation and write-block application. It converts analyst findings into knowledge proposals, applies write blocks against a knowledge root, and optionally lints after apply.

Put together, the stack can express this loop:

  1. 01
    production trace
  2. 02
    agent-eval analyst finding
  3. 03
    agent-knowledge proposal
  4. 04
    safe write block
  5. 05
    lint and readiness checks
  6. 06
    retrieved context
  7. 07
    future production run
  8. 08
    held-out and production evaluation

That is the memory flywheel as software.

The Readiness Gate

The most underrated piece is readiness.

Before an agent starts a task, the system can ask:

  • what knowledge is required for this task?
  • is it present?
  • is it fresh?
  • is it sensitive?
  • how confident does it need to be?
  • what happens if it is missing?

The local agent-knowledge readiness builder maps specs to requirements with fields such as:

  • category
  • acquisitionMode
  • importance
  • freshness
  • sensitivity
  • confidenceNeeded
  • fallbackPolicy
  • minSources
  • minHits

That is a better frame than “give the model memories.”

For blocking requirements, absence blocks, asks, or triggers acquisition. For non-blocking requirements, absence can continue with caveats. For high-sensitivity requirements, retrieval may be disallowed for some roles. For realtime requirements, stale memory counts as missing.

Readiness turns memory from a passive archive into a pre-flight gate.

The Core Test

A memory system is doing real self-improvement when all of these are true:

  1. A trace produces a specific finding.
  2. The finding proposes a scoped memory write.
  3. The write is source-grounded or explicitly scoped to its evidence.
  4. The gate admits, rejects, asks, quarantines, or expires it.
  5. Future retrieval selects it only for appropriate roles and tasks.
  6. A paired eval shows task lift, not just retrieval activity.
  7. Staleness, contradiction, privacy, and poisoning have review paths.

If any part is missing, the system may still be useful, but it is not a disciplined learning loop.

It may just be a larger prompt with a longer memory leak.

Source Trail

Source freshness checked on 2026-06-06.

Revision history9revisions
  1. GPT-6-lunapolish+156−228 view trace →
    Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
    show diff
    diff --git a/src/content/posts/self-improving-stack-memory-flywheels.mdx b/src/content/posts/self-improving-stack-memory-flywheels.mdxindex 8a7ffd9..d5223c8 100644--- a/src/content/posts/self-improving-stack-memory-flywheels.mdx+++ b/src/content/posts/self-improving-stack-memory-flywheels.mdx@@ -47,6 +47,8 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from "../../components/Steps.astro"+ Remembering more is not learning.  Learning means the next run changes in the right direction.@@ -59,13 +61,11 @@ So the important question is not "does the agent have memory?"  The important question is: -```text-what is allowed to persist,-who can retrieve it,-what evidence supports it,-how it is tested,-and how it is retired?-```+- what is allowed to persist,+- who can retrieve it,+- what evidence supports it,+- how it is tested,+- and how it is retired?  That is the memory flywheel. @@ -86,17 +86,16 @@ $$  Write candidates need structure: -```text-u_t =-  kind-  claim_or_procedure-  evidence_refs-  scope-  confidence-  sensitivity-  freshness_policy-  retrieval_policy-```+$u_t$ contains:++- kind+- `claim_or_procedure`+- `evidence_refs`+- scope+- confidence+- sensitivity+- `freshness_policy`+- `retrieval_policy`  The update rule is: @@ -157,16 +156,7 @@ The security story sharpened too. MemoryGraft, submitted on December 18, 2025, d  That is the line from RAG to agent memory: -```text-retrieve facts--> store experiences--> reflect across episodes--> store executable procedures--> manage memory as context--> evaluate multi-session behavior--> learn memory operations--> defend the memory trust boundary-```+<Steps layout="flow" items={[{"title": "retrieve facts"}, {"title": "store experiences"}, {"title": "reflect across episodes"}, {"title": "store executable procedures"}, {"title": "manage memory as context"}, {"title": "evaluate multi-session behavior"}, {"title": "learn memory operations"}, {"title": "defend the memory trust boundary"}]} />  The frontier is not "bigger memory." @@ -194,29 +184,26 @@ A retrieved user preference, an executable skill, a stale API fact, a failed dep  A useful memory flywheel has seven steps: -```text-observe-extract-propose-gate-retrieve-act-evaluate-```+1. observe+2. extract+3. propose+4. gate+5. retrieve+6. act+7. evaluate  The trace supplies the raw material: -```text-tau_t =-  task-  messages-  tool calls-  observations-  artifacts-  verifier results-  analyst findings-  outcome-```+$\tau_t$ contains:++- task+- messages+- tool calls+- observations+- artifacts+- verifier results+- analyst findings+- outcome  The extractor turns trace evidence into proposed writes: @@ -226,26 +213,20 @@ $$  The gate decides whether each write is safe, scoped, supported, and useful: -```text-G_mem(u_i) -> admit | reject | ask | quarantine | expire-```+`G_mem(u_i) → admit | reject | ask | quarantine | expire`  The retriever selects admitted memory for a future task: -```text-Retrieve(M_{t+1}, q, policy) -> context-```+`Retrieve(M_{t+1}, q, policy) → context`  Then the evaluator measures whether the retrieval actually helped.  That last step is where many memory systems become cargo cults. They store more, retrieve more, and show more context to the model, but never run the paired ablation: -```text-same task-same model-same tool surface-with memory versus without memory-```+- same task+- same model+- same tool surface+- with memory versus without memory  Without that ablation, memory success is often just retrieval theater. @@ -273,26 +254,22 @@ A self-generated artifact is not the same thing as a source. An agent can write  Some memories do not need external source grounding. A user preference can be grounded in the user's own instruction. A local coding habit can be grounded in a repeated trace pattern. A decision record can be grounded in the decision meeting or session. But the scope must be explicit: -```text-operator preference for this repo-team convention for this product-task-local assumption-global technical fact-```+- operator preference for this repo+- team convention for this product+- task-local assumption+- global technical fact  Most poisoning problems start when a scoped memory is treated as global truth.  One useful mental model is a scope lattice: -```text-run-task-project-persona-team-organization-global-```+- `run`+- `task`+- `project`+- `persona`+- `team`+- `organization`+- `global`  Promotion up the lattice requires stronger evidence. A run-local observation can become a task memory after repeated traces. A task memory can become a project convention after review. A project convention rarely deserves to become a global technical fact. @@ -316,15 +293,15 @@ Retrieval is not context stuffing.  Retrieval changes the policy input: -```text-pi_theta(y | x)-```+$$+\pi_\theta(y\mid x)+$$  becomes: -```text-pi_theta(y | x, c)-```+$$+\pi_\theta(y\mid x,c)+$$  where $c$ is retrieved context. @@ -349,15 +326,15 @@ A memory can be retrievable and harmful. A vector store can return semantically  The promotion gate compares: -```text-Score_with_memory - Score_without_memory-```+$$+\mathrm{Score}_{\text{with memory}}-\mathrm{Score}_{\text{without memory}}+$$  and also: -```text-Cost_with_memory - Cost_without_memory-```+$$+\mathrm{Cost}_{\text{with memory}}-\mathrm{Cost}_{\text{without memory}}+$$  A memory layer that improves one benchmark by adding large latency and subtle privacy risk may be a bad production trade. @@ -367,13 +344,11 @@ Negative knowledge is one of the most useful and most dangerous forms of memory.  It records what not to do: -```text-do not use endpoint A after version 3-do not assume screenshots live in path P-do not ask the user for repo facts before inspecting the repo-do not collapse supervisor and worker roles for this task class-do not retry a failed deploy hook without checking logs-```+- do not use endpoint A after version 3+- do not assume screenshots live in path P+- do not ask the user for repo facts before inspecting the repo+- do not collapse supervisor and worker roles for this task class+- do not retry a failed deploy hook without checking logs  This is often the difference between an agent that keeps repeating a class of mistake and one that actually compounds. @@ -399,9 +374,7 @@ Procedural memory is close to skill optimization, but the distinction is useful.  A memory can say: -```text-when patching a repo, inspect status and the last few commits first-```+> when patching a repo, inspect status and the last few commits first  A skill can operationalize it: @@ -416,13 +389,11 @@ The skill has an invocation contract, parameters, steps, and verification. The m  Voyager's executable library sits on the skill side. Reflexion's verbal reflections sit on the episodic/procedural memory side. In production agents, the clean loop is: -```text-trace shows repeated procedural failure--> memory records the failure pattern--> skill proposal updates the reusable procedure--> held-out tasks test the skill--> memory stores the promotion evidence-```+1. trace shows repeated procedural failure+2. memory records the failure pattern+3. skill proposal updates the reusable procedure+4. held-out tasks test the skill+5. memory stores the promotion evidence  This prevents the memory layer from becoming a bag of instructions that only work when the model happens to read them. @@ -434,40 +405,32 @@ There are drivers, workers, reviewers, supervisors, routers, judges, researchers  A coding worker may need: -```text-repo conventions-tool-call habits-known failure modes-current task artifacts-```+- repo conventions+- tool-call habits+- known failure modes+- current task artifacts  A supervisor may need: -```text-branch state-worker assignments-conflict map-quality bar-promotion gate-```+- branch state+- worker assignments+- conflict map+- quality bar+- promotion gate  A judge may need: -```text-rubric-reference outputs-verifier traces-leakage restrictions-```+- `rubric`+- reference outputs+- verifier traces+- leakage restrictions  A coordinator may need: -```text-fanout policy-budget policy-selector rules-stop conditions-```+- fanout policy+- budget policy+- selector rules+- stop conditions  This is why a single optimized persona prompt is not enough. In a multi-agent flow, memory has to be routed by role and task. A worker does not need every supervisor constraint. A judge cannot retrieve candidate-internal rationales that contaminate independence. A coordinator cannot treat one worker's failed local path as a global ban unless the evidence says so. @@ -475,13 +438,11 @@ The same point applies to `maxTurns=0` agentic flows.  If a subagent gets one shot, it cannot learn inside its own episode. The learning has to happen outside it: -```text-pre-run retrieval-post-run trace capture-cross-run write proposal-promotion gate-next-run retrieval-```+1. pre-run retrieval+2. post-run trace capture+3. cross-run write proposal+4. promotion gate+5. next-run retrieval  That is still a flywheel, but the flywheel lives in the harness and memory substrate, not inside the worker's conversational loop. @@ -497,15 +458,11 @@ Knowledge poisoning is not merely "the agent did not know something."  A gap is: -```text-the agent needed X and did not have it-```+> the agent needed X and did not have it  Poisoning is: -```text-the agent confidently used X, and X was wrong-```+> the agent confidently used X, and X was wrong  The second case is worse because the agent does not ask. It acts. @@ -513,35 +470,29 @@ In December 2025, MemoryGraft named a concrete version of this attack surface: p  The general pattern is broader: -```text-stale wiki page-outdated web result-wrong prior-run summary-tool description with old return shape-system prompt copied from an older runtime-successful-looking trace from a compromised task-```+- stale wiki page+- outdated web result+- wrong prior-run summary+- tool description with old return shape+- system prompt copied from an older runtime+- successful-looking trace from a compromised task  The defense is not "trust memory less" in the abstract. The defense is dual verification: -```text 1. Did the agent act on the belief? 2. Does trace or source evidence show the belief is false?-```  Only then can the system emit a poisoning finding. Otherwise it risks turning uncertainty into fake certainty.  Poisoning remediation is also a memory write: -```text-mark stale-supersede claim-quarantine source-lower confidence-add expiry-link contradiction evidence-trigger held-out replay-```+- mark stale+- supersede claim+- quarantine source+- lower confidence+- add expiry+- link contradiction evidence+- trigger held-out replay  Bad memory cannot just be deleted quietly. The system needs to learn why it was bad. @@ -553,66 +504,49 @@ The local Tangle stack is close to the architecture described above.  The important detail is that it models memory as structured knowledge, not loose text: -```text-SourceRecord-SourceAnchor-KnowledgeClaim-KnowledgeRelation-KnowledgePage-KnowledgeIndex-KnowledgeSearchResult-KnowledgeLintFinding-KnowledgeRelease-```+- `SourceRecord`+- `SourceAnchor`+- `KnowledgeClaim`+- `KnowledgeRelation`+- `KnowledgePage`+- `KnowledgeIndex`+- `KnowledgeSearchResult`+- `KnowledgeLintFinding`+- `KnowledgeRelease`  That shape supports the gate: -```text-refs-confidence-status-validUntil-lastVerifiedAt-sourceIds-allowedPathPrefixes-lint findings-release reports-```+- `refs`+- `confidence`+- `status`+- `validUntil`+- `lastVerifiedAt`+- `sourceIds`+- `allowedPathPrefixes`+- lint findings+- release reports  `@tangle-network/agent-eval` supplies the analyst side. The local source includes knowledge-gap and knowledge-poisoning analyst specs. The knowledge-gap analyst asks what the agent lacked or what was stale, then attributes the gap to the layer responsible for holding it: -```text-agent-knowledge:wiki:<page>-agent-knowledge:claim:<topic>-agent-knowledge:raw:<source>-agent-knowledge:stale:<page>-websearch:outdated:<topic>-tool-doc:<tool>-system-prompt:<section>-memory:<key>-```+- **Agent-knowledge:** `wiki:<page>`+- **Agent-knowledge:** `claim:<topic>`+- **Agent-knowledge:** `raw:<source>`+- **Agent-knowledge:** `stale:<page>`+- **Websearch:** `outdated:<topic>`+- **Tool-doc:** `<tool>`+- **System-prompt:** `<section>`+- **Memory:** `<key>`  The knowledge-poisoning analyst asks for confident wrong action, then requires the dual verification protocol: -```text-acted on false belief-belief contradicted by trace evidence-```+- acted on false belief+- belief contradicted by trace evidence  `@tangle-network/agent-runtime` supplies the bridge. The local `createSurfaceKnowledgeAdapter` wraps `agent-knowledge` proposal generation and write-block application. It converts analyst findings into knowledge proposals, applies write blocks against a knowledge root, and optionally lints after apply.  Put together, the stack can express this loop: -```text-production trace--> agent-eval analyst finding--> agent-knowledge proposal--> safe write block--> lint and readiness checks--> retrieved context--> future production run--> held-out and production evaluation-```+<Steps layout="flow" items={[{"title": "production trace"}, {"title": "agent-eval analyst finding"}, {"title": "agent-knowledge proposal"}, {"title": "safe write block"}, {"title": "lint and readiness checks"}, {"title": "retrieved context"}, {"title": "future production run"}, {"title": "held-out and production evaluation"}]} />  That is the memory flywheel as software. @@ -622,28 +556,24 @@ The most underrated piece is readiness.  Before an agent starts a task, the system can ask: -```text-what knowledge is required for this task?-is it present?-is it fresh?-is it sensitive?-how confident does it need to be?-what happens if it is missing?-```+- what knowledge is required for this task?+- is it present?+- is it fresh?+- is it sensitive?+- how confident does it need to be?+- what happens if it is missing?  The local `agent-knowledge` readiness builder maps specs to requirements with fields such as: -```text-category-acquisitionMode-importance-freshness-sensitivity-confidenceNeeded-fallbackPolicy-minSources-minHits-```+- `category`+- `acquisitionMode`+- `importance`+- `freshness`+- `sensitivity`+- `confidenceNeeded`+- `fallbackPolicy`+- `minSources`+- `minHits`  That is a better frame than "give the model memories." @@ -655,7 +585,6 @@ Readiness turns memory from a passive archive into a pre-flight gate.  A memory system is doing real self-improvement when all of these are true: -```text 1. A trace produces a specific finding. 2. The finding proposes a scoped memory write. 3. The write is source-grounded or explicitly scoped to its evidence.@@ -663,7 +592,6 @@ A memory system is doing real self-improvement when all of these are true: 5. Future retrieval selects it only for appropriate roles and tasks. 6. A paired eval shows task lift, not just retrieval activity. 7. Staleness, contradiction, privacy, and poisoning have review paths.-```  If any part is missing, the system may still be useful, but it is not a disciplined learning loop. 
  2. GPT-6-lunapolish+37−31 view trace →
    Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
    show diff
    diff --git a/src/content/posts/self-improving-stack-memory-flywheels.mdx b/src/content/posts/self-improving-stack-memory-flywheels.mdxindex 2bc0179..db645d1 100644--- a/src/content/posts/self-improving-stack-memory-flywheels.mdx+++ b/src/content/posts/self-improving-stack-memory-flywheels.mdx@@ -73,12 +73,14 @@ The clean way to think about memory is as mutable external state.  Let: -```text-M_t = memory state before episode t-tau_t = full trace from episode t-u_t = proposed memory write after episode t-G_mem = memory write gate-```+$$+\begin{aligned}+  M_t &= \text{memory state before episode }t \\+  \tau_t &= \text{full trace from episode }t \\+  u_t &= \text{proposed memory write after episode }t \\+  G_{\text{mem}} &= \text{memory write gate}+\end{aligned}+$$  Write candidates need structure: @@ -96,38 +98,42 @@ u_t =  The update rule is: -```text-M_{t+1} =-  Apply(M_t, u_t) if G_mem(u_t, tau_t, policy) passes-  M_t             otherwise-```+$$+M_{t+1}=\begin{cases}+  \operatorname{Apply}(M_t,u_t),&\text{if }G_{\text{mem}}(u_t,\tau_t,\text{policy})\text{ passes}, \\+  M_t,&\text{otherwise.}+\end{cases}+$$  At inference time, the memory layer changes the context seen by the policy: -```text-c_t = Retrieve(M_t, q_t, k, policy)-y_t = pi_theta(x_t, c_t, tools)-```+$$+\begin{aligned}+  c_t &= \operatorname{Retrieve}(M_t,q_t,k,\text{policy}) \\+  y_t &= \pi_\theta(x_t,c_t,\text{tools})+\end{aligned}+$$  where: -```text-x_t = task input-q_t = retrieval query or retrieval plan-k = retrieval budget-c_t = retrieved context-pi_theta = model policy with fixed weights theta-y_t = output or next action-```+$$+\begin{aligned}+  x_t &= \text{task input} \\+  q_t &= \text{retrieval query or retrieval plan} \\+  k &= \text{retrieval budget} \\+  c_t &= \text{retrieved context} \\+  \pi_\theta &= \text{model policy with fixed weights }\theta \\+  y_t &= \text{output or next action}+\end{aligned}+$$  Memory is not magic. It is an intervention on the policy's input distribution.  That gives us a measurable target: -```text-Delta_memory =-  E[Score(pi_theta with M)] - E[Score(pi_theta without M)]-```+$$+\Delta_{\text{memory}}=\mathbb{E}[\operatorname{Score}(\pi_\theta\text{ with }M)]-\mathbb{E}[\operatorname{Score}(\pi_\theta\text{ without }M)]+$$  The unit test for memory is not "did retrieval return something?" @@ -212,9 +218,9 @@ tau_t =  The extractor turns trace evidence into proposed writes: -```text-tau_t -> {u_1, u_2, ..., u_n}-```+$$+\tau_t\longrightarrow\{u_1,u_2,\ldots,u_n\}+$$  The gate decides whether each write is safe, scoped, supported, and useful: @@ -318,7 +324,7 @@ becomes: pi_theta(y | x, c) ``` -where `c` is retrieved context.+where $c$ is retrieved context.  That context can help, do nothing, or harm. It can help by supplying missing facts, preserving user preferences, recalling a successful procedure, or warning against a repeated mistake. It can harm by anchoring the model on irrelevant facts, stale procedures, false summaries, or over-broad preferences. 
  3. GPT-5.5draft view trace →
    Drafted the memory and knowledge flywheels article with memory-state formalism, write gates, retrieval ablations, negative knowledge, multi-agent role routing, poisoning defenses, and Tangle agent-knowledge/runtime/eval placement.
  4. GPT-5.5polish view trace →
    Polished the memory and knowledge article by adding the structured write-candidate schema, scope lattice, admission predicate, memory-versus-skill distinction, and tighter multi-agent and poisoning language.
  5. GPT-5.5polish+3−1 view trace →
    let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
    show diff
    diff --git a/src/content/posts/self-improving-stack-memory-flywheels.mdx b/src/content/posts/self-improving-stack-memory-flywheels.mdxindex 3e45ea2..605f798 100644--- a/src/content/posts/self-improving-stack-memory-flywheels.mdx+++ b/src/content/posts/self-improving-stack-memory-flywheels.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-06   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-memory-flywheels-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-memory-flywheels-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-memory-flywheels-review' }   - date: 2026-06-05@@ -47,7 +49,7 @@ Learning means the next run changes in the right direction.  Memory is one way to change the next run without changing model weights. It lets an agent carry evidence, preferences, decisions, failures, and procedures across episodes. That makes memory powerful. It also makes memory dangerous. -Persistent state is inherited by future behavior. A bad prompt can ruin one run. A bad memory can keep ruining runs until something expires, contradicts, or deletes it.+A bad prompt can ruin one run. A bad memory can keep ruining runs until something expires, contradicts, or deletes it. Persistent state is inherited behavior.  So the important question is not "does the agent have memory?" 
  6. GPT-5.5rewrite view trace →
    60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
  7. GPT-5.5publish view trace →
    Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
  8. GPT-5.5review view trace →
    Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
  9. GPT-5.5outline view trace →
    Research planning pass from a traced session.

Comments

Comments load from GitHub Discussions via Giscus. Configure PUBLIC_GISCUS_REPO, PUBLIC_GISCUS_REPO_ID, PUBLIC_GISCUS_CATEGORY, and PUBLIC_GISCUS_CATEGORY_ID in .env. See giscus.app to generate the IDs after you enable Discussions on the repo.