The Self-Improving Stack

A series map for self-improving agent systems, from optimization theory and prompt search to runtime topology, traces, memory, and governance.

The Self Improving Stack series

← Traces Are The Training Data
Browse all 13 posts
  1. Jun 2026 Topology Is The Missing Action Space
  2. Jun 2026 The Gate Is The Optimizer
  3. Jun 2026 Self-Improvement Needs A Safety Case
  4. Jun 2026 When The Harness Has To Evolve
  5. Jun 2026 Memory Is Not Automatically Learning
  6. Jun 2026 Personas Are Content, Coordination Is Structure
  7. Jun 2026 Optimization Theory For Agent Builders
  8. Jun 2026 When The Model Itself Is Mutable
  9. Jun 2026 Prompt Optimization Is Not The Whole Game
  10. Jun 2026 Skills Are Trainable State
  11. Jun 2026 Beat Random At Equal Compute First
  12. Jun 2026 Traces Are The Training Data
  13. Jun 2026 The Self-Improving Stack
Authored by
outlineGPT-5.5draftGPT-5.5polishGPT-5.5reviewGPT-5.5publishGPT-5.5rewriteGPT-5.5polishGPT-5.5diagramGPT-6-astrapolishGPT-6-luna
Software 1.0: source code. Software 2.0: learned neural network weights. Software 3.0: natural-language prompts.
Three representations of a program, after Andrej Karpathy. Agent systems can combine all three. Source

I can tell a coding agent to parallelize work, and it will often agree with me while still doing one thing at a time.

That failure looks like a prompting problem until you inspect the trace. The sentence “fan out independent subtasks” changed the model’s intention, but it did not create a worker pool, a scheduler, a merge rule, a verifier, or a budget policy. The prompt moved. The action space did not.

That is the category error hiding inside a lot of talk about self-improving agents. We say:

The system optimizes itself.

as if there were one surface called “the system.”

There is not.

There are prompts, skills, tools, traces, memory stores, evaluators, runtime graphs, harnesses, model weights, and release gates. Each one can be optimized. Each one needs a different kind of evidence. Each one can fail in a different way.

Self-improvement only has content after you name the task class. The target is better execution of the work the user is trying to get done: the code change, research answer, design review, deployment, diagnosis, or decision that caused the agent to be invoked in the first place. A system that improves a judge score while making that work slower, less faithful to intent, or harder to audit has optimized a proxy, not the task.

The useful questions are more concrete:

  • Which user task distribution is being improved?
  • What is allowed to change?
  • What evidence shows improvement on that task?
  • How are candidates generated?
  • What gate decides promotion?
  • What can go wrong when that layer changes?

Those six questions are the self-improving stack.

The Loop Behind The Word

A self-improving agent system has a closed loop:

  1. 01
    Run
  2. 02
    Observe
  3. 03
    Diagnose
  4. 04
    Propose
  5. 05
    Validate
  6. 06
    Promote
  7. 07
    Remember
  8. 08
    Govern

The loop is only real when each verb has a concrete implementation.

  • Run: Execute the agent under a sampled user-task scenario.
  • Observe: Capture a full trace, not only a score.
  • Diagnose: Identify failure modes and missing knowledge.
  • Propose: Generate a candidate change.
  • Validate: Test the candidate against the baseline.
  • Promote: Replace the baseline only if the gate passes.
  • Remember: Persist the right lesson for future runs.
  • Govern: Keep the optimizer inside its authority and evidence boundary.

The system is not self-improving because it says “reflect.” It is self-improving when a future run gets better at the class of user tasks it is meant to serve because a previous run produced admissible evidence.

The compact equation is:

Δt=Ex∼Dtask[Eval⁡(Run⁡(ct,x),Run⁡(st,x))]st+1={Promote⁡(st,ct),if Gate⁡(Δt,policy) passes,st,otherwise.\begin{aligned} \Delta_t &= \mathbb{E}_{x\sim D_{\text{task}}}[\operatorname{Eval}(\operatorname{Run}(c_t,x),\operatorname{Run}(s_t,x))] \\ s_{t+1} &= \begin{cases} \operatorname{Promote}(s_t,c_t), & \text{if }\operatorname{Gate}(\Delta_t,\text{policy})\text{ passes}, \\ s_t, & \text{otherwise.} \end{cases} \end{aligned}

where:

st=current system statect=candidate statex∼Dtask(a user task drawn from the target task distribution)Run⁡(⋅)=full agent trajectory under sampled user-task scenariosEval⁡(⋅)=measured evidence on task outcomes and trace behaviorGate⁡(⋅)=promotion rule under policy\begin{aligned} s_t &= \text{current system state} \\ c_t &= \text{candidate state} \\ x &\sim D_{\text{task}}\quad\text{(a user task drawn from the target task distribution)} \\ \operatorname{Run}(\cdot) &= \text{full agent trajectory under sampled user-task scenarios} \\ \operatorname{Eval}(\cdot) &= \text{measured evidence on task outcomes and trace behavior} \\ \operatorname{Gate}(\cdot) &= \text{promotion rule under policy} \end{aligned}

The system improves only when the promoted state performs better on the user-task distribution while staying inside cost, safety, integrity, and governance constraints.

The Stack

The stack has layers because “candidate” can mean many different things, and because every candidate is supposed to improve some named task class.

LayerMutable surfaceSearch operatorTrusted feedbackGate
Optimization theorycandidate state and objectivehill climbing, bandits, evolutionary search, Bayesian searchreward, loss, regret, Pareto frontierstatistical and structural validity
Prompt optimizationinstructions, examples, rubrics, LM program textGEPA, MIPRO, DSPy, AxLLM-style searchtask score, judge score, validation setheld-out prompt eval
Skill optimizationreusable procedures and action policiesreflection, trace-to-skill, skill mutationtransfer success, recurrence reductionskill invocation and transfer tests
Runtime topologydrivers, fanout, reviewers, selectors, turn budgetstopology search, hand-designed patterns, meta-harness mutationfull trajectory score and costtrace integrity plus budget gate
Multi-agent coordinationroles, contracts, supervisors, workersrole decomposition, delegation, debate, votecoordination quality, disagreement resolutionrole isolation and selector audit
Test-time computesamples, branches, retries, verifier callsbest-of-N, tree search, adaptive allocationcompute-matched liftPareto dominance under equal compute
Evaluation gatesscorecards, judges, baselines, release criteriaevaluator design and calibrationpaired deltas, human calibration, deterministic checksfail-closed promotion
Trace systemsspans, artifacts, raw calls, replay recordstrace mining and analyst loopscausal evidence from trajectoriescapture integrity and replayability
Harness evolutionsource code around the agentcode search, meta-harness, worktree variantsbenchmark and production evidencerelease gate outside the mutation surface
Post-trainingmodel weights or adaptersSFT, RLHF, DPO, PPO, GRPO, tool-use RLtraining loss, preferences, verifiable reward, deployment evalmodel release and data-governance gate
Memory and knowledgepersistent state across episodestrace mining, proposal review, retrieval tuningretrieval-conditioned task liftsource, freshness, scope, poisoning checks
Governanceauthority, risk controls, release policysafety-case iterationred-team, audit, incident, outcome evidenceaccountable approval and rollback

This table is the core of the series.

GEPA, SkillOpt, meta-harness, post-training, and memory flywheels all have the same outer skeleton:

Candidate evaluation loop A candidate is proposed from the current program, run and measured, then evaluated at a promotion gate. Acceptance replaces the current program; rejection leaves it unchanged. Evidence from either outcome informs the next proposal. Current program Candidate program Run and measure Promotion gate ACCEPTEDReplace current program REJECTEDKeep current program Evidence informs next proposal Candidate evaluation loop A candidate is proposed from the current program, run and measured, then evaluated at a promotion gate. Acceptance replaces the current program; rejection leaves it unchanged. Evidence from either outcome informs the next proposal. Current program Candidate program Run and measure Promotion gate ACCEPTEDReplace currentprogram REJECTEDKeep currentprogram Evidence informs next proposal

They are not the same system because they mutate different surfaces.

A prompt optimizer cannot add a new sandbox boundary. A skill optimizer cannot guarantee a runtime will invoke the skill. A memory system cannot invent a fanout topology. A harness optimizer cannot safely approve itself if it owns the gate. Post-training can move model behavior, but it also moves the rollback and explanation boundary.

The layer is not an implementation detail.

The layer determines the reachable set.

Layer Confusion

The most common mistake is using the optimizer from one layer to fix a failure in another.

SymptomTempting fixLikely real layer
The agent says “parallelize” but still works seriallyoptimize the promptruntime topology
A worker follows the persona but misses the task contractrewrite the personamulti-agent coordination
The same tool argument bug recursadd another reminderskill or harness evolution
Scores improve while traces get messiertune the judgeevaluation gate and trace integrity
A memory helps one task and poisons anotherretrieve more contextmemory gate and scope policy
A benchmark improves only with more samplesclaim better reasoningtest-time compute accounting

This is the practical reason to map every system by mutable surface first.

Where Prompt Optimization Stops

Prompt optimization is real.

It can discover better instructions, examples, decomposition strategies, rubric wording, and reflection text. Systems like GEPA and DSPy-style optimizers made that point concrete: language program text can be searched.

But prompt search operates inside a fixed runtime.

The objective is roughly:

p∗=argmax⁡pE[R(Run⁡(hfixed,p,x))]p^*=\operatorname*{argmax}_p\mathbb{E}[R(\operatorname{Run}(h_{\text{fixed}},p,x))]

The harness hfixedh_{\text{fixed}} is held constant.

If the runtime lacks a worker pool, the prompt can ask for parallelism but cannot create it. If the tool graph lacks a verifier, the prompt can request verification but cannot execute one. If the evaluator leaks holdout answers, the prompt can overfit beautifully.

This is why the series keeps separating:

  • Better wording
  • Better procedure
  • Better topology
  • Better evaluator
  • Better harness
  • Better model
  • Better memory
  • Better governance

Those are different control surfaces.

Skills Are Trainable State

Skills sit between prompts and code.

A skill is durable procedural memory:

When this task class appears, use this decomposition, with these tools, under these checks, and stop under these conditions.

That makes skills more reusable than a one-off prompt and less rigid than hard-coded application logic.

The hard part is activation. A skill that never triggers is inert. A skill that triggers everywhere becomes a new bug. The gate has to test transfer:

  • Does the skill improve held-out tasks in the intended class?
  • Does it avoid harming nearby tasks outside the class?
  • Does it reduce repeated failures?

That is why skill optimization belongs in the stack but does not replace runtime design.

Topology Is The Action Space

Agent behavior is not only model output.

It is workflow shape:

  • Single shot
  • Refine loop
  • Fanout and vote
  • Planner plus worker
  • Researcher plus coder
  • Supervisor plus reviewer
  • Debate
  • Tree search
  • Human approval gate

The topology defines which actions exist and which observations can influence future actions.

This matters for multi-agent systems. “Persona” is content. “Coordinator,” “reviewer,” “selector,” “budget holder,” and “release approver” are structural roles. You can optimize persona text, but coordination quality usually depends on contracts, routing, isolation, and selection.

For maxTurns=0 worker flows, learning does not happen inside the worker’s conversation. It happens across runs:

  1. 01
    Pre-run retrieval
  2. 02
    Single worker attempt
  3. 03
    Post-run trace capture
  4. 04
    Analyst finding
  5. 05
    Candidate change
  6. 06
    Promotion gate
  7. 07
    Next run

That is still self-improvement, but the loop lives in the harness.

Test-Time Compute Is The Baseline

Before claiming that a new optimizer improved the agent, beat random or naive sampling at equal compute.

A lot of agent improvements are really compute allocation changes:

  • More samples
  • More branches
  • More retries
  • More verifier calls
  • A more expensive judge
  • More time

Those can be useful. They are not free.

The fair comparison is:

ΔQ=Q(candidate)−Q(baseline),ΔC=C(candidate)−C(baseline)\Delta Q = Q(\text{candidate}) - Q(\text{baseline}), \qquad \Delta C = C(\text{candidate}) - C(\text{baseline})

A candidate that wins only by spending more may still be worth shipping, but the claim is different. It is a cost-quality trade, not pure intelligence gain.

The Gate Is The Optimizer

The promotion gate decides what the system becomes.

If the gate rewards shallow style, the system learns shallow style. If the gate leaks the answer, the system learns leakage. If the gate ignores cost, the system learns to spend. If the gate ignores safety, the system learns unsafe shortcuts.

A usable gate is explicit:

promote(c) iff
  paired_delta(c, baseline, holdout) > threshold
  and deterministic_verifiers(c) pass
  and trace_integrity(c) passes
  and cost(c) <= budget
  and safety_regression(c) == false

That gate can be statistical, deterministic, human-reviewed, or all three.

The key is that it is separate from the candidate.

Traces Are The Data

Scores say that something happened.

Traces say what happened.

An agent trace needs enough information to explain the mechanism:

  • Which prompt
  • Which model
  • Which tools and arguments
  • Which observations and retrieved documents
  • Which artifacts and verifier
  • Which failure class and budget
  • Which outcome

Without traces, the system can only hill climb on a lossy projection.

With traces, the system can diagnose:

  • Missing knowledge
  • Bad tool argument
  • Weak verifier
  • Wrong selector
  • Coordination failure
  • Memory poisoning
  • Budget breach
  • Unsafe side effect

That is why traces are not logging decoration. They are the training data for the external-state loop.

Harness Evolution Changes The Machine

When prompt, skill, and topology tuning plateau, the mutable surface may need to be source code.

Harness evolution changes:

  • Planner contracts
  • Tool routers and selectors
  • Trace emitters and verifiers
  • Benchmark adapters
  • Worktree lifecycle
  • Promotion gates
  • Memory write paths

This is powerful because it expands the reachable set.

It is dangerous because the harness may contain the evaluator. The core rule from the governance layer is:

The optimizer cannot own the gate that promotes it.

If the candidate can rewrite the judge or release policy that approves it, the loop is no longer honest.

Post-Training Moves The Model Boundary

Most systems discussed before the post-training layer mutate external state:

  • Prompt
  • Skill
  • Tool documentation
  • Memory
  • Runtime
  • Harness
  • Evaluator

Post-training mutates model behavior itself:

θt+1=Update⁡(θt,data,objective)\theta_{t+1}=\operatorname{Update}(\theta_t,\text{data},\text{objective})

or:

θ′=θbase+Δadapter\theta' = \theta_{\text{base}} + \Delta_{\text{adapter}}

That can generalize better than prompt edits when the signal is strong. It also makes the behavior harder to inspect, partially roll back, and attribute to one trace.

That is why post-training sits near the top of the stack. It is not “more advanced prompt optimization.” It changes the policy.

Memory Is Not Automatically Learning

Memory changes future runs by changing what persists.

The update rule is:

Mt+1={Apply⁡(Mt,ut),if Gmem(ut,trace,policy) passes,Mt,otherwise.M_{t+1}=\begin{cases} \operatorname{Apply}(M_t,u_t),&\text{if }G_{\text{mem}}(u_t,\text{trace},\text{policy})\text{ passes}, \\ M_t,&\text{otherwise.} \end{cases}

A memory system is useful when:

Score⁡(with memory)−Score⁡(without memory)>threshold\operatorname{Score}(\text{with memory}) - \operatorname{Score}(\text{without memory}) > \text{threshold}

under cost, freshness, privacy, and poisoning constraints.

Remembering more is not learning. Learning is remembering the right thing, retrieving it in the right context, and proving it improved behavior.

Governance Closes The Loop

Governance is not the opposite of autonomy.

Governance is what makes autonomy accountable.

A self-improving system needs a safety case:

  • Claim
  • Scope
  • Evidence
  • Residual risk
  • Owner
  • Release gate
  • Rollback path

The system can propose improvements. The gate decides which improvements persist. The owner accepts residual risk.

The minimum release rule is:

ship(candidate) iff
  improves(candidate)
  and evidence_complete(candidate)
  and controls_pass(candidate)
  and owner_accepts_residual_risk(candidate)

Without that rule, self-improvement can become proxy hacking with better branding.

The Practical Test

When someone says their agent improves itself, ask for the layer.

  • What changed, and who proposed it?
  • What evidence was captured, and what baseline was beaten?
  • What held-out set was protected, and what gate approved it?
  • What became more expensive or riskier?
  • What can roll back, and what persisted into the next run?

If those questions have concrete answers, there may be a real loop.

If the answer is only “the model reflected,” there probably is not.

Series Map

Source Trail

Source freshness checked on 2026-06-06.

Revision history11revisions
  1. GPT-6-lunapolish+98−164 view trace →
    Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
    show diff
    diff --git a/src/content/posts/the-self-improving-stack.mdx b/src/content/posts/the-self-improving-stack.mdxindex e30766a..b5c2866 100644--- a/src/content/posts/the-self-improving-stack.mdx+++ b/src/content/posts/the-self-improving-stack.mdx@@ -43,15 +43,15 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro';+ I can tell a coding agent to parallelize work, and it will often agree with me while still doing one thing at a time.  That failure looks like a prompting problem until you inspect the trace. The sentence "fan out independent subtasks" changed the model's intention, but it did not create a worker pool, a scheduler, a merge rule, a verifier, or a budget policy. The prompt moved. The action space did not.  That is the category error hiding inside a lot of talk about self-improving agents. We say: -```text-the system optimizes itself-```+The system optimizes itself.  as if there were one surface called "the system." @@ -63,14 +63,12 @@ Self-improvement only has content after you name the task class. The target is b  The useful questions are more concrete: -```text-what user task distribution is being improved?-what is allowed to change?-what evidence says it improved that task?-how are candidates generated?-what gate decides promotion?-what can go wrong when that layer changes?-```+- Which user task distribution is being improved?+- What is allowed to change?+- What evidence shows improvement on that task?+- How are candidates generated?+- What gate decides promotion?+- What can go wrong when that layer changes?  Those six questions are the self-improving stack. @@ -78,29 +76,18 @@ Those six questions are the self-improving stack.  A self-improving agent system has a closed loop: -```text-run-observe-diagnose-propose-validate-promote-remember-govern-```+<Steps layout="flow" items={[{title:'Run'}, {title:'Observe'}, {title:'Diagnose'}, {title:'Propose'}, {title:'Validate'}, {title:'Promote'}, {title:'Remember'}, {title:'Govern'}]} />  The loop is only real when each verb has a concrete implementation. -```text-run: execute the agent under a sampled user-task scenario-observe: capture a full trace, not only a score-diagnose: identify failure modes and missing knowledge-propose: generate a candidate change-validate: test the candidate against baseline-promote: replace baseline only if the gate passes-remember: persist the right lesson for future runs-govern: keep the optimizer inside its authority and evidence boundary-```+- **Run:** Execute the agent under a sampled user-task scenario.+- **Observe:** Capture a full trace, not only a score.+- **Diagnose:** Identify failure modes and missing knowledge.+- **Propose:** Generate a candidate change.+- **Validate:** Test the candidate against the baseline.+- **Promote:** Replace the baseline only if the gate passes.+- **Remember:** Persist the right lesson for future runs.+- **Govern:** Keep the optimizer inside its authority and evidence boundary.  The system is not self-improving because it says "reflect." It is self-improving when a future run gets better at the class of user tasks it is meant to serve because a previous run produced admissible evidence. @@ -155,12 +142,7 @@ This table is the core of the series.  GEPA, SkillOpt, meta-harness, post-training, and memory flywheels all have the same outer skeleton: -```text-propose candidate-run candidate-measure candidate-promote or reject-```+<Steps layout="flow" items={[{title:'Propose candidate'}, {title:'Run candidate'}, {title:'Measure candidate'}, {title:'Promote or reject'}]} />  They are not the same system because they mutate different surfaces. @@ -205,16 +187,14 @@ If the runtime lacks a worker pool, the prompt can ask for parallelism but canno  This is why the series keeps separating: -```text-better wording-better procedure-better topology-better evaluator-better harness-better model-better memory-better governance-```+- Better wording+- Better procedure+- Better topology+- Better evaluator+- Better harness+- Better model+- Better memory+- Better governance  Those are different control surfaces. @@ -224,23 +204,15 @@ Skills sit between prompts and code.  A skill is durable procedural memory: -```text-when this task class appears,-use this decomposition,-with these tools,-under these checks,-and stop under these conditions-```+> When this task class appears, use this decomposition, with these tools, under these checks, and stop under these conditions.  That makes skills more reusable than a one-off prompt and less rigid than hard-coded application logic.  The hard part is activation. A skill that never triggers is inert. A skill that triggers everywhere becomes a new bug. The gate has to test transfer: -```text-does the skill improve held-out tasks in the intended class?-does it avoid harming nearby tasks outside the class?-does it reduce repeated failures?-```+- Does the skill improve held-out tasks in the intended class?+- Does it avoid harming nearby tasks outside the class?+- Does it reduce repeated failures?  That is why skill optimization belongs in the stack but does not replace runtime design. @@ -250,17 +222,15 @@ Agent behavior is not only model output.  It is workflow shape: -```text-single shot-refine loop-fanout and vote-planner plus worker-researcher plus coder-supervisor plus reviewer-debate-tree search-human approval gate-```+- Single shot+- Refine loop+- Fanout and vote+- Planner plus worker+- Researcher plus coder+- Supervisor plus reviewer+- Debate+- Tree search+- Human approval gate  The topology defines which actions exist and which observations can influence future actions. @@ -268,15 +238,7 @@ This matters for multi-agent systems. "Persona" is content. "Coordinator," "revi  For `maxTurns=0` worker flows, learning does not happen inside the worker's conversation. It happens across runs: -```text-pre-run retrieval-single worker attempt-post-run trace capture-analyst finding-candidate change-promotion gate-next run-```+<Steps layout="flow" items={[{title:'Pre-run retrieval'}, {title:'Single worker attempt'}, {title:'Post-run trace capture'}, {title:'Analyst finding'}, {title:'Candidate change'}, {title:'Promotion gate'}, {title:'Next run'}]} />  That is still self-improvement, but the loop lives in the harness. @@ -286,23 +248,20 @@ Before claiming that a new optimizer improved the agent, beat random or naive sa  A lot of agent improvements are really compute allocation changes: -```text-more samples-more branches-more retries-more verifier calls-more expensive judge-more time-```+- More samples+- More branches+- More retries+- More verifier calls+- A more expensive judge+- More time  Those can be useful. They are not free.  The fair comparison is: -```text-quality(candidate) - quality(baseline)-cost(candidate) - cost(baseline)-```+$$+\Delta Q = Q(\text{candidate}) - Q(\text{baseline}), \qquad \Delta C = C(\text{candidate}) - C(\text{baseline})+$$  A candidate that wins only by spending more may still be worth shipping, but the claim is different. It is a cost-quality trade, not pure intelligence gain. @@ -335,34 +294,26 @@ Traces say what happened.  An agent trace needs enough information to explain the mechanism: -```text-which prompt-which model-which tools-which arguments-which observations-which retrieved documents-which artifacts-which verifier-which failure class-which budget-which outcome-```+- Which prompt+- Which model+- Which tools and arguments+- Which observations and retrieved documents+- Which artifacts and verifier+- Which failure class and budget+- Which outcome  Without traces, the system can only hill climb on a lossy projection.  With traces, the system can diagnose: -```text-missing knowledge-bad tool argument-weak verifier-wrong selector-coordination failure-memory poisoning-budget breach-unsafe side effect-```+- Missing knowledge+- Bad tool argument+- Weak verifier+- Wrong selector+- Coordination failure+- Memory poisoning+- Budget breach+- Unsafe side effect  That is why traces are not logging decoration. They are the training data for the external-state loop. @@ -372,25 +323,19 @@ When prompt, skill, and topology tuning plateau, the mutable surface may need to  Harness evolution changes: -```text-planner contracts-tool routers-selectors-trace emitters-verifiers-benchmark adapters-worktree lifecycle-promotion gates-memory write paths-```+- Planner contracts+- Tool routers and selectors+- Trace emitters and verifiers+- Benchmark adapters+- Worktree lifecycle+- Promotion gates+- Memory write paths  This is powerful because it expands the reachable set.  It is dangerous because the harness may contain the evaluator. The core rule from the governance layer is: -```text-the optimizer cannot own the gate that promotes it-```+The optimizer cannot own the gate that promotes it.  If the candidate can rewrite the judge or release policy that approves it, the loop is no longer honest. @@ -398,15 +343,13 @@ If the candidate can rewrite the judge or release policy that approves it, the l  Most systems discussed before the post-training layer mutate external state: -```text-prompt-skill-tool docs-memory-runtime-harness-evaluator-```+- Prompt+- Skill+- Tool documentation+- Memory+- Runtime+- Harness+- Evaluator  Post-training mutates model behavior itself: @@ -416,9 +359,9 @@ $$  or: -```text-theta' = theta_base + Delta_adapter-```+$$+\theta' = \theta_{\text{base}} + \Delta_{\text{adapter}}+$$  That can generalize better than prompt edits when the signal is strong. It also makes the behavior harder to inspect, partially roll back, and attribute to one trace. @@ -439,9 +382,9 @@ $$  A memory system is useful when: -```text-Score(with memory) - Score(without memory) > threshold-```+$$+\operatorname{Score}(\text{with memory}) - \operatorname{Score}(\text{without memory}) > \text{threshold}+$$  under cost, freshness, privacy, and poisoning constraints. @@ -455,15 +398,13 @@ Governance is what makes autonomy accountable.  A self-improving system needs a safety case: -```text-claim-scope-evidence-residual risk-owner-release gate-rollback path-```+- Claim+- Scope+- Evidence+- Residual risk+- Owner+- Release gate+- Rollback path  The system can propose improvements. The gate decides which improvements persist. The owner accepts residual risk. @@ -483,18 +424,11 @@ Without that rule, self-improvement can become proxy hacking with better brandin  When someone says their agent improves itself, ask for the layer. -```text-What changed?-Who proposed it?-What evidence was captured?-What baseline was beaten?-What held-out set was protected?-What gate approved it?-What got more expensive?-What became riskier?-What can roll back?-What persisted into the next run?-```+- What changed, and who proposed it?+- What evidence was captured, and what baseline was beaten?+- What held-out set was protected, and what gate approved it?+- What became more expensive or riskier?+- What can roll back, and what persisted into the next run?  If those questions have concrete answers, there may be a real loop. 
  2. GPT-6-lunapolish+33−28 view trace →
    Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
    show diff
    diff --git a/src/content/posts/the-self-improving-stack.mdx b/src/content/posts/the-self-improving-stack.mdxindex 03e9f0c..45e176a 100644--- a/src/content/posts/the-self-improving-stack.mdx+++ b/src/content/posts/the-self-improving-stack.mdx@@ -104,25 +104,29 @@ The system is not self-improving because it says "reflect." It is self-improving  The compact equation is: -```text-Delta_t = E_{x ~ D_task}[Eval(Run(c_t, x), Run(s_t, x))]--s_{t+1} =-  Promote(s_t, c_t)-  if Gate(Delta_t, policy) passes-  else s_t-```+$$+\begin{aligned}+  \Delta_t &= \mathbb{E}_{x\sim D_{\text{task}}}[\operatorname{Eval}(\operatorname{Run}(c_t,x),\operatorname{Run}(s_t,x))] \\+  s_{t+1} &=+  \begin{cases}+    \operatorname{Promote}(s_t,c_t), & \text{if }\operatorname{Gate}(\Delta_t,\text{policy})\text{ passes}, \\+    s_t, & \text{otherwise.}+  \end{cases}+\end{aligned}+$$  where: -```text-s_t = current system state-c_t = candidate state-x ~ D_task = a user task drawn from the task distribution the system is meant to serve-Run(.) = full agent trajectory under sampled user-task scenarios-Eval(.) = measured evidence on task outcomes and trace behavior-Gate(.) = promotion rule under policy-```+$$+\begin{aligned}+  s_t &= \text{current system state} \\+  c_t &= \text{candidate state} \\+  x &\sim D_{\text{task}}\quad\text{(a user task drawn from the target task distribution)} \\+  \operatorname{Run}(\cdot) &= \text{full agent trajectory under sampled user-task scenarios} \\+  \operatorname{Eval}(\cdot) &= \text{measured evidence on task outcomes and trace behavior} \\+  \operatorname{Gate}(\cdot) &= \text{promotion rule under policy}+\end{aligned}+$$  The system improves only when the promoted state performs better on the user-task distribution while staying inside cost, safety, integrity, and governance constraints. @@ -189,11 +193,11 @@ But prompt search operates inside a fixed runtime.  The objective is roughly: -```text-p* = argmax_p E[R(Run(h_fixed, p, x))]-```+$$+p^*=\operatorname*{argmax}_p\mathbb{E}[R(\operatorname{Run}(h_{\text{fixed}},p,x))]+$$ -The harness `h_fixed` is held constant.+The harness $h_{\text{fixed}}$ is held constant.  If the runtime lacks a worker pool, the prompt can ask for parallelism but cannot create it. If the tool graph lacks a verifier, the prompt can request verification but cannot execute one. If the evaluator leaks holdout answers, the prompt can overfit beautifully. @@ -404,9 +408,9 @@ evaluator  Post-training mutates model behavior itself: -```text-theta_{t+1} = Update(theta_t, data, objective)-```+$$+\theta_{t+1}=\operatorname{Update}(\theta_t,\text{data},\text{objective})+$$  or: @@ -424,11 +428,12 @@ Memory changes future runs by changing what persists.  The update rule is: -```text-M_{t+1} =-  Apply(M_t, u_t) if G_mem(u_t, trace, policy) passes-  M_t             otherwise-```+$$+M_{t+1}=\begin{cases}+  \operatorname{Apply}(M_t,u_t),&\text{if }G_{\text{mem}}(u_t,\text{trace},\text{policy})\text{ passes}, \\+  M_t,&\text{otherwise.}+\end{cases}+$$  A memory system is useful when: 
  3. GPT-6-astradiagram+5−0 view trace →
    Added a Software 3.0 figure, reused in the article and link preview; prose unchanged. This record is a selected session excerpt; the full source remains private.
    show diff
    diff --git a/src/content/posts/the-self-improving-stack.mdx b/src/content/posts/the-self-improving-stack.mdxindex acdb770..b73c291 100644--- a/src/content/posts/the-self-improving-stack.mdx+++ b/src/content/posts/the-self-improving-stack.mdx@@ -4,6 +4,11 @@ description: 'A series map for self-improving agent systems, from optimization t date: 2026-06-05 tags: ['agents', 'evals', 'systems', 'self-improvement'] draft: false+figure:+  src: '/images/software-3.svg'+  alt: 'Software 1.0: source code. Software 2.0: learned neural network weights. Software 3.0: natural-language prompts.'+  caption: 'Three representations of a program, after Andrej Karpathy. Agent systems can combine all three.'+  source: 'https://www.youtube.com/watch?v=LCEmiRjPEtQ&t=85s' series: 'the-self-improving-stack' outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' human_takeover: 'complete'
  4. GPT-5.5polish+17−9 view trace →
    clarified that self-improvement targets the user-task distribution, tightened the loop equation, and tied evidence to task outcomes
    show diff
    diff --git a/src/content/posts/the-self-improving-stack.mdx b/src/content/posts/the-self-improving-stack.mdxindex 78752d2..a425f89 100644--- a/src/content/posts/the-self-improving-stack.mdx+++ b/src/content/posts/the-self-improving-stack.mdx@@ -16,7 +16,9 @@ authors:   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'polish', date: 2026-06-08 } revisions:+  - { date: 2026-06-08, model: 'gpt-5.5', role: 'polish', note: 'clarified that self-improvement targets the user-task distribution, tightened the loop equation, and tied evidence to task outcomes', commit: '9ea5c8de9b516ce903023be5a14a8098076ddf59', trace_id: '2026-06-08T10-10-44-256Z-gpt-5.5-the-self-improving-stack-polish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-the-self-improving-stack-polish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-the-self-improving-stack-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-publish' }@@ -48,17 +50,20 @@ There is not.  There are prompts, skills, tools, traces, memory stores, evaluators, runtime graphs, harnesses, model weights, and release gates. Each one can be optimized. Each one needs a different kind of evidence. Each one can fail in a different way. +Self-improvement only has content after you name the task class. The target is better execution of the work the user is trying to get done: the code change, research answer, design review, deployment, diagnosis, or decision that caused the agent to be invoked in the first place. A system that improves a judge score while making that work slower, less faithful to intent, or harder to audit has optimized a proxy, not the task.+ The useful questions are more concrete:  ```text+what user task distribution is being improved? what is allowed to change?-what evidence says it improved?+what evidence says it improved that task? how are candidates generated? what gate decides promotion? what can go wrong when that layer changes? ``` -Those five questions are the self-improving stack.+Those six questions are the self-improving stack.  ## The Loop Behind The Word @@ -78,7 +83,7 @@ govern The loop is only real when each verb has a concrete implementation.  ```text-run: execute the agent under a scenario+run: execute the agent under a sampled user-task scenario observe: capture a full trace, not only a score diagnose: identify failure modes and missing knowledge propose: generate a candidate change@@ -88,14 +93,16 @@ remember: persist the right lesson for future runs govern: keep the optimizer inside its authority and evidence boundary ``` -The system is not self-improving because it says "reflect." It is self-improving when a future run changes in the right direction because a previous run produced admissible evidence.+The system is not self-improving because it says "reflect." It is self-improving when a future run gets better at the class of user tasks it is meant to serve because a previous run produced admissible evidence.  The compact equation is:  ```text+Delta_t = E_{x ~ D_task}[Eval(Run(c_t, x), Run(s_t, x))]+ s_{t+1} =   Promote(s_t, c_t)-  if Gate(Eval(Run(c_t), Run(s_t)), policy) passes+  if Gate(Delta_t, policy) passes   else s_t ``` @@ -104,16 +111,17 @@ where: ```text s_t = current system state c_t = candidate state-Run(.) = full agent trajectory under scenarios-Eval(.) = measured evidence+x ~ D_task = a user task drawn from the task distribution the system is meant to serve+Run(.) = full agent trajectory under sampled user-task scenarios+Eval(.) = measured evidence on task outcomes and trace behavior Gate(.) = promotion rule under policy ``` -The system improves only when the promoted state performs better on the right distribution while staying inside cost, safety, integrity, and governance constraints.+The system improves only when the promoted state performs better on the user-task distribution while staying inside cost, safety, integrity, and governance constraints.  ## The Stack -The stack has layers because "candidate" can mean many different things.+The stack has layers because "candidate" can mean many different things, and because every candidate is supposed to improve some named task class.  | Layer | Mutable surface | Search operator | Trusted feedback | Gate | |---|---|---|---|---|
  5. GPT-5.5polish+13−9 view trace →
    let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
    show diff
    diff --git a/src/content/posts/the-self-improving-stack.mdx b/src/content/posts/the-self-improving-stack.mdxindex c213875..23c834f 100644--- a/src/content/posts/the-self-improving-stack.mdx+++ b/src/content/posts/the-self-improving-stack.mdx@@ -15,7 +15,9 @@ authors:   - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-the-self-improving-stack-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Reviewed and dated the source trail, removed remaining temporal language, and marked the source-freshness checkpoint complete.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-review' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'Polished the umbrella article with a layer-confusion diagnostic, tightened promotion-gate phrasing, and verified role-scoped trace capture for separate draft and polish provenance.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-polish' }@@ -29,21 +31,23 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -Self-improvement is not a model property.+I can tell a coding agent to parallelize work, and it will often agree with me while still doing one thing at a time. -It is a system property.+That failure looks like a prompting problem until you inspect the trace. The sentence "fan out independent subtasks" changed the model's intention, but it did not create a worker pool, a scheduler, a merge rule, a verifier, or a budget policy. The prompt moved. The action space did not. -A model can sit inside a self-improving system, but the loop usually lives around it: prompts, skills, tools, traces, memory, evaluators, runtimes, harnesses, and release gates.--That distinction matters because a lot of AI discourse collapses very different loops into one phrase:+That is the category error hiding inside a lot of talk about self-improving agents. We say:  ```text the system optimizes itself ``` -That sentence is too vague.+as if there were one surface called "the system."++There is not.++There are prompts, skills, tools, traces, memory stores, evaluators, runtime graphs, harnesses, model weights, and release gates. Each one can be optimized. Each one needs a different kind of evidence. Each one can fail in a different way. -The useful questions are:+The useful questions are more concrete:  ```text what is allowed to change?@@ -53,9 +57,9 @@ what gate decides promotion? what can go wrong when that layer changes? ``` -Those five questions define the self-improving stack.+Those five questions are the self-improving stack. -## The Loop+## The Loop Behind The Word  A self-improving agent system has a closed loop: 
  6. GPT-5.5rewrite view trace →
    60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
  7. GPT-5.5publish view trace →
    Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
  8. GPT-5.5review view trace →
    Reviewed and dated the source trail, removed remaining temporal language, and marked the source-freshness checkpoint complete.
  9. GPT-5.5polish view trace →
    Polished the umbrella article with a layer-confusion diagnostic, tightened promotion-gate phrasing, and verified role-scoped trace capture for separate draft and polish provenance.
  10. GPT-5.5draft view trace →
    Drafted the umbrella series article with the closed-loop formalism, layer table, practical test, and full series map; replaced outline handoff prose and synced the research overview/status map.
  11. GPT-5.5outline view trace →
    Research planning pass from a traced session.

Comments

Comments load from GitHub Discussions via Giscus. Configure PUBLIC_GISCUS_REPO, PUBLIC_GISCUS_REPO_ID, PUBLIC_GISCUS_CATEGORY, and PUBLIC_GISCUS_CATEGORY_ID in .env. See giscus.app to generate the IDs after you enable Discussions on the repo.