Personas Are Content, Coordination Is Structure
Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
- Created
- Updated
22
Turns
1
Tool calls
14
Files touched
1763m
Duration
Files
src/content/posts/self-improving-stack-prompt-optimization.mdxsrc/content/posts/self-improving-stack-skill-optimization.mdxsrc/content/posts/self-improving-stack-agent-runtime-topology.mdxsrc/content/posts/self-improving-stack-multi-agent-coordination.mdxsrc/content/posts/self-improving-stack-test-time-compute.mdxsrc/content/posts/self-improving-stack-evaluation-gates.mdxsrc/content/posts/self-improving-stack-trace-systems.mdxsrc/content/posts/self-improving-stack-harness-evolution.mdxsrc/content/posts/self-improving-stack-post-training.mdxsrc/content/posts/self-improving-stack-memory-flywheels.mdxsrc/content/posts/self-improving-stack-governance.mdxsrc/content/posts/self-improving-stack-optimization-theory.mdxresearch/self-improving-agent-systems/01-optimization-theory.mdresearch/self-improving-agent-systems/11-memory-knowledge-flywheels.md
Commit
b8fd3db fix(layout): scope global 'aside' rule to .prose — was bleeding the prose-callout border-left into the experiment toc; restore caution tape on /experiment Conversation
22 turns. Full text where captured; older traces show only the first ~280 chars.
- user #1# AGENTS.md instructions for /Users/drew/code/blog <INSTRUCTIONS> ## Shared Agent Defaults Provider-agnostic. Synced to Claude, Codex, and OpenCode installs. ## Repos are alive Multiple agents (Claude, Codex, others) work the same branches and PRs in parallel. Unfamiliar branches, commits you didn't make, in-flight PRs — normal state, not errors. **Orient first, every session:** ``` git status; git log --oneline -10; git reflog | head -20 gh pr list --state open ``` Commit messages tell you what other agents shipped. Five seconds saves a rebase. **Don't ask permission for unexpected state.** …
- user #2I want you to reserach the following and their interplay - hillclimbing msft recent mai work - gepa, dspy, axllm, mipro prompt optimization - skill-opt skill optimization - meta-harness Basically all these things use techqniesu to optimize and hill climb in diff ways, are they doing the same thing? And then hwo does this work across multi agent multi episodic maxTurns=0 agentic flowS? like presumably personas for drivers and supervisors/coordinators of multi agent setups cant simply be so easily gepa optimized without intelligent personification or things like directives that tell subworks to…
- user #3<turn_aborted> The user interrupted the previous turn on purpose. Any running unified exec processes may still be running in the background. If any tools/commands were aborted, they may have partially executed. </turn_aborted>
- user #4And I want you to track this as a block post trace here in this blog
- assistant #5I’ll treat this as supporting research for the blog, not authorship/editing. First I’m orienting the repo state and loading the research workflow so the trace lands in the right place.
- assistant #6Repo is dirty with existing in-flight changes and two Dependabot PRs; nothing looks like a merge/rebase blocker. The relevant repo rule is the blog trace path, so next I’m finding the right post slug and the `pnpm blog research` contract before I start browsing.
- assistant #7The trace command requires a post slug. I’m checking the nearby post titles/frontmatter now; if there’s an obvious draft about self-improving or convergence systems, I’ll attach this research there rather than interrupting you for a slug.
- assistant #8The best fit is `self-improving-ai-systems`: it is a draft explicitly about GEPA/DSPy-style prompt evolution, trace-driven improvement, and where the harness ends and the model begins. I’m starting a supporting-research trace against that slug and will avoid editing the post body.
- assistant #9Local search hit a huge Codex cache file; I’m stopping that scan and narrowing it to skills/repos instead of searching the whole home cache.
- assistant #10The local side is clear enough for `meta-harness`: it is architecture evolution over a measured harness, with trace-reading proposers and Pareto frontier tracking. I’m doing a narrower pass for `skill-opt` specifically, then I’ll synthesize the distinctions instead of dumping source notes.
- assistant #11`skill-opt` appears to be a Microsoft SkillOpt project rather than a local repo skill. I found the public project page and am now tying it to the local `meta-harness` concept plus the agent-runtime pattern around `maxTurns` so the answer matches your stack, not just the papers.
- user #12continue
- assistant #13Continuing with the next highest-leverage gap: publication readiness. I’ll orient first, then check the series for draft status, ordering, provenance completeness, and any remaining publish blockers.
- assistant #14Problem: all 13 series posts are still drafts, and the remaining decision is not writing quality, it is release sequencing and takeover state. Change: I’ll audit every post for provenance/source/anti-pattern readiness, then update the existing research checkpoint with a concrete publication order and blockers rather than flipping `draft: false`. Why long-term right: publishing is an editorial boundary, so the durable move is to make the release gate explicit and leave the irreversible visibility decision to the owner. Cost: likely only the research map changes unless the audit finds a real post defect.
- assistant #15The audit found two real blockers: most leaf posts do not have a standardized dated source-trail section, and a few still use temporal wording like “today” or “recent.” I’m going to fix those across the series before touching publication order.
- assistant #16I’m standardizing the source sections now. The point is not cosmetic: it makes the release gate machine-checkable, so a future pass can tell which posts have dated source provenance without reading the whole essay.
- assistant #17 1 tool
- apply_patch
/Users/drew/code/blog/src/content/posts/self-improving-stack-prompt-optimization.mdxinput
*** Begin Patch *** Update File: /Users/drew/code/blog/src/content/posts/self-improving-stack-prompt-optimization.mdx @@ - note: 'Converted the prompt optimization scaffold into a full technical draft with math, optimizer history, eval protocol, multi-agent boundaries, and Tangle runtime/eval placement.' + note: 'Converted the prompt optimization outline into a full technical draft with math, optimizer history, eval protocol, multi-agent boundaries, and Tangle runtime/eval placement.' @@ -## Sources Checked +## Source Trail + +Source freshness checked on 2026-06-06. *** Update File: /Users/dr…
-
- assistant #18The source sections are now consistent across the series. I’m rerunning the audit to catch anything the first pass surfaced, then I’ll write the publication order into the checkpoint docs.
- assistant #19The release decision is straightforward after the audit: the publication order should follow the conceptual dependency chain, but the posts should stay draft until Drew’s human takeover pass. I’m writing that as the explicit gate now.
- assistant #20Because I edited visible post bodies, I’m attaching review provenance to each touched leaf post. This is the important difference between “research notes changed” and “publication text changed”: post text gets revision rows and trace files.
- assistant #21Trace capture is still attaching review rows across the series. It is slow because each finish command extracts a bounded session window and writes a separate trace per post, but it is the right provenance shape.
- assistant #22Halfway through the review-trace attachment. Once this finishes, I’ll rerun the clean scan and build because frontmatter mutations can break content schema just as easily as body edits.
Diff
No commit diff available — showing current file content (first 80 lines).
---title: 'Prompt Optimization Is Not The Whole Game'description: 'Where GEPA, DSPy, MIPRO, AxLLM, and related prompt optimizers fit inside a larger self-improving agent stack.'date: 2026-06-05tags: ['agents', 'prompts', 'evals', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-05 - model: 'gpt-5.5' role: 'polish' date: 2026-06-05 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-17-233Z-gpt-6-luna-self-improving-stack-prompt-optimization-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-prompt-optimization-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-prompt-optimization-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-prompt-optimization-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-prompt-optimization-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-prompt-optimization-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-05 model: 'gpt-5.5' role: 'draft' note: 'Converted the prompt optimization outline into a full technical draft with math, optimizer history, eval protocol, multi-agent boundaries, and Tangle runtime/eval placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-05 model: 'gpt-5.5' role: 'polish' note: 'Polished post 2 for sharper experimental-design framing, tighter optimizer taxonomy, and clearer runtime/eval boundaries.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro';The first time prompt optimization feels magical is also the moment it starts lying to you.You change a sentence, run the benchmark, and the score moves. GEPA reflects over traces. MIPRO searches instructions and demonstrations. DSPy compiles an LM program. AxLLM packages the optimizer into a TypeScript surface. The intervention is small, the evidence is numeric, and the temptation is to say the agent improved.Sometimes it did.Sometimes the prompt merely learned the evaluator, or found a better wording inside a fixed runtime, or exposed that the real bottleneck was not text at all. The agent might be failing because the tool surface is wrong, retrieval is stale, fanout is unavailable, the judge rewards the wrong behavior, the model is underpowered, the trace is incomplete, or the coordinator is operating with the wrong topology.Prompt optimization is powerful when a text surface has causal leverage over the failure. It is a confound when the missing capability lives outside text.So the question is not "which prompt optimizer is best?" The question is:Which factors are mutable, which stay fixed, and which evaluator is trusted to promote a candidate?Answer that and the ecosystem stops looking like magic. It becomes experimental design.## The ConfoundLet:$$\begin{aligned} p &= \text{prompt artifact or prompt-like text surface} \\ d &= \text{selected demonstrations or exemplars} \\ m &= \text{model or backend} \\ h &= \text{runtime and harness} \\ x &= \text{task sampled from eval distribution }D \\ y &= \text{system output or full trajectory} \\ R &= \text{reward, metric, judge, or scoring function} \\---title: 'Skills Are Trainable State'description: 'How SkillOpt, Voyager-style skill libraries, and agent skills turn durable procedure into an optimization surface.'date: 2026-06-05tags: ['agents', 'skills', 'evals', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-05 - model: 'gpt-5.5' role: 'polish' date: 2026-06-05 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-17-233Z-gpt-6-luna-self-improving-stack-skill-optimization-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-skill-optimization-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-skill-optimization-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-skill-optimization-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-skill-optimization-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-skill-optimization-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-05 model: 'gpt-5.5' role: 'draft' note: 'Converted the skill optimization outline into a full draft covering SkillOpt, Voyager, Trace2Skill, CoEvoSkills, agent skill safety, eval protocol, and Tangle runtime/eval placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-05 model: 'gpt-5.5' role: 'polish' note: 'Polished post 3 for sharper skill-state framing, activation-policy emphasis, and cleaner SkillOpt/runtime/eval boundaries.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro';A prompt can rescue one run.A skill can change the next hundred.That is the opportunity and the danger. A prompt lives in the current context. A skill is procedural state that an agent can reactivate later, possibly across tasks, models, and sessions. If the skill teaches the agent the right repair routine, the system compounds. If it teaches a bad habit, the system inherits it.SkillOpt matters because it treats that procedural state as trainable. It takes a natural-language skill document, runs tasks, reflects on scored trajectories, proposes bounded edits, accepts only validation-improving updates, and exports a deployable `best_skill.md`.The model weights do not move.The agent's operating procedure does.## What Counts As A Skill?The term "skill" is overloaded, so the first move is to separate the artifact from adjacent control surfaces.A skill is a durable, reusable procedure that the agent can load, retrieve, or execute when a task matches some condition. It can be natural language, code, a bundle of instructions plus scripts, or an entry in an executable skill library.It is not identical to a prompt, memory, tool, or runtime.| Artifact | What it stores | How it affects behavior | Main risk || --- | --- | --- | --- || Prompt | in-context instruction | steers the current run | prompt overfit || Memory | facts, preferences, past observations | changes what context is recalled | stale or poisoned state || Tool | executable affordance | expands the action space | unsafe side effects || Skill | reusable procedure | changes how the agent operates across tasks | persistent bad habits || Runtime | loop, topology, budgets, dispatch | controls what actually executes | fake autonomy or hidden confounding |Claude Code skills make this practical: a `SKILL.md` file has frontmatter for discovery and markdown instructions for execution, optionally with supporting files and scripts. Anthropic's docs describe skills as dynamically loaded task procedures, with progressive disclosure so long instructions are loaded only when relevant. That is an operational distinction, not cosmetic packaging.---title: 'Topology Is The Missing Action Space'description: 'Why multi-agent self-improvement needs explicit runtime primitives for fanout, refine, select, parallelism, supervision, budgets, and replay.'date: 2026-06-05tags: ['agents', 'runtime', 'systems', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-05 - model: 'gpt-5.5' role: 'polish' date: 2026-06-05 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-17-233Z-gpt-6-luna-self-improving-stack-agent-runtime-topology-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-agent-runtime-topology-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-agent-runtime-topology-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-agent-runtime-topology-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-agent-runtime-topology-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-agent-runtime-topology-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-05 model: 'gpt-5.5' role: 'draft' note: 'Expanded the runtime topology outline into a full draft with formal topology variables, shipped Tangle runtime primitives, external orchestration context, eval protocol, and failure modes.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-05 model: 'gpt-5.5' role: 'polish' note: 'Polished runtime topology framing, substrate-boundary language, GEPA/topology distinction, and bridge into multi-agent coordination.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro';When I tell a coding agent to "parallelize the work," I am not asking for a different tone.I am asking for a different execution graph: spawn independent executions, cap concurrency, isolate state, collect traces, score results, select or merge outputs, cancel losers, and account for cost. If the runtime cannot express those moves, a prompt can only ask the model to simulate the shape.This is the missing action space in many agent systems.Prompt optimization tunes text. Skill optimization trains durable procedure. Runtime topology optimization changes what can actually happen during execution.## What Topology MeansAn agent runtime topology is the executable shape of the work.It is not the persona. It is not the supervisor prompt. It is not the model's private chain of thought. It is the control structure that decides which agent runs, with which tools, in what order, under what budget, with what state isolation, and with what termination rule.A minimal topology has:- **Nodes:** Agents, tools, validators, selectors, and human gates- **Edges:** Sequence, fanout, handoff, retry, interrupt, and merge- **State:** Traces, memory, artifacts, budgets, and run handles- **Policy:** Planning, selection, cancellation, promotion, and replayThe runtime action space is the set of moves the system can execute:```textA_runtime = { call_tool, call_agent, delegate, fork,---title: 'Personas Are Content, Coordination Is Structure'description: 'How driver, worker, selector, reviewer, analyst, and coordinator roles become reliable multi-agent systems instead of roleplay.'date: 2026-06-05tags: ['agents', 'multi-agent', 'systems', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-05 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-57-880Z-gpt-6-luna-self-improving-stack-multi-agent-coordination-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-multi-agent-coordination-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-multi-agent-coordination-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-multi-agent-coordination-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-multi-agent-coordination-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-multi-agent-coordination-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-05 model: 'gpt-5.5' role: 'draft' note: 'Drafted the multi-agent coordination post with formal role contracts, coordination patterns, disagreement math, Tangle runtime/eval placement, and failure modes.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished with refreshed agent-runtime 0.26.0 and agent-eval 0.34.1 surface review, focused kernel/conversation split, MCP delegation boundary, eval promotion map, and cleaner substrate language.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro'I can give five agents five names and still have one mind making one mistake."Researcher," "critic," "architect," "driver," and "supervisor" are not coordination by themselves. They can be the same model, with the same blind spot, reading the same context, under the same budget, producing five paraphrases of the same failure. The cast list changed. The information structure did not.Multi-agent work becomes real when disagreement becomes useful. That requires separate contracts, state boundaries, tool permissions, selection rules, budgets, and traces.The persona is content.Coordination is structure.## The Coordination SurfaceThe previous post made runtime topology explicit:$$\begin{aligned} g &= \text{executable graph of agents, tools, validators, selectors, handoffs, and gates} \\ \pi &= \text{runtime policy over graph moves}\end{aligned}$$Multi-agent coordination sits one layer above that. It decides what the nodes are supposed to do together.Let:$$\begin{aligned}---title: 'Beat Random At Equal Compute First'description: 'Why best-of-N, self-consistency, verifier reranking, and compute-matched controls are the baseline for agent topology claims.'date: 2026-06-05tags: ['agents', 'evals', 'reasoning', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-57-880Z-gpt-6-luna-self-improving-stack-test-time-compute-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-test-time-compute-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-test-time-compute-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-test-time-compute-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-test-time-compute-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-test-time-compute-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the test-time compute post with compute-matched baselines, selection math, verifier limits, adaptive allocation, and Tangle runtime/eval placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the test-time compute post with the finite-sample pass@k estimator, adaptive compute as a control problem, Pareto dominance language, and stricter promotion criteria.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro'I do not trust a multi-agent system until it beats the boring baseline. More agents is not a strategy. It is a cost increase until it beats blind extra compute.That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity. It spent more budget and mislabeled the budget as architecture.The first gate is simple:> Beat random at equal compute.Not beat one greedy sample. Not beat the weakest baseline. Not beat a single run after quietly raising the turn budget. Beat the best simple use of the same budget.## What Test-Time Compute MeansTest-time compute is extra computation spent after the model weights are fixed and the task is known.It can be spent on:- longer reasoning- repeated sampling- self-consistency- verifier reranking- tree search- iterative refinement- multi-agent fanout- tool use- debate- retrieval- code execution---title: 'The Gate Is The Optimizer'description: 'Why held-out promotion, judge reliability, failure taxonomies, cost ceilings, and confidence intervals decide whether self-improvement is real.'date: 2026-06-05tags: ['agents', 'evals', 'systems', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-37-367Z-gpt-6-luna-self-improving-stack-evaluation-gates-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-evaluation-gates-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-evaluation-gates-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-evaluation-gates-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-evaluation-gates-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-evaluation-gates-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the evaluation-gates post with held-out promotion math, scorecard cells, judge reliability, backend integrity, release confidence, and local Tangle package placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the evaluation-gates post by tightening profile-cell claims against the local AgentProfileCell schema, adding gate pre-registration invariants, clarifying bootstrap interval wording, and preserving the fail-closed promotion model.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---A green score is not a release decision. It is evidence entering a release policy, and in self-improving loops that distinction matters more than the optimizer.GEPA, MIPRO, SkillOpt, topology search, and meta-harness can all generate candidates forever. The gate decides which candidate becomes the system future agents inherit. If the gate is weak, the optimizer learns the gate. If the gate is honest, the optimizer has to improve the product.This is why the gate is not an administrative detail after the interesting work. It is the objective boundary.## What A Gate IsA gate is a promotion policy.Let:$$\begin{aligned} b &= \text{baseline system} \\ c &= \text{candidate system} \\ x &= \text{scenario} \\ p &= \text{agent profile cell} \\ z &= \text{seed or replicate id} \\ R &= \text{task reward or score} \\ C &= \text{measured cost vector} \\ T &= \text{trace integrity predicate} \\ D_{\text{search}} &= \text{search split} \\ D_{\text{holdout}} &= \text{held-out split}\end{aligned}$$The gate is a function:```text---title: 'Traces Are The Training Data'description: 'Why self-improving agents need full trajectories, tool spans, analyst findings, provenance, and replay instead of final scores alone.'date: 2026-06-05tags: ['agents', 'traces', 'evals', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-37-367Z-gpt-6-luna-self-improving-stack-trace-systems-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-trace-systems-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-trace-systems-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-trace-systems-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-trace-systems-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-trace-systems-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the trace-systems post with formal trajectory notation, span ontology, raw provider capture, replay, trace integrity, analyst findings, leakage firewalls, and local Tangle package placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the trace-systems post by adding a trace granularity test, tightening the information-loss claim to a fixed scorer, correcting loop trace event details, and adding trace store surfaces from the local agent-eval audit.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro'The optimizer wants a score. I want the run, because a score only tells you that something happened while a trace preserves enough mechanism to explain what happened.That difference is the difference between tuning a system and optimizing an unidentified projection.A self-improving agent can only improve from the information it preserves. If the run record says "failed, score 0.42," the optimizer can only infer weak global pressure. If the trace says the planner chose the wrong tool, the tool call used a stale argument, the retrieval span returned irrelevant context, the judge penalized a missing artifact, and the retry loop repeated the same action three times, the optimizer has a causal surface.The trace is not decoration around the eval. The trace is the data.## The Information Loss ProblemAn agent run is a trajectory:$$\tau=(x,s_0,a_1,o_1,s_1,\ldots,a_T,o_T,y)$$where:$$\begin{aligned} x &= \text{task} \\ s_t &= \text{internal and external state} \\ a_t &= \text{action} \\ o_t &= \text{observation} \\ y &= \text{outcome}\end{aligned}$$---title: 'When The Harness Has To Evolve'description: 'Why meta-harness, AlphaEvolve-style code search, worktree isolation, and architecture frontiers matter after prompt and skill tuning plateau.'date: 2026-06-05tags: ['agents', 'systems', 'architecture', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.3-codex-spark', role: 'rewrite', date: 2026-06-06 } - { model: 'gpt-5.3-codex-spark', role: 'polish', date: 2026-06-06 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-57-880Z-gpt-6-luna-self-improving-stack-harness-evolution-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-harness-evolution-polish' } - { date: 2026-06-06, model: 'gpt-5.3-codex-spark', role: 'polish', note: 'we have a company website in ~/webb/tangle-website maybe? I want to evaluate which blog posts from this blog we can mirror on that website s · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-06T18-19-05-739Z-gpt-5.3-codex-spark-self-improving-stack-harness-evolution-polish' } - { date: 2026-06-06, model: 'gpt-5.3-codex-spark', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-06T18-19-05-739Z-gpt-5.3-codex-spark-self-improving-stack-harness-evolution-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-harness-evolution-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-harness-evolution-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the harness-evolution post with structural search formalism, meta-harness lifecycle, frontier and gate protocol, worktree isolation, proxy-metric failure modes, maxTurns=0 multi-agent placement, and local Tangle package mapping.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the harness-evolution post by adding a prompt/skill/runtime/harness comparison table, tightening the Tangle package export mapping, and clarifying the local source-version versus dependency-version boundary.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro'When the prompt keeps asking for a capability the runtime cannot express, the next improvement is not a better sentence. It is a different machine.That is the harness-evolution moment. Prompt optimizers can discover better wording, examples, instructions, rubrics, and sometimes better high-level tactics. Skill optimizers can discover reusable procedures. Runtime topology can change how many workers act, who reviews them, and what gets selected.Harness evolution goes one layer higher: it changes the code that defines the agent's reachable behavior.That code might be a planner contract, a driver, a verifier, a budget policy, a benchmark adapter, a trace schema, a replay layer, a selector, a persona manifest, a tool router, or a worktree candidate lifecycle. The harness is not the model. It is the machine around the model that determines which actions exist, which observations are visible, which branches can run, which artifacts count, and which candidate is allowed to become production.So no, GEPA, SkillOpt, AlphaEvolve-style code search, and meta-harness are not all "doing the same thing" in the strong sense. They share an outer loop:<Steps layout="flow" items={[{title: "Propose candidate"}, {title: "Run candidate"}, {title: "Measure candidate"}, {title: "Select survivor"}, {title: "Repeat"}]} />They differ in the mutable surface. That distinction is everything.| Optimizer family | Mutable candidate | Reachable change | Hard limit ||---|---|---|---|| GEPA, MIPRO, DSPy, AxLLM-style prompt search | prompts, demos, instructions, signatures, rubrics | better policy text inside a fixed runtime | cannot add actions the runtime cannot execute || Skill optimization | durable procedures and reusable task policies | better decomposition, tool habits, repair routines | cannot guarantee orchestration unless the runtime invokes the skill || Runtime topology search | driver, fanout, reviewer, selector, budget, turn policy | different execution graph for the same task | cannot safely promote itself without an external gate || Meta-harness and code evolution | source code around runtime, eval, traces, and candidate lifecycle | new action spaces, verifiers, adapters, and promotion protocols | can overfit or capture the evaluator if the outer gate is weak |## The Reachable SetLet a system have a mutable surface $s$.The surface might be:---title: 'When The Model Itself Is Mutable'description: 'How SFT, RLHF, process supervision, tool-use RL, and Microsoft Frontier Tuning differ from public prompt, skill, and harness loops.'date: 2026-06-05tags: ['ai', 'agents', 'models', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-57-880Z-gpt-6-luna-self-improving-stack-post-training-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-post-training-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-post-training-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-post-training-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-post-training-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-post-training-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the post-training article with SFT/RLHF/RLAIF/DPO/process-supervision objectives, verifiable reward, Frontier Tuning placement, external-state versus weight-loop boundaries, data-governance risks, and Tangle RL bridge mapping.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the post-training article by adding PPO and GRPO mechanics, clarifying adapter deltas versus full-weight updates, separating distillation from self-improvement, and tightening the training-boundary language.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro'Most self-improving agent systems that product teams can actually ship do not change model weights. They change prompts, skills, tools, traces, memory, runtime topology, harness code, and promotion gates. That is external-state self-improvement: legible, reversible, and usually cheap enough to iterate.Post-training changes the model itself, which means it changes the power level and the burden of proof.The loop looks familiar:<Steps layout="flow" items={[{title: "Collect behavior"}, {title: "Score behavior"}, {title: "Construct training signal"}, {title: "Update candidate"}, {title: "Evaluate candidate"}, {title: "Promote or reject"}]} />But the mutable surface is no longer a prompt file or a worktree. It is $\theta$, the model parameters, or some parameterized adapter attached to the model.Once $\theta$ moves, the boundary changes. The behavior becomes harder to inspect, harder to patch locally, harder to roll back partially, and harder to explain from a single trace. It can also generalize better than any prompt edit when the signal is strong enough.That is why this layer deserves separate treatment.## The Mutable VariableThe previous posts treated the model as mostly fixed:$$y=\operatorname{model}_{\theta}(\text{prompt},\text{tools},\text{memory},\text{trace\_context})$$External optimization changed everything around $\theta$:- prompt- skill- retrieval corpus---title: 'Memory Is Not Automatically Learning'description: 'How episodic memory, knowledge gates, retrieval evals, negative knowledge, and production trace mining fit into self-improving agent systems.'date: 2026-06-05tags: ['agents', 'memory', 'knowledge', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-37-367Z-gpt-6-luna-self-improving-stack-memory-flywheels-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-memory-flywheels-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-memory-flywheels-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-memory-flywheels-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-memory-flywheels-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-memory-flywheels-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the memory and knowledge flywheels article with memory-state formalism, write gates, retrieval ablations, negative knowledge, multi-agent role routing, poisoning defenses, and Tangle agent-knowledge/runtime/eval placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the memory and knowledge article by adding the structured write-candidate schema, scope lattice, admission predicate, memory-versus-skill distinction, and tighter multi-agent and poisoning language.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from "../../components/Steps.astro"Remembering more is not learning.Learning means the next run changes in the right direction.Memory is one way to change the next run without changing model weights. It lets an agent carry evidence, preferences, decisions, failures, and procedures across episodes. That makes memory powerful. It also makes memory dangerous.A bad prompt can ruin one run. A bad memory can keep ruining runs until something expires, contradicts, or deletes it. Persistent state is inherited behavior.So the important question is not "does the agent have memory?"The important question is:- what is allowed to persist,- who can retrieve it,- what evidence supports it,- how it is tested,- and how it is retired?That is the memory flywheel.## The State VariableThe clean way to think about memory is as mutable external state.Let:$$\begin{aligned}---title: 'Self-Improvement Needs A Safety Case'description: 'Why prompt injection, sandbox boundaries, eval poisoning, provenance, compliance, and release gates are core to any real self-improving agent stack.'date: 2026-06-05tags: ['agents', 'security', 'governance', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-37-367Z-gpt-6-luna-self-improving-stack-governance-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-governance-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-governance-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-governance-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-governance-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-governance-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the governance article with safety-case formalism, threat taxonomy, authority and action-policy gates, eval boundary controls, release and incident-response protocols, public framework mapping, and local Tangle package placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the governance article by adding the controls-by-mutable-surface matrix and tightening the series-closing rule that the optimizer cannot own the gate that promotes it.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---A self-improving agent becomes a governance problem the moment its changes persist.Before that, a bad run is a bad run. After that, the system can change what future agents see, what they believe, which branches run, which outputs are selected, which benchmarks matter, which tools are reachable, and which candidate becomes production.That is not just "better AI." It is an optimizer pointed at its own future behavior. If the loop is well-governed, it compounds. If it is poorly governed, it learns the shortest path through the measurement and then teaches that path to the next run.The last layer in the self-improving stack is not another optimizer. It is the safety case.## The Safety CaseA safety case is not a vibe and not a policy PDF.It is a structured claim with evidence:- **Claim:** this system is acceptably safe for this use- **Scope:** under these users, tools, data, budgets, models, and domains- **Evidence:** evals, traces, red-team results, controls, audits, incidents- **Residual risk:** what can still go wrong- **Owner:** who is accountable- **Gate:** what blocks releaseFor a self-improving system, the safety case has to cover the loop, not only the baseline model.The model may be safe in isolation while the agent is unsafe because it has too much authority. The prompt may be harmless while the tool graph is dangerous. The eval may look honest while the harness leaks holdout tasks. The sandbox may be strong while a delegated worker receives credentials it never needed.The unit of governance is the whole trajectory:$\tau$ contains:- task---title: 'Optimization Theory For Agent Builders'description: 'A compact map from hill climbing and Bayesian search to GEPA, SkillOpt, Frontier Tuning, agent runtimes, and noisy promotion gates.'date: 2026-06-05tags: ['agents', 'evals', 'optimization', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-05 - model: 'gpt-5.5' role: 'polish' date: 2026-06-05 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-57-880Z-gpt-6-luna-self-improving-stack-optimization-theory-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-optimization-theory-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-optimization-theory-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-optimization-theory-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-optimization-theory-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-optimization-theory-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-05 model: 'gpt-5.5' role: 'draft' note: 'Converted the optimization theory outline into a full draft with math, history, optimizer layers, multi-agent caveats, failure modes, and references.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-05 model: 'gpt-5.5' role: 'polish' note: 'Sharpened voice, clarified the same-skeleton/different-layer thesis, reduced generic transitions, and tightened the multi-agent caveat.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro'Every self-improving agent pitch eventually reduces to three questions:- what can change?- what gets scored?- what is allowed to ship?GEPA evolves prompts. MIPRO searches instructions and demonstrations. Ax brings those ideas into a TypeScript agent framework. SkillOpt trains a skill file as if it were an external parameter of a frozen agent. AlphaEvolve mutates code and keeps versions that pass executable tests. Microsoft describes MAI and Frontier Tuning as a hill-climbing machine built around reinforcement learning environments, workflow traces, and model/runtime adaptation.The temptation is to collapse all of that into one phrase: hill climbing.That is half right and half useless.Yes, they share a skeleton:<Steps layout="flow" items={[{title: "Candidate"}, {title: "Roll out"}, {title: "Score"}, {title: "Compare"}, {title: "Keep, reject, or mutate"}]} />No, they are not the same system. They touch different artifacts, trust different evaluators, and carry different safety risks.Self-improvement is search under a budget, with a noisy objective, over a chosen surface. The surface is the part people keep hand-waving away.The better question is not "is this hill climbing?" It is:> What surface is mutable, what score is trusted, and what gate decides promotion?Answer that and GEPA, DSPy, Ax, SkillOpt, meta-harnesses, agent runtimes, and frontier tuning stop looking like disconnected inventions. They become points in the same design space.## The Shared Shape# 01 Optimization TheoryPost: `src/content/posts/self-improving-stack-optimization-theory.mdx`Status: outline checkpointLast updated: 2026-06-05Supporting trace: `2026-06-05T12-08-35-196Z-gpt-5.5`## Core QuestionWhat optimization concepts do prompt, skill, agent, and harness systems borrowwhen they cannot use gradients over model weights?## What To Master- Hill climbing, local search, evolutionary search, and population methods.- Bayesian optimization and bandits for expensive black-box objectives.- Multi-objective optimization and Pareto frontiers.- Credit assignment across prompts, tools, topology, model choice, and memory.- Exploration versus exploitation under limited eval budget.- Noise, variance, confidence intervals, and false promotion.## Table Of Contents1. Cold open: the same search loop wearing different clothes.2. Core abstraction: mutable surface, objective, operator, gate.3. Short history: local search, evolutionary methods, Bayesian optimization, bandits, AutoML/NAS, prompt optimization.4. Why agent optimization is usually black-box optimization.5. Layer map: prompt/program, skill, runtime topology, code, model behavior.6. GEPA, MIPRO, Ax, SkillOpt, MAI, agent-runtime, agent-eval in one ladder.7. Math of promotion: paired deltas, confidence bounds, Pareto dominance.8. Compute-matched baselines: random@k, best-of-N, human edit, stronger model.9. Why multi-agent flows break naive prompt optimization.10. Failure modes: Goodhart, leakage, drift, credit assignment, cost blindness.11. Practical selection rule: optimize the lowest layer that explains the error.12. Ending: the optimizer is not magic; the gate is the governance.## Core Math```texts = mutable surface: prompt, skill, code, topology, weightsD_train = task distribution used to searchD_holdout = task distribution used to decide promotionJ(s) = E_{tau ~ D}[R(run(s, tau))] - lambda * C(s)s' = O(s, traces, feedback, budget)promote = CI_low(J_holdout(s') - J_holdout(s_base)) > epsilonEI(x) = E[max(f(x) - f_best, 0)]```## Layer Map| Layer | Mutable surface | Examples | Caveat || --- | --- | --- | --- || Prompt/program | Instructions, demos, signatures | DSPy MIPROv2, GEPA, AxGEPA/AxMiPRO | Overfits eval phrasing || Skill | Persistent procedural document | SkillOpt, Codex/Claude skills | Can encode brittle or poisoned habits || Runtime topology | Agents, turns, fanout, routing | agent-runtime, agent-eval, meta-harness | Requires runtime-aware search || Code/artifact | Source code, algorithms | AlphaEvolve, OpenEvolve-style loops | Needs real executable gates || Model behavior | Weights, embeddings, runtime policy | Microsoft Frontier Tuning / MAI | Stronger governance and privacy boundary |## Connects To- GEPA: reflective evolutionary search over text.- MIPRO: Bayesian search over instructions and demos.- SkillOpt: bounded textual updates to a procedural artifact.- meta-harness: architectural variants on a Pareto frontier.- agent-eval: held-out gates and confidence-aware promotion.- agent-runtime: runtime topology, fanout, turns, tools, and delegation as a searchable surface.- Microsoft MAI / Frontier Tuning: reinforcement learning environments that tune model/runtime behavior inside a compliance boundary.## Source Trail- Microsoft MAI hill-climbing machine: https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/- Microsoft Frontier Tuning: https://devblogs.microsoft.com/microsoft365dev/frontier-tuning-teaching-ai-to-work-the-way-you-do/- DSPy optimizer docs: https://github.com/stanfordnlp/dspy/blob/main/docs/docs/learn/optimization/optimizers.md- DSPy MIPROv2 docs: https://github.com/stanfordnlp/dspy/blob/main/docs/docs/api/optimizers/MIPROv2.md- AxLLM optimization guide: https://axllm.dev/optimize/- APE: https://arxiv.org/abs/2211.01910- OPRO: https://arxiv.org/abs/2309.03409# 11 Memory And Knowledge FlywheelsPost: `src/content/posts/self-improving-stack-memory-flywheels.mdx`Status: draftedLast updated: 2026-06-06Supporting trace: `2026-06-05T12-08-35-196Z-gpt-5.5`## Core QuestionWhat persists across agent episodes, and how do we know persistence ishelping rather than poisoning future runs?## Core ClaimMemory is not automatically learning. Memory becomes learning only when afuture run changes in the right direction because a gated persistent state waswritten, retrieved, used, and evaluated.Formal loop:```textM_t = memory state before episode ttau_t = full trace from episode tu_t = proposed memory write after episode tG_mem = memory write gateM_{t+1} = Apply(M_t, u_t) if G_mem(u_t, tau_t, policy) passes M_t otherwise```At inference:```textc_t = Retrieve(M_t, q_t, k, policy)y_t = pi_theta(x_t, c_t, tools)```Evaluation target:```textDelta_memory = E[Score(pi_theta with M)] - E[Score(pi_theta without M)]```The paired ablation is mandatory. Retrieval activity is not evidence oflearning.Write candidate schema:```textu_t = kind claim_or_procedure evidence_refs scope confidence sensitivity freshness_policy retrieval_policy```## Map Items### Episodic Memory- Mutable surface: trace, episode summary, decision record.- Objective: preserve causally useful events across runs.- Evaluator: replay, outcome comparison, analyst finding accuracy.- Promotion gate: trace integrity, summary faithfulness, task lift.- Failure mode: compressed summary loses the mechanism that caused success or failure.### Semantic Memory- Mutable surface: source-grounded claim, wiki page, relation, source anchor.- Objective: make stable facts available without retraining the model.- Evaluator: citation coverage, contradiction checks, freshness checks, retrieval-conditioned answer lift.- Promotion gate: source count, confidence, allowed scope, freshness, lint.