When The Model Itself Is Mutable
Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
- Created
- Updated
25
Turns
25
Tool calls
7
Files touched
10m
Duration
Files
src/components/Steps.astrosrc/content/posts/self-improving-stack-multi-agent-coordination.mdx../../components/Steps.astrosrc/content/posts/self-improving-stack-harness-evolution.mdxsrc/content/posts/self-improving-stack-optimization-theory.mdxsrc/content/posts/self-improving-stack-post-training.mdxsrc/content/posts/self-improving-stack-test-time-compute.mdx
Commit
99791a3 fix(blog): replace prose code blocks with compact semantic markup +525 −737 across 6 files
Conversation
25 turns. Full text where captured; older traces show only the first ~280 chars.
- assistant #1 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && pwd && git status --short --branch && git branch --show-current && ls'", max_output_tokens:2000}); text(r.output);
-
- assistant #2 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && realpath AGENTS.md && sed -n \"1,240p\" AGENTS.md && printf \"\\n---VOICE---\\n\" && sed -n \"1,240p\" VOICE.md && printf \"\\n---DESIGN---\\n\" && sed -n \"1,240p\" DESIGN.md && printf \"\\n---FILES---\\n\" && rg --files -g \"agent-work.md\" -g \"*polish*\" /home/drew/code | head -50'", max_output_tokens:8000}); text(r.output);
-
- assistant #3 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && pnpm blog write self-improving-stack-harness-evolution --harness=codex --role=polish'", yield_time_ms:10000, max_output_tokens:3000}); text(r.output);
-
- assistant #4 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && cat SECONDARY_VOICE.md && for slug in self-improving-stack-multi-agent-coordination self-improving-stack-optimization-theory self-improving-stack-post-training self-improving-stack-test-time-compute; do printf \"--- %s ---\\n\" \"$slug\"; pnpm blog write \"$slug\" --harness=codex --role=polish | sed -n \"/Trace marker:/p;/Start your first visible/p;/Include exactly this marker/p;/At the end, run:/p;/BLOG_TRACE_POSTS=/p\"; done'", yield_time_ms:10000, max_output_tokens:4000}); text(r.output);
-
- assistant #5 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nslugs = [\n'self-improving-stack-harness-evolution',\n'self-improving-stack-multi-agent-coordination',\n'self-improving-stack-optimization-theory',\n'self-improving-stack-post-training',\n'self-improving-stack-test-time-compute',\n]\nfor slug in slugs:\n p=Path('src/content/posts')/f'{slug}.mdx'\n lines=p.read_text().splitlines()\n print(f'\\n=== {slug} original={any(x.strip()==\"original: true\" for x in lines[:30])} ===')\n i=0\n while i<len(lines…
-
- assistant #6 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nimport re\nslugs = ['self-improving-stack-harness-evolution','self-improving-stack-multi-agent-coordination','self-improving-stack-optimization-theory','self-improving-stack-post-training','self-improving-stack-test-time-compute']\nfor slug in slugs:\n lines=(Path('src/content/posts')/f'{slug}.mdx').read_text().splitlines()\n print(f'\\n=== {slug} ===')\n i=0\n while i < len(lines):\n m=re.match(r'^ {0,3}(`{3,}|~{3,})(.*)$', lines[i])\n if not m: i+…
-
- assistant #7 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nimport re\nslug='self-improving-stack-harness-evolution'\nlines=(Path('src/content/posts')/f'{slug}.mdx').read_text().splitlines()\ni=0\nwhile i<len(lines):\n m=re.match(r'^ {0,3}(`{3,}|~{3,})(.*)$', lines[i])\n if not m: i+=1; continue\n mark=m.group(1); info=m.group(2).strip(); j=i+1\n while j<len(lines):\n c=re.match(r'^ {0,3}(`+|~+)[ \\t]*$',lines[j])\n if c and c.group(1)[0]==mark[0] and len(c.group(1))>=len(mark): break\n j+=1\n body=lines[i+1…
-
- assistant #8 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nimport re\nslug='self-improving-stack-multi-agent-coordination'\nlines=(Path('src/content/posts')/f'{slug}.mdx').read_text().splitlines()\ni=0\nwhile i<len(lines):\n m=re.match(r'^ {0,3}(`{3,}|~{3,})(.*)$',lines[i])\n if not m: i+=1; continue\n mark=m.group(1); info=m.group(2).strip(); j=i+1\n while j<len(lines):\n c=re.match(r'^ {0,3}(`+|~+)[ \\t]*$',lines[j])\n if c and c.group(1)[0]==mark[0] and len(c.group(1))>=len(mark): break\n j+=1\n if info=…
-
- assistant #9 1 tool
- exec
src/components/Steps.astroinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && sed -n \"1,240p\" src/components/Steps.astro && sed -n \"1,120p\" src/content/posts/self-improving-stack-multi-agent-coordination.mdx'", max_output_tokens:4000}); text(r.output);
-
- assistant #10 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nimport re\nfor slug in ['self-improving-stack-optimization-theory','self-improving-stack-post-training','self-improving-stack-test-time-compute']:\n lines=(Path('src/content/posts')/f'{slug}.mdx').read_text().splitlines(); print(f'\\n=== {slug} ==='); i=0\n while i<len(lines):\n m=re.match(r'^ {0,3}(`{3,}|~{3,})(.*)$',lines[i])\n if not m: i+=1; continue\n mark=m.group(1); info=m.group(2).strip(); j=i+1\n while j<len(lines):\n c=re.match(r'^ {0,3}(`…
-
- assistant #11 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nimport re\nfor slug in ['self-improving-stack-harness-evolution','self-improving-stack-multi-agent-coordination','self-improving-stack-optimization-theory','self-improving-stack-post-training','self-improving-stack-test-time-compute']:\n lines=(Path('src/content/posts')/f'{slug}.mdx').read_text().splitlines(); i=0\n print(f'\\n=== {slug} ===')\n while i<len(lines):\n m=re.match(r'^ {0,3}(`{3,}|~{3,})(.*)$',lines[i])\n if not m: i+=1; continue\n mark…
-
- assistant #12 1 tool
- exec
../../components/Steps.astroinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nimport re, json\n\nbase=Path('src/content/posts')\nslugs=['self-improving-stack-harness-evolution','self-improving-stack-multi-agent-coordination','self-improving-stack-optimization-theory','self-improving-stack-post-training','self-improving-stack-test-time-compute']\n\ndef key(*lines): return tuple(lines)\nflows={\n'self-improving-stack-harness-evolution': {\nkey('propose candidate','run candidate','measure candidate','select survivor','repeat'): …
-
- assistant #13 1 tool
- exec
src/content/posts/self-improving-stack-harness-evolution.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && git diff --check -- src/content/posts/self-improving-stack-harness-evolution.mdx src/content/posts/self-improving-stack-multi-agent-coordination.mdx src/content/posts/self-improving-stack-optimization-theory.mdx src/content/posts/self-improving-stack-post-training.mdx src/content/posts/self-improving-stack-test-time-compute.mdx && git diff --stat -- src/content/posts/self-improving-stack-harness-evolution.mdx src/content/posts/self-improving-stack-multi-agent-coordination.mdx src/content/posts…
-
- assistant #14 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && rg -n \"promote|texttt|allocation\\\\\\\\_policy|operatorname|<Steps layout\" src/content/posts/self-improving-stack-{harness-evolution,multi-agent-coordination,optimization-theory,post-training,test-time-compute}.mdx | head -120'", max_output_tokens:6000}); text(r.output);
-
- assistant #15 1 tool
- exec
src/content/posts/self-improving-stack-multi-agent-coordination.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\np=Path('src/content/posts/self-improving-stack-multi-agent-coordination.mdx')\ns=p.read_text()\ns=s.replace('{title: \"Supervisor assigns\"}, {title: \"Collect\"}', '{title: \"Supervisor\"}, {title: \"Assign\"}, {title: \"Collect\"}')\ns=s.replace('{title: \"Activate specialist\"}, {title: \"Hand off again if needed\"}', '{title: \"Active specialist\"}, {title: \"Maybe hand off again\"}')\ns=s.replace('{title: \"Workers contribute to shared artifact…
-
- assistant #16 1 tool
- exec
src/content/posts/self-improving-stack-multi-agent-coordination.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && sed -n \"438,449p\" src/content/posts/self-improving-stack-multi-agent-coordination.mdx && sed -n \"410,456p\" src/content/posts/self-improving-stack-test-time-compute.mdx'", max_output_tokens:3000}); text(r.output);
-
- assistant #17 1 tool
- exec
src/content/posts/self-improving-stack-multi-agent-coordination.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\np=Path('src/content/posts/self-improving-stack-multi-agent-coordination.mdx')\ns=p.read_text()\nold=next(line for line in s.splitlines() if line.startswith(r'$$\\begin{aligned}\\operatorname{promote}(s_{\\text{multi}})'))\nnew=r'''$$\\begin{aligned}\\operatorname{promote}(s_{\\text{multi}})&\\text{ if: }\\operatorname{LCB}_{95}(\\operatorname{median}(\\text{score}_{\\text{multi}}-\\text{score}_{\\text{baseline}}\\text{ on holdout}))>\\epsilon\\\\&\\…
-
- assistant #18 1 tool
- exec
src/content/posts/self-improving-stack-multi-agent-coordination.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\np=Path('src/content/posts/self-improving-stack-multi-agent-coordination.mdx')\ns=p.read_text()\nold=next(line for line in s.splitlines() if line.startswith(r'$$\\begin{aligned}\\operatorname{promote}(s_{\\text{multi}})'))\nnew=r'''$$\\begin{aligned}\\operatorname{promote}(s_{\\text{multi}})&\\text{ if: }\\operatorname{LCB}_{95}(\\operatorname{median}(\\text{score}_{\\text{multi}}-\\text{score}_{\\text{baseline}}\\text{ on holdout}))>\\epsilon\\\\&\\…
-
- assistant #19 1 tool
- exec
src/content/posts/self-improving-stack-multi-agent-coordination.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && sed -n \"444,447p\" src/content/posts/self-improving-stack-multi-agent-coordination.mdx && sed -n \"432,454p\" src/content/posts/self-improving-stack-test-time-compute.mdx'", max_output_tokens:3000}); text(r.output);
-
- assistant #20 1 tool
- exec
src/content/posts/self-improving-stack-harness-evolution.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\np=Path('src/content/posts/self-improving-stack-harness-evolution.mdx')\ns=p.read_text()\nold='''```text\nagent-runtime:\n execute workflows\n express drivers\n run fanout/refine loops\n declare mutable agent surfaces\n create and finalize candidate worktrees\n\nagent-eval:\n capture traces\n run campaigns\n score profile cells\n analyze failures\n verify artifacts\n maintain frontiers\n gate promotion\n```'''\nnew='''- **agent-runtime**\n - execute …
-
- assistant #21 1 tool
- exec
src/content/posts/self-improving-stack-harness-evolution.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nimport re\nslugs=['self-improving-stack-harness-evolution','self-improving-stack-multi-agent-coordination','self-improving-stack-optimization-theory','self-improving-stack-post-training','self-improving-stack-test-time-compute']\nfor slug in slugs:\n lines=(Path('src/content/posts')/f'{slug}.mdx').read_text().splitlines(); i=0; kept=[]\n while i<len(lines):\n m=re.match(r'^ {0,3}(`{3,}|~{3,})(.*)$',lines[i])\n if not m: i+=1; continue\n mark=m.group…
-
- assistant #22 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nimport re\nslugs=['self-improving-stack-harness-evolution','self-improving-stack-multi-agent-coordination','self-improving-stack-optimization-theory','self-improving-stack-post-training','self-improving-stack-test-time-compute']\nfor slug in slugs:\n lines=(Path('src/content/posts')/f'{slug}.mdx').read_text().splitlines(); i=0; kept=[]\n while i<len(lines):\n m=re.match(r'^ {0,3}(`{3,}|~{3,})(.*)$',lines[i])\n if not m: i+=1; continue\n mark=m.group…
-
- assistant #23 1 tool
- exec
src/content/posts/self-improving-stack-harness-evolution.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && git diff --check -- src/content/posts/self-improving-stack-harness-evolution.mdx src/content/posts/self-improving-stack-multi-agent-coordination.mdx src/content/posts/self-improving-stack-optimization-theory.mdx src/content/posts/self-improving-stack-post-training.mdx src/content/posts/self-improving-stack-test-time-compute.mdx && rg -c \"<Steps layout=\\\"flow\\\"\" src/content/posts/self-improving-stack-{harness-evolution,multi-agent-coordination,optimization-theory,post-training,test-time-c…
-
- assistant #24 1 tool
- exec
../../components/Steps.astroinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cat > /tmp/blog-markup-cleanup-receipt.txt' <<'EOF'\nFocused blog markup cleanup receipt\nDate: 2026-10-02\nCheckout: /home/drew/code/drewstone.github.io\nBranch: fix/article-text-blocks\nScope: five requested self-improving-stack posts only\n\nMarkup conversions (ordinary text fences to Markdown, quote, KaTeX, or Steps flow):\n- self-improving-stack-harness-evolution: 29\n- self-improving-stack-multi-agent-coordination: 21\n- self-improving-stack-optimization-theory: 11\n- self-improving-stack-post-training: 27\n- self-improving-stac…
-
- assistant #25 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cat /tmp/blog-markup-cleanup-receipt.txt'", max_output_tokens:3000}); text(r.output);
-
Diff
Per-file changes from 99791a3.
diff --git a/src/components/Steps.astro b/src/components/Steps.astroindex db21f70..09942a0 100644--- a/src/components/Steps.astro+++ b/src/components/Steps.astro@@ -16,15 +16,16 @@ interface Step { interface Props { items: Step[]+ layout?: 'steps' | 'flow' } -const { items } = Astro.props+const { items, layout = 'steps' } = Astro.props --- -<ol class="steps">+<ol class:list={['steps', { 'steps-flow': layout === 'flow' }]} role="list"> {items.map((s, i) => ( <li class="step">- <span class="step-num">{String(i + 1).padStart(2, '0')}</span>+ {layout === 'steps' && <span class="step-num">{String(i + 1).padStart(2, '0')}</span>} <div class="step-content"> <div class="step-head"> <span class="step-title">{s.title}</span>@@ -115,4 +116,34 @@ const { items } = Astro.props line-height: 1.55; color: var(--fg-muted); }+ .steps-flow {+ display: flex;+ flex-wrap: wrap;+ gap: 0.45rem 1.2rem;+ margin: var(--space-md) 0;+ padding: 0;+ }+ .steps-flow::before { display: none; }+ .steps-flow .step {+ display: flex;+ align-items: baseline;+ gap: 1.2rem;+ padding: 0;+ min-width: 0;+ }+ .steps-flow .step:not(:last-child)::after {+ content: '→';+ color: var(--fg-muted);+ }+ .steps-flow .step-content { padding: 0; }+ .steps-flow .step-head { margin: 0; }+ .steps-flow .step-title {+ font-size: 1rem;+ font-weight: 400;+ line-height: 1.4;+ }+ @media (max-width: 600px) {+ .steps-flow { gap: 0.3rem 0.75rem; }+ .steps-flow .step { gap: 0.75rem; }+ } </style>diff --git a/src/content/posts/self-improving-stack-multi-agent-coordination.mdx b/src/content/posts/self-improving-stack-multi-agent-coordination.mdxindex 69b6dd4..77fddee 100644--- a/src/content/posts/self-improving-stack-multi-agent-coordination.mdx+++ b/src/content/posts/self-improving-stack-multi-agent-coordination.mdx@@ -47,6 +47,9 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'++ I can give five agents five names and still have one mind making one mistake. "Researcher," "critic," "architect," "driver," and "supervisor" are not coordination by themselves. They can be the same model, with the same blind spot, reading the same context, under the same budget, producing five paraphrases of the same failure. The cast list changed. The information structure did not.@@ -108,27 +111,23 @@ That distinction prevents a common category error: evaluating a better persona a A persona can say: -```text-You are a careful supervisor.-You delegate independent work.-You force disagreement before consensus.-You stop when the reviewer passes the artifact.-```+> You are a careful supervisor.+> You delegate independent work.+> You force disagreement before consensus.+> You stop when the reviewer passes the artifact. Those instructions may improve local judgment. They do not grant runtime authority. Authority lives in the structure: -```text-Can this role spawn workers?-Can it choose tools?-Can it read child traces?-Can it cancel branches?-Can it override a reviewer?-Can it spend more budget?-Can it merge artifacts?-Can it promote the result?-```+- Can this role spawn workers?+- Can it choose tools?+- Can it read child traces?+- Can it cancel branches?+- Can it override a reviewer?+- Can it spend more budget?+- Can it merge artifacts?+- Can it promote the result? If the answers are not represented in the runtime, the persona is aspirational. It may be useful text, but it is not a coordination contract. @@ -219,9 +218,7 @@ Mixture-of-Agents made the ensemble structure more layered: agents generate outp The research arc is clear: -```text-sample many -> search branches -> assign roles -> debate -> aggregate -> orchestrate-```+<Steps layout="flow" items={[{title: "Sample many"}, {title: "Search branches"}, {title: "Assign roles"}, {title: "Debate"}, {title: "Aggregate"}, {title: "Orchestrate"}]} /> The open engineering problem is making the orchestration measurable. @@ -233,9 +230,7 @@ The useful patterns are not defined by agent names. They are defined by informat Multiple workers attempt the same task. A selector or verifier picks one. -```text-spawn N -> score each -> return winner-```+<Steps layout="flow" items={[{title: "Spawn N"}, {title: "Score each"}, {title: "Return winner"}]} /> This is strong when outputs are easy to score and independent attempts are cheap. It is weak when scoring is subjective or all workers share the same blind spot. @@ -243,9 +238,7 @@ This is strong when outputs are easy to score and independent attempts are cheap Multiple reasoning paths produce candidate answers. The system chooses the answer supported by the most paths or highest marginal score. -```text-sample paths -> marginalize answers -> choose stable answer-```+<Steps layout="flow" items={[{title: "Sample paths"}, {title: "Marginalize answers"}, {title: "Choose stable answer"}]} /> This helps when the final answer has a stable attractor and errors are diverse. It is less useful for open-ended artifact quality where many incompatible answers can all be plausible. @@ -253,9 +246,7 @@ This helps when the final answer has a stable attractor and errors are diverse. The system expands intermediate states, evaluates partial progress, prunes weak branches, and backtracks. -```text-expand -> evaluate -> select frontier -> continue or backtrack-```+<Steps layout="flow" items={[{title: "Expand"}, {title: "Evaluate"}, {title: "Select frontier"}, {title: "Continue or backtrack"}]} /> This is coordination over thoughts, plans, or artifacts. It requires explicit state and a heuristic good enough to guide search. @@ -263,9 +254,7 @@ This is coordination over thoughts, plans, or artifacts. It requires explicit st Agents expose arguments, counterarguments, and revisions before a final decision. -```text-propose -> critique -> respond -> judge-```+<Steps layout="flow" items={[{title: "Propose"}, {title: "Critique"}, {title: "Respond"}, {title: "Judge"}]} /> Debate is useful when hidden assumptions matter. It fails when agents optimize rhetoric, defer to the strongest voice, or converge before evidence changes. @@ -273,9 +262,7 @@ Debate is useful when hidden assumptions matter. It fails when agents optimize r A central coordinator delegates scoped tasks to workers and keeps authority over the final artifact. -```text-supervisor -> assign -> collect -> merge -> verify-```+<Steps layout="flow" items={[{title: "Supervisor"}, {title: "Assign"}, {title: "Collect"}, {title: "Merge"}, {title: "Verify"}]} /> This is good for work that has clear subdomains. It fails when the supervisor has no real budget, no child trace access, or no merge discipline. @@ -283,9 +270,7 @@ This is good for work that has clear subdomains. It fails when the supervisor ha Control transfers from one agent to another. -```text-triage -> active specialist -> maybe hand off again-```+<Steps layout="flow" items={[{title: "Triage"}, {title: "Active specialist"}, {title: "Maybe hand off again"}]} /> Handoffs are good when the next specialist should own the state and speak directly. They are dangerous when state transfer is implicit or context grows without boundaries. @@ -293,9 +278,7 @@ Handoffs are good when the next specialist should own the state and speak direct Agents write partial results into a shared workspace. Other agents read and improve them. -```text-workers -> shared artifact store -> reviewers -> revised artifact-```+<Steps layout="flow" items={[{title: "Workers"}, {title: "Shared artifact store"}, {title: "Reviewers"}, {title: "Revised artifact"}]} /> This is natural for code, research, planning, and design. It needs locking, provenance, conflict resolution, and traceable authorship. @@ -303,9 +286,7 @@ This is natural for code, research, planning, and design. It needs locking, prov One layer proposes outputs. Later layers aggregate, refine, or route. -```text-proposers -> aggregators -> final selector-```+<Steps layout="flow" items={[{title: "Proposers"}, {title: "Aggregators"}, {title: "Final selector"}]} /> This works when the aggregator can exploit complementary model strengths. It fails when later layers smooth away critical dissent. @@ -355,11 +336,9 @@ The MCP layer is a third shape, not a replacement for either of the first two. ` The separation is important: -```text-runLoop = bounded multi-shot task kernel-conversation = long-horizon participant dialogue-MCP delegation = async specialist work surface-```+- **`runLoop`**: bounded multi-shot task kernel.+- **`conversation`**: long-horizon participant dialogue.+- **MCP delegation**: async specialist work surface. A reliable coordination stack needs the following contracts regardless of which layer hosts them: @@ -402,12 +381,10 @@ In the local `@tangle-network/agent-eval@0.34.1` source, the package is not just For multi-agent coordination, the clean split is: -```text-agent-runtime/conversation decides which participants spoke, under which policy-agent-runtime/loops decides which bounded workers ran-agent-runtime/mcp exposes async specialist delegation-agent-eval decides whether the resulting system was better-```+- agent-runtime/conversation decides which participants spoke, under which policy+- agent-runtime/loops decides which bounded workers ran+- agent-runtime/mcp exposes async specialist delegation+- agent-eval decides whether the resulting system was better Multi-agent coordination without eval is theater. Eval without trace-level runtime evidence is an opinion poll. @@ -455,7 +432,6 @@ Do not ask whether a multi-agent system "feels smarter." Ask whether it beats th Minimum protocol: -```text 1. Define the task distribution and artifact contract. 2. Freeze model set, tools, prompts, skills, dataset, and evaluator where possible. 3. Compare against best single-agent and best-of-N baselines at matched budget.@@ -463,32 +439,23 @@ Minimum protocol: 5. Measure quality, cost, latency, branch failure rate, trace integrity, and human review load. 6. Run ablations: no debate, no shared context, no heterogeneity, deterministic selector. 7. Promote only on held-out lift with acceptable cost, latency, and failure-mode profile.-``` The promotion rule can be written: -```text-promote(s_multi) if:- LCB_95(median(score_multi - score_baseline on holdout)) > epsilon- and median_cost_multi <= cost_ceiling- and median_latency_multi <= latency_ceiling- and trace_integrity == 1- and selector_ablation_delta > 0- and deterministic_failures == 0-```+$$+\begin{aligned}\operatorname{promote}(s_{\text{multi}})&\text{ if: }\operatorname{LCB}_{95}(\operatorname{median}(\text{score}_{\text{multi}}-\text{score}_{\text{baseline}}\text{ on holdout}))>\epsilon\\&\land\operatorname{median\_cost}_{\text{multi}}\le\text{cost\_ceiling}\\&\land\operatorname{median\_latency}_{\text{multi}}\le\text{latency\_ceiling}\\&\land\text{trace\_integrity}=1\\&\land\text{selector\_ablation\_delta}>0\\&\land\text{deterministic\_failures}=0\end{aligned}+$$ The `selector_ablation_delta` term matters. If the multi-agent system still performs the same when the selector is replaced with a trivial rule, the sophisticated coordination may not be doing causal work. For open-ended work, add a disagreement audit: -```text-disagreement_audit:- independent evidence found?- contradictions preserved?- reviewer defects resolved?- final merge cites child lineage?- rejected branches explained?-```+- disagreement_audit:+- independent evidence found?+- contradictions preserved?+- reviewer defects resolved?+- final merge cites child lineage?+- rejected branches explained? Disagreement is useful only when it changes the final decision or improves confidence calibration. @@ -496,43 +463,35 @@ Disagreement is useful only when it changes the final decision or improves confi Prompt optimizers can improve role instructions: -```text-critic prompt-planner prompt-selector rubric-handoff description-reviewer checklist-```+- critic prompt+- planner prompt+- selector rubric+- handoff description+- reviewer checklist Skill optimizers can improve durable role procedure: -```text-how a reviewer inspects a patch-how a researcher triangulates sources-how a coordinator merges conflicting evidence-how an analyst clusters trace failures-```+- how a reviewer inspects a patch+- how a researcher triangulates sources+- how a coordinator merges conflicting evidence+- how an analyst clusters trace failures Runtime topology optimizers can improve execution shape: -```text-fanout width-which roles run in parallel-whether debate happens before or after evidence collection-which selector sees which fields-when branches cancel-how budget is allocated-```+- fanout width+- which roles run in parallel+- whether debate happens before or after evidence collection+- which selector sees which fields+- when branches cancel+- how budget is allocated Harness evolution can change the coordination machine itself: -```text-new driver-new selector implementation-new trace schema-new sandbox isolation model-new promotion gate-```+- new driver+- new selector implementation+- new trace schema+- new sandbox isolation model+- new promotion gate This is where GEPA, MIPRO, SkillOpt, agent-runtime, agent-eval, and meta-harness stop looking like competitors. They operate on different mutable surfaces. @@ -555,13 +514,11 @@ Do not use multiple agents when the only benefit is a richer cast list. The engineering test is simple: -```text-Can the system show why this role existed?-Can it show what information the role had?-Can it show what the role produced?-Can it show how the selector used or rejected that output?-Can it beat a compute-matched single-agent baseline?-```+- Can the system show why this role existed?+- Can it show what information the role had?+- Can it show what the role produced?+- Can it show how the selector used or rejected that output?+- Can it beat a compute-matched single-agent baseline? If not, the coordination is not yet a system property. It is prose. diff --git a/src/content/posts/self-improving-stack-harness-evolution.mdx b/src/content/posts/self-improving-stack-harness-evolution.mdxindex e779d48..fb671a5 100644--- a/src/content/posts/self-improving-stack-harness-evolution.mdx+++ b/src/content/posts/self-improving-stack-harness-evolution.mdx@@ -47,6 +47,9 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'++ When the prompt keeps asking for a capability the runtime cannot express, the next improvement is not a better sentence. It is a different machine. That is the harness-evolution moment. Prompt optimizers can discover better wording, examples, instructions, rubrics, and sometimes better high-level tactics. Skill optimizers can discover reusable procedures. Runtime topology can change how many workers act, who reviews them, and what gets selected.@@ -57,13 +60,7 @@ That code might be a planner contract, a driver, a verifier, a budget policy, a So no, GEPA, SkillOpt, AlphaEvolve-style code search, and meta-harness are not all "doing the same thing" in the strong sense. They share an outer loop: -```text-propose candidate-run candidate-measure candidate-select survivor-repeat-```+<Steps layout="flow" items={[{title: "Propose candidate"}, {title: "Run candidate"}, {title: "Measure candidate"}, {title: "Select survivor"}, {title: "Repeat"}]} /> They differ in the mutable surface. That distinction is everything. @@ -91,9 +88,9 @@ $$ An optimizer has a mutation operator: -```text-M_s(candidate, evidence) -> candidate'-```+$$+\operatorname{M}_s(\text{candidate},\text{evidence})\to\text{candidate}'+$$ The set of systems it can reach after $k$ mutations is: @@ -133,14 +130,9 @@ $$ The promotion rule is not just the objective. It is a gate: -```text-promote(h) iff- quality(h, holdout) > quality(baseline, holdout)- and deterministic_verifiers(h) pass- and trace_integrity(h) passes- and cost(h) is inside budget- and h does not mutate the gate that judged it-```+$$+\begin{aligned}\operatorname{promote}(h)&\iff \operatorname{quality}(h,\text{holdout})>\operatorname{quality}(\text{baseline},\text{holdout})\\&\land \operatorname{deterministic\_verifiers}(h)\text{ pass}\\&\land \operatorname{trace\_integrity}(h)\text{ passes}\\&\land \operatorname{cost}(h)\text{ is inside budget}\\&\land h\text{ does not mutate the gate that judged it}\end{aligned}+$$ That last clause is the dangerous one. @@ -166,13 +158,7 @@ The Darwin Gödel Machine moved the discussion closer to agent harnesses. Instea That is the practical line: -```text-proof-based self-rewrite--> evaluator-driven program search--> LLM-generated code mutation--> archives of self-improving agent harnesses--> production systems with gates, traces, worktrees, and rollback-```+<Steps layout="flow" items={[{title: "Proof-based self-rewrite"}, {title: "Evaluator-driven program search"}, {title: "LLM-generated code mutation"}, {title: "Archives of self-improving agent harnesses"}, {title: "Production systems with gates, traces, worktrees, and rollback"}]} /> The engineering problem is no longer whether code can be searched. It is which parts of the agent system should be mutable, how candidates are isolated, how evidence is preserved, and who prevents the search from learning the wrong gate. @@ -182,32 +168,28 @@ The harness is the code that turns model calls into a system. For an agent, it includes surfaces like: -```text-planner contract-tool routing-memory read and write policy-retrieval policy-driver topology-subagent delegation-supervisor policy-budget ledger-trace emitter-artifact capture-output parser-validator-selector-promotion gate-benchmark adapter-worktree lifecycle-```+- planner contract+- tool routing+- memory read and write policy+- retrieval policy+- driver topology+- subagent delegation+- supervisor policy+- budget ledger+- trace emitter+- artifact capture+- output parser+- validator+- selector+- promotion gate+- benchmark adapter+- worktree lifecycle These are not cosmetic. They define the action space. A prompt can say: -```text-Fan out to three workers, ask one to critique, merge the best answer, and stop after the verifier passes.-```+> Fan out to three workers, ask one to critique, merge the best answer, and stop after the verifier passes. That instruction only works if the runtime exposes fanout, workers, critique, merge, stop, and verifier operations. If the runtime only supports one serial LLM call, the instruction is theater. The model may describe parallelism, but the system did not execute parallelism. @@ -215,17 +197,15 @@ This is why multi-agent optimization cannot be reduced to persona wording. Personas matter. A driver persona that says "act as a strict reviewer" can change behavior. But a real supervisor is more than tone: -```text-observable state-authority to spawn workers-budget allocation-tool access-handoff contract-stop rule-conflict-resolution policy-selection rule-trace obligations-```+- observable state+- authority to spawn workers+- budget allocation+- tool access+- handoff contract+- stop rule+- conflict-resolution policy+- selection rule+- trace obligations If those are not represented in the harness, the optimizer cannot search them as first-class variables. @@ -237,35 +217,23 @@ If an individual worker has no local conversational loop, improvement pressure m The system can still be agentic, but the agency lives in the orchestration layer: -```text-coordinator receives task-coordinator creates worker prompts-workers run bounded episodes-collector parses outputs-verifier scores artifacts-selector chooses candidate-coordinator decides next episode or final answer-```+<Steps layout="flow" items={[{title: "Coordinator receives task"}, {title: "Coordinator creates worker prompts"}, {title: "Workers run bounded episodes"}, {title: "Collector parses outputs"}, {title: "Verifier scores artifacts"}, {title: "Selector chooses candidate"}, {title: "Coordinator decides next episode or final answer"}]} /> GEPA can optimize text inside that flow. It might learn directives like "parallelize independent file reads" or "ask the reviewer to focus on behavioral regressions." But GEPA does not automatically invent a new coordinator unless the coordinator is a mutable candidate representation and the eval rewards the resulting behavior. The rule: -```text If the workflow move is not representable in the candidate, the optimizer cannot select it.-``` So for multi-agent systems, the important question is not "can the prompt mention fanout?" It is: -```text-Can the candidate change the fanout policy?-Can it change the worker mix?-Can it change the supervisor's observable state?-Can it change the verifier?-Can it change the selector?-Can it change the episode boundary?-Can it change how traces and artifacts are passed forward?-```+- Can the candidate change the fanout policy?+- Can it change the worker mix?+- Can it change the supervisor's observable state?+- Can it change the verifier?+- Can it change the selector?+- Can it change the episode boundary?+- Can it change how traces and artifacts are passed forward? That is harness evolution. @@ -277,20 +245,7 @@ It treats the harness as the search object. A good meta-harness loop has the following phases: -```text-discover harness-freeze evals-seed baseline-read traces-propose structural variant-isolate candidate in worktree-smoke test-run full eval-compare against frontier-merge useful lineages-run held-out gate-promote or reject-```+<Steps layout="flow" items={[{title: "Discover harness"}, {title: "Freeze evals"}, {title: "Seed baseline"}, {title: "Read traces"}, {title: "Propose structural variant"}, {title: "Isolate candidate in worktree"}, {title: "Smoke test"}, {title: "Run full eval"}, {title: "Compare against frontier"}, {title: "Merge useful lineages"}, {title: "Run held-out gate"}, {title: "Promote or reject"}]} /> The important word is structural. @@ -298,16 +253,14 @@ Changing $n=8$ to $n=16$ is not harness evolution. Changing a threshold is not h Structural variants change mechanism: -```text-sequential retry -> fanout plus vote-single judge -> deterministic verifier plus semantic judge-summary-only trace -> span tree plus raw provider capture-flat prompt -> declarative persona and tool surfaces-single winner -> Pareto frontier with cost and latency-one agent -> coordinator plus specialist workers-best score -> held-out promotion gate-one code path -> worktree-isolated candidate lifecycle-```+- **sequential retry** → fanout plus vote+- **single judge** → deterministic verifier plus semantic judge+- **summary-only trace** → span tree plus raw provider capture+- **flat prompt** → declarative persona and tool surfaces+- **single winner** → Pareto frontier with cost and latency+- **one agent** → coordinator plus specialist workers+- **best score** → held-out promotion gate+- **one code path** → worktree-isolated candidate lifecycle A meta-harness rejects variants that only tune knobs unless the knob is itself part of a broader mechanism change. The reason is not aesthetic. Knob tuning is cheaper and belongs to ordinary evolution. Meta-harness is expensive because it lets the system rewrite architecture. @@ -317,20 +270,17 @@ Architecture search without a stable baseline is noise wearing a lab coat. At minimum: -```text-baseline_runs >= 3-baseline_value = median(baseline_runs)-spread <= acceptable_noise-```+$$+\begin{aligned}\texttt{baseline\_runs}&\ge 3\\\texttt{baseline\_value}&=\operatorname{median}(\texttt{baseline\_runs})\\\texttt{spread}&\le\texttt{acceptable\_noise}\end{aligned}+$$ If the baseline varies by more than the claimed improvement, the search cannot tell a better harness from a lucky harness. The same applies to candidate variants: -```text-candidate_runs >= 3-candidate_delta = median(candidate_runs - paired_baseline_runs)-```+$$+\begin{aligned}\texttt{candidate\_runs}&\ge 3\\\texttt{candidate\_delta}&=\operatorname{median}(\texttt{candidate\_runs}-\texttt{paired\_baseline\_runs})\end{aligned}+$$ The strongest comparison is paired: @@ -350,12 +300,14 @@ So meta-harness tracks a Pareto frontier. Candidate $a$ dominates candidate $b$ when: -```text-quality_a >= quality_b-cost_a <= cost_b-latency_a <= latency_b-integrity_a >= integrity_b-```+$$+\begin{aligned}+\mathrm{quality}_a &\ge \mathrm{quality}_b \\+\mathrm{cost}_a &\le \mathrm{cost}_b \\+\mathrm{latency}_a &\le \mathrm{latency}_b \\+\mathrm{integrity}_a &\ge \mathrm{integrity}_b+\end{aligned}+$$ with at least one strict improvement. @@ -363,11 +315,9 @@ The frontier is the set of non-dominated candidates. This matters because the next generation may need a lineage merge: -```text-variant A fixes retrieval misses-variant B adds a stronger verifier-variant C combines A and B without inheriting their regressions-```+- variant A fixes retrieval misses+- variant B adds a stronger verifier+- variant C combines A and B without inheriting their regressions Lineage merging is different from picking the current best score. It treats architecture as compositional. The value of a variant is not only its score, but the mechanism it contributes to future candidates. @@ -377,20 +327,18 @@ For harness evolution, candidate isolation is not a workflow nicety. It is part Each candidate carries: -```text-base ref-worktree path-changed files-hypothesis-trace evidence-generation id-parent id-smoke result-eval result-cost ledger-promotion verdict-rollback handle-```+- base ref+- worktree path+- changed files+- hypothesis+- trace evidence+- generation id+- parent id+- smoke result+- eval result+- cost ledger+- promotion verdict+- rollback handle Without isolation, parallel proposers corrupt each other. Without a parent id, lineage is lost. Without a hypothesis, the search cannot learn from failure. Without a rollback handle, promotion is operationally unsafe. @@ -404,27 +352,23 @@ A prompt optimizer can overfit a phrase. A harness optimizer can overfit the ent Examples: -```text-adds a selector that favors judge-friendly wording over correct artifacts-changes the benchmark adapter to drop hard cases-adds retries that hide deterministic failure under higher cost-routes around a verifier instead of satisfying it-improves the aggregate while breaking one high-value persona-creates a worker topology that only works on the search split-reduces latency by skipping trace capture-```+- adds a selector that favors judge-friendly wording over correct artifacts+- changes the benchmark adapter to drop hard cases+- adds retries that hide deterministic failure under higher cost+- routes around a verifier instead of satisfying it+- improves the aggregate while breaking one high-value persona+- creates a worker topology that only works on the search split+- reduces latency by skipping trace capture This is why harness evolution needs outer invariants: -```text-eval definitions are frozen during candidate search-holdout labels are not visible to the candidate-trace capture is mandatory-backend integrity is checked before aggregation-deterministic verifiers run before semantic judges-cost and latency are promotion dimensions-high-value profiles are inspected separately-```+- eval definitions are frozen during candidate search+- holdout labels are not visible to the candidate+- trace capture is mandatory+- backend integrity is checked before aggregation+- deterministic verifiers run before semantic judges+- cost and latency are promotion dimensions+- high-value profiles are inspected separately If the optimizer can edit the gate and then pass the gate, it did not improve the product. It captured the evaluator. @@ -436,83 +380,68 @@ The audited source trees report `@tangle-network/agent-eval` package version `0. `@tangle-network/agent-eval` is the measurement and promotion substrate. The audited local source exposes: -```text-runEvalCampaign-RunRecord-AgentProfileCell-appendScorecard/loadScorecard/diffScorecard-HeldOutGate-assertRealBackend-RawProviderSink-assertRunCaptured-ReplayCache-AnalystRegistry-MultiLayerVerifier-runProductionLoop-runPromptEvolution-runHarnessExperiment-createSandboxCodeMutator-createCompositeMutator-paretoFrontier-```+- `runEvalCampaign`+- `RunRecord`+- `AgentProfileCell`+- `appendScorecard/loadScorecard/diffScorecard`+- `HeldOutGate`+- `assertRealBackend`+- `RawProviderSink`+- `assertRunCaptured`+- `ReplayCache`+- `AnalystRegistry`+- `MultiLayerVerifier`+- `runProductionLoop`+- `runPromptEvolution`+- `runHarnessExperiment`+- `createSandboxCodeMutator`+- `createCompositeMutator`+- `paretoFrontier` That package is where the evaluator, trace, scorecard, analyst, frontier, and gate live. `@tangle-network/agent-runtime` is the execution and candidate-lifecycle substrate. The audited local source exposes: -```text-runLoop-createRefineDriver-createFanoutVoteDriver-LoopTraceEvent-defineAgent-AgentSurfaces-improvementDriver-reflectiveGenerator-agenticGenerator-MCP delegation tools-analyst loop-OTLP export-```+- `runLoop`+- `createRefineDriver`+- `createFanoutVoteDriver`+- `LoopTraceEvent`+- `defineAgent`+- `AgentSurfaces`+- `improvementDriver`+- `reflectiveGenerator`+- `agenticGenerator`+- MCP delegation tools+- analyst loop+- OTLP export The cleanup matters. The runtime improvement surface now has one driver that owns the candidate lifecycle: -```text-create worktree-generate candidate-finalize or discard-repeat for population size-return CodeSurface-```+<Steps layout="flow" items={[{title: "Create worktree"}, {title: "Generate candidate"}, {title: "Finalize or discard"}, {title: "Repeat for population size"}, {title: "Return CodeSurface"}]} /> The generator is the dial: -```text-reflectiveGenerator = cheap patch application from findings-agenticGenerator = coding harness runs inside the candidate worktree-```+- **`reflectiveGenerator`**: cheap patch application from findings.+- **`agenticGenerator`**: coding harness runs inside the candidate worktree. That is a good kernel shape. The lifecycle is centralized, while the candidate producer can vary by cost and depth. The full stack placement is: -```text-agent-runtime:- execute workflows- express drivers- run fanout/refine loops- declare mutable agent surfaces- create and finalize candidate worktrees--agent-eval:- capture traces- run campaigns- score profile cells- analyze failures- verify artifacts- maintain frontiers- gate promotion-```+- **agent-runtime**+ - execute workflows+ - express drivers+ - run fanout/refine loops+ - declare mutable agent surfaces+ - create and finalize candidate worktrees+- **agent-eval**+ - capture traces+ - run campaigns+ - score profile cells+ - analyze failures+ - verify artifacts+ - maintain frontiers+ - gate promotion So meta-harness composes those packages instead of duplicating them. @@ -522,40 +451,34 @@ A practical meta-harness over this stack would use `agent-runtime` to generate a A real harness variant changes at least one of these: -```text-action space-observation space-control flow-candidate representation-verification stack-selection policy-trace ontology-budget policy-promotion policy-rollback path-```+- action space+- observation space+- control flow+- candidate representation+- verification stack+- selection policy+- trace ontology+- budget policy+- promotion policy+- rollback path Examples: -```text-Add raw provider capture and fail-closed replay before judge recalibration.-Replace single worker retry with fanout-vote plus deterministic validator.-Add profile-cell stamping so driver and scorecard identity cannot diverge.-Route analyst findings to declared file surfaces instead of fabricated paths.-Split supervisor persona into authority contract, observation contract, and stop rule.-Add worktree-isolated candidate generation with mandatory discard on failure.-```+- Add raw provider capture and fail-closed replay before judge recalibration.+- Replace single worker retry with fanout-vote plus deterministic validator.+- Add profile-cell stamping so driver and scorecard identity cannot diverge.+- Route analyst findings to declared file surfaces instead of fabricated paths.+- Split supervisor persona into authority contract, observation contract, and stop rule.+- Add worktree-isolated candidate generation with mandatory discard on failure. Non-examples: -```text-Increase population size.-Raise a judge threshold after seeing a candidate.-Add "be rigorous" to the prompt.-Rename a role from reviewer to supervisor.-Let the candidate skip trace capture to reduce latency.-Tune the metric until the candidate wins.-```+- Increase population size.+- Raise a judge threshold after seeing a candidate.+- Add "be rigorous" to the prompt.+- Rename a role from reviewer to supervisor.+- Let the candidate skip trace capture to reduce latency.+- Tune the metric until the candidate wins. The difference is whether the reachable behavior changed. @@ -569,25 +492,21 @@ An optimizer can only select over candidates it can express, run, and evaluate. For example, the instruction "parallelize independent reads" can exist at several levels: -```text-human instruction to a coding agent-prompt directive inside a worker policy-skill file teaching an agent when to fan out-driver topology that actually dispatches concurrent tasks-runtime kernel with maxConcurrency and trace events-meta-harness variant that changes the driver topology-eval gate that rewards equal-quality lower wall time-```+- human instruction to a coding agent+- prompt directive inside a worker policy+- skill file teaching an agent when to fan out+- driver topology that actually dispatches concurrent tasks+- runtime kernel with maxConcurrency and trace events+- meta-harness variant that changes the driver topology+- eval gate that rewards equal-quality lower wall time Those are not equivalent. The higher layers can make the behavior more reliable because they remove dependence on one model remembering one instruction in one context window. The research-level question is representation: -```text-Which workflow dimensions are first-class variables?-Which are merely text?-Which are invisible operator habits?-```+- Which workflow dimensions are first-class variables?+- Which are merely text?+- Which are invisible operator habits? Meta-harness earns its name only when it turns invisible operator habits into explicit mutable surfaces and then tests whether the change generalizes. diff --git a/src/content/posts/self-improving-stack-optimization-theory.mdx b/src/content/posts/self-improving-stack-optimization-theory.mdxindex d1097f0..bb28dc6 100644--- a/src/content/posts/self-improving-stack-optimization-theory.mdx+++ b/src/content/posts/self-improving-stack-optimization-theory.mdx@@ -46,13 +46,14 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'++ Every self-improving agent pitch eventually reduces to three questions: -```text-what can change?-what gets scored?-what is allowed to ship?-```+- what can change?+- what gets scored?+- what is allowed to ship? GEPA evolves prompts. MIPRO searches instructions and demonstrations. Ax brings those ideas into a TypeScript agent framework. SkillOpt trains a skill file as if it were an external parameter of a frozen agent. AlphaEvolve mutates code and keeps versions that pass executable tests. Microsoft describes MAI and Frontier Tuning as a hill-climbing machine built around reinforcement learning environments, workflow traces, and model/runtime adaptation. @@ -62,9 +63,7 @@ That is half right and half useless. Yes, they share a skeleton: -```text-candidate -> rollout -> score -> compare -> keep, reject, or mutate-```+<Steps layout="flow" items={[{title: "Candidate"}, {title: "Roll out"}, {title: "Score"}, {title: "Compare"}, {title: "Keep, reject, or mutate"}]} /> No, they are not the same system. They touch different artifacts, trust different evaluators, and carry different safety risks. @@ -72,9 +71,7 @@ Self-improvement is search under a budget, with a noisy objective, over a chosen The better question is not "is this hill climbing?" It is: -```text-What surface is mutable, what score is trusted, and what gate decides promotion?-```+> What surface is mutable, what score is trusted, and what gate decides promotion? Answer that and GEPA, DSPy, Ax, SkillOpt, meta-harnesses, agent runtimes, and frontier tuning stop looking like disconnected inventions. They become points in the same design space. @@ -111,9 +108,9 @@ $$ The promotion rule separates engineering from theater: -```text-promote(s_new) if CI_low(J_holdout(s_new) - J_holdout(s_base)) > epsilon-```+$$+\operatorname{promote}(s_{\text{new}})\text{ if }\operatorname{CI}_{\text{low}}(J_{\text{holdout}}(s_{\text{new}})-J_{\text{holdout}}(s_{\text{base}}))>\epsilon+$$ Read it as: promote the new artifact only if it beats the baseline on held-out tasks by more than a meaningful margin, after uncertainty is accounted for. @@ -199,14 +196,12 @@ But a multi-agent workflow is not only text. Consider the way a human directs a coding agent: -```text-parallelize independent file reads-inspect the codebase first-do not stop at a plan-run the tests-do not leave sessions running-protect user changes-```+- parallelize independent file reads+- inspect the codebase first+- do not stop at a plan+- run the tests+- do not leave sessions running+- protect user changes Some of that is instruction text. Some of it is runtime capability. "Parallelize independent file reads" only matters if the agent has a tool wrapper that can execute independent calls concurrently. "Run the tests" only matters if the agent can access the repo, install dependencies, execute commands, and read failures. "Do not stop at a plan" only matters if the control loop allows more than one step. A `maxTurns=0` setting is not a personality flaw. It is a control-surface constraint. @@ -218,13 +213,11 @@ But GEPA does not search the full space of possible runtimes unless that space i The practical split is: -```text-Prompt optimizer: improves local behavior expressed in text.-Skill optimizer: improves durable procedure expressed in text.-Runtime optimizer: improves the graph of actions, agents, tools, and budgets.-Harness optimizer: improves the architecture of the system under evaluation.-Model optimizer: changes the policy inside the model itself.-```+- **Prompt optimizer**: improves local behavior expressed in text.+- **Skill optimizer**: improves durable procedure expressed in text.+- **Runtime optimizer**: improves the graph of actions, agents, tools, and budgets.+- **Harness optimizer**: improves the architecture of the system under evaluation.+- **Model optimizer**: changes the policy inside the model itself. That split determines what kind of evidence you need. @@ -245,9 +238,9 @@ If $\operatorname{mean\_delta}$ is positive, the candidate looks better. But "lo So the promotion rule needs uncertainty: -```text-promote if lower_confidence_bound(mean_delta) > epsilon-```+$$+\operatorname{promote}\text{ if }\operatorname{lower\_confidence\_bound}(\operatorname{mean\_delta})>\epsilon+$$ `epsilon` matters. A candidate that improves a benchmark by 0.1 percentage points while increasing cost by 40 percent is not a meaningful improvement for most products. A candidate that improves a rare but high-severity failure mode may be worth it even if the average score barely moves. The margin should reflect the product decision, not only statistical significance. @@ -285,13 +278,11 @@ The same issue appears at inference time. Suppose a new agent workflow performs Optimization claims should be compute-matched whenever possible: -```text-Did the candidate beat random@k?-Did it beat best-of-N?-Did it beat a stronger model at the same cost?-Did it beat the old system with the same retry budget?-Did it win on held-out tasks, not only on the search set?-```+- Did the candidate beat random@k?+- Did it beat best-of-N?+- Did it beat a stronger model at the same cost?+- Did it beat the old system with the same retry budget?+- Did it win on held-out tasks, not only on the search set? Test-time compute is part of the optimizer story. A system can improve by learning a better artifact, by spending more inference budget, or by doing both. Those are different levers. @@ -339,18 +330,16 @@ Good optimization starts by choosing the right coordinate system. Before running an improvement loop, write down: -```text-surface: what can change?-operator: who or what proposes changes?-train tasks: what can the optimizer see?-holdout tasks: what decides promotion?-score: what counts as success?-cost: what gets penalized?-baseline: what is the candidate compared against?-trace: what evidence explains the score?-gate: what prevents false promotion?-rollback: how do we undo a bad promotion?-```+- **surface**: what can change?+- **operator**: who or what proposes changes?+- **train tasks**: what can the optimizer see?+- **holdout tasks**: what decides promotion?+- **score**: what counts as success?+- **cost**: what gets penalized?+- **baseline**: what is the candidate compared against?+- **trace**: what evidence explains the score?+- **gate**: what prevents false promotion?+- **rollback**: how do we undo a bad promotion? If any line is blank, the system is not ready for autonomous improvement. It may still be ready for research. It may be ready for a human-in-the-loop experiment. But it is not ready to let the optimizer promote its own changes. @@ -364,15 +353,11 @@ DSPy and MIPRO made LM programs optimizable. GEPA showed that reflective text ev The through-line is simple: -```text-If you can represent it, vary it, score it, and gate it, you can optimize it.-```+> If you can represent it, vary it, score it, and gate it, you can optimize it. The catch is just as simple: -```text-If you score the wrong thing, you optimize the wrong thing.-```+> If you score the wrong thing, you optimize the wrong thing. For agent builders, that is the whole game. diff --git a/src/content/posts/self-improving-stack-post-training.mdx b/src/content/posts/self-improving-stack-post-training.mdxindex 8afc7ac..20f17de 100644--- a/src/content/posts/self-improving-stack-post-training.mdx+++ b/src/content/posts/self-improving-stack-post-training.mdx@@ -47,20 +47,16 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'++ Most self-improving agent systems that product teams can actually ship do not change model weights. They change prompts, skills, tools, traces, memory, runtime topology, harness code, and promotion gates. That is external-state self-improvement: legible, reversible, and usually cheap enough to iterate. Post-training changes the model itself, which means it changes the power level and the burden of proof. The loop looks familiar: -```text-collect behavior-score behavior-construct training signal-update candidate-evaluate candidate-promote or reject-```+<Steps layout="flow" items={[{title: "Collect behavior"}, {title: "Score behavior"}, {title: "Construct training signal"}, {title: "Update candidate"}, {title: "Evaluate candidate"}, {title: "Promote or reject"}]} /> But the mutable surface is no longer a prompt file or a worktree. It is $\theta$, the model parameters, or some parameterized adapter attached to the model. @@ -72,22 +68,20 @@ That is why this layer deserves separate treatment. The previous posts treated the model as mostly fixed: -```text-y = model_theta(prompt, tools, memory, trace_context)-```+$$+y=\operatorname{model}_{\theta}(\text{prompt},\text{tools},\text{memory},\text{trace\_context})+$$ External optimization changed everything around $\theta$: -```text-prompt-skill-retrieval corpus-tool description-driver topology-verifier-selector-harness-```+- prompt+- skill+- retrieval corpus+- tool description+- driver topology+- verifier+- selector+- harness Post-training changes $\theta$, or a learned delta attached to it: @@ -149,12 +143,10 @@ It is weak when the demonstration only shows the final artifact but not the deci For agents, SFT is often best viewed as initialization: -```text-teach the format-teach the domain language-teach basic tool conventions-teach common interaction patterns-```+- teach the format+- teach the domain language+- teach basic tool conventions+- teach common interaction patterns It is not a full self-improvement loop unless the system keeps collecting new demonstrations, filtering them, validating them, and retraining. @@ -164,19 +156,13 @@ RLHF adds a preference model. The canonical shape is: -```text-1. collect demonstrations-2. train an SFT policy-3. collect preference comparisons-4. train reward model r_phi(x, y)-5. optimize pi_theta against r_phi with a KL penalty-```+<Steps layout="flow" items={[{title: "Collect demonstrations"}, {title: "Train an SFT policy"}, {title: "Collect preference comparisons"}, {title: "Train reward model r_phi(x, y)"}, {title: "Optimize pi_theta against r_phi with a KL penalty"}]} /> Preference data looks like: -```text-(x, y_w, y_l)-```+$$+(x,y_w,y_l)+$$ where $y_w$ is preferred over $y_l$. @@ -235,14 +221,12 @@ If the evaluator has the same blind spot as the policy, the loop can amplify it. For agent systems, RLAIF should be treated like any other judge channel: -```text-pin the judge-calibrate against human labels-test inter-rater reliability-track judge drift-separate judge reward from deterministic verifiers-keep heldout tasks hidden-```+- pin the judge+- calibrate against human labels+- test inter-rater reliability+- track judge drift+- separate judge reward from deterministic verifiers+- keep heldout tasks hidden AI feedback is not automatically objective feedback. @@ -268,10 +252,7 @@ Distillation transfers behavior from one model to another. The usual shape is: -```text-teacher model produces y_teacher-student model trains on (x, y_teacher)-```+<Steps layout="flow" items={[{title: "Teacher model produces y_teacher"}, {title: "Student model trains on (x, y_teacher)"}]} /> That can reduce cost, latency, or deployment size. It can also import the teacher's blind spots, refusals, shortcuts, and style artifacts. @@ -293,9 +274,9 @@ $$ then outcome reward is: -```text-R(tau)-```+$$+R(\tau)+$$ Process reward is: @@ -307,13 +288,11 @@ The 2023 "Let's Verify Step by Step" result made this concrete for math reasonin For agents, process supervision is even more natural. The system already has trace spans: -```text-planner chose wrong tool-retrieval returned stale context-tool argument was invalid-verifier caught missing artifact-retry repeated the same failed action-```+- planner chose wrong tool+- retrieval returned stale context+- tool argument was invalid+- verifier caught missing artifact+- retry repeated the same failed action Those are process labels. @@ -343,21 +322,17 @@ The lesson is not "RL fixes reasoning." The lesson is: -```text-RL becomes much more credible when the reward is verifiable.-```+> RL becomes much more credible when the reward is verifiable. For product agents, that means the best training signals often come from verifiers already used in the harness: -```text-tests-typechecks-schema validation-permission checks-policy checks-live smoke tests-human acceptance events-```+- tests+- typechecks+- schema validation+- permission checks+- policy checks+- live smoke tests+- human acceptance events ## Frontier Tuning @@ -365,11 +340,9 @@ Microsoft's June 2, 2026 Frontier Tuning announcement is the current commercial Microsoft describes Frontier Tuning as applying reinforcement learning inside a customer's compliance boundary using the customer's data, processes, and conventions. Their developer blog says the system has three parts: -```text-managed reinforcement learning environment-customer workflow and domain inputs-tuned output models, skills, and harness-```+- managed reinforcement learning environment+- customer workflow and domain inputs+- tuned output models, skills, and harness It also says the RLE is used for both post-training and inference: during training it learns from workflows, tool usage, and eval signals; at inference it explores multiple frontier and fine-tuned models across turns to find stronger candidate paths before returning an answer. @@ -377,24 +350,20 @@ That is not just fine-tuning a prompt. It is an environment-level adaptation loop: -```text-enterprise workflows -> RLE -> model updates-enterprise traces -> eval signals -> skill and harness updates-enterprise data controls -> compliance boundary -> access-scoped model behavior-```+- **enterprise workflows** → RLE -> model updates+- **enterprise traces** → eval signals -> skill and harness updates+- **enterprise data controls** → compliance boundary -> access-scoped model behavior Microsoft's MAI announcement frames the same direction as a "hill-climbing machine" and explicitly points to traces of real work as valuable training data: steps, decisions, and actions that define how tasks get done inside an organization. That maps directly onto the stack: -```text-traces become training data-evals become reward-tools become environment-skills become reusable policy-harness becomes runtime substrate-weights become another mutable surface-```+- traces become training data+- evals become reward+- tools become environment+- skills become reusable policy+- harness becomes runtime substrate+- weights become another mutable surface The difference is access. Most product teams can ship the first five layers. Frontier labs and large enterprise tuning systems can also move weights. @@ -404,27 +373,23 @@ External-state loops are weaker but auditable. They can edit: -```text-prompt files-skill files-knowledge bases-retrieval corpora-tool manifests-runtime drivers-harness code-eval gates-```+- prompt files+- skill files+- knowledge bases+- retrieval corpora+- tool manifests+- runtime drivers+- harness code+- eval gates They are easy to inspect: -```text-git diff-trace replay-scorecard diff-rollback commit-feature flag-heldout gate-```+- git diff+- trace replay+- scorecard diff+- rollback commit+- feature flag+- heldout gate Weight-level loops are stronger but less locally inspectable. @@ -432,11 +397,9 @@ They can change behavior in ways no prompt diff shows. A model may become better The practical rule: -```text-Keep behavior external when auditability matters more than compression.-Move behavior into weights when the signal is strong, repeated, privacy-safe,-and valuable enough to justify harder inspection.-```+- Keep behavior external when auditability matters more than compression.+- Move behavior into weights when the signal is strong, repeated, privacy-safe,+- and valuable enough to justify harder inspection. Not every successful trace should become a gradient update. @@ -446,31 +409,27 @@ Model training changes the risk surface because training data is not just contex A training-ready record needs: -```text-source provenance-license and ownership-privacy classification-access boundary-split tag-deduplication hash-synthetic-data marker-reward source-verifier version-judge version-contamination status-```+- source provenance+- license and ownership+- privacy classification+- access boundary+- split tag+- deduplication hash+- synthetic-data marker+- reward source+- verifier version+- judge version+- contamination status Generated data is especially tricky. Synthetic examples can help, but recursive training on model outputs can collapse diversity and accumulate errors. The 2024 Nature model-collapse paper shows the failure mode directly: repeated training on generated data can make models forget low-probability events and drift into their own distorted distribution. For self-improving agents, this means: -```text-do not train blindly on your own outputs-do not mix search traces into holdout-do not treat judge rationales as ground truth-do not backpropagate private data outside its boundary-do not let synthetic data lose its label-```+- do not train blindly on your own outputs+- do not mix search traces into holdout+- do not treat judge rationales as ground truth+- do not backpropagate private data outside its boundary+- do not let synthetic data lose its label The post-training layer needs a data firewall, not just a dataset. @@ -482,54 +441,46 @@ They produce the artifacts a training system would need. Local source audit on June 6, 2026: -```text-@tangle-network/agent-eval package source: 0.34.1-@tangle-network/agent-runtime package source: 0.26.0-```+- `@tangle-network/agent-eval package source`: 0.34.1+- `@tangle-network/agent-runtime package source`: 0.26.0 `agent-eval` provides the bridge from eval campaigns to training data: -```text-RunRecord-trialsToRunRecords-verificationReportToRunRecord-extractPreferences-extractVerifiableReward-extractVerifiableRewardsFromRecords-extractStepRewards-prmTrainingPairs-exportRewardModel-off-policy estimators-contamination probes-compute curves-reward-hacking checks-training-data exporters-```+- RunRecord+- trialsToRunRecords+- verificationReportToRunRecord+- extractPreferences+- extractVerifiableReward+- extractVerifiableRewardsFromRecords+- extractStepRewards+- prmTrainingPairs+- exportRewardModel+- off-policy estimators+- contamination probes+- compute curves+- reward-hacking checks+- training-data exporters That is the right boundary. The package is not a training cluster. It converts traces, verifier reports, preferences, and scorecards into clean signals for a downstream trainer. `agent-runtime` provides the execution side: -```text-runLoop-tool and sandbox execution-driver topology-agent surfaces-worktree candidate lifecycle-analyst loop-OTLP export-```+- runLoop+- tool and sandbox execution+- driver topology+- agent surfaces+- worktree candidate lifecycle+- analyst loop+- OTLP export Together they approximate the enterprise RLE shape without moving weights: -```text-runtime executes workflows-eval captures traces and rewards-analysts diagnose failures-external surfaces mutate-gates promote candidates-RL bridge exports training signal-```+- runtime executes workflows+- eval captures traces and rewards+- analysts diagnose failures+- external surfaces mutate+- gates promote candidates+- RL bridge exports training signal The final step, updating $\theta$, remains outside the public product loop unless a real training backend is wired. @@ -541,33 +492,29 @@ A model rollback is a model artifact rollback. That means the promotion gate has to be stricter: -```text-heldout quality lift-profile-cell regression checks-safety regression checks-privacy leak probes-contamination probes-reward-hacking probes-cost and latency checks-calibration checks-domain expert review for high-stakes use-artifact lineage-rollback plan-```+- heldout quality lift+- profile-cell regression checks+- safety regression checks+- privacy leak probes+- contamination probes+- reward-hacking probes+- cost and latency checks+- calibration checks+- domain expert review for high-stakes use+- artifact lineage+- rollback plan The release unit is not a clever answer. It is a model snapshot: -```text-model_id-base_model_id-training_data_manifest-reward_manifest-eval_manifest-policy_manifest-artifact_hash-access_policy-deprecation_plan-```+- `model_id`+- `base_model_id`+- `training_data_manifest`+- `reward_manifest`+- `eval_manifest`+- `policy_manifest`+- `artifact_hash`+- `access_policy`+- `deprecation_plan` Without that lineage, model-level self-improvement is not a controlled system. It is just drift with a training budget. @@ -587,9 +534,7 @@ The difference is reversibility. When the mutable surface is external state, the system remains legible. When the mutable surface is model behavior, the system can become more capable, but the operator owes a stronger data boundary, stronger heldout discipline, stronger artifact lineage, and a clearer answer to one question: -```text Why does this behavior belong in weights instead of in the harness?-``` That is the post-training question. diff --git a/src/content/posts/self-improving-stack-test-time-compute.mdx b/src/content/posts/self-improving-stack-test-time-compute.mdxindex 0da57f8..61d1947 100644--- a/src/content/posts/self-improving-stack-test-time-compute.mdx+++ b/src/content/posts/self-improving-stack-test-time-compute.mdx@@ -47,15 +47,16 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'++ I do not trust a multi-agent system until it beats the boring baseline. More agents is not a strategy. It is a cost increase until it beats blind extra compute. That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity. It spent more budget and mislabeled the budget as architecture. The first gate is simple: -```text-Beat random at equal compute.-```+> Beat random at equal compute. Not beat one greedy sample. Not beat the weakest baseline. Not beat a single run after quietly raising the turn budget. Beat the best simple use of the same budget. @@ -101,19 +102,15 @@ $$ The budget $B$ is not one scalar in practice. It is a vector: -```text-B = {- samples,- tokens,- wall_clock,- model_calls,- tool_calls,- sandbox_minutes,- human_review_minutes,- dollars,- risk_budget-}-```+- samples+- tokens+- wall clock+- model calls+- tool calls+- sandbox minutes+- human review minutes+- dollars+- risk budget A fair comparison fixes the relevant parts of $B$, or it reports the tradeoff instead of pretending the strategy itself improved. @@ -214,13 +211,11 @@ This is where process reward models, unit tests, static analyzers, rubric judges Guided strategies choose where to spend compute: -```text-expand branch-refine branch-spawn another sample-call a verifier-stop early-```+- expand branch+- refine branch+- spawn another sample+- call a verifier+- stop early This includes tree search, branch-and-bound, adaptive sampling, debate, tool-assisted search, and agentic fanout. A guided strategy earns its complexity when it beats the best blind or simple selector baseline under the same budget. @@ -232,15 +227,11 @@ That does not mean repeated sampling solves deployment. Coverage asks: -```text-Did any candidate contain a good answer?-```+> Did any candidate contain a good answer? Selection asks: -```text-Could the system identify that candidate without oracle labels?-```+> Could the system identify that candidate without oracle labels? The two curves can be very different. @@ -271,10 +262,8 @@ The modern reasoning-model era made the same point visible at product scale. Ope The consequence for agent builders is brutal: -```text-If your agent topology beats greedy but loses to simple repeated sampling,-you built an expensive sampler.-```+- If your agent topology beats greedy but loses to simple repeated sampling,+- you built an expensive sampler. That can still be useful. A sampler with a good selector may be exactly what the product needs. But it should be named honestly. @@ -284,33 +273,23 @@ Extra compute has shapes. Parallel sampling: -```text-sample k independent attempts -> select-```+<Steps layout="flow" items={[{title: "Sample k independent attempts"}, {title: "Select"}]} /> Sequential refinement: -```text-attempt -> critique -> revise -> critique -> revise-```+<Steps layout="flow" items={[{title: "Attempt"}, {title: "Critique"}, {title: "Revise"}, {title: "Critique"}, {title: "Revise"}]} /> Tree search: -```text-expand frontier -> score partial states -> allocate next step-```+<Steps layout="flow" items={[{title: "Expand frontier"}, {title: "Score partial states"}, {title: "Allocate next step"}]} /> Debate: -```text-proposal -> criticism -> response -> judge-```+<Steps layout="flow" items={[{title: "Propose"}, {title: "Criticism"}, {title: "Respond"}, {title: "Judge"}]} /> Tool-grounded search: -```text-hypothesis -> tool call -> observation -> update-```+<Steps layout="flow" items={[{title: "Hypothesis"}, {title: "Tool call"}, {title: "Observation"}, {title: "Update"}]} /> None dominates everywhere. @@ -318,28 +297,22 @@ Parallel sampling works when the proposal distribution has enough mass on valid The compute-optimal question is: -```text For this task, model, verifier, and budget, which allocation has highest expected utility?-``` Snell et al. studied this directly for reasoning problems and found that compute-optimal allocation can be much more efficient than plain best-of-N. The important lesson is not one fixed strategy. It is conditional allocation: choose breadth, depth, refinement, or search based on prompt difficulty and verifier behavior. That makes test-time compute a control problem. The controller observes partial evidence: -```text-prompt difficulty estimate-candidate confidence-verifier margin-branch diversity-remaining budget-latency deadline-```+- prompt difficulty estimate+- candidate confidence+- verifier margin+- branch diversity+- remaining budget+- latency deadline and chooses the next action: -```text sample | refine | verify | expand | stop-``` The policy is only good if those observations predict downstream reward. If the difficulty estimate is bad, adaptive compute becomes random budget jitter. If the verifier margin is miscalibrated, early stopping exits on fluent failures. @@ -367,11 +340,9 @@ Unit tests can be incomplete. Typechecks miss semantics. Proof checkers only cov Verifier quality belongs in the objective: -```text-observed_score = V(x, y, trace)-true_score = R(x, y)-verifier_error = observed_score - true_score-```+$$+\begin{aligned}\text{observed\_score}&=V(x,y,\text{trace})\\\text{true\_score}&=R(x,y)\\\text{verifier\_error}&=\text{observed\_score}-\text{true\_score}\end{aligned}+$$ If guided search optimizes $V$ faster than $V$ tracks $R$, the system reward-hacks its own evaluator. @@ -383,26 +354,24 @@ A multi-agent topology is a policy for spending test-time compute. The same budget can be spent as: -```text-8 independent workers-4 workers + 1 verifier-2 workers + 2 rounds of critique-1 worker + 7 refinement turns-1 tree search with frontier size 4 and depth 2-1 tool-heavy agent with expensive environment checks-```+- 8 independent workers+- 4 workers + 1 verifier+- 2 workers + 2 rounds of critique+- 1 worker + 7 refinement turns+- 1 tree search with frontier size 4 and depth 2+- 1 tool-heavy agent with expensive environment checks The topology claim is not: -```text-multi-agent > single-agent-```+$$+\text{multi-agent}>\text{single-agent}+$$ The topology claim is: -```text-allocation_policy_multi(B) > allocation_policy_baseline(B)-```+$$+\operatorname{allocation\_policy}_{\text{multi}}(B)>\operatorname{allocation\_policy}_{\text{baseline}}(B)+$$ at a measured budget $B$. @@ -441,10 +410,8 @@ The refreshed Tangle runtime map makes this concrete. This is the useful split: -```text-runtime spends compute-eval proves whether the spend was worth it-```+- runtime spends compute+- eval proves whether the spend was worth it ## The Equal-Compute Test @@ -465,7 +432,6 @@ budget: Then compare strategies: -```text 1. Single greedy or default run. 2. Random@k under the same model, prompt, and budget. 3. Best-of-N with the deployable selector.@@ -473,48 +439,35 @@ Then compare strategies: 5. Verifier-rerank with the production verifier. 6. Guided topology under the same budget. 7. Adaptive topology with early stopping and budget reallocation.-``` Record: -```text-score-pass_rate-coverage_k-selection_k-selector_loss_k-cost_usd-tokens_in_out-wall_ms-model_calls-tool_calls-branch_failures-trace_integrity-```+- score+- pass_rate+- coverage_k+- selection_k+- selector_loss_k+- cost_usd+- tokens_in_out+- wall_ms+- model_calls+- tool_calls+- branch_failures+- trace_integrity Also report dominance, not just mean score: -```text-strategy_a dominates strategy_b if:- score_a >= score_b- and cost_vector_a <= cost_vector_b componentwise- and at least one inequality is strict-```+$$+\begin{aligned}\text{strategy\_a dominates strategy\_b if: }\text{score\_a}\ge\text{score\_b}\\&\land\text{cost\_vector\_a}\le\text{cost\_vector\_b componentwise}\\&\land\text{at least one inequality is strict}\end{aligned}+$$ A strategy that improves score while increasing cost is not wrong. It is on a tradeoff frontier. A strategy that is worse and more expensive is dead. Promotion should require: -```text-promote(strategy_new) if:- LCB_95(median(score_new - score_baseline on holdout)) > epsilon- and median_cost_new <= cost_ceiling- and median_latency_new <= latency_ceiling- and no baseline Pareto-dominates strategy_new- and trace_integrity == 1- and selector_loss_new <= selector_loss_ceiling- and deterministic_failures == 0-```+$$+\begin{aligned}\operatorname{promote}(\text{strategy\_new})&\text{ if: }\operatorname{LCB}_{95}(\operatorname{median}(\text{score\_new}-\text{score\_baseline on holdout}))>\epsilon\\&\land\text{median\_cost\_new}\le\text{cost\_ceiling}\\&\land\text{median\_latency\_new}\le\text{latency\_ceiling}\\&\land\text{no baseline Pareto-dominates strategy\_new}\\&\land\text{trace\_integrity}=1\\&\land\text{selector\_loss\_new}\le\text{selector\_loss\_ceiling}\\&\land\text{deterministic\_failures}=0\end{aligned}+$$ The baseline should be the strongest simple strategy the product could actually deploy, not a strawman. @@ -578,9 +531,7 @@ Do not use test-time compute when: The serious claim is not "we used more reasoning." It is: -```text-Given the same budget, this allocation policy produced better verified outcomes.-```+> Given the same budget, this allocation policy produced better verified outcomes. That is the first bar for runtime topology, multi-agent coordination, and self-improving harnesses.