GPT-6-luna codex

The Self-Improving Stack

Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.

Created
Updated
27
Turns
27
Tool calls
10
Files touched
11m
Duration

Files

Commit

99791a3 fix(blog): replace prose code blocks with compact semantic markup
+209 −376 across 5 files

Conversation

27 turns. Full text where captured; older traces show only the first ~280 chars.

  1. assistant #1 1 tool
    • exec ../docs/processes/agent-work.md
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && pwd && git status --short --branch && cat AGENTS.md && printf \"\\n--- process ---\\n\" && sed -n \"/^## Delegation/,/^## /p\" ../docs/processes/agent-work.md && printf \"\\n--- guidance ---\\n\" && sed -n \"/^## Use skills and owning guidance/,/^## /p\" ../docs/processes/agent-work.md && printf \"\\n--- VOICE ---\\n\" && cat VOICE.md && printf \"\\n--- DESIGN ---\\n\" && cat DESIGN.md'","yield_time_ms":10000,"max_output_tokens":9000}); text(r.output)
  2. assistant #2 1 tool
    • exec /docs/processes/agent-work.md
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'find /home/drew -path \"*/docs/processes/agent-work.md\" -o -name VOICE.md -o -name DESIGN.md | head -50'","yield_time_ms":10000,"max_output_tokens":2000}); text(r.output)
  3. assistant #3 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && cat VOICE.md && cat DESIGN.md && printf \"\\n--- package ---\\n\" && cat package.json | head -100 && printf \"\\n--- locate process docs ---\\n\" && ls -la /home/drew/code'","yield_time_ms":10000,"max_output_tokens":6500}); text(r.output)
  4. assistant #4 1 tool
    • exec /home/drew/code/drewstone.github.io/AGENTS.md
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'readlink -f /home/drew/code/drewstone.github.io/AGENTS.md && cd /home/drew/code/discovery-lab && sed -n \"/^## Delegation/,/^## /p\" docs/processes/agent-work.md && sed -n \"/^## Use skills and owning guidance/,/^## /p\" docs/processes/agent-work.md'","yield_time_ms":10000,"max_output_tokens":3500}); text(r.output)
  5. assistant #5 1 tool
    • exec /home/drew/code/drewstone.github.io/AGENTS.md
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'ls -l /home/drew/code/drewstone.github.io/AGENTS.md; readlink /home/drew/code/drewstone.github.io/AGENTS.md; ls /home/drew/code/discovery-lab/docs/processes/agent-work.md; rg -n \"Delegation|owning guidance|Use skills|protect the host\" /home/drew/code/discovery-lab/docs/processes/agent-work.md'","yield_time_ms":10000,"max_output_tokens":2000}); text(r.output)
  6. assistant #6 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'ls /home/drew/code/discovery-lab/docs'","yield_time_ms":10000,"max_output_tokens":2000}); text(r.output)
  7. assistant #7 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'ls -la /home/drew/code/discovery-lab/docs/processes && find /home/drew/code/discovery-lab/docs/processes -maxdepth 2 -type f | head'","yield_time_ms":10000,"max_output_tokens":1500}); text(r.output)
  8. assistant #8 1 tool
    • exec /home/drew/code/discovery-lab/docs/processes/agent-work.md
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cat /home/drew/code/discovery-lab/docs/processes/agent-work.md'","yield_time_ms":10000,"max_output_tokens":4500}); text(r.output)
  9. assistant #9 1 tool
    • exec src/content/posts/the-self-improving-stack.mdx
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && for f in src/content/posts/the-self-improving-stack.mdx src/content/posts/self-improving-stack-agent-runtime-topology.mdx src/content/posts/self-improving-stack-prompt-optimization.mdx src/content/posts/self-improving-stack-skill-optimization.mdx; do printf \"--- %s ---\\n\" \"$f\"; sed -n \"1,14p\" \"$f\"; done'","yield_time_ms":10000,"max_output_tokens":2500}); text(r.output)
  10. assistant #10 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 - <<'\"'\"'PY'\"'\"'\nfrom pathlib import Path\nfiles = [\"the-self-improving-stack.mdx\", \"self-improving-stack-agent-runtime-topology.mdx\", \"self-improving-stack-prompt-optimization.mdx\", \"self-improving-stack-skill-optimization.mdx\"]\nfor name in files:\n p = Path(\"src/content/posts\") / name\n lines = p.read_text().splitlines()\n fences=[]; i=0\n while i<len(lines):\n s=lines[i].lstrip()\n if s.startswith(\"```\") or s.startswith(\"~~~\"):\n mark=s[:3]; start=i; lang=s[3:].s…
  11. assistant #11 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && pnpm blog write the-self-improving-stack --harness=codex --role=polish | rg -i \"marker|finish command|pnpm blog finish|voice|anti-pattern\"'","yield_time_ms":10000,"max_output_tokens":3000}); text(r.output)
  12. assistant #12 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && cat SECONDARY_VOICE.md && pnpm blog write self-improving-stack-agent-runtime-topology --harness=codex --role=polish | rg -i \"Trace marker:|Start your first|Include exactly|At the end, run:|Trace boundary:\" && pnpm blog write self-improving-stack-prompt-optimization --harness=codex --role=polish | rg -i \"Trace marker:|Start your first|Include exactly|At the end, run:|Trace boundary:\" && pnpm blog write self-improving-stack-skill-optimization --harness=codex --role=polish | rg -i \"Trace mar…
  13. assistant #13 1 tool
    • exec ../../components/Steps.astro
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 - <<'\"'\"'PY'\"'\"'\nfrom pathlib import Path\n\nfiles = {\n \"the-self-improving-stack.mdx\": {\n \"the system optimizes itself\": \"The system optimizes itself.\",\n \"what user task distribution is being improved?\\nwhat is allowed to change?\\nwhat evidence says it improved that task?\\nhow are candidates generated?\\nwhat gate decides promotion?\\nwhat can go wrong when that layer changes?\": \"- Which user task distribution is being improved?\\n- What is allowed to change?\\n- W…
  14. assistant #14 1 tool
    • exec ../../components/Steps.astro
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nfiles = {\n\"the-self-improving-stack.mdx\": [\n(\"the system optimizes itself\", \"The system optimizes itself.\"),\n(\"what user task distribution is being improved?\\nwhat is allowed to change?\\nwhat evidence says it improved that task?\\nhow are candidates generated?\\nwhat gate decides promotion?\\nwhat can go wrong when that layer changes?\", \"- Which user task distribution is being improved?\\n- What is allowed to change?\\n- What evidence …
  15. assistant #15 1 tool
    • exec ../../components/Steps.astro
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nplans = {\n'the-self-improving-stack.mdx': {\n52: 'The system optimizes itself.',\n66: '- Which user task distribution is being improved?\\n- What is allowed to change?\\n- What evidence shows improvement on that task?\\n- How are candidates generated?\\n- What gate decides promotion?\\n- What can go wrong when that layer changes?',\n81: '<Steps layout=\"flow\" items={[{title:\\'Run\\'}, {title:\\'Observe\\'}, {title:\\'Diagnose\\'}, {title:\\'Propo…
  16. assistant #16 1 tool
    • exec src/content/posts/the-self-improving-stack.mdx
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && git diff -- src/content/posts/the-self-improving-stack.mdx src/content/posts/self-improving-stack-agent-runtime-topology.mdx src/content/posts/self-improving-stack-prompt-optimization.mdx src/content/posts/self-improving-stack-skill-optimization.mdx'","yield_time_ms":10000,"max_output_tokens":14000}); text(r.output)
  17. assistant #17 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nfor name in ['the-self-improving-stack.mdx','self-improving-stack-agent-runtime-topology.mdx','self-improving-stack-prompt-optimization.mdx','self-improving-stack-skill-optimization.mdx']:\n lines=(Path('src/content/posts')/name).read_text().splitlines()\n print('\\n'+name)\n i=0\n while i<len(lines):\n s=lines[i].lstrip()\n if s.startswith('```') or s.startswith('~~~'):\n mark=s[:3]; info=s[3:].strip(); start=i; i+=1; body=[]\n while i<len(lines) a…
  18. assistant #18 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nfor name in ['the-self-improving-stack.mdx','self-improving-stack-agent-runtime-topology.mdx','self-improving-stack-prompt-optimization.mdx','self-improving-stack-skill-optimization.mdx']:\n lines=(Path('src/content/posts')/name).read_text().splitlines()\n print('\\n'+name)\n i=0\n while i<len(lines):\n s=lines[i].lstrip()\n if s.startswith('```') or s.startswith('~~~'):\n mark=s[:3]; info=s[3:].strip(); start=i; i+=1; body=[]\n while i<len(lines) a…
  19. assistant #19 1 tool
    • exec src/content/posts/the-self-improving-stack.mdx
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nroot=Path('src/content/posts/the-self-improving-stack.mdx')\ns=root.read_text()\ns=s.replace(r'$$\\begin{aligned}\\Delta Q &= Q(\\text{candidate}) - Q(\\text{baseline}) \\\\\\n\\Delta C &= C(\\text{candidate}) - C(\\text{baseline})\\end{aligned}$$', r'$$\\Delta Q = Q(\\text{candidate}) - Q(\\text{baseline}), \\qquad \\Delta C = C(\\text{candidate}) - C(\\text{baseline})$$')\ns=s.replace(r\"$$\\theta\\' = \\theta_{\\text{base}} + \\Delta_{\\text{adap…
  20. assistant #20 1 tool
    • exec src/content/posts/the-self-improving-stack.mdx
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nroot=Path('src/content/posts/the-self-improving-stack.mdx')\ns=root.read_text()\ns=s.replace(r'$$\\begin{aligned}\\Delta Q &= Q(\\text{candidate}) - Q(\\text{baseline}) \\\\\\n\\Delta C &= C(\\text{candidate}) - C(\\text{baseline})\\end{aligned}$$', r'$$\\Delta Q = Q(\\text{candidate}) - Q(\\text{baseline}), \\qquad \\Delta C = C(\\text{candidate}) - C(\\text{baseline})$$')\ns=s.replace(r\"$$\\theta\\' = \\theta_{\\text{base}} + \\Delta_{\\text{adap…
  21. assistant #21 1 tool
    • exec src/content/posts/self-improving-stack-prompt-optimization.mdx
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\np=Path('src/content/posts/self-improving-stack-prompt-optimization.mdx')\nlines=p.read_text().splitlines(keepends=True)\nfor i,line in enumerate(lines):\n if '\\\\text{Agent system}' in line:\n lines[i] = r'$$\\begin{gathered}\\text{Agent system} = \\text{text surfaces} \\times \\text{model choices} \\times \\text{tool/action space} \\\\ \\times \\text{memory/retrieval policy} \\times \\text{runtime topology} \\\\ \\times \\text{turn and cost budget…
  22. assistant #22 1 tool
    • exec src/content/posts/self-improving-stack-skill-optimization.mdx
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && sed -n \"218,245p\" src/content/posts/self-improving-stack-skill-optimization.mdx'","yield_time_ms":10000,"max_output_tokens":1000}); text(r.output)
  23. assistant #23 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nfor name in ['the-self-improving-stack.mdx','self-improving-stack-agent-runtime-topology.mdx','self-improving-stack-prompt-optimization.mdx','self-improving-stack-skill-optimization.mdx']:\n path=Path('src/content/posts')/name; lines=path.read_text().splitlines(); fences=[]; i=0\n while i<len(lines):\n s=lines[i].lstrip()\n if s.startswith('```') or s.startswith('~~~'):\n mark=s[:3]; info=s[3:].strip(); start=i; i+=1; body=[]\n while i<len(lines) an…
  24. assistant #24 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 -' <<'PY'\nfrom pathlib import Path\nfor name in ['the-self-improving-stack.mdx','self-improving-stack-agent-runtime-topology.mdx','self-improving-stack-prompt-optimization.mdx','self-improving-stack-skill-optimization.mdx']:\n path=Path('src/content/posts')/name; lines=path.read_text().splitlines(); fences=[]; i=0\n while i<len(lines):\n s=lines[i].lstrip()\n if s.startswith('```') or s.startswith('~~~'):\n mark=s[:3]; info=s[3:].strip(); start=i; i+=1; body=[]\n while i<len(lines) an…
  25. assistant #25 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cat > /tmp/article-text-blocks-audit.txt' <<'EOF'\nFocused article text-fence pass, 2026-10-02\nScope: src/content/posts only; four assigned stack articles.\nNo tests/builds run. Root owns final trace capture, build, commit, and release.\n\nFence accounting (original counts confirmed with a line-state parser):\n- the-self-improving-stack.mdx: 23 total; 21 converted; 2 retained.\n- self-improving-stack-agent-runtime-topology.mdx: 17 total; 12 converted; 5 retained.\n- self-improving-stack-prompt-optimization.mdx: 16 total; 12 converted…
  26. assistant #26 1 tool
    • exec
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cat /tmp/article-text-blocks-audit.txt'","yield_time_ms":10000,"max_output_tokens":2200}); text(r.output)
  27. assistant #27 1 tool
    • exec src/content/posts/the-self-improving-stack.mdx
      input
      const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && git diff --check -- src/content/posts/the-self-improving-stack.mdx src/content/posts/self-improving-stack-agent-runtime-topology.mdx src/content/posts/self-improving-stack-prompt-optimization.mdx src/content/posts/self-improving-stack-skill-optimization.mdx && git status --short --branch'","yield_time_ms":10000,"max_output_tokens":1600}); text(r.output)

Diff

Per-file changes from 99791a3.

AGENTS.md current file (first 80 lines)
# Blog Agent RulesThis Astro and MDX blog distinguishes authorship traces from supporting research traces.For layout, components, or visual examples, read [DESIGN.md](DESIGN.md).For commands and frontmatter fields, inspect `package.json`, `tools/blog-loop.mjs`, and `src/content.config.ts`.## VoiceBefore drafting, rewriting, or polishing any non-original post, read `VOICE.md`.It owns the author voice; this file owns repository rules.## Human-only originalsNever edit the body, add revisions, or capture AI traces for a post with `original: true`.Only fix title, description, tags, or date in its frontmatter when the user explicitly requests that change.Human originals remain distinct from AI-authored posts in revision history and presentation.## Trace lifecycle- Supporting research: use `pnpm blog research <post> --harness=codex|claude-code` to get the prompt. Do not edit the post. Finish with `pnpm blog finish <post> --research --harness=... --note="..."`. This attaches the trace to `supporting_trace_ids`.- AI authorship/editing: use `pnpm blog write <post> --harness=codex|claude-code --role=draft|rewrite|polish|outline` to get the prompt. The command prints a unique trace marker, voice checklist, anti-pattern gates, and the exact finish command. Edit only non-original posts. Finish with the marked command it printed, normally `pnpm blog finish <post> --write --harness=... --role=... --marker="..." --note="..."`. This appends to `revisions[]`.- Human editing: Drew uses `pnpm write <post>` and then `pnpm write <post> --commit --note="..."`. Human revisions use `model: 'human'` and render green.Do not put supporting research traces in `revisions[]`. Do not put authorship traces only in `supporting_trace_ids`.Do not run an unmarked AI write finish unless you are doing an audited recovery capture with `--session=<id>` or `--allow-unmarked`. Unmarked captures can attach stale long-session traces to fresh edits.## Writing Anti-Patterns- Do not leave process labels in post bodies: "first post", "reader hook", "the article should", "keep this compact", "Drew angle to rewrite around", "target audience", "outline notes", or similar note-to-self phrasing.- Do not use em dashes in new prose. Use a comma, colon, parentheses, or a separate sentence.- Do not pad outlines with generic writing advice. Every heading should describe reader-facing content, not a task for a future writer.- Do not mix provenance with the article argument. Trace details belong in frontmatter, trace components, or research notes. If a visible provenance banner is required, keep it short and factual.- Do not publish claims about "SOTA", "latest", or "current" without a dated source trail.- Do not water down technical material for a broad AI audience by removing the math. Explain the symbols instead.- Do not leave placeholders, future-tense instructions, or self-referential scaffolding in a draft unless the file is explicitly a private checklist.- Do not make a taxonomy that only lists tools. Name the mutable surface, objective, evaluator, promotion gate, and failure mode for each layer.
src/content/posts/the-self-improving-stack.mdx +98 −164
diff --git a/src/content/posts/the-self-improving-stack.mdx b/src/content/posts/the-self-improving-stack.mdxindex e30766a..b5c2866 100644--- a/src/content/posts/the-self-improving-stack.mdx+++ b/src/content/posts/the-self-improving-stack.mdx@@ -43,15 +43,15 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro';+ I can tell a coding agent to parallelize work, and it will often agree with me while still doing one thing at a time.  That failure looks like a prompting problem until you inspect the trace. The sentence "fan out independent subtasks" changed the model's intention, but it did not create a worker pool, a scheduler, a merge rule, a verifier, or a budget policy. The prompt moved. The action space did not.  That is the category error hiding inside a lot of talk about self-improving agents. We say: -```text-the system optimizes itself-```+The system optimizes itself.  as if there were one surface called "the system." @@ -63,14 +63,12 @@ Self-improvement only has content after you name the task class. The target is b  The useful questions are more concrete: -```text-what user task distribution is being improved?-what is allowed to change?-what evidence says it improved that task?-how are candidates generated?-what gate decides promotion?-what can go wrong when that layer changes?-```+- Which user task distribution is being improved?+- What is allowed to change?+- What evidence shows improvement on that task?+- How are candidates generated?+- What gate decides promotion?+- What can go wrong when that layer changes?  Those six questions are the self-improving stack. @@ -78,29 +76,18 @@ Those six questions are the self-improving stack.  A self-improving agent system has a closed loop: -```text-run-observe-diagnose-propose-validate-promote-remember-govern-```+<Steps layout="flow" items={[{title:'Run'}, {title:'Observe'}, {title:'Diagnose'}, {title:'Propose'}, {title:'Validate'}, {title:'Promote'}, {title:'Remember'}, {title:'Govern'}]} />  The loop is only real when each verb has a concrete implementation. -```text-run: execute the agent under a sampled user-task scenario-observe: capture a full trace, not only a score-diagnose: identify failure modes and missing knowledge-propose: generate a candidate change-validate: test the candidate against baseline-promote: replace baseline only if the gate passes-remember: persist the right lesson for future runs-govern: keep the optimizer inside its authority and evidence boundary-```+- **Run:** Execute the agent under a sampled user-task scenario.+- **Observe:** Capture a full trace, not only a score.+- **Diagnose:** Identify failure modes and missing knowledge.+- **Propose:** Generate a candidate change.+- **Validate:** Test the candidate against the baseline.+- **Promote:** Replace the baseline only if the gate passes.+- **Remember:** Persist the right lesson for future runs.+- **Govern:** Keep the optimizer inside its authority and evidence boundary.  The system is not self-improving because it says "reflect." It is self-improving when a future run gets better at the class of user tasks it is meant to serve because a previous run produced admissible evidence. @@ -155,12 +142,7 @@ This table is the core of the series.  GEPA, SkillOpt, meta-harness, post-training, and memory flywheels all have the same outer skeleton: -```text-propose candidate-run candidate-measure candidate-promote or reject-```+<Steps layout="flow" items={[{title:'Propose candidate'}, {title:'Run candidate'}, {title:'Measure candidate'}, {title:'Promote or reject'}]} />  They are not the same system because they mutate different surfaces. @@ -205,16 +187,14 @@ If the runtime lacks a worker pool, the prompt can ask for parallelism but canno  This is why the series keeps separating: -```text-better wording-better procedure-better topology-better evaluator-better harness-better model-better memory-better governance-```+- Better wording+- Better procedure+- Better topology+- Better evaluator+- Better harness+- Better model+- Better memory+- Better governance  Those are different control surfaces. @@ -224,23 +204,15 @@ Skills sit between prompts and code.  A skill is durable procedural memory: -```text-when this task class appears,-use this decomposition,-with these tools,-under these checks,-and stop under these conditions-```+> When this task class appears, use this decomposition, with these tools, under these checks, and stop under these conditions.  That makes skills more reusable than a one-off prompt and less rigid than hard-coded application logic.  The hard part is activation. A skill that never triggers is inert. A skill that triggers everywhere becomes a new bug. The gate has to test transfer: -```text-does the skill improve held-out tasks in the intended class?-does it avoid harming nearby tasks outside the class?-does it reduce repeated failures?-```+- Does the skill improve held-out tasks in the intended class?+- Does it avoid harming nearby tasks outside the class?+- Does it reduce repeated failures?  That is why skill optimization belongs in the stack but does not replace runtime design. @@ -250,17 +222,15 @@ Agent behavior is not only model output.  It is workflow shape: -```text-single shot-refine loop-fanout and vote-planner plus worker-researcher plus coder-supervisor plus reviewer-debate-tree search-human approval gate-```+- Single shot+- Refine loop+- Fanout and vote+- Planner plus worker+- Researcher plus coder+- Supervisor plus reviewer+- Debate+- Tree search+- Human approval gate  The topology defines which actions exist and which observations can influence future actions. @@ -268,15 +238,7 @@ This matters for multi-agent systems. "Persona" is content. "Coordinator," "revi  For `maxTurns=0` worker flows, learning does not happen inside the worker's conversation. It happens across runs: -```text-pre-run retrieval-single worker attempt-post-run trace capture-analyst finding-candidate change-promotion gate-next run-```+<Steps layout="flow" items={[{title:'Pre-run retrieval'}, {title:'Single worker attempt'}, {title:'Post-run trace capture'}, {title:'Analyst finding'}, {title:'Candidate change'}, {title:'Promotion gate'}, {title:'Next run'}]} />  That is still self-improvement, but the loop lives in the harness. @@ -286,23 +248,20 @@ Before claiming that a new optimizer improved the agent, beat random or naive sa  A lot of agent improvements are really compute allocation changes: -```text-more samples-more branches-more retries-more verifier calls-more expensive judge-more time-```+- More samples+- More branches+- More retries+- More verifier calls+- A more expensive judge+- More time  Those can be useful. They are not free.  The fair comparison is: -```text-quality(candidate) - quality(baseline)-cost(candidate) - cost(baseline)-```+$$+\Delta Q = Q(\text{candidate}) - Q(\text{baseline}), \qquad \Delta C = C(\text{candidate}) - C(\text{baseline})+$$  A candidate that wins only by spending more may still be worth shipping, but the claim is different. It is a cost-quality trade, not pure intelligence gain. @@ -335,34 +294,26 @@ Traces say what happened.  An agent trace needs enough information to explain the mechanism: -```text-which prompt-which model-which tools-which arguments-which observations-which retrieved documents-which artifacts-which verifier-which failure class-which budget-which outcome-```+- Which prompt+- Which model+- Which tools and arguments+- Which observations and retrieved documents+- Which artifacts and verifier+- Which failure class and budget+- Which outcome  Without traces, the system can only hill climb on a lossy projection.  With traces, the system can diagnose: -```text-missing knowledge-bad tool argument-weak verifier-wrong selector-coordination failure-memory poisoning-budget breach-unsafe side effect-```+- Missing knowledge+- Bad tool argument+- Weak verifier+- Wrong selector+- Coordination failure+- Memory poisoning+- Budget breach+- Unsafe side effect  That is why traces are not logging decoration. They are the training data for the external-state loop. @@ -372,25 +323,19 @@ When prompt, skill, and topology tuning plateau, the mutable surface may need to  Harness evolution changes: -```text-planner contracts-tool routers-selectors-trace emitters-verifiers-benchmark adapters-worktree lifecycle-promotion gates-memory write paths-```+- Planner contracts+- Tool routers and selectors+- Trace emitters and verifiers+- Benchmark adapters+- Worktree lifecycle+- Promotion gates+- Memory write paths  This is powerful because it expands the reachable set.  It is dangerous because the harness may contain the evaluator. The core rule from the governance layer is: -```text-the optimizer cannot own the gate that promotes it-```+The optimizer cannot own the gate that promotes it.  If the candidate can rewrite the judge or release policy that approves it, the loop is no longer honest. @@ -398,15 +343,13 @@ If the candidate can rewrite the judge or release policy that approves it, the l  Most systems discussed before the post-training layer mutate external state: -```text-prompt-skill-tool docs-memory-runtime-harness-evaluator-```+- Prompt+- Skill+- Tool documentation+- Memory+- Runtime+- Harness+- Evaluator  Post-training mutates model behavior itself: @@ -416,9 +359,9 @@ $$  or: -```text-theta' = theta_base + Delta_adapter-```+$$+\theta' = \theta_{\text{base}} + \Delta_{\text{adapter}}+$$  That can generalize better than prompt edits when the signal is strong. It also makes the behavior harder to inspect, partially roll back, and attribute to one trace. @@ -439,9 +382,9 @@ $$  A memory system is useful when: -```text-Score(with memory) - Score(without memory) > threshold-```+$$+\operatorname{Score}(\text{with memory}) - \operatorname{Score}(\text{without memory}) > \text{threshold}+$$  under cost, freshness, privacy, and poisoning constraints. @@ -455,15 +398,13 @@ Governance is what makes autonomy accountable.  A self-improving system needs a safety case: -```text-claim-scope-evidence-residual risk-owner-release gate-rollback path-```+- Claim+- Scope+- Evidence+- Residual risk+- Owner+- Release gate+- Rollback path  The system can propose improvements. The gate decides which improvements persist. The owner accepts residual risk. @@ -483,18 +424,11 @@ Without that rule, self-improvement can become proxy hacking with better brandin  When someone says their agent improves itself, ask for the layer. -```text-What changed?-Who proposed it?-What evidence was captured?-What baseline was beaten?-What held-out set was protected?-What gate approved it?-What got more expensive?-What became riskier?-What can roll back?-What persisted into the next run?-```+- What changed, and who proposed it?+- What evidence was captured, and what baseline was beaten?+- What held-out set was protected, and what gate approved it?+- What became more expensive or riskier?+- What can roll back, and what persisted into the next run?  If those questions have concrete answers, there may be a real loop. 
src/content/posts/self-improving-stack-agent-runtime-topology.mdx +39 −73
diff --git a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx b/src/content/posts/self-improving-stack-agent-runtime-topology.mdxindex 69b3005..6319837 100644--- a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx+++ b/src/content/posts/self-improving-stack-agent-runtime-topology.mdx@@ -46,6 +46,8 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro';+ When I tell a coding agent to "parallelize the work," I am not asking for a different tone.  I am asking for a different execution graph: spawn independent executions, cap concurrency, isolate state, collect traces, score results, select or merge outputs, cancel losers, and account for cost. If the runtime cannot express those moves, a prompt can only ask the model to simulate the shape.@@ -62,12 +64,10 @@ It is not the persona. It is not the supervisor prompt. It is not the model's pr  A minimal topology has: -```text-nodes = agents, tools, validators, selectors, human gates-edges = sequence, fanout, handoff, retry, interrupt, merge-state = trace, memory, artifacts, budgets, run handles-policy = planning, selection, cancellation, promotion, replay-```+- **Nodes:** Agents, tools, validators, selectors, and human gates+- **Edges:** Sequence, fanout, handoff, retry, interrupt, and merge+- **State:** Traces, memory, artifacts, budgets, and run handles+- **Policy:** Planning, selection, cancellation, promotion, and replay  The runtime action space is the set of moves the system can execute: @@ -92,9 +92,7 @@ If `parallel` is not in $A_{\text{runtime}}$, no optimized prompt can make true  The deep question is: -```text-Which topology moves are first-class runtime actions, and which are merely instructions?-```+Which topology moves are first-class runtime actions, and which exist only as instructions?  That one distinction determines whether multi-agent self-improvement is engineering or theater. @@ -144,13 +142,7 @@ The objective is still the same skeleton: propose, run, score, compare, update,  A supervisor prompt can say: -```text-Assign independent subtasks to specialist workers.-Have them work in parallel.-Merge their findings.-Ask a verifier to check the final result.-Stop when the verifier passes.-```+<Steps layout="flow" items={[{title:'Assign independent subtasks'}, {title:'Run specialist workers in parallel'}, {title:'Merge their findings'}, {title:'Verify the result'}, {title:'Stop when the verifier passes'}]} />  That text only has leverage if the runtime has matching actions. @@ -193,27 +185,21 @@ The inspected package surface exposes:  The important split is in the type surface: -```text-kernel owns: iteration accounting, concurrency, aborts, cost, traces-driver owns: topology-validator owns: scoring-output adapter owns: parsing-agent spec owns: executable profile and prompt formatting-```+- **Kernel:** Iteration accounting, concurrency, aborts, cost, and traces+- **Driver:** Topology+- **Validator:** Scoring+- **Output adapter:** Parsing+- **Agent spec:** Executable profile and prompt formatting  The decomposition matters. Swapping a `Driver` changes topology without changing the model, prompt, skill body, validator, or output parser. The execution graph becomes a replaceable runtime object rather than a paragraph inside a supervisor prompt.  For `refine`, the driver emits one task per iteration until the validator accepts the output or a cap is reached: -```text-attempt -> validate -> if invalid, attempt again -> stop on pass or cap-```+<Steps layout="flow" items={[{title:'Attempt'}, {title:'Validate'}, {title:'Retry if invalid'}, {title:'Stop on pass or cap'}]} />  For `fanout-vote`, the driver emits N attempts in the first iteration, lets the kernel run them in parallel subject to `maxConcurrency`, then selects the highest-scoring valid output: -```text-spawn N -> validate each -> select valid winner -> fail if none valid-```+<Steps layout="flow" items={[{title:'Spawn N candidates'}, {title:'Validate each'}, {title:'Select a valid winner'}, {title:'Fail if none pass'}]} />  The MCP delegation layer turns this topology into an agent-callable surface. `delegate_code` can launch specialist coder agents that produce validated patches, return immediately with a `taskId`, and let the caller poll for completion. With `variants > 1`, multiple coder harnesses attempt the task in parallel and the highest-scoring patch wins. `delegate_research` does the same shape for evidence-bearing research, with source diversity, citation density, recency, gap coverage, and namespace isolation in the scoring contract. @@ -225,13 +211,9 @@ Some useful concepts are not present in the inspected local `agent-runtime` pack  The useful conclusion is not "missing feature." It is a sharper map: -```text-shipped in the checked source:-  runLoop, refine, fanout-vote, multi-harness coder/research delegation+**Shipped in the checked source:** `runLoop`, `refine`, `fanout-vote`, and multi-harness coder/research delegation. -obvious next primitives:-  dynamic driver, typed program DSL, supervisor scope, budget ledger, durable replay-```+**Obvious next primitives:** dynamic driver, typed program DSL, supervisor scope, budget ledger, and durable replay.  This boundary is load-bearing. A substrate map is wrong if it treats unshipped names as APIs. If a runtime does not yet expose `Supervisor` or `Scope`, the concept can still be named as a target surface. It cannot be treated as a shipped API. @@ -279,28 +261,22 @@ If candidate A gets one worker and candidate B gets eight workers, B may win bec  A runtime topology benchmark should record: -```text-workers spawned-iterations used-wall-clock latency-tokens in/out-tool calls-LLM calls-failed branches-cancelled branches-human approvals-cost_usd-```+- Workers spawned+- Iterations used+- Wall-clock latency+- Tokens in and out+- Tool and LLM calls+- Failed and cancelled branches+- Human approvals+- `cost_usd`  Then promotion can distinguish: -```text-quality lift at same budget-quality lift for higher budget-latency reduction at same quality-cost reduction at same quality-risk reduction with acceptable quality loss-```+- Quality lift at the same budget+- Quality lift for a higher budget+- Latency reduction at the same quality+- Cost reduction at the same quality+- Risk reduction with acceptable quality loss  Without this ledger, topology search will usually rediscover "try more things" and call it intelligence. @@ -348,16 +324,14 @@ A serious runtime topology eval should treat topology changes as architecture ch  Minimum protocol: -```text-1. Freeze model, prompts, skills, tools, dataset, and evaluator where possible.-2. Register baseline topology hash and candidate topology hash.-3. Run paired scenario/seed comparisons.+1. Freeze the model, prompts, skills, tools, dataset, and evaluator where possible.+2. Register the baseline and candidate topology hashes.+3. Run paired scenario and seed comparisons. 4. Record every branch, tool call, validator result, selector decision, and cost. 5. Compare under at least one compute-matched budget.-6. Run stress cases for timeouts, branch failure, cancellation, and partial results.+6. Stress test timeouts, branch failures, cancellation, and partial results. 7. Reject candidates that hide errors, exceed budget, lose traces, or skip gates.-8. Promote only on held-out lift, cost/latency policy, and trace integrity.-```+8. Promote only when held-out lift, cost and latency policy, and trace integrity pass.  Promotion can look like: @@ -381,23 +355,15 @@ It hides coordination in prose. It spawns workers without isolation. It lets bra  The most common failure is fake fanout: -```text-prompt says: split the task among specialists-runtime does: one model call with specialist names in text-trace says: one branch-```+- **Prompt:** Says to split the task among specialists.+- **Runtime:** Makes one model call with specialist names in the text.+- **Trace:** Records one branch.  The fix is not a better coordinator prompt. The fix is an actual fanout primitive.  Another failure is unpriced parallelism: -```text-candidate topology spawns 8 workers-baseline topology spawns 1 worker-candidate wins by 4 points-candidate costs 12x more-promotion report says "better"-```+The candidate topology spawns eight workers while the baseline spawns one. It wins by four points, costs twelve times more, and the promotion report calls it “better.”  That is not necessarily wrong. It is incomplete. The product decision depends on whether the gain is worth the compute, latency, and operational complexity. 
src/content/posts/self-improving-stack-prompt-optimization.mdx +41 −76
diff --git a/src/content/posts/self-improving-stack-prompt-optimization.mdx b/src/content/posts/self-improving-stack-prompt-optimization.mdxindex 8e27997..453bfb2 100644--- a/src/content/posts/self-improving-stack-prompt-optimization.mdx+++ b/src/content/posts/self-improving-stack-prompt-optimization.mdx@@ -46,6 +46,8 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro';+ The first time prompt optimization feels magical is also the moment it starts lying to you.  You change a sentence, run the benchmark, and the score moves. GEPA reflects over traces. MIPRO searches instructions and demonstrations. DSPy compiles an LM program. AxLLM packages the optimizer into a TypeScript surface. The intervention is small, the evidence is numeric, and the temptation is to say the agent improved.@@ -58,9 +60,7 @@ Prompt optimization is powerful when a text surface has causal leverage over the  So the question is not "which prompt optimizer is best?" The question is: -```text-Which factors are mutable, which factors are held fixed, and which evaluator is trusted enough to promote a candidate?-```+Which factors are mutable, which stay fixed, and which evaluator is trusted to promote a candidate?  Answer that and the ecosystem stops looking like magic. It becomes experimental design. @@ -93,15 +93,11 @@ This is experimental design, not pedantry. A candidate prompt can look better be  Prompt optimization is cleanest when the causal path is: -```text-text surface -> model behavior -> output or trajectory -> score-```+A text surface changes model behavior, which changes the output or trajectory that receives a score.  The measurement becomes contaminated when the actual path is: -```text-text surface -> planner hint -> runtime capability missing -> no action -> evaluator still gives partial credit-```+A text surface can give the planner a hint while the runtime lacks the capability to act on it. No action follows, yet the evaluator may still award partial credit.  The first path is a causal optimization claim. The second is a measurement artifact. @@ -159,14 +155,12 @@ AxLLM brings these ideas into a TypeScript-first programming surface. Ax program  This is the through-line: -```text-APE: generate instruction candidates, score them.-OPRO: use an LLM as the optimizer over scored candidates.-MIPRO: optimize instructions and demos across LM-program modules.-TextGrad: propagate textual feedback through a computation graph.-GEPA: evolve prompts from trace reflection and Pareto selection.-AxLLM: expose these optimization patterns in a typed production surface.-```+- **APE:** Generates instruction candidates and scores them.+- **OPRO:** Uses an LLM to optimize scored candidates.+- **MIPRO:** Optimizes instructions and demos across LM-program modules.+- **TextGrad:** Propagates textual feedback through a computation graph.+- **GEPA:** Evolves prompts through trace reflection and Pareto selection.+- **AxLLM:** Exposes these optimization patterns through a typed production surface.  They share a black-box or language-mediated search skeleton. They differ in candidate representation, proposal operator, feedback channel, and selection rule. @@ -215,16 +209,9 @@ That distinction becomes central for agents. Agent failures are often procedural  A multi-agent system is not a bigger prompt. It is a factored system: -```text-agent system = text surfaces-             x model choices-             x tool/action space-             x memory/retrieval policy-             x runtime topology-             x turn and cost budgets-             x evaluator stack-             x promotion policy-```+$$+\begin{gathered}\text{Agent system} = \text{text surfaces} \times \text{model choices} \times \text{tool/action space} \\ \times \text{memory/retrieval policy} \times \text{runtime topology} \\ \times \text{turn and cost budgets} \times \text{evaluator stack} \times \text{promotion policy}\end{gathered}+$$  Prompt optimizers explore the factors you expose to them. If the candidate encoding only contains a supervisor prompt, the optimizer can only mutate supervisor text. It cannot invent a real fanout executor, persistent memory, sandbox policy, multi-worker merge protocol, or verifier pass unless those dimensions are present in the candidate representation. @@ -232,13 +219,7 @@ This is the exact issue with coordinator and worker personas.  You can optimize a coordinator persona to say: -```text-Fan out independent subtasks.-Run specialists in parallel.-Merge findings.-Escalate disagreement to a supervisor.-Verify before final output.-```+<Steps layout="flow" items={[{title:'Fan out independent subtasks'}, {title:'Run specialists in parallel'}, {title:'Merge findings'}, {title:'Escalate disagreement to a supervisor'}, {title:'Verify before final output'}]} />  That may improve behavior if the runtime already exposes the necessary actions. It will not create those actions. A text instruction to parallelize is operational only if the agent has a tool or runtime primitive that dispatches work concurrently. A text instruction to continue until verified only works if the loop budget and stop semantics allow it. A text instruction to use memory only works if memory is available, scoped, and retrievable. @@ -276,19 +257,12 @@ The Tangle packages fit as substrate, not as another prompt optimizer.  That means the stack split is: -```text-GEPA/MIPRO/Ax/DSPy:-  proposal and search over prompt or LM-program surfaces--agent-runtime:-  execution semantics, topology, tools, loops, surfaces, trace emission--agent-eval:-  scoring, trace analysis, causal attribution, gates, promotion evidence--meta-harness:-  architecture-level search over the system that runs and evaluates agents-```+| System | Responsibility |+| --- | --- |+| GEPA, MIPRO, Ax, DSPy | Propose and search prompt or LM-program surfaces |+| `agent-runtime` | Execution semantics, topology, tools, loops, surfaces, and trace emission |+| `agent-eval` | Scoring, trace analysis, causal attribution, gates, and promotion evidence |+| `meta-harness` | Architecture-level search over the system that runs and evaluates agents |  This composition keeps each layer honest. Prompt optimizers mutate text. Runtime exposes real action and topology. Eval decides promotion under controlled comparisons. Meta-harness touches architecture only when the lower layers plateau or the traces show the wrong surface is being optimized. @@ -298,16 +272,14 @@ A serious prompt optimization run should leave an audit trail strong enough for  Minimum protocol: -```text-1. Freeze model, runtime, toolset, schema, and evaluator.-2. Split examples into search, validation, and holdout.-3. Register baseline prompt/config hash.-4. Generate candidates with stable ids and rationales.-5. Run paired comparisons on identical scenario/seed cells.+1. Freeze the model, runtime, toolset, schema, and evaluator.+2. Split examples into search, validation, and holdout sets.+3. Register the baseline prompt and configuration hash.+4. Generate candidates with stable IDs and rationales.+5. Run paired comparisons on identical scenario and seed cells. 6. Preserve full traces, not only final scores. 7. Reject candidates that violate schema or safety invariants.-8. Promote only on held-out lift, cost budget, and regression checks.-```+8. Promote only when held-out lift, cost budget, and regression checks pass.  For stochastic agents, paired deltas are the basic unit: @@ -330,9 +302,9 @@ The exact statistic can vary. Bootstrap confidence intervals, Wilcoxon signed-ra  For multi-agent systems, add factorial attribution: -```text-cells = model x prompt x topology x scenario x seed-```+$$+\text{cells} = \text{model} \times \text{prompt} \times \text{topology} \times \text{scenario} \times \text{seed}+$$  If the candidate wins, you want to know why. Was the lift from prompt wording, model choice, topology, scenario mix, or an interaction? `agent-eval` has a local causal attribution primitive for exactly this style of factorial decomposition. The point is not academic neatness. It prevents the team from shipping a prompt change while the actual effect came from a model swap or runtime setting. @@ -344,20 +316,15 @@ They overfit benchmark phrasing. They learn judge preferences. They inflate verb  The most dangerous failure is surface misattribution: -```text-observed failure: agent did not verify output-wrong fix: add "always verify" to prompt-right diagnosis: runtime stops after first draft, verifier tool is absent, or eval ignores missing verification-```+- **Observed failure:** The agent did not verify the output.+- **Wrong fix:** Add “always verify” to the prompt.+- **Right diagnosis:** The runtime stops after the first draft, the verifier tool is absent, or the evaluation ignores missing verification.  The text fix may help a little. It is still optimizing the wrong factor.  Another common failure is evaluator coupling: -```text-candidate prompt -> answer phrasing closer to judge rubric -> higher score-candidate prompt -> no real increase in task success-```+A candidate prompt can move answer phrasing closer to the judge rubric and raise the score without increasing task success.  This is why high-stakes surfaces need multiple evaluators: deterministic checks where possible, calibrated LLM judges where needed, human labels for anchor sets, and downstream outcome correlation when production data exists. @@ -388,15 +355,13 @@ Do not use prompt optimization as the primary fix when the traces show a capabil  Robust agent stacks do not pick one optimizer and call it the answer. They route failures to the layer with causal control: -```text-wording failure -> prompt optimizer-procedure failure -> skill optimizer-tool-doc failure -> tool-surface optimizer-memory failure -> knowledge or retrieval update-coordination failure -> runtime/topology search-evaluator failure -> harness and judge repair-model capability failure -> model selection, fine-tuning, or post-training-```+- **Wording failure:** Prompt optimizer+- **Procedure failure:** Skill optimizer+- **Tool documentation failure:** Tool-surface optimizer+- **Memory failure:** Knowledge or retrieval update+- **Coordination failure:** Runtime or topology search+- **Evaluator failure:** Harness and judge repair+- **Model capability failure:** Model selection, fine-tuning, or post-training  That is the core answer to "are they doing the same thing?" 
src/content/posts/self-improving-stack-skill-optimization.mdx +31 −63
diff --git a/src/content/posts/self-improving-stack-skill-optimization.mdx b/src/content/posts/self-improving-stack-skill-optimization.mdxindex a38fd8a..24ac3f1 100644--- a/src/content/posts/self-improving-stack-skill-optimization.mdx+++ b/src/content/posts/self-improving-stack-skill-optimization.mdx@@ -46,6 +46,8 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro';+ A prompt can rescue one run.  A skill can change the next hundred.@@ -80,9 +82,7 @@ The same idea shows up in research systems. Voyager builds an ever-growing libra  The common theme: -```text-skill = persistent procedure + activation condition + operational scope-```+A skill is a persistent procedure together with its activation condition and operational scope.  If the procedure does not persist, it is a prompt. If it persists but only stores observations, it is memory. If it can directly act on the world, it is a tool. If it controls the loop that invokes tools and agents, it is runtime. @@ -121,10 +121,8 @@ That is exactly how `@tangle-network/agent-eval` treats profiles locally: skills  The skill layer has two coupled optimization targets: -```text-skill quality: does the loaded procedure improve the task?-activation quality: does the right skill load at the right time?-```+- **Skill quality:** Does the loaded procedure improve the task?+- **Activation quality:** Does the right skill load at the right time?  Most prompt optimizers focus on the first term and ignore the second. Skill systems cannot. A strong skill that never triggers is dead. A mediocre skill that triggers everywhere becomes global prompt pollution. @@ -146,32 +144,22 @@ As of June 5, 2026, SkillOpt is one of the clearest attempts to make skill train  The loop is: -```text-current skill-  -> rollout on scored tasks-  -> reflect over successes and failures-  -> propose bounded add/delete/replace edits-  -> validate candidate skill-  -> accept only if held-out selection improves-  -> export best_skill.md-```+<Steps layout="flow" items={[{title:'Start with the current skill'}, {title:'Roll out on scored tasks'}, {title:'Reflect on successes and failures'}, {title:'Propose bounded add, delete, or replace edits'}, {title:'Validate the candidate skill'}, {title:'Accept only if held-out selection improves'}, {title:'Export the best skill'}]} />  The target agent is frozen. A separate optimizer model proposes edits. The candidate update is bounded by a textual learning-rate budget, so the skill cannot be rewritten arbitrarily every epoch. Rejected edits are kept as negative evidence. A slow or meta update gives the optimizer longer-horizon memory without bloating deployment. The deployed artifact is the skill file.  That matters for inference cost: -```text-training time: optimizer model + rollouts + gates-deployment time: target model + final skill-```+| Phase | Components |+| --- | --- |+| Training | Optimizer model, rollouts, and gates |+| Deployment | Target model and final skill |  SkillOpt's paper reports best or tied-best performance across 52 evaluated model, benchmark, and harness cells, covering six benchmarks, seven target models, and direct chat, Codex, and Claude Code execution harnesses. It also reports transfer of optimized skill artifacts across model scales, between Codex and Claude Code, and to a nearby benchmark without further optimization.  The exact numbers will age. The durable system claim is: -```text-external procedure can be trained while the policy model stays fixed-```+An external procedure can be trained while the policy model stays fixed.  That is a different layer from prompt optimization. Prompt search asks what text should steer this program. Skill optimization asks what reusable procedure the agent should inherit next time. @@ -185,10 +173,7 @@ A prompt usually disappears when the run ends.  This lifecycle makes the objective different: -```text-prompt success = better behavior on the current eval distribution-skill success = better behavior on future tasks where activation is justified-```+Prompt optimization succeeds when behavior improves on the current evaluation distribution. Skill optimization succeeds when behavior improves on future tasks where activation is justified.  That future-facing property makes skills more powerful and more dangerous. A bad prompt can damage one generation. A bad skill can train the whole operator into a recurring failure mode. @@ -214,14 +199,12 @@ This changes the promotion criteria for skill optimization.  Prompt injection tests are not enough. A skill gate also needs: -```text-registry safety: should this skill be admitted?-selection safety: when does the agent choose it?-content safety: what procedure does it teach?-tool safety: what capabilities does it exercise?-interaction safety: what other skills does it conflict with?-revocation safety: can we disable it and recover?-```+- **Registry safety:** Should this skill be admitted?+- **Selection safety:** When should the agent choose it?+- **Content safety:** What procedure does it teach?+- **Tool safety:** What capabilities does it exercise?+- **Interaction safety:** Which other skills conflict with it?+- **Revocation safety:** Can we disable it and recover?  This is where skill optimization differs from skill generation. Generating a useful skill is only half the problem. Operating a skill ecosystem requires provenance, versioning, linting, evals, trust boundaries, and rollback. @@ -239,19 +222,12 @@ In the Tangle stack, skills belong in the behavior profile and in the improvemen  The architectural split is: -```text-SkillOpt:-  candidate generator for skill text--agent-runtime:-  where skills are declared, loaded, scoped, and paired with tools/runtime--agent-eval:-  profile hashing, scorecards, trace diagnosis, held-out gates, promotion evidence--meta-harness:-  decides when skill edits are insufficient and runtime architecture must change-```+| System | Responsibility |+| --- | --- |+| SkillOpt | Generate candidate skill text |+| `agent-runtime` | Declare, load, and scope skills; pair them with tools and runtime |+| `agent-eval` | Profile hashing, scorecards, trace diagnosis, held-out gates, and promotion evidence |+| `meta-harness` | Decide when skill edits are insufficient and runtime architecture must change |  This prevents a common mistake: treating a skill as capability creation. A skill can teach an agent to use a verifier. It cannot create the verifier tool. A skill can teach a coordinator to fan out work. It cannot create a concurrent dispatch primitive. A skill can teach a coding agent to preserve user changes. It cannot fix a runtime that runs destructive commands. @@ -284,9 +260,7 @@ This matters for the "driver and supervisor" problem. A SkillOpt-style loop can  The correct question is: -```text-Is the failure caused by missing procedure, missing activation, missing affordance, or wrong topology?-```+Is the failure caused by a missing procedure, missing activation, missing affordance, or the wrong topology?  Only the first two are skill optimization problems. @@ -296,18 +270,16 @@ A serious skill optimization protocol should treat a skill change like a behavio  Minimum protocol: -```text-1. Freeze model, prompt version, tools, runtime, and evaluator.-2. Register baseline agent profile hash.+1. Freeze the model, prompt version, tools, runtime, and evaluator.+2. Register the baseline agent profile hash. 3. Split tasks into search, validation, holdout, and transfer sets.-4. Run with and without the skill to estimate marginal effect.+4. Run with and without the skill to estimate its marginal effect. 5. Track activation decisions as first-class events. 6. Preserve full traces, not only task scores. 7. Reject skills that violate tool, security, schema, or scope constraints. 8. Promote only if held-out lift clears uncertainty and cost gates. 9. Run interference tests outside the declared skill scope.-10. Record version, provenance, rejected edits, and rollback path.-```+10. Record the version, provenance, rejected edits, and rollback path.  The core comparison is paired: @@ -346,17 +318,13 @@ They fossilize bad habits. They preserve workarounds after the underlying bug is  The most important failure is procedural poisoning: -```text-one lucky success -> generalized rule -> repeated future failures-```+One lucky success can become a generalized rule that causes repeated failures later.  This is the same shape as overfitting, but more persistent. A prompt overfit tends to stay inside an experiment. A skill overfit becomes part of the operator.  Another failure is activation gaming: -```text-skill description becomes broad -> skill loads often -> eval improves on benchmark -> unrelated tasks degrade-```+A broad skill description can make the skill load more often and improve the benchmark score while degrading unrelated tasks.  This is why the activation policy $a$ belongs in the objective. Optimizing only the body $k$ is incomplete.