The Self-Improving Stack
Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
- Created
- Updated
40
Turns
36
Tool calls
35
Files touched
24m
Duration
Files
/Users/drew/dotfiles/docs/processes/agent-work.md/Users/drew/dotfiles/docs/anti-patterns/Users/drew/dotfiles/docs/anti-patterns/blog-and-research.md;tools/blog-loop.mjs/tmp/audit_blog_math.py/tmp/audit_math_blocks.py/tmp/audit_inline_math.pysrc/content/research/bcww.mdxsrc/content/research/ghz.mdxsrc/content/research/linden-winter.mdxtools/trace-capture.ts/tmp/latex_commands.py/tmp/fence_assignments.py/tmp/more_math.py/tmp/count_display.pysrc/content/posts/self-improving-stack-test-time-compute.mdxsrc/content/posts/self-improving-stack-evaluation-gates.mdx/tmp/count_inline_math.py/tmp/make_math_inventory.py/tmp/blog-math-luna/changed-math.md/tmp/blog-math-luna/audit.mdsrc/content/posts/convergence-as-eval-primitive.mdxsrc/content/posts/self-improving-stack-agent-runtime-topology.mdxsrc/content/posts/self-improving-stack-governance.mdxsrc/content/posts/self-improving-stack-harness-evolution.mdxsrc/content/posts/self-improving-stack-memory-flywheels.mdxsrc/content/posts/self-improving-stack-multi-agent-coordination.mdxsrc/content/posts/self-improving-stack-optimization-theory.mdxsrc/content/posts/self-improving-stack-post-training.mdxsrc/content/posts/self-improving-stack-prompt-optimization.mdxsrc/content/posts/self-improving-stack-skill-optimization.mdxsrc/content/posts/self-improving-stack-trace-systems.mdxsrc/content/posts/the-self-improving-stack.mdx/Users/drew/.codex/sessions/2026/10/02/rollout-2026-10-02T15-42-44-01a0fec8-c092-7552-9714-37c60fda6514.jsonl/Users/drew/.codex/sessions/2026/10/02/rollout-2026-10-02T15-42-44-01a0fec8-c092-7552-9714-37c60fda6514.jsonl\
Commit
a68bc06 fix: render article equations as math +465 −410 across 19 files
Conversation
40 turns. Full text where captured; older traces show only the first ~280 chars.
- user #1# AGENTS.md instructions for /Users/drew/webb/discovery-lab <INSTRUCTIONS> # Shared agent defaults Apply these rules to Claude, Codex, and OpenCode. Resolve installed symlinks before following documentation links. ## Own the outcome Deliver implementation through the requested consumer or live result. Choose the smallest sufficient verification and its stop condition before running checks. Reuse valid results until changed code, dependencies, environment, or a failure invalidates them. Never wait on PR CI; the local gate is the merge gate. Gate on a beelink: merge the base, frozen install, ty…
- assistant #2I’ll audit the blog content on the shared checkout, then make only deliberate math-markup edits in the assigned content directories. I’ll first read the repository and process guidance, and keep an exact inventory of changes and exclusions for the root agent.
- assistant #3 1 tool
- exec
input
const r = await tools.exec_command({cmd:"pwd && rg --files -g 'AGENTS.md' -g 'VOICE.md' -g 'DESIGN.md' -g 'agent-work.md' | head -80 && ls",yield_time_ms:10000,max_output_tokens:3000}); text(r.output)
-
- assistant #4 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && pwd && git status --short --branch && ls'","yield_time_ms":10000,"max_output_tokens":4000}); text(r.output)
-
- assistant #5 1 tool
- exec
/Users/drew/dotfiles/docs/processes/agent-work.mdinput
const r = await tools.exec_command({cmd:"cat /Users/drew/dotfiles/docs/processes/agent-work.md && ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && cat AGENTS.md && cat VOICE.md && cat DESIGN.md'","yield_time_ms":10000,"max_output_tokens":16000}); text(r.output)
-
- assistant #6 1 tool
- exec
/Users/drew/dotfiles/docs/anti-patternsinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'find /home/drew -path \"*/blog-and-research.md\" -o -path \"*/copywriting.md\" -o -path \"*/product-design.md\" 2>/dev/null | head -40; cd /home/drew/code/drewstone.github.io && git log -1 --oneline --decorate && git reflog -5 --oneline && find src/content/posts src/content/research -type f | sort | wc -l' && ls -l /Users/drew/dotfiles/docs/anti-patterns","yield_time_ms":10000,"max_output_tokens":3000}); text(r.output)
-
- assistant #7Problem: many posts appear to show equations as literal code instead of rendered math. Change: audit all assigned Markdown and MDX, then convert only clear mathematical expressions to KaTeX delimiters. Why long-term right: fixes the source content directly while preserving actual code, prose, and human-original bodies. Cost: content-only edits across the shared branch; root owns the combined build, deployment, and rollback. I’m checking the content inventory and trace workflow before editing.
- assistant #8 1 tool
- exec
/Users/drew/dotfiles/docs/anti-patterns/blog-and-research.md;input
const r = await tools.exec_command({cmd:"cat /Users/drew/dotfiles/docs/anti-patterns/blog-and-research.md; ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && git log -1 --oneline --decorate && git reflog -5 --oneline && find src/content/posts src/content/research -type f | sort | wc -l && cat package.json'","yield_time_ms":10000,"max_output_tokens":9000}); text(r.output)
-
- assistant #9 1 tool
- exec
tools/blog-loop.mjsinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && sed -n \"1,260p\" tools/blog-loop.mjs && rg -n \"original: true|^```(text|tex|latex)|`[^`]*(=|\\\\\\\\frac|\\\\\\\\sum|\\\\\\\\int|\\\\\\\\alpha|\\\\\\\\beta|\\\\\\\\times|\\\\\\\\approx|\\\\\\\\to|\\\\\\\\cdot|[A-Za-z0-9]_[A-Za-z0-9]|[A-Za-z0-9]\\\\^[0-9])[^`]*`\" src/content/posts src/content/research'","yield_time_ms":10000,"max_output_tokens":16000}); text(r.output)
-
- assistant #10 1 tool
- exec
tools/blog-loop.mjsinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && sed -n \"1,260p\" tools/blog-loop.mjs'","yield_time_ms":10000,"max_output_tokens":7000}); text(r.output)
-
- assistant #11 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && printf \"%s\\n\" \"POSTS\" && find src/content/posts -type f | sort && printf \"%s\\n\" \"RESEARCH\" && find src/content/research -type f | sort && printf \"%s\\n\" \"ORIGINALS\" && rg -l \"original: true\" src/content/posts src/content/research && printf \"%s\\n\" \"TEXT FENCES\" && rg -n -e \"^```(text|tex|latex)\" src/content/posts src/content/research'","yield_time_ms":10000,"max_output_tokens":8000}); text(r.output)
-
- assistant #12 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && printf POSTS\\n && find src/content/posts -type f | sort && printf RESEARCH\\n && find src/content/research -type f | sort && printf ORIGINALS\\n && rg -l \"original: true\" src/content/posts src/content/research'","yield_time_ms":10000,"max_output_tokens":6000}); text(r.output)
-
- assistant #13 1 tool
- exec
/tmp/audit_blog_math.pyinput
const r = await tools.exec_command({cmd:"cat > /tmp/audit_blog_math.py <<'PY'\nfrom pathlib import Path\nimport re\nfor root in ['src/content/posts','src/content/research']:\n for path in sorted(Path(root).glob('*.mdx')):\n s=path.read_text()\n front=re.match(r'^---\\n(.*?)\\n---\\n',s,re.S)\n orig=bool(front and re.search(r'^original:\\s*true\\s*$',front.group(1),re.M))\n hits=[]\n fenced=False\n for i,line in enumerate(s.splitlines(),1):\n if line.startswith('```'):\n if not fenced:\n fenced=True\n if re.match(r'^```(?:text|tex|latex)\\b',line): hits.append((i,'fence-open',line))\n else: fe…
-
- assistant #14 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && rg -n -i \"latex|tex|equation|frac|sqrt|math|∑|∫|≤|≥|≈|≠|→\" src/content/posts src/content/research'","yield_time_ms":10000,"max_output_tokens":8000}); text(r.output)
-
- assistant #15 1 tool
- exec
/tmp/audit_math_blocks.pyinput
const r = await tools.exec_command({cmd:"cat > /tmp/audit_math_blocks.py <<'PY'\nfrom pathlib import Path\nimport re\nfor root in ['src/content/posts','src/content/research']:\n for path in sorted(Path(root).glob('*.mdx')):\n s=path.read_text()\n if re.search(r'^original:\\s*true\\s*$',s.split('---',2)[1] if s.startswith('---') else '',re.M): continue\n lines=s.splitlines(); in_fence=False; fence=[]; start=0\n for i,line in enumerate(lines,1):\n if line.startswith('```'):\n if not in_fence:\n in_fence=True; fence=[]; start=i; lang=line[3:].strip()\n else:\n text='\\n'.join(fence)\n # keep onl…
-
- assistant #16 1 tool
- exec
input
const r = await tools.exec_command({cmd:"sed -n '360,500p' /tmp/audit_math_blocks.out","yield_time_ms":10000,"max_output_tokens":7000}); text(r.output)
-
- assistant #17 1 tool
- exec
/tmp/audit_inline_math.pyinput
const r = await tools.exec_command({cmd:"cat >/tmp/audit_inline_math.py <<'PY'\nfrom pathlib import Path\nimport re\nfor root in ['src/content/posts','src/content/research']:\n for p in sorted(Path(root).glob('*.mdx')):\n s=p.read_text(); ls=s.splitlines(); in_f=False\n fm=re.match(r'^---\\n(.*?)\\n---',s,re.S); orig=bool(fm and re.search(r'^original:\\s*true\\s*$',fm.group(1),re.M))\n if orig: continue\n for i,line in enumerate(ls,1):\n if line.startswith('```'): in_f=not in_f; continue\n if in_f: continue\n for m in re.finditer(r'`([^`]+)`',line):\n v=m.group(1)\n if (re.search(r'(?<!\\w)(?…
-
- assistant #18 1 tool
- send_message
input
{"target":"/root","message":"gAAAAABqwDQKC6UaT-iQndWBWPc9m0sPhk3FxB4CMObRhlaT0VkYNMz0d97_mrI0FVQcOzKw-DRNJ_xSf5EUM7W3B7fWX-TfFCN6zFwM_wxAJT-3HuzvEoIroHX5sTRfKHv_klmV707E4cjTdncCQ2g2w_d9mIP3h7gy9JtPLBQ2MH9oqVCIdGBlpcJGRlvLl9PWcj2takUDwkBXDENYEF27_tOJUKNS0zX3-gszRasHyzdPz-ZzT77bFxj-RLRtC4mKBMXxcQhEgascjgyYtz9LrdksoOV8Q2RGPnzNN2XUooYSjW4DuUDkmJII6t8O3GdyJnW7I0mhZxXzsvLcJkjRCJQu5rusSkJ63g6MbXvunm3olVebumfgt8blxSmibr4cyzuhOwFkCb0Dzod-6yQ4ndteZN174idxGhE7dLwXE3xwpiq4xE-kby894fIgaGb3eIcAkD6V2HNXhfUfYol2lUV4LIKpTxW-s7ZpLEPVF0Yz9TbxiPr2k18RI1E1M66KyXGOpxu0fYD7jDKnsAtJC-v8UZsT4ezvrN1AsMWAx4UYcecHQQ87zR…
-
- assistant #19 1 tool
- exec
src/content/research/bcww.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && sed -n \"1,110p\" src/content/research/bcww.mdx && sed -n \"1,105p\" src/content/research/ghz.mdx && sed -n \"60,100p\" src/content/research/linden-winter.mdx'","yield_time_ms":10000,"max_output_tokens":8000}); text(r.output)
-
- assistant #20 1 tool
- exec
tools/trace-capture.tsinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && sed -n \"1,260p\" tools/trace-capture.ts && ls traces | head'","yield_time_ms":10000,"max_output_tokens":10000}); text(r.output)
-
- assistant #21 1 tool
- exec
/tmp/latex_commands.pyinput
const r = await tools.exec_command({cmd:"scp /tmp/latex_commands.py beelink1-wsl:/tmp/latex_commands.py && ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 /tmp/latex_commands.py && git diff --check && git diff --name-only -- src/content/research src/content/posts | sort && git status --short'","yield_time_ms":10000,"max_output_tokens":4000}); text(r.output)
-
- assistant #22 1 tool
- exec
/tmp/fence_assignments.pyinput
const r = await tools.exec_command({cmd:"cat >/tmp/fence_assignments.py <<'PY'\nfrom pathlib import Path\nimport re\nfor root in ['src/content/posts','src/content/research']:\n for p in sorted(Path(root).glob('*.mdx')):\n s=p.read_text(); fm=re.match(r'^---\\n(.*?)\\n---\\n',s,re.S); orig=bool(fm and re.search(r'^original:\\s*true\\s*$',fm.group(1),re.M))\n if orig: continue\n fence=False;lang='';body=[];start=0\n for i,l in enumerate(s.splitlines(),1):\n if l.startswith('```'):\n if not fence: fence=True;lang=l[3:].strip();body=[];start=i\n else:\n if lang in ('text','tex','latex',''): \n fo…
-
- assistant #23 1 tool
- exec
/tmp/more_math.pyinput
const r = await tools.exec_command({cmd:"cat > /tmp/more_math.py <<'PY'\nfrom pathlib import Path\nR=Path('/home/drew/code/drewstone.github.io/src/content/posts')\n\ndef replace_fence(slug, old, new):\n p=R/(slug+'.mdx');s=p.read_text();pattern='```text\\n'+old+'\\n```'\n if pattern not in s: raise SystemExit(f'missing fence {slug}: {old[:50]!r}')\n p.write_text(s.replace(pattern,new,1))\n\nreplace_fence('self-improving-stack-test-time-compute', 'pass_hat@k = 1 - C(n - c, k) / C(n, k)', r'''$$\n\\widehat{\\operatorname{pass@k}}=1-\\frac{\\binom{n-c}{k}}{\\binom{n}{k}}\n$$''')\nreplace_fence('…
-
- assistant #24 1 tool
- exec
/tmp/count_display.pyinput
const r = await tools.exec_command({cmd:"cat >/tmp/count_display.py <<'PY'\nfrom pathlib import Path\nimport subprocess\nfiles=subprocess.check_output(['git','diff','--name-only','--','src/content/posts'],text=True).splitlines()\ndef cnt(s): return sum(1 for x in s.splitlines() if x.strip()=='$$')//2\noldtotal=newtotal=0\nfor f in files:\n old=subprocess.check_output(['git','show','HEAD:'+f],text=True);new=Path(f).read_text()\n a,b=cnt(old),cnt(new)\n oldtotal+=a;newtotal+=b\n print(f, 'old',a,'new',b,'delta',b-a)\nprint('TOTAL',oldtotal,newtotal,newtotal-oldtotal)\nPY\nscp /tmp/count_display…
-
- assistant #25 1 tool
- exec
src/content/posts/self-improving-stack-test-time-compute.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 - <<\"PY\"\nimport subprocess,re\nf=\"src/content/posts/self-improving-stack-test-time-compute.mdx\"\na=subprocess.check_output([\"git\",\"show\",\"HEAD:\"+f],text=True)\nb=open(f).read()\nprint(\"passat original code\",a.count(\"`pass@k`\"),\"current inline math\",b.count(r\"$\\\\operatorname{pass@k}$\"))\nPY'","yield_time_ms":10000,"max_output_tokens":1000}); text(r.output)
-
- assistant #26 1 tool
- exec
src/content/posts/self-improving-stack-test-time-compute.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 - <<\"PY\"\np=\"src/content/posts/self-improving-stack-test-time-compute.mdx\";s=open(p).read();print(s.count(\"$\\\\operatorname{pass@k}$\"));print([x for x in s.splitlines() if \"pass@k\" in x][:10])\nPY'","yield_time_ms":10000,"max_output_tokens":1000}); text(r.output)
-
- assistant #27 1 tool
- send_message
input
{"target":"/root","message":"gAAAAABqwDgtteQa8JHN8qGHc01H3vA8ld0JQ1GpvMFwmPmt6VJcqP6sAx8DMqTKgN_AMTXywW410TaF9neOy9PpP5CdJLrqvgVXUou5aXpe9VHWsOsMggC9jb37p5cWC9JiCHB3ZdjHepVsjZtQ023KFd27mVSvKq-7BnSn8EUtqS_yo4YdfuTMDAXKY6-FkBWKApmokhq2IrwA48iJwFZWXhko6Hi_f2bKTm76JDEbObbJKWy3i5chQmp0jSKAAo-o8-AJb7oQHGyCltNk2yFTzgOHrNJJhNbKE76zfRT6h_6uUplqAE8qEn27x21KcnRwWXYOK_38NqesRjJkufMhMJJoUU_PseH92dlvgFZAiTgXbZRJyK6yQexgu5Bz56LIU8uzer5nTsYLn_zCjOM-HP0s2XGPdocB743Vwe8Fs1MRa9Nj3wQqi2f0OllppwwiEAdIbpk-C6pod28NlAK0SnXvKbpzrOxAtED9WAVsvYwCFtvA56y9QKfbCw4TD-6Grl1uPHXV-EzPZq4t9PQgMrBKhgDmIJlYqSA6QBJRYDdKVikcOUFEcC…
-
- assistant #28 1 tool
- exec
src/content/posts/self-improving-stack-evaluation-gates.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 - <<\"PY\"\np=\"src/content/posts/self-improving-stack-evaluation-gates.mdx\";s=open(p).read().replace(\"V(x,y,\\\\tau)\",r\"V(x,y,\\\\text{trace})\");open(p,\"w\").write(s)\nPY'","yield_time_ms":10000,"max_output_tokens":1000}); text(r.output)
-
- assistant #29 1 tool
- exec
src/content/posts/self-improving-stack-evaluation-gates.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && rg -n \"V\\(x,y\" src/content/posts/self-improving-stack-evaluation-gates.mdx'","yield_time_ms":10000,"max_output_tokens":1000}); text(r.output)
-
- assistant #30 1 tool
- exec
src/content/posts/self-improving-stack-evaluation-gates.mdxinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && python3 - <<\"PY\"\np=\"src/content/posts/self-improving-stack-evaluation-gates.mdx\";s=open(p).read().replace(r\"V(x,y,\\\\text{trace})\",r\"V(x,y,\\text{trace})\");open(p,\"w\").write(s)\nPY\nrg -n \"V\\(x,y\" src/content/posts/self-improving-stack-evaluation-gates.mdx'","yield_time_ms":10000,"max_output_tokens":1000}); text(r.output)
-
- assistant #31 1 tool
- exec
/tmp/count_inline_math.pyinput
const r = await tools.exec_command({cmd:"cat >/tmp/count_inline_math.py <<'PY'\nfrom pathlib import Path\nimport subprocess,re,collections\nfiles=subprocess.check_output(['git','diff','--name-only','--','src/content/posts'],text=True).splitlines()\npattern=re.compile(r'(?<!\\$)\\$([^$\\n]+)\\$(?!\\$)')\nold=collections.Counter();new=collections.Counter()\nfor f in files:\n a=subprocess.check_output(['git','show','HEAD:'+f],text=True);b=Path(f).read_text()\n old.update(pattern.findall(a));new.update(pattern.findall(b))\nprint('new inline math span count:',sum((new-old).values()))\nfor k,v in (…
-
- assistant #32 1 tool
- exec
/tmp/make_math_inventory.pyinput
const r = await tools.exec_command({cmd:"cat >/tmp/make_math_inventory.py <<'PY'\nfrom pathlib import Path\nimport re,subprocess,collections\nrepo=Path('/home/drew/code/drewstone.github.io')\nposts=repo/'src/content/posts'\nout=Path('/tmp/blog-math-luna');out.mkdir(exist_ok=True)\nchanged=subprocess.check_output(['git','diff','--name-only','--','src/content/posts'],cwd=repo,text=True).splitlines()\nlines=['# Exact display math changes','']\ncount=0\nfor f in changed:\n p=repo/f\n s=p.read_text()\n blocks=[];inside=False;buf=[];start=0\n for i,line in enumerate(s.splitlines(),1):\n if line.str…
-
- assistant #33 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && pnpm blog write convergence-as-eval-primitive --harness=codex --role=polish | sed -n \"s/^Trace marker: //p\"'","yield_time_ms":10000,"max_output_tokens":500}); text(r.output)
-
- assistant #34 1 tool
- send_message
input
{"target":"/root","message":"gAAAAABqwDiwABymqxDMlGTmqrmi4RUW89n1ukDiFNGKhQLrzj_SycfDgEtYLRuC87GZyuRYmKzZ7ilsPlFji8IIddPha5oyJGSec_CSG0mTCOHPWcuEGI8XzVSjq_XrKSvLcwEQFgQGb1fLxXS4B_VuyZ5qLbpXGDc1t3PFnfIE_DOprAvh5Mh7Bt0qN9wQTNnSjokkLhsKtphSNHP26_k7CiMwZdQDBQ5cJAa-dZZmbcWOBImQdQ4KFTcnMYzTh5tp7ViybCe_bjJyc9VGkK7W15jtNHOQtjawZpKRYMHg1ifU-FbI4X2no_VqqjhGtDe-oGjtbAFgJNdcN1dEO5SayPO0KdzhHub9BUckx8cwzQ-54S3h9HtcixrlW1OkXcqU9PYGMtz6JYLcf78ny6kw-vIka5jAELehGT9lu0pWfDE9bS07EZtpDq8rxasiE8_rRZH6_N0N5T1Jm5b3Pf4-GPCoTUCMpTIFps3BOO-utSwJUEnKKhGKNW7SAAnvi4BCm7k57lN86fwyuRr5P6H6-KkePEARIRYxmiGJJ1Mu_y9zFT9daz0VF5…
-
- assistant #35I’m finishing the requested audit bundle under `/tmp/blog-math-luna`; the repo content is frozen, and this step only writes the summary and evidence inventory.
- assistant #36 1 tool
- exec
/tmp/blog-math-luna/changed-math.mdinput
const r = await tools.exec_command({cmd:"ls -la /tmp/blog-math-luna && sed -n '1,100p' /tmp/blog-math-luna/changed-math.md && sed -n '1,100p' /tmp/blog-math-luna/inline-math-spans.txt",max_output_tokens:2000}); text(r.output);
-
- assistant #37 1 tool
- exec
/tmp/blog-math-luna/changed-math.mdinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'ls -la /tmp/blog-math-luna && sed -n \"1,80p\" /tmp/blog-math-luna/changed-math.md && sed -n \"1,80p\" /tmp/blog-math-luna/inline-math-spans.txt'","max_output_tokens":2500}); text(r.output);
-
- assistant #38 1 tool
- exec
input
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cd /home/drew/code/drewstone.github.io && git status --short -- src/content/posts src/content/research && git diff --stat -- src/content/posts src/content/research && git diff --check -- src/content/posts src/content/research'","max_output_tokens":1400}); text(r.output);
-
- assistant #39 1 tool
- exec
/tmp/blog-math-luna/audit.mdinput
const r = await tools.exec_command({cmd:"ssh beelink1-wsl 'cat > /tmp/blog-math-luna/audit.md <<'\\''EOF'\\''\n# Blog math formatting audit\n\nSource checkout: `/home/drew/code/drewstone.github.io`\nBranch/baseline: `fix/readable-research-layouts` at `e5a9810`\nScope: `src/content/posts` and `src/content/research` only.\n\n## Result\n\nReviewed all post and research MDX content for mathematical notation placed in `text` fences or inline code. Converted clear mathematical expressions to KaTeX math delimiters while preserving meaning. Kept actual code, pseudocode, commands, API identifiers, con…
-
- assistant #40 1 tool
- exec
src/content/posts/convergence-as-eval-primitive.mdxinput
const content = `# Blog math formatting audit Source checkout: \`/home/drew/code/drewstone.github.io\` Branch/baseline: \`fix/readable-research-layouts\` at \`e5a9810\` Scope: \`src/content/posts\` and \`src/content/research\` only. ## Result Reviewed all post and research MDX content for mathematical notation placed in \`text\` fences or inline code. Converted clear mathematical expressions to KaTeX math delimiters while preserving meaning. Kept actual code, pseudocode, commands, API identifiers, configuration/schema examples, and promotion predicates as code. No prose was rewritten and no c…
-
Diff
Per-file changes from a68bc06.
#!/usr/bin/env node/** * blog-loop — lowest-friction lifecycle commands for traced blog work. * * Usage: * pnpm blog research <post> [--harness=codex|claude-code] * pnpm blog write <post> [--harness=codex|claude-code] [--role=draft|rewrite|polish|outline|review|publish] [--marker=<token>] * pnpm blog finish <post> --research --harness=codex --note="source scan" * pnpm blog finish <post> --write --harness=codex --marker=<token> --note="drafted section" */import { existsSync } from 'node:fs'import { readdir, readFile } from 'node:fs/promises'import { spawnSync } from 'node:child_process'import { join, resolve } from 'node:path'import { fileURLToPath } from 'node:url'const REPO = resolve(fileURLToPath(new URL('..', import.meta.url)))const POSTS_DIR = join(REPO, 'src', 'content', 'posts')function parse(argv) { const [cmd, ...rest] = argv const flags = {} const pos = [] for (const a of rest) { if (a.startsWith('--')) { const eq = a.indexOf('=') if (eq > -1) flags[a.slice(2, eq)] = a.slice(eq + 1) else flags[a.slice(2)] = true } else { pos.push(a) } } return { cmd, post: pos.join(' ').trim(), flags }}function field(fm, name) { const m = fm.match(new RegExp(`^${name}:\\s*(.+)$`, 'm')) if (!m) return null return m[1].trim().replace(/^['"]|['"]$/g, '').replace(/''/g, "'")}async function posts() { const files = (await readdir(POSTS_DIR)).filter((f) => f.endsWith('.mdx')).sort() const out = [] for (const file of files) { const path = join(POSTS_DIR, file) const raw = await readFile(path, 'utf8') const m = raw.match(/^---\n([\s\S]*?)\n---\n/) if (!m) continue out.push({ slug: file.replace(/\.mdx$/, ''), title: field(m[1], 'title') ?? file.replace(/\.mdx$/, ''), draft: field(m[1], 'draft') !== 'false', original: field(m[1], 'original') === 'true', }) } return out}function resolvePost(all, query) { if (!query) return null const q = query.toLowerCase() const exact = all.find((p) => p.slug === q) if (exact) return exact const hits = all.filter((p) => p.slug.includes(q) || p.title.toLowerCase().includes(q)) if (hits.length === 1) return hits[0] if (hits.length > 1) return { ambiguous: hits } return null}function defaultMarker(post, role) { const stamp = new Date().toISOString().replace(/[:.]/g, '-') return `BLOGTRACE-${post.slug}-${role}-${stamp}`}function shellEscapeDoubleQuoted(value) { return value.replace(/\\/g, '\\\\').replace(/"/g, '\\"').replace(/\$/g, '\\$').replace(/`/g, '\\`')}function printPrompt(mode, post, harness, role, marker) {---title: 'A five-atom counterexample to BCWW'description: 'An exact classical witness violating equation (4.6) of Bao, Cao, Walter, and Wang, with corrections and retained verification records.'assessed: '2026-10-02'claim: 'Five equally likely bit strings violate a proposed four-party entropy inequality by 0.0529325 bits.'limitation: 'Corpus-origin witness; first discovery is unattributed. The original archived smoke witness was wrong.'record_ids: ['ftqc-zips-2026-08-05-r1', 'fc-audit-smoke-a', 'q-ca-bcww-shareable-pkg']order: 1---import WitnessExplorer from '../../components/WitnessExplorer.astro'import ResearchBibliography from '../../components/ResearchBibliography.astro'Five equally likely bit strings violate the four-party inequality proposed by Bao, Cao, Walter, and Wang [[1]](#ref-1).The violation is exact: its sign reduces to $3^9>2^{14}$.## InequalityLet $S(X)$ denote the entropy of subsystem $X$, with logarithms in base two.Equation (4.6) of [[1]](#ref-1), numbered (4.16) in its first version, proposes $F\leq0$, where$$\begin{aligned}F={}&S(ABD)+S(ABC)+S(BCD)\\&-2S(BD)-2S(BC)+S(CD)\\&-S(AD)-S(AC)-S(AB)+2S(B)+S(A).\end{aligned}$$The counterexample concerns this candidate bound for general quantum states; the paper's holographic entropy results are unaffected.## CounterexampleLet $(A,B,C,D)$ be uniformly distributed on$$\{0001,\;0010,\;0011,\;0111,\;1000\}.$$Each atom has probability $1/5$.The corresponding diagonal quantum state has the same marginal entropies, so a classical violation suffices.The nonzero marginal multiplicities determine all eleven entropy terms:| Subsystem | Multiplicities out of five || --- | --- || $A$, $B$ | $4,1$ || $AB$, $AC$, $AD$ | $3,1,1$ || $BC$, $BD$ | $2,2,1$ || $CD$ | $2,1,1,1$ || $ABC$, $ABD$ | $2,1,1,1$ || $BCD$ | $1,1,1,1,1$ |Substitute $H=\log_2 5-\frac15\sum_j n_j\log_2 n_j$ to obtain$$F=\frac{9\log_2 3-14}{5} =0.052932501298\ldots>0,\qquad3^9=19683>16384=2^{14}.$$The public repository includes a numerical checker and Lean proof material [[2]](#ref-2).The integer comparison proves the sign without floating point; the browser calculation below illustrates it.<WitnessExplorer kind="bcww" />## Correction and attributionThe earlier smoke archive [[4]](#ref-4) used $\{0000,0010,0100,1000,1101\}$, which does **not** violate the bound in the stated party order.The September 3 correction [[5]](#ref-5) records the inconsistency and supplies the atoms above.The widget compares both sets.The corrected point appeared in an external entropy corpus dated August 4, 2026; its generation history is absent from the inspected evidence.First discovery remains unattributed.The August 5 corpus-verification campaign [[3]](#ref-3) retained six Pi session files and a journal recording a director and five children.Its native files report GLM-5.2, contradicting an older DeepSeek attribution.That campaign also checked other quantum-information artifacts; it is a verification trace, not the initial search for this point.The August 13 smoke archive has no native transcript, and the September 3 package retains findings without a recovered complete source run.The retained records therefore do not reconstruct the entire discovery and correction history.The [GHZ-mixture construction](/research/ghz/) reaches $F\approx0.113302$ bits against the same bound.---title: 'A GHZ-mixture witness against BCWW'description: 'An explicit four-qubit state violating the BCWW candidate inequality, with an exact entropy reduction and a replayable verifier.'assessed: '2026-10-02'claim: 'A four-qubit GHZ mixture violates the same proposed inequality by 0.113302 bits.'limitation: 'The relevant marginals are classical. This does not establish a quantum advantage; the original discovery fleet trace is incomplete.'record_ids: ['q-conj-entropy-bcww-repair']order: 2---import WitnessExplorer from '../../components/WitnessExplorer.astro'import ResearchBibliography from '../../components/ResearchBibliography.astro'A mixture of a four-qubit GHZ state and two basis states violates the BCWW candidate inequality [[1]](#ref-1) by $0.113302\ldots$ bits.The relevant marginal entropies are classical, so this construction does not establish a quantum advantage.## State and marginal entropiesDefine$$|\mathrm{GHZ}_4\rangle=\frac{|0000\rangle+|1111\rangle}{\sqrt2},\qquad\rho=\frac34|\mathrm{GHZ}_4\rangle\!\langle\mathrm{GHZ}_4|+\frac18|1010\rangle\!\langle1010|+\frac18|1001\rangle\!\langle1001|.$$The [BCWW functional](/research/bcww/#inequality) uses only proper subsystems.Tracing out a nonempty complement removes the GHZ off-diagonal term.Every required entropy therefore equals the Shannon entropy of the corresponding marginal of$$P(0000)=P(1111)=\frac38,\qquadP(1010)=P(1001)=\frac18.$$With logarithms in base two, substitution gives$$F=\frac{25}{4}-\frac98\log_2 3-\frac{15}{8}\log_2 5 =0.113302008774\ldots>0.$$The exact sign certificate is$$2^{50}=1125899906842624>600677490234375=3^9 5^{15}.$$The retained standard-library verifier reconstructs each marginal and checks the closed form [[2]](#ref-2).Its October 2 replay succeeded.<WitnessExplorer kind="ghz" />The slider assigns weight $p$ to the GHZ projector and $(1-p)/2$ to each basis-state projector.The integer certificate applies at $p=3/4$; other plotted values use floating point.## Attribution and retained attemptThe provenance note credits the fifth arm of an operator-led verification fleet dated August 10, 2026, and records five verification approaches, including blind arithmetic and fresh code.The complete original fleet transcripts were not recovered, so this attribution cannot be reconstructed from a complete discovery trace.The retained pursuit [[3]](#ref-3) began on August 30 and revisited the witness and surrounding claim.Its two child attempts ended down after rate-limit and incomplete-stream failures.The archive contains journal events and the witness artifact, but no native session JSONL.The viewer shows this later repair attempt, not the August 10 fleet; it does not certify complete capture of the original messages, children, or source bytes.This is a second witness to the [same BCWW refutation](/research/bcww/).<ResearchBibliography entries={[ { authors: 'N. Bao, C. Cao, M. Walter, and Z. Wang', title: 'Holographic entropy inequalities and gapped phases of matter', url: 'https://arxiv.org/abs/1507.05650v2', publication: 'J. High Energy Phys. 09 (2015), 203. arXiv:1507.05650v2, Eq. (4.6).', doi: '10.1007/JHEP09(2015)203' }, { title: 'GHZ-mixture entropy verifier', url: '/research/verify-ghz.py', publication: 'Exact retained Python source. Standard-library reconstruction and closed-form check; replayed October 2, 2026.' }, { title: 'BCWW repair pursuit', url: '/research/records/q-conj-entropy-bcww-repair.json', publication: 'Retained event and finding projection, August 30, 2026. A later attempt, distinct from the August 10 verification fleet. Source hashes identify the inspected files.' },]} />---title: 'A family refuting Linden–Winter Conjecture 6'description: 'A six-atom classical family with two exact conditional-independence identities and an unbounded entropy ratio.'assessed: '2026-10-02'claim: 'A six-atom family defeats every finite positive choice of coefficients in the conjectured linear relaxation.'limitation: 'The constrained Linden–Winter theorem survives. The public note is not presented as peer-reviewed, and the discovery history remains partial.'record_ids: ['q-q36-lw6-unbounded-ladder']order: 3---import WitnessExplorer from '../../components/WitnessExplorer.astro'import ResearchBibliography from '../../components/ResearchBibliography.astro'A six-atom classical family refutes Conjecture 6 of Linden and Winter [[1]](#ref-1): no finite positive coefficient triple makes their proposed linear relaxation valid for all four-party quantum states.Their constrained theorem remains valid.## ConjectureLinden and Winter proved a four-party entropy inequality under three linear entropy constraints.Conjecture 6 asks whether positive constants $k_1,k_2,k_3$ can make$$\begin{aligned}f_k={}&k_1 I(A;C\mid B)+k_2 I(C;B\mid A)\\&+k_3 I(A;B\mid D)+I(C;D)-I(C;AB)\geq0\end{aligned}$$hold for every four-party quantum state.Here $I(X;Y\mid Z)=H(XZ)+H(YZ)-H(Z)-H(XYZ)$ denotes conditional mutual information.Classical distributions embed as diagonal quantum states, so a classical family suffices.## Six-atom familyFor an integer $L\geq2$, normalize these weights in $(A,B,C,D)$ order:| Atom | Integer weight || --- | --- || $0000$ | $2\cdot10^{4L+3}$ || $0010$ | $27\cdot10^{4L+3}-27\cdot10^{3L+3}$ || $1101$ | $2\cdot10^{4L}$ || $1111$ | $27\cdot10^{4L}$ || $0101$ | $2\cdot10^{L+3}$ || $0110$ | $27\cdot10^{L+3}$ |Given $D=0$, $A$ is constant; given $D=1$, $B$ is constant.Thus $I(A;B\mid D)=0$.For $B=0$, $A$ is constant; for $B=1$, the identity$$w(1101)w(0110)=w(1111)w(0101)=54\cdot10^{5L+3}$$establishes conditional independence of $A$ and $C$, hence $I(A;C\mid B)=0$.Writing $y_L=I(C;B\mid A)$ and $e_L=I(C;D)-I(C;AB)$ leaves $f_k=k_2y_L+e_L$.The values of $k_1$ and $k_3$ have no effect on this family.## Unbounded ratioSet $\delta=10^{-L}$.The equivalent weights are $2,27(1-\delta),2/1000,27/1000,2\delta^3,27\delta^3$, with sum $Z=29029/1000-27\delta+29\delta^3$.Using natural logarithms, define$$\Phi(a,b)=(a+b)\ln(a+b)-a\ln a-b\ln b.$$Conditional-entropy grouping gives$$\begin{aligned}Zy_L={}&\Phi(2+2\delta^3,27(1-\delta)+27\delta^3)\\&-\Phi(2,27(1-\delta))-\delta^3\Phi(2,27),\\[3pt]Ze_L={}&\Phi(2,27(1-\delta))+\Phi(2/1000,27/1000)\\&+\delta^3\Phi(2,27)-\Phi(2,27(1-\delta)+27\delta^3)\\&-\Phi(2/1000+2\delta^3,27/1000).\end{aligned}$$All arguments of $\Phi$ stay positive near zero, making these expressions analytic there.Their Taylor expansions are#!/usr/bin/env node/** * trace-capture: harness-agnostic session capture for blog revisions. * * Usage: * pnpm tsx tools/trace-capture.ts capture \ * [--harness=claude-code|codex|manual] \ * [--post=<slug>] \ * [--role=outline|draft|rewrite|polish|diagram|review|publish|research] \ * [--session=<session-id>] \ * [--marker="<token>"] \ * [--note=<one-line>] \ * [--commit=<sha>] \ * [--input=<path>] # manual harness only * [--kind=post|series-outline|supporting-research] * [--attach=supporting|revision|none] * [--latest] # choose latest session without requiring post file touch * * pnpm tsx tools/trace-capture.ts capture --auto * # detects from the latest git commit: finds changed posts, matches a * # recent session via ~/.claude/projects or ~/.codex/sessions, writes a * # trace per changed post, appends to frontmatter. * * pnpm tsx tools/trace-capture.ts list * # list existing traces grouped by post. * * pnpm tsx tools/trace-capture.ts show <trace_id> * # dump a trace as JSON. */import { execSync } from 'node:child_process'import { mkdir, readdir, readFile, writeFile } from 'node:fs/promises'import { join } from 'node:path'import ClaudeCodeHarness from './harness/claude-code.js'import CodexHarness from './harness/codex.js'import ManualHarness from './harness/manual.js'import { dedupeAdjacentTurns, type TraceFile, type TraceHarness, type Turn } from './harness/types.js'const ROOT = process.cwd()const POSTS_DIR = join(ROOT, 'src/content/posts')const TRACES_DIR = join(ROOT, 'traces')type Args = Record<string, string | boolean>function parseArgs(argv: string[]): { cmd: string; pos: string[]; flags: Args } { const [cmd, ...rest] = argv const flags: Args = {} const pos: string[] = [] for (const a of rest) { if (a.startsWith('--')) { const eq = a.indexOf('=') if (eq >= 0) flags[a.slice(2, eq)] = a.slice(eq + 1) else flags[a.slice(2)] = true } else pos.push(a) } return { cmd: cmd ?? 'capture', pos, flags }}function git(cmd: string): string { try { return execSync(`git ${cmd}`, { cwd: ROOT, stdio: ['ignore', 'pipe', 'ignore'] }).toString().trim() } catch { return '' }}function headCommit(): string | null { const sha = git('rev-parse HEAD') return sha || null}function changedPostsAtHead(): string[] { const out = git('show --no-renames --name-only --format="" HEAD') return out .split('\n') .map((l) => l.trim()) .filter((l) => l.startsWith('src/content/posts/') && l.endsWith('.mdx')) .map((l) => l.replace('src/content/posts/', '').replace(/\.mdx$/, ''))}diff --git a/src/content/posts/self-improving-stack-test-time-compute.mdx b/src/content/posts/self-improving-stack-test-time-compute.mdxindex c9bd3e9..bff211f 100644--- a/src/content/posts/self-improving-stack-test-time-compute.mdx+++ b/src/content/posts/self-improving-stack-test-time-compute.mdx@@ -79,25 +79,25 @@ The key is that all of these are inference-time choices. They do not change mode Let: -```text-x = task-m = fixed model or model set-h = harness and runtime-s = strategy for spending test-time compute-B = compute budget-R = reward, score, or task success-C = measured cost-```+$$+\begin{aligned}+ x &= \text{task} \\+ m &= \text{fixed model or model set} \\+ h &= \text{harness and runtime} \\+ s &= \text{strategy for spending test-time compute} \\+ B &= \text{compute budget} \\+ R &= \text{reward, score, or task success} \\+ C &= \text{measured cost}+\end{aligned}+$$ The objective is: -```text-J(s | m, h, B) =- E_{x ~ D}[R(run(m, h, s, x, B))]- - lambda * E[C(run(m, h, s, x, B))]-```+$$+J(s\mid m,h,B)=\mathbb{E}_{x\sim D}[R(\operatorname{run}(m,h,s,x,B))]-\lambda\,\mathbb{E}[C(\operatorname{run}(m,h,s,x,B))]+$$ -The budget `B` is not one scalar in practice. It is a vector:+The budget $B$ is not one scalar in practice. It is a vector: ```text B = {@@ -113,7 +113,7 @@ B = { } ``` -A fair comparison fixes the relevant parts of `B`, or it reports the tradeoff instead of pretending the strategy itself improved.+A fair comparison fixes the relevant parts of $B$, or it reports the tradeoff instead of pretending the strategy itself improved. ## The Baseline Ladder @@ -123,70 +123,78 @@ Every topology claim should climb a baseline ladder. The weakest baseline is one ordinary run: -```text-y_1 ~ q_m(y | x)-score = R(y_1)-```+$$+\begin{aligned}+ y_1&\sim q_m(y\mid x) \\+ \text{score}&=R(y_1)+\end{aligned}+$$ Beating this is not enough. Almost any extra compute can beat one sample on tasks with stochastic failures. **Random@k** -Sample `k` candidates from the same model and prompt distribution, then select without additional information or with a fixed production selector:+Sample $k$ candidates from the same model and prompt distribution, then select without additional information or with a fixed production selector: -```text-y_i ~ q_m(y | x), i = 1..k-y_hat = sigma_blind({y_i})-```+$$+\begin{aligned}+ y_i&\sim q_m(y\mid x),\quad i=1,\ldots,k \\+ \hat{y}&=\sigma_{\text{blind}}(\{y_i\})+\end{aligned}+$$ This asks: what happens if we just buy more attempts? **Pass@k** -For verifiable tasks, `pass@k` asks whether any candidate succeeds:+For verifiable tasks, $\operatorname{pass@k}$ asks whether any candidate succeeds: -```text-pass@k = P(max_i R(y_i) = 1)-```+$$+\operatorname{pass@k}=\mathbb{P}\!\left(\max_iR(y_i)=1\right)+$$ -Under independent binary success probability `q`:+Under independent binary success probability $q$: -```text-pass@k = 1 - (1 - q)^k-```+$$+\operatorname{pass@k}=1-(1-q)^k+$$ -That formula is the reason repeated sampling is hard to dismiss. Even a weak model can look strong if the task is verifiable and `k` is large enough.+That formula is the reason repeated sampling is hard to dismiss. Even a weak model can look strong if the task is verifiable and $k$ is large enough. -In code-eval practice, `pass@k` is often estimated from `n` sampled candidates where `c` pass the tests:+In code-eval practice, $\operatorname{pass@k}$ is often estimated from $n$ sampled candidates where $c$ pass the tests: -```text-pass_hat@k = 1 - C(n - c, k) / C(n, k)-```+$$+\widehat{\operatorname{pass@k}}=1-\frac{\binom{n-c}{k}}{\binom{n}{k}}+$$ -where `C(a, b)` is the binomial coefficient, with `C(a, b) = 0` when `a < b`. This estimates the chance that at least one of `k` drawn samples would pass, without pretending the evaluator can deploy the answer key.+where $\binom{a}{b}$ is the binomial coefficient, with $\binom{a}{b}=0$ when $a<b$. This estimates the chance that at least one of $k$ drawn samples would pass, without pretending the evaluator can deploy the answer key. -But `pass@k` is not a deployable policy. It is an oracle coverage metric. It tells you whether a correct answer existed in the sample set, not whether your system could find it without labels.+But $\operatorname{pass@k}$ is not a deployable policy. It is an oracle coverage metric. It tells you whether a correct answer existed in the sample set, not whether your system could find it without labels. **Best-of-N with a selector** Now add a selector: -```text-y_hat = sigma(x, {y_1, ..., y_k})-score = R(y_hat)-```+$$+\begin{aligned}+ \hat{y}&=\sigma(x,\{y_1,\ldots,y_k\}) \\+ \text{score}&=R(\hat{y})+\end{aligned}+$$ -This is production-like only if `sigma` is available at deployment time. A hidden answer key, private unit tests, or human judge may be useful for measurement. It is not a runtime selector unless the product can actually call it.+This is production-like only if $\sigma$ is available at deployment time. A hidden answer key, private unit tests, or human judge may be useful for measurement. It is not a runtime selector unless the product can actually call it. **Self-consistency** Self-consistency samples multiple reasoning paths and chooses the answer supported by the largest mass: -```text-z_i = reasoning path-a_i = final answer extracted from z_i-y_hat = argmax_a count(a_i = a)-```+$$+\begin{aligned}+ z_i&=\text{reasoning path} \\+ a_i&=\text{final answer extracted from }z_i \\+ \hat{y}&=\operatorname*{argmax}_a\operatorname{count}(a_i=a)+\end{aligned}+$$ It is powerful when correct answers are stable attractors and wrong answers are diverse. It is weaker for open-ended artifact quality, where many outputs are plausible and no answer string gets a majority. @@ -194,11 +202,11 @@ It is powerful when correct answers are stable attractors and wrong answers are A verifier or reward model scores candidates: -```text-y_hat = argmax_i V(x, y_i, trace_i)-```+$$+\hat{y}=\operatorname*{argmax}_i V(x,y_i,\operatorname{trace}_i)+$$ -This is where process reward models, unit tests, static analyzers, rubric judges, and domain verifiers enter. The selector becomes useful only to the extent that `V` correlates with true task success and does not overfit superficial features.+This is where process reward models, unit tests, static analyzers, rubric judges, and domain verifiers enter. The selector becomes useful only to the extent that $V$ correlates with true task success and does not overfit superficial features. **Guided compute** @@ -234,16 +242,18 @@ Could the system identify that candidate without oracle labels? The two curves can be very different. -```text-coverage_k = P(exists i: R(y_i) = 1)-selection_k = E[R(sigma({y_i}))]-```+$$+\begin{aligned}+ \operatorname{coverage}_k&=\mathbb{P}(\exists i:R(y_i)=1) \\+ \operatorname{selection}_k&=\mathbb{E}[R(\sigma(\{y_i\}))]+\end{aligned}+$$ The gap is selector loss: -```text-selector_loss_k = coverage_k - selection_k-```+$$+\operatorname{selector\_loss}_k=\operatorname{coverage}_k-\operatorname{selection}_k+$$ A system with high coverage and high selector loss is not ready. It can generate a correct answer somewhere in the pile, but it cannot reliably ship the right artifact. @@ -361,7 +371,7 @@ true_score = R(x, y) verifier_error = observed_score - true_score ``` -If guided search optimizes `V` faster than `V` tracks `R`, the system reward-hacks its own evaluator.+If guided search optimizes $V$ faster than $V$ tracks $R$, the system reward-hacks its own evaluator. This is why the selector must be evaluated, not assumed. @@ -392,7 +402,7 @@ The topology claim is: allocation_policy_multi(B) > allocation_policy_baseline(B) ``` -at a measured budget `B`.+at a measured budget $B$. That is why the previous post insisted that role names are not enough. The role structure matters only if it improves allocation, evidence, selection, or verification under budget. @@ -516,7 +526,7 @@ The candidate wins because it used more samples, turns, tokens, or tools. **Oracle selection** -The paper or eval reports `pass@k`, but the product has no selector that can find the passing candidate.+The paper or eval reports $\operatorname{pass@k}$, but the product has no selector that can find the passing candidate. **Verifier overfit** diff --git a/src/content/posts/self-improving-stack-evaluation-gates.mdx b/src/content/posts/self-improving-stack-evaluation-gates.mdxindex 06564b3..6f14348 100644--- a/src/content/posts/self-improving-stack-evaluation-gates.mdx+++ b/src/content/posts/self-improving-stack-evaluation-gates.mdx@@ -57,18 +57,20 @@ A gate is a promotion policy. Let: -```text-b = baseline system-c = candidate system-x = scenario-p = agent profile cell-z = seed or replicate id-R = task reward or score-C = measured cost vector-T = trace integrity predicate-D_search = search split-D_holdout = held-out split-```+$$+\begin{aligned}+ b &= \text{baseline system} \\+ c &= \text{candidate system} \\+ x &= \text{scenario} \\+ p &= \text{agent profile cell} \\+ z &= \text{seed or replicate id} \\+ R &= \text{task reward or score} \\+ C &= \text{measured cost vector} \\+ T &= \text{trace integrity predicate} \\+ D_{\text{search}} &= \text{search split} \\+ D_{\text{holdout}} &= \text{held-out split}+\end{aligned}+$$ The gate is a function: @@ -129,9 +131,9 @@ Suppose the baseline sees one sample of tasks and the candidate sees another. A The paired comparison fixes that: -```text-delta_i = R(c, x_i, p_i, z_i) - R(b, x_i, p_i, z_i)-```+$$+\Delta_i = R(c,x_i,p_i,z_i)-R(b,x_i,p_i,z_i)+$$ where the candidate and baseline are evaluated on the same scenario, profile, and replicate. Then the question becomes: @@ -143,12 +145,11 @@ The median is useful because agent scores often have heavy tails. One catastroph A practical promotion rule: -```text-n_pairs >= n_min-LCB_95(median(delta)) > epsilon-```+$$+n_{\text{pairs}} \ge n_{\min},\qquad \operatorname{LCB}_{95}(\operatorname{median}(\Delta)) > \epsilon+$$ -`LCB_95` is the lower confidence bound. If the lower bound clears the threshold, the gate has evidence that the lift is not just random luck.+$\operatorname{LCB}_{95}$ is the lower confidence bound. If the lower bound clears the threshold, the gate has evidence that the lift is not just random luck. This is where bootstrap confidence intervals are useful. You resample paired deltas, compute the median for each resample, and inspect the lower quantile: @@ -157,10 +158,12 @@ delta = [delta_1, ..., delta_n] for r in 1..B: sample n deltas with replacement m_r = median(sample)--LCB_95 = quantile({m_r}, 0.025) // two-sided 95 interval ``` +$$+\operatorname{LCB}_{95}=Q_{0.025}(m_1,\ldots,m_B)+$$+ The gate promotes only when the pessimistic estimate is still good enough. ## Search Split Versus Holdout@@ -185,14 +188,18 @@ score(c, holdout) goes down So the gate needs an overfit check: -```text-gap_c = mean_score(c, search) - mean_score(c, holdout)-gap_b = mean_score(b, search) - mean_score(b, holdout)+$$+\begin{aligned}+ \operatorname{gap}_c &= \operatorname{mean\_score}(c,\text{search})-\operatorname{mean\_score}(c,\text{holdout}) \\+ \operatorname{gap}_b &= \operatorname{mean\_score}(b,\text{search})-\operatorname{mean\_score}(b,\text{holdout})+\end{aligned}+$$ +```text reject(c) if gap_c > gap_b + tau ``` -This says the candidate may look better on the search split, but it cannot be much more search-specialized than the baseline. `tau` is slack, not forgiveness. It accounts for sampling noise and legitimate split difficulty differences.+This says the candidate may look better on the search split, but it cannot be much more search-specialized than the baseline. $\tau$ is slack, not forgiveness. It accounts for sampling noise and legitimate split difficulty differences. The gate configuration has to be fixed before the candidate is scored: @@ -289,28 +296,30 @@ then evaluate the judge as a measurement instrument. Let: -```text-V(x, y, trace) = judge score-R(x, y) = true product outcome-```+$$+\begin{aligned}+V(x,y,\text{trace}) &= \text{judge score} \\+R(x,y) &= \text{true product outcome}+\end{aligned}+$$ A judge gate needs calibration: -```text-calibration_error = E[R | V = s] - s-```+$$+\operatorname{calibration\_error}=\mathbb{E}[R\mid V=s]-s+$$ and agreement: -```text-agreement = P(sign(V_a - V_b) = sign(H_a - H_b))-```+$$+\operatorname{agreement}=\mathbb{P}(\operatorname{sign}(V_a-V_b)=\operatorname{sign}(H_a-H_b))+$$ -where `H` is a human preference or trusted adjudicator. For ordinal rubrics, rank correlation is often more useful than raw score correlation:+where $H$ is a human preference or trusted adjudicator. For ordinal rubrics, rank correlation is often more useful than raw score correlation: -```text-rho = Spearman(V, H)-```+$$+\rho=\operatorname{Spearman}(V,H)+$$ A gate should track judge drift over time. If the judge model, rubric, prompt, or examples change, the score distribution can move even when the agent behavior does not. @@ -356,11 +365,15 @@ This prevents aggregate masking. A candidate can improve the mean while regressi Regression detection should combine effect size and statistical confidence: -```text-delta = current - baseline-d = CohenD(current_scores, baseline_scores)-p = WelchT(current_scores, baseline_scores)+$$+\begin{aligned}+ \Delta &= \text{current}-\text{baseline} \\+ d &= \operatorname{CohenD}(\text{current\_scores},\text{baseline\_scores}) \\+ p &= \operatorname{WelchT}(\text{current\_scores},\text{baseline\_scores})+\end{aligned}+$$ +```text regressed if delta < 0 and abs(d) >= d_min and p <= alpha ``` diff --git a/src/content/posts/convergence-as-eval-primitive.mdx b/src/content/posts/convergence-as-eval-primitive.mdxindex 91585ea..1043450 100644--- a/src/content/posts/convergence-as-eval-primitive.mdx+++ b/src/content/posts/convergence-as-eval-primitive.mdx@@ -59,7 +59,7 @@ Two things are now visible that a pass/fail eval cannot show. First, the shape o ## Monotone progress, resumable runs -The second move is making completion% *monotone*. Once a criterion is satisfied at turn `k`, it stays satisfied at turn `k+1`. Progress only moves up.+The second move is making completion% *monotone*. Once a criterion is satisfied at turn `k`, it stays satisfied at turn $k+1$. Progress only moves up. This sounds obvious until you realize most eval harnesses don't do it. They re-check everything on every turn and let earlier wins regress if the agent says something contradictory. That's the wrong model. The run is an accumulation: at any point, the agent has done some subset of the things the scenario asks for, and we want to track that subset growing. @@ -246,7 +246,7 @@ Most eval harnesses treat each turn as a re-measurement. That's fine for one-sho Enforce monotonicity at the harness level. Once a criterion is marked satisfied, it stays satisfied for the duration of the run. The agent cannot un-solve its own progress; the rubric cannot second-guess itself. <Callout tone="ok" title="The small code that buys a lot">-Store satisfied criteria as a `Set<criterionId>` accumulated across turns. On each turn, only call checkers for criteria not already in the set. Completion% is `|satisfied| / |total|` weighted by criterion weight. Cheaper to compute each turn and immune to regression.+Store satisfied criteria as a `Set<criterionId>` accumulated across turns. On each turn, only call checkers for criteria not already in the set. Completion% is $|\text{satisfied}|/|\text{total}|$ weighted by criterion weight. Cheaper to compute each turn and immune to regression. </Callout> ## Resuming a rundiff --git a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx b/src/content/posts/self-improving-stack-agent-runtime-topology.mdxindex 8b1c5da..4d741f2 100644--- a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx+++ b/src/content/posts/self-improving-stack-agent-runtime-topology.mdx@@ -86,7 +86,7 @@ A_runtime = { } ``` -If `parallel` is not in `A_runtime`, no optimized prompt can make true parallelism appear. If `checkpoint` and `replay` are absent, a long-running agent has no durable execution boundary. If `select` is only an LLM preference expressed in prose, the system has no enforceable winner rule.+If `parallel` is not in $A_{\text{runtime}}$, no optimized prompt can make true parallelism appear. If `checkpoint` and `replay` are absent, a long-running agent has no durable execution boundary. If `select` is only an LLM preference expressed in prose, the system has no enforceable winner rule. The deep question is: @@ -100,25 +100,27 @@ That one distinction determines whether multi-agent self-improvement is engineer Let: -```text-g = runtime topology-pi = runtime policy over moves in A_runtime-m = model/backend set-p = prompts and role descriptions-k = active skills-u = tools and external affordances-x = task from distribution D-R = trajectory reward or eval score-C = cost, latency, compute, human review, or risk-```+$$+\begin{aligned}+ g &= \text{runtime topology} \\+ \pi &= \text{runtime policy over moves in } A_{\text{runtime}} \\+ m &= \text{model/backend set} \\+ p &= \text{prompts and role descriptions} \\+ k &= \text{active skills} \\+ u &= \text{tools and external affordances} \\+ x &= \text{task from distribution } D \\+ R &= \text{trajectory reward or eval score} \\+ C &= \text{cost, latency, compute, human review, or risk}+\end{aligned}+$$ Runtime topology optimization estimates: -```text-J(g, pi | m, p, k, u) = E_{x ~ D}[R(run(g, pi, m, p, k, u, x))] - lambda * E[C(run(g, pi, m, p, k, u, x))]-```+$$+J(g, \pi \mid m,p,k,u) = \mathbb{E}_{x\sim D}[R(\operatorname{run}(g,\pi,m,p,k,u,x))] - \lambda\,\mathbb{E}[C(\operatorname{run}(g,\pi,m,p,k,u,x))]+$$ -Prompt and skill optimization usually keep `g` fixed. Topology optimization changes `g`, `pi`, or both.+Prompt and skill optimization usually keep $g$ fixed. Topology optimization changes $g$, $\pi$, or both. This is not gradient descent over a dense parameter tensor. It is discrete search over executable program structure: graph edges, worker counts, branch policies, selectors, validators, budget ledgers, replay boundaries, and promotion gates. diff --git a/src/content/posts/self-improving-stack-governance.mdx b/src/content/posts/self-improving-stack-governance.mdxindex 8562ae9..3d297dd 100644--- a/src/content/posts/self-improving-stack-governance.mdx+++ b/src/content/posts/self-improving-stack-governance.mdx@@ -95,35 +95,40 @@ If any part of that trajectory can mutate future behavior, it belongs in the saf Let: -```text-h = harness and runtime-s = mutable surface-c = candidate change-tau(c) = trace produced by candidate c-R(tau) = utility or quality score-K(tau) = risk vector-A(tau) = authority exercised by the agent-G = promotion gate-```+$$+\begin{aligned}+ h &= \text{harness and runtime} \\+ s &= \text{mutable surface} \\+ c &= \text{candidate change} \\+ \tau(c) &= \text{trace produced by candidate }c \\+ R(\tau) &= \text{utility or quality score} \\+ K(\tau) &= \text{risk vector} \\+ A(\tau) &= \text{authority exercised by the agent} \\+ G &= \text{promotion gate}+\end{aligned}+$$ The naive optimizer wants: -```text-c* = argmax_c E[R(tau(c))]-```+$$+c^*=\operatorname*{argmax}_c\mathbb{E}[R(\tau(c))]+$$ The governed optimizer has a constrained objective: -```text-c* = argmax_c E[R(tau(c))] - lambda^T Cost(tau(c))+$$+c^*=\operatorname*{argmax}_c\left(\mathbb{E}[R(\tau(c))]-\lambda^\top\operatorname{Cost}(\tau(c))\right)+$$ subject to:- K_i(tau(c)) <= risk_limit_i- A(tau(c)) <= authority_cap- trace_integrity(tau(c)) passes- eval_integrity(tau(c)) passes- data_boundary(tau(c)) passes- release_gate(c) passes++```text+K_i(tau(c)) <= risk_limit_i+A(tau(c)) <= authority_cap+trace_integrity(tau(c)) passes+eval_integrity(tau(c)) passes+data_boundary(tau(c)) passes+release_gate(c) passes ``` Promotion is not a single score:diff --git a/src/content/posts/self-improving-stack-harness-evolution.mdx b/src/content/posts/self-improving-stack-harness-evolution.mdxindex e695d3a..af1abb6 100644--- a/src/content/posts/self-improving-stack-harness-evolution.mdx+++ b/src/content/posts/self-improving-stack-harness-evolution.mdx@@ -74,16 +74,18 @@ They differ in the mutable surface. That distinction is everything. ## The Reachable Set -Let a system have a mutable surface `s`.+Let a system have a mutable surface $s$. The surface might be: -```text-s_prompt = prompt text-s_skill = procedural memory-s_runtime = driver topology-s_harness = source code around the agent-```+$$+\begin{aligned}+s_{\text{prompt}} &= \text{prompt text} \\+s_{\text{skill}} &= \text{procedural memory} \\+s_{\text{runtime}} &= \text{driver topology} \\+s_{\text{harness}} &= \text{source code around the agent}+\end{aligned}+$$ An optimizer has a mutation operator: @@ -91,39 +93,41 @@ An optimizer has a mutation operator: M_s(candidate, evidence) -> candidate' ``` -The set of systems it can reach after `k` mutations is:+The set of systems it can reach after $k$ mutations is: -```text-Reach(s_0, M_s, k)-```+$$+\operatorname{Reach}(s_0,M_s,k)+$$ -If `M_s` only edits prompt text, the reachable set does not contain a new worktree isolation protocol, a new fanout scheduler, a new raw-provider capture sink, or a new verifier layer. The prompt can ask for those things. It cannot instantiate them if the runtime has no action that does so.+If $M_s$ only edits prompt text, the reachable set does not contain a new worktree isolation protocol, a new fanout scheduler, a new raw-provider capture sink, or a new verifier layer. The prompt can ask for those things. It cannot instantiate them if the runtime has no action that does so. This is the simplest mathematical reason prompt hill climbing plateaus: -```text-best_prompt = argmax_p E[R(run(h_fixed, p, x))]-```+$$+p_{\text{best}}=\operatorname*{argmax}_p\mathbb{E}[R(\operatorname{run}(h_{\text{fixed}},p,x))]+$$ -The harness `h_fixed` is fixed. The optimizer is searching inside the behavior allowed by that harness.+The harness $h_{\text{fixed}}$ is fixed. The optimizer is searching inside the behavior allowed by that harness. Harness evolution changes the outer variable: -```text-h* = argmax_h E_{x ~ D_holdout, z ~ Z}[R(tau(h, x, z))] - lambda^T C(tau(h, x, z))-```+$$+h^*=\operatorname*{argmax}_h\mathbb{E}_{x\sim D_{\text{holdout}},\,z\sim Z}[R(\tau(h,x,z))]-\lambda^\top C(\tau(h,x,z))+$$ where: -```text-h = harness candidate-x = task or scenario-z = seed, replicate, or profile cell-tau = full trace trajectory-R = reward, score, or verifier result-C = cost vector-lambda = cost weights-```+$$+\begin{aligned}+ h &= \text{harness candidate} \\+ x &= \text{task or scenario} \\+ z &= \text{seed, replicate, or profile cell} \\+ \tau &= \text{full trace trajectory} \\+ R &= \text{reward, score, or verifier result} \\+ C &= \text{cost vector} \\+ \lambda &= \text{cost weights}+\end{aligned}+$$ The promotion rule is not just the objective. It is a gate: @@ -288,7 +292,7 @@ promote or reject The important word is structural. -Changing `n = 8` to `n = 16` is not harness evolution. Changing a threshold is not harness evolution. Adding another sentence to a prompt is not harness evolution.+Changing $n=8$ to $n=16$ is not harness evolution. Changing a threshold is not harness evolution. Adding another sentence to a prompt is not harness evolution. Structural variants change mechanism: @@ -328,9 +332,9 @@ candidate_delta = median(candidate_runs - paired_baseline_runs) The strongest comparison is paired: -```text-delta_i = R(h_candidate, x_i, z_i) - R(h_baseline, x_i, z_i)-```+$$+\Delta_i=R(h_{\text{candidate}},x_i,z_i)-R(h_{\text{baseline}},x_i,z_i)+$$ A candidate deserves promotion only when the paired evidence survives uncertainty, cost, deterministic checks, and holdout. @@ -342,7 +346,7 @@ One candidate may improve correctness and increase cost. Another may reduce late So meta-harness tracks a Pareto frontier. -Candidate `a` dominates candidate `b` when:+Candidate $a$ dominates candidate $b$ when: ```text quality_a >= quality_bdiff --git a/src/content/posts/self-improving-stack-memory-flywheels.mdx b/src/content/posts/self-improving-stack-memory-flywheels.mdxindex 2bc0179..db645d1 100644--- a/src/content/posts/self-improving-stack-memory-flywheels.mdx+++ b/src/content/posts/self-improving-stack-memory-flywheels.mdx@@ -73,12 +73,14 @@ The clean way to think about memory is as mutable external state. Let: -```text-M_t = memory state before episode t-tau_t = full trace from episode t-u_t = proposed memory write after episode t-G_mem = memory write gate-```+$$+\begin{aligned}+ M_t &= \text{memory state before episode }t \\+ \tau_t &= \text{full trace from episode }t \\+ u_t &= \text{proposed memory write after episode }t \\+ G_{\text{mem}} &= \text{memory write gate}+\end{aligned}+$$ Write candidates need structure: @@ -96,38 +98,42 @@ u_t = The update rule is: -```text-M_{t+1} =- Apply(M_t, u_t) if G_mem(u_t, tau_t, policy) passes- M_t otherwise-```+$$+M_{t+1}=\begin{cases}+ \operatorname{Apply}(M_t,u_t),&\text{if }G_{\text{mem}}(u_t,\tau_t,\text{policy})\text{ passes}, \\+ M_t,&\text{otherwise.}+\end{cases}+$$ At inference time, the memory layer changes the context seen by the policy: -```text-c_t = Retrieve(M_t, q_t, k, policy)-y_t = pi_theta(x_t, c_t, tools)-```+$$+\begin{aligned}+ c_t &= \operatorname{Retrieve}(M_t,q_t,k,\text{policy}) \\+ y_t &= \pi_\theta(x_t,c_t,\text{tools})+\end{aligned}+$$ where: -```text-x_t = task input-q_t = retrieval query or retrieval plan-k = retrieval budget-c_t = retrieved context-pi_theta = model policy with fixed weights theta-y_t = output or next action-```+$$+\begin{aligned}+ x_t &= \text{task input} \\+ q_t &= \text{retrieval query or retrieval plan} \\+ k &= \text{retrieval budget} \\+ c_t &= \text{retrieved context} \\+ \pi_\theta &= \text{model policy with fixed weights }\theta \\+ y_t &= \text{output or next action}+\end{aligned}+$$ Memory is not magic. It is an intervention on the policy's input distribution. That gives us a measurable target: -```text-Delta_memory =- E[Score(pi_theta with M)] - E[Score(pi_theta without M)]-```+$$+\Delta_{\text{memory}}=\mathbb{E}[\operatorname{Score}(\pi_\theta\text{ with }M)]-\mathbb{E}[\operatorname{Score}(\pi_\theta\text{ without }M)]+$$ The unit test for memory is not "did retrieval return something?" @@ -212,9 +218,9 @@ tau_t = The extractor turns trace evidence into proposed writes: -```text-tau_t -> {u_1, u_2, ..., u_n}-```+$$+\tau_t\longrightarrow\{u_1,u_2,\ldots,u_n\}+$$ The gate decides whether each write is safe, scoped, supported, and useful: @@ -318,7 +324,7 @@ becomes: pi_theta(y | x, c) ``` -where `c` is retrieved context.+where $c$ is retrieved context. That context can help, do nothing, or harm. It can help by supplying missing facts, preserving user preferences, recalling a successful procedure, or warning against a repeated mistake. It can harm by anchoring the model on irrelevant facts, stale procedures, false summaries, or over-broad preferences. diff --git a/src/content/posts/self-improving-stack-multi-agent-coordination.mdx b/src/content/posts/self-improving-stack-multi-agent-coordination.mdxindex dd76fa1..d086bca 100644--- a/src/content/posts/self-improving-stack-multi-agent-coordination.mdx+++ b/src/content/posts/self-improving-stack-multi-agent-coordination.mdx@@ -59,44 +59,46 @@ Coordination is structure. The previous post made runtime topology explicit: -```text-g = executable graph of agents, tools, validators, selectors, handoffs, and gates-pi = runtime policy over graph moves-```+$$+\begin{aligned}+ g &= \text{executable graph of agents, tools, validators, selectors, handoffs, and gates} \\+ \pi &= \text{runtime policy over graph moves}+\end{aligned}+$$ Multi-agent coordination sits one layer above that. It decides what the nodes are supposed to do together. Let: -```text-r = role contracts-p = persona and instruction content-k = active skills per role-u = tools and permissions per role-c = communication and context-sharing policy-sigma = selector or merger policy-v = verifier and judge stack-b = budget allocation policy-tau = termination and escalation policy-```+$$+\begin{aligned}+ r &= \text{role contracts} \\+ p &= \text{persona and instruction content} \\+ k &= \text{active skills per role} \\+ u &= \text{tools and permissions per role} \\+ c &= \text{communication and context-sharing policy} \\+ \sigma &= \text{selector or merger policy} \\+ v &= \text{verifier and judge stack} \\+ b &= \text{budget allocation policy} \\+ \tau &= \text{termination and escalation policy}+\end{aligned}+$$ A multi-agent system candidate is: -```text-s = (g, r, p, k, u, c, sigma, v, b, tau)-```+$$+s=(g,r,p,k,u,c,\sigma,v,b,\tau)+$$ The optimization target is: -```text-J(s | m, h) =- E_{x ~ D}[R(run(m, h, s, x))]- - lambda * E[C(run(m, h, s, x))]-```+$$+J(s\mid m,h)=\mathbb{E}_{x\sim D}[R(\operatorname{run}(m,h,s,x))]-\lambda\,\mathbb{E}[C(\operatorname{run}(m,h,s,x))]+$$ The notation is only useful because it exposes the coordinate system. -If only `p` changes, the optimizer is tuning role descriptions. If `k` changes, it is tuning durable role procedure. If `g`, `sigma`, `c`, `b`, or `tau` change, it is tuning coordination structure.+If only $p$ changes, the optimizer is tuning role descriptions. If $k$ changes, it is tuning durable role procedure. If $g$, $\sigma$, $c$, $b$, or $\tau$ change, it is tuning coordination structure. That distinction prevents a common category error: evaluating a better persona as if it were a better multi-agent system. @@ -179,18 +181,17 @@ If every worker shares the same prompt, same model, same examples, same context, The idealized independent case is easy: -```text-P(at least one success) = 1 - product_i(1 - q_i)-```+$$+\mathbb{P}(\text{at least one success})=1-\prod_i(1-q_i)+$$ -where `q_i` is the probability that worker `i` independently finds a valid answer.+where $q_i$ is the probability that worker $i$ independently finds a valid answer. But multi-agent LLM systems rarely get independence for free. The useful quantity is not worker count. It is error correlation. -```text-coordination_gain =- E[score_multi at budget B] - E[score_best_single at budget B]-```+$$+\operatorname{coordination\_gain}=\mathbb{E}[\operatorname{score\_multi}\text{ at budget }B]-\mathbb{E}[\operatorname{score\_best\_single}\text{ at budget }B]+$$ If the gain disappears at matched budget, the system did not learn coordination. It bought more samples. diff --git a/src/content/posts/self-improving-stack-optimization-theory.mdx b/src/content/posts/self-improving-stack-optimization-theory.mdxindex 3cef393..ff3aea7 100644--- a/src/content/posts/self-improving-stack-optimization-theory.mdx+++ b/src/content/posts/self-improving-stack-optimization-theory.mdx@@ -84,9 +84,9 @@ For a prompt optimizer, `s` might be a system prompt, a task instruction, a fiel The objective looks like this: -```text-J(s) = E_{tau ~ D}[R(run(s, tau))] - lambda * C(s)-```+$$+J(s)=\mathbb{E}_{\tau\sim D}[R(\operatorname{run}(s,\tau))]-\lambda C(s)+$$ In English: @@ -95,15 +95,15 @@ In English: - `run(s, tau)` is the trace produced when the system with surface `s` attempts the task. - `R` is the reward or score for that trace. - `C` is cost: tokens, latency, dollars, tool calls, human review time, risk.-- `lambda` controls how much you care about cost.+- $\lambda$ controls how much you care about cost. -Most agent teams do not write the equation down. They still live inside it. If a support agent is optimized for accuracy while token cost is ignored, then `lambda` is effectively zero. If a product always picks the cheapest model regardless of success rate, the cost term is dominating the objective. If the judge rewards long answers, the reward function has silently made verbosity a feature.+Most agent teams do not write the equation down. They still live inside it. If a support agent is optimized for accuracy while token cost is ignored, then $\lambda$ is effectively zero. If a product always picks the cheapest model regardless of success rate, the cost term is dominating the objective. If the judge rewards long answers, the reward function has silently made verbosity a feature. An optimizer is an operator: -```text-s_next = O(s_current, traces, feedback, budget)-```+$$+s_{\text{next}}=O(s_{\text{current}},\text{traces},\text{feedback},\text{budget})+$$ `O` can be a human writing a new prompt. It can be Bayesian optimization. It can be an LLM reflecting on failures. It can be evolutionary mutation. It can be reinforcement learning. The operator matters, but a weak gate dominates a clever operator. A sophisticated mutator with a bad score function is just a faster way to overfit. @@ -129,9 +129,9 @@ Simulated annealing adds a different trick: sometimes accept a worse candidate e Bayesian optimization appears when evaluations are expensive. You cannot afford to try every prompt, every model, every tool order, or every hyperparameter combination. So you build a surrogate model of the objective from previous trials and use an acquisition function to choose the next trial. A common acquisition function is expected improvement: -```text-EI(x) = E[max(f(x) - f_best, 0)]-```+$$+\operatorname{EI}(x)=\mathbb{E}[\max(f(x)-f_{\text{best}},0)]+$$ Expected improvement balances exploration and exploitation. It favors candidates that look promising, but it also spends trials in uncertain regions because those trials may reveal a better basin. @@ -151,9 +151,9 @@ The model may be remote. The provider may not expose gradients. The action space The practical problem is closer to: -```text-argmax_s J(s)-```+$$+\operatorname*{argmax}_s J(s)+$$ where `J` is expensive, noisy, partially subjective, and easy to game. @@ -232,12 +232,14 @@ If a prompt changes, a held-out prompt eval may be enough. If a skill changes, y The simplest honest comparison is paired evaluation. Run the baseline and candidate on the same tasks, then compare per-task deltas: -```text-delta_i = score_i(candidate) - score_i(baseline)-mean_delta = (1/n) * sum_i delta_i-```+$$+\begin{aligned}+ \Delta_i &= \operatorname{score}_i(\text{candidate})-\operatorname{score}_i(\text{baseline}) \\+ \operatorname{mean\_delta} &= \frac{1}{n}\sum_i\Delta_i+\end{aligned}+$$ -If `mean_delta` is positive, the candidate looks better. But "looks better" is not enough. With stochastic systems, a candidate can win because it got easier samples, the judge was inconsistent, the model sampled lucky trajectories, or one outlier task dominated the average.+If $\operatorname{mean\_delta}$ is positive, the candidate looks better. But "looks better" is not enough. With stochastic systems, a candidate can win because it got easier samples, the judge was inconsistent, the model sampled lucky trajectories, or one outlier task dominated the average. So the promotion rule needs uncertainty: @@ -249,13 +251,13 @@ promote if lower_confidence_bound(mean_delta) > epsilon For agents, the score is rarely one-dimensional. Quality, cost, latency, robustness, safety, and trace integrity all matter. That leads to Pareto thinking: -```text-candidate A dominates B if:- quality_A >= quality_B- cost_A <= cost_B- latency_A <= latency_B-and at least one inequality is strict-```+$$+A\succ B\quad\text{if}\quad\begin{cases}+ \operatorname{quality}_A\ge\operatorname{quality}_B, \\+ \operatorname{cost}_A\le\operatorname{cost}_B, \\+ \operatorname{latency}_A\le\operatorname{latency}_B,+\end{cases}\quad\text{with at least one strict inequality}+$$ Many real candidates do not dominate each other. One is better and slower. Another is cheaper and less robust. Another is safer but more verbose. This is why GEPA's Pareto framing is natural for compound systems, and why agent-eval style scorecards are more useful than a single magic number. diff --git a/src/content/posts/self-improving-stack-post-training.mdx b/src/content/posts/self-improving-stack-post-training.mdxindex 7d662e2..0f9caf1 100644--- a/src/content/posts/self-improving-stack-post-training.mdx+++ b/src/content/posts/self-improving-stack-post-training.mdx@@ -60,9 +60,9 @@ evaluate candidate promote or reject ``` -But the mutable surface is no longer a prompt file or a worktree. It is `theta`, the model parameters, or some parameterized adapter attached to the model.+But the mutable surface is no longer a prompt file or a worktree. It is $\theta$, the model parameters, or some parameterized adapter attached to the model. -Once `theta` moves, the boundary changes. The behavior becomes harder to inspect, harder to patch locally, harder to roll back partially, and harder to explain from a single trace. It can also generalize better than any prompt edit when the signal is strong enough.+Once $\theta$ moves, the boundary changes. The behavior becomes harder to inspect, harder to patch locally, harder to roll back partially, and harder to explain from a single trace. It can also generalize better than any prompt edit when the signal is strong enough. That is why this layer deserves separate treatment. @@ -74,7 +74,7 @@ The previous posts treated the model as mostly fixed: y = model_theta(prompt, tools, memory, trace_context) ``` -External optimization changed everything around `theta`:+External optimization changed everything around $\theta$: ```text prompt@@ -87,12 +87,14 @@ selector harness ``` -Post-training changes `theta`, or a learned delta attached to it:+Post-training changes $\theta$, or a learned delta attached to it: -```text-theta_{t+1} = Update(theta_t, data, objective)-theta' = theta_base + Delta_adapter-```+$$+\begin{aligned}+ \theta_{t+1}&=\operatorname{Update}(\theta_t,\text{data},\text{objective}) \\+ \theta'&=\theta_{\text{base}}+\Delta_{\text{adapter}}+\end{aligned}+$$ That small notation hides the whole difference. @@ -129,15 +131,15 @@ Supervised fine-tuning is the simplest objective. Given demonstrations: -```text-D = {(x_i, y_i)}-```+$$+D=\{(x_i,y_i)\}+$$ train the model to put probability mass on the demonstrated output: -```text-L_SFT(theta) = - sum_i log pi_theta(y_i | x_i)-```+$$+\mathcal{L}_{\text{SFT}}(\theta)=-\sum_i\log\pi_\theta(y_i\mid x_i)+$$ This is powerful when demonstrations are clean and the task distribution is stable. @@ -174,22 +176,19 @@ Preference data looks like: (x, y_w, y_l) ``` -where `y_w` is preferred over `y_l`.+where $y_w$ is preferred over $y_l$. A common reward-model loss is: -```text-L_RM(phi) =- - log sigma(r_phi(x, y_w) - r_phi(x, y_l))-```+$$+\mathcal{L}_{\text{RM}}(\phi)=-\log\sigma(r_\phi(x,y_w)-r_\phi(x,y_l))+$$ Then the policy objective becomes: -```text-max_theta E_{y ~ pi_theta(. | x)}[- r_phi(x, y) - beta KL(pi_theta(. | x) || pi_ref(. | x))-]-```+$$+\max_\theta\mathbb{E}_{y\sim\pi_\theta(\cdot\mid x)}\left[r_\phi(x,y)-\beta\,D_{\mathrm{KL}}(\pi_\theta(\cdot\mid x)\Vert\pi_{\mathrm{ref}}(\cdot\mid x))\right]+$$ The reward says "move toward preferred behavior." The KL term says "do not drift too far from the reference policy." @@ -203,18 +202,20 @@ The gate has to test the trained policy on held-out tasks, adversarial probes, c PPO is the familiar online optimizer in many RLHF pipelines. It samples from the current policy, scores the samples, estimates an advantage, and updates the policy while clipping large policy-ratio moves: -```text-rho = pi_theta(y | x) / pi_old(y | x)-L_PPO(theta) = E[min(rho A, clip(rho, 1 - eps, 1 + eps) A)]-```+$$+\begin{aligned}+ \rho&=\frac{\pi_\theta(y\mid x)}{\pi_{\text{old}}(y\mid x)} \\+ \mathcal{L}_{\text{PPO}}(\theta)&=\mathbb{E}[\min(\rho A,\operatorname{clip}(\rho,1-\epsilon,1+\epsilon)A)]+\end{aligned}+$$ The clipped objective is an engineering answer to a stability problem: if the policy moves too far in one update, the reward model can be exploited and the policy can drift. GRPO-style training changes the advantage estimate. Instead of training a separate critic to estimate value, sample a group of outputs for the same prompt and normalize rewards inside the group: -```text-A_i = (r_i - mean(r_1, ..., r_G)) / (std(r_1, ..., r_G) + eps)-```+$$+A_i=\frac{r_i-\operatorname{mean}(r_1,\ldots,r_G)}{\operatorname{std}(r_1,\ldots,r_G)+\epsilon}+$$ That group-relative signal is why GRPO has become prominent in reasoning and tool-use RL discussions. It is not a free lunch. The reward still has to be meaningful, the group has to contain useful variation, and the heldout gate still has to catch reward hacking. @@ -247,17 +248,11 @@ AI feedback is not automatically objective feedback. Direct Preference Optimization removes the explicit reward-model and online-RL stages. -For a preferred output `y_w` and rejected output `y_l`, DPO optimizes a classification-style objective over log-probability differences:+For a preferred output $y_w$ and rejected output $y_l$, DPO optimizes a classification-style objective over log-probability differences: -```text-L_DPO(theta) =- - log sigma(- beta [- log pi_theta(y_w | x) - log pi_ref(y_w | x)- - log pi_theta(y_l | x) + log pi_ref(y_l | x)- ]- )-```+$$+\mathcal{L}_{\text{DPO}}(\theta)=-\log\sigma\!\left(\beta\left[\log\frac{\pi_\theta(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)}-\log\frac{\pi_\theta(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}\right]\right)+$$ The reference policy still matters. The preference pair still matters. The quality of the pair dataset becomes the training signal. @@ -290,9 +285,9 @@ Process supervision scores intermediate steps. If an agent run is: -```text-tau = (x, a_1, o_1, a_2, o_2, ..., y)-```+$$+\tau=(x,a_1,o_1,a_2,o_2,\ldots,y)+$$ then outcome reward is: @@ -302,9 +297,9 @@ R(tau) Process reward is: -```text-R_process(tau) = sum_t gamma^t r_t(a_t, o_t, state_t)-```+$$+R_{\text{process}}(\tau)=\sum_t\gamma^t r_t(a_t,o_t,\text{state}_t)+$$ The 2023 "Let's Verify Step by Step" result made this concrete for math reasoning: step-level feedback can outperform final-answer-only feedback when training reward models for hard reasoning. @@ -328,13 +323,15 @@ The strongest RL signals are not vibes. They are checks. For code, math, theorem proving, schemas, and tool tasks, some rewards are decidable: -```text-r = 1[tests_pass]-r = 1[schema_valid]-r = 1[proof_checks]-r = 1[compile_succeeds]-r = fraction_of_assertions_passed-```+$$+\begin{aligned}+r &= \mathbb{1}[\text{tests pass}] \\+r &= \mathbb{1}[\text{schema valid}] \\+r &= \mathbb{1}[\text{proof checks pass}] \\+r &= \mathbb{1}[\text{compilation succeeds}] \\+r &= \text{fraction of assertions passed}+\end{aligned}+$$ This is why tool-use and coding RL are attractive. The environment can score behavior without a subjective judge for at least part of the task. @@ -532,7 +529,7 @@ gates promote candidates RL bridge exports training signal ``` -The final step, updating `theta`, remains outside the public product loop unless a real training backend is wired.+The final step, updating $\theta$, remains outside the public product loop unless a real training backend is wired. ## The Promotion Gate Gets Stricter diff --git a/src/content/posts/self-improving-stack-prompt-optimization.mdx b/src/content/posts/self-improving-stack-prompt-optimization.mdxindex 6dfbb20..da4fd69 100644--- a/src/content/posts/self-improving-stack-prompt-optimization.mdx+++ b/src/content/posts/self-improving-stack-prompt-optimization.mdx@@ -66,24 +66,26 @@ Answer that and the ecosystem stops looking like magic. It becomes experimental Let: -```text-p = prompt artifact or prompt-like text surface-d = selected demonstrations or exemplars-m = model or backend-h = runtime and harness-x = task sampled from the eval distribution D-y = system output or full trajectory-R = reward, metric, judge, or scoring function-C = cost, latency, risk, or resource usage-```+$$+\begin{aligned}+ p &= \text{prompt artifact or prompt-like text surface} \\+ d &= \text{selected demonstrations or exemplars} \\+ m &= \text{model or backend} \\+ h &= \text{runtime and harness} \\+ x &= \text{task sampled from eval distribution }D \\+ y &= \text{system output or full trajectory} \\+ R &= \text{reward, metric, judge, or scoring function} \\+ C &= \text{cost, latency, risk, or resource usage}+\end{aligned}+$$ A prompt optimizer is usually solving: -```text-J(p, d | m, h) = E_{x ~ D}[R(run(m, h, p, d, x))] - lambda * E[C(run(m, h, p, d, x))]-```+$$+J(p,d\mid m,h)=\mathbb{E}_{x\sim D}[R(\operatorname{run}(m,h,p,d,x))]-\lambda\,\mathbb{E}[C(\operatorname{run}(m,h,p,d,x))]+$$ -The conditional bar is the whole argument. It says the optimizer is improving `p` and maybe `d` while model `m` and runtime `h` are treated as fixed. If the model, toolset, router, memory, worker count, turn budget, or evaluator also changes, the experiment no longer estimates a pure prompt effect. It estimates a confounded system effect.+The conditional bar is the whole argument. It says the optimizer is improving $p$ and maybe $d$ while model $m$ and runtime $h$ are treated as fixed. If the model, toolset, router, memory, worker count, turn budget, or evaluator also changes, the experiment no longer estimates a pure prompt effect. It estimates a confounded system effect. This is experimental design, not pedantry. A candidate prompt can look better because the model changed. A model can look better because the prompt changed. A workflow can look better because the judge prompt became easier. A multi-agent coordinator can look better because a runtime silently raised its turn budget. @@ -238,7 +240,7 @@ Verify before final output. That may improve behavior if the runtime already exposes the necessary actions. It will not create those actions. A text instruction to parallelize is operational only if the agent has a tool or runtime primitive that dispatches work concurrently. A text instruction to continue until verified only works if the loop budget and stop semantics allow it. A text instruction to use memory only works if memory is available, scoped, and retrievable. -`maxTurns` is a runtime semantic, not a style preference. If zero means "no autonomous turns," then the prompt cannot create a multi-step process. If zero means "unbounded," the optimization problem becomes budget and safety control. Either way, the parameter belongs to `h`, the runtime/harness side of `J(p, d | m, h)`.+`maxTurns` is a runtime semantic, not a style preference. If zero means "no autonomous turns," then the prompt cannot create a multi-step process. If zero means "unbounded," the optimization problem becomes budget and safety control. Either way, the parameter belongs to $h$, the runtime/harness side of $J(p,d\mid m,h)$. The correct formal move is to expand the candidate: @@ -252,10 +254,12 @@ s = { budget_policy, evaluator_config }--J(s) = E_{x ~ D}[R(run(s, x))] - lambda_cost*C - lambda_risk*K ``` +$$+J(s)=\mathbb{E}_{x\sim D}[R(\operatorname{run}(s,x))]-\lambda_{\text{cost}}C-\lambda_{\text{risk}}K+$$+ Now the loop is no longer pure prompt optimization. It is system optimization. GEPA-style reflection may still be part of the proposal operator, but the mutable surface is larger than prompts. This is where meta-harnesses enter. A meta-harness does not merely tune wording inside one harness. It searches over code, architecture, eval plumbing, candidate generators, and workflow topology. Its failure mode is also larger: it can improve the harness while weakening the real product task. That is why trace integrity and held-out promotion become stricter as the mutable surface expands.@@ -305,9 +309,9 @@ Minimum protocol: For stochastic agents, paired deltas are the basic unit: -```text-delta_i = score_i(candidate) - score_i(baseline)-```+$$+\Delta_i=\operatorname{score}_i(\text{candidate})-\operatorname{score}_i(\text{baseline})+$$ Promotion should look more like: diff --git a/src/content/posts/self-improving-stack-skill-optimization.mdx b/src/content/posts/self-improving-stack-skill-optimization.mdxindex 78d07a4..e73894b 100644--- a/src/content/posts/self-improving-stack-skill-optimization.mdx+++ b/src/content/posts/self-improving-stack-skill-optimization.mdx@@ -90,22 +90,24 @@ That middle layer is where a lot of real agent learning will happen, because it Let: -```text-k = skill artifact-a = activation policy or retrieval policy for the skill-m = model or backend-h = harness and runtime-p = ordinary prompt context-x = task from distribution D-R = task reward or eval score-C = cost, latency, token load, or risk-```+$$+\begin{aligned}+ k &= \text{skill artifact} \\+ a &= \text{activation policy or retrieval policy for the skill} \\+ m &= \text{model or backend} \\+ h &= \text{harness and runtime} \\+ p &= \text{ordinary prompt context} \\+ x &= \text{task from distribution }D \\+ R &= \text{task reward or eval score} \\+ C &= \text{cost, latency, token load, or risk}+\end{aligned}+$$ Skill optimization estimates: -```text-J(k, a | m, h, p) = E_{x ~ D}[R(run(m, h, p, k, a, x))] - lambda * E[C(run(m, h, p, k, a, x))]-```+$$+J(k,a\mid m,h,p)=\mathbb{E}_{x\sim D}[R(\operatorname{run}(m,h,p,k,a,x))]-\lambda\,\mathbb{E}[C(\operatorname{run}(m,h,p,k,a,x))]+$$ The conditional variables matter. If a skill is evaluated while the model, tools, runtime, evaluator, or prompt also change, the measured lift is not a clean skill effect. The correct cell is an agent profile: @@ -307,16 +309,18 @@ Minimum protocol: The core comparison is paired: -```text-delta_i = score_i(profile_with_skill) - score_i(profile_without_skill)-```+$$+\Delta_i=\operatorname{score}_i(\text{profile with skill})-\operatorname{score}_i(\text{profile without skill})+$$ For activation: -```text-precision = relevant_loaded / loaded-recall = relevant_loaded / relevant_tasks-```+$$+\begin{aligned}+ \operatorname{precision}&=\frac{\text{relevant loaded}}{\text{loaded}} \\+ \operatorname{recall}&=\frac{\text{relevant loaded}}{\text{relevant tasks}}+\end{aligned}+$$ For promotion: @@ -352,7 +356,7 @@ Another failure is activation gaming: skill description becomes broad -> skill loads often -> eval improves on benchmark -> unrelated tasks degrade ``` -This is why the activation policy `a` belongs in the objective. Optimizing only the body `k` is incomplete.+This is why the activation policy $a$ belongs in the objective. Optimizing only the body $k$ is incomplete. ## When Skills Are The Right Surface diff --git a/src/content/posts/self-improving-stack-trace-systems.mdx b/src/content/posts/self-improving-stack-trace-systems.mdxindex 3274500..27e26f0 100644--- a/src/content/posts/self-improving-stack-trace-systems.mdx+++ b/src/content/posts/self-improving-stack-trace-systems.mdx@@ -57,25 +57,27 @@ The trace is not decoration around the eval. The trace is the data. An agent run is a trajectory: -```text-tau = (x, s_0, a_1, o_1, s_1, ..., a_T, o_T, y)-```+$$+\tau=(x,s_0,a_1,o_1,s_1,\ldots,a_T,o_T,y)+$$ where: -```text-x = task-s_t = internal and external state-a_t = action-o_t = observation-y = outcome-```+$$+\begin{aligned}+ x &= \text{task} \\+ s_t &= \text{internal and external state} \\+ a_t &= \text{action} \\+ o_t &= \text{observation} \\+ y &= \text{outcome}+\end{aligned}+$$ A final score from a fixed scorer is a projection: -```text-score = R(tau)-```+$$+\text{score}=R(\tau)+$$ That projection is intentionally lossy. It collapses a long sequence of decisions, calls, costs, artifacts, and observations into one number. @@ -83,11 +85,11 @@ Optimization needs the lost variables. The crude information-theory version: -```text-I(tau; failure_cause) >= I(score; failure_cause)-```+$$+I(\tau;\text{failure cause})\ge I(\text{score};\text{failure cause})+$$ -When `score = R(tau)` and the scorer is fixed, the score is a deterministic projection of the trajectory. By the data processing inequality, that projection cannot contain more information about the failure cause than the trajectory itself. Usually it contains dramatically less.+When $\text{score}=R(\tau)$ and the scorer is fixed, the score is a deterministic projection of the trajectory. By the data processing inequality, that projection cannot contain more information about the failure cause than the trajectory itself. Usually it contains dramatically less. This does not mean every byte is equally useful. It means the system must preserve the variables that can explain responsible mechanism: diff --git a/src/content/posts/the-self-improving-stack.mdx b/src/content/posts/the-self-improving-stack.mdxindex 03e9f0c..45e176a 100644--- a/src/content/posts/the-self-improving-stack.mdx+++ b/src/content/posts/the-self-improving-stack.mdx@@ -104,25 +104,29 @@ The system is not self-improving because it says "reflect." It is self-improving The compact equation is: -```text-Delta_t = E_{x ~ D_task}[Eval(Run(c_t, x), Run(s_t, x))]--s_{t+1} =- Promote(s_t, c_t)- if Gate(Delta_t, policy) passes- else s_t-```+$$+\begin{aligned}+ \Delta_t &= \mathbb{E}_{x\sim D_{\text{task}}}[\operatorname{Eval}(\operatorname{Run}(c_t,x),\operatorname{Run}(s_t,x))] \\+ s_{t+1} &=+ \begin{cases}+ \operatorname{Promote}(s_t,c_t), & \text{if }\operatorname{Gate}(\Delta_t,\text{policy})\text{ passes}, \\+ s_t, & \text{otherwise.}+ \end{cases}+\end{aligned}+$$ where: -```text-s_t = current system state-c_t = candidate state-x ~ D_task = a user task drawn from the task distribution the system is meant to serve-Run(.) = full agent trajectory under sampled user-task scenarios-Eval(.) = measured evidence on task outcomes and trace behavior-Gate(.) = promotion rule under policy-```+$$+\begin{aligned}+ s_t &= \text{current system state} \\+ c_t &= \text{candidate state} \\+ x &\sim D_{\text{task}}\quad\text{(a user task drawn from the target task distribution)} \\+ \operatorname{Run}(\cdot) &= \text{full agent trajectory under sampled user-task scenarios} \\+ \operatorname{Eval}(\cdot) &= \text{measured evidence on task outcomes and trace behavior} \\+ \operatorname{Gate}(\cdot) &= \text{promotion rule under policy}+\end{aligned}+$$ The system improves only when the promoted state performs better on the user-task distribution while staying inside cost, safety, integrity, and governance constraints. @@ -189,11 +193,11 @@ But prompt search operates inside a fixed runtime. The objective is roughly: -```text-p* = argmax_p E[R(Run(h_fixed, p, x))]-```+$$+p^*=\operatorname*{argmax}_p\mathbb{E}[R(\operatorname{Run}(h_{\text{fixed}},p,x))]+$$ -The harness `h_fixed` is held constant.+The harness $h_{\text{fixed}}$ is held constant. If the runtime lacks a worker pool, the prompt can ask for parallelism but cannot create it. If the tool graph lacks a verifier, the prompt can request verification but cannot execute one. If the evaluator leaks holdout answers, the prompt can overfit beautifully. @@ -404,9 +408,9 @@ evaluator Post-training mutates model behavior itself: -```text-theta_{t+1} = Update(theta_t, data, objective)-```+$$+\theta_{t+1}=\operatorname{Update}(\theta_t,\text{data},\text{objective})+$$ or: @@ -424,11 +428,12 @@ Memory changes future runs by changing what persists. The update rule is: -```text-M_{t+1} =- Apply(M_t, u_t) if G_mem(u_t, trace, policy) passes- M_t otherwise-```+$$+M_{t+1}=\begin{cases}+ \operatorname{Apply}(M_t,u_t),&\text{if }G_{\text{mem}}(u_t,\text{trace},\text{policy})\text{ passes}, \\+ M_t,&\text{otherwise.}+\end{cases}+$$ A memory system is useful when: