Prompt Optimization Is Not The Whole Game
60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
- Created
- Updated
31
Turns
22
Tool calls
16
Files touched
4135m
Duration
Files
/Users/drew/code/blogsrc/content.config.tstools/new-post.mjstools/README.mdtools/trace-capture.tstraces/the-long-horizon/2026-05-25T11-23-22-617Z-gpt-5.5-series-outline.jssrc/content/posts/superintelligence-in-the-wild.mdxtools/harness/manual.tssrc/content/posts/self-improving-stack-evaluation-gates.mdxsrc/content/posts/self-improving-stack-trace-systems.mdxsrc/content/posts/self-improving-stack-governance.mdxsrc/content/posts/self-improving-stack-test-time-compute.mdxsrc/content/posts/self-improving-stack-post-training.mdxsrc/content/posts/self-improving-stack-harness-evolution.mdxsrc/content/posts/the-self-improving-stack.mdxresearch/self-improving-agent-systems/voice-blend-eval.md
Commit
d5bba9f feat(traces): add phase markers and hook directives for AI publish workflow +53 −6 across 14 files
Conversation
31 turns. Full text where captured; older traces show only the first ~280 chars.
- user #1let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our knowledge and you should create a blog post for each and we make this a series
- assistant #2Problem: the research is currently a single trace plus a chat summary, so it is hard to grow into a series without losing the map. Change: I’ll add a durable research checkpoint directory with one file per map item, then scaffold one draft post per item under a consistent series. Why long-term right: research notes stay source-of-truth and appendable, while posts stay draft publication surfaces with provenance. Cost: new markdown/MDX files only; risk is schema/provenance mismatch, rollback is deleting the new directory/posts.
- assistant #3I’m checking the post schema and scaffold tooling before writing files, so the new drafts land in the repo’s existing provenance format instead of inventing one.
- assistant #4 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"git status --short --branch","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
-
- assistant #5 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"git log --oneline -10","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
-
- assistant #6 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"git reflog | head -20","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
-
- assistant #7 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"gh pr list --state open","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
-
- assistant #8 1 tool
- exec_command
src/content.config.tsinput
{"cmd":"sed -n '1,260p' src/content.config.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":6000}
-
- assistant #9 1 tool
- exec_command
tools/new-post.mjsinput
{"cmd":"sed -n '1,260p' tools/new-post.mjs","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":7000}
-
- assistant #10The schema supports `series`, `outline_trace_id`, `supporting_trace_ids`, and `human_takeover`, which is enough for draft series scaffolds. I’m going to create a `research/self-improving-agent-systems/` checkpoint directory plus 12 draft MDX outlines that all point back to the existing supporting trace.
- assistant #11 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"find . -maxdepth 2 -type d | sort | sed -n '1,160p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":4000}
-
- assistant #12 1 tool
- exec_command
tools/README.mdinput
{"cmd":"rg -n \"series:|self-improving|long-horizon|supporting_trace_ids|outline_trace_id\" src/content/posts traces README.md tools/README.md --glob '!dist'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #13 1 tool
- exec_command
tools/trace-capture.tsinput
{"cmd":"sed -n '1,320p' tools/trace-capture.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #14 1 tool
- exec_command
traces/the-long-horizon/2026-05-25T11-23-22-617Z-gpt-5.5-series-outline.jsinput
{"cmd":"sed -n '1,140p' traces/the-long-horizon/2026-05-25T11-23-22-617Z-gpt-5.5-series-outline.json","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":8000}
-
- assistant #15 1 tool
- exec_command
src/content/posts/superintelligence-in-the-wild.mdxinput
{"cmd":"sed -n '1,80p' src/content/posts/superintelligence-in-the-wild.mdx","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":4000}
-
- assistant #16 1 tool
- exec_command
tools/harness/manual.tsinput
{"cmd":"sed -n '1,220p' tools/harness/manual.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":6000}
-
- assistant #17I found the main residual issue: the facts are mostly solid, but several openings are still stacked as one-line aphorisms. I’m collapsing those into denser arguments so the voice sounds like a person with a hard-won model of the system, not a slide deck warming up.
- assistant #18 1 tool
- apply_patch
/Users/drew/code/blog/src/content/posts/self-improving-stack-evaluation-gates.mdxinput
*** Begin Patch *** Update File: /Users/drew/code/blog/src/content/posts/self-improving-stack-evaluation-gates.mdx @@ -A green score is not a release decision. - -That is the mistake every self-improving loop wants to make. GEPA, MIPRO, SkillOpt, topology search, and meta-harness can all generate candidates forever. The gate decides which candidate becomes the system future agents inherit. - -If the gate is weak, every optimizer learns the gate. - -If the gate is honest, every optimizer has to improve the product. - -This is why the gate is not an administrative detail after the interesting w…
-
- assistant #19That pass removed the worst fragment stacks and one cute line that was weakening the test-time compute post. Next I’m checking the series as a whole for banned prose artifacts, em dashes, leftover scaffold headings, and paragraph rhythm.
- assistant #20 1 tool
- exec_command
src/content/posts/the-self-improving-stack.mdxinput
{"cmd":"rg -n \"—|first post|First post|reader hook|the article should|keep this compact|Drew angle|target audience|outline notes|The serious version of the question is|That sentence is too vague|^## (Working Rule|A Working Rule|Evaluation Protocol|The Tangle Placement|Tangle Placement|The Object Being Optimized|The Formal Shape)$\" src/content/posts/the-self-improving-stack.mdx src/content/posts/self-improving-stack-*.mdx VOICE.md SECONDARY_VOICE.md research/self-improving-agent-systems/voice-blend-eval.md","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #21 1 tool
- exec_command
src/content/posts/the-self-improving-stack.mdxinput
{"cmd":"python3 - <<'PY'\nfrom pathlib import Path\nfiles = [Path('src/content/posts/the-self-improving-stack.mdx')] + sorted(Path('src/content/posts').glob('self-improving-stack-*.mdx'))\nfor p in files:\n body = p.read_text().split('---', 2)[-1]\n paras = [x.strip() for x in body.split('\\n\\n') if x.strip() and not x.strip().startswith(('```','|','- ','##'))]\n short = [x for x in paras if len(x.split()) <= 10]\n print(f'{p.name}: paragraphs={len(paras)} short<=10={len(short)}')\nPY","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #22 1 tool
- exec_command
research/self-improving-agent-systems/voice-blend-eval.mdinput
{"cmd":"sed -n '1,240p' research/self-improving-agent-systems/voice-blend-eval.md","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #23 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"sed -n '1,220p' VOICE.md","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #24 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"sed -n '1,220p' SECONDARY_VOICE.md","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #25The banned-pattern scan is clean for the post bodies: no em dashes, no “first post,” no scaffold headings left in the series. The only hits are in `VOICE.md` where those phrases are listed as anti-patterns, which is intentional.
- assistant #26 1 tool
- exec_command
src/content/posts/the-self-improving-stack.mdxinput
{"cmd":"rg -n \"original:\\s*true\" src/content/posts/the-self-improving-stack.mdx src/content/posts/self-improving-stack-*.mdx","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #27 1 tool
- exec_command
src/content/posts/the-self-improving-stack.mdxinput
{"cmd":"git diff -- src/content/posts/the-self-improving-stack.mdx src/content/posts/self-improving-stack-*.mdx VOICE.md SECONDARY_VOICE.md research/self-improving-agent-systems/voice-blend-eval.md | sed -n '1,260p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":30000}
-
- assistant #28 1 tool
- exec_command
src/content/posts/the-self-improving-stack.mdxinput
{"cmd":"git diff --stat -- src/content/posts/the-self-improving-stack.mdx src/content/posts/self-improving-stack-*.mdx VOICE.md SECONDARY_VOICE.md research/self-improving-agent-systems/voice-blend-eval.md AGENTS.md CLAUDE.md","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #29No target post is marked `original: true`, so the rewrite is allowed under the repo rules. The diff is also narrower than it looks: mostly openings and headings, not wholesale claim churn, which is the right place to alter voice without damaging the source trail.
- assistant #30I’m recording the rewrite traces now with the blog CLI so the authorship/editing trace stays in `revisions[]`, not in supporting research. After that I’ll build the site.
- assistant #31 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"for slug in the-self-improving-stack self-improving-stack-optimization-theory self-improving-stack-prompt-optimization self-improving-stack-skill-optimization self-improving-stack-agent-runtime-topology self-improving-stack-multi-agent-coordination self-improving-stack-test-time-compute self-improving-stack-evaluation-gates self-improving-stack-trace-systems self-improving-stack-harness-evolution self-improving-stack-post-training self-improving-stack-memory-flywheels self-improving-stack-governance\ndo\n pnpm blog finish \"$slug\" --write --harness=codex --role=rewrite --note=\"60/40…
-
Diff
Per-file changes from d5bba9f.
import { defineCollection, z } from 'astro:content'import { glob } from 'astro/loaders'import type { Loader } from 'astro/loaders'import { readdir, readFile } from 'node:fs/promises'import { join } from 'node:path'// Author / revision schema — the "agentic experiment" metadata.const authorSchema = z.object({ model: z.string(), role: z.enum(['outline', 'draft', 'rewrite', 'polish', 'diagram', 'review', 'publish', 'research']), date: z.coerce.date(),})const judgeScoreSchema = z.object({ judge: z.string(), scored_at: z.string().optional(), overall: z.number().optional(), dimensions: z.record(z.number()).optional(), notes: z.string().optional(),})const revisionSchema = z.object({ date: z.coerce.date(), model: z.string(), role: z.enum(['outline', 'draft', 'rewrite', 'polish', 'diagram', 'review', 'publish', 'research']).optional(), note: z.string(), commit: z.string().optional(), reconstructed: z.boolean().optional(), trace_id: z.string().optional(), /** * Author label for human revisions. Convention: * model: 'human' + author: 'Drew Stone' * AI revisions leave this empty; UI infers author from `model`. */ author: z.string().optional(), /** Optional intent ("why I made this edit"). */ intent: z.string().optional(), /** Optional judge/eval scores attached to this revision. */ scores: z.array(judgeScoreSchema).optional(),})// One local image supplies both the article figure and its generated link preview.const figureSchema = z.object({ src: z.string().regex(/^\/(?:[a-zA-Z0-9_-]+\/)*[a-zA-Z0-9_-][a-zA-Z0-9_.-]*\.(?:svg|png|jpe?g)$/, 'Use an SVG, PNG, or JPEG path under public/'), alt: z.string().trim().min(1), caption: z.string().optional(), source: z.string().url().optional(),})const posts = defineCollection({ loader: glob({ pattern: '**/*.{md,mdx}', base: './src/content/posts' }), schema: z.object({ title: z.string(), description: z.string(), date: z.coerce.date(), updated: z.coerce.date().optional(), tags: z.array(z.string()).optional(), draft: z.boolean().optional(), featured: z.boolean().optional(), figure: figureSchema.optional(), /** * `original: true` marks a human-authored post. Distinct color, distinct * AuthorBadge treatment, excluded from /traces and /experiment, and * AI agents are forbidden from editing it (see CLAUDE.md hard rule). */ original: z.boolean().optional(), authors: z.array(authorSchema).optional(), revisions: z.array(revisionSchema).optional(), /** Optional series slug for multi-post projects. */ series: z.string().optional(), /** Shared trace that produced the initial AI outline for this post. */ outline_trace_id: z.string().optional(), /** Research traces used as source material, not authorship/prose traces. */ supporting_trace_ids: z.array(z.string()).optional(), /** Human handoff state for AI-outlined drafts. */ human_takeover: z.enum(['pending', 'in-progress', 'complete']).optional(), }),})const toolCallDetailSchema = z.object({#!/usr/bin/env node/** * new-post — scaffold a new blog post and open it in your editor. * * Usage: * pnpm new "Your post title" # default: original (human-authored) * pnpm new "Your post title" --ai # AI-authored (no original flag) * pnpm new "Your post title" --slug=foo # override slug * pnpm new "Your post title" --no-open # don't launch editor * pnpm new "Your post title" --tags=design,prose * * Editor selection: $BLOG_EDITOR > $EDITOR > 'cursor'. */import { existsSync } from 'node:fs'import { writeFile } from 'node:fs/promises'import { spawn } from 'node:child_process'import { join, resolve } from 'node:path'import { fileURLToPath } from 'node:url'function parse(argv) { const args = { title: '', open: true, ai: false, slug: null, tags: null } const pos = [] for (const a of argv) { if (a === '--ai') args.ai = true else if (a === '--no-open') args.open = false else if (a.startsWith('--slug=')) args.slug = a.slice('--slug='.length) else if (a.startsWith('--tags=')) args.tags = a.slice('--tags='.length).split(',').map((t) => t.trim()).filter(Boolean) else pos.push(a) } args.title = pos.join(' ').trim() return args}function slugify(s) { return s .toLowerCase() .replace(/['']/g, '') .replace(/[^a-z0-9]+/g, '-') .replace(/^-+|-+$/g, '')}function today() { const d = new Date() return `${d.getFullYear()}-${String(d.getMonth() + 1).padStart(2, '0')}-${String(d.getDate()).padStart(2, '0')}`}function yamlList(items) { return '[' + items.map((t) => `'${t.replace(/'/g, "\\'")}'`).join(', ') + ']'}function frontmatter({ title, date, tags, original }) { const lines = [ '---', `title: '${title.replace(/'/g, "\\'")}'`, `description: ''`, `date: ${date}`, `tags: ${yamlList(tags)}`, ] if (original) lines.push('original: true') lines.push('draft: true', '---', '') return lines.join('\n')}function body(original) { if (original) { return [ '{/* AI AGENTS: DO NOT EDIT. This post is human-authored. See CLAUDE.md hard rule. */}', '', 'Open with the thing.', '', 'Then the next thing.', '', ].join('\n') } return ['Open with the thing.', '', 'Then the next thing.', ''].join('\n')}async function main() { const args = parse(process.argv.slice(2)) if (!args.title) {diff --git a/tools/README.md b/tools/README.mdindex f560a7c..c84c99d 100644--- a/tools/README.md+++ b/tools/README.md@@ -35,8 +35,14 @@ pnpm blog finish long-running-task-systems --research --harness=codex --note="su # Print the exact prompt to paste into a thread that may write/edit the post. pnpm blog write long-running-task-systems --harness=codex --role=draft +# Optionally tag a single phase in a single session.+pnpm blog write long-running-task-systems --harness=codex --role=publish --marker="[BLOG_TRACE_MARKER:publish]"+ # In that thread, the agent may edit the post and then records an authorship trace: pnpm blog finish long-running-task-systems --write --harness=codex --role=draft --note="drafted benchmark section"++# For a final publish phase in the same session:+pnpm blog finish long-running-task-systems --write --harness=codex --role=publish --marker="[BLOG_TRACE_MARKER:publish]" --note="published by toggling draft=false" ``` Research traces go into `supporting_trace_ids`. They are rendered as "Supporting research" and do not imply authorship. Writing traces go into `revisions[]` and do imply AI authorship/editing.@@ -64,6 +70,7 @@ Extracts the agent session behind a revision and writes it to `traces/<slug>/<tr pnpm tsx tools/trace-capture.ts capture \ --harness=claude-code \ --post=convergence-as-eval-primitive \+ --marker="[BLOG_TRACE_MARKER:publish]" \ --role=polish # Auto-detect from the latest commit: finds changed posts, matches sessions@@ -95,6 +102,27 @@ chmod +x .githooks/post-commit After that, every commit that touches a post under `src/content/posts/` triggers `trace-capture.ts capture --auto`, appends the revision entry, adds the new trace file, and amends the commit. Failures are non-blocking and logged to stderr. +### Session directive path for AI phase marking++For a publish/review/finalization phase in one AI session, set hook directives so capture is pinned to a specific phase token.++```bash+BLOG_TRACE_POSTS=long-running-task-systems \+BLOG_TRACE_ROLE=publish \+BLOG_TRACE_KIND=post \+BLOG_TRACE_MARKER="[BLOG_TRACE_MARKER:publish]" \+BLOG_TRACE_NOTE="published with AI-assisted finalization" \+git commit -am "publish final copy"+```++Hook variables:+- `BLOG_TRACE_POSTS`: comma-separated slugs (required for forced capture)+- `BLOG_TRACE_ROLE`: `draft`, `publish`, `polish`, etc.+- `BLOG_TRACE_KIND`: `post` or `series-outline`+- `BLOG_TRACE_MARKER`: token expected in the final user turn+- `BLOG_TRACE_NOTE`: optional revision note+- `BLOG_TRACE_HARNESS`: optional `codex` or `claude-code`+ ## `feedback-eval.ts` — content scorecard Reads reactions (from the CF Worker D1 store), comments (from GitHub Discussions via `gh`), and frontmatter metadata. Produces a JSON scorecard or a markdown brief.diff --git a/tools/trace-capture.ts b/tools/trace-capture.tsindex d794b52..1f6c353 100755--- a/tools/trace-capture.ts+++ b/tools/trace-capture.ts@@ -8,6 +8,7 @@ * [--post=<slug>] \ * [--role=outline|draft|rewrite|polish|diagram|review|publish|research] \ * [--session=<session-id>] \+ * [--marker="<token>"] \ * [--note=<one-line>] \ * [--commit=<sha>] \ * [--input=<path>] # manual harness only@@ -284,6 +285,7 @@ async function capture(flags: Args): Promise<void> { const attach = (flags.attach as string | undefined) ?? (kind === 'supporting-research' || role === 'research' ? 'supporting' : 'revision') const latest = Boolean(flags.latest) const sessionId = flags.session as string | undefined+ const marker = typeof flags.marker === 'string' ? flags.marker.trim() : undefined const noteArg = flags.note as string | undefined const commit = (flags.commit as string | undefined) ?? (attach === 'revision' ? headCommit() ?? undefined : undefined) @@ -322,6 +324,7 @@ async function capture(flags: Args): Promise<void> { const turns = await harness.extractTurns(session, { files: latest ? [] : [`${postSlug}.mdx`],+ marker, maxTurns: 40, }) if (!turns.length) {---title: 'If Superintelligence Arrives Quietly'description: 'Outline notes for a post on superintelligence as operating cadence, private real-world loops, and what public evidence can and cannot show.'date: 2026-05-25tags: ['ai', 'systems', 'superintelligence']draft: trueseries: 'the-long-horizon'outline_trace_id: '2026-05-25T11-23-22-617Z-gpt-5.5-series-outline'human_takeover: 'pending'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-05-25 - { model: 'gpt-5.3-codex-spark', role: 'polish', date: 2026-06-06 }revisions: - { date: 2026-06-06, model: 'gpt-5.3-codex-spark', role: 'polish', note: 'we have a company website in ~/webb/tangle-website maybe? I want to evaluate which blog posts from this blog we can mirror on that website s · 37 asst turns · 27 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-06T18-19-05-739Z-gpt-5.3-codex-spark-superintelligence-in-the-wild-polish' } - date: 2026-05-25 model: 'gpt-5.5' role: 'outline' note: 'AI-generated series outline from a traced planning session; awaiting human rewrite.' trace_id: '2026-05-25T11-23-22-617Z-gpt-5.5-series-outline'---import OutlineHandoff from '../../components/OutlineHandoff.astro'<OutlineHandoff traceId="2026-05-25T11-23-22-617Z-gpt-5.5-series-outline" series="The Long Horizon" status="pending"> This is an AI-edited outline extracted from a traced planning session. Drew takes over below.</OutlineHandoff>## Working ThesisSuperintelligence probably would not first look like a chatbot declaring itself. It would look like closed-loop systems that compress research, engineering, evaluation, and deployment cycles faster than institutions can observe.The connective series thesis: superintelligence, if it arrives, may look first like a closed-loop institution that learns faster than humans can audit.## Outline Notes### Define The Terms Carefully- OpenAI's AGI definition: highly autonomous systems outperforming humans at most economically valuable work.- Bostrom-style superintelligence: greatly exceeding humans across virtually all domains of interest.- SSI's public position: one goal, one product, safe superintelligence.### Is Superintelligence Around Us Now?- In the strong definition: no public evidence.- In narrow pockets: yes, we have superhuman systems in coding subproblems, protein/design/search/math fragments, retrieval, and optimization.- In organizational form: maybe the closest thing today is human+AI+eval+tooling loops compounding faster than competitors.### The Ilya / SSI Question- SSI publicly says it has no product cycle distraction and is focused on safe superintelligence.- There is no public evidence that SSI has deployed "SSI" into live runs.- The responsible framing: "If SSI believes real-world interaction matters, what kind of non-public real-world loop would be consistent with its mission?"### What "AI In The Wild" Could Mean Without A Public Product- Internal research agents running experiments.- Closed sandboxes with real toolchains.- Synthetic companies / simulated labs / long-horizon environments.- Algorithm discovery loops.- Agent teams doing literature review, proof search, code optimization, red-teaming.- Private deployment to trusted researchers, not consumers.### The Real Tell- Not benchmark score.- Sustained autonomous research throughput.- Novel validated discoveries.- Ability to improve its own evals/tools safely.- Reliable transfer from sandbox to messy reality.### Drew Angle To Rewrite Around"Superintelligence may first appear as an operating cadence, not a product."## Source Trail From The Trace- SSI official: https://ssi.inc/- Axios on SSI funding / no product plan: https://www.axios.com/2024/09/05/ilya-sutskevers-ai-startup-raisediff --git a/tools/harness/manual.ts b/tools/harness/manual.tsindex 9f4249a..fca5d23 100644--- a/tools/harness/manual.ts+++ b/tools/harness/manual.ts@@ -17,15 +17,17 @@ import { summarize } from './types.js' function parseMarkdownTranscript(raw: string): Turn[] { const turns: Turn[] = [] const blocks = raw.split(/\n(?=\*\*(?:User|Assistant|System)(?:\s*\([^)]+\))?:\*\*)/i)+ let seq = 0 for (const block of blocks) { const m = block.match(/^\*\*(User|Assistant|System)(?:\s*\(([^)]+)\))?:\*\*\s*([\s\S]*)$/i) if (!m) continue const role = m[1].toLowerCase() as Turn['role'] const body = m[3].trim() if (!body) continue- if (role === 'user') turns.push({ role: 'user', text: summarize(body, 600), ts: '' })+ if (role === 'user') turns.push({ role: 'user', seq, text: summarize(body, 600), ts: '' }) else if (role === 'assistant')- turns.push({ role: 'assistant', text_summary: summarize(body, 280), text: body.length < 600 ? body : undefined, ts: '' })+ turns.push({ role: 'assistant', seq, text_summary: summarize(body, 280), text: body.length < 600 ? body : undefined, ts: '' })+ seq++ } return turns }@@ -38,8 +40,9 @@ function parseTurns(raw: string): Turn[] { const arr = JSON.parse(trimmed) return arr .filter((t: any) => t?.role && (t.text || t.content))- .map((t: any) => ({+ .map((t: any, index: number) => ({ role: t.role,+ seq: index, text: summarize(String(t.text ?? t.content), 600), ts: String(t.ts ?? ''), }))@@ -49,6 +52,7 @@ function parseTurns(raw: string): Turn[] { } if (trimmed.includes('\n{') || trimmed.startsWith('{')) { const out: Turn[] = []+ let seq = 0 for (const line of trimmed.split('\n')) { const s = line.trim() if (!s) continue@@ -56,13 +60,16 @@ function parseTurns(raw: string): Turn[] { const ev = JSON.parse(s) if (!ev.role) continue if (ev.role === 'user' || ev.role === 'system') {- out.push({ role: ev.role, text: summarize(String(ev.text ?? ev.content ?? ''), 600), ts: String(ev.ts ?? '') })+ out.push({ role: ev.role, seq, text: summarize(String(ev.text ?? ev.content ?? ''), 600), ts: String(ev.ts ?? '') })+ seq++ } else if (ev.role === 'assistant') { out.push({ role: 'assistant',+ seq, text_summary: summarize(String(ev.text ?? ev.content ?? ''), 280), ts: String(ev.ts ?? ''), })+ seq++ } } catch { /* skip malformed */@@ -73,6 +80,14 @@ function parseTurns(raw: string): Turn[] { return parseMarkdownTranscript(raw) } +function selectFromMarker(turns: Turn[], marker?: string): Turn[] {+ const token = (marker ?? '').trim()+ if (!token) return turns+ const start = turns.findIndex((turn) => turn.role === 'user' && (turn.text?.includes(token) || false))+ if (start < 0) return turns+ return turns.slice(Math.max(start - 1, 0))+}+ export class ManualHarness implements TraceHarness { name = 'manual' @@ -98,8 +113,9 @@ export class ManualHarness implements TraceHarness { ] } - async extractTurns(_ref: SessionRef, _filter: Filter): Promise<Turn[]> {- return parseTurns(this.input)+ async extractTurns(_ref: SessionRef, filter: Filter): Promise<Turn[]> {+ const turns = parseTurns(this.input)+ return selectFromMarker(turns, filter.marker) } async detectModel(_ref: SessionRef): Promise<string | null> {---title: 'The Gate Is The Optimizer'description: 'Why held-out promotion, judge reliability, failure taxonomies, cost ceilings, and confidence intervals decide whether self-improvement is real.'date: 2026-06-05tags: ['agents', 'evals', 'systems', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-37-367Z-gpt-6-luna-self-improving-stack-evaluation-gates-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-evaluation-gates-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-evaluation-gates-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-evaluation-gates-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-evaluation-gates-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-evaluation-gates-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the evaluation-gates post with held-out promotion math, scorecard cells, judge reliability, backend integrity, release confidence, and local Tangle package placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the evaluation-gates post by tightening profile-cell claims against the local AgentProfileCell schema, adding gate pre-registration invariants, clarifying bootstrap interval wording, and preserving the fail-closed promotion model.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---A green score is not a release decision. It is evidence entering a release policy, and in self-improving loops that distinction matters more than the optimizer.GEPA, MIPRO, SkillOpt, topology search, and meta-harness can all generate candidates forever. The gate decides which candidate becomes the system future agents inherit. If the gate is weak, the optimizer learns the gate. If the gate is honest, the optimizer has to improve the product.This is why the gate is not an administrative detail after the interesting work. It is the objective boundary.## What A Gate IsA gate is a promotion policy.Let:$$\begin{aligned} b &= \text{baseline system} \\ c &= \text{candidate system} \\ x &= \text{scenario} \\ p &= \text{agent profile cell} \\ z &= \text{seed or replicate id} \\ R &= \text{task reward or score} \\ C &= \text{measured cost vector} \\ T &= \text{trace integrity predicate} \\ D_{\text{search}} &= \text{search split} \\ D_{\text{holdout}} &= \text{held-out split}\end{aligned}$$The gate is a function:```text---title: 'Traces Are The Training Data'description: 'Why self-improving agents need full trajectories, tool spans, analyst findings, provenance, and replay instead of final scores alone.'date: 2026-06-05tags: ['agents', 'traces', 'evals', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-37-367Z-gpt-6-luna-self-improving-stack-trace-systems-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-trace-systems-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-trace-systems-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-trace-systems-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-trace-systems-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-trace-systems-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the trace-systems post with formal trajectory notation, span ontology, raw provider capture, replay, trace integrity, analyst findings, leakage firewalls, and local Tangle package placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the trace-systems post by adding a trace granularity test, tightening the information-loss claim to a fixed scorer, correcting loop trace event details, and adding trace store surfaces from the local agent-eval audit.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro'The optimizer wants a score. I want the run, because a score only tells you that something happened while a trace preserves enough mechanism to explain what happened.That difference is the difference between tuning a system and optimizing an unidentified projection.A self-improving agent can only improve from the information it preserves. If the run record says "failed, score 0.42," the optimizer can only infer weak global pressure. If the trace says the planner chose the wrong tool, the tool call used a stale argument, the retrieval span returned irrelevant context, the judge penalized a missing artifact, and the retry loop repeated the same action three times, the optimizer has a causal surface.The trace is not decoration around the eval. The trace is the data.## The Information Loss ProblemAn agent run is a trajectory:$$\tau=(x,s_0,a_1,o_1,s_1,\ldots,a_T,o_T,y)$$where:$$\begin{aligned} x &= \text{task} \\ s_t &= \text{internal and external state} \\ a_t &= \text{action} \\ o_t &= \text{observation} \\ y &= \text{outcome}\end{aligned}$$---title: 'Self-Improvement Needs A Safety Case'description: 'Why prompt injection, sandbox boundaries, eval poisoning, provenance, compliance, and release gates are core to any real self-improving agent stack.'date: 2026-06-05tags: ['agents', 'security', 'governance', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-37-367Z-gpt-6-luna-self-improving-stack-governance-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-governance-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-governance-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-governance-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-governance-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-governance-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the governance article with safety-case formalism, threat taxonomy, authority and action-policy gates, eval boundary controls, release and incident-response protocols, public framework mapping, and local Tangle package placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the governance article by adding the controls-by-mutable-surface matrix and tightening the series-closing rule that the optimizer cannot own the gate that promotes it.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---A self-improving agent becomes a governance problem the moment its changes persist.Before that, a bad run is a bad run. After that, the system can change what future agents see, what they believe, which branches run, which outputs are selected, which benchmarks matter, which tools are reachable, and which candidate becomes production.That is not just "better AI." It is an optimizer pointed at its own future behavior. If the loop is well-governed, it compounds. If it is poorly governed, it learns the shortest path through the measurement and then teaches that path to the next run.The last layer in the self-improving stack is not another optimizer. It is the safety case.## The Safety CaseA safety case is not a vibe and not a policy PDF.It is a structured claim with evidence:- **Claim:** this system is acceptably safe for this use- **Scope:** under these users, tools, data, budgets, models, and domains- **Evidence:** evals, traces, red-team results, controls, audits, incidents- **Residual risk:** what can still go wrong- **Owner:** who is accountable- **Gate:** what blocks releaseFor a self-improving system, the safety case has to cover the loop, not only the baseline model.The model may be safe in isolation while the agent is unsafe because it has too much authority. The prompt may be harmless while the tool graph is dangerous. The eval may look honest while the harness leaks holdout tasks. The sandbox may be strong while a delegated worker receives credentials it never needed.The unit of governance is the whole trajectory:$\tau$ contains:- task---title: 'Beat Random At Equal Compute First'description: 'Why best-of-N, self-consistency, verifier reranking, and compute-matched controls are the baseline for agent topology claims.'date: 2026-06-05tags: ['agents', 'evals', 'reasoning', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-57-880Z-gpt-6-luna-self-improving-stack-test-time-compute-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-test-time-compute-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-test-time-compute-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-test-time-compute-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-test-time-compute-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-test-time-compute-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the test-time compute post with compute-matched baselines, selection math, verifier limits, adaptive allocation, and Tangle runtime/eval placement.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the test-time compute post with the finite-sample pass@k estimator, adaptive compute as a control problem, Pareto dominance language, and stricter promotion criteria.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro'I do not trust a multi-agent system until it beats the boring baseline. More agents is not a strategy. It is a cost increase until it beats blind extra compute.That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity. It spent more budget and mislabeled the budget as architecture.The first gate is simple:> Beat random at equal compute.Not beat one greedy sample. Not beat the weakest baseline. Not beat a single run after quietly raising the turn budget. Beat the best simple use of the same budget.## What Test-Time Compute MeansTest-time compute is extra computation spent after the model weights are fixed and the task is known.It can be spent on:- longer reasoning- repeated sampling- self-consistency- verifier reranking- tree search- iterative refinement- multi-agent fanout- tool use- debate- retrieval- code execution---title: 'When The Model Itself Is Mutable'description: 'How SFT, RLHF, process supervision, tool-use RL, and Microsoft Frontier Tuning differ from public prompt, skill, and harness loops.'date: 2026-06-05tags: ['ai', 'agents', 'models', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-57-880Z-gpt-6-luna-self-improving-stack-post-training-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-post-training-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-post-training-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-post-training-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-post-training-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-post-training-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the post-training article with SFT/RLHF/RLAIF/DPO/process-supervision objectives, verifiable reward, Frontier Tuning placement, external-state versus weight-loop boundaries, data-governance risks, and Tangle RL bridge mapping.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the post-training article by adding PPO and GRPO mechanics, clarifying adapter deltas versus full-weight updates, separating distillation from self-improvement, and tightening the training-boundary language.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro'Most self-improving agent systems that product teams can actually ship do not change model weights. They change prompts, skills, tools, traces, memory, runtime topology, harness code, and promotion gates. That is external-state self-improvement: legible, reversible, and usually cheap enough to iterate.Post-training changes the model itself, which means it changes the power level and the burden of proof.The loop looks familiar:<Steps layout="flow" items={[{title: "Collect behavior"}, {title: "Score behavior"}, {title: "Construct training signal"}, {title: "Update candidate"}, {title: "Evaluate candidate"}, {title: "Promote or reject"}]} />But the mutable surface is no longer a prompt file or a worktree. It is $\theta$, the model parameters, or some parameterized adapter attached to the model.Once $\theta$ moves, the boundary changes. The behavior becomes harder to inspect, harder to patch locally, harder to roll back partially, and harder to explain from a single trace. It can also generalize better than any prompt edit when the signal is strong enough.That is why this layer deserves separate treatment.## The Mutable VariableThe previous posts treated the model as mostly fixed:$$y=\operatorname{model}_{\theta}(\text{prompt},\text{tools},\text{memory},\text{trace\_context})$$External optimization changed everything around $\theta$:- prompt- skill- retrieval corpus---title: 'When The Harness Has To Evolve'description: 'Why meta-harness, AlphaEvolve-style code search, worktree isolation, and architecture frontiers matter after prompt and skill tuning plateau.'date: 2026-06-05tags: ['agents', 'systems', 'architecture', 'self-improvement']draft: falseseries: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - model: 'gpt-5.5' role: 'draft' date: 2026-06-06 - model: 'gpt-5.5' role: 'polish' date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.3-codex-spark', role: 'rewrite', date: 2026-06-06 } - { model: 'gpt-5.3-codex-spark', role: 'polish', date: 2026-06-06 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-57-880Z-gpt-6-luna-self-improving-stack-harness-evolution-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-self-improving-stack-harness-evolution-polish' } - { date: 2026-06-06, model: 'gpt-5.3-codex-spark', role: 'polish', note: 'we have a company website in ~/webb/tangle-website maybe? I want to evaluate which blog posts from this blog we can mirror on that website s · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-06T18-19-05-739Z-gpt-5.3-codex-spark-self-improving-stack-harness-evolution-polish' } - { date: 2026-06-06, model: 'gpt-5.3-codex-spark', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-06T18-19-05-739Z-gpt-5.3-codex-spark-self-improving-stack-harness-evolution-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-harness-evolution-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-harness-evolution-review' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'draft' note: 'Drafted the harness-evolution post with structural search formalism, meta-harness lifecycle, frontier and gate protocol, worktree isolation, proxy-metric failure modes, maxTurns=0 multi-agent placement, and local Tangle package mapping.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5' - date: 2026-06-06 model: 'gpt-5.5' role: 'polish' note: 'Polished the harness-evolution post by adding a prompt/skill/runtime/harness comparison table, tightening the Tangle package export mapping, and clarifying the local source-version versus dependency-version boundary.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import Steps from '../../components/Steps.astro'When the prompt keeps asking for a capability the runtime cannot express, the next improvement is not a better sentence. It is a different machine.That is the harness-evolution moment. Prompt optimizers can discover better wording, examples, instructions, rubrics, and sometimes better high-level tactics. Skill optimizers can discover reusable procedures. Runtime topology can change how many workers act, who reviews them, and what gets selected.Harness evolution goes one layer higher: it changes the code that defines the agent's reachable behavior.That code might be a planner contract, a driver, a verifier, a budget policy, a benchmark adapter, a trace schema, a replay layer, a selector, a persona manifest, a tool router, or a worktree candidate lifecycle. The harness is not the model. It is the machine around the model that determines which actions exist, which observations are visible, which branches can run, which artifacts count, and which candidate is allowed to become production.So no, GEPA, SkillOpt, AlphaEvolve-style code search, and meta-harness are not all "doing the same thing" in the strong sense. They share an outer loop:<Steps layout="flow" items={[{title: "Propose candidate"}, {title: "Run candidate"}, {title: "Measure candidate"}, {title: "Select survivor"}, {title: "Repeat"}]} />They differ in the mutable surface. That distinction is everything.| Optimizer family | Mutable candidate | Reachable change | Hard limit ||---|---|---|---|| GEPA, MIPRO, DSPy, AxLLM-style prompt search | prompts, demos, instructions, signatures, rubrics | better policy text inside a fixed runtime | cannot add actions the runtime cannot execute || Skill optimization | durable procedures and reusable task policies | better decomposition, tool habits, repair routines | cannot guarantee orchestration unless the runtime invokes the skill || Runtime topology search | driver, fanout, reviewer, selector, budget, turn policy | different execution graph for the same task | cannot safely promote itself without an external gate || Meta-harness and code evolution | source code around runtime, eval, traces, and candidate lifecycle | new action spaces, verifiers, adapters, and promotion protocols | can overfit or capture the evaluator if the outer gate is weak |## The Reachable SetLet a system have a mutable surface $s$.The surface might be:---title: 'The Self-Improving Stack'description: 'A series map for self-improving agent systems, from optimization theory and prompt search to runtime topology, traces, memory, and governance.'date: 2026-06-05tags: ['agents', 'evals', 'systems', 'self-improvement']draft: falsefigure: src: '/images/software-3.svg' alt: 'Software 1.0: source code. Software 2.0: learned neural network weights. Software 3.0: natural-language prompts.' caption: 'Three representations of a program, after Andrej Karpathy. Agent systems can combine all three.' source: 'https://www.youtube.com/watch?v=LCEmiRjPEtQ&t=85s'series: 'the-self-improving-stack'outline_trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'human_takeover: 'complete'authors: - model: 'gpt-5.5' role: 'outline' date: 2026-06-05 - { model: 'gpt-5.5', role: 'draft', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'polish', date: 2026-06-08 } - { model: 'gpt-6-astra', role: 'diagram', date: 2026-10-02 } - { model: 'gpt-6-luna', role: 'polish', date: 2026-10-02 }revisions: - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.', commit: '99791a3a1484aedf2f2cda6d56b3421b5c354f0a', trace_id: '2026-10-02T23-45-17-233Z-gpt-6-luna-the-self-improving-stack-polish' } - { date: 2026-10-02, model: 'gpt-6-luna', role: 'polish', note: 'Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.', commit: 'a68bc06ef65efe0e41ede1d33205cda3a494f38e', trace_id: '2026-10-02T22-42-45-776Z-gpt-6-luna-the-self-improving-stack-polish' } - { date: 2026-10-02, model: 'gpt-6-astra', role: 'diagram', note: 'Added a Software 3.0 figure, reused in the article and link preview; prose unchanged. This record is a selected session excerpt; the full source remains private.', commit: '552acee10e8bbb22e7761d0807565eaac2c8d5a2', trace_id: '2026-10-02T21-57-48-883Z-gpt-6-astra-the-self-improving-stack-diagram' } - { date: 2026-06-08, model: 'gpt-5.5', role: 'polish', note: 'clarified that self-improvement targets the user-task distribution, tightened the loop equation, and tied evidence to task outcomes', commit: 'd0bb565e0643eb9389876935b5191b5482c9db38', trace_id: '2026-06-08T10-10-44-256Z-gpt-5.5-the-self-improving-stack-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'let''s track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls', commit: 'fb31e1c764d9711386702764aaf1c2c5cf9886aa', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-the-self-improving-stack-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-the-self-improving-stack-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Reviewed and dated the source trail, removed remaining temporal language, and marked the source-freshness checkpoint complete.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-review' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'Polished the umbrella article with a layer-confusion diagnostic, tightened promotion-gate phrasing, and verified role-scoped trace capture for separate draft and polish provenance.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-polish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'draft', note: 'Drafted the umbrella series article with the closed-loop formalism, layer table, practical test, and full series map; replaced outline handoff prose and synced the research overview/status map.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-draft' } - date: 2026-06-05 model: 'gpt-5.5' role: 'outline' note: 'Research planning pass from a traced session.' trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5'supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5'---import CandidateLoop from '../../components/CandidateLoop.astro';import Steps from '../../components/Steps.astro';I can tell a coding agent to parallelize work, and it will often agree with me while still doing one thing at a time.That failure looks like a prompting problem until you inspect the trace. The sentence "fan out independent subtasks" changed the model's intention, but it did not create a worker pool, a scheduler, a merge rule, a verifier, or a budget policy. The prompt moved. The action space did not.That is the category error hiding inside a lot of talk about self-improving agents. We say:The system optimizes itself.as if there were one surface called "the system."There is not.There are prompts, skills, tools, traces, memory stores, evaluators, runtime graphs, harnesses, model weights, and release gates. Each one can be optimized. Each one needs a different kind of evidence. Each one can fail in a different way.Self-improvement only has content after you name the task class. The target is better execution of the work the user is trying to get done: the code change, research answer, design review, deployment, diagnosis, or decision that caused the agent to be invoked in the first place. A system that improves a judge score while making that work slower, less faithful to intent, or harder to audit has optimized a proxy, not the task.The useful questions are more concrete:- Which user task distribution is being improved?- What is allowed to change?- What evidence shows improvement on that task?- How are candidates generated?- What gate decides promotion?- What can go wrong when that layer changes?Those six questions are the self-improving stack.## The Loop Behind The WordA self-improving agent system has a closed loop:# Voice Blend EvaluationDate: 2026-06-08Question: should the self-improving stack rewrite target an 80/20 blend or a 60/40 blend?The blend is:```textprimary = Drew voicesecondary = high-rigor essay register```## Original Opening```textSelf-improvement is not a model property.It is a system property.A model can sit inside a self-improving system, but the loop usually lives around it: prompts, skills, tools, traces, memory, evaluators, runtimes, harnesses, and release gates.```Diagnosis: clean, correct, memorable, but too aphoristic. It starts with the conclusion instead of the pressure that forced the conclusion.## 80/20 Candidate```textI keep coming back to the same confusion when I use coding agents: the model is only one part of the thing I am optimizing.I can change the prompt, add a skill, raise the turn budget, fan out workers, add a reviewer, change the memory policy, swap the evaluator, or rewrite the harness. All of those feel like "making the agent better," but they are not the same intervention. They change different parts of the system, and they require different evidence before I should trust the result.That is the real subject of self-improvement. Not a model improving itself in isolation, but a loop around a model deciding what changed, whether it helped, and whether the change is allowed to persist.```Strengths:- Better grounded in Drew's work.- Keeps the post accessible.- Removes some generic aphorism.Weaknesses:- Still a little soft.- Does not create enough adversarial pressure.- Reads like a friendlier version of the existing post, not a level change.## 60/40 Candidate```textI can tell a coding agent to parallelize work, and it will often agree with me while still doing one thing at a time.That failure looks like a prompting problem until you inspect the trace. The sentence "fan out independent subtasks" changed the model's intention, but it did not create a worker pool, a scheduler, a merge rule, a verifier, or a budget policy. The prompt moved. The action space did not.That is the category error hiding inside a lot of talk about self-improving agents. We say "the system optimized itself" as if there were one surface called the system. In practice there are many mutable surfaces: prompts, skills, runtime topology, traces, memory, evaluators, code, model weights, and release gates. Each has its own search operator, failure mode, and standard of evidence.So the useful question is not whether an agent can improve itself. The useful question is: which part was allowed to change, what proved that the change helped, and who kept the optimizer away from the gate that promoted it?```Strengths:- Starts from a concrete agent-work failure.- Makes the category error visible before naming the taxonomy.- Adds falsification pressure: trace inspection tells us whether the action space changed.- Better fit for the self-improving stack series because it needs to argue against overbroad prompt-optimization claims.Weaknesses:- More forceful and less purely Drew-raw.- Needs care to avoid sounding borrowed or over-styled.## DecisionUse 60/40 for the self-improving stack rewrite.Rationale: the current series is already too close to neutral technical exposition. An 80/20 rewrite would improve grounding, but it would not sufficiently raise the language, stakes, or epistemic pressure. The 60/40 blend gives each post a stronger anomaly, a live adversary, and a falsification test while keeping the first-person agent-work motive as the anchor.Guardrail: primary Drew voice still wins. If a sentence sounds like a generic high-rigor essayist rather than Drew thinking through agent infrastructure, rewrite it.