GPT-5.5 codex

Topology Is The Missing Action Space

let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls

Created
Updated
32
Turns
23
Tool calls
18
Files touched
4161m
Duration

Files

Commit

fb31e1c chore(workspace): cleanup merged in-progress content and trace artifacts
+208 −51 across 14 files

Conversation

32 turns. Full text where captured; older traces show only the first ~280 chars.

  1. user #1
    let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our knowledge and you should create a blog post for each and we make this a series
  2. assistant #2
    Problem: the research is currently a single trace plus a chat summary, so it is hard to grow into a series without losing the map. Change: I’ll add a durable research checkpoint directory with one file per map item, then scaffold one draft post per item under a consistent series. Why long-term right: research notes stay source-of-truth and appendable, while posts stay draft publication surfaces with provenance. Cost: new markdown/MDX files only; risk is schema/provenance mismatch, rollback is deleting the new directory/posts.
  3. assistant #3
    I’m checking the post schema and scaffold tooling before writing files, so the new drafts land in the repo’s existing provenance format instead of inventing one.
  4. assistant #4 1 tool
    • exec_command /Users/drew/code/blog
      input
      {"cmd":"git status --short --branch","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
  5. assistant #5 1 tool
    • exec_command /Users/drew/code/blog
      input
      {"cmd":"git log --oneline -10","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
  6. assistant #6 1 tool
    • exec_command /Users/drew/code/blog
      input
      {"cmd":"git reflog | head -20","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
  7. assistant #7 1 tool
    • exec_command /Users/drew/code/blog
      input
      {"cmd":"gh pr list --state open","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
  8. assistant #8 1 tool
    • exec_command src/content.config.ts
      input
      {"cmd":"sed -n '1,260p' src/content.config.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":6000}
  9. assistant #9 1 tool
    • exec_command tools/new-post.mjs
      input
      {"cmd":"sed -n '1,260p' tools/new-post.mjs","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":7000}
  10. assistant #10
    The schema supports `series`, `outline_trace_id`, `supporting_trace_ids`, and `human_takeover`, which is enough for draft series scaffolds. I’m going to create a `research/self-improving-agent-systems/` checkpoint directory plus 12 draft MDX outlines that all point back to the existing supporting trace.
  11. assistant #11 1 tool
    • exec_command /Users/drew/code/blog
      input
      {"cmd":"find . -maxdepth 2 -type d | sort | sed -n '1,160p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":4000}
  12. assistant #12 1 tool
    • exec_command tools/README.md
      input
      {"cmd":"rg -n \"series:|self-improving|long-horizon|supporting_trace_ids|outline_trace_id\" src/content/posts traces README.md tools/README.md --glob '!dist'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
  13. assistant #13 1 tool
    • exec_command tools/trace-capture.ts
      input
      {"cmd":"sed -n '1,320p' tools/trace-capture.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
  14. assistant #14 1 tool
    • exec_command traces/the-long-horizon/2026-05-25T11-23-22-617Z-gpt-5.5-series-outline.js
      input
      {"cmd":"sed -n '1,140p' traces/the-long-horizon/2026-05-25T11-23-22-617Z-gpt-5.5-series-outline.json","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":8000}
  15. assistant #15 1 tool
    • exec_command src/content/posts/superintelligence-in-the-wild.mdx
      input
      {"cmd":"sed -n '1,80p' src/content/posts/superintelligence-in-the-wild.mdx","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":4000}
  16. assistant #16 1 tool
    • exec_command tools/harness/manual.ts
      input
      {"cmd":"sed -n '1,220p' tools/harness/manual.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":6000}
  17. assistant #17 1 tool
    • exec_command traces/the-self-improving-stack/2026-06-05T12-35-48-868Z-gpt-5.5-the-self-improving-stack-rewrite.js
      input
      {"cmd":"jq '.turns[0], .turns[-1]' traces/the-self-improving-stack/2026-06-05T12-35-48-868Z-gpt-5.5-the-self-improving-stack-rewrite.json","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":20000}
  18. assistant #18 1 tool
    • exec_command tools/trace-capture.ts
      input
      {"cmd":"sed -n '260,460p' tools/trace-capture.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":24000}
  19. assistant #19 1 tool
    • exec_command tools/harness/types.ts
      input
      {"cmd":"sed -n '1,170p' tools/harness/types.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":20000}
  20. assistant #20
    The prose review has two levels: publishing infrastructure first, then style. The publishing bug is higher priority because the blog now visibly reports a false-ish revision lineage even though the articles build.
  21. assistant #21 1 tool
    • exec_command src/content/posts/the-self-improving-stack.mdx
      input
      {"cmd":"nl -ba src/content/posts/the-self-improving-stack.mdx | sed -n '1,70p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
  22. assistant #22 1 tool
    • exec_command src/content/posts/self-improving-stack-harness-evolution.mdx
      input
      {"cmd":"nl -ba src/content/posts/self-improving-stack-harness-evolution.mdx | sed -n '1,75p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
  23. assistant #23 1 tool
    • exec_command src/content/posts/self-improving-stack-trace-systems.mdx
      input
      {"cmd":"nl -ba src/content/posts/self-improving-stack-trace-systems.mdx | sed -n '40,75p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
  24. assistant #24 1 tool
    • exec_command src/content/posts/self-improving-stack-agent-runtime-topology.mdx
      input
      {"cmd":"nl -ba src/content/posts/self-improving-stack-agent-runtime-topology.mdx | sed -n '50,70p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
  25. assistant #25 1 tool
    • exec_command src/content/posts/self-improving-stack-test-time-compute.mdx
      input
      {"cmd":"nl -ba src/content/posts/self-improving-stack-test-time-compute.mdx | sed -n '385,405p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
  26. assistant #26 1 tool
    • exec_command src/content/posts/self-improving-stack-evaluation-gates.mdx
      input
      {"cmd":"nl -ba src/content/posts/self-improving-stack-evaluation-gates.mdx | sed -n '475,485p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
  27. assistant #27
    The writing itself is materially better than before, but I would not call the package publish-clean yet. The top blocker is trace accuracy; the second issue is a remaining repeated contrast cadence that the new voice guide explicitly warns against.
  28. assistant #28 1 tool
    • exec_command /Users/drew/code/blog
      input
      {"cmd":"pnpm build","workdir":"/Users/drew/code/blog","yield_time_ms":30000,"max_output_tokens":30000}
  29. assistant #29 1 tool
    • exec_command src/content/posts/the-self-improving-stack.mdx
      input
      {"cmd":"git diff --stat -- src/content/posts/the-self-improving-stack.mdx src/content/posts/self-improving-stack-*.mdx VOICE.md SECONDARY_VOICE.md research/self-improving-agent-systems/voice-blend-eval.md .evolve/skill-runs.jsonl AGENTS.md CLAUDE.md","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":20000}
  30. assistant #30
    Review complete. I’m not patching in this pass because the biggest issue is provenance policy, not a one-line prose typo; it needs a deliberate trace cleanup rather than quietly editing metadata by hand.
  31. assistant #31
    **Findings** High: the rewrite provenance is wrong enough that I would not ship this as-is. The 60/40 rewrite happened on June 8, but the posts now show rewrite authorship dated June 5 or June 6 with older model labels and trace IDs, for example [the-self-improving-stack.mdx](/Users/drew/code/blog/src/content/posts/the-self-improving-stack.mdx:18) and [self-improving-stack-harness-evolution.mdx](/Users/drew/code/blog/src/content/posts/self-improving-stack-harness-evolution.mdx:22). The trace payload confirms it captured an older long session starting June 5, with the June 8 finish command only appearing at the tail. That breaks the blog’s trace integrity story. Medium: the prose is much better, but the series still overuses the exact contrast cadence the new voice guide warns against. Examples: [trace-systems](/Users/drew/code/blog/src/content/posts/self-improving-stack-trace-systems.mdx:52), [runtime-topology](/Users/drew/code/blog/src/content/posts/self-improving-stack-agent-runtime-topology.mdx:58), [harness-evolution](/Users/drew/code/blog/src/content/posts/self-improving-stack-harness-evolution.mdx:64). It is not fatal, but it makes the series feel more “AI-polished” than Drew-authored in places. Low: the rewrite improves the openings, but some posts now have a strong first paragraph and then fall back into framework exposition quickly. The umbrella, prompt optimization, skill optimization, and topology posts are closest. Evaluation gates, traces, governance, and post-training still read more like excellent internal docs than essays. Build status: `pnpm build` passes. Same unrelated warning remains: unused `stat` import in `tools/og-render.ts`. My honest take: content quality is around 8/10 now, up from maybe 6.5/10. Publish-quality requires fixing provenance first, then doing one narrower style pass to reduce repeated “not X, it is Y” moves and add more lived connective tissue in the middle sections. The technical argument is strong; the trace layer is currently the blocker.
  32. user #32
    ok then beef up the prompt (your prompt) and ensure it uses the voice and improve this

Diff

Per-file changes from fb31e1c.

src/content.config.ts +1 −0
diff --git a/src/content.config.ts b/src/content.config.tsindex 9b45af5..b460bec 100644--- a/src/content.config.ts+++ b/src/content.config.ts@@ -76,6 +76,7 @@ const toolCallDetailSchema = z.object({ })  const turnSchema = z.object({+  seq: z.number().optional(),   role: z.enum(['user', 'assistant', 'system', 'tool']),   text: z.string().optional(),   text_summary: z.string().optional(),
tools/new-post.mjs current file (first 80 lines)
#!/usr/bin/env node/** * new-post — scaffold a new blog post and open it in your editor. * * Usage: *   pnpm new "Your post title"             # default: original (human-authored) *   pnpm new "Your post title" --ai        # AI-authored (no original flag) *   pnpm new "Your post title" --slug=foo  # override slug *   pnpm new "Your post title" --no-open   # don't launch editor *   pnpm new "Your post title" --tags=design,prose * * Editor selection: $BLOG_EDITOR > $EDITOR > 'cursor'. */import { existsSync } from 'node:fs'import { writeFile } from 'node:fs/promises'import { spawn } from 'node:child_process'import { join, resolve } from 'node:path'import { fileURLToPath } from 'node:url'function parse(argv) {  const args = { title: '', open: true, ai: false, slug: null, tags: null }  const pos = []  for (const a of argv) {    if (a === '--ai') args.ai = true    else if (a === '--no-open') args.open = false    else if (a.startsWith('--slug=')) args.slug = a.slice('--slug='.length)    else if (a.startsWith('--tags=')) args.tags = a.slice('--tags='.length).split(',').map((t) => t.trim()).filter(Boolean)    else pos.push(a)  }  args.title = pos.join(' ').trim()  return args}function slugify(s) {  return s    .toLowerCase()    .replace(/['']/g, '')    .replace(/[^a-z0-9]+/g, '-')    .replace(/^-+|-+$/g, '')}function today() {  const d = new Date()  return `${d.getFullYear()}-${String(d.getMonth() + 1).padStart(2, '0')}-${String(d.getDate()).padStart(2, '0')}`}function yamlList(items) {  return '[' + items.map((t) => `'${t.replace(/'/g, "\\'")}'`).join(', ') + ']'}function frontmatter({ title, date, tags, original }) {  const lines = [    '---',    `title: '${title.replace(/'/g, "\\'")}'`,    `description: ''`,    `date: ${date}`,    `tags: ${yamlList(tags)}`,  ]  if (original) lines.push('original: true')  lines.push('draft: true', '---', '')  return lines.join('\n')}function body(original) {  if (original) {    return [      '{/* AI AGENTS: DO NOT EDIT. This post is human-authored. See CLAUDE.md hard rule. */}',      '',      'Open with the thing.',      '',      'Then the next thing.',      '',    ].join('\n')  }  return ['Open with the thing.', '', 'Then the next thing.', ''].join('\n')}async function main() {  const args = parse(process.argv.slice(2))  if (!args.title) {
tools/README.md current file (first 80 lines)
# tools/Scripts that capture, shape, and evaluate the blog's agentic data.## `new-post.mjs` / `edit-post.mjs` — draft and human edit helpers```bash# Create a human-authored draft and open it.pnpm new "Post title" --tags=agents,systems# Create an AI-assisted draft.pnpm new "Post title" --ai --tags=agents,systems# Open an existing post by slug or title substring.pnpm write long-running-task-systems# After editing, commit and let the post-commit hook record the green human revision.pnpm write long-running-task-systems --commit --note="rewrote the outline into a first human draft"# Mark an AI-outline handoff as complete and publish.pnpm write long-running-task-systems --done --publish --commit --note="publish human rewrite"```## `blog-loop.mjs` — traced AI lifecycleUse this when starting a clean AI thread.```bash# Print the exact prompt to paste into a clean research thread.pnpm blog research long-running-task-systems --harness=codex# In that thread, the agent does not edit the post. At the end it runs:pnpm blog finish long-running-task-systems --research --harness=codex --note="surveyed long-horizon benchmarks"# Print the exact prompt to paste into a thread that may write/edit the post.pnpm blog write long-running-task-systems --harness=codex --role=draft# The generated prompt includes a trace marker, voice checklist, anti-pattern gates,# and the exact finish command. Keep the marker in the first assistant update and# in the final response so trace capture can isolate the current phase.# In that thread, the agent may edit the post and then records an authorship trace:pnpm blog finish long-running-task-systems --write --harness=codex --role=draft --marker="BLOGTRACE-..." --note="drafted benchmark section"# For a final publish phase in the same session:pnpm blog finish long-running-task-systems --write --harness=codex --role=publish --marker="[BLOG_TRACE_MARKER:publish]" --note="published by toggling draft=false"```Research traces go into `supporting_trace_ids`. They are rendered as "Supporting research" and do not imply authorship. Writing traces go into `revisions[]` and do imply AI authorship/editing. Unmarked write finishes are refused by `pnpm blog finish`; use `--session=<id>` or `--allow-unmarked` only for audited recovery captures.If a thread started before you decided the target post, tell the agent:```textThis thread is supporting research for <post-slug>. Do not mark it as authorship. Attach this session as supporting research using the blog lifecycle.```Then the agent should run:```bashpnpm blog finish <post-slug> --research --harness=codex --note="supporting research"```## `trace-capture.ts` — harness-agnostic session captureExtracts the agent session behind a revision and writes it to `traces/<slug>/<trace_id>.json`. Appends a revisions entry to the post's frontmatter that links back.### Commands```bash# Capture a specific post with an explicit harnesspnpm tsx tools/trace-capture.ts capture \  --harness=claude-code \  --post=convergence-as-eval-primitive \  --marker="[BLOG_TRACE_MARKER:publish]" \  --role=polish# Auto-detect from the latest commit: finds changed posts, matches sessions# via ~/.claude/projects/ or ~/.codex/sessions/, writes traces + appends# frontmatter entries.pnpm tsx tools/trace-capture.ts capture --auto
tools/trace-capture.ts current file (first 80 lines)
#!/usr/bin/env node/** * trace-capture: harness-agnostic session capture for blog revisions. * * Usage: *   pnpm tsx tools/trace-capture.ts capture \ *     [--harness=claude-code|codex|manual] \ *     [--post=<slug>] \ *     [--role=outline|draft|rewrite|polish|diagram|review|publish|research] \ *     [--session=<session-id>] \ *     [--marker="<token>"] \ *     [--note=<one-line>] \ *     [--commit=<sha>] \ *     [--input=<path>]            # manual harness only *     [--kind=post|series-outline|supporting-research] *     [--attach=supporting|revision|none] *     [--latest]                  # choose latest session without requiring post file touch * *   pnpm tsx tools/trace-capture.ts capture --auto *     # detects from the latest git commit: finds changed posts, matches a *     # recent session via ~/.claude/projects or ~/.codex/sessions, writes a *     # trace per changed post, appends to frontmatter. * *   pnpm tsx tools/trace-capture.ts list *     # list existing traces grouped by post. * *   pnpm tsx tools/trace-capture.ts show <trace_id> *     # dump a trace as JSON. */import { execSync } from 'node:child_process'import { mkdir, readdir, readFile, writeFile } from 'node:fs/promises'import { join } from 'node:path'import ClaudeCodeHarness from './harness/claude-code.js'import CodexHarness from './harness/codex.js'import ManualHarness from './harness/manual.js'import { dedupeAdjacentTurns, type TraceFile, type TraceHarness, type Turn } from './harness/types.js'const ROOT = process.cwd()const POSTS_DIR = join(ROOT, 'src/content/posts')const TRACES_DIR = join(ROOT, 'traces')type Args = Record<string, string | boolean>function parseArgs(argv: string[]): { cmd: string; pos: string[]; flags: Args } {  const [cmd, ...rest] = argv  const flags: Args = {}  const pos: string[] = []  for (const a of rest) {    if (a.startsWith('--')) {      const eq = a.indexOf('=')      if (eq >= 0) flags[a.slice(2, eq)] = a.slice(eq + 1)      else flags[a.slice(2)] = true    } else pos.push(a)  }  return { cmd: cmd ?? 'capture', pos, flags }}function git(cmd: string): string {  try {    return execSync(`git ${cmd}`, { cwd: ROOT, stdio: ['ignore', 'pipe', 'ignore'] }).toString().trim()  } catch {    return ''  }}function headCommit(): string | null {  const sha = git('rev-parse HEAD')  return sha || null}function changedPostsAtHead(): string[] {  const out = git('show --no-renames --name-only --format="" HEAD')  return out    .split('\n')    .map((l) => l.trim())    .filter((l) => l.startsWith('src/content/posts/') && l.endsWith('.mdx'))    .map((l) => l.replace('src/content/posts/', '').replace(/\.mdx$/, ''))}
src/content/posts/superintelligence-in-the-wild.mdx +82 −0
diff --git a/src/content/posts/superintelligence-in-the-wild.mdx b/src/content/posts/superintelligence-in-the-wild.mdxnew file mode 100644index 0000000..e879076--- /dev/null+++ b/src/content/posts/superintelligence-in-the-wild.mdx@@ -0,0 +1,82 @@+---+title: 'If Superintelligence Arrives Quietly'+description: 'Outline notes for a post on superintelligence as operating cadence, private real-world loops, and what public evidence can and cannot show.'+date: 2026-05-25+tags: ['ai', 'systems', 'superintelligence']+draft: true+series: 'the-long-horizon'+outline_trace_id: '2026-05-25T11-23-22-617Z-gpt-5.5-series-outline'+human_takeover: 'pending'+authors:+  - model: 'gpt-5.5'+    role: 'outline'+    date: 2026-05-25+revisions:+  - date: 2026-05-25+    model: 'gpt-5.5'+    role: 'outline'+    note: 'AI-generated series outline from a traced planning session; awaiting human rewrite.'+    trace_id: '2026-05-25T11-23-22-617Z-gpt-5.5-series-outline'+---++import OutlineHandoff from '../../components/OutlineHandoff.astro'++<OutlineHandoff traceId="2026-05-25T11-23-22-617Z-gpt-5.5-series-outline" series="The Long Horizon" status="pending">+  This is an AI-edited outline extracted from a traced planning session. Drew takes over below.+</OutlineHandoff>++## Working Thesis++Superintelligence probably would not first look like a chatbot declaring itself. It would look like closed-loop systems that compress research, engineering, evaluation, and deployment cycles faster than institutions can observe.++The connective series thesis: superintelligence, if it arrives, may look first like a closed-loop institution that learns faster than humans can audit.++## Outline Notes++### Define The Terms Carefully++- OpenAI's AGI definition: highly autonomous systems outperforming humans at most economically valuable work.+- Bostrom-style superintelligence: greatly exceeding humans across virtually all domains of interest.+- SSI's public position: one goal, one product, safe superintelligence.++### Is Superintelligence Around Us Now?++- In the strong definition: no public evidence.+- In narrow pockets: yes, we have superhuman systems in coding subproblems, protein/design/search/math fragments, retrieval, and optimization.+- In organizational form: maybe the closest thing today is human+AI+eval+tooling loops compounding faster than competitors.++### The Ilya / SSI Question++- SSI publicly says it has no product cycle distraction and is focused on safe superintelligence.+- There is no public evidence that SSI has deployed "SSI" into live runs.+- The responsible framing: "If SSI believes real-world interaction matters, what kind of non-public real-world loop would be consistent with its mission?"++### What "AI In The Wild" Could Mean Without A Public Product++- Internal research agents running experiments.+- Closed sandboxes with real toolchains.+- Synthetic companies / simulated labs / long-horizon environments.+- Algorithm discovery loops.+- Agent teams doing literature review, proof search, code optimization, red-teaming.+- Private deployment to trusted researchers, not consumers.++### The Real Tell++- Not benchmark score.+- Sustained autonomous research throughput.+- Novel validated discoveries.+- Ability to improve its own evals/tools safely.+- Reliable transfer from sandbox to messy reality.++### Drew Angle To Rewrite Around++"Superintelligence may first appear as an operating cadence, not a product."++## Source Trail From The Trace++- SSI official: https://ssi.inc/+- Axios on SSI funding / no product plan: https://www.axios.com/2024/09/05/ilya-sutskevers-ai-startup-raise+- OpenAI Charter AGI definition: https://openai.com/charter/+- Bostrom definition: https://nickbostrom.com/superintelligence+- Dwarkesh/Ilya episode summary: https://www.tapesearch.com/episode/dwarkesh-and-ilya-sutskever-on-what-comes-after-scaling/9H4vn7L2gPKLaenJWdWnM4+- Anthropic RSP / frontier safety framing: https://www.anthropic.com/news/responsible-scaling-policy-v3
tools/harness/manual.ts current file (first 80 lines)
/** * Manual harness — read a JSONL or plain-text transcript from a file or stdin. * * Useful when a revision comes from a harness we don't have an adapter for * (web Claude, API SDK, another CLI). Supply the turns directly. * * Accepted formats (auto-detected): *   1) JSONL with {role, text|content, ts?} per line *   2) Markdown with **User:** / **Assistant:** headers *   3) JSON array of {role, text} objects */import { readFile } from 'node:fs/promises'import type { FindOpts, Filter, SessionRef, Turn, TraceHarness } from './types.js'import { summarize } from './types.js'function parseMarkdownTranscript(raw: string): Turn[] {  const turns: Turn[] = []  const blocks = raw.split(/\n(?=\*\*(?:User|Assistant|System)(?:\s*\([^)]+\))?:\*\*)/i)  let seq = 0  for (const block of blocks) {    const m = block.match(/^\*\*(User|Assistant|System)(?:\s*\(([^)]+)\))?:\*\*\s*([\s\S]*)$/i)    if (!m) continue    const role = m[1].toLowerCase() as Turn['role']    const body = m[3].trim()    if (!body) continue    if (role === 'user') turns.push({ role: 'user', seq, text: summarize(body, 600), ts: '' })    else if (role === 'assistant')      turns.push({ role: 'assistant', seq, text_summary: summarize(body, 280), text: body.length < 600 ? body : undefined, ts: '' })    seq++  }  return turns}function parseTurns(raw: string): Turn[] {  const trimmed = raw.trim()  if (!trimmed) return []  if (trimmed.startsWith('[')) {    try {      const arr = JSON.parse(trimmed)      return arr        .filter((t: any) => t?.role && (t.text || t.content))        .map((t: any, index: number) => ({          role: t.role,          seq: index,          text: summarize(String(t.text ?? t.content), 600),          ts: String(t.ts ?? ''),        }))    } catch {      /* fall through */    }  }  if (trimmed.includes('\n{') || trimmed.startsWith('{')) {    const out: Turn[] = []    let seq = 0    for (const line of trimmed.split('\n')) {      const s = line.trim()      if (!s) continue      try {        const ev = JSON.parse(s)        if (!ev.role) continue        if (ev.role === 'user' || ev.role === 'system') {          out.push({ role: ev.role, seq, text: summarize(String(ev.text ?? ev.content ?? ''), 600), ts: String(ev.ts ?? '') })          seq++        } else if (ev.role === 'assistant') {          out.push({            role: 'assistant',            seq,            text_summary: summarize(String(ev.text ?? ev.content ?? ''), 280),            ts: String(ev.ts ?? ''),          })          seq++        }      } catch {        /* skip malformed */      }    }    if (out.length) return out  }  return parseMarkdownTranscript(raw)
tools/harness/types.ts current file (first 80 lines)
/** * Shared types for session trace extraction across harnesses. * * Every harness (Claude Code, Codex, manual, future) produces a list of * normalized {@link Turn}s. The orchestrator stitches those turns into a trace * file that lives alongside the post it describes. *//** Detailed record of a single tool invocation inside a turn. */export type ToolCallDetail = {  name: string  /** Truncated rendering of the tool input (Bash command, file path, edit args, etc). */  input_preview?: string  /** File the tool wrote/edited, when applicable. */  file_path?: string  /** Truncated rendering of the tool result, when captured. */  result_preview?: string}/** A single turn of agent activity, lossy-summarized for repo storage. */export type Turn = {  role: 'user' | 'assistant' | 'system' | 'tool'  /** Stable sequence index in the source session (0-based, when available). */  seq?: number  /** Full text (no truncation). For very long content (>8k) the harness may still trim. */  text?: string  /** First ~280 chars of assistant prose — kept for compact list views. */  text_summary?: string  /** Rough count of tool calls issued by this turn (assistant only). */  tool_calls?: number  /** Names of tools invoked, truncated to first 6. */  tool_names?: string[]  /** Detailed per-tool records (preferred over `tool_names` going forward). */  tool_call_details?: ToolCallDetail[]  /** Files mutated (Edit/Write/MultiEdit) in this turn. */  files_touched?: string[]  /** Whether this turn contained a thinking block (for assistant). */  had_thinking?: boolean  ts: string}/** Pointer to a session file on disk; harness-specific metadata lives in meta. */export type SessionRef = {  id: string  harness: string  path: string  started_at?: string  ended_at?: string  cwd?: string  files_touched?: string[]  meta?: Record<string, unknown>}export type FindOpts = {  since?: Date  until?: Date  cwd?: string  filesTouched?: string[]  /** Prefer the most recent session whose touched files include any of these. */  limit?: number}export type Filter = {  /** Only include turns that touched these files (or surrounding turns). */  files?: string[]  /** Optional marker that must appear in a user turn to delimit the captured block. */  marker?: string  /** Include at most this many turns; head and tail preserved. */  maxTurns?: number}export interface TraceHarness {  name: string  findSessions(opts: FindOpts): Promise<SessionRef[]>  extractTurns(ref: SessionRef, filter: Filter): Promise<Turn[]>  detectModel(ref: SessionRef): Promise<string | null>}/** A per-file diff stat (additions/deletions) computed from `git show --numstat`. */export type FileDiffStat = {
src/content/posts/the-self-improving-stack.mdx +13 −9
diff --git a/src/content/posts/the-self-improving-stack.mdx b/src/content/posts/the-self-improving-stack.mdxindex c213875..23c834f 100644--- a/src/content/posts/the-self-improving-stack.mdx+++ b/src/content/posts/the-self-improving-stack.mdx@@ -15,7 +15,9 @@ authors:   - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-the-self-improving-stack-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Reviewed and dated the source trail, removed remaining temporal language, and marked the source-freshness checkpoint complete.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-review' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'Polished the umbrella article with a layer-confusion diagnostic, tightened promotion-gate phrasing, and verified role-scoped trace capture for separate draft and polish provenance.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-polish' }@@ -29,21 +31,23 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -Self-improvement is not a model property.+I can tell a coding agent to parallelize work, and it will often agree with me while still doing one thing at a time. -It is a system property.+That failure looks like a prompting problem until you inspect the trace. The sentence "fan out independent subtasks" changed the model's intention, but it did not create a worker pool, a scheduler, a merge rule, a verifier, or a budget policy. The prompt moved. The action space did not. -A model can sit inside a self-improving system, but the loop usually lives around it: prompts, skills, tools, traces, memory, evaluators, runtimes, harnesses, and release gates.--That distinction matters because a lot of AI discourse collapses very different loops into one phrase:+That is the category error hiding inside a lot of talk about self-improving agents. We say:  ```text the system optimizes itself ``` -That sentence is too vague.+as if there were one surface called "the system."++There is not.++There are prompts, skills, tools, traces, memory stores, evaluators, runtime graphs, harnesses, model weights, and release gates. Each one can be optimized. Each one needs a different kind of evidence. Each one can fail in a different way. -The useful questions are:+The useful questions are more concrete:  ```text what is allowed to change?@@ -53,9 +57,9 @@ what gate decides promotion? what can go wrong when that layer changes? ``` -Those five questions define the self-improving stack.+Those five questions are the self-improving stack. -## The Loop+## The Loop Behind The Word  A self-improving agent system has a closed loop: 
src/content/posts/self-improving-stack-harness-evolution.mdx +7 −13
diff --git a/src/content/posts/self-improving-stack-harness-evolution.mdx b/src/content/posts/self-improving-stack-harness-evolution.mdxindex 094b4a6..22768de 100644--- a/src/content/posts/self-improving-stack-harness-evolution.mdx+++ b/src/content/posts/self-improving-stack-harness-evolution.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-06   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.3-codex-spark', role: 'rewrite', date: 2026-06-06 } revisions:+  - { date: 2026-06-06, model: 'gpt-5.3-codex-spark', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-06T18-19-05-739Z-gpt-5.3-codex-spark-self-improving-stack-harness-evolution-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-harness-evolution-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-harness-evolution-review' }   - date: 2026-06-05@@ -41,21 +43,15 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -When the prompt keeps asking for a capability the runtime cannot express, stop optimizing the prompt and change the machine.+When the prompt keeps asking for a capability the runtime cannot express, the next improvement is not a better sentence. It is a different machine. -That is the harness-evolution moment.+That is the harness-evolution moment. Prompt optimizers can discover better wording, examples, instructions, rubrics, and sometimes better high-level tactics. Skill optimizers can discover reusable procedures. Runtime topology can change how many workers act, who reviews them, and what gets selected. -Prompt optimizers can discover better wording, examples, instructions, rubrics, and sometimes better high-level tactics. Skill optimizers can discover reusable procedures. Runtime topology can change how many workers act, who reviews them, and what gets selected.--Harness evolution goes one layer higher.--It changes the code that defines the agent's reachable behavior.+Harness evolution goes one layer higher: it changes the code that defines the agent's reachable behavior.  That code might be a planner contract, a driver, a verifier, a budget policy, a benchmark adapter, a trace schema, a replay layer, a selector, a persona manifest, a tool router, or a worktree candidate lifecycle. The harness is not the model. It is the machine around the model that determines which actions exist, which observations are visible, which branches can run, which artifacts count, and which candidate is allowed to become production. -So no, GEPA, SkillOpt, AlphaEvolve-style code search, and meta-harness are not all "doing the same thing" in the strong sense.--They share an outer loop:+So no, GEPA, SkillOpt, AlphaEvolve-style code search, and meta-harness are not all "doing the same thing" in the strong sense. They share an outer loop:  ```text propose candidate@@ -65,9 +61,7 @@ select survivor repeat ``` -They differ in the mutable surface.--That distinction is everything.+They differ in the mutable surface. That distinction is everything.  | Optimizer family | Mutable candidate | Reachable change | Hard limit | |---|---|---|---|
src/content/posts/self-improving-stack-trace-systems.mdx +6 −8
diff --git a/src/content/posts/self-improving-stack-trace-systems.mdx b/src/content/posts/self-improving-stack-trace-systems.mdxindex 5f2bcce..9a9e631 100644--- a/src/content/posts/self-improving-stack-trace-systems.mdx+++ b/src/content/posts/self-improving-stack-trace-systems.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-06   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-trace-systems-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-trace-systems-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-trace-systems-review' }   - date: 2026-06-05@@ -41,17 +43,13 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -Scores tell you that something happened.--Traces tell you what happened.+The optimizer wants a score. I want the run, because a score only tells you that something happened while a trace preserves enough mechanism to explain what happened.  That difference is the difference between tuning a system and optimizing an unidentified projection.  A self-improving agent can only improve from the information it preserves. If the run record says "failed, score 0.42," the optimizer can only infer weak global pressure. If the trace says the planner chose the wrong tool, the tool call used a stale argument, the retrieval span returned irrelevant context, the judge penalized a missing artifact, and the retry loop repeated the same action three times, the optimizer has a causal surface. -The trace is not decoration around the eval.--The trace is the data.+The trace is not decoration around the eval. The trace is the data.  ## The Information Loss Problem @@ -600,7 +598,7 @@ eval preserves and analyzes behavior gates decide whether behavior can ship ``` -## Failure Modes+## How Trace Systems Lie  **Score-only learning** @@ -638,7 +636,7 @@ Sensitive fields are removed without recording what was removed, destroying the  Run records, traces, scorecard cells, and analyst findings use different ids, so the evidence cannot be joined. -## Working Rule+## When A Trace Is Training Data  Do not optimize from final scores alone. 
src/content/posts/self-improving-stack-agent-runtime-topology.mdx +7 −5
diff --git a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx b/src/content/posts/self-improving-stack-agent-runtime-topology.mdxindex 1599922..960a85e 100644--- a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx+++ b/src/content/posts/self-improving-stack-agent-runtime-topology.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-05   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-agent-runtime-topology-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-agent-runtime-topology-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-agent-runtime-topology-review' }   - date: 2026-06-05@@ -41,9 +43,9 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -"Parallelize the work" is not a style preference.+When I tell a coding agent to "parallelize the work," I am not asking for a different tone. -For a human operator talking to a coding agent, it is a request for a different execution graph: spawn independent executions, cap concurrency, isolate state, collect traces, score results, select or merge outputs, cancel losers, and account for cost. If the runtime cannot express those moves, a prompt can only ask the model to simulate the shape.+I am asking for a different execution graph: spawn independent executions, cap concurrency, isolate state, collect traces, score results, select or merge outputs, cancel losers, and account for cost. If the runtime cannot express those moves, a prompt can only ask the model to simulate the shape.  This is the missing action space in many agent systems. @@ -335,7 +337,7 @@ The trace must show:  Without that trace, you cannot tell whether the topology helped, whether one branch got lucky, or whether the selector quietly ignored the evidence. -## Evaluation Protocol+## The Topology Test  A serious runtime topology eval should treat topology changes as architecture changes. @@ -366,7 +368,7 @@ promote(g_new) if:  For topology, `trace_integrity` is not optional. A candidate that wins while losing branch traces, skipping validator spans, or hiding failed children is not a better runtime. It is an unobservable runtime. -## Failure Modes+## How Topology Lies  Runtime topology fails in recognizable ways. @@ -394,7 +396,7 @@ promotion report says "better"  That is not necessarily wrong. It is incomplete. The product decision depends on whether the gain is worth the compute, latency, and operational complexity. -## A Working Rule+## When Topology Is The Right Surface  Use prompt optimization when the failure is wording. 
src/content/posts/self-improving-stack-test-time-compute.mdx +8 −8
diff --git a/src/content/posts/self-improving-stack-test-time-compute.mdx b/src/content/posts/self-improving-stack-test-time-compute.mdxindex 2ec4d75..a716b11 100644--- a/src/content/posts/self-improving-stack-test-time-compute.mdx+++ b/src/content/posts/self-improving-stack-test-time-compute.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-06   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-test-time-compute-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-test-time-compute-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-test-time-compute-review' }   - date: 2026-06-05@@ -41,11 +43,9 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -More agents is not a strategy.+I do not trust a multi-agent system until it beats the boring baseline. More agents is not a strategy. It is a cost increase until it beats blind extra compute. -It is a cost increase until it beats blind extra compute.--That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity.+That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity. It spent more budget and mislabeled the budget as architecture.  The first gate is simple: @@ -394,7 +394,7 @@ at a measured budget `B`.  That is why the previous post insisted that role names are not enough. The role structure matters only if it improves allocation, evidence, selection, or verification under budget. -## Tangle Placement+## Where The Local Stack Fits  The refreshed Tangle runtime map makes this concrete. @@ -432,7 +432,7 @@ runtime spends compute eval proves whether the spend was worth it ``` -## Evaluation Protocol+## The Equal-Compute Test  A serious test-time compute eval starts with a budget table. @@ -504,7 +504,7 @@ promote(strategy_new) if:  The baseline should be the strongest simple strategy the product could actually deploy, not a strawman. -## Failure Modes+## How More Compute Lies  Test-time compute fails in predictable ways. @@ -540,7 +540,7 @@ The token budget is matched, but one strategy uses much more sandbox time, brows  The system reports that one candidate succeeded somewhere in the batch but cannot select it reliably. -## Working Rule+## When More Compute Is Evidence  Do not evaluate an agent topology against one sample. 
src/content/posts/self-improving-stack-evaluation-gates.mdx +6 −8
diff --git a/src/content/posts/self-improving-stack-evaluation-gates.mdx b/src/content/posts/self-improving-stack-evaluation-gates.mdxindex da8d0d0..52f7a8b 100644--- a/src/content/posts/self-improving-stack-evaluation-gates.mdx+++ b/src/content/posts/self-improving-stack-evaluation-gates.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-06   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-evaluation-gates-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-evaluation-gates-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-evaluation-gates-review' }   - date: 2026-06-05@@ -41,15 +43,11 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -An optimizer can propose forever.+A green score is not a release decision. It is evidence entering a release policy, and in self-improving loops that distinction matters more than the optimizer. -The gate decides what becomes the system.+GEPA, MIPRO, SkillOpt, topology search, and meta-harness can all generate candidates forever. The gate decides which candidate becomes the system future agents inherit. If the gate is weak, the optimizer learns the gate. If the gate is honest, the optimizer has to improve the product. -That is why the gate is not an administrative detail after the interesting work. It is the objective boundary. GEPA, MIPRO, SkillOpt, runtime topology search, and meta-harness can all generate candidates. The gate decides which candidate is allowed to replace the baseline.--If the gate is weak, every optimizer learns the gate.--If the gate is honest, every optimizer has to improve the product.+This is why the gate is not an administrative detail after the interesting work. It is the objective boundary.  ## What A Gate Is @@ -574,7 +572,7 @@ The final output is scored, but the branch, tool, verifier, and selector evidenc  All failures collapse into one scalar, so the next optimizer has no diagnostic direction. -## Working Rule+## When A Candidate Can Ship  The optimizer proposes. 
research/self-improving-agent-systems/voice-blend-eval.md +78 −0
diff --git a/research/self-improving-agent-systems/voice-blend-eval.md b/research/self-improving-agent-systems/voice-blend-eval.mdnew file mode 100644index 0000000..6930fb5--- /dev/null+++ b/research/self-improving-agent-systems/voice-blend-eval.md@@ -0,0 +1,78 @@+# Voice Blend Evaluation++Date: 2026-06-08++Question: should the self-improving stack rewrite target an 80/20 blend or a 60/40 blend?++The blend is:++```text+primary = Drew voice+secondary = high-rigor essay register+```++## Original Opening++```text+Self-improvement is not a model property.++It is a system property.++A model can sit inside a self-improving system, but the loop usually lives around it: prompts, skills, tools, traces, memory, evaluators, runtimes, harnesses, and release gates.+```++Diagnosis: clean, correct, memorable, but too aphoristic. It starts with the conclusion instead of the pressure that forced the conclusion.++## 80/20 Candidate++```text+I keep coming back to the same confusion when I use coding agents: the model is only one part of the thing I am optimizing.++I can change the prompt, add a skill, raise the turn budget, fan out workers, add a reviewer, change the memory policy, swap the evaluator, or rewrite the harness. All of those feel like "making the agent better," but they are not the same intervention. They change different parts of the system, and they require different evidence before I should trust the result.++That is the real subject of self-improvement. Not a model improving itself in isolation, but a loop around a model deciding what changed, whether it helped, and whether the change is allowed to persist.+```++Strengths:++- Better grounded in Drew's work.+- Keeps the post accessible.+- Removes some generic aphorism.++Weaknesses:++- Still a little soft.+- Does not create enough adversarial pressure.+- Reads like a friendlier version of the existing post, not a level change.++## 60/40 Candidate++```text+I can tell a coding agent to parallelize work, and it will often agree with me while still doing one thing at a time.++That failure looks like a prompting problem until you inspect the trace. The sentence "fan out independent subtasks" changed the model's intention, but it did not create a worker pool, a scheduler, a merge rule, a verifier, or a budget policy. The prompt moved. The action space did not.++That is the category error hiding inside a lot of talk about self-improving agents. We say "the system optimized itself" as if there were one surface called the system. In practice there are many mutable surfaces: prompts, skills, runtime topology, traces, memory, evaluators, code, model weights, and release gates. Each has its own search operator, failure mode, and standard of evidence.++So the useful question is not whether an agent can improve itself. The useful question is: which part was allowed to change, what proved that the change helped, and who kept the optimizer away from the gate that promoted it?+```++Strengths:++- Starts from a concrete agent-work failure.+- Makes the category error visible before naming the taxonomy.+- Adds falsification pressure: trace inspection tells us whether the action space changed.+- Better fit for the self-improving stack series because it needs to argue against overbroad prompt-optimization claims.++Weaknesses:++- More forceful and less purely Drew-raw.+- Needs care to avoid sounding borrowed or over-styled.++## Decision++Use 60/40 for the self-improving stack rewrite.++Rationale: the current series is already too close to neutral technical exposition. An 80/20 rewrite would improve grounding, but it would not sufficiently raise the language, stakes, or epistemic pressure. The 60/40 blend gives each post a stronger anomaly, a live adversary, and a falsification test while keeping the first-person agent-work motive as the anchor.++Guardrail: primary Drew voice still wins. If a sentence sounds like a generic high-rigor essayist rather than Drew thinking through agent infrastructure, rewrite it.