Prompt Optimization Is Not The Whole Game
let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
- Created
- Updated
32
Turns
23
Tool calls
12
Files touched
4156m
Duration
Files
/Users/drew/code/blogsrc/content.config.tstools/new-post.mjstools/README.mdtools/trace-capture.tstraces/the-long-horizon/2026-05-25T11-23-22-617Z-gpt-5.5-series-outline.jssrc/content/posts/superintelligence-in-the-wild.mdxtools/harness/manual.ts.evolve/skill-runs.jssrc/content/posts/the-self-improving-stack.mdx.evolve/skill-runs.jsonlresearch/self-improving-agent-systems/voice-blend-eval.md
Commit
fb31e1c chore(workspace): cleanup merged in-progress content and trace artifacts +174 −9 across 8 files
Conversation
32 turns. Full text where captured; older traces show only the first ~280 chars.
- user #1let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our knowledge and you should create a blog post for each and we make this a series
- assistant #2Problem: the research is currently a single trace plus a chat summary, so it is hard to grow into a series without losing the map. Change: I’ll add a durable research checkpoint directory with one file per map item, then scaffold one draft post per item under a consistent series. Why long-term right: research notes stay source-of-truth and appendable, while posts stay draft publication surfaces with provenance. Cost: new markdown/MDX files only; risk is schema/provenance mismatch, rollback is deleting the new directory/posts.
- assistant #3I’m checking the post schema and scaffold tooling before writing files, so the new drafts land in the repo’s existing provenance format instead of inventing one.
- assistant #4 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"git status --short --branch","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
-
- assistant #5 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"git log --oneline -10","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
-
- assistant #6 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"git reflog | head -20","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
-
- assistant #7 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"gh pr list --state open","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":2000}
-
- assistant #8 1 tool
- exec_command
src/content.config.tsinput
{"cmd":"sed -n '1,260p' src/content.config.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":6000}
-
- assistant #9 1 tool
- exec_command
tools/new-post.mjsinput
{"cmd":"sed -n '1,260p' tools/new-post.mjs","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":7000}
-
- assistant #10The schema supports `series`, `outline_trace_id`, `supporting_trace_ids`, and `human_takeover`, which is enough for draft series scaffolds. I’m going to create a `research/self-improving-agent-systems/` checkpoint directory plus 12 draft MDX outlines that all point back to the existing supporting trace.
- assistant #11 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"find . -maxdepth 2 -type d | sort | sed -n '1,160p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":4000}
-
- assistant #12 1 tool
- exec_command
tools/README.mdinput
{"cmd":"rg -n \"series:|self-improving|long-horizon|supporting_trace_ids|outline_trace_id\" src/content/posts traces README.md tools/README.md --glob '!dist'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #13 1 tool
- exec_command
tools/trace-capture.tsinput
{"cmd":"sed -n '1,320p' tools/trace-capture.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #14 1 tool
- exec_command
traces/the-long-horizon/2026-05-25T11-23-22-617Z-gpt-5.5-series-outline.jsinput
{"cmd":"sed -n '1,140p' traces/the-long-horizon/2026-05-25T11-23-22-617Z-gpt-5.5-series-outline.json","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":8000}
-
- assistant #15 1 tool
- exec_command
src/content/posts/superintelligence-in-the-wild.mdxinput
{"cmd":"sed -n '1,80p' src/content/posts/superintelligence-in-the-wild.mdx","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":4000}
-
- assistant #16 1 tool
- exec_command
tools/harness/manual.tsinput
{"cmd":"sed -n '1,220p' tools/harness/manual.ts","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":6000}
-
- assistant #17 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"pnpm build","workdir":"/Users/drew/code/blog","yield_time_ms":30000,"max_output_tokens":30000}
-
- assistant #18`pnpm build` passes. There is one existing Vite warning about an unused `stat` import in `tools/og-render.ts`; it does not block the build and is unrelated to this rewrite.
- assistant #19 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"ls -la .evolve","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #20 1 tool
- exec_command
.evolve/skill-runs.jsinput
{"cmd":"test -f .evolve/skill-runs.jsonl && tail -20 .evolve/skill-runs.jsonl || true","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #21 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"git diff --stat","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":20000}
-
- assistant #22 1 tool
- exec_command
src/content/posts/the-self-improving-stack.mdxinput
{"cmd":"git diff -- src/content/posts/the-self-improving-stack.mdx | sed -n '1,220p'","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":20000}
-
- assistant #23The polish skill keeps a local run ledger, so I’m adding one line for this voice rewrite pass. That keeps the repo’s internal process state consistent with the blog trace lifecycle.
- assistant #24 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"date -u +%Y-%m-%dT%H:%M:%SZ","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #25 1 tool
- apply_patch
/Users/drew/code/blog/.evolve/skill-runs.jsonlinput
*** Begin Patch *** Update File: /Users/drew/code/blog/.evolve/skill-runs.jsonl @@ {"skill":"/governor","ts":"2026-06-06T17:50:44Z","project":"blog","target":"self-improving-stack-series","operatorPrompt":"ok whats next governor","durationMin":4,"verdict":"STOP","dispatchedTo":"/human-takeover","operatorOverride":null,"transcriptPath":null,"traceDir":null} +{"skill":"/polish","ts":"2026-06-08T09:34:12Z","project":"blog","target":"self-improving-stack-series","operatorPrompt":"ok now let's do 60/40 vs 80/20 and you compare and evaluate wdyt? you should honestly rewrite it all","durationMin":65…
-
- assistant #26One small process fix landed too: `.evolve/skill-runs.jsonl` now has a polish entry for this 60/40 rewrite run. I’m doing the final worktree and ledger checks now.
- assistant #27 1 tool
- exec_command
.evolve/skill-runs.jsinput
{"cmd":"tail -1 .evolve/skill-runs.jsonl | jq .","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #28 1 tool
- exec_command
/Users/drew/code/bloginput
{"cmd":"git status --short","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":30000}
-
- assistant #29 1 tool
- exec_command
.evolve/skill-runs.jsinput
{"cmd":"git diff --stat -- .evolve/skill-runs.jsonl AGENTS.md CLAUDE.md src/content/posts/the-self-improving-stack.mdx src/content/posts/self-improving-stack-*.mdx","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":20000}
-
- assistant #30 1 tool
- exec_command
src/content/posts/the-self-improving-stack.mdxinput
{"cmd":"rg -n \"—|first post|First post|reader hook|the article should|keep this compact|Drew angle|target audience|outline notes|The serious version of the question is|That sentence is too vague|^## (Working Rule|A Working Rule|Evaluation Protocol|The Tangle Placement|Tangle Placement|The Object Being Optimized|The Formal Shape)$\" src/content/posts/the-self-improving-stack.mdx src/content/posts/self-improving-stack-*.mdx VOICE.md SECONDARY_VOICE.md research/self-improving-agent-systems/voice-blend-eval.md","workdir":"/Users/drew/code/blog","yield_time_ms":1000,"max_output_tokens":12000}
-
- assistant #31Done. I chose `60/40` over `80/20`. 80/20 made the posts warmer, but not structurally better enough. The defect was not just “sounds too AI.” It was that the series often opened from clean abstraction instead of a live systems failure. 60/40 forced the rewrite to start from concrete agent-work pressure, then climb into formalism, gates, traces, and falsification. I rewrote the voice layer and series pass across all 13 self-improving stack posts: - Added [VOICE.md](/Users/drew/code/blog/VOICE.md) and [SECONDARY_VOICE.md](/Users/drew/code/blog/SECONDARY_VOICE.md) - Added the comparison note at [voice-blend-eval.md](/Users/drew/code/blog/research/self-improving-agent-systems/voice-blend-eval.md) - Reworked openings and scaffold headings across `the-self-improving-stack` plus the 12 series posts - Recorded rewrite revisions and created 13 rewrite trace JSON files - Updated `.evolve/skill-runs.jsonl` - Updated `AGENTS.md` and `CLAUDE.md` so future agents read `VOICE.md` Verification: - `pnpm build` passes - Banned prose scan is clean for article bodies - No em dashes in the rewritten post prose - All 13 target posts now have the `60/40 voice rewrite` revision note - Build has one unrelated existing warning: unused `stat` import in `tools/og-render.ts` Technical explanation: I treated “voice” like an optimization objective with a regularizer. Drew voice is the objective anchor: motive, lived agent workflow, speculative reach. Secondary voice is the regularizer: denser reasoning, falsification pressure, less smooth survey prose. 80/20 was a small prompt edit. 60/40 changed the representation: each post now asks what mutable surface changed, what evidence promotes it, and what would prove the improvement was fake. FYI these edits landed on PR branch `feat/ui-soft-night-mode-home-intro`, which is currently PR #23 about the UI/theme work.
- user #32ok review it again?
Diff
Per-file changes from fb31e1c.
diff --git a/src/content.config.ts b/src/content.config.tsindex 9b45af5..b460bec 100644--- a/src/content.config.ts+++ b/src/content.config.ts@@ -76,6 +76,7 @@ const toolCallDetailSchema = z.object({ }) const turnSchema = z.object({+ seq: z.number().optional(), role: z.enum(['user', 'assistant', 'system', 'tool']), text: z.string().optional(), text_summary: z.string().optional(),#!/usr/bin/env node/** * new-post — scaffold a new blog post and open it in your editor. * * Usage: * pnpm new "Your post title" # default: original (human-authored) * pnpm new "Your post title" --ai # AI-authored (no original flag) * pnpm new "Your post title" --slug=foo # override slug * pnpm new "Your post title" --no-open # don't launch editor * pnpm new "Your post title" --tags=design,prose * * Editor selection: $BLOG_EDITOR > $EDITOR > 'cursor'. */import { existsSync } from 'node:fs'import { writeFile } from 'node:fs/promises'import { spawn } from 'node:child_process'import { join, resolve } from 'node:path'import { fileURLToPath } from 'node:url'function parse(argv) { const args = { title: '', open: true, ai: false, slug: null, tags: null } const pos = [] for (const a of argv) { if (a === '--ai') args.ai = true else if (a === '--no-open') args.open = false else if (a.startsWith('--slug=')) args.slug = a.slice('--slug='.length) else if (a.startsWith('--tags=')) args.tags = a.slice('--tags='.length).split(',').map((t) => t.trim()).filter(Boolean) else pos.push(a) } args.title = pos.join(' ').trim() return args}function slugify(s) { return s .toLowerCase() .replace(/['']/g, '') .replace(/[^a-z0-9]+/g, '-') .replace(/^-+|-+$/g, '')}function today() { const d = new Date() return `${d.getFullYear()}-${String(d.getMonth() + 1).padStart(2, '0')}-${String(d.getDate()).padStart(2, '0')}`}function yamlList(items) { return '[' + items.map((t) => `'${t.replace(/'/g, "\\'")}'`).join(', ') + ']'}function frontmatter({ title, date, tags, original }) { const lines = [ '---', `title: '${title.replace(/'/g, "\\'")}'`, `description: ''`, `date: ${date}`, `tags: ${yamlList(tags)}`, ] if (original) lines.push('original: true') lines.push('draft: true', '---', '') return lines.join('\n')}function body(original) { if (original) { return [ '{/* AI AGENTS: DO NOT EDIT. This post is human-authored. See CLAUDE.md hard rule. */}', '', 'Open with the thing.', '', 'Then the next thing.', '', ].join('\n') } return ['Open with the thing.', '', 'Then the next thing.', ''].join('\n')}async function main() { const args = parse(process.argv.slice(2)) if (!args.title) {# tools/Scripts that capture, shape, and evaluate the blog's agentic data.## `new-post.mjs` / `edit-post.mjs` — draft and human edit helpers```bash# Create a human-authored draft and open it.pnpm new "Post title" --tags=agents,systems# Create an AI-assisted draft.pnpm new "Post title" --ai --tags=agents,systems# Open an existing post by slug or title substring.pnpm write long-running-task-systems# After editing, commit and let the post-commit hook record the green human revision.pnpm write long-running-task-systems --commit --note="rewrote the outline into a first human draft"# Mark an AI-outline handoff as complete and publish.pnpm write long-running-task-systems --done --publish --commit --note="publish human rewrite"```## `blog-loop.mjs` — traced AI lifecycleUse this when starting a clean AI thread.```bash# Print the exact prompt to paste into a clean research thread.pnpm blog research long-running-task-systems --harness=codex# In that thread, the agent does not edit the post. At the end it runs:pnpm blog finish long-running-task-systems --research --harness=codex --note="surveyed long-horizon benchmarks"# Print the exact prompt to paste into a thread that may write/edit the post.pnpm blog write long-running-task-systems --harness=codex --role=draft# The generated prompt includes a trace marker, voice checklist, anti-pattern gates,# and the exact finish command. Keep the marker in the first assistant update and# in the final response so trace capture can isolate the current phase.# In that thread, the agent may edit the post and then records an authorship trace:pnpm blog finish long-running-task-systems --write --harness=codex --role=draft --marker="BLOGTRACE-..." --note="drafted benchmark section"# For a final publish phase in the same session:pnpm blog finish long-running-task-systems --write --harness=codex --role=publish --marker="[BLOG_TRACE_MARKER:publish]" --note="published by toggling draft=false"```Research traces go into `supporting_trace_ids`. They are rendered as "Supporting research" and do not imply authorship. Writing traces go into `revisions[]` and do imply AI authorship/editing. Unmarked write finishes are refused by `pnpm blog finish`; use `--session=<id>` or `--allow-unmarked` only for audited recovery captures.If a thread started before you decided the target post, tell the agent:```textThis thread is supporting research for <post-slug>. Do not mark it as authorship. Attach this session as supporting research using the blog lifecycle.```Then the agent should run:```bashpnpm blog finish <post-slug> --research --harness=codex --note="supporting research"```## `trace-capture.ts` — harness-agnostic session captureExtracts the agent session behind a revision and writes it to `traces/<slug>/<trace_id>.json`. Appends a revisions entry to the post's frontmatter that links back.### Commands```bash# Capture a specific post with an explicit harnesspnpm tsx tools/trace-capture.ts capture \ --harness=claude-code \ --post=convergence-as-eval-primitive \ --marker="[BLOG_TRACE_MARKER:publish]" \ --role=polish# Auto-detect from the latest commit: finds changed posts, matches sessions# via ~/.claude/projects/ or ~/.codex/sessions/, writes traces + appends# frontmatter entries.pnpm tsx tools/trace-capture.ts capture --auto#!/usr/bin/env node/** * trace-capture: harness-agnostic session capture for blog revisions. * * Usage: * pnpm tsx tools/trace-capture.ts capture \ * [--harness=claude-code|codex|manual] \ * [--post=<slug>] \ * [--role=outline|draft|rewrite|polish|diagram|review|publish|research] \ * [--session=<session-id>] \ * [--marker="<token>"] \ * [--note=<one-line>] \ * [--commit=<sha>] \ * [--input=<path>] # manual harness only * [--kind=post|series-outline|supporting-research] * [--attach=supporting|revision|none] * [--latest] # choose latest session without requiring post file touch * * pnpm tsx tools/trace-capture.ts capture --auto * # detects from the latest git commit: finds changed posts, matches a * # recent session via ~/.claude/projects or ~/.codex/sessions, writes a * # trace per changed post, appends to frontmatter. * * pnpm tsx tools/trace-capture.ts list * # list existing traces grouped by post. * * pnpm tsx tools/trace-capture.ts show <trace_id> * # dump a trace as JSON. */import { execSync } from 'node:child_process'import { mkdir, readdir, readFile, writeFile } from 'node:fs/promises'import { join } from 'node:path'import ClaudeCodeHarness from './harness/claude-code.js'import CodexHarness from './harness/codex.js'import ManualHarness from './harness/manual.js'import { dedupeAdjacentTurns, type TraceFile, type TraceHarness, type Turn } from './harness/types.js'const ROOT = process.cwd()const POSTS_DIR = join(ROOT, 'src/content/posts')const TRACES_DIR = join(ROOT, 'traces')type Args = Record<string, string | boolean>function parseArgs(argv: string[]): { cmd: string; pos: string[]; flags: Args } { const [cmd, ...rest] = argv const flags: Args = {} const pos: string[] = [] for (const a of rest) { if (a.startsWith('--')) { const eq = a.indexOf('=') if (eq >= 0) flags[a.slice(2, eq)] = a.slice(eq + 1) else flags[a.slice(2)] = true } else pos.push(a) } return { cmd: cmd ?? 'capture', pos, flags }}function git(cmd: string): string { try { return execSync(`git ${cmd}`, { cwd: ROOT, stdio: ['ignore', 'pipe', 'ignore'] }).toString().trim() } catch { return '' }}function headCommit(): string | null { const sha = git('rev-parse HEAD') return sha || null}function changedPostsAtHead(): string[] { const out = git('show --no-renames --name-only --format="" HEAD') return out .split('\n') .map((l) => l.trim()) .filter((l) => l.startsWith('src/content/posts/') && l.endsWith('.mdx')) .map((l) => l.replace('src/content/posts/', '').replace(/\.mdx$/, ''))}diff --git a/src/content/posts/superintelligence-in-the-wild.mdx b/src/content/posts/superintelligence-in-the-wild.mdxnew file mode 100644index 0000000..e879076--- /dev/null+++ b/src/content/posts/superintelligence-in-the-wild.mdx@@ -0,0 +1,82 @@+---+title: 'If Superintelligence Arrives Quietly'+description: 'Outline notes for a post on superintelligence as operating cadence, private real-world loops, and what public evidence can and cannot show.'+date: 2026-05-25+tags: ['ai', 'systems', 'superintelligence']+draft: true+series: 'the-long-horizon'+outline_trace_id: '2026-05-25T11-23-22-617Z-gpt-5.5-series-outline'+human_takeover: 'pending'+authors:+ - model: 'gpt-5.5'+ role: 'outline'+ date: 2026-05-25+revisions:+ - date: 2026-05-25+ model: 'gpt-5.5'+ role: 'outline'+ note: 'AI-generated series outline from a traced planning session; awaiting human rewrite.'+ trace_id: '2026-05-25T11-23-22-617Z-gpt-5.5-series-outline'+---++import OutlineHandoff from '../../components/OutlineHandoff.astro'++<OutlineHandoff traceId="2026-05-25T11-23-22-617Z-gpt-5.5-series-outline" series="The Long Horizon" status="pending">+ This is an AI-edited outline extracted from a traced planning session. Drew takes over below.+</OutlineHandoff>++## Working Thesis++Superintelligence probably would not first look like a chatbot declaring itself. It would look like closed-loop systems that compress research, engineering, evaluation, and deployment cycles faster than institutions can observe.++The connective series thesis: superintelligence, if it arrives, may look first like a closed-loop institution that learns faster than humans can audit.++## Outline Notes++### Define The Terms Carefully++- OpenAI's AGI definition: highly autonomous systems outperforming humans at most economically valuable work.+- Bostrom-style superintelligence: greatly exceeding humans across virtually all domains of interest.+- SSI's public position: one goal, one product, safe superintelligence.++### Is Superintelligence Around Us Now?++- In the strong definition: no public evidence.+- In narrow pockets: yes, we have superhuman systems in coding subproblems, protein/design/search/math fragments, retrieval, and optimization.+- In organizational form: maybe the closest thing today is human+AI+eval+tooling loops compounding faster than competitors.++### The Ilya / SSI Question++- SSI publicly says it has no product cycle distraction and is focused on safe superintelligence.+- There is no public evidence that SSI has deployed "SSI" into live runs.+- The responsible framing: "If SSI believes real-world interaction matters, what kind of non-public real-world loop would be consistent with its mission?"++### What "AI In The Wild" Could Mean Without A Public Product++- Internal research agents running experiments.+- Closed sandboxes with real toolchains.+- Synthetic companies / simulated labs / long-horizon environments.+- Algorithm discovery loops.+- Agent teams doing literature review, proof search, code optimization, red-teaming.+- Private deployment to trusted researchers, not consumers.++### The Real Tell++- Not benchmark score.+- Sustained autonomous research throughput.+- Novel validated discoveries.+- Ability to improve its own evals/tools safely.+- Reliable transfer from sandbox to messy reality.++### Drew Angle To Rewrite Around++"Superintelligence may first appear as an operating cadence, not a product."++## Source Trail From The Trace++- SSI official: https://ssi.inc/+- Axios on SSI funding / no product plan: https://www.axios.com/2024/09/05/ilya-sutskevers-ai-startup-raise+- OpenAI Charter AGI definition: https://openai.com/charter/+- Bostrom definition: https://nickbostrom.com/superintelligence+- Dwarkesh/Ilya episode summary: https://www.tapesearch.com/episode/dwarkesh-and-ilya-sutskever-on-what-comes-after-scaling/9H4vn7L2gPKLaenJWdWnM4+- Anthropic RSP / frontier safety framing: https://www.anthropic.com/news/responsible-scaling-policy-v3/** * Manual harness — read a JSONL or plain-text transcript from a file or stdin. * * Useful when a revision comes from a harness we don't have an adapter for * (web Claude, API SDK, another CLI). Supply the turns directly. * * Accepted formats (auto-detected): * 1) JSONL with {role, text|content, ts?} per line * 2) Markdown with **User:** / **Assistant:** headers * 3) JSON array of {role, text} objects */import { readFile } from 'node:fs/promises'import type { FindOpts, Filter, SessionRef, Turn, TraceHarness } from './types.js'import { summarize } from './types.js'function parseMarkdownTranscript(raw: string): Turn[] { const turns: Turn[] = [] const blocks = raw.split(/\n(?=\*\*(?:User|Assistant|System)(?:\s*\([^)]+\))?:\*\*)/i) let seq = 0 for (const block of blocks) { const m = block.match(/^\*\*(User|Assistant|System)(?:\s*\(([^)]+)\))?:\*\*\s*([\s\S]*)$/i) if (!m) continue const role = m[1].toLowerCase() as Turn['role'] const body = m[3].trim() if (!body) continue if (role === 'user') turns.push({ role: 'user', seq, text: summarize(body, 600), ts: '' }) else if (role === 'assistant') turns.push({ role: 'assistant', seq, text_summary: summarize(body, 280), text: body.length < 600 ? body : undefined, ts: '' }) seq++ } return turns}function parseTurns(raw: string): Turn[] { const trimmed = raw.trim() if (!trimmed) return [] if (trimmed.startsWith('[')) { try { const arr = JSON.parse(trimmed) return arr .filter((t: any) => t?.role && (t.text || t.content)) .map((t: any, index: number) => ({ role: t.role, seq: index, text: summarize(String(t.text ?? t.content), 600), ts: String(t.ts ?? ''), })) } catch { /* fall through */ } } if (trimmed.includes('\n{') || trimmed.startsWith('{')) { const out: Turn[] = [] let seq = 0 for (const line of trimmed.split('\n')) { const s = line.trim() if (!s) continue try { const ev = JSON.parse(s) if (!ev.role) continue if (ev.role === 'user' || ev.role === 'system') { out.push({ role: ev.role, seq, text: summarize(String(ev.text ?? ev.content ?? ''), 600), ts: String(ev.ts ?? '') }) seq++ } else if (ev.role === 'assistant') { out.push({ role: 'assistant', seq, text_summary: summarize(String(ev.text ?? ev.content ?? ''), 280), ts: String(ev.ts ?? ''), }) seq++ } } catch { /* skip malformed */ } } if (out.length) return out } return parseMarkdownTranscript(raw)diff --git a/src/content/posts/the-self-improving-stack.mdx b/src/content/posts/the-self-improving-stack.mdxindex c213875..23c834f 100644--- a/src/content/posts/the-self-improving-stack.mdx+++ b/src/content/posts/the-self-improving-stack.mdx@@ -15,7 +15,9 @@ authors: - { model: 'gpt-5.5', role: 'polish', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+ - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+ - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-the-self-improving-stack-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Reviewed and dated the source trail, removed remaining temporal language, and marked the source-freshness checkpoint complete.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-review' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'polish', note: 'Polished the umbrella article with a layer-confusion diagnostic, tightened promotion-gate phrasing, and verified role-scoped trace capture for separate draft and polish provenance.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-the-self-improving-stack-polish' }@@ -29,21 +31,23 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -Self-improvement is not a model property.+I can tell a coding agent to parallelize work, and it will often agree with me while still doing one thing at a time. -It is a system property.+That failure looks like a prompting problem until you inspect the trace. The sentence "fan out independent subtasks" changed the model's intention, but it did not create a worker pool, a scheduler, a merge rule, a verifier, or a budget policy. The prompt moved. The action space did not. -A model can sit inside a self-improving system, but the loop usually lives around it: prompts, skills, tools, traces, memory, evaluators, runtimes, harnesses, and release gates.--That distinction matters because a lot of AI discourse collapses very different loops into one phrase:+That is the category error hiding inside a lot of talk about self-improving agents. We say: ```text the system optimizes itself ``` -That sentence is too vague.+as if there were one surface called "the system."++There is not.++There are prompts, skills, tools, traces, memory stores, evaluators, runtime graphs, harnesses, model weights, and release gates. Each one can be optimized. Each one needs a different kind of evidence. Each one can fail in a different way. -The useful questions are:+The useful questions are more concrete: ```text what is allowed to change?@@ -53,9 +57,9 @@ what gate decides promotion? what can go wrong when that layer changes? ``` -Those five questions define the self-improving stack.+Those five questions are the self-improving stack. -## The Loop+## The Loop Behind The Word A self-improving agent system has a closed loop: diff --git a/research/self-improving-agent-systems/voice-blend-eval.md b/research/self-improving-agent-systems/voice-blend-eval.mdnew file mode 100644index 0000000..6930fb5--- /dev/null+++ b/research/self-improving-agent-systems/voice-blend-eval.md@@ -0,0 +1,78 @@+# Voice Blend Evaluation++Date: 2026-06-08++Question: should the self-improving stack rewrite target an 80/20 blend or a 60/40 blend?++The blend is:++```text+primary = Drew voice+secondary = high-rigor essay register+```++## Original Opening++```text+Self-improvement is not a model property.++It is a system property.++A model can sit inside a self-improving system, but the loop usually lives around it: prompts, skills, tools, traces, memory, evaluators, runtimes, harnesses, and release gates.+```++Diagnosis: clean, correct, memorable, but too aphoristic. It starts with the conclusion instead of the pressure that forced the conclusion.++## 80/20 Candidate++```text+I keep coming back to the same confusion when I use coding agents: the model is only one part of the thing I am optimizing.++I can change the prompt, add a skill, raise the turn budget, fan out workers, add a reviewer, change the memory policy, swap the evaluator, or rewrite the harness. All of those feel like "making the agent better," but they are not the same intervention. They change different parts of the system, and they require different evidence before I should trust the result.++That is the real subject of self-improvement. Not a model improving itself in isolation, but a loop around a model deciding what changed, whether it helped, and whether the change is allowed to persist.+```++Strengths:++- Better grounded in Drew's work.+- Keeps the post accessible.+- Removes some generic aphorism.++Weaknesses:++- Still a little soft.+- Does not create enough adversarial pressure.+- Reads like a friendlier version of the existing post, not a level change.++## 60/40 Candidate++```text+I can tell a coding agent to parallelize work, and it will often agree with me while still doing one thing at a time.++That failure looks like a prompting problem until you inspect the trace. The sentence "fan out independent subtasks" changed the model's intention, but it did not create a worker pool, a scheduler, a merge rule, a verifier, or a budget policy. The prompt moved. The action space did not.++That is the category error hiding inside a lot of talk about self-improving agents. We say "the system optimized itself" as if there were one surface called the system. In practice there are many mutable surfaces: prompts, skills, runtime topology, traces, memory, evaluators, code, model weights, and release gates. Each has its own search operator, failure mode, and standard of evidence.++So the useful question is not whether an agent can improve itself. The useful question is: which part was allowed to change, what proved that the change helped, and who kept the optimizer away from the gate that promoted it?+```++Strengths:++- Starts from a concrete agent-work failure.+- Makes the category error visible before naming the taxonomy.+- Adds falsification pressure: trace inspection tells us whether the action space changed.+- Better fit for the self-improving stack series because it needs to argue against overbroad prompt-optimization claims.++Weaknesses:++- More forceful and less purely Drew-raw.+- Needs care to avoid sounding borrowed or over-styled.++## Decision++Use 60/40 for the self-improving stack rewrite.++Rationale: the current series is already too close to neutral technical exposition. An 80/20 rewrite would improve grounding, but it would not sufficiently raise the language, stakes, or epistemic pressure. The 60/40 blend gives each post a stronger anomaly, a live adversary, and a falsification test while keeping the first-person agent-work motive as the anchor.++Guardrail: primary Drew voice still wins. If a sentence sounds like a generic high-rigor essayist rather than Drew thinking through agent infrastructure, rewrite it.