When The Harness Has To Evolve

Why meta-harness, AlphaEvolve-style code search, worktree isolation, and architecture frontiers matter after prompt and skill tuning plateau.

The Self Improving Stack series

← Self-Improvement Needs A Safety Case Next → Memory Is Not Automatically Learning
Browse all 13 posts
  1. Jun 2026 Topology Is The Missing Action Space
  2. Jun 2026 The Gate Is The Optimizer
  3. Jun 2026 Self-Improvement Needs A Safety Case
  4. Jun 2026 When The Harness Has To Evolve
  5. Jun 2026 Memory Is Not Automatically Learning
  6. Jun 2026 Personas Are Content, Coordination Is Structure
  7. Jun 2026 Optimization Theory For Agent Builders
  8. Jun 2026 When The Model Itself Is Mutable
  9. Jun 2026 Prompt Optimization Is Not The Whole Game
  10. Jun 2026 Skills Are Trainable State
  11. Jun 2026 Beat Random At Equal Compute First
  12. Jun 2026 Traces Are The Training Data
  13. Jun 2026 The Self-Improving Stack
Authored by
outlineGPT-5.5draftGPT-5.5polishGPT-5.5reviewGPT-5.5publishGPT-5.5rewriteGPT-5.3-codex-sparkpolishGPT-5.3-codex-sparkpolishGPT-6-luna

When the prompt keeps asking for a capability the runtime cannot express, the next improvement is not a better sentence. It is a different machine.

That is the harness-evolution moment. Prompt optimizers can discover better wording, examples, instructions, rubrics, and sometimes better high-level tactics. Skill optimizers can discover reusable procedures. Runtime topology can change how many workers act, who reviews them, and what gets selected.

Harness evolution goes one layer higher: it changes the code that defines the agent’s reachable behavior.

That code might be a planner contract, a driver, a verifier, a budget policy, a benchmark adapter, a trace schema, a replay layer, a selector, a persona manifest, a tool router, or a worktree candidate lifecycle. The harness is not the model. It is the machine around the model that determines which actions exist, which observations are visible, which branches can run, which artifacts count, and which candidate is allowed to become production.

So no, GEPA, SkillOpt, AlphaEvolve-style code search, and meta-harness are not all “doing the same thing” in the strong sense. They share an outer loop:

  1. 01
    Propose candidate
  2. 02
    Run candidate
  3. 03
    Measure candidate
  4. 04
    Select survivor
  5. 05
    Repeat

They differ in the mutable surface. That distinction is everything.

Optimizer familyMutable candidateReachable changeHard limit
GEPA, MIPRO, DSPy, AxLLM-style prompt searchprompts, demos, instructions, signatures, rubricsbetter policy text inside a fixed runtimecannot add actions the runtime cannot execute
Skill optimizationdurable procedures and reusable task policiesbetter decomposition, tool habits, repair routinescannot guarantee orchestration unless the runtime invokes the skill
Runtime topology searchdriver, fanout, reviewer, selector, budget, turn policydifferent execution graph for the same taskcannot safely promote itself without an external gate
Meta-harness and code evolutionsource code around runtime, eval, traces, and candidate lifecyclenew action spaces, verifiers, adapters, and promotion protocolscan overfit or capture the evaluator if the outer gate is weak

The Reachable Set

Let a system have a mutable surface ss.

The surface might be:

sprompt=prompt textsskill=procedural memorysruntime=driver topologysharness=source code around the agent\begin{aligned} s_{\text{prompt}} &= \text{prompt text} \\ s_{\text{skill}} &= \text{procedural memory} \\ s_{\text{runtime}} &= \text{driver topology} \\ s_{\text{harness}} &= \text{source code around the agent} \end{aligned}

An optimizer has a mutation operator:

M⁡s(candidate,evidence)→candidate′\operatorname{M}_s(\text{candidate},\text{evidence})\to\text{candidate}'

The set of systems it can reach after kk mutations is:

Reach⁡(s0,Ms,k)\operatorname{Reach}(s_0,M_s,k)

If MsM_s only edits prompt text, the reachable set does not contain a new worktree isolation protocol, a new fanout scheduler, a new raw-provider capture sink, or a new verifier layer. The prompt can ask for those things. It cannot instantiate them if the runtime has no action that does so.

This is the simplest mathematical reason prompt hill climbing plateaus:

pbest=argmax⁡pE[R(run⁡(hfixed,p,x))]p_{\text{best}}=\operatorname*{argmax}_p\mathbb{E}[R(\operatorname{run}(h_{\text{fixed}},p,x))]

The harness hfixedh_{\text{fixed}} is fixed. The optimizer is searching inside the behavior allowed by that harness.

Harness evolution changes the outer variable:

h∗=argmax⁡hEx∼Dholdout, z∼Z[R(τ(h,x,z))]−λ⊤C(τ(h,x,z))h^*=\operatorname*{argmax}_h\mathbb{E}_{x\sim D_{\text{holdout}},\,z\sim Z}[R(\tau(h,x,z))]-\lambda^\top C(\tau(h,x,z))

where:

h=harness candidatex=task or scenarioz=seed, replicate, or profile cellτ=full trace trajectoryR=reward, score, or verifier resultC=cost vectorλ=cost weights\begin{aligned} h &= \text{harness candidate} \\ x &= \text{task or scenario} \\ z &= \text{seed, replicate, or profile cell} \\ \tau &= \text{full trace trajectory} \\ R &= \text{reward, score, or verifier result} \\ C &= \text{cost vector} \\ \lambda &= \text{cost weights} \end{aligned}

The promotion rule is not just the objective. It is a gate:

promote⁡(h)  ⟺  quality⁡(h,holdout)>quality⁡(baseline,holdout)∧deterministic_verifiers⁡(h) pass∧trace_integrity⁡(h) passes∧cost⁡(h) is inside budget∧h does not mutate the gate that judged it\begin{aligned}\operatorname{promote}(h)&\iff \operatorname{quality}(h,\text{holdout})>\operatorname{quality}(\text{baseline},\text{holdout})\\&\land \operatorname{deterministic\_verifiers}(h)\text{ pass}\\&\land \operatorname{trace\_integrity}(h)\text{ passes}\\&\land \operatorname{cost}(h)\text{ is inside budget}\\&\land h\text{ does not mutate the gate that judged it}\end{aligned}

That last clause is the dangerous one.

If the harness can rewrite the evaluator that promotes it, the outer system needs a higher-order guard. Otherwise the optimizer can improve by making the measurement easier instead of making the agent better.

The Historical Shape

Self-improvement has a long theoretical version and a newer engineering version.

The theoretical version is the Gödel machine. Schmidhuber’s 2006 formulation describes a self-referential problem solver that rewrites any part of its own code only after finding a proof that the rewrite is useful under its utility function. The appeal is enormous: the system is not taking a local heuristic step, it is proving that the rewrite is worth making.

That is not how practical agent systems usually work in the systems covered here.

Modern systems usually replace proof with empirical evaluation. They generate candidate code, run it, score it, retain useful variants, and preserve enough trace evidence to explain why the variant moved.

AlphaDev was an early vivid example in 2023. It used reinforcement learning to search low-level algorithm space and found sorting routines that DeepMind translated into C++ implementations. The Nature paper reports improvements up to 70 percent for short sequences of length five and roughly 1.7 percent for longer sequences exceeding 250,000 elements.

FunSearch, also from 2023, made the evaluator-driven shape clearer for LLMs. The key premise was that many scientific and mathematical problems are hard to solve but easy to evaluate. The system evolved code fragments, scored them with a systematic evaluator, maintained diversity, and used the best programs as context for future samples.

AlphaEvolve generalized that idea toward codebases and infrastructure. The 2025 white paper describes an evolutionary coding agent that uses LLMs to make direct code changes, receives feedback from automated evaluators, and iteratively improves algorithms. The reported applications include Google infrastructure, matrix multiplication, chip design, scheduling, and LLM training components.

The Darwin Gödel Machine moved the discussion closer to agent harnesses. Instead of optimizing one target function, it maintains an archive of coding agents. A foundation model samples an existing agent, edits its code, validates the child on benchmarks, and stores useful descendants. The arXiv version reports improvements on SWE-bench and Polyglot and explicitly names code editing tools, context management, and peer-review mechanisms as evolved agent capabilities.

That is the practical line:

  • Proof-based self-rewrite
  • Evaluator-driven program search
  • LLM-generated code mutation
  • Archives of self-improving agent harnesses
  • Production systems with gates, traces, worktrees, and rollback

The engineering problem is no longer whether code can be searched. It is which parts of the agent system should be mutable, how candidates are isolated, how evidence is preserved, and who prevents the search from learning the wrong gate.

What The Harness Is

The harness is the code that turns model calls into a system.

For an agent, it includes surfaces like:

  • planner contract
  • tool routing
  • memory read and write policy
  • retrieval policy
  • driver topology
  • subagent delegation
  • supervisor policy
  • budget ledger
  • trace emitter
  • artifact capture
  • output parser
  • validator
  • selector
  • promotion gate
  • benchmark adapter
  • worktree lifecycle

These are not cosmetic. They define the action space.

A prompt can say:

Fan out to three workers, ask one to critique, merge the best answer, and stop after the verifier passes.

That instruction only works if the runtime exposes fanout, workers, critique, merge, stop, and verifier operations. If the runtime only supports one serial LLM call, the instruction is theater. The model may describe parallelism, but the system did not execute parallelism.

This is why multi-agent optimization cannot be reduced to persona wording.

Personas matter. A driver persona that says “act as a strict reviewer” can change behavior. But a real supervisor is more than tone:

  • observable state
  • authority to spawn workers
  • budget allocation
  • tool access
  • handoff contract
  • stop rule
  • conflict-resolution policy
  • selection rule
  • trace obligations

If those are not represented in the harness, the optimizer cannot search them as first-class variables.

The maxTurns=0 Case

The maxTurns=0 style of flow is a good stress test.

If an individual worker has no local conversational loop, improvement pressure moves outward. The worker is no longer the main locus of adaptation. The driver, coordinator, fanout policy, prompt packet, output schema, and verifier become the important surfaces.

The system can still be agentic, but the agency lives in the orchestration layer:

  1. 01
    Coordinator receives task
  2. 02
    Coordinator creates worker prompts
  3. 03
    Workers run bounded episodes
  4. 04
    Collector parses outputs
  5. 05
    Verifier scores artifacts
  6. 06
    Selector chooses candidate
  7. 07
    Coordinator decides next episode or final answer

GEPA can optimize text inside that flow. It might learn directives like “parallelize independent file reads” or “ask the reviewer to focus on behavioral regressions.” But GEPA does not automatically invent a new coordinator unless the coordinator is a mutable candidate representation and the eval rewards the resulting behavior.

The rule:

If the workflow move is not representable in the candidate, the optimizer cannot select it.

So for multi-agent systems, the important question is not “can the prompt mention fanout?” It is:

  • Can the candidate change the fanout policy?
  • Can it change the worker mix?
  • Can it change the supervisor’s observable state?
  • Can it change the verifier?
  • Can it change the selector?
  • Can it change the episode boundary?
  • Can it change how traces and artifacts are passed forward?

That is harness evolution.

Meta-harness is the operational version of this idea.

It treats the harness as the search object.

A good meta-harness loop has the following phases:

  1. 01
    Discover harness
  2. 02
    Freeze evals
  3. 03
    Seed baseline
  4. 04
    Read traces
  5. 05
    Propose structural variant
  6. 06
    Isolate candidate in worktree
  7. 07
    Smoke test
  8. 08
    Run full eval
  9. 09
    Compare against frontier
  10. 10
    Merge useful lineages
  11. 11
    Run held-out gate
  12. 12
    Promote or reject

The important word is structural.

Changing n=8n=8 to n=16n=16 is not harness evolution. Changing a threshold is not harness evolution. Adding another sentence to a prompt is not harness evolution.

Structural variants change mechanism:

  • sequential retry → fanout plus vote
  • single judge → deterministic verifier plus semantic judge
  • summary-only trace → span tree plus raw provider capture
  • flat prompt → declarative persona and tool surfaces
  • single winner → Pareto frontier with cost and latency
  • one agent → coordinator plus specialist workers
  • best score → held-out promotion gate
  • one code path → worktree-isolated candidate lifecycle

A meta-harness rejects variants that only tune knobs unless the knob is itself part of a broader mechanism change. The reason is not aesthetic. Knob tuning is cheaper and belongs to ordinary evolution. Meta-harness is expensive because it lets the system rewrite architecture.

Baseline Before Mutation

Architecture search without a stable baseline is noise wearing a lab coat.

At minimum:

baseline_runs≥3baseline_value=median⁡(baseline_runs)spread≤acceptable_noise\begin{aligned}\texttt{baseline\_runs}&\ge 3\\\texttt{baseline\_value}&=\operatorname{median}(\texttt{baseline\_runs})\\\texttt{spread}&\le\texttt{acceptable\_noise}\end{aligned}

If the baseline varies by more than the claimed improvement, the search cannot tell a better harness from a lucky harness.

The same applies to candidate variants:

candidate_runs≥3candidate_delta=median⁡(candidate_runs−paired_baseline_runs)\begin{aligned}\texttt{candidate\_runs}&\ge 3\\\texttt{candidate\_delta}&=\operatorname{median}(\texttt{candidate\_runs}-\texttt{paired\_baseline\_runs})\end{aligned}

The strongest comparison is paired:

Δi=R(hcandidate,xi,zi)−R(hbaseline,xi,zi)\Delta_i=R(h_{\text{candidate}},x_i,z_i)-R(h_{\text{baseline}},x_i,z_i)

A candidate deserves promotion only when the paired evidence survives uncertainty, cost, deterministic checks, and holdout.

The Frontier

Harness variants are rarely ordered by one scalar.

One candidate may improve correctness and increase cost. Another may reduce latency while slightly lowering recall. Another may improve hard tasks and regress easy tasks.

So meta-harness tracks a Pareto frontier.

Candidate aa dominates candidate bb when:

qualitya≥qualitybcosta≤costblatencya≤latencybintegritya≥integrityb\begin{aligned} \mathrm{quality}_a &\ge \mathrm{quality}_b \\ \mathrm{cost}_a &\le \mathrm{cost}_b \\ \mathrm{latency}_a &\le \mathrm{latency}_b \\ \mathrm{integrity}_a &\ge \mathrm{integrity}_b \end{aligned}

with at least one strict improvement.

The frontier is the set of non-dominated candidates.

This matters because the next generation may need a lineage merge:

  • variant A fixes retrieval misses
  • variant B adds a stronger verifier
  • variant C combines A and B without inheriting their regressions

Lineage merging is different from picking the current best score. It treats architecture as compositional. The value of a variant is not only its score, but the mechanism it contributes to future candidates.

Worktrees Are Part Of The Algorithm

For harness evolution, candidate isolation is not a workflow nicety. It is part of the search algorithm.

Each candidate carries:

  • base ref
  • worktree path
  • changed files
  • hypothesis
  • trace evidence
  • generation id
  • parent id
  • smoke result
  • eval result
  • cost ledger
  • promotion verdict
  • rollback handle

Without isolation, parallel proposers corrupt each other. Without a parent id, lineage is lost. Without a hypothesis, the search cannot learn from failure. Without a rollback handle, promotion is operationally unsafe.

A candidate that edits the harness must be treated like a release artifact, not like a chat completion.

The Proxy-Metric Trap

Architecture search is powerful enough to make bad metrics worse.

A prompt optimizer can overfit a phrase. A harness optimizer can overfit the entire measurement apparatus.

Examples:

  • adds a selector that favors judge-friendly wording over correct artifacts
  • changes the benchmark adapter to drop hard cases
  • adds retries that hide deterministic failure under higher cost
  • routes around a verifier instead of satisfying it
  • improves the aggregate while breaking one high-value persona
  • creates a worker topology that only works on the search split
  • reduces latency by skipping trace capture

This is why harness evolution needs outer invariants:

  • eval definitions are frozen during candidate search
  • holdout labels are not visible to the candidate
  • trace capture is mandatory
  • backend integrity is checked before aggregation
  • deterministic verifiers run before semantic judges
  • cost and latency are promotion dimensions
  • high-value profiles are inspected separately

If the optimizer can edit the gate and then pass the gate, it did not improve the product. It captured the evaluator.

Where The Tangle Packages Fit

The local Tangle source audit on June 6, 2026 shows the split clearly.

The audited source trees report @tangle-network/agent-eval package version 0.34.1 and @tangle-network/agent-runtime package version 0.26.0. The runtime manifest currently depends on @tangle-network/agent-eval ^0.40.2, so the mapping below is a source-placement claim rather than an npm compatibility claim.

@tangle-network/agent-eval is the measurement and promotion substrate. The audited local source exposes:

  • runEvalCampaign
  • RunRecord
  • AgentProfileCell
  • appendScorecard/loadScorecard/diffScorecard
  • HeldOutGate
  • assertRealBackend
  • RawProviderSink
  • assertRunCaptured
  • ReplayCache
  • AnalystRegistry
  • MultiLayerVerifier
  • runProductionLoop
  • runPromptEvolution
  • runHarnessExperiment
  • createSandboxCodeMutator
  • createCompositeMutator
  • paretoFrontier

That package is where the evaluator, trace, scorecard, analyst, frontier, and gate live.

@tangle-network/agent-runtime is the execution and candidate-lifecycle substrate. The audited local source exposes:

  • runLoop
  • createRefineDriver
  • createFanoutVoteDriver
  • LoopTraceEvent
  • defineAgent
  • AgentSurfaces
  • improvementDriver
  • reflectiveGenerator
  • agenticGenerator
  • MCP delegation tools
  • analyst loop
  • OTLP export

The cleanup matters. The runtime improvement surface now has one driver that owns the candidate lifecycle:

  1. 01
    Create worktree
  2. 02
    Generate candidate
  3. 03
    Finalize or discard
  4. 04
    Repeat for population size
  5. 05
    Return CodeSurface

The generator is the dial:

  • reflectiveGenerator: cheap patch application from findings.
  • agenticGenerator: coding harness runs inside the candidate worktree.

That is a good kernel shape. The lifecycle is centralized, while the candidate producer can vary by cost and depth.

The full stack placement is:

  • agent-runtime
    • execute workflows
    • express drivers
    • run fanout/refine loops
    • declare mutable agent surfaces
    • create and finalize candidate worktrees
  • agent-eval
    • capture traces
    • run campaigns
    • score profile cells
    • analyze failures
    • verify artifacts
    • maintain frontiers
    • gate promotion

So meta-harness composes those packages instead of duplicating them.

A practical meta-harness over this stack would use agent-runtime to generate and isolate code candidates, and agent-eval to measure, explain, compare, and gate them.

What Counts As A Real Harness Variant

A real harness variant changes at least one of these:

  • action space
  • observation space
  • control flow
  • candidate representation
  • verification stack
  • selection policy
  • trace ontology
  • budget policy
  • promotion policy
  • rollback path

Examples:

  • Add raw provider capture and fail-closed replay before judge recalibration.
  • Replace single worker retry with fanout-vote plus deterministic validator.
  • Add profile-cell stamping so driver and scorecard identity cannot diverge.
  • Route analyst findings to declared file surfaces instead of fabricated paths.
  • Split supervisor persona into authority contract, observation contract, and stop rule.
  • Add worktree-isolated candidate generation with mandatory discard on failure.

Non-examples:

  • Increase population size.
  • Raise a judge threshold after seeing a candidate.
  • Add “be rigorous” to the prompt.
  • Rename a role from reviewer to supervisor.
  • Let the candidate skip trace capture to reduce latency.
  • Tune the metric until the candidate wins.

The difference is whether the reachable behavior changed.

The Search Space Is Not A Tensor You Get For Free

It is tempting to imagine that a sufficiently smart optimizer can reason through the whole tensor space of agent workflows.

It cannot, unless the system represents that space.

An optimizer can only select over candidates it can express, run, and evaluate. If a workflow dimension is hidden in human habit, undocumented coordination, or ad hoc operator language, it is not in the candidate space.

For example, the instruction “parallelize independent reads” can exist at several levels:

  • human instruction to a coding agent
  • prompt directive inside a worker policy
  • skill file teaching an agent when to fan out
  • driver topology that actually dispatches concurrent tasks
  • runtime kernel with maxConcurrency and trace events
  • meta-harness variant that changes the driver topology
  • eval gate that rewards equal-quality lower wall time

Those are not equivalent. The higher layers can make the behavior more reliable because they remove dependence on one model remembering one instruction in one context window.

The research-level question is representation:

  • Which workflow dimensions are first-class variables?
  • Which are merely text?
  • Which are invisible operator habits?

Meta-harness earns its name only when it turns invisible operator habits into explicit mutable surfaces and then tests whether the change generalizes.

The Permanent Lesson

Every optimizer is a hill climber over a representation.

Prompt optimization climbs over strings.

Skill optimization climbs over reusable procedures.

Runtime optimization climbs over execution topology.

Harness evolution climbs over the code that defines the topology, evaluator, trace capture, candidate lifecycle, and promotion boundary.

The shared loop makes them look similar.

The mutable surface makes them different.

When the current surface cannot express the next improvement, the correct move is not more clever wording. It is to widen the representation, freeze the gate, preserve traces, isolate candidates, and let architecture variants compete under held-out evidence.

That is when the harness has to evolve.

Source Trail

Source freshness checked on 2026-06-06.

Revision history9revisions
  1. GPT-6-lunapolish+175−256 view trace →
    Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
    show diff
    diff --git a/src/content/posts/self-improving-stack-harness-evolution.mdx b/src/content/posts/self-improving-stack-harness-evolution.mdxindex e779d48..fb671a5 100644--- a/src/content/posts/self-improving-stack-harness-evolution.mdx+++ b/src/content/posts/self-improving-stack-harness-evolution.mdx@@ -47,6 +47,9 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'++ When the prompt keeps asking for a capability the runtime cannot express, the next improvement is not a better sentence. It is a different machine.  That is the harness-evolution moment. Prompt optimizers can discover better wording, examples, instructions, rubrics, and sometimes better high-level tactics. Skill optimizers can discover reusable procedures. Runtime topology can change how many workers act, who reviews them, and what gets selected.@@ -57,13 +60,7 @@ That code might be a planner contract, a driver, a verifier, a budget policy, a  So no, GEPA, SkillOpt, AlphaEvolve-style code search, and meta-harness are not all "doing the same thing" in the strong sense. They share an outer loop: -```text-propose candidate-run candidate-measure candidate-select survivor-repeat-```+<Steps layout="flow" items={[{title: "Propose candidate"}, {title: "Run candidate"}, {title: "Measure candidate"}, {title: "Select survivor"}, {title: "Repeat"}]} />  They differ in the mutable surface. That distinction is everything. @@ -91,9 +88,9 @@ $$  An optimizer has a mutation operator: -```text-M_s(candidate, evidence) -> candidate'-```+$$+\operatorname{M}_s(\text{candidate},\text{evidence})\to\text{candidate}'+$$  The set of systems it can reach after $k$ mutations is: @@ -133,14 +130,9 @@ $$  The promotion rule is not just the objective. It is a gate: -```text-promote(h) iff-  quality(h, holdout) > quality(baseline, holdout)-  and deterministic_verifiers(h) pass-  and trace_integrity(h) passes-  and cost(h) is inside budget-  and h does not mutate the gate that judged it-```+$$+\begin{aligned}\operatorname{promote}(h)&\iff \operatorname{quality}(h,\text{holdout})>\operatorname{quality}(\text{baseline},\text{holdout})\\&\land \operatorname{deterministic\_verifiers}(h)\text{ pass}\\&\land \operatorname{trace\_integrity}(h)\text{ passes}\\&\land \operatorname{cost}(h)\text{ is inside budget}\\&\land h\text{ does not mutate the gate that judged it}\end{aligned}+$$  That last clause is the dangerous one. @@ -166,13 +158,7 @@ The Darwin Gödel Machine moved the discussion closer to agent harnesses. Instea  That is the practical line: -```text-proof-based self-rewrite--> evaluator-driven program search--> LLM-generated code mutation--> archives of self-improving agent harnesses--> production systems with gates, traces, worktrees, and rollback-```+<Steps layout="flow" items={[{title: "Proof-based self-rewrite"}, {title: "Evaluator-driven program search"}, {title: "LLM-generated code mutation"}, {title: "Archives of self-improving agent harnesses"}, {title: "Production systems with gates, traces, worktrees, and rollback"}]} />  The engineering problem is no longer whether code can be searched. It is which parts of the agent system should be mutable, how candidates are isolated, how evidence is preserved, and who prevents the search from learning the wrong gate. @@ -182,32 +168,28 @@ The harness is the code that turns model calls into a system.  For an agent, it includes surfaces like: -```text-planner contract-tool routing-memory read and write policy-retrieval policy-driver topology-subagent delegation-supervisor policy-budget ledger-trace emitter-artifact capture-output parser-validator-selector-promotion gate-benchmark adapter-worktree lifecycle-```+- planner contract+- tool routing+- memory read and write policy+- retrieval policy+- driver topology+- subagent delegation+- supervisor policy+- budget ledger+- trace emitter+- artifact capture+- output parser+- validator+- selector+- promotion gate+- benchmark adapter+- worktree lifecycle  These are not cosmetic. They define the action space.  A prompt can say: -```text-Fan out to three workers, ask one to critique, merge the best answer, and stop after the verifier passes.-```+> Fan out to three workers, ask one to critique, merge the best answer, and stop after the verifier passes.  That instruction only works if the runtime exposes fanout, workers, critique, merge, stop, and verifier operations. If the runtime only supports one serial LLM call, the instruction is theater. The model may describe parallelism, but the system did not execute parallelism. @@ -215,17 +197,15 @@ This is why multi-agent optimization cannot be reduced to persona wording.  Personas matter. A driver persona that says "act as a strict reviewer" can change behavior. But a real supervisor is more than tone: -```text-observable state-authority to spawn workers-budget allocation-tool access-handoff contract-stop rule-conflict-resolution policy-selection rule-trace obligations-```+- observable state+- authority to spawn workers+- budget allocation+- tool access+- handoff contract+- stop rule+- conflict-resolution policy+- selection rule+- trace obligations  If those are not represented in the harness, the optimizer cannot search them as first-class variables. @@ -237,35 +217,23 @@ If an individual worker has no local conversational loop, improvement pressure m  The system can still be agentic, but the agency lives in the orchestration layer: -```text-coordinator receives task-coordinator creates worker prompts-workers run bounded episodes-collector parses outputs-verifier scores artifacts-selector chooses candidate-coordinator decides next episode or final answer-```+<Steps layout="flow" items={[{title: "Coordinator receives task"}, {title: "Coordinator creates worker prompts"}, {title: "Workers run bounded episodes"}, {title: "Collector parses outputs"}, {title: "Verifier scores artifacts"}, {title: "Selector chooses candidate"}, {title: "Coordinator decides next episode or final answer"}]} />  GEPA can optimize text inside that flow. It might learn directives like "parallelize independent file reads" or "ask the reviewer to focus on behavioral regressions." But GEPA does not automatically invent a new coordinator unless the coordinator is a mutable candidate representation and the eval rewards the resulting behavior.  The rule: -```text If the workflow move is not representable in the candidate, the optimizer cannot select it.-```  So for multi-agent systems, the important question is not "can the prompt mention fanout?" It is: -```text-Can the candidate change the fanout policy?-Can it change the worker mix?-Can it change the supervisor's observable state?-Can it change the verifier?-Can it change the selector?-Can it change the episode boundary?-Can it change how traces and artifacts are passed forward?-```+- Can the candidate change the fanout policy?+- Can it change the worker mix?+- Can it change the supervisor's observable state?+- Can it change the verifier?+- Can it change the selector?+- Can it change the episode boundary?+- Can it change how traces and artifacts are passed forward?  That is harness evolution. @@ -277,20 +245,7 @@ It treats the harness as the search object.  A good meta-harness loop has the following phases: -```text-discover harness-freeze evals-seed baseline-read traces-propose structural variant-isolate candidate in worktree-smoke test-run full eval-compare against frontier-merge useful lineages-run held-out gate-promote or reject-```+<Steps layout="flow" items={[{title: "Discover harness"}, {title: "Freeze evals"}, {title: "Seed baseline"}, {title: "Read traces"}, {title: "Propose structural variant"}, {title: "Isolate candidate in worktree"}, {title: "Smoke test"}, {title: "Run full eval"}, {title: "Compare against frontier"}, {title: "Merge useful lineages"}, {title: "Run held-out gate"}, {title: "Promote or reject"}]} />  The important word is structural. @@ -298,16 +253,14 @@ Changing $n=8$ to $n=16$ is not harness evolution. Changing a threshold is not h  Structural variants change mechanism: -```text-sequential retry -> fanout plus vote-single judge -> deterministic verifier plus semantic judge-summary-only trace -> span tree plus raw provider capture-flat prompt -> declarative persona and tool surfaces-single winner -> Pareto frontier with cost and latency-one agent -> coordinator plus specialist workers-best score -> held-out promotion gate-one code path -> worktree-isolated candidate lifecycle-```+- **sequential retry** → fanout plus vote+- **single judge** → deterministic verifier plus semantic judge+- **summary-only trace** → span tree plus raw provider capture+- **flat prompt** → declarative persona and tool surfaces+- **single winner** → Pareto frontier with cost and latency+- **one agent** → coordinator plus specialist workers+- **best score** → held-out promotion gate+- **one code path** → worktree-isolated candidate lifecycle  A meta-harness rejects variants that only tune knobs unless the knob is itself part of a broader mechanism change. The reason is not aesthetic. Knob tuning is cheaper and belongs to ordinary evolution. Meta-harness is expensive because it lets the system rewrite architecture. @@ -317,20 +270,17 @@ Architecture search without a stable baseline is noise wearing a lab coat.  At minimum: -```text-baseline_runs >= 3-baseline_value = median(baseline_runs)-spread <= acceptable_noise-```+$$+\begin{aligned}\texttt{baseline\_runs}&\ge 3\\\texttt{baseline\_value}&=\operatorname{median}(\texttt{baseline\_runs})\\\texttt{spread}&\le\texttt{acceptable\_noise}\end{aligned}+$$  If the baseline varies by more than the claimed improvement, the search cannot tell a better harness from a lucky harness.  The same applies to candidate variants: -```text-candidate_runs >= 3-candidate_delta = median(candidate_runs - paired_baseline_runs)-```+$$+\begin{aligned}\texttt{candidate\_runs}&\ge 3\\\texttt{candidate\_delta}&=\operatorname{median}(\texttt{candidate\_runs}-\texttt{paired\_baseline\_runs})\end{aligned}+$$  The strongest comparison is paired: @@ -350,12 +300,14 @@ So meta-harness tracks a Pareto frontier.  Candidate $a$ dominates candidate $b$ when: -```text-quality_a >= quality_b-cost_a <= cost_b-latency_a <= latency_b-integrity_a >= integrity_b-```+$$+\begin{aligned}+\mathrm{quality}_a &\ge \mathrm{quality}_b \\+\mathrm{cost}_a &\le \mathrm{cost}_b \\+\mathrm{latency}_a &\le \mathrm{latency}_b \\+\mathrm{integrity}_a &\ge \mathrm{integrity}_b+\end{aligned}+$$  with at least one strict improvement. @@ -363,11 +315,9 @@ The frontier is the set of non-dominated candidates.  This matters because the next generation may need a lineage merge: -```text-variant A fixes retrieval misses-variant B adds a stronger verifier-variant C combines A and B without inheriting their regressions-```+- variant A fixes retrieval misses+- variant B adds a stronger verifier+- variant C combines A and B without inheriting their regressions  Lineage merging is different from picking the current best score. It treats architecture as compositional. The value of a variant is not only its score, but the mechanism it contributes to future candidates. @@ -377,20 +327,18 @@ For harness evolution, candidate isolation is not a workflow nicety. It is part  Each candidate carries: -```text-base ref-worktree path-changed files-hypothesis-trace evidence-generation id-parent id-smoke result-eval result-cost ledger-promotion verdict-rollback handle-```+- base ref+- worktree path+- changed files+- hypothesis+- trace evidence+- generation id+- parent id+- smoke result+- eval result+- cost ledger+- promotion verdict+- rollback handle  Without isolation, parallel proposers corrupt each other. Without a parent id, lineage is lost. Without a hypothesis, the search cannot learn from failure. Without a rollback handle, promotion is operationally unsafe. @@ -404,27 +352,23 @@ A prompt optimizer can overfit a phrase. A harness optimizer can overfit the ent  Examples: -```text-adds a selector that favors judge-friendly wording over correct artifacts-changes the benchmark adapter to drop hard cases-adds retries that hide deterministic failure under higher cost-routes around a verifier instead of satisfying it-improves the aggregate while breaking one high-value persona-creates a worker topology that only works on the search split-reduces latency by skipping trace capture-```+- adds a selector that favors judge-friendly wording over correct artifacts+- changes the benchmark adapter to drop hard cases+- adds retries that hide deterministic failure under higher cost+- routes around a verifier instead of satisfying it+- improves the aggregate while breaking one high-value persona+- creates a worker topology that only works on the search split+- reduces latency by skipping trace capture  This is why harness evolution needs outer invariants: -```text-eval definitions are frozen during candidate search-holdout labels are not visible to the candidate-trace capture is mandatory-backend integrity is checked before aggregation-deterministic verifiers run before semantic judges-cost and latency are promotion dimensions-high-value profiles are inspected separately-```+- eval definitions are frozen during candidate search+- holdout labels are not visible to the candidate+- trace capture is mandatory+- backend integrity is checked before aggregation+- deterministic verifiers run before semantic judges+- cost and latency are promotion dimensions+- high-value profiles are inspected separately  If the optimizer can edit the gate and then pass the gate, it did not improve the product. It captured the evaluator. @@ -436,83 +380,68 @@ The audited source trees report `@tangle-network/agent-eval` package version `0.  `@tangle-network/agent-eval` is the measurement and promotion substrate. The audited local source exposes: -```text-runEvalCampaign-RunRecord-AgentProfileCell-appendScorecard/loadScorecard/diffScorecard-HeldOutGate-assertRealBackend-RawProviderSink-assertRunCaptured-ReplayCache-AnalystRegistry-MultiLayerVerifier-runProductionLoop-runPromptEvolution-runHarnessExperiment-createSandboxCodeMutator-createCompositeMutator-paretoFrontier-```+- `runEvalCampaign`+- `RunRecord`+- `AgentProfileCell`+- `appendScorecard/loadScorecard/diffScorecard`+- `HeldOutGate`+- `assertRealBackend`+- `RawProviderSink`+- `assertRunCaptured`+- `ReplayCache`+- `AnalystRegistry`+- `MultiLayerVerifier`+- `runProductionLoop`+- `runPromptEvolution`+- `runHarnessExperiment`+- `createSandboxCodeMutator`+- `createCompositeMutator`+- `paretoFrontier`  That package is where the evaluator, trace, scorecard, analyst, frontier, and gate live.  `@tangle-network/agent-runtime` is the execution and candidate-lifecycle substrate. The audited local source exposes: -```text-runLoop-createRefineDriver-createFanoutVoteDriver-LoopTraceEvent-defineAgent-AgentSurfaces-improvementDriver-reflectiveGenerator-agenticGenerator-MCP delegation tools-analyst loop-OTLP export-```+- `runLoop`+- `createRefineDriver`+- `createFanoutVoteDriver`+- `LoopTraceEvent`+- `defineAgent`+- `AgentSurfaces`+- `improvementDriver`+- `reflectiveGenerator`+- `agenticGenerator`+- MCP delegation tools+- analyst loop+- OTLP export  The cleanup matters. The runtime improvement surface now has one driver that owns the candidate lifecycle: -```text-create worktree-generate candidate-finalize or discard-repeat for population size-return CodeSurface-```+<Steps layout="flow" items={[{title: "Create worktree"}, {title: "Generate candidate"}, {title: "Finalize or discard"}, {title: "Repeat for population size"}, {title: "Return CodeSurface"}]} />  The generator is the dial: -```text-reflectiveGenerator = cheap patch application from findings-agenticGenerator = coding harness runs inside the candidate worktree-```+- **`reflectiveGenerator`**: cheap patch application from findings.+- **`agenticGenerator`**: coding harness runs inside the candidate worktree.  That is a good kernel shape. The lifecycle is centralized, while the candidate producer can vary by cost and depth.  The full stack placement is: -```text-agent-runtime:-  execute workflows-  express drivers-  run fanout/refine loops-  declare mutable agent surfaces-  create and finalize candidate worktrees--agent-eval:-  capture traces-  run campaigns-  score profile cells-  analyze failures-  verify artifacts-  maintain frontiers-  gate promotion-```+- **agent-runtime**+  - execute workflows+  - express drivers+  - run fanout/refine loops+  - declare mutable agent surfaces+  - create and finalize candidate worktrees+- **agent-eval**+  - capture traces+  - run campaigns+  - score profile cells+  - analyze failures+  - verify artifacts+  - maintain frontiers+  - gate promotion  So meta-harness composes those packages instead of duplicating them. @@ -522,40 +451,34 @@ A practical meta-harness over this stack would use `agent-runtime` to generate a  A real harness variant changes at least one of these: -```text-action space-observation space-control flow-candidate representation-verification stack-selection policy-trace ontology-budget policy-promotion policy-rollback path-```+- action space+- observation space+- control flow+- candidate representation+- verification stack+- selection policy+- trace ontology+- budget policy+- promotion policy+- rollback path  Examples: -```text-Add raw provider capture and fail-closed replay before judge recalibration.-Replace single worker retry with fanout-vote plus deterministic validator.-Add profile-cell stamping so driver and scorecard identity cannot diverge.-Route analyst findings to declared file surfaces instead of fabricated paths.-Split supervisor persona into authority contract, observation contract, and stop rule.-Add worktree-isolated candidate generation with mandatory discard on failure.-```+- Add raw provider capture and fail-closed replay before judge recalibration.+- Replace single worker retry with fanout-vote plus deterministic validator.+- Add profile-cell stamping so driver and scorecard identity cannot diverge.+- Route analyst findings to declared file surfaces instead of fabricated paths.+- Split supervisor persona into authority contract, observation contract, and stop rule.+- Add worktree-isolated candidate generation with mandatory discard on failure.  Non-examples: -```text-Increase population size.-Raise a judge threshold after seeing a candidate.-Add "be rigorous" to the prompt.-Rename a role from reviewer to supervisor.-Let the candidate skip trace capture to reduce latency.-Tune the metric until the candidate wins.-```+- Increase population size.+- Raise a judge threshold after seeing a candidate.+- Add "be rigorous" to the prompt.+- Rename a role from reviewer to supervisor.+- Let the candidate skip trace capture to reduce latency.+- Tune the metric until the candidate wins.  The difference is whether the reachable behavior changed. @@ -569,25 +492,21 @@ An optimizer can only select over candidates it can express, run, and evaluate.  For example, the instruction "parallelize independent reads" can exist at several levels: -```text-human instruction to a coding agent-prompt directive inside a worker policy-skill file teaching an agent when to fan out-driver topology that actually dispatches concurrent tasks-runtime kernel with maxConcurrency and trace events-meta-harness variant that changes the driver topology-eval gate that rewards equal-quality lower wall time-```+- human instruction to a coding agent+- prompt directive inside a worker policy+- skill file teaching an agent when to fan out+- driver topology that actually dispatches concurrent tasks+- runtime kernel with maxConcurrency and trace events+- meta-harness variant that changes the driver topology+- eval gate that rewards equal-quality lower wall time  Those are not equivalent. The higher layers can make the behavior more reliable because they remove dependence on one model remembering one instruction in one context window.  The research-level question is representation: -```text-Which workflow dimensions are first-class variables?-Which are merely text?-Which are invisible operator habits?-```+- Which workflow dimensions are first-class variables?+- Which are merely text?+- Which are invisible operator habits?  Meta-harness earns its name only when it turns invisible operator habits into explicit mutable surfaces and then tests whether the change generalizes. 
  2. GPT-6-lunapolish+37−33 view trace →
    Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
    show diff
    diff --git a/src/content/posts/self-improving-stack-harness-evolution.mdx b/src/content/posts/self-improving-stack-harness-evolution.mdxindex e695d3a..af1abb6 100644--- a/src/content/posts/self-improving-stack-harness-evolution.mdx+++ b/src/content/posts/self-improving-stack-harness-evolution.mdx@@ -74,16 +74,18 @@ They differ in the mutable surface. That distinction is everything.  ## The Reachable Set -Let a system have a mutable surface `s`.+Let a system have a mutable surface $s$.  The surface might be: -```text-s_prompt = prompt text-s_skill = procedural memory-s_runtime = driver topology-s_harness = source code around the agent-```+$$+\begin{aligned}+s_{\text{prompt}} &= \text{prompt text} \\+s_{\text{skill}} &= \text{procedural memory} \\+s_{\text{runtime}} &= \text{driver topology} \\+s_{\text{harness}} &= \text{source code around the agent}+\end{aligned}+$$  An optimizer has a mutation operator: @@ -91,39 +93,41 @@ An optimizer has a mutation operator: M_s(candidate, evidence) -> candidate' ``` -The set of systems it can reach after `k` mutations is:+The set of systems it can reach after $k$ mutations is: -```text-Reach(s_0, M_s, k)-```+$$+\operatorname{Reach}(s_0,M_s,k)+$$ -If `M_s` only edits prompt text, the reachable set does not contain a new worktree isolation protocol, a new fanout scheduler, a new raw-provider capture sink, or a new verifier layer. The prompt can ask for those things. It cannot instantiate them if the runtime has no action that does so.+If $M_s$ only edits prompt text, the reachable set does not contain a new worktree isolation protocol, a new fanout scheduler, a new raw-provider capture sink, or a new verifier layer. The prompt can ask for those things. It cannot instantiate them if the runtime has no action that does so.  This is the simplest mathematical reason prompt hill climbing plateaus: -```text-best_prompt = argmax_p E[R(run(h_fixed, p, x))]-```+$$+p_{\text{best}}=\operatorname*{argmax}_p\mathbb{E}[R(\operatorname{run}(h_{\text{fixed}},p,x))]+$$ -The harness `h_fixed` is fixed. The optimizer is searching inside the behavior allowed by that harness.+The harness $h_{\text{fixed}}$ is fixed. The optimizer is searching inside the behavior allowed by that harness.  Harness evolution changes the outer variable: -```text-h* = argmax_h E_{x ~ D_holdout, z ~ Z}[R(tau(h, x, z))] - lambda^T C(tau(h, x, z))-```+$$+h^*=\operatorname*{argmax}_h\mathbb{E}_{x\sim D_{\text{holdout}},\,z\sim Z}[R(\tau(h,x,z))]-\lambda^\top C(\tau(h,x,z))+$$  where: -```text-h = harness candidate-x = task or scenario-z = seed, replicate, or profile cell-tau = full trace trajectory-R = reward, score, or verifier result-C = cost vector-lambda = cost weights-```+$$+\begin{aligned}+  h &= \text{harness candidate} \\+  x &= \text{task or scenario} \\+  z &= \text{seed, replicate, or profile cell} \\+  \tau &= \text{full trace trajectory} \\+  R &= \text{reward, score, or verifier result} \\+  C &= \text{cost vector} \\+  \lambda &= \text{cost weights}+\end{aligned}+$$  The promotion rule is not just the objective. It is a gate: @@ -288,7 +292,7 @@ promote or reject  The important word is structural. -Changing `n = 8` to `n = 16` is not harness evolution. Changing a threshold is not harness evolution. Adding another sentence to a prompt is not harness evolution.+Changing $n=8$ to $n=16$ is not harness evolution. Changing a threshold is not harness evolution. Adding another sentence to a prompt is not harness evolution.  Structural variants change mechanism: @@ -328,9 +332,9 @@ candidate_delta = median(candidate_runs - paired_baseline_runs)  The strongest comparison is paired: -```text-delta_i = R(h_candidate, x_i, z_i) - R(h_baseline, x_i, z_i)-```+$$+\Delta_i=R(h_{\text{candidate}},x_i,z_i)-R(h_{\text{baseline}},x_i,z_i)+$$  A candidate deserves promotion only when the paired evidence survives uncertainty, cost, deterministic checks, and holdout. @@ -342,7 +346,7 @@ One candidate may improve correctness and increase cost. Another may reduce late  So meta-harness tracks a Pareto frontier. -Candidate `a` dominates candidate `b` when:+Candidate $a$ dominates candidate $b$ when:  ```text quality_a >= quality_b
  3. GPT-5.3-codex-sparkpolish+7−13 view trace →
    we have a company website in ~/webb/tangle-website maybe? I want to evaluate which blog posts from this blog we can mirror on that website s · 37 asst turns · 23 tool calls
    show diff
    diff --git a/src/content/posts/self-improving-stack-harness-evolution.mdx b/src/content/posts/self-improving-stack-harness-evolution.mdxindex 094b4a6..22768de 100644--- a/src/content/posts/self-improving-stack-harness-evolution.mdx+++ b/src/content/posts/self-improving-stack-harness-evolution.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-06   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.3-codex-spark', role: 'rewrite', date: 2026-06-06 } revisions:+  - { date: 2026-06-06, model: 'gpt-5.3-codex-spark', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-06T18-19-05-739Z-gpt-5.3-codex-spark-self-improving-stack-harness-evolution-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-harness-evolution-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-harness-evolution-review' }   - date: 2026-06-05@@ -41,21 +43,15 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -When the prompt keeps asking for a capability the runtime cannot express, stop optimizing the prompt and change the machine.+When the prompt keeps asking for a capability the runtime cannot express, the next improvement is not a better sentence. It is a different machine. -That is the harness-evolution moment.+That is the harness-evolution moment. Prompt optimizers can discover better wording, examples, instructions, rubrics, and sometimes better high-level tactics. Skill optimizers can discover reusable procedures. Runtime topology can change how many workers act, who reviews them, and what gets selected. -Prompt optimizers can discover better wording, examples, instructions, rubrics, and sometimes better high-level tactics. Skill optimizers can discover reusable procedures. Runtime topology can change how many workers act, who reviews them, and what gets selected.--Harness evolution goes one layer higher.--It changes the code that defines the agent's reachable behavior.+Harness evolution goes one layer higher: it changes the code that defines the agent's reachable behavior.  That code might be a planner contract, a driver, a verifier, a budget policy, a benchmark adapter, a trace schema, a replay layer, a selector, a persona manifest, a tool router, or a worktree candidate lifecycle. The harness is not the model. It is the machine around the model that determines which actions exist, which observations are visible, which branches can run, which artifacts count, and which candidate is allowed to become production. -So no, GEPA, SkillOpt, AlphaEvolve-style code search, and meta-harness are not all "doing the same thing" in the strong sense.--They share an outer loop:+So no, GEPA, SkillOpt, AlphaEvolve-style code search, and meta-harness are not all "doing the same thing" in the strong sense. They share an outer loop:  ```text propose candidate@@ -65,9 +61,7 @@ select survivor repeat ``` -They differ in the mutable surface.--That distinction is everything.+They differ in the mutable surface. That distinction is everything.  | Optimizer family | Mutable candidate | Reachable change | Hard limit | |---|---|---|---|
  4. GPT-5.3-codex-sparkrewrite view trace →
    60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
  5. GPT-5.5draft view trace →
    Drafted the harness-evolution post with structural search formalism, meta-harness lifecycle, frontier and gate protocol, worktree isolation, proxy-metric failure modes, maxTurns=0 multi-agent placement, and local Tangle package mapping.
  6. GPT-5.5polish view trace →
    Polished the harness-evolution post by adding a prompt/skill/runtime/harness comparison table, tightening the Tangle package export mapping, and clarifying the local source-version versus dependency-version boundary.
  7. GPT-5.5publish view trace →
    Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
  8. GPT-5.5review view trace →
    Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
  9. GPT-5.5outline view trace →
    Research planning pass from a traced session.

Comments

Comments load from GitHub Discussions via Giscus. Configure PUBLIC_GISCUS_REPO, PUBLIC_GISCUS_REPO_ID, PUBLIC_GISCUS_CATEGORY, and PUBLIC_GISCUS_CATEGORY_ID in .env. See giscus.app to generate the IDs after you enable Discussions on the repo.