Topology Is The Missing Action Space

Why multi-agent self-improvement needs explicit runtime primitives for fanout, refine, select, parallelism, supervision, budgets, and replay.

The Self Improving Stack series

Next → The Gate Is The Optimizer
Browse all 13 posts
  1. Jun 2026 Topology Is The Missing Action Space
  2. Jun 2026 The Gate Is The Optimizer
  3. Jun 2026 Self-Improvement Needs A Safety Case
  4. Jun 2026 When The Harness Has To Evolve
  5. Jun 2026 Memory Is Not Automatically Learning
  6. Jun 2026 Personas Are Content, Coordination Is Structure
  7. Jun 2026 Optimization Theory For Agent Builders
  8. Jun 2026 When The Model Itself Is Mutable
  9. Jun 2026 Prompt Optimization Is Not The Whole Game
  10. Jun 2026 Skills Are Trainable State
  11. Jun 2026 Beat Random At Equal Compute First
  12. Jun 2026 Traces Are The Training Data
  13. Jun 2026 The Self-Improving Stack
Authored by
outlineGPT-5.5draftGPT-5.5polishGPT-5.5reviewGPT-5.5publishGPT-5.5rewriteGPT-5.5polishGPT-6-luna

When I tell a coding agent to “parallelize the work,” I am not asking for a different tone.

I am asking for a different execution graph: spawn independent executions, cap concurrency, isolate state, collect traces, score results, select or merge outputs, cancel losers, and account for cost. If the runtime cannot express those moves, a prompt can only ask the model to simulate the shape.

This is the missing action space in many agent systems.

Prompt optimization tunes text. Skill optimization trains durable procedure. Runtime topology optimization changes what can actually happen during execution.

What Topology Means

An agent runtime topology is the executable shape of the work.

It is not the persona. It is not the supervisor prompt. It is not the model’s private chain of thought. It is the control structure that decides which agent runs, with which tools, in what order, under what budget, with what state isolation, and with what termination rule.

A minimal topology has:

  • Nodes: Agents, tools, validators, selectors, and human gates
  • Edges: Sequence, fanout, handoff, retry, interrupt, and merge
  • State: Traces, memory, artifacts, budgets, and run handles
  • Policy: Planning, selection, cancellation, promotion, and replay

The runtime action space is the set of moves the system can execute:

A_runtime = {
  call_tool,
  call_agent,
  delegate,
  fork,
  parallel,
  refine,
  select,
  merge,
  interrupt,
  abort,
  checkpoint,
  replay
}

If parallel is not in AruntimeA_{\text{runtime}}, no optimized prompt can make true parallelism appear. If checkpoint and replay are absent, a long-running agent has no durable execution boundary. If select is only an LLM preference expressed in prose, the system has no enforceable winner rule.

The deep question is:

Which topology moves are first-class runtime actions, and which exist only as instructions?

That one distinction determines whether multi-agent self-improvement is engineering or theater.

The Optimization Problem

Let:

g=runtime topologyπ=runtime policy over moves in Aruntimem=model/backend setp=prompts and role descriptionsk=active skillsu=tools and external affordancesx=task from distribution DR=trajectory reward or eval scoreC=cost, latency, compute, human review, or risk\begin{aligned} g &= \text{runtime topology} \\ \pi &= \text{runtime policy over moves in } A_{\text{runtime}} \\ m &= \text{model/backend set} \\ p &= \text{prompts and role descriptions} \\ k &= \text{active skills} \\ u &= \text{tools and external affordances} \\ x &= \text{task from distribution } D \\ R &= \text{trajectory reward or eval score} \\ C &= \text{cost, latency, compute, human review, or risk} \end{aligned}

Runtime topology optimization estimates:

J(g,π∣m,p,k,u)=Ex∼D[R(run⁡(g,π,m,p,k,u,x))]−λ E[C(run⁡(g,π,m,p,k,u,x))]J(g, \pi \mid m,p,k,u) = \mathbb{E}_{x\sim D}[R(\operatorname{run}(g,\pi,m,p,k,u,x))] - \lambda\,\mathbb{E}[C(\operatorname{run}(g,\pi,m,p,k,u,x))]

Prompt and skill optimization usually keep gg fixed. Topology optimization changes gg, π\pi, or both.

This is not gradient descent over a dense parameter tensor. It is discrete search over executable program structure: graph edges, worker counts, branch policies, selectors, validators, budget ledgers, replay boundaries, and promotion gates.

This is a larger search space. It includes:

  • whether to solve sequentially or in parallel
  • how many workers to spawn
  • which worker profiles to use
  • whether to refine, vote, merge, or hand off
  • when to stop
  • what budget to enforce
  • which verifier is authoritative
  • whether failures retry, abort, or escalate
  • whether state is shared, forked, or isolated

The objective is still the same skeleton: propose, run, score, compare, update, promote. The mutable surface is now the control program.

Why Persona Prompts Are Too Weak

A supervisor prompt can say:

  1. 01
    Assign independent subtasks
  2. 02
    Run specialist workers in parallel
  3. 03
    Merge their findings
  4. 04
    Verify the result
  5. 05
    Stop when the verifier passes

That text only has leverage if the runtime has matching actions.

The prompt can choose among available tools. It cannot create an async task queue. It cannot isolate worktrees. It cannot enforce max concurrency. It cannot guarantee all worker traces are captured. It cannot make a verifier’s decision final if the loop ignores that decision. It cannot resume after process death if the execution was never journaled.

This is why maxTurns, maxIterations, maxConcurrency, timeout, and budget are not writing advice. They are part of the policy surface.

maxTurns = 0 is especially revealing. Depending on the harness, it can mean no autonomous continuation, unbounded continuation, or delegation to an outer driver that owns turn accounting. Those are three different systems. A prompt optimizer cannot infer the intended semantics unless the harness makes them explicit and the traces expose the result.

The same logic applies to supervisors and coordinators. Personification can shape priors, tone, and local judgment. It does not grant authority. The runtime decides whether supervision is a real control point.

The Current Runtime Map

As of June 5, 2026, the agent framework ecosystem is converging on the same idea: topology is moving out of prose and into runtime primitives.

LangGraph describes itself as a low-level orchestration framework and runtime for long-running, stateful agents, with persistence, fault tolerance, streaming, interrupts, memory, subgraphs, human-in-the-loop, and tracing. Its core point is not a better prompt template. It is executable graph state.

AutoGen AgentChat exposes teams and named multi-agent patterns: selector group chat, swarm, Magentic-One, and GraphFlow workflows over directed graphs of agents. Again, the interesting object is not the agent bio. It is the coordination pattern.

OpenAI’s Agents SDK documentation separates orchestration via LLM from orchestration via code. It names handoffs, agents as tools, evaluator loops, and parallel execution as distinct patterns. The docs make the tradeoff explicit: code-level orchestration is more deterministic and predictable in speed, cost, and performance.

Temporal is not an agent framework, but it matters for the runtime conversation because it gives workflow systems durable primitives: workflows, activities, workers, child workflows, cancellation, timers, versioning, and message passing. Long-running agents need the same class of execution guarantees when outputs matter.

No single framework settles the question. The shared signal is that serious systems are making execution shape explicit.

The Tangle Runtime Surface

@tangle-network/agent-runtime belongs exactly at this layer.

The inspected package surface exposes:

  • runLoop: topology-agnostic loop kernel over sandbox executions.
  • Driver: the object that owns topology through plan() and decide().
  • createRefineDriver: single-task iterative refinement until validator pass or cap.
  • createFanoutVoteDriver: N parallel attempts, score valid outputs, pick winner or fail.
  • AgentRunSpec: profile plus task-to-prompt formatter for one runnable agent.
  • OutputAdapter: event stream to typed output.
  • Validator: output to score and pass/fail verdict.
  • MCP delegation tools: delegate_code, delegate_research, delegation_status, delegation_history, and delegate_feedback.

The important split is in the type surface:

  • Kernel: Iteration accounting, concurrency, aborts, cost, and traces
  • Driver: Topology
  • Validator: Scoring
  • Output adapter: Parsing
  • Agent spec: Executable profile and prompt formatting

The decomposition matters. Swapping a Driver changes topology without changing the model, prompt, skill body, validator, or output parser. The execution graph becomes a replaceable runtime object rather than a paragraph inside a supervisor prompt.

For refine, the driver emits one task per iteration until the validator accepts the output or a cap is reached:

  1. 01
    Attempt
  2. 02
    Validate
  3. 03
    Retry if invalid
  4. 04
    Stop on pass or cap

For fanout-vote, the driver emits N attempts in the first iteration, lets the kernel run them in parallel subject to maxConcurrency, then selects the highest-scoring valid output:

  1. 01
    Spawn N candidates
  2. 02
    Validate each
  3. 03
    Select a valid winner
  4. 04
    Fail if none pass

The MCP delegation layer turns this topology into an agent-callable surface. delegate_code can launch specialist coder agents that produce validated patches, return immediately with a taskId, and let the caller poll for completion. With variants > 1, multiple coder harnesses attempt the task in parallel and the highest-scoring patch wins. delegate_research does the same shape for evidence-bearing research, with source diversity, citation density, recency, gap coverage, and namespace isolation in the scoring contract.

That is executable topology. A prompt that says “parallelize” is not.

The Boundary Matters

Some useful concepts are not present in the inspected local agent-runtime package surface:

The useful conclusion is not “missing feature.” It is a sharper map:

Shipped in the checked source: runLoop, refine, fanout-vote, and multi-harness coder/research delegation.

Obvious next primitives: dynamic driver, typed program DSL, supervisor scope, budget ledger, and durable replay.

This boundary is load-bearing. A substrate map is wrong if it treats unshipped names as APIs. If a runtime does not yet expose Supervisor or Scope, the concept can still be named as a target surface. It cannot be treated as a shipped API.

The next layer would make recursive execution explicit:

scope = {
  budget_remaining,
  depth_remaining,
  allowed_tools,
  allowed_agents,
  trace_parent,
  cancellation_token
}

spawn(scope, child_spec) -> child_scope

That is what a supervisor needs to be more than a persona. It needs a scoped ability to allocate work, conserve budget, observe child traces, merge results, and abort or retry branches.

Selector Versus Judge

Topology creates a new failure mode: confusing the selector with the judge.

A judge scores an output. A selector chooses the next branch, winner, or action. In simple systems they can be the same function. In serious agent systems they should be separated.

A judge is epistemic. It estimates quality. A selector is executive. It spends budget, cancels branches, returns artifacts, and determines which trajectory becomes system behavior.

judge(output, trace) -> score, dimensions, rationale
selector(candidates, scores, policy) -> next branch or winner

Why separate them?

Because selection is a runtime authority. It decides which branch continues, which worker wins, which patch is returned, which tool call is blocked, and when the system stops spending money. If an LLM judge gives a score, but the runtime selector ignores cost, safety, or deterministic failures, the topology is weak even if the judge is smart.

This is why agent-eval belongs next to agent-runtime. Runtime decides what work actually ran. Eval decides whether the output, trace, and cost evidence are strong enough to promote the topology or its parameters.

Compute Matching

Topology optimization has to control compute.

If candidate A gets one worker and candidate B gets eight workers, B may win because it spent more, not because its topology is better. That can still be a valid product decision, but it is not a compute-matched comparison.

A runtime topology benchmark should record:

  • Workers spawned
  • Iterations used
  • Wall-clock latency
  • Tokens in and out
  • Tool and LLM calls
  • Failed and cancelled branches
  • Human approvals
  • cost_usd

Then promotion can distinguish:

  • Quality lift at the same budget
  • Quality lift for a higher budget
  • Latency reduction at the same quality
  • Cost reduction at the same quality
  • Risk reduction with acceptable quality loss

Without this ledger, topology search will usually rediscover “try more things” and call it intelligence.

Multi-Agent Optimization

A multi-agent candidate is not one prompt. It is a structured object:

s = {
  topology,
  role_prompts,
  active_skills_by_role,
  agent_profiles,
  tool_policy,
  memory_policy,
  validator_stack,
  selector_policy,
  budget_policy
}

If only role_prompts are mutable, GEPA can improve personas. If active_skills_by_role and skill_bodies are mutable, SkillOpt-style methods can improve durable procedure. If topology and selector_policy are mutable, the optimizer is now searching runtime architecture.

This is the answer to the “can GEPA optimize the whole multi-agent workflow?” question. It can optimize text surfaces that influence the workflow. It searches topology only if topology is serialized as a candidate, executed by the runtime, observed in traces, and scored by an evaluator. Otherwise it can discover better instructions about coordination, not better coordination mechanisms.

That expanded action space is powerful, but it raises the bar for evidence.

The trace must show:

  • which branches were spawned
  • what each branch saw
  • which tools each branch used
  • which state was shared or isolated
  • which validator scored each output
  • why the selector picked the winner
  • which branches were cancelled
  • what budget was consumed
  • whether replay would produce the same control path

Without that trace, you cannot tell whether the topology helped, whether one branch got lucky, or whether the selector quietly ignored the evidence.

The Topology Test

A serious runtime topology eval should treat topology changes as architecture changes.

Minimum protocol:

  1. Freeze the model, prompts, skills, tools, dataset, and evaluator where possible.
  2. Register the baseline and candidate topology hashes.
  3. Run paired scenario and seed comparisons.
  4. Record every branch, tool call, validator result, selector decision, and cost.
  5. Compare under at least one compute-matched budget.
  6. Stress test timeouts, branch failures, cancellation, and partial results.
  7. Reject candidates that hide errors, exceed budget, lose traces, or skip gates.
  8. Promote only when held-out lift, cost and latency policy, and trace integrity pass.

Promotion can look like:

promote(g_new) if:
  LCB_95(median(score_new - score_base on holdout)) > epsilon
  and median_cost_new <= cost_ceiling
  and median_latency_new <= latency_ceiling
  and trace_integrity == 1
  and deterministic_failures == 0
  and safety_regressions == 0

For topology, trace_integrity is not optional. A candidate that wins while losing branch traces, skipping validator spans, or hiding failed children is not a better runtime. It is an unobservable runtime.

How Topology Lies

Runtime topology fails in recognizable ways.

It hides coordination in prose. It spawns workers without isolation. It lets branches clobber the same files. It votes with a weak judge. It retries the same failure shape. It spends more compute and calls that progress. It cancels useful branches too early. It never cancels losing branches. It loses traces. It lets a supervisor override deterministic failures. It treats human approval as final approval rather than a scoped interrupt. It has no replay semantics, so every bug is a rumor.

The most common failure is fake fanout:

  • Prompt: Says to split the task among specialists.
  • Runtime: Makes one model call with specialist names in the text.
  • Trace: Records one branch.

The fix is not a better coordinator prompt. The fix is an actual fanout primitive.

Another failure is unpriced parallelism:

The candidate topology spawns eight workers while the baseline spawns one. It wins by four points, costs twelve times more, and the promotion report calls it “better.”

That is not necessarily wrong. It is incomplete. The product decision depends on whether the gain is worth the compute, latency, and operational complexity.

When Topology Is The Right Surface

Use prompt optimization when the failure is wording.

Use skill optimization when the failure is recurring procedure.

Use runtime topology optimization when the failure is execution shape:

  • The system needs parallel branches.
  • The system needs a verifier loop.
  • The system needs specialist delegation.
  • The system needs human interruption and resume.
  • The system needs state isolation.
  • The system needs branch cancellation.
  • The system needs compute-matched selection.
  • The system needs replayable traces.

Do not hide these requirements in a persona.

If “parallelize” matters, make it a runtime action. If “supervise” matters, give the supervisor authority over scoped branches and budgets. If “verify” matters, make the validator a gate. If “stop” matters, make termination a policy, not a vibe.

The operational test is simple: if the desired improvement would change the system’s sequence diagram, it belongs in topology. If it only changes how an existing node thinks or writes, it may belong in a prompt or skill.

Topology is where agent systems stop being advice and become execution.

Source Trail

Source freshness checked on 2026-06-06.

Revision history9revisions
  1. GPT-6-lunapolish+39−73 view trace →
    Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
    show diff
    diff --git a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx b/src/content/posts/self-improving-stack-agent-runtime-topology.mdxindex 69b3005..6319837 100644--- a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx+++ b/src/content/posts/self-improving-stack-agent-runtime-topology.mdx@@ -46,6 +46,8 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro';+ When I tell a coding agent to "parallelize the work," I am not asking for a different tone.  I am asking for a different execution graph: spawn independent executions, cap concurrency, isolate state, collect traces, score results, select or merge outputs, cancel losers, and account for cost. If the runtime cannot express those moves, a prompt can only ask the model to simulate the shape.@@ -62,12 +64,10 @@ It is not the persona. It is not the supervisor prompt. It is not the model's pr  A minimal topology has: -```text-nodes = agents, tools, validators, selectors, human gates-edges = sequence, fanout, handoff, retry, interrupt, merge-state = trace, memory, artifacts, budgets, run handles-policy = planning, selection, cancellation, promotion, replay-```+- **Nodes:** Agents, tools, validators, selectors, and human gates+- **Edges:** Sequence, fanout, handoff, retry, interrupt, and merge+- **State:** Traces, memory, artifacts, budgets, and run handles+- **Policy:** Planning, selection, cancellation, promotion, and replay  The runtime action space is the set of moves the system can execute: @@ -92,9 +92,7 @@ If `parallel` is not in $A_{\text{runtime}}$, no optimized prompt can make true  The deep question is: -```text-Which topology moves are first-class runtime actions, and which are merely instructions?-```+Which topology moves are first-class runtime actions, and which exist only as instructions?  That one distinction determines whether multi-agent self-improvement is engineering or theater. @@ -144,13 +142,7 @@ The objective is still the same skeleton: propose, run, score, compare, update,  A supervisor prompt can say: -```text-Assign independent subtasks to specialist workers.-Have them work in parallel.-Merge their findings.-Ask a verifier to check the final result.-Stop when the verifier passes.-```+<Steps layout="flow" items={[{title:'Assign independent subtasks'}, {title:'Run specialist workers in parallel'}, {title:'Merge their findings'}, {title:'Verify the result'}, {title:'Stop when the verifier passes'}]} />  That text only has leverage if the runtime has matching actions. @@ -193,27 +185,21 @@ The inspected package surface exposes:  The important split is in the type surface: -```text-kernel owns: iteration accounting, concurrency, aborts, cost, traces-driver owns: topology-validator owns: scoring-output adapter owns: parsing-agent spec owns: executable profile and prompt formatting-```+- **Kernel:** Iteration accounting, concurrency, aborts, cost, and traces+- **Driver:** Topology+- **Validator:** Scoring+- **Output adapter:** Parsing+- **Agent spec:** Executable profile and prompt formatting  The decomposition matters. Swapping a `Driver` changes topology without changing the model, prompt, skill body, validator, or output parser. The execution graph becomes a replaceable runtime object rather than a paragraph inside a supervisor prompt.  For `refine`, the driver emits one task per iteration until the validator accepts the output or a cap is reached: -```text-attempt -> validate -> if invalid, attempt again -> stop on pass or cap-```+<Steps layout="flow" items={[{title:'Attempt'}, {title:'Validate'}, {title:'Retry if invalid'}, {title:'Stop on pass or cap'}]} />  For `fanout-vote`, the driver emits N attempts in the first iteration, lets the kernel run them in parallel subject to `maxConcurrency`, then selects the highest-scoring valid output: -```text-spawn N -> validate each -> select valid winner -> fail if none valid-```+<Steps layout="flow" items={[{title:'Spawn N candidates'}, {title:'Validate each'}, {title:'Select a valid winner'}, {title:'Fail if none pass'}]} />  The MCP delegation layer turns this topology into an agent-callable surface. `delegate_code` can launch specialist coder agents that produce validated patches, return immediately with a `taskId`, and let the caller poll for completion. With `variants > 1`, multiple coder harnesses attempt the task in parallel and the highest-scoring patch wins. `delegate_research` does the same shape for evidence-bearing research, with source diversity, citation density, recency, gap coverage, and namespace isolation in the scoring contract. @@ -225,13 +211,9 @@ Some useful concepts are not present in the inspected local `agent-runtime` pack  The useful conclusion is not "missing feature." It is a sharper map: -```text-shipped in the checked source:-  runLoop, refine, fanout-vote, multi-harness coder/research delegation+**Shipped in the checked source:** `runLoop`, `refine`, `fanout-vote`, and multi-harness coder/research delegation. -obvious next primitives:-  dynamic driver, typed program DSL, supervisor scope, budget ledger, durable replay-```+**Obvious next primitives:** dynamic driver, typed program DSL, supervisor scope, budget ledger, and durable replay.  This boundary is load-bearing. A substrate map is wrong if it treats unshipped names as APIs. If a runtime does not yet expose `Supervisor` or `Scope`, the concept can still be named as a target surface. It cannot be treated as a shipped API. @@ -279,28 +261,22 @@ If candidate A gets one worker and candidate B gets eight workers, B may win bec  A runtime topology benchmark should record: -```text-workers spawned-iterations used-wall-clock latency-tokens in/out-tool calls-LLM calls-failed branches-cancelled branches-human approvals-cost_usd-```+- Workers spawned+- Iterations used+- Wall-clock latency+- Tokens in and out+- Tool and LLM calls+- Failed and cancelled branches+- Human approvals+- `cost_usd`  Then promotion can distinguish: -```text-quality lift at same budget-quality lift for higher budget-latency reduction at same quality-cost reduction at same quality-risk reduction with acceptable quality loss-```+- Quality lift at the same budget+- Quality lift for a higher budget+- Latency reduction at the same quality+- Cost reduction at the same quality+- Risk reduction with acceptable quality loss  Without this ledger, topology search will usually rediscover "try more things" and call it intelligence. @@ -348,16 +324,14 @@ A serious runtime topology eval should treat topology changes as architecture ch  Minimum protocol: -```text-1. Freeze model, prompts, skills, tools, dataset, and evaluator where possible.-2. Register baseline topology hash and candidate topology hash.-3. Run paired scenario/seed comparisons.+1. Freeze the model, prompts, skills, tools, dataset, and evaluator where possible.+2. Register the baseline and candidate topology hashes.+3. Run paired scenario and seed comparisons. 4. Record every branch, tool call, validator result, selector decision, and cost. 5. Compare under at least one compute-matched budget.-6. Run stress cases for timeouts, branch failure, cancellation, and partial results.+6. Stress test timeouts, branch failures, cancellation, and partial results. 7. Reject candidates that hide errors, exceed budget, lose traces, or skip gates.-8. Promote only on held-out lift, cost/latency policy, and trace integrity.-```+8. Promote only when held-out lift, cost and latency policy, and trace integrity pass.  Promotion can look like: @@ -381,23 +355,15 @@ It hides coordination in prose. It spawns workers without isolation. It lets bra  The most common failure is fake fanout: -```text-prompt says: split the task among specialists-runtime does: one model call with specialist names in text-trace says: one branch-```+- **Prompt:** Says to split the task among specialists.+- **Runtime:** Makes one model call with specialist names in the text.+- **Trace:** Records one branch.  The fix is not a better coordinator prompt. The fix is an actual fanout primitive.  Another failure is unpriced parallelism: -```text-candidate topology spawns 8 workers-baseline topology spawns 1 worker-candidate wins by 4 points-candidate costs 12x more-promotion report says "better"-```+The candidate topology spawns eight workers while the baseline spawns one. It wins by four points, costs twelve times more, and the promotion report calls it “better.”  That is not necessarily wrong. It is incomplete. The product decision depends on whether the gain is worth the compute, latency, and operational complexity. 
  2. GPT-6-lunapolish+18−16 view trace →
    Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
    show diff
    diff --git a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx b/src/content/posts/self-improving-stack-agent-runtime-topology.mdxindex 8b1c5da..4d741f2 100644--- a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx+++ b/src/content/posts/self-improving-stack-agent-runtime-topology.mdx@@ -86,7 +86,7 @@ A_runtime = { } ``` -If `parallel` is not in `A_runtime`, no optimized prompt can make true parallelism appear. If `checkpoint` and `replay` are absent, a long-running agent has no durable execution boundary. If `select` is only an LLM preference expressed in prose, the system has no enforceable winner rule.+If `parallel` is not in $A_{\text{runtime}}$, no optimized prompt can make true parallelism appear. If `checkpoint` and `replay` are absent, a long-running agent has no durable execution boundary. If `select` is only an LLM preference expressed in prose, the system has no enforceable winner rule.  The deep question is: @@ -100,25 +100,27 @@ That one distinction determines whether multi-agent self-improvement is engineer  Let: -```text-g = runtime topology-pi = runtime policy over moves in A_runtime-m = model/backend set-p = prompts and role descriptions-k = active skills-u = tools and external affordances-x = task from distribution D-R = trajectory reward or eval score-C = cost, latency, compute, human review, or risk-```+$$+\begin{aligned}+  g &= \text{runtime topology} \\+  \pi &= \text{runtime policy over moves in } A_{\text{runtime}} \\+  m &= \text{model/backend set} \\+  p &= \text{prompts and role descriptions} \\+  k &= \text{active skills} \\+  u &= \text{tools and external affordances} \\+  x &= \text{task from distribution } D \\+  R &= \text{trajectory reward or eval score} \\+  C &= \text{cost, latency, compute, human review, or risk}+\end{aligned}+$$  Runtime topology optimization estimates: -```text-J(g, pi | m, p, k, u) = E_{x ~ D}[R(run(g, pi, m, p, k, u, x))] - lambda * E[C(run(g, pi, m, p, k, u, x))]-```+$$+J(g, \pi \mid m,p,k,u) = \mathbb{E}_{x\sim D}[R(\operatorname{run}(g,\pi,m,p,k,u,x))] - \lambda\,\mathbb{E}[C(\operatorname{run}(g,\pi,m,p,k,u,x))]+$$ -Prompt and skill optimization usually keep `g` fixed. Topology optimization changes `g`, `pi`, or both.+Prompt and skill optimization usually keep $g$ fixed. Topology optimization changes $g$, $\pi$, or both.  This is not gradient descent over a dense parameter tensor. It is discrete search over executable program structure: graph edges, worker counts, branch policies, selectors, validators, budget ledgers, replay boundaries, and promotion gates. 
  3. GPT-5.5polish+7−5 view trace →
    let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
    show diff
    diff --git a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx b/src/content/posts/self-improving-stack-agent-runtime-topology.mdxindex 1599922..960a85e 100644--- a/src/content/posts/self-improving-stack-agent-runtime-topology.mdx+++ b/src/content/posts/self-improving-stack-agent-runtime-topology.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-05   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-agent-runtime-topology-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-agent-runtime-topology-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-agent-runtime-topology-review' }   - date: 2026-06-05@@ -41,9 +43,9 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -"Parallelize the work" is not a style preference.+When I tell a coding agent to "parallelize the work," I am not asking for a different tone. -For a human operator talking to a coding agent, it is a request for a different execution graph: spawn independent executions, cap concurrency, isolate state, collect traces, score results, select or merge outputs, cancel losers, and account for cost. If the runtime cannot express those moves, a prompt can only ask the model to simulate the shape.+I am asking for a different execution graph: spawn independent executions, cap concurrency, isolate state, collect traces, score results, select or merge outputs, cancel losers, and account for cost. If the runtime cannot express those moves, a prompt can only ask the model to simulate the shape.  This is the missing action space in many agent systems. @@ -335,7 +337,7 @@ The trace must show:  Without that trace, you cannot tell whether the topology helped, whether one branch got lucky, or whether the selector quietly ignored the evidence. -## Evaluation Protocol+## The Topology Test  A serious runtime topology eval should treat topology changes as architecture changes. @@ -366,7 +368,7 @@ promote(g_new) if:  For topology, `trace_integrity` is not optional. A candidate that wins while losing branch traces, skipping validator spans, or hiding failed children is not a better runtime. It is an unobservable runtime. -## Failure Modes+## How Topology Lies  Runtime topology fails in recognizable ways. @@ -394,7 +396,7 @@ promotion report says "better"  That is not necessarily wrong. It is incomplete. The product decision depends on whether the gain is worth the compute, latency, and operational complexity. -## A Working Rule+## When Topology Is The Right Surface  Use prompt optimization when the failure is wording. 
  4. GPT-5.5rewrite view trace →
    60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
  5. GPT-5.5publish view trace →
    Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
  6. GPT-5.5review view trace →
    Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
  7. GPT-5.5outline view trace →
    Research planning pass from a traced session.
  8. GPT-5.5draft view trace →
    Expanded the runtime topology outline into a full draft with formal topology variables, shipped Tangle runtime primitives, external orchestration context, eval protocol, and failure modes.
  9. GPT-5.5polish view trace →
    Polished runtime topology framing, substrate-boundary language, GEPA/topology distinction, and bridge into multi-agent coordination.

Comments

Comments load from GitHub Discussions via Giscus. Configure PUBLIC_GISCUS_REPO, PUBLIC_GISCUS_REPO_ID, PUBLIC_GISCUS_CATEGORY, and PUBLIC_GISCUS_CATEGORY_ID in .env. See giscus.app to generate the IDs after you enable Discussions on the repo.