Personas Are Content, Coordination Is Structure
How driver, worker, selector, reviewer, analyst, and coordinator roles become reliable multi-agent systems instead of roleplay.
The Self Improving Stack series
Browse all 13 posts
- Topology Is The Missing Action Space
- The Gate Is The Optimizer
- Self-Improvement Needs A Safety Case
- When The Harness Has To Evolve
- Memory Is Not Automatically Learning
- Personas Are Content, Coordination Is Structure
- Optimization Theory For Agent Builders
- When The Model Itself Is Mutable
- Prompt Optimization Is Not The Whole Game
- Skills Are Trainable State
- Beat Random At Equal Compute First
- Traces Are The Training Data
- The Self-Improving Stack
I can give five agents five names and still have one mind making one mistake.
“Researcher,” “critic,” “architect,” “driver,” and “supervisor” are not coordination by themselves. They can be the same model, with the same blind spot, reading the same context, under the same budget, producing five paraphrases of the same failure. The cast list changed. The information structure did not.
Multi-agent work becomes real when disagreement becomes useful. That requires separate contracts, state boundaries, tool permissions, selection rules, budgets, and traces.
The persona is content.
Coordination is structure.
The Coordination Surface
The previous post made runtime topology explicit:
Multi-agent coordination sits one layer above that. It decides what the nodes are supposed to do together.
Let:
A multi-agent system candidate is:
The optimization target is:
The notation is only useful because it exposes the coordinate system.
If only changes, the optimizer is tuning role descriptions. If changes, it is tuning durable role procedure. If , , , , or change, it is tuning coordination structure.
That distinction prevents a common category error: evaluating a better persona as if it were a better multi-agent system.
Persona Is Not Authority
A persona can say:
You are a careful supervisor. You delegate independent work. You force disagreement before consensus. You stop when the reviewer passes the artifact.
Those instructions may improve local judgment. They do not grant runtime authority.
Authority lives in the structure:
- Can this role spawn workers?
- Can it choose tools?
- Can it read child traces?
- Can it cancel branches?
- Can it override a reviewer?
- Can it spend more budget?
- Can it merge artifacts?
- Can it promote the result?
If the answers are not represented in the runtime, the persona is aspirational. It may be useful text, but it is not a coordination contract.
This is why “supervisor” is an overloaded word. In one system, it means a prompt that asks an LLM to choose the next agent. In another, it means a graph node that routes state. In a third, it means a durable workflow controller with scoped budget, cancellation, replay, and audit authority. These are not interchangeable.
Roles As Contracts
A role becomes real when it has an input contract, an output contract, authority, state access, and accountability.
The minimum useful taxonomy:
| Role | Contract | Failure mode |
|---|---|---|
| Driver | Choose the next runtime move | private reasoning controls execution without trace |
| Planner | Decompose goal into scoped tasks | decomposition creates fake parallelism or missing dependencies |
| Worker | Produce an artifact under a task contract | hidden assumptions leak into final output |
| Reviewer | Find defects against a rubric | reviewer becomes vague taste instead of a gate |
| Judge | Score output or trajectory | judge labels are reused as training signal without holdout |
| Selector | Pick branch, winner, or continuation | selector optimizes agreement, not correctness |
| Coordinator | Allocate work and merge artifacts | merge erases dissent and provenance |
| Analyst | Convert traces into findings | analyst lists symptoms instead of causal failure modes |
The same LLM can occupy several roles, but the role contracts still need separation. A selector that also writes the candidate can rationalize its own work. A reviewer that never blocks promotion is commentary. A coordinator that cannot see child traces is guessing.
In code, the role split looks less like a cast list and more like a typed interface:
worker(task, context, tools, budget) -> artifact, trace
reviewer(artifact, rubric, trace) -> defects, pass
judge(artifact, task, trace) -> score, dimensions
selector(candidates, scores, budget_policy) -> winner | continue | abort
coordinator(goal, children, traces) -> merged_artifact, lineage
The point is not bureaucracy. The point is falsifiability. If a role has no contract, its contribution cannot be tested.
The Disagreement Problem
Multi-agent systems are usually sold as specialization. The stronger reason is controlled disagreement.
For an ensemble to help, at least one of these must be true:
- agents see different evidence
- agents use different tools
- agents use different models
- agents use different skills
- agents explore different branches
- agents are scored by an independent verifier
- the selector can preserve dissent instead of forcing consensus
If every worker shares the same prompt, same model, same examples, same context, same decoding parameters, and same evaluator, the ensemble is highly correlated.
The idealized independent case is easy:
where is the probability that worker independently finds a valid answer.
But multi-agent LLM systems rarely get independence for free. The useful quantity is not worker count. It is error correlation.
If the gain disappears at matched budget, the system did not learn coordination. It bought more samples.
If the gain disappears when workers use isolated context, the system may have been copying. If the gain disappears when the selector is replaced with a deterministic verifier, the selector may have been rewarding style. If the gain disappears on held-out tasks, the role split overfit the benchmark.
This is the central test: does the coordination policy create useful diversity, or just more tokens?
Lineage: From Sampling To Societies
The modern multi-agent conversation did not appear fully formed. It grew out of several older ideas.
Self-consistency sampled multiple reasoning paths and selected the most consistent answer. That is not multi-agent in the social sense, but it is the simplest version of a coordination move: generate diverse candidates, aggregate them with a rule, and beat greedy decoding on reasoning tasks.
Tree of Thoughts made the search structure more explicit. It explored coherent intermediate thoughts, evaluated them, and allowed lookahead and backtracking. Again, the key object is not a persona. It is a search policy over branches.
CAMEL pushed role-playing into agent cooperation. Its important contribution is not that agents had names. It is that role-conditioned communicative agents could be studied as a cooperative system, with inception prompting used to keep the interaction on task.
AutoGen framed multi-agent applications as configurable conversable agents, where interaction behavior can be programmed in natural language or code. That shifted attention from one prompt to conversation protocol.
Multiagent debate showed another route: multiple model instances propose and critique answers over rounds, improving factuality and reasoning in some settings. But debate also reveals the danger. More discussion is not automatically more truth. It can become persuasion, anchoring, or convergence to a fluent wrong answer unless the selector and verifier are strong.
Mixture-of-Agents made the ensemble structure more layered: agents generate outputs, later agents consume previous outputs as auxiliary information, and an aggregator improves the final response. This is close to a production pattern: proposers, aggregators, and selectors are distinct roles.
The research arc is clear:
- 01 Sample many
- 02 Search branches
- 03 Assign roles
- 04 Debate
- 05 Aggregate
- 06 Orchestrate
The open engineering problem is making the orchestration measurable.
Coordination Patterns
The useful patterns are not defined by agent names. They are defined by information flow and authority.
Best-of-N
Multiple workers attempt the same task. A selector or verifier picks one.
- 01 Spawn N
- 02 Score each
- 03 Return winner
This is strong when outputs are easy to score and independent attempts are cheap. It is weak when scoring is subjective or all workers share the same blind spot.
Self-consistency
Multiple reasoning paths produce candidate answers. The system chooses the answer supported by the most paths or highest marginal score.
- 01 Sample paths
- 02 Marginalize answers
- 03 Choose stable answer
This helps when the final answer has a stable attractor and errors are diverse. It is less useful for open-ended artifact quality where many incompatible answers can all be plausible.
Tree search
The system expands intermediate states, evaluates partial progress, and prunes weak branches. It selects a frontier to continue or backtracks to another branch.
This is coordination over thoughts, plans, or artifacts. It requires explicit state and a heuristic good enough to guide search.
Debate
Agents expose arguments, counterarguments, and revisions before a final decision.
- 01 Propose
- 02 Critique
- 03 Respond
- 04 Judge
Debate is useful when hidden assumptions matter. It fails when agents optimize rhetoric, defer to the strongest voice, or converge before evidence changes.
Supervisor
A central coordinator delegates scoped tasks to workers and keeps authority over the final artifact.
- 01 Supervisor
- 02 Assign
- 03 Collect
- 04 Merge
- 05 Verify
This is good for work that has clear subdomains. It fails when the supervisor has no real budget, no child trace access, or no merge discipline.
Handoff
Control transfers from one agent to another.
- 01 Triage
- 02 Active specialist
- 03 Maybe hand off again
Handoffs are good when the next specialist should own the state and speak directly. They are dangerous when state transfer is implicit or context grows without boundaries.
Blackboard
Agents write partial results into a shared workspace. Other agents read and improve them.
This is natural for code, research, planning, and design. It needs locking, provenance, conflict resolution, and traceable authorship.
Layered mixture
One layer proposes outputs. Later layers aggregate, refine, or route.
- 01 Proposers
- 02 Aggregators
- 03 Final selector
This works when the aggregator can exploit complementary model strengths. It fails when later layers smooth away critical dissent.
The Framework Map
As of June 5, 2026, major agent frameworks expose this distinction directly.
OpenAI’s Agents SDK documentation defines orchestration as which agents run, in what order, and how that decision is made. It separates LLM-driven orchestration from code-driven orchestration, names agents-as-tools and handoffs as common patterns, and explicitly says code orchestration is more deterministic and predictable for speed, cost, and performance.
AutoGen AgentChat exposes teams and multi-agent design patterns, including Selector Group Chat, Swarm, Magentic-One, and GraphFlow. The documentation names selectors, shared context, localized tool-based routing, and directed graphs as first-class concepts.
LangChain and LangGraph documentation describes multi-agent systems as coordination among specialized components, while warning that a single agent with the right tools and prompt can often be enough. Its handoff docs are especially concrete: behavior changes through state, agents can be distinct graph nodes, and context engineering determines what messages cross agent boundaries.
The shared direction is not “more personas.” It is explicit control over routing, handoff, state, context, and observability.
Where The Local Stack Fits
In the local @tangle-network/agent-runtime@0.26.0 source, the surface separates three coordination shapes.
The first layer is the focused multi-shot kernel:
runLoop: a topology-agnostic kernel over sandbox executions.Driver: the topology object throughplan()anddecide().createRefineDriver: serial attempt, validate, retry until pass or cap.createFanoutVoteDriver: parallel attempts with scored winner selection.AgentRunSpec: profile plus task-to-prompt formatter.OutputAdapter: sandbox event stream to typed output.Validator: typed output to score and pass/fail verdict.
This kernel is intentionally narrow. It owns iteration accounting, bounded concurrency, abort propagation, cost aggregation, and trace emission. It does not own persona, domain policy, output scoring, or topology. That shape is good for optimization: the mutable coordinate is the driver and profile set, not a hidden monolith.
The second layer is the multi-agent conversation substrate:
defineConversation: declares participants and policy before execution.runConversation/runConversationStream: drive speaker turns and event streams.createConversationBackend: lets a whole conversation become a participant in a larger conversation.ConversationPolicy:maxTurns,maxCreditsCents, turn order, halt predicate, default call policy.ConversationParticipant.authSource: per-participant billing identity, either forward the user or use agent-owned credentials.ConversationJournal: resumable transcript storage, with in-memory, file, and SQL implementations.turnId: deterministic per-turn id for retries and trace stitching.buildForwardHeadersandDEFAULT_MAX_DEPTH: cross-gateway run, turn, parent-turn, speaker, authorization, and recursion-depth propagation.CircuitBreakerStateand call policy: per-participant deadlines, retries, backoff, and circuit breaking.
This is a more serious coordination surface. It makes long-running multi-agent dialogue a runtime object with durability, economics, recursion bounds, and trace correlation.
The MCP layer is a third shape, not a replacement for either of the first two. delegate_code, delegate_research, delegate_feedback, delegation_status, and delegation_history expose async fire-and-poll delegation to agents. The runtime owns the queue, feedback store, schemas, and tool projection. The product supplies the delegates. The default coder delegate is shipped through coderProfile and multiHarnessCoderFanout; researcher delegation is peer-backed through @tangle-network/agent-knowledge or an injected ResearcherDelegate, not a top-level agent-runtime/profiles export in the inspected source.
The separation is important:
runLoop: bounded multi-shot task kernel.conversation: long-horizon participant dialogue.- MCP delegation: async specialist work surface.
A reliable coordination stack needs the following contracts regardless of which layer hosts them:
scope:
budget
allowed tools
allowed agents
state visibility
cancellation authority
trace parent
assignment:
task
role contract
input artifacts
expected output
verifier
deadline
selection:
candidates
scores
cost
risk
lineage
decision rationale
In the local @tangle-network/agent-eval@0.34.1 source, the package is not just a judge wrapper. It is a promotion and analysis system:
AgentProfileCell,AGENT_PROFILE_KINDS,buildSandboxAgentProfileCell, andtoAgentProfileJson: stable cells for model, prompt, tool, skill, runtime, and harness variation.runEvalCampaign: variant by scenario campaign runner with raw-provider capture and profile-cell checks.HeldOutGate: paired promotion gate, now with a cost ceiling so lift cannot ignore budget.- scorecards and release confidence: longitudinal evidence, paired deltas, overfit gaps, release reports.
runProductionLoop: production trace clusters to candidate improvement to held-out gate to PR.runIntentMatchJudge, failure taxonomy, semantic judges, and multi-layer verifiers: scoring beyond one rubric prompt.AnalystRegistrywithDEFAULT_TRACE_ANALYST_KINDS: failure-mode, knowledge-gap, knowledge-poisoning, and improvement analysts over trace stores.- focused subpaths such as
/optimization,/reporting,/control,/rl,/traces,/pipelines,/meta-eval,/prm,/builder-eval,/governance, and/knowledge.
For multi-agent coordination, the clean split is:
- agent-runtime/conversation decides which participants spoke, under which policy
- agent-runtime/loops decides which bounded workers ran
- agent-runtime/mcp exposes async specialist delegation
- agent-eval decides whether the resulting system was better
Multi-agent coordination without eval is theater. Eval without trace-level runtime evidence is an opinion poll.
Why More Agents Often Make Things Worse
The failure modes are predictable.
Correlated blind spots
Five agents using the same model and context may agree because they share the same missing fact.
Consensus collapse
Agents converge on the first plausible answer because nobody has authority or incentive to preserve dissent.
Selector overfitting
The selector learns to prefer fluent, long, confident, or rubric-shaped outputs instead of correct ones.
Unpriced compute
The multi-agent variant wins because it used 8 workers against a single-worker baseline.
Context contamination
A worker sees another worker’s answer before producing its own, so the supposed independent samples are not independent.
Merge loss
The coordinator combines outputs but drops provenance, uncertainty, and unresolved contradictions.
Authority confusion
The reviewer finds a hard failure, but the supervisor treats it as advisory feedback and ships anyway.
Trace gaps
The final answer looks good, but the system cannot show which child saw which context, used which tools, or caused which decision.
These are not edge cases. They are the default unless the coordination structure prevents them.
The Coordination Test
Do not ask whether a multi-agent system “feels smarter.” Ask whether it beats the right baseline.
Minimum protocol:
- Define the task distribution and artifact contract.
- Freeze model set, tools, prompts, skills, dataset, and evaluator where possible.
- Compare against best single-agent and best-of-N baselines at matched budget.
- Record child context, tool calls, artifacts, scores, selector decisions, and merge lineage.
- Measure quality, cost, latency, branch failure rate, trace integrity, and human review load.
- Run ablations: no debate, no shared context, no heterogeneity, deterministic selector.
- Promote only on held-out lift with acceptable cost, latency, and failure-mode profile.
The promotion rule can be written:
The selector_ablation_delta term matters. If the multi-agent system still performs the same when the selector is replaced with a trivial rule, the sophisticated coordination may not be doing causal work.
For open-ended work, add a disagreement audit:
- disagreement_audit:
- independent evidence found?
- contradictions preserved?
- reviewer defects resolved?
- final merge cites child lineage?
- rejected branches explained?
Disagreement is useful only when it changes the final decision or improves confidence calibration.
What Optimizers Can And Cannot Do
Prompt optimizers can improve role instructions:
- critic prompt
- planner prompt
- selector rubric
- handoff description
- reviewer checklist
Skill optimizers can improve durable role procedure:
- how a reviewer inspects a patch
- how a researcher triangulates sources
- how a coordinator merges conflicting evidence
- how an analyst clusters trace failures
Runtime topology optimizers can improve execution shape:
- fanout width
- which roles run in parallel
- whether debate happens before or after evidence collection
- which selector sees which fields
- when branches cancel
- how budget is allocated
Harness evolution can change the coordination machine itself:
- new driver
- new selector implementation
- new trace schema
- new sandbox isolation model
- new promotion gate
This is where GEPA, MIPRO, SkillOpt, agent-runtime, agent-eval, and meta-harness stop looking like competitors. They operate on different mutable surfaces.
The mistake is asking one optimizer to search a surface it cannot execute.
When More Agents Are Worth It
Use multiple agents when the work needs at least one of these:
- independent evidence gathering
- heterogeneous tools or models
- decomposable subtasks with real parallelism
- adversarial review
- branch search with pruning
- scoped handoff to a specialist
- artifact merge with provenance
- trace analysis by multiple lenses
Do not use multiple agents when the only benefit is a richer cast list.
The engineering test is simple:
- Can the system show why this role existed?
- Can it show what information the role had?
- Can it show what the role produced?
- Can it show how the selector used or rejected that output?
- Can it beat a compute-matched single-agent baseline?
If not, the coordination is not yet a system property. It is prose.
Personas can help agents think in different local modes. Coordination decides whether those modes become useful work.
Source Trail
Source freshness checked on 2026-06-06.
- Self-Consistency Improves Chain of Thought Reasoning in Language Models, checked June 5, 2026.
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models, checked June 5, 2026.
- CAMEL: Communicative Agents for “Mind” Exploration of Large Language Model Society, checked June 5, 2026.
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, checked June 5, 2026.
- Improving Factuality and Reasoning in Language Models through Multiagent Debate, checked June 5, 2026.
- Mixture-of-Agents Enhances Large Language Model Capabilities, checked June 5, 2026.
- OpenAI Agents SDK orchestration docs, checked June 5, 2026.
- AutoGen AgentChat docs, checked June 5, 2026.
- LangChain multi-agent docs, checked June 5, 2026.
- Local
@tangle-network/agent-runtime@0.26.0source audit:runLoop,Driver,createRefineDriver,createFanoutVoteDriver,defineConversation,runConversation,createConversationBackend,ConversationJournal,authSource, cross-gateway headers, MCP delegation tools, June 6, 2026. - Local
@tangle-network/agent-eval@0.34.1source audit:AgentProfileCell,AGENT_PROFILE_KINDS,buildSandboxAgentProfileCell,runEvalCampaign,HeldOutGate, release confidence, scorecard, intent-match judge, failure taxonomy,AnalystRegistry,DEFAULT_TRACE_ANALYST_KINDS, June 6, 2026.
Revision history9revisions
- Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
show diff
diff --git a/src/content/posts/self-improving-stack-multi-agent-coordination.mdx b/src/content/posts/self-improving-stack-multi-agent-coordination.mdxindex 69b6dd4..77fddee 100644--- a/src/content/posts/self-improving-stack-multi-agent-coordination.mdx+++ b/src/content/posts/self-improving-stack-multi-agent-coordination.mdx@@ -47,6 +47,9 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'++ I can give five agents five names and still have one mind making one mistake. "Researcher," "critic," "architect," "driver," and "supervisor" are not coordination by themselves. They can be the same model, with the same blind spot, reading the same context, under the same budget, producing five paraphrases of the same failure. The cast list changed. The information structure did not.@@ -108,27 +111,23 @@ That distinction prevents a common category error: evaluating a better persona a A persona can say: -```text-You are a careful supervisor.-You delegate independent work.-You force disagreement before consensus.-You stop when the reviewer passes the artifact.-```+> You are a careful supervisor.+> You delegate independent work.+> You force disagreement before consensus.+> You stop when the reviewer passes the artifact. Those instructions may improve local judgment. They do not grant runtime authority. Authority lives in the structure: -```text-Can this role spawn workers?-Can it choose tools?-Can it read child traces?-Can it cancel branches?-Can it override a reviewer?-Can it spend more budget?-Can it merge artifacts?-Can it promote the result?-```+- Can this role spawn workers?+- Can it choose tools?+- Can it read child traces?+- Can it cancel branches?+- Can it override a reviewer?+- Can it spend more budget?+- Can it merge artifacts?+- Can it promote the result? If the answers are not represented in the runtime, the persona is aspirational. It may be useful text, but it is not a coordination contract. @@ -219,9 +218,7 @@ Mixture-of-Agents made the ensemble structure more layered: agents generate outp The research arc is clear: -```text-sample many -> search branches -> assign roles -> debate -> aggregate -> orchestrate-```+<Steps layout="flow" items={[{title: "Sample many"}, {title: "Search branches"}, {title: "Assign roles"}, {title: "Debate"}, {title: "Aggregate"}, {title: "Orchestrate"}]} /> The open engineering problem is making the orchestration measurable. @@ -233,9 +230,7 @@ The useful patterns are not defined by agent names. They are defined by informat Multiple workers attempt the same task. A selector or verifier picks one. -```text-spawn N -> score each -> return winner-```+<Steps layout="flow" items={[{title: "Spawn N"}, {title: "Score each"}, {title: "Return winner"}]} /> This is strong when outputs are easy to score and independent attempts are cheap. It is weak when scoring is subjective or all workers share the same blind spot. @@ -243,9 +238,7 @@ This is strong when outputs are easy to score and independent attempts are cheap Multiple reasoning paths produce candidate answers. The system chooses the answer supported by the most paths or highest marginal score. -```text-sample paths -> marginalize answers -> choose stable answer-```+<Steps layout="flow" items={[{title: "Sample paths"}, {title: "Marginalize answers"}, {title: "Choose stable answer"}]} /> This helps when the final answer has a stable attractor and errors are diverse. It is less useful for open-ended artifact quality where many incompatible answers can all be plausible. @@ -253,9 +246,7 @@ This helps when the final answer has a stable attractor and errors are diverse. The system expands intermediate states, evaluates partial progress, prunes weak branches, and backtracks. -```text-expand -> evaluate -> select frontier -> continue or backtrack-```+<Steps layout="flow" items={[{title: "Expand"}, {title: "Evaluate"}, {title: "Select frontier"}, {title: "Continue or backtrack"}]} /> This is coordination over thoughts, plans, or artifacts. It requires explicit state and a heuristic good enough to guide search. @@ -263,9 +254,7 @@ This is coordination over thoughts, plans, or artifacts. It requires explicit st Agents expose arguments, counterarguments, and revisions before a final decision. -```text-propose -> critique -> respond -> judge-```+<Steps layout="flow" items={[{title: "Propose"}, {title: "Critique"}, {title: "Respond"}, {title: "Judge"}]} /> Debate is useful when hidden assumptions matter. It fails when agents optimize rhetoric, defer to the strongest voice, or converge before evidence changes. @@ -273,9 +262,7 @@ Debate is useful when hidden assumptions matter. It fails when agents optimize r A central coordinator delegates scoped tasks to workers and keeps authority over the final artifact. -```text-supervisor -> assign -> collect -> merge -> verify-```+<Steps layout="flow" items={[{title: "Supervisor"}, {title: "Assign"}, {title: "Collect"}, {title: "Merge"}, {title: "Verify"}]} /> This is good for work that has clear subdomains. It fails when the supervisor has no real budget, no child trace access, or no merge discipline. @@ -283,9 +270,7 @@ This is good for work that has clear subdomains. It fails when the supervisor ha Control transfers from one agent to another. -```text-triage -> active specialist -> maybe hand off again-```+<Steps layout="flow" items={[{title: "Triage"}, {title: "Active specialist"}, {title: "Maybe hand off again"}]} /> Handoffs are good when the next specialist should own the state and speak directly. They are dangerous when state transfer is implicit or context grows without boundaries. @@ -293,9 +278,7 @@ Handoffs are good when the next specialist should own the state and speak direct Agents write partial results into a shared workspace. Other agents read and improve them. -```text-workers -> shared artifact store -> reviewers -> revised artifact-```+<Steps layout="flow" items={[{title: "Workers"}, {title: "Shared artifact store"}, {title: "Reviewers"}, {title: "Revised artifact"}]} /> This is natural for code, research, planning, and design. It needs locking, provenance, conflict resolution, and traceable authorship. @@ -303,9 +286,7 @@ This is natural for code, research, planning, and design. It needs locking, prov One layer proposes outputs. Later layers aggregate, refine, or route. -```text-proposers -> aggregators -> final selector-```+<Steps layout="flow" items={[{title: "Proposers"}, {title: "Aggregators"}, {title: "Final selector"}]} /> This works when the aggregator can exploit complementary model strengths. It fails when later layers smooth away critical dissent. @@ -355,11 +336,9 @@ The MCP layer is a third shape, not a replacement for either of the first two. ` The separation is important: -```text-runLoop = bounded multi-shot task kernel-conversation = long-horizon participant dialogue-MCP delegation = async specialist work surface-```+- **`runLoop`**: bounded multi-shot task kernel.+- **`conversation`**: long-horizon participant dialogue.+- **MCP delegation**: async specialist work surface. A reliable coordination stack needs the following contracts regardless of which layer hosts them: @@ -402,12 +381,10 @@ In the local `@tangle-network/agent-eval@0.34.1` source, the package is not just For multi-agent coordination, the clean split is: -```text-agent-runtime/conversation decides which participants spoke, under which policy-agent-runtime/loops decides which bounded workers ran-agent-runtime/mcp exposes async specialist delegation-agent-eval decides whether the resulting system was better-```+- agent-runtime/conversation decides which participants spoke, under which policy+- agent-runtime/loops decides which bounded workers ran+- agent-runtime/mcp exposes async specialist delegation+- agent-eval decides whether the resulting system was better Multi-agent coordination without eval is theater. Eval without trace-level runtime evidence is an opinion poll. @@ -455,7 +432,6 @@ Do not ask whether a multi-agent system "feels smarter." Ask whether it beats th Minimum protocol: -```text 1. Define the task distribution and artifact contract. 2. Freeze model set, tools, prompts, skills, dataset, and evaluator where possible. 3. Compare against best single-agent and best-of-N baselines at matched budget.@@ -463,32 +439,23 @@ Minimum protocol: 5. Measure quality, cost, latency, branch failure rate, trace integrity, and human review load. 6. Run ablations: no debate, no shared context, no heterogeneity, deterministic selector. 7. Promote only on held-out lift with acceptable cost, latency, and failure-mode profile.-``` The promotion rule can be written: -```text-promote(s_multi) if:- LCB_95(median(score_multi - score_baseline on holdout)) > epsilon- and median_cost_multi <= cost_ceiling- and median_latency_multi <= latency_ceiling- and trace_integrity == 1- and selector_ablation_delta > 0- and deterministic_failures == 0-```+$$+\begin{aligned}\operatorname{promote}(s_{\text{multi}})&\text{ if: }\operatorname{LCB}_{95}(\operatorname{median}(\text{score}_{\text{multi}}-\text{score}_{\text{baseline}}\text{ on holdout}))>\epsilon\\&\land\operatorname{median\_cost}_{\text{multi}}\le\text{cost\_ceiling}\\&\land\operatorname{median\_latency}_{\text{multi}}\le\text{latency\_ceiling}\\&\land\text{trace\_integrity}=1\\&\land\text{selector\_ablation\_delta}>0\\&\land\text{deterministic\_failures}=0\end{aligned}+$$ The `selector_ablation_delta` term matters. If the multi-agent system still performs the same when the selector is replaced with a trivial rule, the sophisticated coordination may not be doing causal work. For open-ended work, add a disagreement audit: -```text-disagreement_audit:- independent evidence found?- contradictions preserved?- reviewer defects resolved?- final merge cites child lineage?- rejected branches explained?-```+- disagreement_audit:+- independent evidence found?+- contradictions preserved?+- reviewer defects resolved?+- final merge cites child lineage?+- rejected branches explained? Disagreement is useful only when it changes the final decision or improves confidence calibration. @@ -496,43 +463,35 @@ Disagreement is useful only when it changes the final decision or improves confi Prompt optimizers can improve role instructions: -```text-critic prompt-planner prompt-selector rubric-handoff description-reviewer checklist-```+- critic prompt+- planner prompt+- selector rubric+- handoff description+- reviewer checklist Skill optimizers can improve durable role procedure: -```text-how a reviewer inspects a patch-how a researcher triangulates sources-how a coordinator merges conflicting evidence-how an analyst clusters trace failures-```+- how a reviewer inspects a patch+- how a researcher triangulates sources+- how a coordinator merges conflicting evidence+- how an analyst clusters trace failures Runtime topology optimizers can improve execution shape: -```text-fanout width-which roles run in parallel-whether debate happens before or after evidence collection-which selector sees which fields-when branches cancel-how budget is allocated-```+- fanout width+- which roles run in parallel+- whether debate happens before or after evidence collection+- which selector sees which fields+- when branches cancel+- how budget is allocated Harness evolution can change the coordination machine itself: -```text-new driver-new selector implementation-new trace schema-new sandbox isolation model-new promotion gate-```+- new driver+- new selector implementation+- new trace schema+- new sandbox isolation model+- new promotion gate This is where GEPA, MIPRO, SkillOpt, agent-runtime, agent-eval, and meta-harness stop looking like competitors. They operate on different mutable surfaces. @@ -555,13 +514,11 @@ Do not use multiple agents when the only benefit is a richer cast list. The engineering test is simple: -```text-Can the system show why this role existed?-Can it show what information the role had?-Can it show what the role produced?-Can it show how the selector used or rejected that output?-Can it beat a compute-matched single-agent baseline?-```+- Can the system show why this role existed?+- Can it show what information the role had?+- Can it show what the role produced?+- Can it show how the selector used or rejected that output?+- Can it beat a compute-matched single-agent baseline? If not, the coordination is not yet a system property. It is prose. - Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
show diff
diff --git a/src/content/posts/self-improving-stack-multi-agent-coordination.mdx b/src/content/posts/self-improving-stack-multi-agent-coordination.mdxindex dd76fa1..d086bca 100644--- a/src/content/posts/self-improving-stack-multi-agent-coordination.mdx+++ b/src/content/posts/self-improving-stack-multi-agent-coordination.mdx@@ -59,44 +59,46 @@ Coordination is structure. The previous post made runtime topology explicit: -```text-g = executable graph of agents, tools, validators, selectors, handoffs, and gates-pi = runtime policy over graph moves-```+$$+\begin{aligned}+ g &= \text{executable graph of agents, tools, validators, selectors, handoffs, and gates} \\+ \pi &= \text{runtime policy over graph moves}+\end{aligned}+$$ Multi-agent coordination sits one layer above that. It decides what the nodes are supposed to do together. Let: -```text-r = role contracts-p = persona and instruction content-k = active skills per role-u = tools and permissions per role-c = communication and context-sharing policy-sigma = selector or merger policy-v = verifier and judge stack-b = budget allocation policy-tau = termination and escalation policy-```+$$+\begin{aligned}+ r &= \text{role contracts} \\+ p &= \text{persona and instruction content} \\+ k &= \text{active skills per role} \\+ u &= \text{tools and permissions per role} \\+ c &= \text{communication and context-sharing policy} \\+ \sigma &= \text{selector or merger policy} \\+ v &= \text{verifier and judge stack} \\+ b &= \text{budget allocation policy} \\+ \tau &= \text{termination and escalation policy}+\end{aligned}+$$ A multi-agent system candidate is: -```text-s = (g, r, p, k, u, c, sigma, v, b, tau)-```+$$+s=(g,r,p,k,u,c,\sigma,v,b,\tau)+$$ The optimization target is: -```text-J(s | m, h) =- E_{x ~ D}[R(run(m, h, s, x))]- - lambda * E[C(run(m, h, s, x))]-```+$$+J(s\mid m,h)=\mathbb{E}_{x\sim D}[R(\operatorname{run}(m,h,s,x))]-\lambda\,\mathbb{E}[C(\operatorname{run}(m,h,s,x))]+$$ The notation is only useful because it exposes the coordinate system. -If only `p` changes, the optimizer is tuning role descriptions. If `k` changes, it is tuning durable role procedure. If `g`, `sigma`, `c`, `b`, or `tau` change, it is tuning coordination structure.+If only $p$ changes, the optimizer is tuning role descriptions. If $k$ changes, it is tuning durable role procedure. If $g$, $\sigma$, $c$, $b$, or $\tau$ change, it is tuning coordination structure. That distinction prevents a common category error: evaluating a better persona as if it were a better multi-agent system. @@ -179,18 +181,17 @@ If every worker shares the same prompt, same model, same examples, same context, The idealized independent case is easy: -```text-P(at least one success) = 1 - product_i(1 - q_i)-```+$$+\mathbb{P}(\text{at least one success})=1-\prod_i(1-q_i)+$$ -where `q_i` is the probability that worker `i` independently finds a valid answer.+where $q_i$ is the probability that worker $i$ independently finds a valid answer. But multi-agent LLM systems rarely get independence for free. The useful quantity is not worker count. It is error correlation. -```text-coordination_gain =- E[score_multi at budget B] - E[score_best_single at budget B]-```+$$+\operatorname{coordination\_gain}=\mathbb{E}[\operatorname{score\_multi}\text{ at budget }B]-\mathbb{E}[\operatorname{score\_best\_single}\text{ at budget }B]+$$ If the gain disappears at matched budget, the system did not learn coordination. It bought more samples. - Polished with refreshed agent-runtime 0.26.0 and agent-eval 0.34.1 surface review, focused kernel/conversation split, MCP delegation boundary, eval promotion map, and cleaner substrate language.
- let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
show diff
diff --git a/src/content/posts/self-improving-stack-multi-agent-coordination.mdx b/src/content/posts/self-improving-stack-multi-agent-coordination.mdxindex a665c19..28d3a0f 100644--- a/src/content/posts/self-improving-stack-multi-agent-coordination.mdx+++ b/src/content/posts/self-improving-stack-multi-agent-coordination.mdx@@ -19,7 +19,9 @@ authors: date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+ - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+ - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-multi-agent-coordination-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-multi-agent-coordination-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-multi-agent-coordination-review' } - date: 2026-06-05@@ -41,17 +43,17 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -The hard part is not naming the agents.+I can give five agents five names and still have one mind making one mistake. -The hard part is making their disagreement useful.+"Researcher," "critic," "architect," "driver," and "supervisor" are not coordination by themselves. They can be the same model, with the same blind spot, reading the same context, under the same budget, producing five paraphrases of the same failure. The cast list changed. The information structure did not. -A "researcher," "critic," "architect," "driver," and "supervisor" can all be the same model, with the same blind spot, reading the same context, under the same budget, producing five versions of one mistake. That is not a multi-agent system in the meaningful sense. It is a single correlated policy wearing role labels.+Multi-agent work becomes real when disagreement becomes useful. That requires separate contracts, state boundaries, tool permissions, selection rules, budgets, and traces. -Multi-agent coordination becomes real when roles have separate contracts, state boundaries, tool permissions, selection rules, budgets, and traces.+The persona is content. -The persona is content. Coordination is structure.+Coordination is structure. -## The Object Being Optimized+## The Coordination Surface The previous post made runtime topology explicit: @@ -90,7 +92,7 @@ J(s | m, h) = - lambda * E[C(run(m, h, s, x))] ``` -The important part is not the notation. It is the coordinate system.+The notation is only useful because it exposes the coordinate system. If only `p` changes, the optimizer is tuning role descriptions. If `k` changes, it is tuning durable role procedure. If `g`, `sigma`, `c`, `b`, or `tau` change, it is tuning coordination structure. @@ -314,7 +316,7 @@ LangChain and LangGraph documentation describes multi-agent systems as coordinat The shared direction is not "more personas." It is explicit control over routing, handoff, state, context, and observability. -## The Tangle Placement+## Where The Local Stack Fits In the local `@tangle-network/agent-runtime@0.26.0` source, the surface separates three coordination shapes. @@ -442,7 +444,7 @@ The final answer looks good, but the system cannot show which child saw which co These are not edge cases. They are the default unless the coordination structure prevents them. -## Evaluation Protocol+## The Coordination Test Do not ask whether a multi-agent system "feels smarter." Ask whether it beats the right baseline. @@ -531,7 +533,7 @@ This is where GEPA, MIPRO, SkillOpt, agent-runtime, agent-eval, and meta-harness The mistake is asking one optimizer to search a surface it cannot execute. -## Working Rule+## When More Agents Are Worth It Use multiple agents when the work needs at least one of these: - 60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
- Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
- Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
- Research planning pass from a traced session.
- Drafted the multi-agent coordination post with formal role contracts, coordination patterns, disagreement math, Tangle runtime/eval placement, and failure modes.
Comments
PUBLIC_GISCUS_REPO,PUBLIC_GISCUS_REPO_ID,PUBLIC_GISCUS_CATEGORY, andPUBLIC_GISCUS_CATEGORY_IDin.env. See giscus.app to generate the IDs after you enable Discussions on the repo.