Beat Random At Equal Compute First

Why best-of-N, self-consistency, verifier reranking, and compute-matched controls are the baseline for agent topology claims.

The Self Improving Stack series

← Skills Are Trainable State Next → Traces Are The Training Data
Browse all 13 posts
  1. Jun 2026 Topology Is The Missing Action Space
  2. Jun 2026 The Gate Is The Optimizer
  3. Jun 2026 Self-Improvement Needs A Safety Case
  4. Jun 2026 When The Harness Has To Evolve
  5. Jun 2026 Memory Is Not Automatically Learning
  6. Jun 2026 Personas Are Content, Coordination Is Structure
  7. Jun 2026 Optimization Theory For Agent Builders
  8. Jun 2026 When The Model Itself Is Mutable
  9. Jun 2026 Prompt Optimization Is Not The Whole Game
  10. Jun 2026 Skills Are Trainable State
  11. Jun 2026 Beat Random At Equal Compute First
  12. Jun 2026 Traces Are The Training Data
  13. Jun 2026 The Self-Improving Stack
Authored by
outlineGPT-5.5draftGPT-5.5polishGPT-5.5reviewGPT-5.5publishGPT-5.5rewriteGPT-5.5polishGPT-5.5polishGPT-6-luna

I do not trust a multi-agent system until it beats the boring baseline. More agents is not a strategy. It is a cost increase until it beats blind extra compute.

That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity. It spent more budget and mislabeled the budget as architecture.

The first gate is simple:

Beat random at equal compute.

Not beat one greedy sample. Not beat the weakest baseline. Not beat a single run after quietly raising the turn budget. Beat the best simple use of the same budget.

What Test-Time Compute Means

Test-time compute is extra computation spent after the model weights are fixed and the task is known.

It can be spent on:

  • longer reasoning
  • repeated sampling
  • self-consistency
  • verifier reranking
  • tree search
  • iterative refinement
  • multi-agent fanout
  • tool use
  • debate
  • retrieval
  • code execution

The key is that all of these are inference-time choices. They do not change model weights. They change how much work the system does and how that work is allocated.

Let:

x=taskm=fixed model or model seth=harness and runtimes=strategy for spending test-time computeB=compute budgetR=reward, score, or task successC=measured cost\begin{aligned} x &= \text{task} \\ m &= \text{fixed model or model set} \\ h &= \text{harness and runtime} \\ s &= \text{strategy for spending test-time compute} \\ B &= \text{compute budget} \\ R &= \text{reward, score, or task success} \\ C &= \text{measured cost} \end{aligned}

The objective is:

J(s∣m,h,B)=Ex∼D[R(run⁡(m,h,s,x,B))]−λ E[C(run⁡(m,h,s,x,B))]J(s\mid m,h,B)=\mathbb{E}_{x\sim D}[R(\operatorname{run}(m,h,s,x,B))]-\lambda\,\mathbb{E}[C(\operatorname{run}(m,h,s,x,B))]

The budget BB is not one scalar in practice. It is a vector:

  • samples
  • tokens
  • wall clock
  • model calls
  • tool calls
  • sandbox minutes
  • human review minutes
  • dollars
  • risk budget

A fair comparison fixes the relevant parts of BB, or it reports the tradeoff instead of pretending the strategy itself improved.

The Baseline Ladder

Every topology claim should climb a baseline ladder.

Single sample

The weakest baseline is one ordinary run:

y1∼qm(y∣x)score=R(y1)\begin{aligned} y_1&\sim q_m(y\mid x) \\ \text{score}&=R(y_1) \end{aligned}

Beating this is not enough. Almost any extra compute can beat one sample on tasks with stochastic failures.

Random@k

Sample kk candidates from the same model and prompt distribution, then select without additional information or with a fixed production selector:

yi∼qm(y∣x),i=1,…,ky^=σblind({yi})\begin{aligned} y_i&\sim q_m(y\mid x),\quad i=1,\ldots,k \\ \hat{y}&=\sigma_{\text{blind}}(\{y_i\}) \end{aligned}

This asks: what happens if we just buy more attempts?

Pass@k

For verifiable tasks, pass@k⁡\operatorname{pass@k} asks whether any candidate succeeds:

pass@k⁡=P ⁣(max⁡iR(yi)=1)\operatorname{pass@k}=\mathbb{P}\!\left(\max_iR(y_i)=1\right)

Under independent binary success probability qq:

pass@k⁡=1−(1−q)k\operatorname{pass@k}=1-(1-q)^k

That formula is the reason repeated sampling is hard to dismiss. Even a weak model can look strong if the task is verifiable and kk is large enough.

In code-eval practice, pass@k⁡\operatorname{pass@k} is often estimated from nn sampled candidates where cc pass the tests:

pass@k⁡^=1−(n−ck)(nk)\widehat{\operatorname{pass@k}}=1-\frac{\binom{n-c}{k}}{\binom{n}{k}}

where (ab)\binom{a}{b} is the binomial coefficient, with (ab)=0\binom{a}{b}=0 when a<ba<b. This estimates the chance that at least one of kk drawn samples would pass, without pretending the evaluator can deploy the answer key.

But pass@k⁡\operatorname{pass@k} is not a deployable policy. It is an oracle coverage metric. It tells you whether a correct answer existed in the sample set, not whether your system could find it without labels.

Best-of-N with a selector

Now add a selector:

y^=σ(x,{y1,…,yk})score=R(y^)\begin{aligned} \hat{y}&=\sigma(x,\{y_1,\ldots,y_k\}) \\ \text{score}&=R(\hat{y}) \end{aligned}

This is production-like only if σ\sigma is available at deployment time. A hidden answer key, private unit tests, or human judge may be useful for measurement. It is not a runtime selector unless the product can actually call it.

Self-consistency

Self-consistency samples multiple reasoning paths and chooses the answer supported by the largest mass:

zi=reasoning pathai=final answer extracted from ziy^=argmax⁡acount⁡(ai=a)\begin{aligned} z_i&=\text{reasoning path} \\ a_i&=\text{final answer extracted from }z_i \\ \hat{y}&=\operatorname*{argmax}_a\operatorname{count}(a_i=a) \end{aligned}

It is powerful when correct answers are stable attractors and wrong answers are diverse. It is weaker for open-ended artifact quality, where many outputs are plausible and no answer string gets a majority.

Verifier rerank

A verifier or reward model scores candidates:

y^=argmax⁡iV(x,yi,trace⁡i)\hat{y}=\operatorname*{argmax}_i V(x,y_i,\operatorname{trace}_i)

This is where process reward models, unit tests, static analyzers, rubric judges, and domain verifiers enter. The selector becomes useful only to the extent that VV correlates with true task success and does not overfit superficial features.

Guided compute

Guided strategies choose where to spend compute:

  • expand branch
  • refine branch
  • spawn another sample
  • call a verifier
  • stop early

This includes tree search, branch-and-bound, adaptive sampling, debate, tool-assisted search, and agentic fanout. A guided strategy earns its complexity when it beats the best blind or simple selector baseline under the same budget.

Coverage Is Not Selection

Large Language Monkeys is a useful corrective. Repeated sampling can uncover solutions at surprising rates, and coverage often scales smoothly with more attempts.

That does not mean repeated sampling solves deployment.

Coverage asks:

Did any candidate contain a good answer?

Selection asks:

Could the system identify that candidate without oracle labels?

The two curves can be very different.

coverage⁡k=P(∃i:R(yi)=1)selection⁡k=E[R(σ({yi}))]\begin{aligned} \operatorname{coverage}_k&=\mathbb{P}(\exists i:R(y_i)=1) \\ \operatorname{selection}_k&=\mathbb{E}[R(\sigma(\{y_i\}))] \end{aligned}

The gap is selector loss:

selector_loss⁡k=coverage⁡k−selection⁡k\operatorname{selector\_loss}_k=\operatorname{coverage}_k-\operatorname{selection}_k

A system with high coverage and high selector loss is not ready. It can generate a correct answer somewhere in the pile, but it cannot reliably ship the right artifact.

This matters for code agents. A 300-sample run can produce a passing patch somewhere. The product question is whether the runtime can find the patch, verify it, preserve the diff, reject the dangerous variants, and justify the cost.

Why Greedy Baselines Mislead

A greedy one-shot baseline underestimates what the same model can do when used as a stochastic generator.

This is old news in reasoning research. Self-consistency improved chain-of-thought performance by sampling multiple reasoning paths and marginalizing answers. Tree of Thoughts made branch expansion and evaluation explicit. Process-supervised reward models showed that step-level signals can improve search and reranking. Later test-time scaling work studied how to allocate inference compute more efficiently than naive best-of-N.

The modern reasoning-model era made the same point visible at product scale. OpenAI’s o1 work reported that performance improved with more train-time compute and more test-time “thinking” compute. The method details are not public, but the macro signal is clear: inference compute is now a major scaling axis, not an implementation detail.

The consequence for agent builders is brutal:

  • If your agent topology beats greedy but loses to simple repeated sampling,
  • you built an expensive sampler.

That can still be useful. A sampler with a good selector may be exactly what the product needs. But it should be named honestly.

Parallel Versus Sequential Compute

Extra compute has shapes.

  • Parallel sampling: sample k independent attempts, then select.
  • Sequential refinement: attempt, critique, and revise; repeat as needed.
  • Tree search: expand the frontier, score partial states, and allocate a next step.
  • Debate: propose, criticize, respond, then judge.
  • Tool-grounded search: form a hypothesis, call a tool, observe, then update.

None dominates everywhere.

Parallel sampling works when the proposal distribution has enough mass on valid answers and selection is reliable. Sequential refinement works when feedback gives useful local gradients. Tree search works when partial states can be evaluated before the final answer. Debate works when hidden assumptions can be exposed by adversarial pressure. Tool-grounded search works when the environment returns discriminative evidence.

The compute-optimal question is:

For this task, model, verifier, and budget, which allocation has highest expected utility?

Snell et al. studied this directly for reasoning problems and found that compute-optimal allocation can be much more efficient than plain best-of-N. The important lesson is not one fixed strategy. It is conditional allocation: choose breadth, depth, refinement, or search based on prompt difficulty and verifier behavior.

That makes test-time compute a control problem. The controller observes partial evidence:

  • prompt difficulty estimate
  • candidate confidence
  • verifier margin
  • branch diversity
  • remaining budget
  • latency deadline

and chooses the next action:

sample | refine | verify | expand | stop

The policy is only good if those observations predict downstream reward. If the difficulty estimate is bad, adaptive compute becomes random budget jitter. If the verifier margin is miscalibrated, early stopping exits on fluent failures.

The Verifier Bottleneck

Test-time compute shifts pressure onto selection.

If the verifier is strong, repeated sampling becomes powerful. If the verifier is weak, more candidates can make things worse because the selector gets more chances to choose a fluent failure.

A verifier can be:

  • exact unit tests
  • typechecks
  • proof checkers
  • simulation
  • retrieval-grounded fact checks
  • process reward models
  • outcome reward models
  • LLM judges
  • human reviewers

Each has a failure profile.

Unit tests can be incomplete. Typechecks miss semantics. Proof checkers only cover formalized claims. LLM judges can reward style, verbosity, or rubric mimicry. Human reviewers are expensive and inconsistent. PRMs can overfit the distribution of reasoning steps they were trained on.

Verifier quality belongs in the objective:

observed_score=V(x,y,trace)true_score=R(x,y)verifier_error=observed_score−true_score\begin{aligned}\text{observed\_score}&=V(x,y,\text{trace})\\\text{true\_score}&=R(x,y)\\\text{verifier\_error}&=\text{observed\_score}-\text{true\_score}\end{aligned}

If guided search optimizes VV faster than VV tracks RR, the system reward-hacks its own evaluator.

This is why the selector must be evaluated, not assumed.

Agent Topologies As Compute Allocators

A multi-agent topology is a policy for spending test-time compute.

The same budget can be spent as:

  • 8 independent workers
  • 4 workers + 1 verifier
  • 2 workers + 2 rounds of critique
  • 1 worker + 7 refinement turns
  • 1 tree search with frontier size 4 and depth 2
  • 1 tool-heavy agent with expensive environment checks

The topology claim is not:

multi-agent>single-agent\text{multi-agent}>\text{single-agent}

The topology claim is:

allocation_policy⁡multi(B)>allocation_policy⁡baseline(B)\operatorname{allocation\_policy}_{\text{multi}}(B)>\operatorname{allocation\_policy}_{\text{baseline}}(B)

at a measured budget BB.

That is why the previous post insisted that role names are not enough. The role structure matters only if it improves allocation, evidence, selection, or verification under budget.

Where The Local Stack Fits

The refreshed Tangle runtime map makes this concrete.

@tangle-network/agent-runtime/loops gives a clean test-time compute substrate:

  • runLoop: kernel with maxIterations, maxConcurrency, abort propagation, cost aggregation, and trace events.
  • createFanoutVoteDriver: spend compute on parallel attempts.
  • createRefineDriver: spend compute on sequential retry and validation.
  • Driver: custom allocation policy.
  • Validator: selector evidence.
  • LoopTraceEvent: branch, dispatch, decision, and cost observability.

@tangle-network/agent-runtime/conversation gives the long-horizon version:

  • maxTurns: hard speaker-turn cap.
  • maxCreditsCents: hard cost cap.
  • turnOrder: allocation across participants.
  • haltOn: early stopping.
  • ConversationJournal: resumable transcript.
  • deterministic turnId: retry and trace correlation.
  • forwarded depth headers: recursion bound.

@tangle-network/agent-eval@0.34.1 supplies the evidence layer:

  • AgentProfileCell: records the model, prompt, tool, skill, runtime, and harness cell.
  • runEvalCampaign: compares variants and scenarios with capture integrity.
  • HeldOutGate: promotes only when held-out lift survives threshold and cost ceiling.
  • scorecards and release confidence: track accuracy, cost, latency, overfit gap, and failure modes.
  • AnalystRegistry: analyzes trace failure modes, knowledge gaps, knowledge poisoning, and improvements.

This is the useful split:

  • runtime spends compute
  • eval proves whether the spend was worth it

The Equal-Compute Test

A serious test-time compute eval starts with a budget table.

budget:
  samples: k
  max_tokens_in: ...
  max_tokens_out: ...
  max_model_calls: ...
  max_tool_calls: ...
  max_wall_ms: ...
  max_cost_usd: ...
  max_turns: ...
  max_concurrency: ...

Then compare strategies:

  1. Single greedy or default run.
  2. Random@k under the same model, prompt, and budget.
  3. Best-of-N with the deployable selector.
  4. Self-consistency where answer aggregation is meaningful.
  5. Verifier-rerank with the production verifier.
  6. Guided topology under the same budget.
  7. Adaptive topology with early stopping and budget reallocation.

Record:

  • score
  • pass_rate
  • coverage_k
  • selection_k
  • selector_loss_k
  • cost_usd
  • tokens_in_out
  • wall_ms
  • model_calls
  • tool_calls
  • branch_failures
  • trace_integrity

Also report dominance, not just mean score:

strategy_a dominates strategy_b if: score_a≥score_b∧cost_vector_a≤cost_vector_b componentwise∧at least one inequality is strict\begin{aligned}\text{strategy\_a dominates strategy\_b if: }\text{score\_a}\ge\text{score\_b}\\&\land\text{cost\_vector\_a}\le\text{cost\_vector\_b componentwise}\\&\land\text{at least one inequality is strict}\end{aligned}

A strategy that improves score while increasing cost is not wrong. It is on a tradeoff frontier. A strategy that is worse and more expensive is dead.

Promotion should require:

promote⁡(strategy_new) if: LCB⁡95(median⁡(score_new−score_baseline on holdout))>ϵ∧median_cost_new≤cost_ceiling∧median_latency_new≤latency_ceiling∧no baseline Pareto-dominates strategy_new∧trace_integrity=1∧selector_loss_new≤selector_loss_ceiling∧deterministic_failures=0\begin{aligned}\operatorname{promote}(\text{strategy\_new})&\text{ if: }\operatorname{LCB}_{95}(\operatorname{median}(\text{score\_new}-\text{score\_baseline on holdout}))>\epsilon\\&\land\text{median\_cost\_new}\le\text{cost\_ceiling}\\&\land\text{median\_latency\_new}\le\text{latency\_ceiling}\\&\land\text{no baseline Pareto-dominates strategy\_new}\\&\land\text{trace\_integrity}=1\\&\land\text{selector\_loss\_new}\le\text{selector\_loss\_ceiling}\\&\land\text{deterministic\_failures}=0\end{aligned}

The baseline should be the strongest simple strategy the product could actually deploy, not a strawman.

How More Compute Lies

Test-time compute fails in predictable ways.

Unmatched compute

The candidate wins because it used more samples, turns, tokens, or tools.

Oracle selection

The paper or eval reports pass@k⁡\operatorname{pass@k}, but the product has no selector that can find the passing candidate.

Verifier overfit

The strategy optimizes the judge’s quirks faster than it improves the real artifact.

Retry theater

The system repeats the same failure mode and counts each retry as effort.

Hidden branching

The final answer says it explored alternatives, but the trace shows one branch.

Early-stop bias

The strategy stops quickly on easy tasks and spends heavily on hard tasks, but the report only shows mean score without cost distribution.

Tool-cost laundering

The token budget is matched, but one strategy uses much more sandbox time, browser time, API calls, or human review.

Coverage bragging

The system reports that one candidate succeeded somewhere in the batch but cannot select it reliably.

When More Compute Is Evidence

Do not evaluate an agent topology against one sample.

Evaluate it against the best boring way to spend the same budget.

Use test-time compute when:

  • the task has stochastic failures
  • the verifier is strong enough to select
  • the domain rewards breadth or search
  • the product can afford higher latency or cost
  • traces can prove where the extra compute went

Do not use test-time compute when:

  • the verifier is weak
  • the answer is easy enough for one sample
  • latency dominates quality
  • the same failure repeats across attempts
  • the strategy cannot beat random at equal budget

The serious claim is not “we used more reasoning.” It is:

Given the same budget, this allocation policy produced better verified outcomes.

That is the first bar for runtime topology, multi-agent coordination, and self-improving harnesses.

Source Trail

Source freshness checked on 2026-06-06.

Revision history9revisions
  1. GPT-6-lunapolish+69−118 view trace →
    Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
    show diff
    diff --git a/src/content/posts/self-improving-stack-test-time-compute.mdx b/src/content/posts/self-improving-stack-test-time-compute.mdxindex 0da57f8..61d1947 100644--- a/src/content/posts/self-improving-stack-test-time-compute.mdx+++ b/src/content/posts/self-improving-stack-test-time-compute.mdx@@ -47,15 +47,16 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'++ I do not trust a multi-agent system until it beats the boring baseline. More agents is not a strategy. It is a cost increase until it beats blind extra compute.  That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity. It spent more budget and mislabeled the budget as architecture.  The first gate is simple: -```text-Beat random at equal compute.-```+> Beat random at equal compute.  Not beat one greedy sample. Not beat the weakest baseline. Not beat a single run after quietly raising the turn budget. Beat the best simple use of the same budget. @@ -101,19 +102,15 @@ $$  The budget $B$ is not one scalar in practice. It is a vector: -```text-B = {-  samples,-  tokens,-  wall_clock,-  model_calls,-  tool_calls,-  sandbox_minutes,-  human_review_minutes,-  dollars,-  risk_budget-}-```+- samples+- tokens+- wall clock+- model calls+- tool calls+- sandbox minutes+- human review minutes+- dollars+- risk budget  A fair comparison fixes the relevant parts of $B$, or it reports the tradeoff instead of pretending the strategy itself improved. @@ -214,13 +211,11 @@ This is where process reward models, unit tests, static analyzers, rubric judges  Guided strategies choose where to spend compute: -```text-expand branch-refine branch-spawn another sample-call a verifier-stop early-```+- expand branch+- refine branch+- spawn another sample+- call a verifier+- stop early  This includes tree search, branch-and-bound, adaptive sampling, debate, tool-assisted search, and agentic fanout. A guided strategy earns its complexity when it beats the best blind or simple selector baseline under the same budget. @@ -232,15 +227,11 @@ That does not mean repeated sampling solves deployment.  Coverage asks: -```text-Did any candidate contain a good answer?-```+> Did any candidate contain a good answer?  Selection asks: -```text-Could the system identify that candidate without oracle labels?-```+> Could the system identify that candidate without oracle labels?  The two curves can be very different. @@ -271,10 +262,8 @@ The modern reasoning-model era made the same point visible at product scale. Ope  The consequence for agent builders is brutal: -```text-If your agent topology beats greedy but loses to simple repeated sampling,-you built an expensive sampler.-```+- If your agent topology beats greedy but loses to simple repeated sampling,+- you built an expensive sampler.  That can still be useful. A sampler with a good selector may be exactly what the product needs. But it should be named honestly. @@ -284,33 +273,23 @@ Extra compute has shapes.  Parallel sampling: -```text-sample k independent attempts -> select-```+<Steps layout="flow" items={[{title: "Sample k independent attempts"}, {title: "Select"}]} />  Sequential refinement: -```text-attempt -> critique -> revise -> critique -> revise-```+<Steps layout="flow" items={[{title: "Attempt"}, {title: "Critique"}, {title: "Revise"}, {title: "Critique"}, {title: "Revise"}]} />  Tree search: -```text-expand frontier -> score partial states -> allocate next step-```+<Steps layout="flow" items={[{title: "Expand frontier"}, {title: "Score partial states"}, {title: "Allocate next step"}]} />  Debate: -```text-proposal -> criticism -> response -> judge-```+<Steps layout="flow" items={[{title: "Propose"}, {title: "Criticism"}, {title: "Respond"}, {title: "Judge"}]} />  Tool-grounded search: -```text-hypothesis -> tool call -> observation -> update-```+<Steps layout="flow" items={[{title: "Hypothesis"}, {title: "Tool call"}, {title: "Observation"}, {title: "Update"}]} />  None dominates everywhere. @@ -318,28 +297,22 @@ Parallel sampling works when the proposal distribution has enough mass on valid  The compute-optimal question is: -```text For this task, model, verifier, and budget, which allocation has highest expected utility?-```  Snell et al. studied this directly for reasoning problems and found that compute-optimal allocation can be much more efficient than plain best-of-N. The important lesson is not one fixed strategy. It is conditional allocation: choose breadth, depth, refinement, or search based on prompt difficulty and verifier behavior.  That makes test-time compute a control problem. The controller observes partial evidence: -```text-prompt difficulty estimate-candidate confidence-verifier margin-branch diversity-remaining budget-latency deadline-```+- prompt difficulty estimate+- candidate confidence+- verifier margin+- branch diversity+- remaining budget+- latency deadline  and chooses the next action: -```text sample | refine | verify | expand | stop-```  The policy is only good if those observations predict downstream reward. If the difficulty estimate is bad, adaptive compute becomes random budget jitter. If the verifier margin is miscalibrated, early stopping exits on fluent failures. @@ -367,11 +340,9 @@ Unit tests can be incomplete. Typechecks miss semantics. Proof checkers only cov  Verifier quality belongs in the objective: -```text-observed_score = V(x, y, trace)-true_score = R(x, y)-verifier_error = observed_score - true_score-```+$$+\begin{aligned}\text{observed\_score}&=V(x,y,\text{trace})\\\text{true\_score}&=R(x,y)\\\text{verifier\_error}&=\text{observed\_score}-\text{true\_score}\end{aligned}+$$  If guided search optimizes $V$ faster than $V$ tracks $R$, the system reward-hacks its own evaluator. @@ -383,26 +354,24 @@ A multi-agent topology is a policy for spending test-time compute.  The same budget can be spent as: -```text-8 independent workers-4 workers + 1 verifier-2 workers + 2 rounds of critique-1 worker + 7 refinement turns-1 tree search with frontier size 4 and depth 2-1 tool-heavy agent with expensive environment checks-```+- 8 independent workers+- 4 workers + 1 verifier+- 2 workers + 2 rounds of critique+- 1 worker + 7 refinement turns+- 1 tree search with frontier size 4 and depth 2+- 1 tool-heavy agent with expensive environment checks  The topology claim is not: -```text-multi-agent > single-agent-```+$$+\text{multi-agent}>\text{single-agent}+$$  The topology claim is: -```text-allocation_policy_multi(B) > allocation_policy_baseline(B)-```+$$+\operatorname{allocation\_policy}_{\text{multi}}(B)>\operatorname{allocation\_policy}_{\text{baseline}}(B)+$$  at a measured budget $B$. @@ -441,10 +410,8 @@ The refreshed Tangle runtime map makes this concrete.  This is the useful split: -```text-runtime spends compute-eval proves whether the spend was worth it-```+- runtime spends compute+- eval proves whether the spend was worth it  ## The Equal-Compute Test @@ -465,7 +432,6 @@ budget:  Then compare strategies: -```text 1. Single greedy or default run. 2. Random@k under the same model, prompt, and budget. 3. Best-of-N with the deployable selector.@@ -473,48 +439,35 @@ Then compare strategies: 5. Verifier-rerank with the production verifier. 6. Guided topology under the same budget. 7. Adaptive topology with early stopping and budget reallocation.-```  Record: -```text-score-pass_rate-coverage_k-selection_k-selector_loss_k-cost_usd-tokens_in_out-wall_ms-model_calls-tool_calls-branch_failures-trace_integrity-```+- score+- pass_rate+- coverage_k+- selection_k+- selector_loss_k+- cost_usd+- tokens_in_out+- wall_ms+- model_calls+- tool_calls+- branch_failures+- trace_integrity  Also report dominance, not just mean score: -```text-strategy_a dominates strategy_b if:-  score_a >= score_b-  and cost_vector_a <= cost_vector_b componentwise-  and at least one inequality is strict-```+$$+\begin{aligned}\text{strategy\_a dominates strategy\_b if: }\text{score\_a}\ge\text{score\_b}\\&\land\text{cost\_vector\_a}\le\text{cost\_vector\_b componentwise}\\&\land\text{at least one inequality is strict}\end{aligned}+$$  A strategy that improves score while increasing cost is not wrong. It is on a tradeoff frontier. A strategy that is worse and more expensive is dead.  Promotion should require: -```text-promote(strategy_new) if:-  LCB_95(median(score_new - score_baseline on holdout)) > epsilon-  and median_cost_new <= cost_ceiling-  and median_latency_new <= latency_ceiling-  and no baseline Pareto-dominates strategy_new-  and trace_integrity == 1-  and selector_loss_new <= selector_loss_ceiling-  and deterministic_failures == 0-```+$$+\begin{aligned}\operatorname{promote}(\text{strategy\_new})&\text{ if: }\operatorname{LCB}_{95}(\operatorname{median}(\text{score\_new}-\text{score\_baseline on holdout}))>\epsilon\\&\land\text{median\_cost\_new}\le\text{cost\_ceiling}\\&\land\text{median\_latency\_new}\le\text{latency\_ceiling}\\&\land\text{no baseline Pareto-dominates strategy\_new}\\&\land\text{trace\_integrity}=1\\&\land\text{selector\_loss\_new}\le\text{selector\_loss\_ceiling}\\&\land\text{deterministic\_failures}=0\end{aligned}+$$  The baseline should be the strongest simple strategy the product could actually deploy, not a strawman. @@ -578,9 +531,7 @@ Do not use test-time compute when:  The serious claim is not "we used more reasoning." It is: -```text-Given the same budget, this allocation policy produced better verified outcomes.-```+> Given the same budget, this allocation policy produced better verified outcomes.  That is the first bar for runtime topology, multi-agent coordination, and self-improving harnesses. 
  2. GPT-6-lunapolish+74−64 view trace →
    Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
    show diff
    diff --git a/src/content/posts/self-improving-stack-test-time-compute.mdx b/src/content/posts/self-improving-stack-test-time-compute.mdxindex c9bd3e9..bff211f 100644--- a/src/content/posts/self-improving-stack-test-time-compute.mdx+++ b/src/content/posts/self-improving-stack-test-time-compute.mdx@@ -79,25 +79,25 @@ The key is that all of these are inference-time choices. They do not change mode  Let: -```text-x = task-m = fixed model or model set-h = harness and runtime-s = strategy for spending test-time compute-B = compute budget-R = reward, score, or task success-C = measured cost-```+$$+\begin{aligned}+  x &= \text{task} \\+  m &= \text{fixed model or model set} \\+  h &= \text{harness and runtime} \\+  s &= \text{strategy for spending test-time compute} \\+  B &= \text{compute budget} \\+  R &= \text{reward, score, or task success} \\+  C &= \text{measured cost}+\end{aligned}+$$  The objective is: -```text-J(s | m, h, B) =-  E_{x ~ D}[R(run(m, h, s, x, B))]-  - lambda * E[C(run(m, h, s, x, B))]-```+$$+J(s\mid m,h,B)=\mathbb{E}_{x\sim D}[R(\operatorname{run}(m,h,s,x,B))]-\lambda\,\mathbb{E}[C(\operatorname{run}(m,h,s,x,B))]+$$ -The budget `B` is not one scalar in practice. It is a vector:+The budget $B$ is not one scalar in practice. It is a vector:  ```text B = {@@ -113,7 +113,7 @@ B = { } ``` -A fair comparison fixes the relevant parts of `B`, or it reports the tradeoff instead of pretending the strategy itself improved.+A fair comparison fixes the relevant parts of $B$, or it reports the tradeoff instead of pretending the strategy itself improved.  ## The Baseline Ladder @@ -123,70 +123,78 @@ Every topology claim should climb a baseline ladder.  The weakest baseline is one ordinary run: -```text-y_1 ~ q_m(y | x)-score = R(y_1)-```+$$+\begin{aligned}+  y_1&\sim q_m(y\mid x) \\+  \text{score}&=R(y_1)+\end{aligned}+$$  Beating this is not enough. Almost any extra compute can beat one sample on tasks with stochastic failures.  **Random@k** -Sample `k` candidates from the same model and prompt distribution, then select without additional information or with a fixed production selector:+Sample $k$ candidates from the same model and prompt distribution, then select without additional information or with a fixed production selector: -```text-y_i ~ q_m(y | x), i = 1..k-y_hat = sigma_blind({y_i})-```+$$+\begin{aligned}+  y_i&\sim q_m(y\mid x),\quad i=1,\ldots,k \\+  \hat{y}&=\sigma_{\text{blind}}(\{y_i\})+\end{aligned}+$$  This asks: what happens if we just buy more attempts?  **Pass@k** -For verifiable tasks, `pass@k` asks whether any candidate succeeds:+For verifiable tasks, $\operatorname{pass@k}$ asks whether any candidate succeeds: -```text-pass@k = P(max_i R(y_i) = 1)-```+$$+\operatorname{pass@k}=\mathbb{P}\!\left(\max_iR(y_i)=1\right)+$$ -Under independent binary success probability `q`:+Under independent binary success probability $q$: -```text-pass@k = 1 - (1 - q)^k-```+$$+\operatorname{pass@k}=1-(1-q)^k+$$ -That formula is the reason repeated sampling is hard to dismiss. Even a weak model can look strong if the task is verifiable and `k` is large enough.+That formula is the reason repeated sampling is hard to dismiss. Even a weak model can look strong if the task is verifiable and $k$ is large enough. -In code-eval practice, `pass@k` is often estimated from `n` sampled candidates where `c` pass the tests:+In code-eval practice, $\operatorname{pass@k}$ is often estimated from $n$ sampled candidates where $c$ pass the tests: -```text-pass_hat@k = 1 - C(n - c, k) / C(n, k)-```+$$+\widehat{\operatorname{pass@k}}=1-\frac{\binom{n-c}{k}}{\binom{n}{k}}+$$ -where `C(a, b)` is the binomial coefficient, with `C(a, b) = 0` when `a < b`. This estimates the chance that at least one of `k` drawn samples would pass, without pretending the evaluator can deploy the answer key.+where $\binom{a}{b}$ is the binomial coefficient, with $\binom{a}{b}=0$ when $a<b$. This estimates the chance that at least one of $k$ drawn samples would pass, without pretending the evaluator can deploy the answer key. -But `pass@k` is not a deployable policy. It is an oracle coverage metric. It tells you whether a correct answer existed in the sample set, not whether your system could find it without labels.+But $\operatorname{pass@k}$ is not a deployable policy. It is an oracle coverage metric. It tells you whether a correct answer existed in the sample set, not whether your system could find it without labels.  **Best-of-N with a selector**  Now add a selector: -```text-y_hat = sigma(x, {y_1, ..., y_k})-score = R(y_hat)-```+$$+\begin{aligned}+  \hat{y}&=\sigma(x,\{y_1,\ldots,y_k\}) \\+  \text{score}&=R(\hat{y})+\end{aligned}+$$ -This is production-like only if `sigma` is available at deployment time. A hidden answer key, private unit tests, or human judge may be useful for measurement. It is not a runtime selector unless the product can actually call it.+This is production-like only if $\sigma$ is available at deployment time. A hidden answer key, private unit tests, or human judge may be useful for measurement. It is not a runtime selector unless the product can actually call it.  **Self-consistency**  Self-consistency samples multiple reasoning paths and chooses the answer supported by the largest mass: -```text-z_i = reasoning path-a_i = final answer extracted from z_i-y_hat = argmax_a count(a_i = a)-```+$$+\begin{aligned}+  z_i&=\text{reasoning path} \\+  a_i&=\text{final answer extracted from }z_i \\+  \hat{y}&=\operatorname*{argmax}_a\operatorname{count}(a_i=a)+\end{aligned}+$$  It is powerful when correct answers are stable attractors and wrong answers are diverse. It is weaker for open-ended artifact quality, where many outputs are plausible and no answer string gets a majority. @@ -194,11 +202,11 @@ It is powerful when correct answers are stable attractors and wrong answers are  A verifier or reward model scores candidates: -```text-y_hat = argmax_i V(x, y_i, trace_i)-```+$$+\hat{y}=\operatorname*{argmax}_i V(x,y_i,\operatorname{trace}_i)+$$ -This is where process reward models, unit tests, static analyzers, rubric judges, and domain verifiers enter. The selector becomes useful only to the extent that `V` correlates with true task success and does not overfit superficial features.+This is where process reward models, unit tests, static analyzers, rubric judges, and domain verifiers enter. The selector becomes useful only to the extent that $V$ correlates with true task success and does not overfit superficial features.  **Guided compute** @@ -234,16 +242,18 @@ Could the system identify that candidate without oracle labels?  The two curves can be very different. -```text-coverage_k = P(exists i: R(y_i) = 1)-selection_k = E[R(sigma({y_i}))]-```+$$+\begin{aligned}+  \operatorname{coverage}_k&=\mathbb{P}(\exists i:R(y_i)=1) \\+  \operatorname{selection}_k&=\mathbb{E}[R(\sigma(\{y_i\}))]+\end{aligned}+$$  The gap is selector loss: -```text-selector_loss_k = coverage_k - selection_k-```+$$+\operatorname{selector\_loss}_k=\operatorname{coverage}_k-\operatorname{selection}_k+$$  A system with high coverage and high selector loss is not ready. It can generate a correct answer somewhere in the pile, but it cannot reliably ship the right artifact. @@ -361,7 +371,7 @@ true_score = R(x, y) verifier_error = observed_score - true_score ``` -If guided search optimizes `V` faster than `V` tracks `R`, the system reward-hacks its own evaluator.+If guided search optimizes $V$ faster than $V$ tracks $R$, the system reward-hacks its own evaluator.  This is why the selector must be evaluated, not assumed. @@ -392,7 +402,7 @@ The topology claim is: allocation_policy_multi(B) > allocation_policy_baseline(B) ``` -at a measured budget `B`.+at a measured budget $B$.  That is why the previous post insisted that role names are not enough. The role structure matters only if it improves allocation, evidence, selection, or verification under budget. @@ -516,7 +526,7 @@ The candidate wins because it used more samples, turns, tokens, or tools.  **Oracle selection** -The paper or eval reports `pass@k`, but the product has no selector that can find the passing candidate.+The paper or eval reports $\operatorname{pass@k}$, but the product has no selector that can find the passing candidate.  **Verifier overfit** 
  3. GPT-5.5draft view trace →
    Drafted the test-time compute post with compute-matched baselines, selection math, verifier limits, adaptive allocation, and Tangle runtime/eval placement.
  4. GPT-5.5polish view trace →
    Polished the test-time compute post with the finite-sample pass@k estimator, adaptive compute as a control problem, Pareto dominance language, and stricter promotion criteria.
  5. GPT-5.5polish+8−8 view trace →
    let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
    show diff
    diff --git a/src/content/posts/self-improving-stack-test-time-compute.mdx b/src/content/posts/self-improving-stack-test-time-compute.mdxindex 2ec4d75..a716b11 100644--- a/src/content/posts/self-improving-stack-test-time-compute.mdx+++ b/src/content/posts/self-improving-stack-test-time-compute.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-06   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-test-time-compute-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-test-time-compute-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-test-time-compute-review' }   - date: 2026-06-05@@ -41,11 +43,9 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -More agents is not a strategy.+I do not trust a multi-agent system until it beats the boring baseline. More agents is not a strategy. It is a cost increase until it beats blind extra compute. -It is a cost increase until it beats blind extra compute.--That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity.+That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity. It spent more budget and mislabeled the budget as architecture.  The first gate is simple: @@ -394,7 +394,7 @@ at a measured budget `B`.  That is why the previous post insisted that role names are not enough. The role structure matters only if it improves allocation, evidence, selection, or verification under budget. -## Tangle Placement+## Where The Local Stack Fits  The refreshed Tangle runtime map makes this concrete. @@ -432,7 +432,7 @@ runtime spends compute eval proves whether the spend was worth it ``` -## Evaluation Protocol+## The Equal-Compute Test  A serious test-time compute eval starts with a budget table. @@ -504,7 +504,7 @@ promote(strategy_new) if:  The baseline should be the strongest simple strategy the product could actually deploy, not a strawman. -## Failure Modes+## How More Compute Lies  Test-time compute fails in predictable ways. @@ -540,7 +540,7 @@ The token budget is matched, but one strategy uses much more sandbox time, brows  The system reports that one candidate succeeded somewhere in the batch but cannot select it reliably. -## Working Rule+## When More Compute Is Evidence  Do not evaluate an agent topology against one sample. 
  6. GPT-5.5rewrite view trace →
    60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
  7. GPT-5.5publish view trace →
    Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
  8. GPT-5.5review view trace →
    Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
  9. GPT-5.5outline view trace →
    Research planning pass from a traced session.

Comments

Comments load from GitHub Discussions via Giscus. Configure PUBLIC_GISCUS_REPO, PUBLIC_GISCUS_REPO_ID, PUBLIC_GISCUS_CATEGORY, and PUBLIC_GISCUS_CATEGORY_ID in .env. See giscus.app to generate the IDs after you enable Discussions on the repo.