Beat Random At Equal Compute First
Why best-of-N, self-consistency, verifier reranking, and compute-matched controls are the baseline for agent topology claims.
The Self Improving Stack series
Browse all 13 posts
- Topology Is The Missing Action Space
- The Gate Is The Optimizer
- Self-Improvement Needs A Safety Case
- When The Harness Has To Evolve
- Memory Is Not Automatically Learning
- Personas Are Content, Coordination Is Structure
- Optimization Theory For Agent Builders
- When The Model Itself Is Mutable
- Prompt Optimization Is Not The Whole Game
- Skills Are Trainable State
- Beat Random At Equal Compute First
- Traces Are The Training Data
- The Self-Improving Stack
I do not trust a multi-agent system until it beats the boring baseline. More agents is not a strategy. It is a cost increase until it beats blind extra compute.
That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity. It spent more budget and mislabeled the budget as architecture.
The first gate is simple:
Beat random at equal compute.
Not beat one greedy sample. Not beat the weakest baseline. Not beat a single run after quietly raising the turn budget. Beat the best simple use of the same budget.
What Test-Time Compute Means
Test-time compute is extra computation spent after the model weights are fixed and the task is known.
It can be spent on:
- longer reasoning
- repeated sampling
- self-consistency
- verifier reranking
- tree search
- iterative refinement
- multi-agent fanout
- tool use
- debate
- retrieval
- code execution
The key is that all of these are inference-time choices. They do not change model weights. They change how much work the system does and how that work is allocated.
Let:
The objective is:
The budget is not one scalar in practice. It is a vector:
- samples
- tokens
- wall clock
- model calls
- tool calls
- sandbox minutes
- human review minutes
- dollars
- risk budget
A fair comparison fixes the relevant parts of , or it reports the tradeoff instead of pretending the strategy itself improved.
The Baseline Ladder
Every topology claim should climb a baseline ladder.
Single sample
The weakest baseline is one ordinary run:
Beating this is not enough. Almost any extra compute can beat one sample on tasks with stochastic failures.
Random@k
Sample candidates from the same model and prompt distribution, then select without additional information or with a fixed production selector:
This asks: what happens if we just buy more attempts?
Pass@k
For verifiable tasks, asks whether any candidate succeeds:
Under independent binary success probability :
That formula is the reason repeated sampling is hard to dismiss. Even a weak model can look strong if the task is verifiable and is large enough.
In code-eval practice, is often estimated from sampled candidates where pass the tests:
where is the binomial coefficient, with when . This estimates the chance that at least one of drawn samples would pass, without pretending the evaluator can deploy the answer key.
But is not a deployable policy. It is an oracle coverage metric. It tells you whether a correct answer existed in the sample set, not whether your system could find it without labels.
Best-of-N with a selector
Now add a selector:
This is production-like only if is available at deployment time. A hidden answer key, private unit tests, or human judge may be useful for measurement. It is not a runtime selector unless the product can actually call it.
Self-consistency
Self-consistency samples multiple reasoning paths and chooses the answer supported by the largest mass:
It is powerful when correct answers are stable attractors and wrong answers are diverse. It is weaker for open-ended artifact quality, where many outputs are plausible and no answer string gets a majority.
Verifier rerank
A verifier or reward model scores candidates:
This is where process reward models, unit tests, static analyzers, rubric judges, and domain verifiers enter. The selector becomes useful only to the extent that correlates with true task success and does not overfit superficial features.
Guided compute
Guided strategies choose where to spend compute:
- expand branch
- refine branch
- spawn another sample
- call a verifier
- stop early
This includes tree search, branch-and-bound, adaptive sampling, debate, tool-assisted search, and agentic fanout. A guided strategy earns its complexity when it beats the best blind or simple selector baseline under the same budget.
Coverage Is Not Selection
Large Language Monkeys is a useful corrective. Repeated sampling can uncover solutions at surprising rates, and coverage often scales smoothly with more attempts.
That does not mean repeated sampling solves deployment.
Coverage asks:
Did any candidate contain a good answer?
Selection asks:
Could the system identify that candidate without oracle labels?
The two curves can be very different.
The gap is selector loss:
A system with high coverage and high selector loss is not ready. It can generate a correct answer somewhere in the pile, but it cannot reliably ship the right artifact.
This matters for code agents. A 300-sample run can produce a passing patch somewhere. The product question is whether the runtime can find the patch, verify it, preserve the diff, reject the dangerous variants, and justify the cost.
Why Greedy Baselines Mislead
A greedy one-shot baseline underestimates what the same model can do when used as a stochastic generator.
This is old news in reasoning research. Self-consistency improved chain-of-thought performance by sampling multiple reasoning paths and marginalizing answers. Tree of Thoughts made branch expansion and evaluation explicit. Process-supervised reward models showed that step-level signals can improve search and reranking. Later test-time scaling work studied how to allocate inference compute more efficiently than naive best-of-N.
The modern reasoning-model era made the same point visible at product scale. OpenAI’s o1 work reported that performance improved with more train-time compute and more test-time “thinking” compute. The method details are not public, but the macro signal is clear: inference compute is now a major scaling axis, not an implementation detail.
The consequence for agent builders is brutal:
- If your agent topology beats greedy but loses to simple repeated sampling,
- you built an expensive sampler.
That can still be useful. A sampler with a good selector may be exactly what the product needs. But it should be named honestly.
Parallel Versus Sequential Compute
Extra compute has shapes.
- Parallel sampling: sample k independent attempts, then select.
- Sequential refinement: attempt, critique, and revise; repeat as needed.
- Tree search: expand the frontier, score partial states, and allocate a next step.
- Debate: propose, criticize, respond, then judge.
- Tool-grounded search: form a hypothesis, call a tool, observe, then update.
None dominates everywhere.
Parallel sampling works when the proposal distribution has enough mass on valid answers and selection is reliable. Sequential refinement works when feedback gives useful local gradients. Tree search works when partial states can be evaluated before the final answer. Debate works when hidden assumptions can be exposed by adversarial pressure. Tool-grounded search works when the environment returns discriminative evidence.
The compute-optimal question is:
For this task, model, verifier, and budget, which allocation has highest expected utility?
Snell et al. studied this directly for reasoning problems and found that compute-optimal allocation can be much more efficient than plain best-of-N. The important lesson is not one fixed strategy. It is conditional allocation: choose breadth, depth, refinement, or search based on prompt difficulty and verifier behavior.
That makes test-time compute a control problem. The controller observes partial evidence:
- prompt difficulty estimate
- candidate confidence
- verifier margin
- branch diversity
- remaining budget
- latency deadline
and chooses the next action:
sample | refine | verify | expand | stop
The policy is only good if those observations predict downstream reward. If the difficulty estimate is bad, adaptive compute becomes random budget jitter. If the verifier margin is miscalibrated, early stopping exits on fluent failures.
The Verifier Bottleneck
Test-time compute shifts pressure onto selection.
If the verifier is strong, repeated sampling becomes powerful. If the verifier is weak, more candidates can make things worse because the selector gets more chances to choose a fluent failure.
A verifier can be:
- exact unit tests
- typechecks
- proof checkers
- simulation
- retrieval-grounded fact checks
- process reward models
- outcome reward models
- LLM judges
- human reviewers
Each has a failure profile.
Unit tests can be incomplete. Typechecks miss semantics. Proof checkers only cover formalized claims. LLM judges can reward style, verbosity, or rubric mimicry. Human reviewers are expensive and inconsistent. PRMs can overfit the distribution of reasoning steps they were trained on.
Verifier quality belongs in the objective:
If guided search optimizes faster than tracks , the system reward-hacks its own evaluator.
This is why the selector must be evaluated, not assumed.
Agent Topologies As Compute Allocators
A multi-agent topology is a policy for spending test-time compute.
The same budget can be spent as:
- 8 independent workers
- 4 workers + 1 verifier
- 2 workers + 2 rounds of critique
- 1 worker + 7 refinement turns
- 1 tree search with frontier size 4 and depth 2
- 1 tool-heavy agent with expensive environment checks
The topology claim is not:
The topology claim is:
at a measured budget .
That is why the previous post insisted that role names are not enough. The role structure matters only if it improves allocation, evidence, selection, or verification under budget.
Where The Local Stack Fits
The refreshed Tangle runtime map makes this concrete.
@tangle-network/agent-runtime/loops gives a clean test-time compute substrate:
runLoop: kernel withmaxIterations,maxConcurrency, abort propagation, cost aggregation, and trace events.createFanoutVoteDriver: spend compute on parallel attempts.createRefineDriver: spend compute on sequential retry and validation.Driver: custom allocation policy.Validator: selector evidence.LoopTraceEvent: branch, dispatch, decision, and cost observability.
@tangle-network/agent-runtime/conversation gives the long-horizon version:
maxTurns: hard speaker-turn cap.maxCreditsCents: hard cost cap.turnOrder: allocation across participants.haltOn: early stopping.ConversationJournal: resumable transcript.- deterministic
turnId: retry and trace correlation. - forwarded depth headers: recursion bound.
@tangle-network/agent-eval@0.34.1 supplies the evidence layer:
AgentProfileCell: records the model, prompt, tool, skill, runtime, and harness cell.runEvalCampaign: compares variants and scenarios with capture integrity.HeldOutGate: promotes only when held-out lift survives threshold and cost ceiling.- scorecards and release confidence: track accuracy, cost, latency, overfit gap, and failure modes.
AnalystRegistry: analyzes trace failure modes, knowledge gaps, knowledge poisoning, and improvements.
This is the useful split:
- runtime spends compute
- eval proves whether the spend was worth it
The Equal-Compute Test
A serious test-time compute eval starts with a budget table.
budget:
samples: k
max_tokens_in: ...
max_tokens_out: ...
max_model_calls: ...
max_tool_calls: ...
max_wall_ms: ...
max_cost_usd: ...
max_turns: ...
max_concurrency: ...
Then compare strategies:
- Single greedy or default run.
- Random@k under the same model, prompt, and budget.
- Best-of-N with the deployable selector.
- Self-consistency where answer aggregation is meaningful.
- Verifier-rerank with the production verifier.
- Guided topology under the same budget.
- Adaptive topology with early stopping and budget reallocation.
Record:
- score
- pass_rate
- coverage_k
- selection_k
- selector_loss_k
- cost_usd
- tokens_in_out
- wall_ms
- model_calls
- tool_calls
- branch_failures
- trace_integrity
Also report dominance, not just mean score:
A strategy that improves score while increasing cost is not wrong. It is on a tradeoff frontier. A strategy that is worse and more expensive is dead.
Promotion should require:
The baseline should be the strongest simple strategy the product could actually deploy, not a strawman.
How More Compute Lies
Test-time compute fails in predictable ways.
Unmatched compute
The candidate wins because it used more samples, turns, tokens, or tools.
Oracle selection
The paper or eval reports , but the product has no selector that can find the passing candidate.
Verifier overfit
The strategy optimizes the judge’s quirks faster than it improves the real artifact.
Retry theater
The system repeats the same failure mode and counts each retry as effort.
Hidden branching
The final answer says it explored alternatives, but the trace shows one branch.
Early-stop bias
The strategy stops quickly on easy tasks and spends heavily on hard tasks, but the report only shows mean score without cost distribution.
Tool-cost laundering
The token budget is matched, but one strategy uses much more sandbox time, browser time, API calls, or human review.
Coverage bragging
The system reports that one candidate succeeded somewhere in the batch but cannot select it reliably.
When More Compute Is Evidence
Do not evaluate an agent topology against one sample.
Evaluate it against the best boring way to spend the same budget.
Use test-time compute when:
- the task has stochastic failures
- the verifier is strong enough to select
- the domain rewards breadth or search
- the product can afford higher latency or cost
- traces can prove where the extra compute went
Do not use test-time compute when:
- the verifier is weak
- the answer is easy enough for one sample
- latency dominates quality
- the same failure repeats across attempts
- the strategy cannot beat random at equal budget
The serious claim is not “we used more reasoning.” It is:
Given the same budget, this allocation policy produced better verified outcomes.
That is the first bar for runtime topology, multi-agent coordination, and self-improving harnesses.
Source Trail
Source freshness checked on 2026-06-06.
- Evaluating Large Language Models Trained on Code, checked June 6, 2026.
- Self-Consistency Improves Chain of Thought Reasoning in Language Models, checked June 6, 2026.
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models, checked June 6, 2026.
- Let’s Verify Step by Step, checked June 6, 2026.
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling, checked June 6, 2026.
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, checked June 6, 2026.
- Learning to reason with LLMs, checked June 6, 2026.
- Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs, checked June 6, 2026.
- Local
@tangle-network/agent-runtime@0.26.0source audit:runLoop, drivers, conversation policy, journals, MCP delegation, June 6, 2026. - Local
@tangle-network/agent-eval@0.34.1source audit:AgentProfileCell,runEvalCampaign,HeldOutGate, scorecards, release confidence,AnalystRegistry, June 6, 2026.
Revision history9revisions
- Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
show diff
diff --git a/src/content/posts/self-improving-stack-test-time-compute.mdx b/src/content/posts/self-improving-stack-test-time-compute.mdxindex 0da57f8..61d1947 100644--- a/src/content/posts/self-improving-stack-test-time-compute.mdx+++ b/src/content/posts/self-improving-stack-test-time-compute.mdx@@ -47,15 +47,16 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'++ I do not trust a multi-agent system until it beats the boring baseline. More agents is not a strategy. It is a cost increase until it beats blind extra compute. That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity. It spent more budget and mislabeled the budget as architecture. The first gate is simple: -```text-Beat random at equal compute.-```+> Beat random at equal compute. Not beat one greedy sample. Not beat the weakest baseline. Not beat a single run after quietly raising the turn budget. Beat the best simple use of the same budget. @@ -101,19 +102,15 @@ $$ The budget $B$ is not one scalar in practice. It is a vector: -```text-B = {- samples,- tokens,- wall_clock,- model_calls,- tool_calls,- sandbox_minutes,- human_review_minutes,- dollars,- risk_budget-}-```+- samples+- tokens+- wall clock+- model calls+- tool calls+- sandbox minutes+- human review minutes+- dollars+- risk budget A fair comparison fixes the relevant parts of $B$, or it reports the tradeoff instead of pretending the strategy itself improved. @@ -214,13 +211,11 @@ This is where process reward models, unit tests, static analyzers, rubric judges Guided strategies choose where to spend compute: -```text-expand branch-refine branch-spawn another sample-call a verifier-stop early-```+- expand branch+- refine branch+- spawn another sample+- call a verifier+- stop early This includes tree search, branch-and-bound, adaptive sampling, debate, tool-assisted search, and agentic fanout. A guided strategy earns its complexity when it beats the best blind or simple selector baseline under the same budget. @@ -232,15 +227,11 @@ That does not mean repeated sampling solves deployment. Coverage asks: -```text-Did any candidate contain a good answer?-```+> Did any candidate contain a good answer? Selection asks: -```text-Could the system identify that candidate without oracle labels?-```+> Could the system identify that candidate without oracle labels? The two curves can be very different. @@ -271,10 +262,8 @@ The modern reasoning-model era made the same point visible at product scale. Ope The consequence for agent builders is brutal: -```text-If your agent topology beats greedy but loses to simple repeated sampling,-you built an expensive sampler.-```+- If your agent topology beats greedy but loses to simple repeated sampling,+- you built an expensive sampler. That can still be useful. A sampler with a good selector may be exactly what the product needs. But it should be named honestly. @@ -284,33 +273,23 @@ Extra compute has shapes. Parallel sampling: -```text-sample k independent attempts -> select-```+<Steps layout="flow" items={[{title: "Sample k independent attempts"}, {title: "Select"}]} /> Sequential refinement: -```text-attempt -> critique -> revise -> critique -> revise-```+<Steps layout="flow" items={[{title: "Attempt"}, {title: "Critique"}, {title: "Revise"}, {title: "Critique"}, {title: "Revise"}]} /> Tree search: -```text-expand frontier -> score partial states -> allocate next step-```+<Steps layout="flow" items={[{title: "Expand frontier"}, {title: "Score partial states"}, {title: "Allocate next step"}]} /> Debate: -```text-proposal -> criticism -> response -> judge-```+<Steps layout="flow" items={[{title: "Propose"}, {title: "Criticism"}, {title: "Respond"}, {title: "Judge"}]} /> Tool-grounded search: -```text-hypothesis -> tool call -> observation -> update-```+<Steps layout="flow" items={[{title: "Hypothesis"}, {title: "Tool call"}, {title: "Observation"}, {title: "Update"}]} /> None dominates everywhere. @@ -318,28 +297,22 @@ Parallel sampling works when the proposal distribution has enough mass on valid The compute-optimal question is: -```text For this task, model, verifier, and budget, which allocation has highest expected utility?-``` Snell et al. studied this directly for reasoning problems and found that compute-optimal allocation can be much more efficient than plain best-of-N. The important lesson is not one fixed strategy. It is conditional allocation: choose breadth, depth, refinement, or search based on prompt difficulty and verifier behavior. That makes test-time compute a control problem. The controller observes partial evidence: -```text-prompt difficulty estimate-candidate confidence-verifier margin-branch diversity-remaining budget-latency deadline-```+- prompt difficulty estimate+- candidate confidence+- verifier margin+- branch diversity+- remaining budget+- latency deadline and chooses the next action: -```text sample | refine | verify | expand | stop-``` The policy is only good if those observations predict downstream reward. If the difficulty estimate is bad, adaptive compute becomes random budget jitter. If the verifier margin is miscalibrated, early stopping exits on fluent failures. @@ -367,11 +340,9 @@ Unit tests can be incomplete. Typechecks miss semantics. Proof checkers only cov Verifier quality belongs in the objective: -```text-observed_score = V(x, y, trace)-true_score = R(x, y)-verifier_error = observed_score - true_score-```+$$+\begin{aligned}\text{observed\_score}&=V(x,y,\text{trace})\\\text{true\_score}&=R(x,y)\\\text{verifier\_error}&=\text{observed\_score}-\text{true\_score}\end{aligned}+$$ If guided search optimizes $V$ faster than $V$ tracks $R$, the system reward-hacks its own evaluator. @@ -383,26 +354,24 @@ A multi-agent topology is a policy for spending test-time compute. The same budget can be spent as: -```text-8 independent workers-4 workers + 1 verifier-2 workers + 2 rounds of critique-1 worker + 7 refinement turns-1 tree search with frontier size 4 and depth 2-1 tool-heavy agent with expensive environment checks-```+- 8 independent workers+- 4 workers + 1 verifier+- 2 workers + 2 rounds of critique+- 1 worker + 7 refinement turns+- 1 tree search with frontier size 4 and depth 2+- 1 tool-heavy agent with expensive environment checks The topology claim is not: -```text-multi-agent > single-agent-```+$$+\text{multi-agent}>\text{single-agent}+$$ The topology claim is: -```text-allocation_policy_multi(B) > allocation_policy_baseline(B)-```+$$+\operatorname{allocation\_policy}_{\text{multi}}(B)>\operatorname{allocation\_policy}_{\text{baseline}}(B)+$$ at a measured budget $B$. @@ -441,10 +410,8 @@ The refreshed Tangle runtime map makes this concrete. This is the useful split: -```text-runtime spends compute-eval proves whether the spend was worth it-```+- runtime spends compute+- eval proves whether the spend was worth it ## The Equal-Compute Test @@ -465,7 +432,6 @@ budget: Then compare strategies: -```text 1. Single greedy or default run. 2. Random@k under the same model, prompt, and budget. 3. Best-of-N with the deployable selector.@@ -473,48 +439,35 @@ Then compare strategies: 5. Verifier-rerank with the production verifier. 6. Guided topology under the same budget. 7. Adaptive topology with early stopping and budget reallocation.-``` Record: -```text-score-pass_rate-coverage_k-selection_k-selector_loss_k-cost_usd-tokens_in_out-wall_ms-model_calls-tool_calls-branch_failures-trace_integrity-```+- score+- pass_rate+- coverage_k+- selection_k+- selector_loss_k+- cost_usd+- tokens_in_out+- wall_ms+- model_calls+- tool_calls+- branch_failures+- trace_integrity Also report dominance, not just mean score: -```text-strategy_a dominates strategy_b if:- score_a >= score_b- and cost_vector_a <= cost_vector_b componentwise- and at least one inequality is strict-```+$$+\begin{aligned}\text{strategy\_a dominates strategy\_b if: }\text{score\_a}\ge\text{score\_b}\\&\land\text{cost\_vector\_a}\le\text{cost\_vector\_b componentwise}\\&\land\text{at least one inequality is strict}\end{aligned}+$$ A strategy that improves score while increasing cost is not wrong. It is on a tradeoff frontier. A strategy that is worse and more expensive is dead. Promotion should require: -```text-promote(strategy_new) if:- LCB_95(median(score_new - score_baseline on holdout)) > epsilon- and median_cost_new <= cost_ceiling- and median_latency_new <= latency_ceiling- and no baseline Pareto-dominates strategy_new- and trace_integrity == 1- and selector_loss_new <= selector_loss_ceiling- and deterministic_failures == 0-```+$$+\begin{aligned}\operatorname{promote}(\text{strategy\_new})&\text{ if: }\operatorname{LCB}_{95}(\operatorname{median}(\text{score\_new}-\text{score\_baseline on holdout}))>\epsilon\\&\land\text{median\_cost\_new}\le\text{cost\_ceiling}\\&\land\text{median\_latency\_new}\le\text{latency\_ceiling}\\&\land\text{no baseline Pareto-dominates strategy\_new}\\&\land\text{trace\_integrity}=1\\&\land\text{selector\_loss\_new}\le\text{selector\_loss\_ceiling}\\&\land\text{deterministic\_failures}=0\end{aligned}+$$ The baseline should be the strongest simple strategy the product could actually deploy, not a strawman. @@ -578,9 +531,7 @@ Do not use test-time compute when: The serious claim is not "we used more reasoning." It is: -```text-Given the same budget, this allocation policy produced better verified outcomes.-```+> Given the same budget, this allocation policy produced better verified outcomes. That is the first bar for runtime topology, multi-agent coordination, and self-improving harnesses. - Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
show diff
diff --git a/src/content/posts/self-improving-stack-test-time-compute.mdx b/src/content/posts/self-improving-stack-test-time-compute.mdxindex c9bd3e9..bff211f 100644--- a/src/content/posts/self-improving-stack-test-time-compute.mdx+++ b/src/content/posts/self-improving-stack-test-time-compute.mdx@@ -79,25 +79,25 @@ The key is that all of these are inference-time choices. They do not change mode Let: -```text-x = task-m = fixed model or model set-h = harness and runtime-s = strategy for spending test-time compute-B = compute budget-R = reward, score, or task success-C = measured cost-```+$$+\begin{aligned}+ x &= \text{task} \\+ m &= \text{fixed model or model set} \\+ h &= \text{harness and runtime} \\+ s &= \text{strategy for spending test-time compute} \\+ B &= \text{compute budget} \\+ R &= \text{reward, score, or task success} \\+ C &= \text{measured cost}+\end{aligned}+$$ The objective is: -```text-J(s | m, h, B) =- E_{x ~ D}[R(run(m, h, s, x, B))]- - lambda * E[C(run(m, h, s, x, B))]-```+$$+J(s\mid m,h,B)=\mathbb{E}_{x\sim D}[R(\operatorname{run}(m,h,s,x,B))]-\lambda\,\mathbb{E}[C(\operatorname{run}(m,h,s,x,B))]+$$ -The budget `B` is not one scalar in practice. It is a vector:+The budget $B$ is not one scalar in practice. It is a vector: ```text B = {@@ -113,7 +113,7 @@ B = { } ``` -A fair comparison fixes the relevant parts of `B`, or it reports the tradeoff instead of pretending the strategy itself improved.+A fair comparison fixes the relevant parts of $B$, or it reports the tradeoff instead of pretending the strategy itself improved. ## The Baseline Ladder @@ -123,70 +123,78 @@ Every topology claim should climb a baseline ladder. The weakest baseline is one ordinary run: -```text-y_1 ~ q_m(y | x)-score = R(y_1)-```+$$+\begin{aligned}+ y_1&\sim q_m(y\mid x) \\+ \text{score}&=R(y_1)+\end{aligned}+$$ Beating this is not enough. Almost any extra compute can beat one sample on tasks with stochastic failures. **Random@k** -Sample `k` candidates from the same model and prompt distribution, then select without additional information or with a fixed production selector:+Sample $k$ candidates from the same model and prompt distribution, then select without additional information or with a fixed production selector: -```text-y_i ~ q_m(y | x), i = 1..k-y_hat = sigma_blind({y_i})-```+$$+\begin{aligned}+ y_i&\sim q_m(y\mid x),\quad i=1,\ldots,k \\+ \hat{y}&=\sigma_{\text{blind}}(\{y_i\})+\end{aligned}+$$ This asks: what happens if we just buy more attempts? **Pass@k** -For verifiable tasks, `pass@k` asks whether any candidate succeeds:+For verifiable tasks, $\operatorname{pass@k}$ asks whether any candidate succeeds: -```text-pass@k = P(max_i R(y_i) = 1)-```+$$+\operatorname{pass@k}=\mathbb{P}\!\left(\max_iR(y_i)=1\right)+$$ -Under independent binary success probability `q`:+Under independent binary success probability $q$: -```text-pass@k = 1 - (1 - q)^k-```+$$+\operatorname{pass@k}=1-(1-q)^k+$$ -That formula is the reason repeated sampling is hard to dismiss. Even a weak model can look strong if the task is verifiable and `k` is large enough.+That formula is the reason repeated sampling is hard to dismiss. Even a weak model can look strong if the task is verifiable and $k$ is large enough. -In code-eval practice, `pass@k` is often estimated from `n` sampled candidates where `c` pass the tests:+In code-eval practice, $\operatorname{pass@k}$ is often estimated from $n$ sampled candidates where $c$ pass the tests: -```text-pass_hat@k = 1 - C(n - c, k) / C(n, k)-```+$$+\widehat{\operatorname{pass@k}}=1-\frac{\binom{n-c}{k}}{\binom{n}{k}}+$$ -where `C(a, b)` is the binomial coefficient, with `C(a, b) = 0` when `a < b`. This estimates the chance that at least one of `k` drawn samples would pass, without pretending the evaluator can deploy the answer key.+where $\binom{a}{b}$ is the binomial coefficient, with $\binom{a}{b}=0$ when $a<b$. This estimates the chance that at least one of $k$ drawn samples would pass, without pretending the evaluator can deploy the answer key. -But `pass@k` is not a deployable policy. It is an oracle coverage metric. It tells you whether a correct answer existed in the sample set, not whether your system could find it without labels.+But $\operatorname{pass@k}$ is not a deployable policy. It is an oracle coverage metric. It tells you whether a correct answer existed in the sample set, not whether your system could find it without labels. **Best-of-N with a selector** Now add a selector: -```text-y_hat = sigma(x, {y_1, ..., y_k})-score = R(y_hat)-```+$$+\begin{aligned}+ \hat{y}&=\sigma(x,\{y_1,\ldots,y_k\}) \\+ \text{score}&=R(\hat{y})+\end{aligned}+$$ -This is production-like only if `sigma` is available at deployment time. A hidden answer key, private unit tests, or human judge may be useful for measurement. It is not a runtime selector unless the product can actually call it.+This is production-like only if $\sigma$ is available at deployment time. A hidden answer key, private unit tests, or human judge may be useful for measurement. It is not a runtime selector unless the product can actually call it. **Self-consistency** Self-consistency samples multiple reasoning paths and chooses the answer supported by the largest mass: -```text-z_i = reasoning path-a_i = final answer extracted from z_i-y_hat = argmax_a count(a_i = a)-```+$$+\begin{aligned}+ z_i&=\text{reasoning path} \\+ a_i&=\text{final answer extracted from }z_i \\+ \hat{y}&=\operatorname*{argmax}_a\operatorname{count}(a_i=a)+\end{aligned}+$$ It is powerful when correct answers are stable attractors and wrong answers are diverse. It is weaker for open-ended artifact quality, where many outputs are plausible and no answer string gets a majority. @@ -194,11 +202,11 @@ It is powerful when correct answers are stable attractors and wrong answers are A verifier or reward model scores candidates: -```text-y_hat = argmax_i V(x, y_i, trace_i)-```+$$+\hat{y}=\operatorname*{argmax}_i V(x,y_i,\operatorname{trace}_i)+$$ -This is where process reward models, unit tests, static analyzers, rubric judges, and domain verifiers enter. The selector becomes useful only to the extent that `V` correlates with true task success and does not overfit superficial features.+This is where process reward models, unit tests, static analyzers, rubric judges, and domain verifiers enter. The selector becomes useful only to the extent that $V$ correlates with true task success and does not overfit superficial features. **Guided compute** @@ -234,16 +242,18 @@ Could the system identify that candidate without oracle labels? The two curves can be very different. -```text-coverage_k = P(exists i: R(y_i) = 1)-selection_k = E[R(sigma({y_i}))]-```+$$+\begin{aligned}+ \operatorname{coverage}_k&=\mathbb{P}(\exists i:R(y_i)=1) \\+ \operatorname{selection}_k&=\mathbb{E}[R(\sigma(\{y_i\}))]+\end{aligned}+$$ The gap is selector loss: -```text-selector_loss_k = coverage_k - selection_k-```+$$+\operatorname{selector\_loss}_k=\operatorname{coverage}_k-\operatorname{selection}_k+$$ A system with high coverage and high selector loss is not ready. It can generate a correct answer somewhere in the pile, but it cannot reliably ship the right artifact. @@ -361,7 +371,7 @@ true_score = R(x, y) verifier_error = observed_score - true_score ``` -If guided search optimizes `V` faster than `V` tracks `R`, the system reward-hacks its own evaluator.+If guided search optimizes $V$ faster than $V$ tracks $R$, the system reward-hacks its own evaluator. This is why the selector must be evaluated, not assumed. @@ -392,7 +402,7 @@ The topology claim is: allocation_policy_multi(B) > allocation_policy_baseline(B) ``` -at a measured budget `B`.+at a measured budget $B$. That is why the previous post insisted that role names are not enough. The role structure matters only if it improves allocation, evidence, selection, or verification under budget. @@ -516,7 +526,7 @@ The candidate wins because it used more samples, turns, tokens, or tools. **Oracle selection** -The paper or eval reports `pass@k`, but the product has no selector that can find the passing candidate.+The paper or eval reports $\operatorname{pass@k}$, but the product has no selector that can find the passing candidate. **Verifier overfit** - Drafted the test-time compute post with compute-matched baselines, selection math, verifier limits, adaptive allocation, and Tangle runtime/eval placement.
- Polished the test-time compute post with the finite-sample pass@k estimator, adaptive compute as a control problem, Pareto dominance language, and stricter promotion criteria.
- let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
show diff
diff --git a/src/content/posts/self-improving-stack-test-time-compute.mdx b/src/content/posts/self-improving-stack-test-time-compute.mdxindex 2ec4d75..a716b11 100644--- a/src/content/posts/self-improving-stack-test-time-compute.mdx+++ b/src/content/posts/self-improving-stack-test-time-compute.mdx@@ -19,7 +19,9 @@ authors: date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+ - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+ - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-test-time-compute-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-test-time-compute-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-test-time-compute-review' } - date: 2026-06-05@@ -41,11 +43,9 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -More agents is not a strategy.+I do not trust a multi-agent system until it beats the boring baseline. More agents is not a strategy. It is a cost increase until it beats blind extra compute. -It is a cost increase until it beats blind extra compute.--That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity.+That is the baseline every agent topology has to face. If a supervisor, debate loop, reflection loop, or specialist fanout wins only because it spent more samples, more tokens, more wall-clock, or more tool calls, the structure has not yet earned its complexity. It spent more budget and mislabeled the budget as architecture. The first gate is simple: @@ -394,7 +394,7 @@ at a measured budget `B`. That is why the previous post insisted that role names are not enough. The role structure matters only if it improves allocation, evidence, selection, or verification under budget. -## Tangle Placement+## Where The Local Stack Fits The refreshed Tangle runtime map makes this concrete. @@ -432,7 +432,7 @@ runtime spends compute eval proves whether the spend was worth it ``` -## Evaluation Protocol+## The Equal-Compute Test A serious test-time compute eval starts with a budget table. @@ -504,7 +504,7 @@ promote(strategy_new) if: The baseline should be the strongest simple strategy the product could actually deploy, not a strawman. -## Failure Modes+## How More Compute Lies Test-time compute fails in predictable ways. @@ -540,7 +540,7 @@ The token budget is matched, but one strategy uses much more sandbox time, brows The system reports that one candidate succeeded somewhere in the batch but cannot select it reliably. -## Working Rule+## When More Compute Is Evidence Do not evaluate an agent topology against one sample. - 60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
- Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
- Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
- Research planning pass from a traced session.
Comments
PUBLIC_GISCUS_REPO,PUBLIC_GISCUS_REPO_ID,PUBLIC_GISCUS_CATEGORY, andPUBLIC_GISCUS_CATEGORY_IDin.env. See giscus.app to generate the IDs after you enable Discussions on the repo.