The Gate Is The Optimizer

Why held-out promotion, judge reliability, failure taxonomies, cost ceilings, and confidence intervals decide whether self-improvement is real.

The Self Improving Stack series

← Topology Is The Missing Action Space Next → Self-Improvement Needs A Safety Case
Browse all 13 posts
  1. Jun 2026 Topology Is The Missing Action Space
  2. Jun 2026 The Gate Is The Optimizer
  3. Jun 2026 Self-Improvement Needs A Safety Case
  4. Jun 2026 When The Harness Has To Evolve
  5. Jun 2026 Memory Is Not Automatically Learning
  6. Jun 2026 Personas Are Content, Coordination Is Structure
  7. Jun 2026 Optimization Theory For Agent Builders
  8. Jun 2026 When The Model Itself Is Mutable
  9. Jun 2026 Prompt Optimization Is Not The Whole Game
  10. Jun 2026 Skills Are Trainable State
  11. Jun 2026 Beat Random At Equal Compute First
  12. Jun 2026 Traces Are The Training Data
  13. Jun 2026 The Self-Improving Stack
Authored by
outlineGPT-5.5draftGPT-5.5polishGPT-5.5reviewGPT-5.5publishGPT-5.5rewriteGPT-5.5polishGPT-5.5polishGPT-6-luna

A green score is not a release decision. It is evidence entering a release policy, and in self-improving loops that distinction matters more than the optimizer.

GEPA, MIPRO, SkillOpt, topology search, and meta-harness can all generate candidates forever. The gate decides which candidate becomes the system future agents inherit. If the gate is weak, the optimizer learns the gate. If the gate is honest, the optimizer has to improve the product.

This is why the gate is not an administrative detail after the interesting work. It is the objective boundary.

What A Gate Is

A gate is a promotion policy.

Let:

b=baseline systemc=candidate systemx=scenariop=agent profile cellz=seed or replicate idR=task reward or scoreC=measured cost vectorT=trace integrity predicateDsearch=search splitDholdout=held-out split\begin{aligned} b &= \text{baseline system} \\ c &= \text{candidate system} \\ x &= \text{scenario} \\ p &= \text{agent profile cell} \\ z &= \text{seed or replicate id} \\ R &= \text{task reward or score} \\ C &= \text{measured cost vector} \\ T &= \text{trace integrity predicate} \\ D_{\text{search}} &= \text{search split} \\ D_{\text{holdout}} &= \text{held-out split} \end{aligned}

The gate is a function:

G(c, b, D_holdout, p, z) -> {promote, reject}

It is allowed to inspect evidence. It is not allowed to move the goalpost after seeing the candidate.

The simplest version is:

promote(c) iff
  quality(c, holdout) > quality(b, holdout)
  and cost(c) <= cost_ceiling
  and latency(c) <= latency_ceiling
  and deterministic_failures(c) = 0
  and trace_integrity(c) = 1

That version is readable, but still too loose. A real gate needs paired observations, uncertainty, split discipline, judge reliability, and regression protection.

The Gate Is Not A Leaderboard

A leaderboard asks:

Which system had the highest score?

A gate asks:

Should this candidate replace this baseline for this product?

Those are different questions.

Leaderboards are useful for orientation. They are bad promotion policies. They compress context, cost, latency, tool availability, profile differences, data leakage, and failure severity into one rank.

Agent systems make this worse because candidates can improve the metric while damaging the workflow:

  • better aggregate score, worse high-value persona
  • better judge score, worse deterministic verifier
  • better single-shot result, worse cost
  • better easy tasks, worse hard tasks
  • better search split, worse holdout
  • better answer style, worse intent match
  • better visible output, broken trace capture

The gate has to preserve the baseline unless the candidate earns replacement.

Paired Evidence

Unpaired averages are fragile.

Suppose the baseline sees one sample of tasks and the candidate sees another. A mean difference can be task mix, not improvement.

The paired comparison fixes that:

Δi=R(c,xi,pi,zi)−R(b,xi,pi,zi)\Delta_i = R(c,x_i,p_i,z_i)-R(b,x_i,p_i,z_i)

where the candidate and baseline are evaluated on the same scenario, profile, and replicate. Then the question becomes:

Is median(delta_i) reliably positive on held-out items?

The median is useful because agent scores often have heavy tails. One catastrophic failure or one lucky success can distort a mean. The median asks whether the typical paired task improved.

A practical promotion rule:

npairs≥nmin⁡,LCB⁡95(median⁡(Δ))>ϵn_{\text{pairs}} \ge n_{\min},\qquad \operatorname{LCB}_{95}(\operatorname{median}(\Delta)) > \epsilon

LCB⁡95\operatorname{LCB}_{95} is the lower confidence bound. If the lower bound clears the threshold, the gate has evidence that the lift is not just random luck.

This is where bootstrap confidence intervals are useful. You resample paired deltas, compute the median for each resample, and inspect the lower quantile:

delta = [delta_1, ..., delta_n]
for r in 1..B:
  sample n deltas with replacement
  m_r = median(sample)
LCB⁡95=Q0.025(m1,…,mB)\operatorname{LCB}_{95}=Q_{0.025}(m_1,\ldots,m_B)

The gate promotes only when the pessimistic estimate is still good enough.

Search Split Versus Holdout

Optimizers need data to search.

Gates need data the optimizer did not tune against.

That gives two split families:

  • DsearchD_{\text{search}}: used to propose, mutate, rank, debug, and iterate.
  • DholdoutD_{\text{holdout}}: used to decide promotion.

The failure mode is:

  • Search score increases.
  • Holdout score decreases.

So the gate needs an overfit check:

gap⁡c=mean_score⁡(c,search)−mean_score⁡(c,holdout)gap⁡b=mean_score⁡(b,search)−mean_score⁡(b,holdout)\begin{aligned} \operatorname{gap}_c &= \operatorname{mean\_score}(c,\text{search})-\operatorname{mean\_score}(c,\text{holdout}) \\ \operatorname{gap}_b &= \operatorname{mean\_score}(b,\text{search})-\operatorname{mean\_score}(b,\text{holdout}) \end{aligned}
reject(c) if gap_c > gap_b + tau

This says the candidate may look better on the search split, but it cannot be much more search-specialized than the baseline. τ\tau is slack, not forgiveness. It accounts for sampling noise and legitimate split difficulty differences.

The gate configuration has to be fixed before the candidate is scored:

  • scenario ids
  • split ids
  • metric weights
  • deterministic checks
  • judge versions
  • budget ceilings
  • minimum paired runs
  • epsilon
  • alpha

If those change after seeing the candidate, the run is a new experiment. It is not the same gate.

Cost Is Part Of Correctness

The previous post argued that test-time compute has to beat random at equal budget. Evaluation gates enforce the product version of the same rule.

Cost is a vector:

C = {
  dollars,
  input_tokens,
  output_tokens,
  wall_ms,
  model_calls,
  tool_calls,
  sandbox_minutes,
  human_review_minutes,
  risk_budget
}

A candidate that improves score by 2 points and costs 20 times more is not automatically bad. It might be worth it. But it is not the same product.

So the gate should separate quality from efficiency:

quality_gate(c, b) = LCB_95(median(delta)) > epsilon
efficiency_gate(c) = median_cost(c) <= cost_ceiling
latency_gate(c) = p95_wall_ms(c) <= latency_ceiling

Then promotion is conjunctive:

promote(c) iff
  quality_gate(c, b)
  and efficiency_gate(c)
  and latency_gate(c)

When the candidate clears quality but fails cost, that is valuable evidence. It says the optimizer found a better behavior that is too expensive for the current product envelope.

Deterministic Failures Dominate Judges

LLM judges are useful. They are also weak authorities.

If a patch fails tests, a judge saying “looks good” does not rescue it. If a workflow violates permissions, a rubric score does not excuse it. If the trace has no real backend activity, a pass rate is meaningless.

Gate precedence is:

  1. deterministic verifier
  2. trace integrity
  3. backend integrity
  4. cost and latency policy
  5. calibrated semantic judge
  6. aggregate score

The order matters. A judge is allowed to score ambiguous quality. It is not allowed to override hard evidence.

This is the difference between an eval and a release gate. An eval reports. A release gate refuses.

Judge Reliability

Open-ended agent work needs semantic evaluation. Exact unit tests do not cover “did the agent satisfy the user’s intent” or “is the answer useful enough to ship.”

LLM-as-judge research made this practical. G-Eval showed that prompted GPT-4 style evaluators could align better with human judgments on natural-language tasks than older automatic metrics. MT-Bench and Chatbot Arena showed that strong model judges can approximate human preference for open-ended dialogue, while also documenting position bias, verbosity bias, self-enhancement bias, and limited reasoning.

The lesson is not “use an LLM judge and trust it.”

The lesson is:

  • Use judges where deterministic verification is unavailable,
  • then evaluate the judge as a measurement instrument.

Let:

V(x,y,trace)=judge scoreR(x,y)=true product outcome\begin{aligned} V(x,y,\text{trace}) &= \text{judge score} \\ R(x,y) &= \text{true product outcome} \end{aligned}

A judge gate needs calibration:

calibration_error⁡=E[R∣V=s]−s\operatorname{calibration\_error}=\mathbb{E}[R\mid V=s]-s

and agreement:

agreement⁡=P(sign⁡(Va−Vb)=sign⁡(Ha−Hb))\operatorname{agreement}=\mathbb{P}(\operatorname{sign}(V_a-V_b)=\operatorname{sign}(H_a-H_b))

where HH is a human preference or trusted adjudicator. For ordinal rubrics, rank correlation is often more useful than raw score correlation:

ρ=Spearman⁡(V,H)\rho=\operatorname{Spearman}(V,H)

A gate should track judge drift over time. If the judge model, rubric, prompt, or examples change, the score distribution can move even when the agent behavior does not.

That is why judge identity belongs in the profile cell.

Scorecards Are Cells, Not Averages

HELM pushed a simple but important idea: language model evaluation should expose multiple metrics and scenarios, not only one headline number.

Agent evaluation needs the same idea, but with runtime context.

A scorecard cell is keyed by:

cell = (scenario_id, profile_hash)

The profile material covers:

  • model
  • prompt hash
  • harness
  • source profile hash
  • dimensions

Tool surface, skill surface, runtime topology, judge version, backend, and tenant or persona metadata belong in the source profile or dimensions. The important property is not where each field lives. The important property is that behaviorally different runs land in different cells.

Then every commit appends a new observation:

scorecard[cell].timeline += {
  commit,
  scores,
  composite,
  per_dimension,
  run_ids
}

This prevents aggregate masking. A candidate can improve the mean while regressing one persona, one tool boundary, one model backend, or one scenario family.

Regression detection should combine effect size and statistical confidence:

Δ=current−baselined=CohenD⁡(current_scores,baseline_scores)p=WelchT⁡(current_scores,baseline_scores)\begin{aligned} \Delta &= \text{current}-\text{baseline} \\ d &= \operatorname{CohenD}(\text{current\_scores},\text{baseline\_scores}) \\ p &= \operatorname{WelchT}(\text{current\_scores},\text{baseline\_scores}) \end{aligned}
regressed if delta < 0 and abs(d) >= d_min and p <= alpha

If sample size is too small for statistics, use a conservative raw-delta threshold and mark the evidence weak.

Backend Integrity

One of the easiest ways to get a false eval is to never call the model.

The system can run every scenario, produce every row, and still be blind if the backend was a stub, a bridge was down, auth failed, or the provider route silently returned canned output.

Backend integrity has to be a gate, not a warning.

The minimal fingerprint:

real_backend(record) iff
  token_usage.input > 0
  or token_usage.output > 0

Then:

  • reject if every record is stub
  • reject_or_quarantine if records are mixed real and stub
  • flag if output tokens exist but cost is zero

The mixed case matters. A partial backend failure is missing data, not agent failure. Treating missing data as bad agent behavior poisons the optimizer. It teaches the system to “fix” a candidate that was never actually evaluated.

Semantic Fulfillment

For agents, the central question is often not:

Did the output look fluent?

It is:

Did the system do what the user asked?

That needs an intent-match layer. The evaluator should compare user request, available context, trace, and final artifact. It should distinguish:

  • solved the requested task
  • solved a nearby task
  • gave generic advice
  • refused incorrectly
  • changed forbidden files
  • skipped the hard part
  • produced plausible but ungrounded work

This is especially important for multi-agent systems. A coordinator can produce a clean final answer while worker traces show that the crucial evidence never arrived.

Failure Taxonomies

A scalar score tells the optimizer which candidate won.

A failure taxonomy tells the next optimizer where to search.

Examples:

  • reasoning_error
  • tool_selection_error
  • tool_argument_error
  • bad_retrieval
  • missing_codebase_context
  • missing_credentials
  • integration_auth_expired
  • budget_exceeded
  • format_drift
  • insufficient_evidence
  • ambiguous_user_intent
  • knowledge_readiness_blocked

The taxonomy turns eval into diagnosis. A prompt optimizer can respond to instruction-following failures. A skill optimizer can respond to repeated procedure failures. A runtime topology optimizer can respond to tool selection, budget, or missing-context failures. A knowledge system can respond to bad retrieval or stale external data.

Without a taxonomy, every loss becomes “make the prompt better.”

Release Confidence

The release gate is broader than the held-out gate.

The held-out gate answers:

Did candidate beat baseline on held-out paired evidence?

The release confidence layer asks:

Is there enough evidence to ship this change?

A useful release confidence scorecard has five axes:

  • Corpus: scenarios, split coverage, manifest integrity
  • Quality: pass rate, mean score, deterministic verifier status
  • Generalization: holdout runs, search-holdout gap, paired gate decision
  • Diagnostics: failure rows have actionable side information
  • Efficiency: mean cost, p95 wall time, budget compliance

The important phrase is fail closed. Missing corpus, missing holdout, missing traces, missing backend evidence, or missing diagnostics is not neutral. It is a reason to reject promotion until the evidence exists.

In compact form:

release_promote(c) iff
  held_out_gate(c, b) = promote
  and corpus_axis(c) = pass
  and quality_axis(c) = pass
  and generalization_axis(c) = pass
  and diagnostics_axis(c) = pass
  and efficiency_axis(c) = pass

The release gate composes evidence. It does not average away missing evidence.

Where Tangle Fits

Local package audit on June 6, 2026:

  • @tangle-network/agent-eval@0.34.1
  • @tangle-network/agent-runtime@0.26.0

agent-runtime spends compute. agent-eval decides whether the spend earned promotion.

The relevant agent-eval surface:

  • runEvalCampaign: runs variant by scenario by seed matrices, requires non-empty variants and scenarios, requires commitSha, fingerprints campaign inputs, and captures run integrity.
  • HeldOutGate: evaluates paired holdout deltas against a named baseline, bootstrap lower confidence bound, overfit gap, productive-run minimum, and optional cost ceiling.
  • evaluateReleaseConfidence: composes corpus, quality, generalization, diagnostics, and efficiency into a production-facing scorecard.
  • recordRunsToScorecard, loadScorecard, diffScorecard: maintain an append-only scenario by profile timeline and detect regressions.
  • AgentProfileCell: records profile id, source profile hash, harness, model, prompt hash, and dimensions so score changes are attributable to a stable run identity.
  • assertRealBackend: rejects blind evals where every run has zero token activity.
  • runIntentMatchJudge: evaluates semantic fulfillment.
  • FAILURE_CLASSES: gives failure analysis a typed ontology.
  • AnalystRegistry with DEFAULT_TRACE_ANALYST_KINDS: chains failure-mode, knowledge-gap, knowledge-poisoning, and improvement analysts.
  • runProductionLoop: ties observed traces, clustered failures, mutation, held-out gate, release confidence, and candidate promotion into one loop.

The relevant agent-runtime surface:

  • runLoop: executes bounded refine or fanout loops with cost aggregation and trace events.
  • createRefineDriver: spends compute sequentially.
  • createFanoutVoteDriver: spends compute across parallel variants.
  • Validator: supplies selector evidence.
  • conversation: enforces maxTurns, maxCreditsCents, turnOrder, haltOn, journals, and deterministic turn ids.
  • MCP delegation: makes specialist code and research work observable rather than hidden inside one prompt.

The boundary is clean:

  • runtime creates candidate behavior
  • eval determines whether behavior can replace the baseline

Gate Failure Modes

Holdout leakage

The optimizer sees examples, labels, judge rationales, or traces from the promotion split.

Unpaired comparisons

Candidate and baseline see different tasks, profiles, or seeds.

Mean-only promotion

The aggregate improves while a high-value cell regresses.

Judge monoculture

One judge model, one rubric, and one prompt define the whole objective.

Deterministic override

A judge passes an artifact that failed tests, permissions, build, or safety policy.

Backend blindness

The eval ran against stubs, missing auth, or a dead bridge.

Cost laundering

The score improves only by spending more hidden compute, tools, sandbox time, or human review.

Trace amnesia

The final output is scored, but the branch, tool, verifier, and selector evidence is missing.

Failure flattening

All failures collapse into one scalar, so the next optimizer has no diagnostic direction.

When A Candidate Can Ship

The optimizer proposes.

The gate governs.

A self-improving agent system needs a gate that can answer:

  • What changed?
  • Which baseline did it beat?
  • Which held-out tasks did it beat it on?
  • How much uncertainty remains?
  • Which profiles regressed?
  • Which deterministic checks failed?
  • Which judge scored it?
  • Was the backend real?
  • What did it cost?
  • What failure modes remain?
  • Can the trace prove all of that?

If the answer is missing, the candidate stays a candidate.

The gate is where self-improvement stops being a story about better prompts and becomes a release discipline.

Source Trail

Source freshness checked on 2026-06-06.

Revision history9revisions
  1. GPT-6-lunapolish+69−118 view trace →
    Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
    show diff
    diff --git a/src/content/posts/self-improving-stack-evaluation-gates.mdx b/src/content/posts/self-improving-stack-evaluation-gates.mdxindex 77dbf77..4e2c5ae 100644--- a/src/content/posts/self-improving-stack-evaluation-gates.mdx+++ b/src/content/posts/self-improving-stack-evaluation-gates.mdx@@ -99,15 +99,11 @@ That version is readable, but still too loose. A real gate needs paired observat  A leaderboard asks: -```text-Which system had the highest score?-```+> Which system had the highest score?  A gate asks: -```text-Should this candidate replace this baseline for this product?-```+> Should this candidate replace this baseline for this product?  Those are different questions. @@ -139,9 +135,7 @@ $$  where the candidate and baseline are evaluated on the same scenario, profile, and replicate. Then the question becomes: -```text-Is median(delta_i) reliably positive on held-out items?-```+> Is median(delta_i) reliably positive on held-out items?  The median is useful because agent scores often have heavy tails. One catastrophic failure or one lucky success can distort a mean. The median asks whether the typical paired task improved. @@ -176,17 +170,13 @@ Gates need data the optimizer did not tune against.  That gives two split families: -```text-D_search  = used to propose, mutate, rank, debug, and iterate-D_holdout = used to decide promotion-```+- $D_{\text{search}}$: used to propose, mutate, rank, debug, and iterate.+- $D_{\text{holdout}}$: used to decide promotion.  The failure mode is: -```text-score(c, search) goes up-score(c, holdout) goes down-```+- Search score increases.+- Holdout score decreases.  So the gate needs an overfit check: @@ -205,17 +195,15 @@ This says the candidate may look better on the search split, but it cannot be mu  The gate configuration has to be fixed before the candidate is scored: -```text-scenario ids-split ids-metric weights-deterministic checks-judge versions-budget ceilings-minimum paired runs-epsilon-alpha-```+- scenario ids+- split ids+- metric weights+- deterministic checks+- judge versions+- budget ceilings+- minimum paired runs+- `epsilon`+- `alpha`  If those change after seeing the candidate, the run is a new experiment. It is not the same gate. @@ -268,14 +256,12 @@ If a patch fails tests, a judge saying "looks good" does not rescue it. If a wor  Gate precedence is: -```text-deterministic verifier-  > trace integrity-  > backend integrity-  > cost and latency policy-  > calibrated semantic judge-  > aggregate score-```+1. deterministic verifier+2. trace integrity+3. backend integrity+4. cost and latency policy+5. calibrated semantic judge+6. aggregate score  The order matters. A judge is allowed to score ambiguous quality. It is not allowed to override hard evidence. @@ -291,10 +277,8 @@ The lesson is not "use an LLM judge and trust it."  The lesson is: -```text-Use judges where deterministic verification is unavailable,-then evaluate the judge as a measurement instrument.-```+- Use judges where deterministic verification is unavailable,+- then evaluate the judge as a measurement instrument.  Let: @@ -335,19 +319,15 @@ Agent evaluation needs the same idea, but with runtime context.  A scorecard cell is keyed by: -```text-cell = (scenario_id, profile_hash)-```+`cell = (scenario_id, profile_hash)`  The profile material covers: -```text-model-prompt hash-harness-source profile hash-dimensions-```+- `model`+- prompt hash+- `harness`+- source profile hash+- `dimensions`  Tool surface, skill surface, runtime topology, judge version, backend, and tenant or persona metadata belong in the source profile or dimensions. The important property is not where each field lives. The important property is that behaviorally different runs land in different cells. @@ -399,11 +379,9 @@ real_backend(record) iff  Then: -```text-reject if every record is stub-reject_or_quarantine if records are mixed real and stub-flag if output tokens exist but cost is zero-```+- reject if every record is stub+- reject_or_quarantine if records are mixed real and stub+- flag if output tokens exist but cost is zero  The mixed case matters. A partial backend failure is missing data, not agent failure. Treating missing data as bad agent behavior poisons the optimizer. It teaches the system to "fix" a candidate that was never actually evaluated. @@ -411,15 +389,11 @@ The mixed case matters. A partial backend failure is missing data, not agent fai  For agents, the central question is often not: -```text-Did the output look fluent?-```+> Did the output look fluent?  It is: -```text-Did the system do what the user asked?-```+> Did the system do what the user asked?  That needs an intent-match layer. The evaluator should compare user request, available context, trace, and final artifact. It should distinguish: @@ -441,20 +415,18 @@ A failure taxonomy tells the next optimizer where to search.  Examples: -```text-reasoning_error-tool_selection_error-tool_argument_error-bad_retrieval-missing_codebase_context-missing_credentials-integration_auth_expired-budget_exceeded-format_drift-insufficient_evidence-ambiguous_user_intent-knowledge_readiness_blocked-```+- `reasoning_error`+- `tool_selection_error`+- `tool_argument_error`+- `bad_retrieval`+- `missing_codebase_context`+- `missing_credentials`+- `integration_auth_expired`+- `budget_exceeded`+- `format_drift`+- `insufficient_evidence`+- `ambiguous_user_intent`+- `knowledge_readiness_blocked`  The taxonomy turns eval into diagnosis. A prompt optimizer can respond to instruction-following failures. A skill optimizer can respond to repeated procedure failures. A runtime topology optimizer can respond to tool selection, budget, or missing-context failures. A knowledge system can respond to bad retrieval or stale external data. @@ -466,34 +438,19 @@ The release gate is broader than the held-out gate.  The held-out gate answers: -```text-Did candidate beat baseline on held-out paired evidence?-```+> Did candidate beat baseline on held-out paired evidence?  The release confidence layer asks: -```text-Is there enough evidence to ship this change?-```+> Is there enough evidence to ship this change?  A useful release confidence scorecard has five axes: -```text-corpus:-  scenarios, split coverage, manifest integrity--quality:-  pass rate, mean score, deterministic verifier status--generalization:-  holdout runs, search-holdout gap, paired gate decision--diagnostics:-  failure rows have actionable side information--efficiency:-  mean cost, p95 wall time, budget compliance-```+- **Corpus:** scenarios, split coverage, manifest integrity+- **Quality:** pass rate, mean score, deterministic verifier status+- **Generalization:** holdout runs, search-holdout gap, paired gate decision+- **Diagnostics:** failure rows have actionable side information+- **Efficiency:** mean cost, p95 wall time, budget compliance  The important phrase is fail closed. Missing corpus, missing holdout, missing traces, missing backend evidence, or missing diagnostics is not neutral. It is a reason to reject promotion until the evidence exists. @@ -515,10 +472,8 @@ The release gate composes evidence. It does not average away missing evidence.  Local package audit on June 6, 2026: -```text-@tangle-network/agent-eval@0.34.1-@tangle-network/agent-runtime@0.26.0-```+- `@tangle-network/agent-eval@0.34.1`+- `@tangle-network/agent-runtime@0.26.0`  `agent-runtime` spends compute. `agent-eval` decides whether the spend earned promotion. @@ -546,10 +501,8 @@ The relevant `agent-runtime` surface:  The boundary is clean: -```text-runtime creates candidate behavior-eval determines whether behavior can replace the baseline-```+- runtime creates candidate behavior+- eval determines whether behavior can replace the baseline  ## Gate Failure Modes @@ -597,19 +550,17 @@ The gate governs.  A self-improving agent system needs a gate that can answer: -```text-What changed?-Which baseline did it beat?-Which held-out tasks did it beat it on?-How much uncertainty remains?-Which profiles regressed?-Which deterministic checks failed?-Which judge scored it?-Was the backend real?-What did it cost?-What failure modes remain?-Can the trace prove all of that?-```+- What changed?+- Which baseline did it beat?+- Which held-out tasks did it beat it on?+- How much uncertainty remains?+- Which profiles regressed?+- Which deterministic checks failed?+- Which judge scored it?+- Was the backend real?+- What did it cost?+- What failure modes remain?+- Can the trace prove all of that?  If the answer is missing, the candidate stays a candidate. 
  2. GPT-6-lunapolish+57−44 view trace →
    Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
    show diff
    diff --git a/src/content/posts/self-improving-stack-evaluation-gates.mdx b/src/content/posts/self-improving-stack-evaluation-gates.mdxindex 06564b3..6f14348 100644--- a/src/content/posts/self-improving-stack-evaluation-gates.mdx+++ b/src/content/posts/self-improving-stack-evaluation-gates.mdx@@ -57,18 +57,20 @@ A gate is a promotion policy.  Let: -```text-b = baseline system-c = candidate system-x = scenario-p = agent profile cell-z = seed or replicate id-R = task reward or score-C = measured cost vector-T = trace integrity predicate-D_search = search split-D_holdout = held-out split-```+$$+\begin{aligned}+  b &= \text{baseline system} \\+  c &= \text{candidate system} \\+  x &= \text{scenario} \\+  p &= \text{agent profile cell} \\+  z &= \text{seed or replicate id} \\+  R &= \text{task reward or score} \\+  C &= \text{measured cost vector} \\+  T &= \text{trace integrity predicate} \\+  D_{\text{search}} &= \text{search split} \\+  D_{\text{holdout}} &= \text{held-out split}+\end{aligned}+$$  The gate is a function: @@ -129,9 +131,9 @@ Suppose the baseline sees one sample of tasks and the candidate sees another. A  The paired comparison fixes that: -```text-delta_i = R(c, x_i, p_i, z_i) - R(b, x_i, p_i, z_i)-```+$$+\Delta_i = R(c,x_i,p_i,z_i)-R(b,x_i,p_i,z_i)+$$  where the candidate and baseline are evaluated on the same scenario, profile, and replicate. Then the question becomes: @@ -143,12 +145,11 @@ The median is useful because agent scores often have heavy tails. One catastroph  A practical promotion rule: -```text-n_pairs >= n_min-LCB_95(median(delta)) > epsilon-```+$$+n_{\text{pairs}} \ge n_{\min},\qquad \operatorname{LCB}_{95}(\operatorname{median}(\Delta)) > \epsilon+$$ -`LCB_95` is the lower confidence bound. If the lower bound clears the threshold, the gate has evidence that the lift is not just random luck.+$\operatorname{LCB}_{95}$ is the lower confidence bound. If the lower bound clears the threshold, the gate has evidence that the lift is not just random luck.  This is where bootstrap confidence intervals are useful. You resample paired deltas, compute the median for each resample, and inspect the lower quantile: @@ -157,10 +158,12 @@ delta = [delta_1, ..., delta_n] for r in 1..B:   sample n deltas with replacement   m_r = median(sample)--LCB_95 = quantile({m_r}, 0.025)  // two-sided 95 interval ``` +$$+\operatorname{LCB}_{95}=Q_{0.025}(m_1,\ldots,m_B)+$$+ The gate promotes only when the pessimistic estimate is still good enough.  ## Search Split Versus Holdout@@ -185,14 +188,18 @@ score(c, holdout) goes down  So the gate needs an overfit check: -```text-gap_c = mean_score(c, search) - mean_score(c, holdout)-gap_b = mean_score(b, search) - mean_score(b, holdout)+$$+\begin{aligned}+  \operatorname{gap}_c &= \operatorname{mean\_score}(c,\text{search})-\operatorname{mean\_score}(c,\text{holdout}) \\+  \operatorname{gap}_b &= \operatorname{mean\_score}(b,\text{search})-\operatorname{mean\_score}(b,\text{holdout})+\end{aligned}+$$ +```text reject(c) if gap_c > gap_b + tau ``` -This says the candidate may look better on the search split, but it cannot be much more search-specialized than the baseline. `tau` is slack, not forgiveness. It accounts for sampling noise and legitimate split difficulty differences.+This says the candidate may look better on the search split, but it cannot be much more search-specialized than the baseline. $\tau$ is slack, not forgiveness. It accounts for sampling noise and legitimate split difficulty differences.  The gate configuration has to be fixed before the candidate is scored: @@ -289,28 +296,30 @@ then evaluate the judge as a measurement instrument.  Let: -```text-V(x, y, trace) = judge score-R(x, y) = true product outcome-```+$$+\begin{aligned}+V(x,y,\text{trace}) &= \text{judge score} \\+R(x,y) &= \text{true product outcome}+\end{aligned}+$$  A judge gate needs calibration: -```text-calibration_error = E[R | V = s] - s-```+$$+\operatorname{calibration\_error}=\mathbb{E}[R\mid V=s]-s+$$  and agreement: -```text-agreement = P(sign(V_a - V_b) = sign(H_a - H_b))-```+$$+\operatorname{agreement}=\mathbb{P}(\operatorname{sign}(V_a-V_b)=\operatorname{sign}(H_a-H_b))+$$ -where `H` is a human preference or trusted adjudicator. For ordinal rubrics, rank correlation is often more useful than raw score correlation:+where $H$ is a human preference or trusted adjudicator. For ordinal rubrics, rank correlation is often more useful than raw score correlation: -```text-rho = Spearman(V, H)-```+$$+\rho=\operatorname{Spearman}(V,H)+$$  A gate should track judge drift over time. If the judge model, rubric, prompt, or examples change, the score distribution can move even when the agent behavior does not. @@ -356,11 +365,15 @@ This prevents aggregate masking. A candidate can improve the mean while regressi  Regression detection should combine effect size and statistical confidence: -```text-delta = current - baseline-d = CohenD(current_scores, baseline_scores)-p = WelchT(current_scores, baseline_scores)+$$+\begin{aligned}+  \Delta &= \text{current}-\text{baseline} \\+  d &= \operatorname{CohenD}(\text{current\_scores},\text{baseline\_scores}) \\+  p &= \operatorname{WelchT}(\text{current\_scores},\text{baseline\_scores})+\end{aligned}+$$ +```text regressed if delta < 0 and abs(d) >= d_min and p <= alpha ``` 
  3. GPT-5.5draft view trace →
    Drafted the evaluation-gates post with held-out promotion math, scorecard cells, judge reliability, backend integrity, release confidence, and local Tangle package placement.
  4. GPT-5.5polish view trace →
    Polished the evaluation-gates post by tightening profile-cell claims against the local AgentProfileCell schema, adding gate pre-registration invariants, clarifying bootstrap interval wording, and preserving the fail-closed promotion model.
  5. GPT-5.5polish+6−8 view trace →
    let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
    show diff
    diff --git a/src/content/posts/self-improving-stack-evaluation-gates.mdx b/src/content/posts/self-improving-stack-evaluation-gates.mdxindex da8d0d0..52f7a8b 100644--- a/src/content/posts/self-improving-stack-evaluation-gates.mdx+++ b/src/content/posts/self-improving-stack-evaluation-gates.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-06   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-evaluation-gates-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-evaluation-gates-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-evaluation-gates-review' }   - date: 2026-06-05@@ -41,15 +43,11 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -An optimizer can propose forever.+A green score is not a release decision. It is evidence entering a release policy, and in self-improving loops that distinction matters more than the optimizer. -The gate decides what becomes the system.+GEPA, MIPRO, SkillOpt, topology search, and meta-harness can all generate candidates forever. The gate decides which candidate becomes the system future agents inherit. If the gate is weak, the optimizer learns the gate. If the gate is honest, the optimizer has to improve the product. -That is why the gate is not an administrative detail after the interesting work. It is the objective boundary. GEPA, MIPRO, SkillOpt, runtime topology search, and meta-harness can all generate candidates. The gate decides which candidate is allowed to replace the baseline.--If the gate is weak, every optimizer learns the gate.--If the gate is honest, every optimizer has to improve the product.+This is why the gate is not an administrative detail after the interesting work. It is the objective boundary.  ## What A Gate Is @@ -574,7 +572,7 @@ The final output is scored, but the branch, tool, verifier, and selector evidenc  All failures collapse into one scalar, so the next optimizer has no diagnostic direction. -## Working Rule+## When A Candidate Can Ship  The optimizer proposes. 
  6. GPT-5.5rewrite view trace →
    60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
  7. GPT-5.5publish view trace →
    Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
  8. GPT-5.5review view trace →
    Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
  9. GPT-5.5outline view trace →
    Research planning pass from a traced session.

Comments

Comments load from GitHub Discussions via Giscus. Configure PUBLIC_GISCUS_REPO, PUBLIC_GISCUS_REPO_ID, PUBLIC_GISCUS_CATEGORY, and PUBLIC_GISCUS_CATEGORY_ID in .env. See giscus.app to generate the IDs after you enable Discussions on the repo.