The Gate Is The Optimizer
Why held-out promotion, judge reliability, failure taxonomies, cost ceilings, and confidence intervals decide whether self-improvement is real.
The Self Improving Stack series
Browse all 13 posts
- Topology Is The Missing Action Space
- The Gate Is The Optimizer
- Self-Improvement Needs A Safety Case
- When The Harness Has To Evolve
- Memory Is Not Automatically Learning
- Personas Are Content, Coordination Is Structure
- Optimization Theory For Agent Builders
- When The Model Itself Is Mutable
- Prompt Optimization Is Not The Whole Game
- Skills Are Trainable State
- Beat Random At Equal Compute First
- Traces Are The Training Data
- The Self-Improving Stack
A green score is not a release decision. It is evidence entering a release policy, and in self-improving loops that distinction matters more than the optimizer.
GEPA, MIPRO, SkillOpt, topology search, and meta-harness can all generate candidates forever. The gate decides which candidate becomes the system future agents inherit. If the gate is weak, the optimizer learns the gate. If the gate is honest, the optimizer has to improve the product.
This is why the gate is not an administrative detail after the interesting work. It is the objective boundary.
What A Gate Is
A gate is a promotion policy.
Let:
The gate is a function:
G(c, b, D_holdout, p, z) -> {promote, reject}
It is allowed to inspect evidence. It is not allowed to move the goalpost after seeing the candidate.
The simplest version is:
promote(c) iff
quality(c, holdout) > quality(b, holdout)
and cost(c) <= cost_ceiling
and latency(c) <= latency_ceiling
and deterministic_failures(c) = 0
and trace_integrity(c) = 1
That version is readable, but still too loose. A real gate needs paired observations, uncertainty, split discipline, judge reliability, and regression protection.
The Gate Is Not A Leaderboard
A leaderboard asks:
Which system had the highest score?
A gate asks:
Should this candidate replace this baseline for this product?
Those are different questions.
Leaderboards are useful for orientation. They are bad promotion policies. They compress context, cost, latency, tool availability, profile differences, data leakage, and failure severity into one rank.
Agent systems make this worse because candidates can improve the metric while damaging the workflow:
- better aggregate score, worse high-value persona
- better judge score, worse deterministic verifier
- better single-shot result, worse cost
- better easy tasks, worse hard tasks
- better search split, worse holdout
- better answer style, worse intent match
- better visible output, broken trace capture
The gate has to preserve the baseline unless the candidate earns replacement.
Paired Evidence
Unpaired averages are fragile.
Suppose the baseline sees one sample of tasks and the candidate sees another. A mean difference can be task mix, not improvement.
The paired comparison fixes that:
where the candidate and baseline are evaluated on the same scenario, profile, and replicate. Then the question becomes:
Is median(delta_i) reliably positive on held-out items?
The median is useful because agent scores often have heavy tails. One catastrophic failure or one lucky success can distort a mean. The median asks whether the typical paired task improved.
A practical promotion rule:
is the lower confidence bound. If the lower bound clears the threshold, the gate has evidence that the lift is not just random luck.
This is where bootstrap confidence intervals are useful. You resample paired deltas, compute the median for each resample, and inspect the lower quantile:
delta = [delta_1, ..., delta_n]
for r in 1..B:
sample n deltas with replacement
m_r = median(sample)
The gate promotes only when the pessimistic estimate is still good enough.
Search Split Versus Holdout
Optimizers need data to search.
Gates need data the optimizer did not tune against.
That gives two split families:
- : used to propose, mutate, rank, debug, and iterate.
- : used to decide promotion.
The failure mode is:
- Search score increases.
- Holdout score decreases.
So the gate needs an overfit check:
reject(c) if gap_c > gap_b + tau
This says the candidate may look better on the search split, but it cannot be much more search-specialized than the baseline. is slack, not forgiveness. It accounts for sampling noise and legitimate split difficulty differences.
The gate configuration has to be fixed before the candidate is scored:
- scenario ids
- split ids
- metric weights
- deterministic checks
- judge versions
- budget ceilings
- minimum paired runs
epsilonalpha
If those change after seeing the candidate, the run is a new experiment. It is not the same gate.
Cost Is Part Of Correctness
The previous post argued that test-time compute has to beat random at equal budget. Evaluation gates enforce the product version of the same rule.
Cost is a vector:
C = {
dollars,
input_tokens,
output_tokens,
wall_ms,
model_calls,
tool_calls,
sandbox_minutes,
human_review_minutes,
risk_budget
}
A candidate that improves score by 2 points and costs 20 times more is not automatically bad. It might be worth it. But it is not the same product.
So the gate should separate quality from efficiency:
quality_gate(c, b) = LCB_95(median(delta)) > epsilon
efficiency_gate(c) = median_cost(c) <= cost_ceiling
latency_gate(c) = p95_wall_ms(c) <= latency_ceiling
Then promotion is conjunctive:
promote(c) iff
quality_gate(c, b)
and efficiency_gate(c)
and latency_gate(c)
When the candidate clears quality but fails cost, that is valuable evidence. It says the optimizer found a better behavior that is too expensive for the current product envelope.
Deterministic Failures Dominate Judges
LLM judges are useful. They are also weak authorities.
If a patch fails tests, a judge saying “looks good” does not rescue it. If a workflow violates permissions, a rubric score does not excuse it. If the trace has no real backend activity, a pass rate is meaningless.
Gate precedence is:
- deterministic verifier
- trace integrity
- backend integrity
- cost and latency policy
- calibrated semantic judge
- aggregate score
The order matters. A judge is allowed to score ambiguous quality. It is not allowed to override hard evidence.
This is the difference between an eval and a release gate. An eval reports. A release gate refuses.
Judge Reliability
Open-ended agent work needs semantic evaluation. Exact unit tests do not cover “did the agent satisfy the user’s intent” or “is the answer useful enough to ship.”
LLM-as-judge research made this practical. G-Eval showed that prompted GPT-4 style evaluators could align better with human judgments on natural-language tasks than older automatic metrics. MT-Bench and Chatbot Arena showed that strong model judges can approximate human preference for open-ended dialogue, while also documenting position bias, verbosity bias, self-enhancement bias, and limited reasoning.
The lesson is not “use an LLM judge and trust it.”
The lesson is:
- Use judges where deterministic verification is unavailable,
- then evaluate the judge as a measurement instrument.
Let:
A judge gate needs calibration:
and agreement:
where is a human preference or trusted adjudicator. For ordinal rubrics, rank correlation is often more useful than raw score correlation:
A gate should track judge drift over time. If the judge model, rubric, prompt, or examples change, the score distribution can move even when the agent behavior does not.
That is why judge identity belongs in the profile cell.
Scorecards Are Cells, Not Averages
HELM pushed a simple but important idea: language model evaluation should expose multiple metrics and scenarios, not only one headline number.
Agent evaluation needs the same idea, but with runtime context.
A scorecard cell is keyed by:
cell = (scenario_id, profile_hash)
The profile material covers:
model- prompt hash
harness- source profile hash
dimensions
Tool surface, skill surface, runtime topology, judge version, backend, and tenant or persona metadata belong in the source profile or dimensions. The important property is not where each field lives. The important property is that behaviorally different runs land in different cells.
Then every commit appends a new observation:
scorecard[cell].timeline += {
commit,
scores,
composite,
per_dimension,
run_ids
}
This prevents aggregate masking. A candidate can improve the mean while regressing one persona, one tool boundary, one model backend, or one scenario family.
Regression detection should combine effect size and statistical confidence:
regressed if delta < 0 and abs(d) >= d_min and p <= alpha
If sample size is too small for statistics, use a conservative raw-delta threshold and mark the evidence weak.
Backend Integrity
One of the easiest ways to get a false eval is to never call the model.
The system can run every scenario, produce every row, and still be blind if the backend was a stub, a bridge was down, auth failed, or the provider route silently returned canned output.
Backend integrity has to be a gate, not a warning.
The minimal fingerprint:
real_backend(record) iff
token_usage.input > 0
or token_usage.output > 0
Then:
- reject if every record is stub
- reject_or_quarantine if records are mixed real and stub
- flag if output tokens exist but cost is zero
The mixed case matters. A partial backend failure is missing data, not agent failure. Treating missing data as bad agent behavior poisons the optimizer. It teaches the system to “fix” a candidate that was never actually evaluated.
Semantic Fulfillment
For agents, the central question is often not:
Did the output look fluent?
It is:
Did the system do what the user asked?
That needs an intent-match layer. The evaluator should compare user request, available context, trace, and final artifact. It should distinguish:
- solved the requested task
- solved a nearby task
- gave generic advice
- refused incorrectly
- changed forbidden files
- skipped the hard part
- produced plausible but ungrounded work
This is especially important for multi-agent systems. A coordinator can produce a clean final answer while worker traces show that the crucial evidence never arrived.
Failure Taxonomies
A scalar score tells the optimizer which candidate won.
A failure taxonomy tells the next optimizer where to search.
Examples:
reasoning_errortool_selection_errortool_argument_errorbad_retrievalmissing_codebase_contextmissing_credentialsintegration_auth_expiredbudget_exceededformat_driftinsufficient_evidenceambiguous_user_intentknowledge_readiness_blocked
The taxonomy turns eval into diagnosis. A prompt optimizer can respond to instruction-following failures. A skill optimizer can respond to repeated procedure failures. A runtime topology optimizer can respond to tool selection, budget, or missing-context failures. A knowledge system can respond to bad retrieval or stale external data.
Without a taxonomy, every loss becomes “make the prompt better.”
Release Confidence
The release gate is broader than the held-out gate.
The held-out gate answers:
Did candidate beat baseline on held-out paired evidence?
The release confidence layer asks:
Is there enough evidence to ship this change?
A useful release confidence scorecard has five axes:
- Corpus: scenarios, split coverage, manifest integrity
- Quality: pass rate, mean score, deterministic verifier status
- Generalization: holdout runs, search-holdout gap, paired gate decision
- Diagnostics: failure rows have actionable side information
- Efficiency: mean cost, p95 wall time, budget compliance
The important phrase is fail closed. Missing corpus, missing holdout, missing traces, missing backend evidence, or missing diagnostics is not neutral. It is a reason to reject promotion until the evidence exists.
In compact form:
release_promote(c) iff
held_out_gate(c, b) = promote
and corpus_axis(c) = pass
and quality_axis(c) = pass
and generalization_axis(c) = pass
and diagnostics_axis(c) = pass
and efficiency_axis(c) = pass
The release gate composes evidence. It does not average away missing evidence.
Where Tangle Fits
Local package audit on June 6, 2026:
@tangle-network/agent-eval@0.34.1@tangle-network/agent-runtime@0.26.0
agent-runtime spends compute. agent-eval decides whether the spend earned promotion.
The relevant agent-eval surface:
runEvalCampaign: runs variant by scenario by seed matrices, requires non-empty variants and scenarios, requirescommitSha, fingerprints campaign inputs, and captures run integrity.HeldOutGate: evaluates paired holdout deltas against a named baseline, bootstrap lower confidence bound, overfit gap, productive-run minimum, and optional cost ceiling.evaluateReleaseConfidence: composes corpus, quality, generalization, diagnostics, and efficiency into a production-facing scorecard.recordRunsToScorecard,loadScorecard,diffScorecard: maintain an append-only scenario by profile timeline and detect regressions.AgentProfileCell: records profile id, source profile hash, harness, model, prompt hash, and dimensions so score changes are attributable to a stable run identity.assertRealBackend: rejects blind evals where every run has zero token activity.runIntentMatchJudge: evaluates semantic fulfillment.FAILURE_CLASSES: gives failure analysis a typed ontology.AnalystRegistrywithDEFAULT_TRACE_ANALYST_KINDS: chains failure-mode, knowledge-gap, knowledge-poisoning, and improvement analysts.runProductionLoop: ties observed traces, clustered failures, mutation, held-out gate, release confidence, and candidate promotion into one loop.
The relevant agent-runtime surface:
runLoop: executes bounded refine or fanout loops with cost aggregation and trace events.createRefineDriver: spends compute sequentially.createFanoutVoteDriver: spends compute across parallel variants.Validator: supplies selector evidence.conversation: enforcesmaxTurns,maxCreditsCents,turnOrder,haltOn, journals, and deterministic turn ids.- MCP delegation: makes specialist code and research work observable rather than hidden inside one prompt.
The boundary is clean:
- runtime creates candidate behavior
- eval determines whether behavior can replace the baseline
Gate Failure Modes
Holdout leakage
The optimizer sees examples, labels, judge rationales, or traces from the promotion split.
Unpaired comparisons
Candidate and baseline see different tasks, profiles, or seeds.
Mean-only promotion
The aggregate improves while a high-value cell regresses.
Judge monoculture
One judge model, one rubric, and one prompt define the whole objective.
Deterministic override
A judge passes an artifact that failed tests, permissions, build, or safety policy.
Backend blindness
The eval ran against stubs, missing auth, or a dead bridge.
Cost laundering
The score improves only by spending more hidden compute, tools, sandbox time, or human review.
Trace amnesia
The final output is scored, but the branch, tool, verifier, and selector evidence is missing.
Failure flattening
All failures collapse into one scalar, so the next optimizer has no diagnostic direction.
When A Candidate Can Ship
The optimizer proposes.
The gate governs.
A self-improving agent system needs a gate that can answer:
- What changed?
- Which baseline did it beat?
- Which held-out tasks did it beat it on?
- How much uncertainty remains?
- Which profiles regressed?
- Which deterministic checks failed?
- Which judge scored it?
- Was the backend real?
- What did it cost?
- What failure modes remain?
- Can the trace prove all of that?
If the answer is missing, the candidate stays a candidate.
The gate is where self-improvement stops being a story about better prompts and becomes a release discipline.
Source Trail
Source freshness checked on 2026-06-06.
- Holistic Evaluation of Language Models, checked June 6, 2026.
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, checked June 6, 2026.
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, checked June 6, 2026.
- Evaluating Large Language Models Trained on Code, checked June 6, 2026.
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, checked June 6, 2026.
- Local
@tangle-network/agent-eval@0.34.1source audit:HeldOutGate,runEvalCampaign,evaluateReleaseConfidence, scorecard,AgentProfileCell,assertRealBackend,runIntentMatchJudge,FAILURE_CLASSES,AnalystRegistry,DEFAULT_TRACE_ANALYST_KINDS,runProductionLoop, June 6, 2026. - Local
@tangle-network/agent-runtime@0.26.0source audit:runLoop,createRefineDriver,createFanoutVoteDriver,Validator, conversation policy, journals, MCP delegation, June 6, 2026.
Revision history9revisions
- Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
show diff
diff --git a/src/content/posts/self-improving-stack-evaluation-gates.mdx b/src/content/posts/self-improving-stack-evaluation-gates.mdxindex 77dbf77..4e2c5ae 100644--- a/src/content/posts/self-improving-stack-evaluation-gates.mdx+++ b/src/content/posts/self-improving-stack-evaluation-gates.mdx@@ -99,15 +99,11 @@ That version is readable, but still too loose. A real gate needs paired observat A leaderboard asks: -```text-Which system had the highest score?-```+> Which system had the highest score? A gate asks: -```text-Should this candidate replace this baseline for this product?-```+> Should this candidate replace this baseline for this product? Those are different questions. @@ -139,9 +135,7 @@ $$ where the candidate and baseline are evaluated on the same scenario, profile, and replicate. Then the question becomes: -```text-Is median(delta_i) reliably positive on held-out items?-```+> Is median(delta_i) reliably positive on held-out items? The median is useful because agent scores often have heavy tails. One catastrophic failure or one lucky success can distort a mean. The median asks whether the typical paired task improved. @@ -176,17 +170,13 @@ Gates need data the optimizer did not tune against. That gives two split families: -```text-D_search = used to propose, mutate, rank, debug, and iterate-D_holdout = used to decide promotion-```+- $D_{\text{search}}$: used to propose, mutate, rank, debug, and iterate.+- $D_{\text{holdout}}$: used to decide promotion. The failure mode is: -```text-score(c, search) goes up-score(c, holdout) goes down-```+- Search score increases.+- Holdout score decreases. So the gate needs an overfit check: @@ -205,17 +195,15 @@ This says the candidate may look better on the search split, but it cannot be mu The gate configuration has to be fixed before the candidate is scored: -```text-scenario ids-split ids-metric weights-deterministic checks-judge versions-budget ceilings-minimum paired runs-epsilon-alpha-```+- scenario ids+- split ids+- metric weights+- deterministic checks+- judge versions+- budget ceilings+- minimum paired runs+- `epsilon`+- `alpha` If those change after seeing the candidate, the run is a new experiment. It is not the same gate. @@ -268,14 +256,12 @@ If a patch fails tests, a judge saying "looks good" does not rescue it. If a wor Gate precedence is: -```text-deterministic verifier- > trace integrity- > backend integrity- > cost and latency policy- > calibrated semantic judge- > aggregate score-```+1. deterministic verifier+2. trace integrity+3. backend integrity+4. cost and latency policy+5. calibrated semantic judge+6. aggregate score The order matters. A judge is allowed to score ambiguous quality. It is not allowed to override hard evidence. @@ -291,10 +277,8 @@ The lesson is not "use an LLM judge and trust it." The lesson is: -```text-Use judges where deterministic verification is unavailable,-then evaluate the judge as a measurement instrument.-```+- Use judges where deterministic verification is unavailable,+- then evaluate the judge as a measurement instrument. Let: @@ -335,19 +319,15 @@ Agent evaluation needs the same idea, but with runtime context. A scorecard cell is keyed by: -```text-cell = (scenario_id, profile_hash)-```+`cell = (scenario_id, profile_hash)` The profile material covers: -```text-model-prompt hash-harness-source profile hash-dimensions-```+- `model`+- prompt hash+- `harness`+- source profile hash+- `dimensions` Tool surface, skill surface, runtime topology, judge version, backend, and tenant or persona metadata belong in the source profile or dimensions. The important property is not where each field lives. The important property is that behaviorally different runs land in different cells. @@ -399,11 +379,9 @@ real_backend(record) iff Then: -```text-reject if every record is stub-reject_or_quarantine if records are mixed real and stub-flag if output tokens exist but cost is zero-```+- reject if every record is stub+- reject_or_quarantine if records are mixed real and stub+- flag if output tokens exist but cost is zero The mixed case matters. A partial backend failure is missing data, not agent failure. Treating missing data as bad agent behavior poisons the optimizer. It teaches the system to "fix" a candidate that was never actually evaluated. @@ -411,15 +389,11 @@ The mixed case matters. A partial backend failure is missing data, not agent fai For agents, the central question is often not: -```text-Did the output look fluent?-```+> Did the output look fluent? It is: -```text-Did the system do what the user asked?-```+> Did the system do what the user asked? That needs an intent-match layer. The evaluator should compare user request, available context, trace, and final artifact. It should distinguish: @@ -441,20 +415,18 @@ A failure taxonomy tells the next optimizer where to search. Examples: -```text-reasoning_error-tool_selection_error-tool_argument_error-bad_retrieval-missing_codebase_context-missing_credentials-integration_auth_expired-budget_exceeded-format_drift-insufficient_evidence-ambiguous_user_intent-knowledge_readiness_blocked-```+- `reasoning_error`+- `tool_selection_error`+- `tool_argument_error`+- `bad_retrieval`+- `missing_codebase_context`+- `missing_credentials`+- `integration_auth_expired`+- `budget_exceeded`+- `format_drift`+- `insufficient_evidence`+- `ambiguous_user_intent`+- `knowledge_readiness_blocked` The taxonomy turns eval into diagnosis. A prompt optimizer can respond to instruction-following failures. A skill optimizer can respond to repeated procedure failures. A runtime topology optimizer can respond to tool selection, budget, or missing-context failures. A knowledge system can respond to bad retrieval or stale external data. @@ -466,34 +438,19 @@ The release gate is broader than the held-out gate. The held-out gate answers: -```text-Did candidate beat baseline on held-out paired evidence?-```+> Did candidate beat baseline on held-out paired evidence? The release confidence layer asks: -```text-Is there enough evidence to ship this change?-```+> Is there enough evidence to ship this change? A useful release confidence scorecard has five axes: -```text-corpus:- scenarios, split coverage, manifest integrity--quality:- pass rate, mean score, deterministic verifier status--generalization:- holdout runs, search-holdout gap, paired gate decision--diagnostics:- failure rows have actionable side information--efficiency:- mean cost, p95 wall time, budget compliance-```+- **Corpus:** scenarios, split coverage, manifest integrity+- **Quality:** pass rate, mean score, deterministic verifier status+- **Generalization:** holdout runs, search-holdout gap, paired gate decision+- **Diagnostics:** failure rows have actionable side information+- **Efficiency:** mean cost, p95 wall time, budget compliance The important phrase is fail closed. Missing corpus, missing holdout, missing traces, missing backend evidence, or missing diagnostics is not neutral. It is a reason to reject promotion until the evidence exists. @@ -515,10 +472,8 @@ The release gate composes evidence. It does not average away missing evidence. Local package audit on June 6, 2026: -```text-@tangle-network/agent-eval@0.34.1-@tangle-network/agent-runtime@0.26.0-```+- `@tangle-network/agent-eval@0.34.1`+- `@tangle-network/agent-runtime@0.26.0` `agent-runtime` spends compute. `agent-eval` decides whether the spend earned promotion. @@ -546,10 +501,8 @@ The relevant `agent-runtime` surface: The boundary is clean: -```text-runtime creates candidate behavior-eval determines whether behavior can replace the baseline-```+- runtime creates candidate behavior+- eval determines whether behavior can replace the baseline ## Gate Failure Modes @@ -597,19 +550,17 @@ The gate governs. A self-improving agent system needs a gate that can answer: -```text-What changed?-Which baseline did it beat?-Which held-out tasks did it beat it on?-How much uncertainty remains?-Which profiles regressed?-Which deterministic checks failed?-Which judge scored it?-Was the backend real?-What did it cost?-What failure modes remain?-Can the trace prove all of that?-```+- What changed?+- Which baseline did it beat?+- Which held-out tasks did it beat it on?+- How much uncertainty remains?+- Which profiles regressed?+- Which deterministic checks failed?+- Which judge scored it?+- Was the backend real?+- What did it cost?+- What failure modes remain?+- Can the trace prove all of that? If the answer is missing, the candidate stays a candidate. - Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
show diff
diff --git a/src/content/posts/self-improving-stack-evaluation-gates.mdx b/src/content/posts/self-improving-stack-evaluation-gates.mdxindex 06564b3..6f14348 100644--- a/src/content/posts/self-improving-stack-evaluation-gates.mdx+++ b/src/content/posts/self-improving-stack-evaluation-gates.mdx@@ -57,18 +57,20 @@ A gate is a promotion policy. Let: -```text-b = baseline system-c = candidate system-x = scenario-p = agent profile cell-z = seed or replicate id-R = task reward or score-C = measured cost vector-T = trace integrity predicate-D_search = search split-D_holdout = held-out split-```+$$+\begin{aligned}+ b &= \text{baseline system} \\+ c &= \text{candidate system} \\+ x &= \text{scenario} \\+ p &= \text{agent profile cell} \\+ z &= \text{seed or replicate id} \\+ R &= \text{task reward or score} \\+ C &= \text{measured cost vector} \\+ T &= \text{trace integrity predicate} \\+ D_{\text{search}} &= \text{search split} \\+ D_{\text{holdout}} &= \text{held-out split}+\end{aligned}+$$ The gate is a function: @@ -129,9 +131,9 @@ Suppose the baseline sees one sample of tasks and the candidate sees another. A The paired comparison fixes that: -```text-delta_i = R(c, x_i, p_i, z_i) - R(b, x_i, p_i, z_i)-```+$$+\Delta_i = R(c,x_i,p_i,z_i)-R(b,x_i,p_i,z_i)+$$ where the candidate and baseline are evaluated on the same scenario, profile, and replicate. Then the question becomes: @@ -143,12 +145,11 @@ The median is useful because agent scores often have heavy tails. One catastroph A practical promotion rule: -```text-n_pairs >= n_min-LCB_95(median(delta)) > epsilon-```+$$+n_{\text{pairs}} \ge n_{\min},\qquad \operatorname{LCB}_{95}(\operatorname{median}(\Delta)) > \epsilon+$$ -`LCB_95` is the lower confidence bound. If the lower bound clears the threshold, the gate has evidence that the lift is not just random luck.+$\operatorname{LCB}_{95}$ is the lower confidence bound. If the lower bound clears the threshold, the gate has evidence that the lift is not just random luck. This is where bootstrap confidence intervals are useful. You resample paired deltas, compute the median for each resample, and inspect the lower quantile: @@ -157,10 +158,12 @@ delta = [delta_1, ..., delta_n] for r in 1..B: sample n deltas with replacement m_r = median(sample)--LCB_95 = quantile({m_r}, 0.025) // two-sided 95 interval ``` +$$+\operatorname{LCB}_{95}=Q_{0.025}(m_1,\ldots,m_B)+$$+ The gate promotes only when the pessimistic estimate is still good enough. ## Search Split Versus Holdout@@ -185,14 +188,18 @@ score(c, holdout) goes down So the gate needs an overfit check: -```text-gap_c = mean_score(c, search) - mean_score(c, holdout)-gap_b = mean_score(b, search) - mean_score(b, holdout)+$$+\begin{aligned}+ \operatorname{gap}_c &= \operatorname{mean\_score}(c,\text{search})-\operatorname{mean\_score}(c,\text{holdout}) \\+ \operatorname{gap}_b &= \operatorname{mean\_score}(b,\text{search})-\operatorname{mean\_score}(b,\text{holdout})+\end{aligned}+$$ +```text reject(c) if gap_c > gap_b + tau ``` -This says the candidate may look better on the search split, but it cannot be much more search-specialized than the baseline. `tau` is slack, not forgiveness. It accounts for sampling noise and legitimate split difficulty differences.+This says the candidate may look better on the search split, but it cannot be much more search-specialized than the baseline. $\tau$ is slack, not forgiveness. It accounts for sampling noise and legitimate split difficulty differences. The gate configuration has to be fixed before the candidate is scored: @@ -289,28 +296,30 @@ then evaluate the judge as a measurement instrument. Let: -```text-V(x, y, trace) = judge score-R(x, y) = true product outcome-```+$$+\begin{aligned}+V(x,y,\text{trace}) &= \text{judge score} \\+R(x,y) &= \text{true product outcome}+\end{aligned}+$$ A judge gate needs calibration: -```text-calibration_error = E[R | V = s] - s-```+$$+\operatorname{calibration\_error}=\mathbb{E}[R\mid V=s]-s+$$ and agreement: -```text-agreement = P(sign(V_a - V_b) = sign(H_a - H_b))-```+$$+\operatorname{agreement}=\mathbb{P}(\operatorname{sign}(V_a-V_b)=\operatorname{sign}(H_a-H_b))+$$ -where `H` is a human preference or trusted adjudicator. For ordinal rubrics, rank correlation is often more useful than raw score correlation:+where $H$ is a human preference or trusted adjudicator. For ordinal rubrics, rank correlation is often more useful than raw score correlation: -```text-rho = Spearman(V, H)-```+$$+\rho=\operatorname{Spearman}(V,H)+$$ A gate should track judge drift over time. If the judge model, rubric, prompt, or examples change, the score distribution can move even when the agent behavior does not. @@ -356,11 +365,15 @@ This prevents aggregate masking. A candidate can improve the mean while regressi Regression detection should combine effect size and statistical confidence: -```text-delta = current - baseline-d = CohenD(current_scores, baseline_scores)-p = WelchT(current_scores, baseline_scores)+$$+\begin{aligned}+ \Delta &= \text{current}-\text{baseline} \\+ d &= \operatorname{CohenD}(\text{current\_scores},\text{baseline\_scores}) \\+ p &= \operatorname{WelchT}(\text{current\_scores},\text{baseline\_scores})+\end{aligned}+$$ +```text regressed if delta < 0 and abs(d) >= d_min and p <= alpha ``` - Drafted the evaluation-gates post with held-out promotion math, scorecard cells, judge reliability, backend integrity, release confidence, and local Tangle package placement.
- Polished the evaluation-gates post by tightening profile-cell claims against the local AgentProfileCell schema, adding gate pre-registration invariants, clarifying bootstrap interval wording, and preserving the fail-closed promotion model.
- let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
show diff
diff --git a/src/content/posts/self-improving-stack-evaluation-gates.mdx b/src/content/posts/self-improving-stack-evaluation-gates.mdxindex da8d0d0..52f7a8b 100644--- a/src/content/posts/self-improving-stack-evaluation-gates.mdx+++ b/src/content/posts/self-improving-stack-evaluation-gates.mdx@@ -19,7 +19,9 @@ authors: date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+ - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+ - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-evaluation-gates-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-evaluation-gates-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-evaluation-gates-review' } - date: 2026-06-05@@ -41,15 +43,11 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -An optimizer can propose forever.+A green score is not a release decision. It is evidence entering a release policy, and in self-improving loops that distinction matters more than the optimizer. -The gate decides what becomes the system.+GEPA, MIPRO, SkillOpt, topology search, and meta-harness can all generate candidates forever. The gate decides which candidate becomes the system future agents inherit. If the gate is weak, the optimizer learns the gate. If the gate is honest, the optimizer has to improve the product. -That is why the gate is not an administrative detail after the interesting work. It is the objective boundary. GEPA, MIPRO, SkillOpt, runtime topology search, and meta-harness can all generate candidates. The gate decides which candidate is allowed to replace the baseline.--If the gate is weak, every optimizer learns the gate.--If the gate is honest, every optimizer has to improve the product.+This is why the gate is not an administrative detail after the interesting work. It is the objective boundary. ## What A Gate Is @@ -574,7 +572,7 @@ The final output is scored, but the branch, tool, verifier, and selector evidenc All failures collapse into one scalar, so the next optimizer has no diagnostic direction. -## Working Rule+## When A Candidate Can Ship The optimizer proposes. - 60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
- Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
- Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
- Research planning pass from a traced session.
Comments
PUBLIC_GISCUS_REPO,PUBLIC_GISCUS_REPO_ID,PUBLIC_GISCUS_CATEGORY, andPUBLIC_GISCUS_CATEGORY_IDin.env. See giscus.app to generate the IDs after you enable Discussions on the repo.