Traces Are The Training Data
Why self-improving agents need full trajectories, tool spans, analyst findings, provenance, and replay instead of final scores alone.
The Self Improving Stack series
Browse all 13 posts
- Topology Is The Missing Action Space
- The Gate Is The Optimizer
- Self-Improvement Needs A Safety Case
- When The Harness Has To Evolve
- Memory Is Not Automatically Learning
- Personas Are Content, Coordination Is Structure
- Optimization Theory For Agent Builders
- When The Model Itself Is Mutable
- Prompt Optimization Is Not The Whole Game
- Skills Are Trainable State
- Beat Random At Equal Compute First
- Traces Are The Training Data
- The Self-Improving Stack
The optimizer wants a score. I want the run, because a score only tells you that something happened while a trace preserves enough mechanism to explain what happened.
That difference is the difference between tuning a system and optimizing an unidentified projection.
A self-improving agent can only improve from the information it preserves. If the run record says “failed, score 0.42,” the optimizer can only infer weak global pressure. If the trace says the planner chose the wrong tool, the tool call used a stale argument, the retrieval span returned irrelevant context, the judge penalized a missing artifact, and the retry loop repeated the same action three times, the optimizer has a causal surface.
The trace is not decoration around the eval. The trace is the data.
The Information Loss Problem
An agent run is a trajectory:
where:
A final score from a fixed scorer is a projection:
That projection is intentionally lossy. It collapses a long sequence of decisions, calls, costs, artifacts, and observations into one number.
Optimization needs the lost variables.
The crude information-theory version:
When and the scorer is fixed, the score is a deterministic projection of the trajectory. By the data processing inequality, that projection cannot contain more information about the failure cause than the trajectory itself. Usually it contains dramatically less.
This does not mean every byte is equally useful. It means the system must preserve the variables that can explain responsible mechanism:
- which model answered
- which prompt was used
- which branch ran
- which tool was called
- which arguments were passed
- which observations came back
- which artifact changed
- which verifier judged it
- which budget was spent
- which failure class was assigned
Without those variables, the optimizer is moving an unidentified intervention against an unidentified mechanism.
Trace Versus Summary
A summary says:
The agent tried to use the API, failed, and produced a partial answer.
A trace says:
tool span:
toolName = github.search
args = { query: "HeldOutGate cost ceiling" }
result.count = 0
llm span:
model = ...
promptSha = ...
output = "No such API exists"
retrieval span:
hits = [...]
judge span:
dimension = source_grounding
score = 0.2
targetSpanId = ...
The summary is readable. The trace is inspectable.
Summaries are useful outputs of traces. They are not replacements for traces. Once the mechanism is compressed away, no later analyst can recover it.
What A Trace Must Capture
A useful agent trace has several layers.
Run identity
runIdscenarioIdcandidateIddatasetVersioncodeShapromptShamodelFingerprintseedenvFingerprintparentRunIdprojectIdchatIdlayer
This makes the run attributable. Without identity, the score cannot be tied to a candidate, commit, profile, prompt, model, or scenario.
Span tree
- agent span
- llm span
- tool span
- retrieval span
- judge span
- sandbox span
- custom span
The span tree gives causality and nesting. A tool call can be under a planner branch. A judge can target a specific span. A sandbox failure can be tied to the code artifact it ran.
Events
budget_decrementbudget_breachstate_mutationpolicy_violationredaction_appliederrorcustom
Events capture point-in-time facts that are not whole spans.
Budget ledger
tokenswallMscallsusdremainingbreached
This lets the evaluator distinguish a smarter policy from a more expensive one.
Artifacts
diffsfileslogsscreenshots- test reports
- retrieved documents
- judge reports
Artifacts make traces material. A span saying “patched file” is weaker than an artifact hash and storage pointer for the patch.
Outcome
scorepassfailureClassnotes
The outcome is still necessary. It is the label. It is just not enough by itself.
Trace Granularity
The trace is detailed enough when it can answer a counterfactual:
If this action, observation, tool result, verifier result, or budget event had changed, would the outcome have changed?
Too coarse:
agent failed at research
This does not identify whether the failure was query formation, source choice, stale retrieval, missing credentials, synthesis, or judge mismatch.
Too fine:
every token, cursor movement, and private secret copied into a permanent record
This increases cost and risk without necessarily improving diagnosis.
The target is sufficient structure:
- enough fields to localize the responsible mechanism
- enough ids to join evidence across run record, trace, artifact, scorecard, and finding
- enough redaction to preserve privacy and auditability
Raw Provider Capture
Structured LLM spans record intent.
Raw provider capture records what actually went over the wire.
That distinction matters. A proxy can report a different model name than the model that answered. A streaming parser can drop a field. Token usage can be missing. A retry can produce the final answer while the span only shows the last attempt. A judge can run on stale output.
So the trace system needs both:
LlmSpan:
model
messages
output
token usage
cost
RawProviderEvent:
request body
response body
endpoint
baseUrl
provider
model
attemptIndex
statusCode
durationMs
redactedFields
The raw event is not for dashboards. It is for forensics, replay, and audit.
The rule:
Every LLM span that affects a score needs matching raw request evidence.
If the structured span exists but the raw provider event is missing, the run is not launch-grade evidence.
Replay
Raw capture turns old runs into reusable experimental material.
If a run has recorded request and response events, a replay cache can map:
Canonical request → captured response.
That enables:
- judge replay without new model calls
- rubric comparison on identical outputs
- determinism audits
- failure triage without spending fresh tokens
- regression analysis across new evaluators
Replay is especially important for judge calibration. If two judges score different fresh samples, disagreement may be sampling noise. If two judges score the same replayed outputs, disagreement is evaluator behavior.
The replay miss policy matters:
throwfallback_to_networkfail_closed
For determinism audits and promotion gates, fail closed. A silent network fallback turns replay into a new experiment.
Trace Integrity
Trace capture is not binary. It can fail partially.
The integrity check is:
- run exists
- llm span count >= minimum
- tool span count >= minimum, when tools are expected
- judge span count >= minimum, when judges are expected
- raw provider events exist
- raw provider events cover llm spans
- outcome exists
In compact form:
trace_integrity(tau) =
run_present
and expected_spans_present
and raw_coverage_ok
and outcome_present
Promotion gates treat missing trace evidence as missing evidence, not as a neutral value.
The highest-cost failure is an orphan LLM span:
- structured llm span exists
- raw request is missing
That usually means capture was wired to the wrong sink, the call bypassed the instrumented client, or the route changed under the harness.
Backend Integrity
Trace integrity asks whether the run was captured.
Backend integrity asks whether the backend was real.
The minimal signal:
stub_record = tokenUsage.input == 0 and tokenUsage.output == 0
Then:
- all stub records -> reject
- mixed real and stub records -> quarantine or reject in CI
- real tokens with zero cost -> cost ledger bug
This is not a small bookkeeping issue. If a campaign runs against stubs, every downstream statistic is corrupted: scorecard deltas, held-out gates, analyst findings, and optimizer decisions.
The right interpretation of a stub campaign is:
We did not evaluate the agent.
not:
The agent failed every task.
Analyst Findings
The trace is raw material. Analysts turn it into structured diagnosis.
A useful finding has:
finding_idanalyst_id- severity
- area
- claim
- rationale
evidence_refsrecommended_actionvalidation_plan- confidence
- subject
The evidence_refs field is the key. A finding without a span, event, artifact, metric, or prior finding reference is an unsupported assertion.
The analyst layer supports multiple lenses:
failure-modeknowledge-gapknowledge-poisoningimprovement
Those lenses answer different questions:
- What failed?
- What knowledge was missing?
- What knowledge was wrong or harmful?
- What change is worth testing next?
That is how traces become optimizer input.
tau → findings → candidate mutation → eval → gate
The output is not “a summary of the run.” The output is a set of attributed hypotheses with validation plans.
The Leakage Firewall
Traces can create hidden oracles.
The system has to separate runtime observations from evaluation labels.
Allowed at runtime:
- tool outputs
- compiler errors
- test failures available to the product
- retrieval results
- user feedback
- budget remaining
- branch status
Forbidden as runtime steering signals:
- holdout labels
- private judge scores
- answer keys
- post-hoc evaluator rationales
- promotion decisions
- human review notes unavailable in production
The rule:
If the production system cannot observe it, the runtime policy cannot use it.
The optimizer can train from eval traces after the run. The runtime cannot peek at the gate during the run.
This matters for GEPA-style prompt optimization, skill optimization, and topology search. They may use trace-derived feedback to propose candidates. They may not smuggle holdout answers into the candidate.
Privacy And Redaction
Traces are powerful because they are detailed.
That also makes them dangerous.
A trace can contain credentials, user data, private documents, file paths, source code, prompts, screenshots, and integration responses.
The trace system needs two simultaneous properties:
- enough detail for causality
- enough redaction for safety
Redaction has to happen at capture time for obvious secrets:
AuthorizationX-Api-KeyCookiepasswordsecrettokenaccess_tokenrefresh_token
But redaction is not just deletion. It records what was removed:
redactedFields = [...]
That lets a reviewer distinguish “the tool never sent auth” from “auth existed but was redacted.”
Over-redaction destroys causality. Under-redaction leaks data. The practical compromise is typed redaction plus artifact-level access control.
Where OpenTelemetry Fits
OpenTelemetry is the right outer shape for distributed traces.
As of June 6, 2026, the OpenTelemetry generative AI semantic conventions are marked Development and include GenAI model spans, agent spans, events, metrics, exceptions, OpenAI conventions, Anthropic conventions, and MCP conventions.
That is useful common infrastructure. It gives agent traces a path into existing collectors, dashboards, retention policies, and incident tooling.
But agent self-improvement needs more than generic spans. It needs first-class concepts that ordinary service traces do not enforce:
- candidate id
- scenario id
- split tag
- prompt hash
- config hash
- failure class
- judge verdict
- artifact hash
- budget ledger
- profile cell
- raw provider event
- promotion decision
- analyst finding
The right design is not “OTel or agent schema.” It is:
- OTel-compatible transport
- agent-specific schema
- promotion-grade integrity checks
Where Tangle Fits
Local package audit on June 6, 2026:
@tangle-network/agent-eval@0.34.1@tangle-network/agent-runtime@0.26.0
agent-eval provides the trace and analysis layer:
TraceSchema v1:Run,Span,TraceEvent,BudgetLedgerEntry,Artifact, and failure taxonomy.TraceEmitter: run lifecycle, hierarchical span helpers, run-complete hooks.RawProviderSink: request, response, and error capture with redaction and retry attempt indexes.assertRunCaptured: span, raw event, raw coverage, and outcome integrity.ReplayCacheandcreateReplayFetch: replay captured provider calls.RunRecord: promotion-grade analysis row with model snapshot, prompt hash, config hash, commit, cost, token usage, split tag, outcome, and optionalAgentProfileCell.InMemoryTraceStoreandFileSystemTraceStore: in-process and append-only trace persistence.AnalystRegistry: runs isolated analysts over trace stores, run records, artifacts, judge inputs, or custom inputs.AnalystFinding: content-addressed findings with evidence references and validation plans.OtlpFileTraceStore: production trace consumption path.
agent-runtime provides execution traces:
runLoop: emits topology events for loop start, iteration start, iteration dispatch, iteration end, decisions, and loop end.LoopTraceEvent: records driver, agent run names, task hashes, iteration placement, history length, winner, cost, duration, and iteration count.createRefineDriverandcreateFanoutVoteDriver: create different trace shapes for sequential and parallel compute.- conversation journals: preserve multi-turn runs, halt reasons, cost caps, turn order, deterministic turn ids, and resumability.
- OTLP exporter: ships loop trace events to an OpenTelemetry collector without adding a full SDK dependency.
This split matters:
- 01 runtime emits behavior
- 02 eval preserves and analyzes behavior
- 03 gates decide whether behavior can ship
How Trace Systems Lie
Score-only learning
The optimizer sees a scalar and no mechanism.
Summary collapse
An LLM summary replaces the span tree, tool arguments, artifacts, and raw events.
Orphan spans
Structured spans exist without raw provider evidence.
Backend blindness
The campaign ran against stubs or partial backend failure.
Trace amnesia
The final answer is captured, but failed branches, retries, and rejected candidates are missing.
Judge leakage
Private evaluator labels or rationales leak into runtime policy.
Artifact loss
The trace references a patch, screenshot, log, or retrieved document that no longer exists.
Redaction erasure
Sensitive fields are removed without recording what was removed, destroying the ability to diagnose missing auth or missing tenant scope.
Unjoined identity
Run records, traces, scorecard cells, and analyst findings use different ids, so the evidence cannot be joined.
When A Trace Is Training Data
Do not optimize from final scores alone.
Preserve the trajectory.
An agent trace must answer:
- Who ran?
- Against which scenario?
- With which model and prompt?
- Through which topology?
- Which actions were taken?
- Which observations came back?
- Which artifacts changed?
- Which verifier judged them?
- What did it cost?
- What failed?
- Which evidence supports that diagnosis?
- Can the run be replayed?
- Can the gate trust the capture?
If it cannot answer those questions, it is not training data for a self-improving agent. It is an anecdote about a run.
The trace is where agent behavior becomes learnable.
Source Trail
Source freshness checked on 2026-06-06.
- ReAct: Synergizing Reasoning and Acting in Language Models, checked June 6, 2026.
- Reflexion: Language Agents with Verbal Reinforcement Learning, checked June 6, 2026.
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing, checked June 6, 2026.
- Can Large Language Models Really Improve by Self-critiquing Their Own Plans?, checked June 6, 2026.
- OpenTelemetry semantic conventions for generative AI systems, checked June 6, 2026.
- OpenTelemetry semantic conventions for generative client AI spans, checked June 6, 2026.
- OpenTelemetry semantic conventions for GenAI agent and framework spans, checked June 6, 2026.
- Local
@tangle-network/agent-eval@0.34.1source audit: trace schema,TraceEmitter,RawProviderSink,assertRunCaptured, replay,RunRecord,AnalystRegistry,AnalystFinding,OtlpFileTraceStore, June 6, 2026. - Local
@tangle-network/agent-runtime@0.26.0source audit:runLoop,LoopTraceEvent, refine/fanout drivers, conversation journals, OTLP exporter, June 6, 2026.
Revision history9revisions
- Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
show diff
diff --git a/src/content/posts/self-improving-stack-trace-systems.mdx b/src/content/posts/self-improving-stack-trace-systems.mdxindex 953b091..569e40d 100644--- a/src/content/posts/self-improving-stack-trace-systems.mdx+++ b/src/content/posts/self-improving-stack-trace-systems.mdx@@ -47,6 +47,8 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'+ The optimizer wants a score. I want the run, because a score only tells you that something happened while a trace preserves enough mechanism to explain what happened. That difference is the difference between tuning a system and optimizing an unidentified projection.@@ -95,18 +97,16 @@ When $\text{score}=R(\tau)$ and the scorer is fixed, the score is a deterministi This does not mean every byte is equally useful. It means the system must preserve the variables that can explain responsible mechanism: -```text-which model answered-which prompt was used-which branch ran-which tool was called-which arguments were passed-which observations came back-which artifact changed-which verifier judged it-which budget was spent-which failure class was assigned-```+- which model answered+- which prompt was used+- which branch ran+- which tool was called+- which arguments were passed+- which observations came back+- which artifact changed+- which verifier judged it+- which budget was spent+- which failure class was assigned Without those variables, the optimizer is moving an unidentified intervention against an unidentified mechanism. @@ -114,9 +114,7 @@ Without those variables, the optimizer is moving an unidentified intervention ag A summary says: -```text-The agent tried to use the API, failed, and produced a partial answer.-```+> The agent tried to use the API, failed, and produced a partial answer. A trace says: @@ -150,87 +148,75 @@ A useful agent trace has several layers. **Run identity** -```text-runId-scenarioId-candidateId-datasetVersion-codeSha-promptSha-modelFingerprint-seed-envFingerprint-parentRunId-projectId-chatId-layer-```+- `runId`+- `scenarioId`+- `candidateId`+- `datasetVersion`+- `codeSha`+- `promptSha`+- `modelFingerprint`+- `seed`+- `envFingerprint`+- `parentRunId`+- `projectId`+- `chatId`+- `layer` This makes the run attributable. Without identity, the score cannot be tied to a candidate, commit, profile, prompt, model, or scenario. **Span tree** -```text-agent span-llm span-tool span-retrieval span-judge span-sandbox span-custom span-```+- agent span+- llm span+- tool span+- retrieval span+- judge span+- sandbox span+- custom span The span tree gives causality and nesting. A tool call can be under a planner branch. A judge can target a specific span. A sandbox failure can be tied to the code artifact it ran. **Events** -```text-budget_decrement-budget_breach-state_mutation-policy_violation-redaction_applied-error-custom-```+- `budget_decrement`+- `budget_breach`+- `state_mutation`+- `policy_violation`+- `redaction_applied`+- `error`+- `custom` Events capture point-in-time facts that are not whole spans. **Budget ledger** -```text-tokens-wallMs-calls-usd-remaining-breached-```+- `tokens`+- `wallMs`+- `calls`+- `usd`+- `remaining`+- `breached` This lets the evaluator distinguish a smarter policy from a more expensive one. **Artifacts** -```text-diffs-files-logs-screenshots-test reports-retrieved documents-judge reports-```+- `diffs`+- `files`+- `logs`+- `screenshots`+- test reports+- retrieved documents+- judge reports Artifacts make traces material. A span saying "patched file" is weaker than an artifact hash and storage pointer for the patch. **Outcome** -```text-score-pass-failureClass-notes-```+- `score`+- `pass`+- `failureClass`+- `notes` The outcome is still necessary. It is the label. It is just not enough by itself. @@ -238,34 +224,25 @@ The outcome is still necessary. It is the label. It is just not enough by itself The trace is detailed enough when it can answer a counterfactual: -```text-If this action, observation, tool result, verifier result, or budget event had changed,-would the outcome have changed?-```+> If this action, observation, tool result, verifier result, or budget event had changed, would the outcome have changed? Too coarse: -```text-agent failed at research-```+> agent failed at research This does not identify whether the failure was query formation, source choice, stale retrieval, missing credentials, synthesis, or judge mismatch. Too fine: -```text-every token, cursor movement, and private secret copied into a permanent record-```+> every token, cursor movement, and private secret copied into a permanent record This increases cost and risk without necessarily improving diagnosis. The target is sufficient structure: -```text-enough fields to localize the responsible mechanism-enough ids to join evidence across run record, trace, artifact, scorecard, and finding-enough redaction to preserve privacy and auditability-```+- enough fields to localize the responsible mechanism+- enough ids to join evidence across run record, trace, artifact, scorecard, and finding+- enough redaction to preserve privacy and auditability ## Raw Provider Capture @@ -302,9 +279,7 @@ The raw event is not for dashboards. It is for forensics, replay, and audit. The rule: -```text-Every LLM span that affects a score needs matching raw request evidence.-```+> Every LLM span that affects a score needs matching raw request evidence. If the structured span exists but the raw provider event is missing, the run is not launch-grade evidence. @@ -314,9 +289,7 @@ Raw capture turns old runs into reusable experimental material. If a run has recorded request and response events, a replay cache can map: -```text-canonical_request -> captured_response-```+**Canonical request → captured response.** That enables: @@ -330,11 +303,9 @@ Replay is especially important for judge calibration. If two judges score differ The replay miss policy matters: -```text-throw-fallback_to_network-fail_closed-```+- `throw`+- `fallback_to_network`+- `fail_closed` For determinism audits and promotion gates, fail closed. A silent network fallback turns replay into a new experiment. @@ -344,15 +315,13 @@ Trace capture is not binary. It can fail partially. The integrity check is: -```text-run exists-llm span count >= minimum-tool span count >= minimum, when tools are expected-judge span count >= minimum, when judges are expected-raw provider events exist-raw provider events cover llm spans-outcome exists-```+- run exists+- llm span count >= minimum+- tool span count >= minimum, when tools are expected+- judge span count >= minimum, when judges are expected+- raw provider events exist+- raw provider events cover llm spans+- outcome exists In compact form: @@ -368,10 +337,8 @@ Promotion gates treat missing trace evidence as missing evidence, not as a neutr The highest-cost failure is an orphan LLM span: -```text-structured llm span exists-raw request is missing-```+- structured llm span exists+- raw request is missing That usually means capture was wired to the wrong sink, the call bypassed the instrumented client, or the route changed under the harness. @@ -389,25 +356,19 @@ stub_record = tokenUsage.input == 0 and tokenUsage.output == 0 Then: -```text-all stub records -> reject-mixed real and stub records -> quarantine or reject in CI-real tokens with zero cost -> cost ledger bug-```+- all stub records -> reject+- mixed real and stub records -> quarantine or reject in CI+- real tokens with zero cost -> cost ledger bug This is not a small bookkeeping issue. If a campaign runs against stubs, every downstream statistic is corrupted: scorecard deltas, held-out gates, analyst findings, and optimizer decisions. The right interpretation of a stub campaign is: -```text-We did not evaluate the agent.-```+> We did not evaluate the agent. not: -```text-The agent failed every task.-```+> The agent failed every task. ## Analyst Findings @@ -415,30 +376,26 @@ The trace is raw material. Analysts turn it into structured diagnosis. A useful finding has: -```text-finding_id-analyst_id-severity-area-claim-rationale-evidence_refs-recommended_action-validation_plan-confidence-subject-```+- `finding_id`+- `analyst_id`+- severity+- area+- claim+- rationale+- `evidence_refs`+- `recommended_action`+- `validation_plan`+- confidence+- subject The `evidence_refs` field is the key. A finding without a span, event, artifact, metric, or prior finding reference is an unsupported assertion. The analyst layer supports multiple lenses: -```text-failure-mode-knowledge-gap-knowledge-poisoning-improvement-```+- `failure-mode`+- `knowledge-gap`+- `knowledge-poisoning`+- `improvement` Those lenses answer different questions: @@ -449,9 +406,7 @@ Those lenses answer different questions: That is how traces become optimizer input. -```text-tau -> findings -> candidate mutation -> eval -> gate-```+`tau` → findings → candidate mutation → eval → gate The output is not "a summary of the run." The output is a set of attributed hypotheses with validation plans. @@ -463,32 +418,26 @@ The system has to separate runtime observations from evaluation labels. Allowed at runtime: -```text-tool outputs-compiler errors-test failures available to the product-retrieval results-user feedback-budget remaining-branch status-```+- tool outputs+- compiler errors+- test failures available to the product+- retrieval results+- user feedback+- budget remaining+- branch status Forbidden as runtime steering signals: -```text-holdout labels-private judge scores-answer keys-post-hoc evaluator rationales-promotion decisions-human review notes unavailable in production-```+- holdout labels+- private judge scores+- answer keys+- post-hoc evaluator rationales+- promotion decisions+- human review notes unavailable in production The rule: -```text-If the production system cannot observe it, the runtime policy cannot use it.-```+> If the production system cannot observe it, the runtime policy cannot use it. The optimizer can train from eval traces after the run. The runtime cannot peek at the gate during the run. @@ -504,29 +453,23 @@ A trace can contain credentials, user data, private documents, file paths, sourc The trace system needs two simultaneous properties: -```text-enough detail for causality-enough redaction for safety-```+- enough detail for causality+- enough redaction for safety Redaction has to happen at capture time for obvious secrets: -```text-Authorization-X-Api-Key-Cookie-password-secret-token-access_token-refresh_token-```+- `Authorization`+- `X-Api-Key`+- `Cookie`+- `password`+- `secret`+- `token`+- `access_token`+- `refresh_token` But redaction is not just deletion. It records what was removed: -```text-redactedFields = [...]-```+`redactedFields = [...]` That lets a reviewer distinguish "the tool never sent auth" from "auth existed but was redacted." @@ -542,38 +485,32 @@ That is useful common infrastructure. It gives agent traces a path into existing But agent self-improvement needs more than generic spans. It needs first-class concepts that ordinary service traces do not enforce: -```text-candidate id-scenario id-split tag-prompt hash-config hash-failure class-judge verdict-artifact hash-budget ledger-profile cell-raw provider event-promotion decision-analyst finding-```+- candidate id+- scenario id+- split tag+- prompt hash+- config hash+- failure class+- judge verdict+- artifact hash+- budget ledger+- profile cell+- raw provider event+- promotion decision+- analyst finding The right design is not "OTel or agent schema." It is: -```text-OTel-compatible transport-agent-specific schema-promotion-grade integrity checks-```+- OTel-compatible transport+- agent-specific schema+- promotion-grade integrity checks ## Where Tangle Fits Local package audit on June 6, 2026: -```text-@tangle-network/agent-eval@0.34.1-@tangle-network/agent-runtime@0.26.0-```+- `@tangle-network/agent-eval@0.34.1`+- `@tangle-network/agent-runtime@0.26.0` `agent-eval` provides the trace and analysis layer: @@ -598,11 +535,11 @@ Local package audit on June 6, 2026: This split matters: -```text-runtime emits behavior-eval preserves and analyzes behavior-gates decide whether behavior can ship-```+<Steps layout='flow' items={[+ { title: 'runtime emits behavior' },+ { title: 'eval preserves and analyzes behavior' },+ { title: 'gates decide whether behavior can ship' }+]} /> ## How Trace Systems Lie @@ -650,21 +587,19 @@ Preserve the trajectory. An agent trace must answer: -```text-Who ran?-Against which scenario?-With which model and prompt?-Through which topology?-Which actions were taken?-Which observations came back?-Which artifacts changed?-Which verifier judged them?-What did it cost?-What failed?-Which evidence supports that diagnosis?-Can the run be replayed?-Can the gate trust the capture?-```+- Who ran?+- Against which scenario?+- With which model and prompt?+- Through which topology?+- Which actions were taken?+- Which observations came back?+- Which artifacts changed?+- Which verifier judged them?+- What did it cost?+- What failed?+- Which evidence supports that diagnosis?+- Can the run be replayed?+- Can the gate trust the capture? If it cannot answer those questions, it is not training data for a self-improving agent. It is an anecdote about a run. - Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
show diff
diff --git a/src/content/posts/self-improving-stack-trace-systems.mdx b/src/content/posts/self-improving-stack-trace-systems.mdxindex 3274500..27e26f0 100644--- a/src/content/posts/self-improving-stack-trace-systems.mdx+++ b/src/content/posts/self-improving-stack-trace-systems.mdx@@ -57,25 +57,27 @@ The trace is not decoration around the eval. The trace is the data. An agent run is a trajectory: -```text-tau = (x, s_0, a_1, o_1, s_1, ..., a_T, o_T, y)-```+$$+\tau=(x,s_0,a_1,o_1,s_1,\ldots,a_T,o_T,y)+$$ where: -```text-x = task-s_t = internal and external state-a_t = action-o_t = observation-y = outcome-```+$$+\begin{aligned}+ x &= \text{task} \\+ s_t &= \text{internal and external state} \\+ a_t &= \text{action} \\+ o_t &= \text{observation} \\+ y &= \text{outcome}+\end{aligned}+$$ A final score from a fixed scorer is a projection: -```text-score = R(tau)-```+$$+\text{score}=R(\tau)+$$ That projection is intentionally lossy. It collapses a long sequence of decisions, calls, costs, artifacts, and observations into one number. @@ -83,11 +85,11 @@ Optimization needs the lost variables. The crude information-theory version: -```text-I(tau; failure_cause) >= I(score; failure_cause)-```+$$+I(\tau;\text{failure cause})\ge I(\text{score};\text{failure cause})+$$ -When `score = R(tau)` and the scorer is fixed, the score is a deterministic projection of the trajectory. By the data processing inequality, that projection cannot contain more information about the failure cause than the trajectory itself. Usually it contains dramatically less.+When $\text{score}=R(\tau)$ and the scorer is fixed, the score is a deterministic projection of the trajectory. By the data processing inequality, that projection cannot contain more information about the failure cause than the trajectory itself. Usually it contains dramatically less. This does not mean every byte is equally useful. It means the system must preserve the variables that can explain responsible mechanism: - Drafted the trace-systems post with formal trajectory notation, span ontology, raw provider capture, replay, trace integrity, analyst findings, leakage firewalls, and local Tangle package placement.
- Polished the trace-systems post by adding a trace granularity test, tightening the information-loss claim to a fixed scorer, correcting loop trace event details, and adding trace store surfaces from the local agent-eval audit.
- let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
show diff
diff --git a/src/content/posts/self-improving-stack-trace-systems.mdx b/src/content/posts/self-improving-stack-trace-systems.mdxindex 5f2bcce..9a9e631 100644--- a/src/content/posts/self-improving-stack-trace-systems.mdx+++ b/src/content/posts/self-improving-stack-trace-systems.mdx@@ -19,7 +19,9 @@ authors: date: 2026-06-06 - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 } - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+ - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+ - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-trace-systems-rewrite' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-trace-systems-publish' } - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-trace-systems-review' } - date: 2026-06-05@@ -41,17 +43,13 @@ supporting_trace_ids: - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -Scores tell you that something happened.--Traces tell you what happened.+The optimizer wants a score. I want the run, because a score only tells you that something happened while a trace preserves enough mechanism to explain what happened. That difference is the difference between tuning a system and optimizing an unidentified projection. A self-improving agent can only improve from the information it preserves. If the run record says "failed, score 0.42," the optimizer can only infer weak global pressure. If the trace says the planner chose the wrong tool, the tool call used a stale argument, the retrieval span returned irrelevant context, the judge penalized a missing artifact, and the retry loop repeated the same action three times, the optimizer has a causal surface. -The trace is not decoration around the eval.--The trace is the data.+The trace is not decoration around the eval. The trace is the data. ## The Information Loss Problem @@ -600,7 +598,7 @@ eval preserves and analyzes behavior gates decide whether behavior can ship ``` -## Failure Modes+## How Trace Systems Lie **Score-only learning** @@ -638,7 +636,7 @@ Sensitive fields are removed without recording what was removed, destroying the Run records, traces, scorecard cells, and analyst findings use different ids, so the evidence cannot be joined. -## Working Rule+## When A Trace Is Training Data Do not optimize from final scores alone. - 60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
- Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
- Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
- Research planning pass from a traced session.
Comments
PUBLIC_GISCUS_REPO,PUBLIC_GISCUS_REPO_ID,PUBLIC_GISCUS_CATEGORY, andPUBLIC_GISCUS_CATEGORY_IDin.env. See giscus.app to generate the IDs after you enable Discussions on the repo.