Traces Are The Training Data

Why self-improving agents need full trajectories, tool spans, analyst findings, provenance, and replay instead of final scores alone.

The Self Improving Stack series

← Beat Random At Equal Compute First Next → The Self-Improving Stack
Browse all 13 posts
  1. Jun 2026 Topology Is The Missing Action Space
  2. Jun 2026 The Gate Is The Optimizer
  3. Jun 2026 Self-Improvement Needs A Safety Case
  4. Jun 2026 When The Harness Has To Evolve
  5. Jun 2026 Memory Is Not Automatically Learning
  6. Jun 2026 Personas Are Content, Coordination Is Structure
  7. Jun 2026 Optimization Theory For Agent Builders
  8. Jun 2026 When The Model Itself Is Mutable
  9. Jun 2026 Prompt Optimization Is Not The Whole Game
  10. Jun 2026 Skills Are Trainable State
  11. Jun 2026 Beat Random At Equal Compute First
  12. Jun 2026 Traces Are The Training Data
  13. Jun 2026 The Self-Improving Stack
Authored by
outlineGPT-5.5draftGPT-5.5polishGPT-5.5reviewGPT-5.5publishGPT-5.5rewriteGPT-5.5polishGPT-5.5polishGPT-6-luna

The optimizer wants a score. I want the run, because a score only tells you that something happened while a trace preserves enough mechanism to explain what happened.

That difference is the difference between tuning a system and optimizing an unidentified projection.

A self-improving agent can only improve from the information it preserves. If the run record says “failed, score 0.42,” the optimizer can only infer weak global pressure. If the trace says the planner chose the wrong tool, the tool call used a stale argument, the retrieval span returned irrelevant context, the judge penalized a missing artifact, and the retry loop repeated the same action three times, the optimizer has a causal surface.

The trace is not decoration around the eval. The trace is the data.

The Information Loss Problem

An agent run is a trajectory:

τ=(x,s0,a1,o1,s1,…,aT,oT,y)\tau=(x,s_0,a_1,o_1,s_1,\ldots,a_T,o_T,y)

where:

x=taskst=internal and external stateat=actionot=observationy=outcome\begin{aligned} x &= \text{task} \\ s_t &= \text{internal and external state} \\ a_t &= \text{action} \\ o_t &= \text{observation} \\ y &= \text{outcome} \end{aligned}

A final score from a fixed scorer is a projection:

score=R(τ)\text{score}=R(\tau)

That projection is intentionally lossy. It collapses a long sequence of decisions, calls, costs, artifacts, and observations into one number.

Optimization needs the lost variables.

The crude information-theory version:

I(τ;failure cause)≥I(score;failure cause)I(\tau;\text{failure cause})\ge I(\text{score};\text{failure cause})

When score=R(τ)\text{score}=R(\tau) and the scorer is fixed, the score is a deterministic projection of the trajectory. By the data processing inequality, that projection cannot contain more information about the failure cause than the trajectory itself. Usually it contains dramatically less.

This does not mean every byte is equally useful. It means the system must preserve the variables that can explain responsible mechanism:

  • which model answered
  • which prompt was used
  • which branch ran
  • which tool was called
  • which arguments were passed
  • which observations came back
  • which artifact changed
  • which verifier judged it
  • which budget was spent
  • which failure class was assigned

Without those variables, the optimizer is moving an unidentified intervention against an unidentified mechanism.

Trace Versus Summary

A summary says:

The agent tried to use the API, failed, and produced a partial answer.

A trace says:

tool span:
  toolName = github.search
  args = { query: "HeldOutGate cost ceiling" }
  result.count = 0

llm span:
  model = ...
  promptSha = ...
  output = "No such API exists"

retrieval span:
  hits = [...]

judge span:
  dimension = source_grounding
  score = 0.2
  targetSpanId = ...

The summary is readable. The trace is inspectable.

Summaries are useful outputs of traces. They are not replacements for traces. Once the mechanism is compressed away, no later analyst can recover it.

What A Trace Must Capture

A useful agent trace has several layers.

Run identity

  • runId
  • scenarioId
  • candidateId
  • datasetVersion
  • codeSha
  • promptSha
  • modelFingerprint
  • seed
  • envFingerprint
  • parentRunId
  • projectId
  • chatId
  • layer

This makes the run attributable. Without identity, the score cannot be tied to a candidate, commit, profile, prompt, model, or scenario.

Span tree

  • agent span
  • llm span
  • tool span
  • retrieval span
  • judge span
  • sandbox span
  • custom span

The span tree gives causality and nesting. A tool call can be under a planner branch. A judge can target a specific span. A sandbox failure can be tied to the code artifact it ran.

Events

  • budget_decrement
  • budget_breach
  • state_mutation
  • policy_violation
  • redaction_applied
  • error
  • custom

Events capture point-in-time facts that are not whole spans.

Budget ledger

  • tokens
  • wallMs
  • calls
  • usd
  • remaining
  • breached

This lets the evaluator distinguish a smarter policy from a more expensive one.

Artifacts

  • diffs
  • files
  • logs
  • screenshots
  • test reports
  • retrieved documents
  • judge reports

Artifacts make traces material. A span saying “patched file” is weaker than an artifact hash and storage pointer for the patch.

Outcome

  • score
  • pass
  • failureClass
  • notes

The outcome is still necessary. It is the label. It is just not enough by itself.

Trace Granularity

The trace is detailed enough when it can answer a counterfactual:

If this action, observation, tool result, verifier result, or budget event had changed, would the outcome have changed?

Too coarse:

agent failed at research

This does not identify whether the failure was query formation, source choice, stale retrieval, missing credentials, synthesis, or judge mismatch.

Too fine:

every token, cursor movement, and private secret copied into a permanent record

This increases cost and risk without necessarily improving diagnosis.

The target is sufficient structure:

  • enough fields to localize the responsible mechanism
  • enough ids to join evidence across run record, trace, artifact, scorecard, and finding
  • enough redaction to preserve privacy and auditability

Raw Provider Capture

Structured LLM spans record intent.

Raw provider capture records what actually went over the wire.

That distinction matters. A proxy can report a different model name than the model that answered. A streaming parser can drop a field. Token usage can be missing. A retry can produce the final answer while the span only shows the last attempt. A judge can run on stale output.

So the trace system needs both:

LlmSpan:
  model
  messages
  output
  token usage
  cost

RawProviderEvent:
  request body
  response body
  endpoint
  baseUrl
  provider
  model
  attemptIndex
  statusCode
  durationMs
  redactedFields

The raw event is not for dashboards. It is for forensics, replay, and audit.

The rule:

Every LLM span that affects a score needs matching raw request evidence.

If the structured span exists but the raw provider event is missing, the run is not launch-grade evidence.

Replay

Raw capture turns old runs into reusable experimental material.

If a run has recorded request and response events, a replay cache can map:

Canonical request → captured response.

That enables:

  • judge replay without new model calls
  • rubric comparison on identical outputs
  • determinism audits
  • failure triage without spending fresh tokens
  • regression analysis across new evaluators

Replay is especially important for judge calibration. If two judges score different fresh samples, disagreement may be sampling noise. If two judges score the same replayed outputs, disagreement is evaluator behavior.

The replay miss policy matters:

  • throw
  • fallback_to_network
  • fail_closed

For determinism audits and promotion gates, fail closed. A silent network fallback turns replay into a new experiment.

Trace Integrity

Trace capture is not binary. It can fail partially.

The integrity check is:

  • run exists
  • llm span count >= minimum
  • tool span count >= minimum, when tools are expected
  • judge span count >= minimum, when judges are expected
  • raw provider events exist
  • raw provider events cover llm spans
  • outcome exists

In compact form:

trace_integrity(tau) =
  run_present
  and expected_spans_present
  and raw_coverage_ok
  and outcome_present

Promotion gates treat missing trace evidence as missing evidence, not as a neutral value.

The highest-cost failure is an orphan LLM span:

  • structured llm span exists
  • raw request is missing

That usually means capture was wired to the wrong sink, the call bypassed the instrumented client, or the route changed under the harness.

Backend Integrity

Trace integrity asks whether the run was captured.

Backend integrity asks whether the backend was real.

The minimal signal:

stub_record = tokenUsage.input == 0 and tokenUsage.output == 0

Then:

  • all stub records -> reject
  • mixed real and stub records -> quarantine or reject in CI
  • real tokens with zero cost -> cost ledger bug

This is not a small bookkeeping issue. If a campaign runs against stubs, every downstream statistic is corrupted: scorecard deltas, held-out gates, analyst findings, and optimizer decisions.

The right interpretation of a stub campaign is:

We did not evaluate the agent.

not:

The agent failed every task.

Analyst Findings

The trace is raw material. Analysts turn it into structured diagnosis.

A useful finding has:

  • finding_id
  • analyst_id
  • severity
  • area
  • claim
  • rationale
  • evidence_refs
  • recommended_action
  • validation_plan
  • confidence
  • subject

The evidence_refs field is the key. A finding without a span, event, artifact, metric, or prior finding reference is an unsupported assertion.

The analyst layer supports multiple lenses:

  • failure-mode
  • knowledge-gap
  • knowledge-poisoning
  • improvement

Those lenses answer different questions:

  • What failed?
  • What knowledge was missing?
  • What knowledge was wrong or harmful?
  • What change is worth testing next?

That is how traces become optimizer input.

tau → findings → candidate mutation → eval → gate

The output is not “a summary of the run.” The output is a set of attributed hypotheses with validation plans.

The Leakage Firewall

Traces can create hidden oracles.

The system has to separate runtime observations from evaluation labels.

Allowed at runtime:

  • tool outputs
  • compiler errors
  • test failures available to the product
  • retrieval results
  • user feedback
  • budget remaining
  • branch status

Forbidden as runtime steering signals:

  • holdout labels
  • private judge scores
  • answer keys
  • post-hoc evaluator rationales
  • promotion decisions
  • human review notes unavailable in production

The rule:

If the production system cannot observe it, the runtime policy cannot use it.

The optimizer can train from eval traces after the run. The runtime cannot peek at the gate during the run.

This matters for GEPA-style prompt optimization, skill optimization, and topology search. They may use trace-derived feedback to propose candidates. They may not smuggle holdout answers into the candidate.

Privacy And Redaction

Traces are powerful because they are detailed.

That also makes them dangerous.

A trace can contain credentials, user data, private documents, file paths, source code, prompts, screenshots, and integration responses.

The trace system needs two simultaneous properties:

  • enough detail for causality
  • enough redaction for safety

Redaction has to happen at capture time for obvious secrets:

  • Authorization
  • X-Api-Key
  • Cookie
  • password
  • secret
  • token
  • access_token
  • refresh_token

But redaction is not just deletion. It records what was removed:

redactedFields = [...]

That lets a reviewer distinguish “the tool never sent auth” from “auth existed but was redacted.”

Over-redaction destroys causality. Under-redaction leaks data. The practical compromise is typed redaction plus artifact-level access control.

Where OpenTelemetry Fits

OpenTelemetry is the right outer shape for distributed traces.

As of June 6, 2026, the OpenTelemetry generative AI semantic conventions are marked Development and include GenAI model spans, agent spans, events, metrics, exceptions, OpenAI conventions, Anthropic conventions, and MCP conventions.

That is useful common infrastructure. It gives agent traces a path into existing collectors, dashboards, retention policies, and incident tooling.

But agent self-improvement needs more than generic spans. It needs first-class concepts that ordinary service traces do not enforce:

  • candidate id
  • scenario id
  • split tag
  • prompt hash
  • config hash
  • failure class
  • judge verdict
  • artifact hash
  • budget ledger
  • profile cell
  • raw provider event
  • promotion decision
  • analyst finding

The right design is not “OTel or agent schema.” It is:

  • OTel-compatible transport
  • agent-specific schema
  • promotion-grade integrity checks

Where Tangle Fits

Local package audit on June 6, 2026:

  • @tangle-network/agent-eval@0.34.1
  • @tangle-network/agent-runtime@0.26.0

agent-eval provides the trace and analysis layer:

  • TraceSchema v1: Run, Span, TraceEvent, BudgetLedgerEntry, Artifact, and failure taxonomy.
  • TraceEmitter: run lifecycle, hierarchical span helpers, run-complete hooks.
  • RawProviderSink: request, response, and error capture with redaction and retry attempt indexes.
  • assertRunCaptured: span, raw event, raw coverage, and outcome integrity.
  • ReplayCache and createReplayFetch: replay captured provider calls.
  • RunRecord: promotion-grade analysis row with model snapshot, prompt hash, config hash, commit, cost, token usage, split tag, outcome, and optional AgentProfileCell.
  • InMemoryTraceStore and FileSystemTraceStore: in-process and append-only trace persistence.
  • AnalystRegistry: runs isolated analysts over trace stores, run records, artifacts, judge inputs, or custom inputs.
  • AnalystFinding: content-addressed findings with evidence references and validation plans.
  • OtlpFileTraceStore: production trace consumption path.

agent-runtime provides execution traces:

  • runLoop: emits topology events for loop start, iteration start, iteration dispatch, iteration end, decisions, and loop end.
  • LoopTraceEvent: records driver, agent run names, task hashes, iteration placement, history length, winner, cost, duration, and iteration count.
  • createRefineDriver and createFanoutVoteDriver: create different trace shapes for sequential and parallel compute.
  • conversation journals: preserve multi-turn runs, halt reasons, cost caps, turn order, deterministic turn ids, and resumability.
  • OTLP exporter: ships loop trace events to an OpenTelemetry collector without adding a full SDK dependency.

This split matters:

  1. 01
    runtime emits behavior
  2. 02
    eval preserves and analyzes behavior
  3. 03
    gates decide whether behavior can ship

How Trace Systems Lie

Score-only learning

The optimizer sees a scalar and no mechanism.

Summary collapse

An LLM summary replaces the span tree, tool arguments, artifacts, and raw events.

Orphan spans

Structured spans exist without raw provider evidence.

Backend blindness

The campaign ran against stubs or partial backend failure.

Trace amnesia

The final answer is captured, but failed branches, retries, and rejected candidates are missing.

Judge leakage

Private evaluator labels or rationales leak into runtime policy.

Artifact loss

The trace references a patch, screenshot, log, or retrieved document that no longer exists.

Redaction erasure

Sensitive fields are removed without recording what was removed, destroying the ability to diagnose missing auth or missing tenant scope.

Unjoined identity

Run records, traces, scorecard cells, and analyst findings use different ids, so the evidence cannot be joined.

When A Trace Is Training Data

Do not optimize from final scores alone.

Preserve the trajectory.

An agent trace must answer:

  • Who ran?
  • Against which scenario?
  • With which model and prompt?
  • Through which topology?
  • Which actions were taken?
  • Which observations came back?
  • Which artifacts changed?
  • Which verifier judged them?
  • What did it cost?
  • What failed?
  • Which evidence supports that diagnosis?
  • Can the run be replayed?
  • Can the gate trust the capture?

If it cannot answer those questions, it is not training data for a self-improving agent. It is an anecdote about a run.

The trace is where agent behavior becomes learnable.

Source Trail

Source freshness checked on 2026-06-06.

Revision history9revisions
  1. GPT-6-lunapolish+159−224 view trace →
    Replaced prose code blocks with lists, equations, and compact flows. Audited recovery: selected public tool inputs from the October 2 editing session, not the complete rollout. Parent integration corrected MDX, display math, and responsive layout; full native records remain private.
    show diff
    diff --git a/src/content/posts/self-improving-stack-trace-systems.mdx b/src/content/posts/self-improving-stack-trace-systems.mdxindex 953b091..569e40d 100644--- a/src/content/posts/self-improving-stack-trace-systems.mdx+++ b/src/content/posts/self-improving-stack-trace-systems.mdx@@ -47,6 +47,8 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- +import Steps from '../../components/Steps.astro'+ The optimizer wants a score. I want the run, because a score only tells you that something happened while a trace preserves enough mechanism to explain what happened.  That difference is the difference between tuning a system and optimizing an unidentified projection.@@ -95,18 +97,16 @@ When $\text{score}=R(\tau)$ and the scorer is fixed, the score is a deterministi  This does not mean every byte is equally useful. It means the system must preserve the variables that can explain responsible mechanism: -```text-which model answered-which prompt was used-which branch ran-which tool was called-which arguments were passed-which observations came back-which artifact changed-which verifier judged it-which budget was spent-which failure class was assigned-```+- which model answered+- which prompt was used+- which branch ran+- which tool was called+- which arguments were passed+- which observations came back+- which artifact changed+- which verifier judged it+- which budget was spent+- which failure class was assigned  Without those variables, the optimizer is moving an unidentified intervention against an unidentified mechanism. @@ -114,9 +114,7 @@ Without those variables, the optimizer is moving an unidentified intervention ag  A summary says: -```text-The agent tried to use the API, failed, and produced a partial answer.-```+> The agent tried to use the API, failed, and produced a partial answer.  A trace says: @@ -150,87 +148,75 @@ A useful agent trace has several layers.  **Run identity** -```text-runId-scenarioId-candidateId-datasetVersion-codeSha-promptSha-modelFingerprint-seed-envFingerprint-parentRunId-projectId-chatId-layer-```+- `runId`+- `scenarioId`+- `candidateId`+- `datasetVersion`+- `codeSha`+- `promptSha`+- `modelFingerprint`+- `seed`+- `envFingerprint`+- `parentRunId`+- `projectId`+- `chatId`+- `layer`  This makes the run attributable. Without identity, the score cannot be tied to a candidate, commit, profile, prompt, model, or scenario.  **Span tree** -```text-agent span-llm span-tool span-retrieval span-judge span-sandbox span-custom span-```+- agent span+- llm span+- tool span+- retrieval span+- judge span+- sandbox span+- custom span  The span tree gives causality and nesting. A tool call can be under a planner branch. A judge can target a specific span. A sandbox failure can be tied to the code artifact it ran.  **Events** -```text-budget_decrement-budget_breach-state_mutation-policy_violation-redaction_applied-error-custom-```+- `budget_decrement`+- `budget_breach`+- `state_mutation`+- `policy_violation`+- `redaction_applied`+- `error`+- `custom`  Events capture point-in-time facts that are not whole spans.  **Budget ledger** -```text-tokens-wallMs-calls-usd-remaining-breached-```+- `tokens`+- `wallMs`+- `calls`+- `usd`+- `remaining`+- `breached`  This lets the evaluator distinguish a smarter policy from a more expensive one.  **Artifacts** -```text-diffs-files-logs-screenshots-test reports-retrieved documents-judge reports-```+- `diffs`+- `files`+- `logs`+- `screenshots`+- test reports+- retrieved documents+- judge reports  Artifacts make traces material. A span saying "patched file" is weaker than an artifact hash and storage pointer for the patch.  **Outcome** -```text-score-pass-failureClass-notes-```+- `score`+- `pass`+- `failureClass`+- `notes`  The outcome is still necessary. It is the label. It is just not enough by itself. @@ -238,34 +224,25 @@ The outcome is still necessary. It is the label. It is just not enough by itself  The trace is detailed enough when it can answer a counterfactual: -```text-If this action, observation, tool result, verifier result, or budget event had changed,-would the outcome have changed?-```+> If this action, observation, tool result, verifier result, or budget event had changed, would the outcome have changed?  Too coarse: -```text-agent failed at research-```+> agent failed at research  This does not identify whether the failure was query formation, source choice, stale retrieval, missing credentials, synthesis, or judge mismatch.  Too fine: -```text-every token, cursor movement, and private secret copied into a permanent record-```+> every token, cursor movement, and private secret copied into a permanent record  This increases cost and risk without necessarily improving diagnosis.  The target is sufficient structure: -```text-enough fields to localize the responsible mechanism-enough ids to join evidence across run record, trace, artifact, scorecard, and finding-enough redaction to preserve privacy and auditability-```+- enough fields to localize the responsible mechanism+- enough ids to join evidence across run record, trace, artifact, scorecard, and finding+- enough redaction to preserve privacy and auditability  ## Raw Provider Capture @@ -302,9 +279,7 @@ The raw event is not for dashboards. It is for forensics, replay, and audit.  The rule: -```text-Every LLM span that affects a score needs matching raw request evidence.-```+> Every LLM span that affects a score needs matching raw request evidence.  If the structured span exists but the raw provider event is missing, the run is not launch-grade evidence. @@ -314,9 +289,7 @@ Raw capture turns old runs into reusable experimental material.  If a run has recorded request and response events, a replay cache can map: -```text-canonical_request -> captured_response-```+**Canonical request → captured response.**  That enables: @@ -330,11 +303,9 @@ Replay is especially important for judge calibration. If two judges score differ  The replay miss policy matters: -```text-throw-fallback_to_network-fail_closed-```+- `throw`+- `fallback_to_network`+- `fail_closed`  For determinism audits and promotion gates, fail closed. A silent network fallback turns replay into a new experiment. @@ -344,15 +315,13 @@ Trace capture is not binary. It can fail partially.  The integrity check is: -```text-run exists-llm span count >= minimum-tool span count >= minimum, when tools are expected-judge span count >= minimum, when judges are expected-raw provider events exist-raw provider events cover llm spans-outcome exists-```+- run exists+- llm span count >= minimum+- tool span count >= minimum, when tools are expected+- judge span count >= minimum, when judges are expected+- raw provider events exist+- raw provider events cover llm spans+- outcome exists  In compact form: @@ -368,10 +337,8 @@ Promotion gates treat missing trace evidence as missing evidence, not as a neutr  The highest-cost failure is an orphan LLM span: -```text-structured llm span exists-raw request is missing-```+- structured llm span exists+- raw request is missing  That usually means capture was wired to the wrong sink, the call bypassed the instrumented client, or the route changed under the harness. @@ -389,25 +356,19 @@ stub_record = tokenUsage.input == 0 and tokenUsage.output == 0  Then: -```text-all stub records -> reject-mixed real and stub records -> quarantine or reject in CI-real tokens with zero cost -> cost ledger bug-```+- all stub records -> reject+- mixed real and stub records -> quarantine or reject in CI+- real tokens with zero cost -> cost ledger bug  This is not a small bookkeeping issue. If a campaign runs against stubs, every downstream statistic is corrupted: scorecard deltas, held-out gates, analyst findings, and optimizer decisions.  The right interpretation of a stub campaign is: -```text-We did not evaluate the agent.-```+> We did not evaluate the agent.  not: -```text-The agent failed every task.-```+> The agent failed every task.  ## Analyst Findings @@ -415,30 +376,26 @@ The trace is raw material. Analysts turn it into structured diagnosis.  A useful finding has: -```text-finding_id-analyst_id-severity-area-claim-rationale-evidence_refs-recommended_action-validation_plan-confidence-subject-```+- `finding_id`+- `analyst_id`+- severity+- area+- claim+- rationale+- `evidence_refs`+- `recommended_action`+- `validation_plan`+- confidence+- subject  The `evidence_refs` field is the key. A finding without a span, event, artifact, metric, or prior finding reference is an unsupported assertion.  The analyst layer supports multiple lenses: -```text-failure-mode-knowledge-gap-knowledge-poisoning-improvement-```+- `failure-mode`+- `knowledge-gap`+- `knowledge-poisoning`+- `improvement`  Those lenses answer different questions: @@ -449,9 +406,7 @@ Those lenses answer different questions:  That is how traces become optimizer input. -```text-tau -> findings -> candidate mutation -> eval -> gate-```+`tau` → findings → candidate mutation → eval → gate  The output is not "a summary of the run." The output is a set of attributed hypotheses with validation plans. @@ -463,32 +418,26 @@ The system has to separate runtime observations from evaluation labels.  Allowed at runtime: -```text-tool outputs-compiler errors-test failures available to the product-retrieval results-user feedback-budget remaining-branch status-```+- tool outputs+- compiler errors+- test failures available to the product+- retrieval results+- user feedback+- budget remaining+- branch status  Forbidden as runtime steering signals: -```text-holdout labels-private judge scores-answer keys-post-hoc evaluator rationales-promotion decisions-human review notes unavailable in production-```+- holdout labels+- private judge scores+- answer keys+- post-hoc evaluator rationales+- promotion decisions+- human review notes unavailable in production  The rule: -```text-If the production system cannot observe it, the runtime policy cannot use it.-```+> If the production system cannot observe it, the runtime policy cannot use it.  The optimizer can train from eval traces after the run. The runtime cannot peek at the gate during the run. @@ -504,29 +453,23 @@ A trace can contain credentials, user data, private documents, file paths, sourc  The trace system needs two simultaneous properties: -```text-enough detail for causality-enough redaction for safety-```+- enough detail for causality+- enough redaction for safety  Redaction has to happen at capture time for obvious secrets: -```text-Authorization-X-Api-Key-Cookie-password-secret-token-access_token-refresh_token-```+- `Authorization`+- `X-Api-Key`+- `Cookie`+- `password`+- `secret`+- `token`+- `access_token`+- `refresh_token`  But redaction is not just deletion. It records what was removed: -```text-redactedFields = [...]-```+`redactedFields = [...]`  That lets a reviewer distinguish "the tool never sent auth" from "auth existed but was redacted." @@ -542,38 +485,32 @@ That is useful common infrastructure. It gives agent traces a path into existing  But agent self-improvement needs more than generic spans. It needs first-class concepts that ordinary service traces do not enforce: -```text-candidate id-scenario id-split tag-prompt hash-config hash-failure class-judge verdict-artifact hash-budget ledger-profile cell-raw provider event-promotion decision-analyst finding-```+- candidate id+- scenario id+- split tag+- prompt hash+- config hash+- failure class+- judge verdict+- artifact hash+- budget ledger+- profile cell+- raw provider event+- promotion decision+- analyst finding  The right design is not "OTel or agent schema." It is: -```text-OTel-compatible transport-agent-specific schema-promotion-grade integrity checks-```+- OTel-compatible transport+- agent-specific schema+- promotion-grade integrity checks  ## Where Tangle Fits  Local package audit on June 6, 2026: -```text-@tangle-network/agent-eval@0.34.1-@tangle-network/agent-runtime@0.26.0-```+- `@tangle-network/agent-eval@0.34.1`+- `@tangle-network/agent-runtime@0.26.0`  `agent-eval` provides the trace and analysis layer: @@ -598,11 +535,11 @@ Local package audit on June 6, 2026:  This split matters: -```text-runtime emits behavior-eval preserves and analyzes behavior-gates decide whether behavior can ship-```+<Steps layout='flow' items={[+  { title: 'runtime emits behavior' },+  { title: 'eval preserves and analyzes behavior' },+  { title: 'gates decide whether behavior can ship' }+]} />  ## How Trace Systems Lie @@ -650,21 +587,19 @@ Preserve the trajectory.  An agent trace must answer: -```text-Who ran?-Against which scenario?-With which model and prompt?-Through which topology?-Which actions were taken?-Which observations came back?-Which artifacts changed?-Which verifier judged them?-What did it cost?-What failed?-Which evidence supports that diagnosis?-Can the run be replayed?-Can the gate trust the capture?-```+- Who ran?+- Against which scenario?+- With which model and prompt?+- Through which topology?+- Which actions were taken?+- Which observations came back?+- Which artifacts changed?+- Which verifier judged them?+- What did it cost?+- What failed?+- Which evidence supports that diagnosis?+- Can the run be replayed?+- Can the gate trust the capture?  If it cannot answer those questions, it is not training data for a self-improving agent. It is an anecdote about a run. 
  2. GPT-6-lunapolish+19−17 view trace →
    Rendered existing equations with KaTeX. Audited recovery from the Luna editing session: selected public messages and tool-input previews, not the complete rollout. Parent integration review corrected prime notation.
    show diff
    diff --git a/src/content/posts/self-improving-stack-trace-systems.mdx b/src/content/posts/self-improving-stack-trace-systems.mdxindex 3274500..27e26f0 100644--- a/src/content/posts/self-improving-stack-trace-systems.mdx+++ b/src/content/posts/self-improving-stack-trace-systems.mdx@@ -57,25 +57,27 @@ The trace is not decoration around the eval. The trace is the data.  An agent run is a trajectory: -```text-tau = (x, s_0, a_1, o_1, s_1, ..., a_T, o_T, y)-```+$$+\tau=(x,s_0,a_1,o_1,s_1,\ldots,a_T,o_T,y)+$$  where: -```text-x = task-s_t = internal and external state-a_t = action-o_t = observation-y = outcome-```+$$+\begin{aligned}+  x &= \text{task} \\+  s_t &= \text{internal and external state} \\+  a_t &= \text{action} \\+  o_t &= \text{observation} \\+  y &= \text{outcome}+\end{aligned}+$$  A final score from a fixed scorer is a projection: -```text-score = R(tau)-```+$$+\text{score}=R(\tau)+$$  That projection is intentionally lossy. It collapses a long sequence of decisions, calls, costs, artifacts, and observations into one number. @@ -83,11 +85,11 @@ Optimization needs the lost variables.  The crude information-theory version: -```text-I(tau; failure_cause) >= I(score; failure_cause)-```+$$+I(\tau;\text{failure cause})\ge I(\text{score};\text{failure cause})+$$ -When `score = R(tau)` and the scorer is fixed, the score is a deterministic projection of the trajectory. By the data processing inequality, that projection cannot contain more information about the failure cause than the trajectory itself. Usually it contains dramatically less.+When $\text{score}=R(\tau)$ and the scorer is fixed, the score is a deterministic projection of the trajectory. By the data processing inequality, that projection cannot contain more information about the failure cause than the trajectory itself. Usually it contains dramatically less.  This does not mean every byte is equally useful. It means the system must preserve the variables that can explain responsible mechanism: 
  3. GPT-5.5draft view trace →
    Drafted the trace-systems post with formal trajectory notation, span ontology, raw provider capture, replay, trace integrity, analyst findings, leakage firewalls, and local Tangle package placement.
  4. GPT-5.5polish view trace →
    Polished the trace-systems post by adding a trace granularity test, tightening the information-loss claim to a fixed scorer, correcting loop trace event details, and adding trace store surfaces from the local agent-eval audit.
  5. GPT-5.5polish+6−8 view trace →
    let's track a section for each of these map items, and eventually a full article too, but i want a directory we can use to checkpoint our kn · 37 asst turns · 23 tool calls
    show diff
    diff --git a/src/content/posts/self-improving-stack-trace-systems.mdx b/src/content/posts/self-improving-stack-trace-systems.mdxindex 5f2bcce..9a9e631 100644--- a/src/content/posts/self-improving-stack-trace-systems.mdx+++ b/src/content/posts/self-improving-stack-trace-systems.mdx@@ -19,7 +19,9 @@ authors:     date: 2026-06-06   - { model: 'gpt-5.5', role: 'review', date: 2026-06-05 }   - { model: 'gpt-5.5', role: 'publish', date: 2026-06-05 }+  - { model: 'gpt-5.5', role: 'rewrite', date: 2026-06-05 } revisions:+  - { date: 2026-06-05, model: 'gpt-5.5', role: 'rewrite', note: '60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.', commit: 'd5bba9f0c633e5d2794e9b8e062ab48b15bbd1f5', trace_id: '2026-06-05T12-35-48-868Z-gpt-5.5-self-improving-stack-trace-systems-rewrite' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'publish', note: 'Published the self-improving stack series at Drew''s request, marking human takeover complete and flipping the post live.', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-trace-systems-publish' }   - { date: 2026-06-05, model: 'gpt-5.5', role: 'review', note: 'Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.', commit: 'b8fd3dbe812dd9ddd73865ae65fcc0d381b59d69', trace_id: '2026-06-05T12-08-35-196Z-gpt-5.5-self-improving-stack-trace-systems-review' }   - date: 2026-06-05@@ -41,17 +43,13 @@ supporting_trace_ids:   - '2026-06-05T12-08-35-196Z-gpt-5.5' --- -Scores tell you that something happened.--Traces tell you what happened.+The optimizer wants a score. I want the run, because a score only tells you that something happened while a trace preserves enough mechanism to explain what happened.  That difference is the difference between tuning a system and optimizing an unidentified projection.  A self-improving agent can only improve from the information it preserves. If the run record says "failed, score 0.42," the optimizer can only infer weak global pressure. If the trace says the planner chose the wrong tool, the tool call used a stale argument, the retrieval span returned irrelevant context, the judge penalized a missing artifact, and the retry loop repeated the same action three times, the optimizer has a causal surface. -The trace is not decoration around the eval.--The trace is the data.+The trace is not decoration around the eval. The trace is the data.  ## The Information Loss Problem @@ -600,7 +598,7 @@ eval preserves and analyzes behavior gates decide whether behavior can ship ``` -## Failure Modes+## How Trace Systems Lie  **Score-only learning** @@ -638,7 +636,7 @@ Sensitive fields are removed without recording what was removed, destroying the  Run records, traces, scorecard cells, and analyst findings use different ids, so the evidence cannot be joined. -## Working Rule+## When A Trace Is Training Data  Do not optimize from final scores alone. 
  6. GPT-5.5rewrite view trace →
    60/40 voice rewrite: grounded openings in concrete agent-work failures, removed scaffold headings, added falsification pressure, and tightened paragraph rhythm while preserving source trails.
  7. GPT-5.5publish view trace →
    Published the self-improving stack series at Drew's request, marking human takeover complete and flipping the post live.
  8. GPT-5.5review view trace →
    Standardized the source-trail section, dated source freshness, and removed remaining temporal or process wording from publication-visible text.
  9. GPT-5.5outline view trace →
    Research planning pass from a traced session.

Comments

Comments load from GitHub Discussions via Giscus. Configure PUBLIC_GISCUS_REPO, PUBLIC_GISCUS_REPO_ID, PUBLIC_GISCUS_CATEGORY, and PUBLIC_GISCUS_CATEGORY_ID in .env. See giscus.app to generate the IDs after you enable Discussions on the repo.