Repository navigation
Conversation
E2E tests are runningAuthorization passed for this commit. See the E2E Tests workflow for results. |
PR Summary by QodoCapture retry prompts and tool arguments in Level 3 telemetry
AI Description
Diagram
High-Level Assessment
Files changed (19)
|
|
🤖 Finished Review · ✅ Success · Started 6:30 PM UTC · Completed 6:53 PM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $9.85 |
Site previewPreview: https://ba1f3afa-site.fullsend-ai.workers.dev Commit: |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
Code Review by Qodo
1.
|
|
Risk Assessment: moderate (2/5) DetailsSubstantial diff size and active churn across runner telemetry components are balanced by zero protected paths, no workflow or dependency modifications, a 0.38 test-file ratio, and established author history, maintaining a moderate risk rating. Previous runRisk Assessment: moderate (2/5) DetailsSubstantial diff size and active churn across runner telemetry components are balanced by zero protected paths, no workflow or dependency modifications, a 0.35 test-file ratio, and established author history, resulting in a moderate risk rating. Previous run (2)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 2,695-line telemetry feature across actively churned runner paths, mitigated by zero protected-path or dependency changes, an established author, a 36% test-file ratio with unit tests, and no regression reverts, preserving the prior moderate score following minor corrective commits. Previous run (3)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 2,691-line telemetry feature across actively churned runner paths, mitigated by zero protected-path or dependency changes, an established author, a 36% test-file ratio with unit tests, and no regression reverts, preserving the prior moderate score following minor corrective commits. Previous run (4)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 2,607-line telemetry feature across actively churned runner paths, mitigated by zero protected-path or dependency changes, an established author, a 36% test-file ratio with unit tests, and no regression reverts, preserving the prior moderate score. Previous run (5)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 2,497-line telemetry feature across actively churned runner paths, mitigated by zero protected-path or dependency changes, an increased 36% test-file ratio with added unit tests, established author status, and lack of regression reverts, preserving the prior moderate score. Previous run (6)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 2,012-line telemetry feature across actively churned runner paths, mitigated by zero protected-path or dependency changes, a 32% test-file ratio, established author status, and lack of regression reverts, preserving the prior moderate score. Previous run (7)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 1,932-line telemetry feature across actively churned runner paths, mitigated by zero protected-path or dependency changes, a 32% test-file ratio, established author status, and lack of regression reverts, preserving the prior moderate score. Previous run (8)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 1,916-line telemetry feature across actively churned runner paths, mitigated by zero protected-path or dependency changes, a 32% test-file ratio, established author status, and lack of regression reverts, preserving the prior moderate score. Previous run (9)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 1,887-line telemetry feature across actively churned runner paths, mitigated by zero protected-path or dependency changes, solid test coverage (32%), established author status, and lack of regression reverts, preserving the prior moderate score. Previous run (10)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 1,750-line telemetry feature across actively churned runner paths, mitigated by zero protected-path or dependency changes, solid test coverage (32%), established author status, and lack of regression reverts, preserving the prior moderate score. Previous run (11)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 1,557-line telemetry feature across actively churned runner paths, mitigated by opt-in environment variable gating, strong test coverage additions, absence of protected paths or dependency changes, and alignment with the authorized predecessor scope. Previous run (12)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 1,426-line telemetry feature across actively churned runner paths, mitigated by opt-in environment variable gating, strong test coverage additions, absence of protected paths or dependency changes, and alignment with the authorized predecessor scope. Previous run (13)Risk Assessment: moderate (2/5) DetailsModerate risk reflects a substantial 1,279-line feature with active git churn across CLI telemetry paths, mitigated by no protected-path, workflow, or dependency changes, test additions, scope alignment with predecessor PR #6603, and opt-in environment-variable gating. Previous run (14)Risk Assessment: moderate (2/5) DetailsLarge diff (1,280 lines across 22 files) is offset by zero protected-path or dependency changes and adequate test coverage, keeping Tier 1 at 1.75/5, while CLI file churn and regression history place Tier 2 at 2.57/5, yielding a 62/38 weighted composite of ~2.06 that preserves the prior moderate score of 2. Previous run (15)Risk Assessment: moderate (2/5) DetailsLarge diff (1,280 lines across 22 files) is offset by zero protected-path or dependency changes and adequate test coverage, keeping Tier 1 at 1.75/5, while CLI file churn and regression history place Tier 2 at 2.57/5, yielding a 62/38 weighted composite of ~2.06 that preserves the prior moderate score of 2. Previous run (16)Risk Assessment: moderate (2/5) DetailsNo linked issue, so weight redistributes to 62% Tier 1 / 38% Tier 2. Tier 1 unchanged from prior review (1.75/5). Tier 2 recomputed fresh at current head (~2.6/5). Composite ~2.09, rounding to 2 (moderate); preserved per re-review anchoring since neither tier crossed a rounding boundary. Previous run (17)Risk Assessment: moderate (2/5) DetailsNo linked issue, so weight redistributes to 62% Tier 1 / 38% Tier 2. Tier 1 is essentially unchanged from the prior review (1.75/5: large blast radius/line-count bucket, zero protected/security/CI/dependency changes, non-bot non-first-time author, ~35% test ratio). Tier 2 rose slightly to 3.0/5 (from 2.57) on high fix/revert-grep counts and change-coupling on hotspot files, offset by low code age and zero true revert/sentiment hits. Composite 0.62x1.75 + 0.38x3.0 = 2.23, rounding to 2 (moderate) — consistent with the prior assessment (score preserved per re-review anchoring; the Tier 2 increase did not cross a rounding boundary). Previous run (18)Risk Assessment: moderate (2/5) DetailsLarge diff (901 lines/19 files, blast_radius=large) is offset by no protected/security-path or CI/dependency changes, a non-first-time non-bot author, and adequate test coverage (32%), keeping Tier 1 low (1.75/5); Tier 2 is pulled to moderate (2.57/5) mainly by heavy recent fix-commit history and churn on hotspot files with some unaddressed change-coupling; no formally linked issue redistributes weight to 62% Tier 1 / 38% Tier 2, yielding a moderate composite score of 2. |
|
Looks good to me Earlier findings
Previous runReviewFindingsHigh
Medium
Low
Next steps:
Previous run (2)ReviewReason: ambiguous-findings This PR was NOT reviewed. Do not count this as an approval. Previous run (3)ReviewFindingsMedium
Earlier findings
Next steps:
Previous run (4)ReviewFindingsHigh
Medium
Earlier findings
Next steps:
Previous run (5)ReviewFindingsMedium
Earlier findings
Next steps:
Previous run (6)ReviewFindingsHigh
Medium
Earlier findings
Next steps:
Previous run (7)ReviewFindingsMedium
Next steps:
Previous run (8)ReviewFindingsMedium
Next steps:
Previous run (9)ReviewFindingsMedium
Low
Next steps:
Previous run (10)ReviewFindingsMedium
Low
Next steps:
Previous run (11)ReviewFindingsMedium
Low
Next steps:
Previous run (12)ReviewFindingsMedium
Labels: The PR extends credential redaction on exported telemetry and has a remaining secret-exposure finding. Next steps:
Previous run (13)ReviewFindingsHigh
Low
Next steps:
Previous run (14)ReviewFindingsHigh
Low
Labels: The PR changes Go code implementing agent-runner telemetry capture. Next steps:
Previous run (15)ReviewFindingsLow
No other findings met the review threshold. Previous run (16)ReviewFindingsLow
No other findings met the review threshold. Previous run (17)ReviewFindingsLow
Previous run (18)ReviewFindingsMedium
Low
Next steps:
Previous run (19)ReviewFindingsMedium
Low
Next steps:
|
|
🤖 Finished Review · ✅ Success · Started 6:58 PM UTC · Completed 7:21 PM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $10.30 |
Under the Level 3 content gate, a retry iteration that composes a validation-feedback prompt (feedback_mode: append) now records it on the iteration's agent span as gen_ai.input.messages: one user message with one text part. The first iteration and retries without feedback send the runtime's default prompt and record nothing. The recorded copy passes through the collector's redaction, and its findings join the iteration's; Result no longer returns early when an iteration has findings or dropped bytes but no parts. The encoded input is charged against maxEncodedContentBytes so the two content attributes of one span stay within the proven size together. Signed-off-by: Dharit Shah <dhshah@redhat.com>
Under the Level 3 content gate a tool_call part now carries the call's arguments. ToolUseEvent gains Arguments — the input as the stream carried it, from both Claude Code parser paths; pi, codex and OpenCode leave it empty (fullsend-ai#7414) and the renderer does not print it. The collector decodes the value, redacts each string and object key on its own, and encodes it again: the redactor's patterns are written for plain text, and over serialised JSON they miss an assignment that opens a string or follows an escaped newline, a value behind escaped quotes, and JSON nested in a string, while Unicode folding can close a string early. Each string member is scanned once more, already redacted, beside its key, so the member-name pattern still sees the pair; only that pattern's finding is kept from the second scan. Arguments are dropped whole, charged to the dropped bytes, and the part marked fullsend.truncated (it keeps its id, name and summary) when their redacted encoding exceeds maxToolArgumentsBytes (8 KiB; a cut object is not JSON), when the text is not one JSON value, or when two keys of one object redact to the same string. A call without a name carries none. On the three measured live review streams 0-3 calls per iteration exceed the bound, each an Agent dispatch prompt. The encoded-ceiling property test generates and counts arguments. Regression pins, green on arrival: execute_tool spans carry nothing of a call's arguments, and with the gate off arguments and results stay out of run-telemetry.jsonl while the tool span is still written. Signed-off-by: Dharit Shah <dhshah@redhat.com>
Level 3 now records tool-call arguments and, on a retry that carries validation feedback, the prompt the runner composed. The reference, the dev guide, the runtime matrix and the two user guides say what is recorded, what is still not (the rest of the model's input), how arguments are redacted, bounded and counted, and that the input message is cut upstream and not described by the truncation markers. The size figures are re-measured with arguments sharing the total: on the three captured review runs the shipped bounds evict 31-63% of tool results (28-56% before), a 1 MiB total still evicts none, and encoding adds 8-10% at these bounds. Signed-off-by: Dharit Shah <dhshah@redhat.com>
…arguments Dated annotations only; no Decision text changes. ADR 0050 records that the retry prompt and tool arguments now ride the Level 3 record and what of the model's input is still not recorded. ADR 0108 records that its "full arguments next" is in place on the message record up to a per-call bound, while execute_tool spans stay metadata only. Signed-off-by: Dharit Shah <dhshah@redhat.com>
Result cleared the accepted arguments of a call whose name redacted to nothing without counting them. They are now added to the dropped bytes, and the span and the part — kept only when it has a summary — are marked truncated, as for every other way a call loses its arguments. Signed-off-by: Dharit Shah <dhshah@redhat.com>
The cap was checked before each delta was appended, so the delta that crossed it went in whole and the input could reach twice the cap. Now that the accumulated input travels on ToolUseEvent.Arguments, cut the crossing delta at the cap. Signed-off-by: Dharit Shah <dhshah@redhat.com>
The collector redacted by pattern only, so a credential with no recognizable shape — the case the literal pass in redactFeedback exists for — reached the record unchanged. Move that pass into replaceEnvSecrets, walked longest value first, and run it in the collector's redact ahead of the output pipeline: text, reasoning, tool results, arguments, names, summaries and the input message. An id that holds a value is dropped, like an id with any other finding. Each replacement counts in fullsend.content.redactions. Signed-off-by: Dharit Shah <dhshah@redhat.com>
3d19a5f ran the literal pass ahead of the output pipeline only. The pipeline's normalizer joins a value the stream split with an invisible character or spelled in compatibility forms, so that value reached the record, the input message included. Run the pass on the pipeline's result as well. The pass ahead of it stays: it sees the value as the env has it, and keeps a pattern from masking part of it. Not covered, and pinned by a test for the assignment case: a value written that way which a pattern also recognises is masked by that pattern. A call loses its summary when redaction found a secret in its arguments: the parser cut the summary out of them before anything scanned it. A number leaf in tool arguments is scanned as its digits. Signed-off-by: Dharit Shah <dhshah@redhat.com>
…DR annotations The reference says where the pass runs, how it counts, what it does not cover, when a call loses its summary, and the limit for runtimes that report no arguments; it adds the fourth way a call loses its arguments. The dev guide says a marked tool_call part lost its arguments whole, never in part, and that tool span names and ids do not get the literal pass. The contributing guide names replaceEnvSecrets. The two 2026-09-18 ADR annotations repeated what the reference holds; they are a short note and a link now. Signed-off-by: Dharit Shah <dhshah@redhat.com>
3d19a5f to
7944734
Compare
|
🤖 Finished Review · ✅ Success · Started 9:10 PM UTC · Completed 9:31 PM UTC Commit: Runtime: pi · Model: sonnet → claude-sonnet-5 · Effort: high · Cost: $9.49 |
Superseded by updated review
|
🤖 Review · Commit: |
An exact value over the lines of an array was sought in the strings joined by a line break, then joined by nothing as well, since a line of a notebook cell keeps its own break; a Windows line ending or a trailing space would have been the next case, and each case a view of its own. It is now sought once, in the strings in order with white space set aside, which reads the value as a reader of the lines does, however a line ends and wherever the value was cut. The private key block keeps the line-break join its pattern reads over. RuntimeSecrets lists the registered values for it. Signed-off-by: Dharit Shah <dhshah@redhat.com>
|
@waynesun09 one decision needed on the argument-redaction bar, so the review loop can end. As of f713132 the collector judges each key, string and number on its own; values under a secret-named member by that name; and across strings only two things: a private-key block (the strings joined by a line break) and an exact runner-env or runtime value (the strings in order, white space aside). That is more than Sentry or the secret scanners do; none reconstruct pattern-shaped secrets across fields. The review bot keeps flagging further cross-field reconstructions: an assignment, header or connection string whose value begins in the next string; a token split in pieces; a member-name pair a string's own quote closes. Proposal: those are documented limits ( OK as the bar, or do you want any of them in? Each would be a small monotone addition. |
|
🤖 Finished Review · ✅ Success · Started 2:38 PM UTC · Completed 2:57 PM UTC Commit: Runtime: pi · Model: openai/gpt-6.1-sol → gpt-6.1-sol · Effort: high · Cost: $2.66 |
A key masked whole names a secret whatever it was called, but the check read the key as scanned, before the document it holds could mask it: a key holding a document that names a secret in a spelling only decoding reads was masked, and the value under it kept. The check now reads the key as it is exported. Signed-off-by: Dharit Shah <dhshah@redhat.com>
|
🤖 Finished Review · ✅ Success · Started 3:05 PM UTC · Completed 3:19 PM UTC Commit: Runtime: pi · Model: openai/gpt-6.1-sol → gpt-6.1-sol · Effort: high · Cost: $1.83 |
waynesun09
left a comment
There was a problem hiding this comment.
Three findings from review of head fa2f0b4, posted inline.
…ules The cross-string judgement masked as little as it could: it located each key, string and number in the text as written with a scanner of its own, judged their joins by a line break and by nothing against match offsets, and judged a held document both as scanned and as written. Each view had an edge, and each review round found the input past it. Each level is now judged in three coarse ways, every one of which only masks or drops more. Each key, string and number is redacted on its own, as before. A string holding a JSON document is judged as the arguments are: decoded from the string as written, walked under the same member, encoded again, so the structure a reader decodes is the structure that was judged, at the cost of the document's layout. One the normalizer changes at all, one held past the depth bound, one whose keys redact alike, text shaped like a document that does not parse, and one whose text holds an exact value with structure of its own are masked whole. Of two members of one name decoding keeps the last and the record keeps what was decoded, so the duplicate-name walk is gone. The atoms of the arguments, their strings and numbers decoded in the order written by the standard decoder's token walk, a held document's in its place, are judged joined by a line break for the private key block, as decoded and as the normalizer renders them, and in order with white space set aside for an exact runtime or runner env value, sought in the text as written as well when it holds structure of its own; either drops the arguments. A private key inside one string drops them too, rather than masking the block in place. The exported set shrinks or holds in every case the tests had. The fidelity changes: a held document is encoded again; a document with a compatibility character is masked whole; a private key or an exact value over the lines of a held document drops the arguments instead of masking the string. The shelved rewrite's leak rows replay as before: only the documented limits stay red (a member-name pair a string's own quote closes, an assignment whose value is the next number). Signed-off-by: Dharit Shah <dhshah@redhat.com>
|
🤖 Finished Review · ❌ Failure (post-script /home/runner/work/fullsend/fullsend/.fullsend/.fullsend-cache/resources/sha256/f05af62aba180f85f38250bc8336353830efb200682f9e3c56e42caca037fb2f/scripts/post-review.sh failed: exit status 1) · Started 4:25 PM UTC · Completed 4:43 PM UTC Commit: Runtime: pi · Model: openai/gpt-6.1-sol → gpt-6.1-sol · Effort: high · Cost: $2.51 |
|
Update, since the mechanism above has changed: as of bb50b21 the collector is simpler and the question is the same. Each key, string and number is still judged on its own; a string holding a JSON document is decoded, walked and encoded again (one the normalizer changes at all is masked whole); and across strings it judges the same two things, a private-key block and an exact runner-env or runtime value, over the strings in written order by the standard decoder's token walk, with no match-offset bookkeeping. One behaviour change: a private key inside one string now drops the arguments rather than masking the block in place. The documented limits and the bot's two open findings are unchanged. Decision asked for is unchanged: OK as the bar, or any of them in? |
Review found the arguments path accreting judgements without converging: a per-leaf scan, a member-name scan with a stand-in name, held documents walked to a depth, duplicate-name and normalizer-change detection, a fail-closed mask on document-shaped text that caught ordinary Edit fragments and cost them their summary, and a cross-string judgement in two renderings. Each round added a rule, and the known gaps stayed. The contract is now closed by construction. A tool_call part records the members of the call's input that name what was called — paths, patterns, commands, modes and bounds, the list in recordedArguments — each redacted as text, and nothing else of the input: a file body, an edit, a prompt, a notebook source, a todo list, and any member whose value is an object or an array is dropped, scanned first so a secret in it still counts, charged as the redacted text it was scanned as, and the part marked. No fail-closed mask is left: a dropped member raises no finding of its own, keeps the summary, and adds nothing to fullsend.content.redactions. A string the normalizer stripped an escape sequence or tag characters from is still masked whole. The raw bound goes with the walk whose cost it limited; the encoded bound stays. The redactor's match-location and runtime-secret helpers, added for the cross-string judgement, are gone with it; envLiterals keeps its one caller. The guides and the two ADR annotations state the list and the one-sentence guarantee. Signed-off-by: Dharit Shah <dhshah@redhat.com>
|
Superseded by the review's third finding: the allowlist contract landed in c8324df (the record holds the listed members of a call, each redacted as text, and nothing else of its input), and the two deferred findings are resolved as not applicable. No decision is pending here. |
|
🤖 Finished Review · ✅ Success · Started 6:01 PM UTC · Completed 6:19 PM UTC Commit: Runtime: pi · Model: openai/gpt-6.1-sol → gpt-6.1-sol · Effort: high · Cost: $3.78 |
… private A number under a listed member of a tool call's arguments was copied without the scan the previous head gave it: a runner env value sent as a number was exported whole. It is now scanned as its digits like every kept string, and the mask is exported as the string it is; a clean number stays a number. decodeJSON accepted anything after the first JSON value once the raw bound and its json.Valid gate went. It now requires white space at most after the value; arguments with trailing input take the existing path: scanned as text so a secret in the tail counts, dropped whole, charged. The telemetry file is created 0600: under the content gate it holds prompts and tool arguments, like the feedback audit file the runner keeps at 0600. Docs: the user guide no longer says tool arguments include file contents; the dev guide drops the two sentences describing the deleted raw bound and key-collision rule and says kept numbers are scanned. Review findings on c8324df: 4222621667 (high), 4222621676 (medium), 4222621683 (low), and the telemetry file mode from the summary. Signed-off-by: Dharit Shah <dhshah@redhat.com>
|
🤖 Finished Review · ✅ Success · Started 6:57 PM UTC · Completed 7:13 PM UTC Commit: Runtime: pi · Model: openai/gpt-6.1-sol → gpt-6.1-sol · Effort: high · Cost: $2.46 |
Third and last PR of the ADR 0050 Level 3 series, after #6429 (gate, collector, budget) and #6603 (tool results,
execute_toolspans, ADR 0108). It delivers the two items #6603's description named as next: the runner-composed retry prompt asgen_ai.input.messages, and tool-callargumentson the message record. The carrier is unchanged — the record on theagentspan, as ADR 0108 decided;execute_toolspans stay metadata only. Still one env var, no new configuration.What this does
internal/cli)validation_loop.feedback_mode: append, the promptbuildFeedbackPromptcomposed is recorded on that iteration'sagentspan asgen_ai.input.messages: oneusermessage, one text part (semconv v1.37.0 input-messages schema). It is set right after the span starts, so the cancel, error and normal finalize paths carry it. The first iteration and retries without feedback send the runtime's fixed default prompt and record nothing.internal/runtime)ToolUseEventgainsArguments: the call's input as the stream carried it, from both Claude Code paths (assistant line;stream_eventdeltas). Unchecked and unredacted at this layer, never rendered. Input assembled from deltas is cut at the parser's 1 MiBmaxToolInputSize. pi and codex leave it empty (#7414); so does OpenCode, a stub runtime.internal/cli)tool_callparts carryarguments: the members of the call's input that name what was called — paths, patterns, commands, modes and bounds, the list inrecordedArguments(file_path,pattern,path,command,description,url, the notebook, Grep and bound members) — each decoded and redacted as text on its own, then encoded again, and nothing else of the input: a file body, an edit, a prompt, a notebook source, a todo list, and any member whose value is an object or an array is dropped, scanned first so a secret in it still counts, charged as the redacted text it was scanned as, and the part markedfullsend.truncated. The guarantee is one sentence: the record holds these members of a call, each redacted as text, and nothing else of its input. A string the normalizer stripped an escape sequence or tag characters from is masked whole (stripping can take a token's first letter). No fail-closed mask is left: a dropped member raises no finding of its own, keeps the summary, and adds nothing tofullsend.content.redactions. Done once, when the event is handled. A call without a name — or whose name redacts to nothing — carries none; arguments lost to a name that redacted away are charged and marked like the dropped ones below.maxToolArgumentsBytes = 8 KiB, measured on the redacted encoding of the kept members. Over it — or when the text is not one JSON object — the arguments are dropped whole (a cut object is not JSON): the part keeps id, name and summary, is markedfullsend.truncated, and the bytes are charged tofullsend.content.dropped_bytes. The walk is one pass over the top-level members, so no raw bound is needed ahead of it. The input message's encoded size is charged against the existing 255,000-byte ceiling, so the two content attributes of one span stay within the proven size together.GH_WORKFLOW_TOKEN, #6649), 8 bytes or longer, with[REDACTED:<key>]—replaceEnvSecrets, one longest-first pass over both sources, shared withredactFeedback. It applies to all Level 3 content, the text and tool results #6429 and #6603 shipped included, and toexecute_toolspan names and ids. A call loses itssummarywhen redaction found a secret in its arguments; a number leaf in arguments is scanned as its digits.Why arguments are not scanned as serialized text
The redactor's patterns are written for plain text. A table test run first against the simple approach (scan the JSON text, keep it if still valid) failed 4 of the 8 shapes it started with (the first eight rows of
TestContentCollector_RedactsSecretsInArguments; the table has since grown): an assignment that opens a string ({"command":"DEPLOY_TOKEN=… make"}— the pattern anchors on line start or whitespace, and sees a quote), an assignment after an escaped newline, a value behind escaped quotes, and JSON nested in a string. Unicode folding over the text also turned fullwidth quotation marks into ones that closed the string and added a member. Decoding first gives the redactor the text it was written for; the per-member scan keeps what the serialized form did catch ("password": "…"), where the leaf alone has no context. That scan runs on every string, number and key under a secret-named member, at any depth, not only clean ones — a zero-width character, an accent or a second token in the value would otherwise switch it off — it goes to the pattern stage alone (both strings are already normalized, and a second normalizer pass over the pair can strip an escape sequence the first left open, and the value with it), and its probe text puts a comma where a double quote was: the quote would end the pattern's quoted run early, and the comma — like the quote — is outside every pattern's token class, so it joins no two runs into a token a prefix pattern would mask first (an underscore would: it sits inside most token classes).Measured
Three captured live review streams (the ones behind #6603's figures), sub-agent calls included:
Every call over 8 KiB is an
Agentdispatch prompt. Replayed through the real collector: with arguments sharing the 256 KiB total, the shipped bounds evict 31–63% of tool results (28–56% before — the baseline replay gives #6603's published 55% / 56% / 28%, in this table's run order); a 1 MiB total still evicts none (records of 456–1,011 KB); encoding adds 8–10% at these bounds. Replayed again through the collector atc8324dfd: all 513 calls keepargumentsand their summary, with no finding; the one member dropped by name isAgent.prompt(20 calls, marked). An agent that writes files is measured in the Evidence run below. Raising the total stays #7415.Decisions
tool_callparts keep the property feat(telemetry): capture tool results in Level 3 content and emit execute_tool spans #6603 gave them — never cut, only dropped — so the exact-accounting tests stand; the property test now generates arguments and counts them. What is recorded is decided by member name, not by size or by what a scan finds in it: aWriteorEditkeeps itsfile_pathand nothing of the body, aBashkeeps itscommand, anAgentdispatch keeps itsdescriptionand not its prompt. The cost on the three measured review runs: the 20Agentprompts are not recorded (the 0–3 per iteration over 8 KiB were not before either). Review of head fa2f0b4 asked for a contract closed by construction rather than another rule; this is the allowlist option of the two it offered.redactFeedbackscans the feedback before it is sanitized and framed, and sanitizing can join a token that scan saw split, so the recorded copy is pattern-scanned again (the prompt the agent is sent is not changed); the pass also folds compatibility characters the prompt kept. Folding can grow the copy about elevenfold (U+FDFA, 3 bytes to 33): worst case about 113 KB for a 10 KiB feedback (113,095 bytes measured), which is why the size is charged to the ceiling rather than assumed small; a test builds that case and pins that input and output together stay within 255,000 bytes.summarystays besidearguments: it is what survives when arguments are dropped, and what pi, codex and OpenCode have.fullsend.content.redactions;Resultno longer returns early when an iteration has findings or dropped bytes and no parts (a retry that fails before emitting anything). The input's pattern scan repeatsredactFeedback's when sanitizing changed nothing; a connection-string mask left by that first scan matches its own pattern and counts a finding. SecretsredactFeedbackalready masked are recorded as masks and are not counted (it discards its findings), so a retry prompt carryingghp_...adds 0 tofullsend.content.redactions;TestAttachInput_RedactsAndCountsFindingsfeedsattachInputan unmasked token, the case where this scan is the first to see it.gen_ai.input.messagescarries one message, the runner'suserprompt. The convention's worked example puts client-executed tool results in that attribute underrole:"tool"; here they stay ongen_ai.output.messages, so the role-placement deviation feat(telemetry): capture tool results in Level 3 content and emit execute_tool spans #6603 documented atcontentMessageis unchanged, for the reasons given there and in ADR 0108: parts keep stream order, and that record is the scorer contract.runner_envcredentials, but the record leaves the runner, and the contribution guide's first redaction invariant asks for the literal pass wherever content does. The pass runs ahead of the pipeline, on the text as written, so a pattern does not mask part of a value and keep its first bytes, and again on the pipeline's result, because its normalizer joins a value the stream split with an invisible character or spelled in compatibility forms. Not covered, and pinned by a test for the assignment case: a value written that way which a pattern also recognises — in an assignment, a header, a secret-named field or a connection string, or by its own prefix — is masked by that pattern, and shows what its mask shows. Longer values go first, so a value that contains another is replaced whole and the text does not depend on map order (this changesredactFeedbacktoo). Each replaced key counts once per pass per scanned string infullsend.content.redactions, and again where a pattern then masks the marker (an assignment, a header, a secret-named field); an id that holds a value is dropped, like an id with any other finding. The parser cuts a call'ssummaryout of its arguments before anything scans it, so a call loses the summary when redaction found a secret — a runner env value or a pattern's — in its arguments; pi, codex and OpenCode report a summary and no arguments, and there a value that straddles the cut stays in part.execute_toolspan names and ids get the same redaction (redactText, shared by the collector and the tool span tracker).parent_tool_use_idis still dropped at decode, so sub-agent calls and their arguments sit flat and unattributed on the record; it stays ADR 0050's deferred item 1, as ADR 0108's Consequences records — no issue filed yet), and the rest of the model's input: the fixed default prompt is a constant that carries no task and is left out by choice; the requests the runtime builds are not visible to fullsend, which reads the runtime's stream, not its API requests.execute_toolspans EM-002 reads. Whether that is enough for those rules is for eval-measure scores trace health but not run health — add a deterministic tool-call defect scorer #7246 to decide.Evidence
Local gated run of this branch at
c8324dfd, 2026-10-08, file sink only — no OTLP endpoint was set, so this shows the record, not backend acceptance. A throwaway agent fromfullsend agent new --validation-loopthat writes and edits files on purpose:Writea README and a Python module,Editboth,Writea notebook,Readeach,Bash ls;feedback_mode: appendand a validator that rejects iteration 1 on purpose;claude-opus-4-6on Vertex; exit 0, validation passed on iteration 2. Trace4b4238c7d2e203c81d5fb2e502a6884a, 31 spans. The earlier run ate343f872(2026-09-21, a read-only agent) measured the input message the same way: 592 bytes for the retry prompt, 10 of 10 and 12 of 12 parts carrying arguments.agentspanagentspangen_ai.input.messagesusermessage, one text part — the default prompt, the framing, and the validator's line inside the<validation-output>fencegen_ai.output.messagesfinish_reason: stopstoptool_callparts carryingargumentsBash, 4Write, 2Edit, 3Read)Bash, 5Write, 3Read)Write/Edit/NotebookEditparts:argumentskept,summarykept, markedfile_path(andreplace_all) and their summary, all 6 marked; 1,316 bytes of bodies and edits droppedfile_pathand their summary, all 5 marked; 543 bytes droppedheld_documentfindings, redactions, truncationfullsend.content.truncated: true(dropped members only)true(dropped members only)A
Writepart as recorded:{"type":"tool_call","id":"toolu_vrtx_01FxxaAhPEqPkz6tTWCS7zoX","name":"Write","summary":"/sandbox/workspace/target/notes/README.md","arguments":{"file_path":"/sandbox/workspace/target/notes/README.md"},"fullsend.truncated":true}. Theexecute_toolspans carrygen_ai.operation.name,gen_ai.tool.name,gen_ai.tool.call.idand, on failed calls,error.type— nothing of the arguments. The gate-off case is covered by the pinned test, not by a live run.Tests
Red first, except where this paragraph says otherwise: three tests were green on arrival, and the
attachInputcall site has no unit test. 46 hand mutants over the new behaviour (charge, redaction, findings, role and part type, theResultguard, invalid-JSON handling,UseNumber, null, the bound and its boundary, charge unit, key redaction and single key scan, the per-member scan — removed, gated on a clean value, keeping any finding, uncounted, quotes kept, run through the whole pipeline — key collisions, sorted walk, nameless calls, array walk, accounting, single scan, serialized-scan regression, span leakage, both parser paths, renderer) — all killed. Pre-push review over three rounds found one high defect (the member-name scan was skipped when the value raised any other finding) and two regressions in its fixes (an underscore as the quote stand-in; the scanned pair going back through the normalizer), plus wording and accounting errors in the docs and comments — all fixed in these commits. Three tests were green on arrival, with no production code changed under them: arguments absent fromexecute_toolspans, and arguments and results absent fromrun-telemetry.jsonlwith the gate off while the tool span is still written (both labelled as pins in their commit), and the renderer not printing arguments (renderer.gois untouched). The one line inrunAgentthat callsattachInputhas no unit test, as withattachContent's call site; the evidence run above exercises it. After opening, review led to six fixes, each red first. From Qodo: arguments lost to a redacted-away name were not charged or marked (99f13991), the delta crossing the parser's input cap went in whole (1c124ceb), and the collector had no literal pass for runner env values (c697b84b). From the review agent: that pass ran only ahead of the normalizer, so a value in compatibility forms reached the record (a1c1851a, which also drops the summary of a call whose arguments held a secret and scans number leaves). From the review agent's later rounds: a value nested in an array or object under a secret-named member, and then a key there, kept an opaque credential; provider-only keys missed the literal pass; andexecute_toolnames and ids had none (3ecc9b50,1fd2b7d4,e7c1e44e). Keys are masked like values, field names of eight characters or more included, so an object such as{"auth":{"username":…,"password":…}}collides and its call's arguments are dropped, charged and marked — the review bot's suggested fix, taken over numbered placeholders that would have added an output convention. Hand mutants over these fixes are all killed except three that cannot change behaviour and therunAgentline that hands the tracker the runner env, untested like theattachInputcall site. From the review agent's later round andwaynesun09: a private key over the lines of a direct array and a fold that moves a value between members of a held document (fixed as the bot suggested — a cross-string judgement, and masking whole a held document the normalizer changes — in938a07c5, with the stripped-escape rule above); and the per-value scan under a long secret-named key running before the bound (6ba0c639). From the review agent's round on938a07c5: a key-collision drop short-circuited the cross-string judgement (so its finding could not cost the summary), a key masked whole lost its name, and the cross-string pass did not know runner env values — fixed inb1eb40b2, each by judging more; then an exact value over two lines that keep their own breaks (2b07a78c), generalised to the white-space-aside search (f713132b) so no line ending or cut point is a new case; and a key masked whole as a document it holds now names a secret too (fa2f0b43). Each fix masks more, never differently: the fix policy from here is to add a view or widen a mask, so the exported set only shrinks. Two heavier designs for the last fix — a pipeline stage between the normalizer and the patterns, and masking a value's beginning at the end of a string — were built and dropped after review: the first let a combining mark after a value defeat it and its mask hid the keyword a pattern needs; the second masked ordinary words and dropped call ids. A step-back review on 10-08 — field practice is per-field scrubbing or one flat scan of the serialized payload, never a reconstruction across fields — found the precision layer over-built: a scanner of its own for the atoms, match offsets, two join views, a held document judged as scanned and as written.bb50b213replaces it with the standard decoder's token walk and three coarse rules that only mask or drop more (code in the section 265 → 217 lines); the shelved rewrite's leak rows replay the same, only the documented limits red. Human review of head fa2f0b4 (waynesun09, three findings): ordinaryEditfragments masked whole by the document-shaped fail-closed rule, losing their summary and inflating the redaction count; the design accreting judgements without converging, with two simpler contracts offered; and the evidence measured only on review streams at an older head.c8324dfdtakes the allowlist contract — the record holds the listed members of a call, each redacted as text, and nothing else of its input — which removes the held-document, cross-string, member-name, collision and fail-closed rules together (the section is 78 code lines, from 265), leaves no non-secret finding to cost a summary or count as a redaction, and is measured above on the three review streams and on a fix-style run at that head. The review agent's round onc8324dfdfound two regressions of that commit — a number under a listed member copied unscanned, and trailing input after the JSON value accepted once the raw bound's validity gate went — plus a stale user-guide sentence and the telemetry file's 0644 mode: fixed incbfe80e0, each red first (a kept number is scanned as its digits, one JSON value then white space at most, the file created 0600, the dev guide's two sentences on the deleted raw bound and collision rule removed).