{"id":"20708003-95b6-4906-9b57-eb9f63c332ac","arxiv_id":"2608.02302","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Agent-declared causal-hypothesis boundaries yield variable-length semantic phases that stay attributable after declaration scrubbing, but the resulting DPO preference signal is construction-bound and does not transfer to adversarial items.","lead":"An LLM coding agent is asked to declare the causal hypothesis it is pursuing at each stage of a task; those declarations split long trajectories into semantic phases. This paper tests whether those self-declared segments are real by seeing how much survives when the declarations are removed, and finds them coherent and not cheaply reproducible, though the fitted preference signal transfers weakly.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing no-declaration control: the declaration prompt may induce the phase structure the external-recovery tests attribute to the agent's own epistemic states.","rationale":"The reader's weakest assumption already identifies the same gap: declarations may not faithfully mark causal belief states, and the paper provides no evidence that the protocol leaves search behavior invariant. My reading sharpens this into a specific failure mode: the declaration prompt may itself induce the phase coherence that Section 5 measures. The high attribution accuracy is therefore not decisive between 'declarations expose existing phases' and 'declarations create those phases.' This matters because contribution 3 claims the cuts 'follow the agent's reasoning' and are 'not an artefact of the logging'; a behavioral intervention is arguably worse than a logging artefact because it changes the distribution from which training units are drawn. The paper is unusually honest about the missing reviewer ablation, but that honesty does not fill the evidential gap. A no-declaration collection arm is the natural and direct test. If it shows behaviorally identical trajectories and comparable attribution lift, the concern is resolved; if not, the conditional verdict should remain and the central claim needs re-scoping to 'declaration-induced semantic phases,' which could still be useful but is not the same claim as 'the policy's own epistemic commitment.' I therefore keep the reader's CONDITIONAL verdict unchanged: the paper's other evidence—the paired window comparison, the label-permutation collapse, the mechanical-rule failure, and the honest treatment of downstream null results—supports a promising but not yet settled core claim.","tokens_in":19634,"tokens_out":6490,"duration_ms":68352,"concrete_test":"Collect at least 100 matched task instances under the declaration-on and declaration-off protocols, holding seed, scaffold, reviewer, and budgets fixed; in the off arm remove only the declaration sentence from the prompt. Compare per-task success, trajectory length, and tool-call distributions between arms. Then, on the declaration-off trajectories, use a post-hoc boundary predictor (e.g., the §5.1 attribution matcher with sliding-window optimization) and measure the same attribution lift over chance. If the off arm is behaviorally different, or if its attribution lift is materially lower than the declared arm's, the recovered coherence is protocol-induced and the central unit is not validated as a faithful epistemic boundary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing gap is the absence of a no-declaration control. Every trajectory in the corpus was collected under a prompt that instructs the agent to declare causal hypotheses at adoption events (§3.1). The external-recovery tests in §5 delete the declaration text and scrub hypothesis strings, but they cannot delete the behavioral effect of having been prompted to declare: an agent asked to maintain a named conjecture may organize subsequent edits and tests so that they are consistent with that conjecture, producing action blocks that are attributable at 2.18x chance precisely because the protocol induced the coherence. The 9.8% token-cost statistic (§1.3) measures verbosity, not invariance. The paper is explicit that the critical ablation was not run: 'we did not run the reviewer without declarations' (§3.1), and no trajectory was collected without the declaration contract. Consequently the central claim that the boundaries are 'the policy's own epistemic commitment' rather than 'prompt compliance' is unsupported: the §5 results establish coherence of the instrumented distribution, not the existence of natural semantic phases that declarations merely expose. A post-hoc-restatement failure would lower attribution; a behavioral-induction failure would raise it, so the current high attribution is consistent with the alternative explanation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a collection-time protocol in which a coding agent is prompted to declare falsifiable causal hypotheses during a rollout, and the interval governed by each hypothesis is treated as a variable-length 'semantic phase' that serves as a training unit. The central claim is that these declared boundaries are real and non-trivial: after deleting the declaration text and scrubbing hypothesis strings, an external model can still attribute action blocks to their governing hypothesis at about twice chance, better than equal-length windows; a code-blind human annotator matches boundaries above random; a mechanical test-event rule does not. The paper also constructs phase-boundary preference pairs, trains a DPO adapter, and reports that on adversarial held-out items no decisions change while on matched-construction items 4/60 change, with controls changing none. The writing is unusually candid about limitations, including the absence of a no-declaration control and the post-hoc nature of some evaluation choices.","tokens_in":19950,"tokens_out":7644,"duration_ms":65990,"significance":"If the central claim holds, the protocol provides a low-cost, dependency-free segmentation unit for long-horizon agent trajectories, with potential applications to credit assignment, process supervision, segment-level preference learning, and context compression. The paper's strengths include a concrete data/code release, multiple control conditions for the attribution test (label permutation, lexical baseline, equal-length windows, offset decay), and a clear separation between what is claimed and what is not. However, the load-bearing evidence for the reality of the boundaries is weakened by the absence of a no-declaration control, by an acknowledged filter asymmetry in the main paired comparison, and by the use of an interested, post-hoc human annotation. These issues are fixable in principle and do not invalidate the overall idea, but they do require revision before the central claims can be accepted.","major_comments":[{"comment":"Missing no-declaration control. The central claim (contribution 3, §1.4) is that declared boundaries track the policy's own epistemic commitments and 'are not an artefact of the logging.' The §5 tests scrub the declaration text and hypothesis strings, but they cannot remove the behavioral effect of having been told to maintain named conjectures. Every trajectory in the corpus was collected under the declaration contract; an agent prompted to keep a conjecture may organize subsequent edits and tests to be consistent with that conjecture, which would produce exactly the attribution and boundary-placement signal observed. The paper itself admits the critical ablation was not run: 'we did not run the reviewer without declarations' (§3.1), and §9.1 states 'the boundary exists because the agent was asked to declare it.' The 9.8% token-cost statistic (§1.3) measures verbosity, not behavioral in","section":"Table 3, §5.1"},{"comment":"The main paired comparison is biased by a filter that favors the declared arm. The caption states that blocks under 200 rendered characters are skipped in both arms, but it also states that this drops more declared than equal-length blocks and 'the filter runs in the declared arm's favour.' This asymmetry undermines the paired sign test (p=0.0002) and the disjoint intervals (2.38x vs 2.07x). To support the claim that declared blocks are better than equal-length blocks over the same trajectories, the comparison should be reported with the filter applied symmetrically (or with no filter), and the results should be shown to be insensitive to the filter threshold. As reported, the headline attribution advantage may be an artifact of which blocks were deleted.","section":"Table 3, §5.1"},{"comment":"The 'code-blind' annotation is an interested, post-hoc floor. The annotator is the first author, knew the segmentation criterion, and the mark budget and primary statistic were chosen after the metrics were computed; the paper acknowledges the bias runs toward over-recovery. The permutation result (24 vs 11.5 expected, p<0.0001) is suggestive, but it is not independent evidence. To carry the 'not cheaply reproducible' claim, the annotation should be repeated by an independent annotator who does not know the criterion, with a pre-registered mark budget and a pre-specified primary metric. The current result is a lower bound, as stated, but it cannot serve as a principal piece of evidence for the reality of the boundaries.","section":"§5.2, §9.2"},{"comment":"The downstream preference result is not statistically distinguishable from noise. On the matched-construction set, 4/60 changes has exact McNemar p=0.125, instance-level p=0.125, and shift-direction p=0.093; the adversarial set shows 0/91 changes. The paper's conclusion that 'the fitted preference is real but bound to how the pairs were written' is stronger than the evidence: all interval estimates include the null. This does not undermine the segmentation claim, but it should be reported as a null/underpowered result, not as evidence that the boundary is the bottleneck. The sentence in the Conclusion ('the next step is a more diverse pair corpus, not a different boundary') goes beyond what the data show.","section":"§6, Table 4, §9.4"}],"minor_comments":[{"comment":"The header 'Chance×chance' is confusing; it should read '×chance' or 'multiple of chance' to match the values in the table.","section":"Table 3"},{"comment":"The phrase 'A blinded annotator finds the spans usable' initially suggests a human annotator, but the text later reveals it is a model annotator. Please call it a 'blinded model annotator' on first mention to avoid ambiguity.","section":"§4"},{"comment":"The sentence 'the identical keep-categories with no phase structure at all—one global bucket, one global eight-hit cap—retain exactly as many edits and tests in 79% of the record's size' is hard to parse. Please rephrase to state explicitly what budget the ablation row is compared against.","section":"§7"},{"comment":"The caption 'as declared boundary slid by (messages)' should be reworded, e.g., 'attribution as the boundary is slid by N messages from the declared position.'","section":"Figure 3 caption"},{"comment":"The offset curve shows a maximum at -2 messages, but the paper reports no interval and does not test that difference. The text already notes this, but the figure and surrounding discussion should not imply that the declared position is optimal.","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-written and unusually honest, and the data/code release is a strength. The decisive gap is the missing no-declaration control: every trajectory was collected under the declaration contract, so the recovery tests cannot distinguish 'epistemic state exposed by the declaration' from 'phase structure induced by the declaration prompt.' If the authors can collect a matched corpus without the declaration prompt and show the recovery signal persists (or substantially adjust the claims), the paper could be publishable. The filter asymmetry in the main attribution comparison and the non-significant downstream preference results are additional concerns that should be addressed directly. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The first is that this is a genuinely new idea: the acting agent declares its causal hypothesis at adoption time, and the resulting phases are used as training units, with no gold patch, milestone vocabulary, or retrospective segmenter. The second is that the paper is unusually candid about its own soft spots — it says in Section 3.1 that the reviewer without declarations was never run, and in Section 9 that the construction attribution doesn't reach significance. That candor is real, and it makes the paper engageable.\n\nWhat the paper does well: the attribution evidence is substantial. Deleting declaration text and scrubbing hypothesis strings, a model attributes action blocks to their governing hypothesis at 2.38x chance in the paired run, beats equal-length blocks with a sign test p=0.0002, survives a model-free TF-IDF control, and collapses under label permutation. The mechanical test-event rule fails at both ends of the sweep. The paper ships code and data, and the release notes are unusually precise about what is excluded. Those are real points in its favor.\n\nThe load-bearing soft spot is exactly the one the stress-test names: every trajectory was collected under a prompt instructing the agent to declare named hypotheses. The recovery tests scrub the declaration and the hypothesis strings, but they cannot scrub the behavioral effect of having been asked to maintain a named conjecture. An agent prompted to organize its work around hypotheses may produce action blocks that are coherent because the prompt induced that coherence. The 9.8% token-cost figure measures verbosity, not behavioral invariance. The authors admit the critical ablation was never run. So the strong reading — that the boundaries are the policy's own epistemic commitments — is unsupported. What is supported is that the instrumented distribution is coherent and not cheaply reproducible by a parser. That's still worth something, but it's a weaker claim than the abstract's language suggests.\n\nSmaller issues, all acknowledged in the text: the Table 3 filter drops more declared than window blocks and runs in the declared arm's favor; the human boundary annotation was done by the interested first author with the statistic fixed after the fact; and the downstream DPO result is construction-bound, with exact McNemar p=0.125 on the matched set. None of these are hidden. Together they mean the central claim should be treated as promising rather than settled.\n\nWho should read this: anyone working on credit assignment, process supervision, or segment-level preference learning for agents. I would send it to peer review. A serious referee can push for the no-declaration control or, failing that, for reframing the central claim to 'coherence under the declaration protocol.' I'd cite it and I'd bring it to a reading group.","headline":"Novel collection-time self-segmentation with honest controls; the missing no-declaration arm leaves the epistemic-state claim unproven, but the paper is worth a serious referee.","tokens_in":20378,"tokens_out":3176,"would_cite":true,"duration_ms":27857,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Self-declared hypotheses turn agent trajectories into real training phases","keywords":["semantic self-segmentation","agent-declared boundaries","causal hypotheses","semantic phases","credit assignment","preference optimization","coding-agent trajectories","phase-boundary training"],"falsifier":"Collect matched task budgets under the declaration contract and under a no-declaration contract and compare the distributions of tool calls, actions, and outcomes; a significant divergence would mean the prompt itself changed search, so the boundary is compliance rather than epistemic state. As a second check, generate plausible hypotheses after the trajectory and rerun the attribution test—if post-hoc restatements attribute as well as live declarations, then the boundary position rather than the declaration carries the signal.","tokens_in":19534,"feed_emoji":"🧩","tokens_out":8139,"duration_ms":66728,"temperature":0.7,"pith_summary":"Long-horizon coding-agent trajectories are hard to turn into training data because neither single actions, episode labels, nor fixed windows align with what actually matters—the stretch of work governed by one causal hypothesis. The paper proposes that the acting agent declare each such hypothesis as it works, so a trajectory segments itself into variable-length semantic phases at collection time. The declaration costs one prompt instruction and no gold patch, milestone vocabulary, replayable environment, teacher logits, or retrospective segmenter. Most of the paper tests whether these boundaries are real: with every declaration scrubbed, an independent observer still attributes action blocks to their governing hypothesis at over twice chance and better than equal-length windows, while a mechanical test-event rule fails. One collection yields four supervised targets, and a standard preference optimizer trained on the resulting pairs moves a small number of held-out decisions from wrong to right, which the paper reads as construction-bound fitting rather than proof of generalization.","feed_headline":"Agent-declared phases beat fixed windows in attribution 2.38x","feed_subtitle":"One prompt line makes coding agents expose their own causal phases—no gold patch, milestone list, or segmenter needed.","key_machinery":"The semantic phase is the central object: the variable-length interval governed by one declared, falsifiable causal hypothesis about the cause of the problem, the search region, or the mechanism a repair must change. Its phase record carries the trajectory prefix before the decision, the hypothesis, the bound actions and evidence, the closing boundary, and a retrospective audit verdict. The declaration protocol is the instrument that produces it at collection time; the name of the hypothesis is what makes the phase an addressable handle, letting a reviewer negate one conjecture, later training targets condition on which direction was eliminated, and a boundary mark a deployment-faithful stat","core_discovery":"The paper's central claim: a falsifiable causal hypothesis declared by the acting agent during a tool-using rollout is a legitimate training-data boundary. Each declaration opens a semantic phase; actions and evidence bind to it until the next declaration or trajectory end. Because the agent names its conjecture, a reviewer can negate it by name, producing wrong-cause-then-correction transitions. The decisive test deletes every declaration: action blocks are still attributed to their governing hypothesis at 2.38 times chance and better than equal-length windows (paired sign test p = 0.0002), while a mechanical test-event rule fails. The boundary is thus legible but not derivable; the paper d","pith_inferences":["Editorial extension: if declarations faithfully mark epistemic state, the same protocol could apply beyond coding to any tool-using agent whose work is organized by checkable causal conjectures, turning trajectory logging into an instrumented experiment.","Editorial extension: because the paper did not run the reviewer without declarations, a direct test is whether post-hoc-inferred hypotheses (or no hypotheses at all) give the same attribution and downstream properties; if they do, the declaration is a convenience rather than the source of the boundary's meaning.","Editorial extension: the matched-construction DPO result points to pair diversity as the next lever; one concrete test is to build preference pairs from several independent generators and see whether held-out transfer improves, which would separate construction effects from boundary effects.","Editorial extension: the compression experiment measures retention of artifacts, not policy performance; running a deployed policy with the compressed ledger as its context would test whether addressability translates into better search, a step the paper identifies but leaves open."],"forward_implications":["With the declaration contract in place before collection, one rollout yields four supervised targets—audit judgments (2,721 rows), proposed next hypotheses (1,405), localized fixes (553), and phase-boundary preference pairs (2,551)—including audit supervision drawn from exactly the failed spans an episode label discards.","The declared boundary tracks behavior rather than format: with declarations scrubbed, attribution to the governing hypothesis reaches 2.38 times chance and beats equal-length blocks on the same trajectories (paired sign test p = 0.0002), so an outside observer recovers a substantial part of the structure.","The boundary is not cheaply reproducible: a mechanical test-event rule matches no more boundaries than random placement at either a strict or a permissive threshold, and the code-blind annotator recovers boundaries only by over-segmenting.","Preference optimization on phase-boundary pairs yields an orientation-sensitive signal that a label-inverted arm does not reproduce, and on matched-construction held-out items the adapter changes four of sixty decisions, all from wrong to right, where two controls change none.","The phase-level credit assigner, normalized within a trajectory, gives opposite-signed advantages to phases of differing verdict class in 68% of multi-phase resolved episodes—a sign structure an episode label cannot express."],"fun_headline_variants":["Self-declared phases beat fixed windows 2.38x","Agent-declared boundaries: 2.38x attribution without declarations","Coding agents' own hypotheses out-perform fixed windows 2.38x","No declarations, no gold patches: phases still map 2.38x"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that an agent prompted to declare its current hypothesis emits declarations that faithfully mark genuine changes in its causal belief state without changing its search behavior; the paper notes the reviewer-without-declarations ablation was never run, and the 9.8% token-cost measurement does not measure behavioral invariance.","fun_headline_variants_meta":{"raw":{"variants":["Self-declared phases beat fixed windows 2.38x","Agent-declared boundaries: 2.38x attribution without declarations","Coding agents' own hypotheses out-perform fixed windows 2.38x","No declarations, no gold patches: phases still map 2.38x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00143,"raw_usage":{"total_tokens":5668,"prompt_tokens":873,"completion_tokens":4795,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":4714}},"tokens_in":617,"tokens_out":4795,"duration_ms":29757,"temperature":1.0,"reasoning_tokens":4714,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T09:39:31.594899+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect matched task budgets under the declaration contract and under a no-declaration contract and compare the distributions of tool calls, actions, and outcomes; a significant divergence would mean the prompt itself changed search, so the boundary is compliance rather than epistemic state. As a second check, generate plausible hypotheses after the trajectory and rerun the attribution test—if post-hoc restatements attribute as well as live declarations, then the boundary position rather than the declaration carries the signal.","supporting_citations":[],"review_version":1}