{"id":"bee48901-cc3b-46fe-9225-1db049542a45","arxiv_id":"2608.01604","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Post-training Qwen3.5-122B-A10B on 363 office workflow tasks improved SWE-Bench Pro pass@1 by 5.8 points, with trajectory analysis attributing the gain to four general goal-directed behaviors.","lead":"The authors post-trained a large language model on office-style multi-tool tasks and found that its score on a software engineering benchmark improved by 5.8 points, despite no coding tasks in training. The paper interprets the gain as evidence that long-horizon training strengthens general goal-directed behaviors that transfer across domains.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single unseeded training run leaves the +5.8pp SWE-Bench Pro gain without a variance estimate; the transfer claim needs multi-seed replication before the GDE account can be evaluated.","rationale":"The paper is unusually careful: it labels the causal mechanism as a hypothesis (Section 7), lists single-run and outcome-conditioned analysis as limitations (Section 8), and does not overclaim. The measured +5.8pp on SWE-Bench Pro is a real, directionally striking observation, and the paired trajectory examples are concrete. My concern is not about integrity or data provenance per se; I did not find internal evidence that LHMTA contained SWE-Bench Pro material, and the authors explicitly assert isolation. The weakest point is statistical: every headline number comes from one training run. The paper gives no error bars, and the behavioral analysis is conditioned on the 77 newly-passed tasks, with the implied roughly 35 regressions unexamined. A multi-seed replication is the single check that would establish whether the transfer is a property of the training intervention or of one stochastic draw. This does not invalidate the preprint, but it keeps the verdict at conditional rather than full acceptance.","tokens_in":16756,"tokens_out":8571,"duration_ms":85103,"concrete_test":"Retrain the full pipeline (SFT warm-up plus GSPO) from the same base model with at least three independent random seeds, evaluate each final checkpoint on SWE-Bench Pro with the identical harness, and report per-seed pass@1 plus a confidence interval for the improvement. If the interval includes zero, the central transfer claim is not established. In the same runs, also report the number of pass-to-fail regressions to check whether the GDE gains are monotonic.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is a quantitative comparison between one base checkpoint and one post-trained checkpoint. The paper reports no seed variability for the training pipeline (LoRA initialization, rollout sampling in GSPO, and optimizer noise are all stochastic), no confidence interval, and no significance test. A naive two-proportion test on n=731 would give SE of roughly 2.2 pp, but that treats the benchmark as the only noise source and ignores training-seed variance, which can be comparable to or larger than the 5.8 pp effect. The paper itself states in Section 8 that \"All results come from one base model and one post-training run\" and that reproducibility across seeds is not established. If the effect is seed-dependent, the claimed cross-domain transfer and the GDE interpretation built on it do not follow. Additionally, the reported 77 newly-passed SWE-Bench Pro cases imply roughly 35 pass-to-fail regressions (net +42 on 731 tasks); these regressions are never analyzed, so the directional \"improvement in all four GDE behaviors\" narrative is incomplete. The load-bearing point is that the transfer effect is currently a single point estimate.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a post-training experiment in which Qwen3.5-122B-A10B is trained on 363 Long-Horizon Multi-Tool Agent (LHMTA) office-workflow tasks, with no software-engineering content, and then evaluated on SWE-Bench Pro. The central empirical claim is a pass@1 improvement from 20.5% to 26.3% (+5.8 percentage points) under greedy decoding, which the authors interpret as evidence that long-horizon post-training strengthens a domain-general capability they call goal-directed execution (GDE). GDE is operationalized through four behaviors—goal formation, state construction, goal stability, and verification—and the paper presents paired trajectory analyses and aggregate SWE-Bench Pro metrics as evidence that these behaviors improved after training in both office and software domains. Section 8 explicitly concedes that the experiment is a single model and a single post-training run, that the behavioral framework was refined iteratively on the same trajectories used to illustrate it, that the case analysis is outcome-conditioned, and that no causal or domain-matched counterfactual was run.","tokens_in":16946,"tokens_out":4160,"duration_ms":38935,"significance":"If the +5.8pp transfer is robust, the result is significant: it would show that long-horizon post-training on non-software tasks can improve software-engineering performance, with practical implications for data selection and for behavioral accounts of agent post-training. The paper's strengths are the use of an external benchmark with no fitted parameters, the detailed and deterministic definitions of the aggregate behavioral metrics in Appendix A, and the paired-trajectory methodology that grounds the qualitative analysis in concrete tool calls and artifacts. However, the central transfer claim is a single point estimate with no variance estimate, and the behavioral explanation is partly circular because the four capabilities were developed and demonstrated on the same outcome-conditioned trajectories. The significance of the result therefore hinges on reproducibility and on validation of the behavioral coding that the current manuscript does not yet provide.","major_comments":[{"comment":"The central transfer claim rests on a single post-training run; the paper provides no variance estimate, confidence interval, or significance test for the +5.8pp SWE-Bench Pro gain. LoRA initialization, GSPO rollout sampling, and optimizer noise are stochastic, so training-seed variance could be comparable to or larger than the reported effect. Section 8 concedes this ('one base model and one post-training run'), but it remains load-bearing: without multi-seed replication or at least a reproducibility check, the cross-domain transfer and the GDE interpretation built on it are not yet established.","section":"Section 5.2 and Table 3"},{"comment":"The four GDE capabilities were refined iteratively through qualitative analysis of the same outcome-conditioned trajectories used to demonstrate them, so the claim that 'matched trajectory analysis shows gains in all four GDE behaviors' is partly circular. The framework was not fixed before inspecting the data, and the cases shown were selected for explanatory clarity rather than randomly. An independent, pre-registered coding protocol with blinded annotators and inter-rater reliability, or validation on held-out trajectories, is needed for the behavioral claim to support the transfer account.","section":"Section 6.1 and Section 8"},{"comment":"The aggregate behavioral analysis covers 731 paired SWE-Bench Pro tasks but reports no analysis of the approximately 35 pass-to-fail regressions implied by 77 newly passed tasks and a net gain of 42. The paper's directional narrative—improvement in all four GDE behaviors—is incomplete without understanding whether regressions exhibit the same GDE failures or different ones; this is directly relevant to whether post-training strengthened GDE rather than merely shifted the model's success set.","section":"Section 6.3 and Table 5"},{"comment":"The claimed isolation of training from SWE-Bench Pro is asserted but not documented. Because the transfer interpretation depends on no benchmark instances, graders, checkpoint-selection signal, or hyperparameter-tuning signal from SWE-Bench Pro, the paper should provide a provenance audit or contamination check: dataset hashes, exact task lists, overlap tests with SWE-Bench Pro repositories and issues, and a statement of how the checkpoint was selected. This is a factual premise that cannot be independently verified from the current text.","section":"Section 5.2"}],"minor_comments":[{"comment":"There are missing spaces in 'capabilitygoal-directed execution' and 'we usegoal-directed execution'; please fix the typography.","section":"Abstract and Section 1"},{"comment":"The paired figures are illustrative and were selected after observing outcomes; each figure should explicitly state that the cases are not randomly sampled and should include case identifiers or repository links to support independent audit.","section":"Section 6.2 and Figures 5-8"},{"comment":"The definition of 'mean first-test position' depends on commands classified as the assistant's shell commands; consider also reporting the fraction of runs with no formal test separately, since normalizing only among tested runs can mask a bimodal distribution.","section":"Table 5 and Appendix A"},{"comment":"The related-work discussion mentions HiAgent and ReCAP in a single sentence; a slightly fuller comparison would clarify what the goal-loop representation adds over these existing hierarchical working-memory and recursive-planning approaches.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper would benefit from a data-availability statement that releases the LHMTA task list, the SWE-Bench Pro contamination check, and the exact checkpoint-selection procedure. I also note that the manuscript relies substantially on unpublished work from the authors' own group (Ritchie et al. 2026; Mehta et al. 2026a,b); the editor may wish to ensure the claimed novelty relative to those papers is clear. The single-run nature of the central result is the main risk to the paper's contribution, and it is fixable only by additional experiments or a strong reproducibility argument."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — this paper asks a good question: does long-horizon post-training on non-software tasks transfer to software engineering? The headline result is a 5.8pp pass@1 gain on SWE-Bench Pro from training on office workflows, and that result is genuinely interesting. If it holds, it changes how we think about post-training data selection. The authors also deserve credit for a clean benchmark setup: the training collection appears free of SWE content, and they state explicitly that no benchmark feedback was used. And the limitations section is unusually candid — they admit the single-run design up front.\n\nThe soft spots are in proportion: the central claim is a single point estimate. One model, one training run, no seeds, no confidence interval. The transfer gain of 5.8pp on 731 tasks corresponds to about 42 net new passes. The paper reports that the trained model newly passed 77 SWE-Bench Pro tasks. Those two numbers only reconcile if there were roughly 35 regressions, and regressions are never mentioned. That is a real gap. It matters because the behavioral story — 'gains in all four GDE behaviors' — is built on the newly-passed cases only. If 35 tasks got worse, the paper owes the reader an account of how common regressions were and what they looked like.\n\nThe GDE framework itself is a reasonable qualitative vocabulary, but it is not new in substance: TOTE, Soar, CoALA, and the failure-taxonomy literature say much the same thing. The authors cite those lines, so no one is being misled. The risk is that the framework does no explanatory work beyond labeling; the transfer result is the empirical payload, and it is currently under-verified.\n\nThe data provenance — no SWE instances in LHMTA, no grader feedback — is asserted but not independently verifiable. That's not a flaw in the writing; it's a limitation that artifact release would address.\n\nMy recommendation: send it to a serious referee, but condition acceptance on replication evidence or at least variance estimates across seeds, a reconciliation of the 77 vs +5.8pp numbers, and a discussion of regressions. The work is honest and the question is important. It just isn't proven yet.","headline":"The +5.8pp transfer is plausible but rests on a single unseeded run; the GDE framing adds interpretation rather than evidence, and the paper needs multi-seed replication and a reconciliation of its reported counts before the result is trustworthy.","tokens_in":17500,"tokens_out":2921,"would_cite":false,"duration_ms":25626,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Office-workflow training lifts coding benchmark by 5.8 points","keywords":["goal-directed execution","cross-domain transfer","long-horizon post-training","SWE-Bench Pro","office workflows","reinforcement learning","agent trajectories","behavioral analysis"],"falsifier":"An audit would settle it: if the 363-task LHMTA snapshot or the supervised teacher trajectories contain repository content, tests, or grader outputs from SWE-Bench Pro tasks, or if retraining with the same data but selecting the checkpoint on the LHMTA holdout alone fails to reproduce the roughly 5.8-point SWE-Bench Pro gain, the transfer claim as stated would not survive. A cleaner experiment is to repeat the full two-stage run with a held-out copy of SWE-Bench Pro never opened during development and pre-register the checkpoint-selection rule.","tokens_in":1616,"feed_emoji":"🤖","tokens_out":3910,"duration_ms":82740,"temperature":0.7,"pith_summary":"The paper tries to establish that long-horizon post-training on tasks from one domain can strengthen a domain-general capability, which it calls goal-directed execution (GDE), rather than only carving task-specific skills. It post-trains a 122-billion-parameter model on 363 office-workflow tasks that contain no software-engineering content, then measures the model on the SWE-Bench Pro coding benchmark: first-attempt success rises from 20.5% to 26.3%. Paired trajectory analysis of tasks that the base model failed and the trained model passed shows gains in all four GDE behaviors—goal formation, state construction, goal stability, and verification—in both office and software settings. The causal link to long-horizon task structure is presented as a hypothesis, but if the empirical transfer holds, it implies that data-curation and post-training decisions should account for the structural demands a task exercises, not just its subject matter.","feed_headline":"Office-workflow training lifts coding benchmark by 5.8 points","feed_subtitle":"A 122B model trained on 363 office tasks, no coding content, improved software pass@1 from 20.5% to 26.3%.","key_machinery":"The carrying mechanism is a recursive goal-loop model of agent behavior, formalized as goal-directed execution (GDE): the agent forms a goal from its parent objective and working state, acts, updates its working state from environment feedback, and verifies whether the state satisfies the goal, decomposing into nested loops when actions are too abstract. The four capabilities named by GDE—goal formation, state construction, goal stability, and verification—are the interpretive grid used to compare base and trained trajectories. Training itself runs through a two-stage post-training recipe on 363 LHMTA office tasks: a supervised warm-up on high-scoring teacher trajectories, followed by GSPO reinforcement learning with a dense reward equal to the fraction of grader criteria satisfied per trajectory. The design isolates the structural demand of long-horizon work because the training tasks share no content with software engineering.","core_discovery":"The central discovery is that a model trained entirely on office workflows—documents, spreadsheets, web research, file manipulation, and scheduling—improves at resolving software-repository issues. The trained checkpoint scores 26.3% pass@1 on SWE-Bench Pro under greedy decoding versus 20.5% for the base model, a gain of 5.8 percentage points, despite the training collection containing no software-engineering tasks, graders, or benchmark feedback. The paper argues that the gain is best explained behaviorally: post-training strengthens four observable capabilities that transfer across domains—forming the right next goal, constructing and maintaining task-relevant state, preserving higher-level requirements under local pressure, and verifying completion against the environment. Matched trajectory analysis of 103 newly-passing tasks shows these differences in both domains, and aggregate SWE-Bench Pro metrics shift in the same direction: less repeated retrieval, more contact with reference-patch files, far smaller patches, and nearly double the share of runs that execute a formal test.","pith_inferences":["If the transfer is causal, it predicts similar gains from other non-software long-horizon domains, such as customer support, scientific data analysis, or healthcare administration, whenever tasks share deep decomposition, parallel synthesis, entangled constraints, and long dependent chains.","A natural ablation would hold training data volume constant and vary one task demand at a time; the paper's account predicts that deep decomposition and long dependent chains are the largest drivers because they most directly exercise goal formation and verification.","The aggregate metrics suggest a practical test for production pipelines: monitor retrieval repetition, patch size, and test timing as inexpensive proxies for goal-directed execution during post-training, even when the target domain is software engineering."],"forward_implications":["Post-training data value should be measured by the behavioral demands it exercises, not only by topic: a domain with no coding content can improve a coding benchmark.","The two-stage recipe—supervised warm-up plus GSPO with dense criterion-level rewards—produces transfer beyond the training environment: LHMTA holdout +17.5pp, Toolathlon +9.6pp, BFCL-V4 +3.5pp, and SWE-Bench Pro +5.8pp.","The trained model adds roughly one quarter as many lines (111.7 versus 415.5) while touching more reference-patch files, indicating more targeted implementation rather than brute-force exploration.","Formal-test runs nearly double, from 37.5% to 73.3%, with earlier first-test position, suggesting long-horizon post-training shifts agents toward verification-heavy execution.","GDE provides a shared vocabulary for failures across domains: a goal-stability breakdown in a scheduling task and a stale-test override in a repository are the same capability failing at different surfaces."],"supporting_citations":[{"why":"Provides the SWE-Bench Pro target benchmark whose pass@1 defines the measured +5.8pp transfer.","marker":"[Deng et al., 2025]"},{"why":"Supplies the GSPO sequence-level estimator used in the RL stage that produced the trained checkpoint.","marker":"[Zheng et al., 2025]"},{"why":"Supplies the LoRA adaptation method used to train the model in the two-stage recipe.","marker":"[Hu et al., 2022]"},{"why":"Defines the Model Context Protocol through which the LHMTA office-workflow environments are exposed to the agent.","marker":"[Anthropic, 2024]"}],"fun_headline_variants":["Office training, zero code, lifts coding score by 5.8","No coding tasks in data, yet coding up 5.8 points","Office-workflow training betters code benchmark by 5.8","Office tasks only, yet code repair improves 5.8 points"],"cache_read_input_tokens":19712,"weakest_assumption_plain":"The load-bearing premise is that the SWE-Bench Pro coding benchmark never influenced the training run: no benchmark tasks or grader outputs entered the 363-task collection, and no checkpoint, hyperparameter, or reward choice was made using benchmark scores.","fun_headline_variants_meta":{"raw":{"variants":["Office training, zero code, lifts coding score by 5.8","No coding tasks in data, yet coding up 5.8 points","Office-workflow training betters code benchmark by 5.8","Office tasks only, yet code repair improves 5.8 points"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001717,"raw_usage":{"total_tokens":6799,"prompt_tokens":958,"completion_tokens":5841,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":5764}},"tokens_in":574,"tokens_out":5841,"duration_ms":33703,"temperature":1.0,"reasoning_tokens":5764,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:06:09.223084+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An audit would settle it: if the 363-task LHMTA snapshot or the supervised teacher trajectories contain repository content, tests, or grader outputs from SWE-Bench Pro tasks, or if retraining with the same data but selecting the checkpoint on the LHMTA holdout alone fails to reproduce the roughly 5.8-point SWE-Bench Pro gain, the transfer claim as stated would not survive. A cleaner experiment is to repeat the full two-stage run with a held-out copy of SWE-Bench Pro never opened during development and pre-register the checkpoint-selection rule.","supporting_citations":[],"review_version":1}