{"id":"4d207e7a-e75f-47eb-b2d4-7fa0efc5aa10","arxiv_id":"2608.12960","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A GPT-based model trained on RL trajectories refines and fuses tract-specific tractography policies, improving Dice and overlap on public datasets while depending on atlas-derived reference fibers.","lead":"This thesis builds two deep-reinforcement-learning pipelines that reconstruct individual white-matter tracts of the brain directly from diffusion MRI, using a GPT-style transformer to refine or combine multiple RL agents. The authors report improved overlap and lower false positives on public benchmarks and argue this reduces the need for manually labeled fibers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed improvements over base RL policies are often within one standard deviation and overreach is not consistently reduced; the central 'improving' claim is not established without significance testing.","rationale":"The paper proposes two GPT-based frameworks that refine and fuse RL policies for tract-specific tractography, claiming improved Dice/overlap with reduced overreach and generalization without ground-truth fibers. For that claim to hold, the models must demonstrate clear, consistent improvements over the base policies. The reported results do not establish this: in numerous configurations the differences are smaller than the reported standard deviations (e.g., HCP AF 55.5±4.8 vs 54.6±4.7; HCP CST 69.4±1.3 vs 69.2±0.8), and overreach is sometimes higher than baselines (e.g., HCP AF 12.3 vs DET 10.3). No statistical tests accompany these numbers. The reader's weakest assumption concerned atlas-derived streamlines functioning as ground truth; I agree that this is a genuine issue, but it is secondary to the quantitative weakness, since even the stated performance improvements are not robustly demonstrated. The concrete paired test across all tract-dataset configurations will determine whether the improvements are real. The verdict should therefore remain conditional, pending that evidence.","tokens_in":45019,"tokens_out":8611,"duration_ms":82594,"concrete_test":"Recompute per-subject Dice/Overlap/Overreach from the saved models and run paired bootstrap or Wilcoxon signed-rank tests between FusionNet (or T-RLF) and each base policy (TD3, SAC, DDPG, and the best classical baseline), using identical seeds and test subjects as Tables 3.3-3.4 and 4.3-4.4. Also count how many of the 20+ tract-dataset configurations show an improvement over the best base policy exceeding one pooled standard deviation, and whether overreach improves or worsens in each. If the majority of comparisons are not significant (e.g., p>0.05) or overreach is worse in a substantial fraction, the headline improvement claim should be withdrawn.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is that Tract-RLFormer and FusionNet improve Dice/overlap and reduce overreach relative to the base policies they refine or fuse. The reported tables do not support this with the strength claimed. In Table 4.3 (HCP), AF Dice for FusionNet is 55.5±4.8 vs SAC 54.6±4.7 and vs DET 53.4±5.0; CG Dice in Table 4.4 is 56.6±2.0 vs SAC 56.2±1.9, and CST is 69.4±1.3 vs πavg 69.2±0.8. On TractoInferno, AF Dice is 53.2±8.9 vs SAC 52.3±8.9 and CG is 64.0±6.4 vs πmaxQ 63.5±6.3. Similar sub-point differences appear in Chapter 3: Tables 3.3-3.4 show T-RLF vs TD3 Dice deltas of only 0.2-2.0 points. Moreover, overreach is not consistently reduced: HCP AF OR is 12.3 for FusionNet vs 10.3 for DET; HCP CC OR is 44.8 vs DDPG's 37.3; TtoI AF OR is 39.1 vs TD3's 37.6. No p-values, confidence intervals, or paired tests are reported anywhere. Since the headline contribution is stated as improving these metrics, the lack of demonstrated statistical significance leaves the central claim unsupported. This is independent of, and more fundamental than, the atlas-prior concern: even if atlas priors were acceptable, the claimed policy improvement may be within noise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The thesis proposes two tract-specific white-matter tractography frameworks built on a GPT decoder-only sequence model. Tract-RLFormer first trains a TD3 policy inside tract-specific masks generated by a Mask Refinement Module, collects rollouts as (return-to-go, state, action) trajectories, and trains a GPT model with mixed-tract pretraining followed by tract-specific finetuning to obtain a refined tracking policy. TractRLFusion extends this idea to fuse three policies (TD3, SAC, DDPG), using Episodic Data Selection to retain anatomically plausible high-Q trajectories, a FusionNet trained on those selected rollouts, and Multi-Critic Policy Fine-Tuning to refine the fused policy. The paper claims improvements in Dice, overlap, and overreach relative to the base and ensemble policies, and reports generalization across TractoInferno, HCP, and ISMRM datasets, all under a stated premise of not relying on ground-truth fibers for training.","tokens_in":45431,"tokens_out":8125,"duration_ms":83130,"significance":"If the reported improvements were statistically supported, the framework would be a useful contribution: it shows a way to reuse RL rollouts and anatomical atlases to obtain tract-specific tracking policies, and to combine complementary RL policies without a supervised fiber dataset for the final policy. The final evaluation is against external reference tracts, so the headline result is not definitionally circular. The paper also has concrete strengths: extensive validation on public benchmarks, comparison with classical, supervised, and RL baselines, and ablations of MRM, the two-stage training, EDS, and MCPFT. The central weakness is quantitative: most reported differences are small relative to the shown variability, no significance testing is provided, and overreach is not consistently reduced. The no-ground-truth claim also needs qualification because atlas-derived reference streamlines are used for mask supervision, trajectory selection, and final cleaning.","major_comments":[{"comment":"The claim that FusionNet consistently improves Dice and better balances overlap versus overreach is not established by the reported numbers. Several headline differences are within one standard deviation: HCP AF Dice is 55.5±4.8 for FusionNet versus 54.6±4.7 for SAC; HCP CST Dice is 69.4±1.3 versus 69.2±0.8 for πavg; TtoI AF Dice is 53.2±8.9 versus 52.3±8.9 for SAC. Overreach is also often higher for FusionNet than for the least-overreaching baseline: HCP AF OR is 12.3 versus 10.3 for DET, HCP CC OR is 44.8 versus 37.3 for DDPG, TtoI AF OR is 39.1 versus 37.6 for TD3, and TtoI CG OR is 41.6 versus 30.4 for DDPG. No p-values, confidence intervals, or paired tests are reported anywhere, and the Chapter 3 tables report no variability at all. Since the abstract and Section 1.5 state improvement as the contribution, the current evidence is insufficient to support the central claim.","section":"§4.3.1, Tables 4.3–4.4"},{"comment":"The claim of training 'without ground-truth fibers' and 'without ground-truth annotations' is misleading as stated. The Mask Refinement Module is trained with binary cross-entropy against a voxel-wise ground truth derived from RecobundlesX atlas reference streamlines, with the text explicitly calling this 'the ground truth for each voxel'. Episodic Data Selection filters trajectories by MDF distance to atlas reference streamlines, and the final tract cleaning uses Fast Streamline Search against atlas reference tracts. The accurate statement is that the method does not use subject-specific ground-truth fiber sets, not that it avoids reference/annotation data entirely. Because the same atlas prior appears in mask generation, trajectory selection, and cleaning, it needs a sensitivity analysis or an ablation that removes FSS/atlas-based selection to demonstrate that the reported gains are not largely attributable to the atlas prior rather than to learned policy refinement.","section":"Abstract; §1.5; §3.3.1.2; §4.2.3; §3.3.4"},{"comment":"Tract-RLFormer is described as consistently outperforming the TD3 policy from which it was trained, but the reported gaps are very small and no standard deviations or tests are given. Examples include HCP left CG Dice 53.3 versus 53.0 for TD3, HCP right CG Dice 45.6 versus 45.2, and TtoI AF Dice 52.7 versus 51.8. Overreach is not consistently reduced: T-RLF has higher OR than TD3 for TtoI CG left (28.6 versus 27.3), AF right (49.8 versus 46.9), PYT left (17.2 versus 15.9), and CC (32.6 versus 26.1). In several rows classical DET/PROB also exceed T-RLF. These results do not support the Section 3.5 summary statement that the framework improves performance and reduces false positives.","section":"§3.4.3, Tables 3.3–3.4"},{"comment":"The fusion mechanism assumes that Q-values from TD3, SAC, and DDPG are comparable across policies, since EDS selects the policy with the maximum expected Q-value and MCPFT aggregates the three critics through Eq. (4.2.3). The three policies were trained with different discount factors, learning rates, and exploration/entropy settings (Table 4.1), so their critic outputs are not calibrated to a common scale. The paper does not discuss this cross-policy comparability issue, which is load-bearing for both the selection and the combined actor loss. A concise experimental justification, for example a study of Q-value distributions per policy or a normalized variant of the selection criterion, is needed.","section":"§4.2.3 and §4.2.5"}],"minor_comments":[{"comment":"The text says seven principal tracts are used, but Section 3.4 then adds the Optical Radius tract as an eighth tract; the abbreviation list also defines OR as 'Optical Radius', whereas the standard term is optic radiation.","section":"§3.4, §3.4.2"},{"comment":"There is a typo in the HCP DET overreach entry ('21..5'), and several other numbers use a non-standard notation for standard deviations in Chapter 4 tables; these should be cleaned up.","section":"Table 3.3"},{"comment":"The caption refers to 'Section 3.2.1.2', which does not exist; this cross-reference should be corrected.","section":"§3.4.4, Table 3.5"},{"comment":"The note explaining why PROB is omitted from some TractoInferno rows is informative, but it should be moved into the table caption or stated before the first table that uses this exclusion.","section":"§4.3.1"},{"comment":"The caption contains the misspelling 'Tract-RLForemer' and should be corrected.","section":"Figure 3.5"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know: this is a competent engineering thesis that combines GPT-based sequence modeling with RL tractography, but the headline claim that it improves over base policies is not actually supported by the numbers. The stress-test note is right—most reported gains are within one standard deviation, and overreach is not consistently reduced.\n\nWhat is genuinely new: the integration of Decision-Transformer-style sequence modeling into tract-specific RL tractography, and the specific fusion machinery (EDS, MCPFT) for combining TD3, SAC, and DDPG policies. I don't know of another paper doing exactly this. The thesis is also thorough in experimental scope: three public datasets, standard Dice/OL/OR metrics, ablations for MRM, EDS, MCPFT, and comparisons against classical, supervised, RL, and ensemble baselines. The architecture and training details are described clearly enough to reimplement, though no code is released.\n\nThe soft spots are proportionate to the central claim. First, statistical: HCP AF Dice is 55.5±4.8 for FusionNet vs 54.6±4.7 for SAC; CG is 56.6±2.0 vs 56.2±1.9; many Table 3.3–3.4 deltas are 0.2–2.0 points. There are no p-values, confidence intervals, or paired tests anywhere, and overreach sometimes goes the wrong direction (HCP AF OR 12.3 vs DET 10.3; CC OR 44.8 vs DDPG 37.3). So \"improving\" is asserted, not demonstrated. Second, the \"no ground-truth fibers for training\" claim is overstated: atlas reference streamlines from RecobundlesX are used in MRM mask training, in EDS via MDF-distance selection, and in FSS cleaning. That is a reasonable anatomical prior, but it is still reference data. Third, there is an internal contradiction: Section 3.4.2 says tract-specific TD3 cannot be tested on HCP or ISMRM, yet Tables 3.3 and 3.4 report TD3 scores on exactly those datasets.\n\nWho is this for? Someone working on RL-based tractography or on using sequence models to refine/fuse policies would find the architecture ideas useful. It deserves a serious referee because the problem is relevant and the experiments are extensive, but the referee should require significance testing, a clearer statement about the role of atlas priors, and a fix of the contradictory claims before the headline conclusion is accepted. My recommendation: send it to peer review, but expect major revisions.","headline":"Solid engineering thesis with a plausible but statistically unsupported claim of improvement over base RL policies; worth refereeing with major revisions.","tokens_in":45914,"tokens_out":3648,"would_cite":false,"duration_ms":36802,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-based sequence models can refine and fuse RL tracking policies for tract-specific brain tractography without ground-truth fibers.","keywords":["white matter tractography","reinforcement learning","policy refinement","policy fusion","GPT sequence modeling","tract-specific tractography","diffusion MRI","no ground-truth training"],"falsifier":"Train the framework under two conditions: once with the true atlas reference streamlines and once with the same atlas rotated, warped, or replaced by a different bundle atlas, keeping all RL rollouts and training hyperparameters fixed. If Dice, overlap, and overreach on held-out subjects do not degrade substantially in the corrupted-atlas condition, the no-ground-truth claim stands; if they do, the reported gains are artifacts of atlas priors rather than evidence of learned policy refinement.","tokens_in":44796,"feed_emoji":"🧠","tokens_out":6877,"duration_ms":64094,"temperature":0.7,"pith_summary":"Tractography reconstructs white-matter pathways from diffusion MRI, but whole-brain tracking produces many false positives and needs a separate segmentation step, while supervised training requires ground-truth fibers that are rarely available. This thesis claims that a tract-specific hybrid of reinforcement learning and supervised sequence modeling can address both problems at once. In the proposed frameworks, a GPT-style decoder-only transformer is trained offline on rollouts from RL tracking agents, represented as state, action, and return-to-go tuples, first on a mixed set of tracts and then fine-tuned per tract, without ever seeing ground-truth streamlines. The refined policy, called Tract-RLFormer, and the fused policy, called TractRLFusion, combining TD3, SAC, and DDPG, are reported to improve Dice and overlap while reducing overreach compared with base policies and whole-brain methods, and to generalize from TractoInferno to HCP and ISMRM data. If the claim is right, tract-specific tractography no longer needs labeled fibers or an explicit segmentation stage, only atlas-derived masks and RL experience.","feed_headline":"GPT-trained policies refine brain tractography without ground truth","feed_subtitle":"Two frameworks turn RL rollouts into tract-specific trackers that beat whole-brain baselines on three public datasets.","key_machinery":"The central object is the trajectory token sequence: each timestep contributes a return-to-go scalar, a 334-dimensional state built from spherical-harmonic coefficients of a voxel and its neighbors, mask values, and the last four tracking directions, plus a 3-dimensional action, all processed by a causal decoder-only transformer with a 40-token context. This is the mechanism by which offline RL experience becomes a tract-specific policy: the transformer is pre-trained on mixed-tract trajectories and fine-tuned per tract, and at inference the return-to-go is fixed to a high expert value of 300 so the model generates actions conditioned on the promised return. Around this core, the Mask Refinement Module prunes dilated atlas masks into subject-specific tracking regions, Episodic Data Selection curates trajectories by MDF distance and Q-value, and Multi-Critic Policy Fine-Tuning anchors the fused actor with the original policies' critics.","core_discovery":"The central claim is that a trajectory-level sequence model can refine and fuse RL tracking policies more effectively than the policies can perform on their own. In Tract-RLFormer, a TD3 agent is trained per tract inside masks produced by the Mask Refinement Module; its rollouts are converted into return-to-go, state, and action trajectories; three decoder layers are pre-trained on a mixed-tract dataset of 150,000 trajectories, and a fourth layer is fine-tuned per tract with a five-step cosine angular loss. TractRLFusion repeats this idea across three policies: Episodic Data Selection keeps trajectories whose streamlines are within a 5 mm mean direct-flip distance of atlas reference streamlines and, across policies, selects the trajectories with the highest predicted Q-value; a GPT-based FusionNet is trained on the curated data and then refined by Multi-Critic Policy Fine-Tuning, in which the TD3, SAC, and DDPG critics add Q-value gradients to the five-step loss. The paper reports the highest Dice among the compared methods on nearly every tract and dataset, with lower overreach than the exploratory SAC policy and higher overlap than conservative TD3 and DDPG, and it attributes the gain to the balance the fused policy strikes between overlap and overreach.","pith_inferences":["If the no-ground-truth claim holds, the same offline trajectory-refinement recipe should transfer to other RL-based curve-tracing tasks in medical imaging, such as vessel or airway delineation, where annotated centerlines are scarce.","The method's dependence on atlas reference streamlines means its no-ground-truth claim is really no subject-specific ground truth; a testable extension would be to measure how much of the reported Dice gain disappears when the atlas is swapped or misaligned.","Because FusionNet outperforms TractSeg even inside TractSeg's own masks, the learned fused policy rather than the superior mask may be doing part of the work; this could be isolated by running the base policies inside the same masks.","The return-to-go conditioning, fixed at 300 during inference, is a plausible control knob for the overlap–overreach trade-off; the paper does not explore it, but varying this value could offer a simple user-facing sensitivity dial."],"forward_implications":["Tract-specific RL policies can be improved without ground-truth fibers, using only atlas-derived masks, RL rollouts, and a return-to-go conditioned sequence model.","The same pre-trained GPT backbone, fine-tuned per tract, transfers from TractoInferno training data to HCP and ISMRM test data, so a single foundation model could cover many bundles.","Fusing a conservative policy such as TD3 or DDPG with an exploratory policy such as SAC produces a better overlap–overreach balance than any single policy or decision-level voting or averaging.","Tract-specific masks alone raise the Dice of even untrained classical DET and PROB trackers, so mask quality is a major lever on tractography accuracy.","Because training is offline on rollouts, the framework can be extended to new RL policies or new datasets by re-running trajectory collection and fine-tuning, without re-training from scratch."],"supporting_citations":[{"why":"Supplies the RL environment for tractography, including the state, action, reward, and termination conditions that the policies interact with.","marker":"[48]"},{"why":"Provides the RL-based tractography baselines and documents the overlap–overreach trade-off that the fusion framework targets.","marker":"[49]"},{"why":"Source of the TractoInferno training and test subjects and of the supervised-learning baseline scores the paper compares against.","marker":"[42]"},{"why":"Supplies the HCP reference tract segmentations used for generalization evaluation and the TractSeg masks and baselines.","marker":"[28]"},{"why":"Provides the atlas reference streamlines that drive the Mask Refinement Module, Episodic Data Selection, and tract cleaning.","marker":"[69]"},{"why":"Establishes the return-to-go conditioned sequence-modeling recipe that Tract-RLFormer and FusionNet adopt for offline policy learning.","marker":"[70]"},{"why":"Defines the TD3 algorithm used as the level-1 policy in Tract-RLFormer and as one of the three fused policies.","marker":"[56]"},{"why":"Defines SAC, the exploratory policy whose high overlap and high overreach behavior is balanced by the fusion.","marker":"[57]"},{"why":"Defines DDPG, the conservative policy fused alongside TD3 and SAC.","marker":"[55]"},{"why":"Provides the Fast Streamline Search used in post-processing to remove false-positive fibers against atlas references.","marker":"[71]"}],"fun_headline_variants":["No ground truth needed: RL plus GPT refines brain tractography","Tract-specific RL tracker fuses policies to cut false positives","Hybrid RL and supervised learning beat whole-brain tractography","Fusing multiple RL policies yields robust white-matter mapping"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that atlas-derived reference streamlines are accurate enough, after registration, to stand in for ground truth when generating masks, selecting training trajectories, and cleaning tracts; if they are wrong for a subject or tract, the claimed gains could come from the atlas prior rather than from learned policy improvement.","fun_headline_variants_meta":{"raw":{"variants":["No ground truth needed: RL plus GPT refines brain tractography","Tract-specific RL tracker fuses policies to cut false positives","Hybrid RL and supervised learning beat whole-brain tractography","Fusing multiple RL policies yields robust white-matter mapping"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00075,"raw_usage":{"total_tokens":3408,"prompt_tokens":1079,"completion_tokens":2329,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":695,"completion_tokens_details":{"reasoning_tokens":2259}},"tokens_in":695,"tokens_out":2329,"duration_ms":17941,"temperature":1.0,"reasoning_tokens":2259,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:42:10.720480+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the framework under two conditions: once with the true atlas reference streamlines and once with the same atlas rotated, warped, or replaced by a different bundle atlas, keeping all RL rollouts and training hyperparameters fixed. If Dice, overlap, and overreach on held-out subjects do not degrade substantially in the corrupted-atlas condition, the no-ground-truth claim stands; if they do, the reported gains are artifacts of atlas priors rather than evidence of learned policy refinement.","supporting_citations":[{"cited_title":"Track-to-learn: A general framework for tractography with deep reinforcement learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the RL environment for tractography, including the state, action, reward, and termination conditions that the policies interact with."},{"cited_title":"What matters in reinforcement learning for tractography,","cited_arxiv_id":null,"evidence_quote":"Provides the RL-based tractography baselines and documents the overlap–overreach trade-off that the fusion framework targets."},{"cited_title":"Tractoinferno-a large-scale, open-source, multi-site database for machine learning dmri tractography,","cited_arxiv_id":null,"evidence_quote":"Source of the TractoInferno training and test subjects and of the supervised-learning baseline scores the paper compares against."},{"cited_title":"Tractseg-fast and accurate white matter tract segmentation,","cited_arxiv_id":null,"evidence_quote":"Supplies the HCP reference tract segmentations used for generalization evaluation and the TractSeg masks and baselines."},{"cited_title":"Population average atlas for recobundlesx,","cited_arxiv_id":null,"evidence_quote":"Provides the atlas reference streamlines that drive the Mask Refinement Module, Episodic Data Selection, and tract cleaning."},{"cited_title":"Decision transformer: Reinforcement learning via sequence modeling,","cited_arxiv_id":null,"evidence_quote":"Establishes the return-to-go conditioned sequence-modeling recipe that Tract-RLFormer and FusionNet adopt for offline policy learning."},{"cited_title":"Addressing function approximation error in actor-critic methods,","cited_arxiv_id":null,"evidence_quote":"Defines the TD3 algorithm used as the level-1 policy in Tract-RLFormer and as one of the three fused policies."},{"cited_title":"Soft actor-critic: Off-policy maximumentropydeepreinforcementlearningwithastochasticactor,","cited_arxiv_id":null,"evidence_quote":"Defines SAC, the exploratory policy whose high overlap and high overreach behavior is balanced by the fusion."},{"cited_title":"Fast streamline search: An exact technique for diffusion mri tractography,","cited_arxiv_id":null,"evidence_quote":"Provides the Fast Streamline Search used in post-processing to remove false-positive fibers against atlas references."}],"review_version":1}