{"id":"288d4e10-2a72-459d-af22-3ffab4d7b846","arxiv_id":"2508.00496","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"LesiOnTime segments small breast lesions in longitudinal DCE-MRI using temporal attention and BI-RADS consistency regularization, reporting a 5% Dice gain over baselines.","lead":"A new deep learning method, LesiOnTime, combines information from previous breast MRI scans and radiologist BI-RADS scores to segment small breast lesions. If the reported 5% Dice improvement holds, it could help radiologists detect emerging cancers earlier in high-risk screening programs.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The submitted full text is a different paper (HannesImitation), so the LesiOnTime 5% Dice claim has no inspectable evidence in this manuscript; the central result is unverifiable as submitted.","rationale":"The reader correctly identified that the supplied full text is a different paper and therefore returned UNVERDICTED. My stress-test agrees with that outcome: the central claim is unsupported because no methods or experiments for LesiOnTime are present in this submission. However, the reader framed the weakest assumption as BI-RADS label reliability, which presupposes that the LesiOnTime method is available to inspect. The more load-bearing concern is more basic: the submission provides no LesiOnTime text at all, so neither the TPA block, the BCR loss, the dataset, nor the Dice comparison can be evaluated. The proposed check therefore targets the immediate barrier—mismatched full text—and then, if the correct text is retrieved, the reproducibility of the headline result via the public code. I do not see grounds to reject the underlying scientific claim on its merits, because no merits are visible; similarly, there is no evidence to accept it. The reader's UNVERDICTED verdict should stand unchanged. Credit is given where due: the abstract states the code is public, which is a positive signal, but a code link alone cannot substitute for the missing methods and results in the manuscript body.","tokens_in":4437,"tokens_out":2997,"duration_ms":32333,"concrete_test":"Fetch the current arXiv record for 2508.00496 and check whether the PDF body corresponds to LesiOnTime; if it does not, request the corrected manuscript and re-review Sections 3–4. Once the correct text is available, independently rerun the LesiOnTime training and evaluation using the public code and dataset; if the code or data are unavailable, the claimed 5% Dice improvement remains unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that LesiOnTime outperforms state-of-the-art single-timepoint and longitudinal baselines by 5% in Dice—cannot be checked because the body of arXiv:2508.00496 is a different manuscript, 'HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning' (arXiv:2508.00491). No LesiOnTime methods, dataset details, baseline definitions, or experimental tables appear in the submission, so the comparison and the BCR loss are entirely unverifiable. The only potentially checkable artifact is the claimed public repository, but without the correct manuscript text the reader cannot tell what the code is supposed to reproduce. If the body were corrected, the next weakest point would be the BCR loss's reliance on BI-RADS labels: the abstract does not state whether BI-RADS scores are needed at inference, and no inter-rater agreement is reported, so the regularization could encode label noise or a test-time shortcut. As submitted, however, the primary barrier is evidential, not methodological.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract of arXiv:2508.00496 announces LesiOnTime, a 3D segmentation method for small breast lesions in longitudinal DCE-MRI that combines a Temporal Prior Attention (TPA) block with a BI-RADS Consistency Regularization (BCR) loss, and reports a 5% Dice improvement over single-timepoint and longitudinal baselines on a curated in-house dataset of high-risk patients. However, the submitted full text is not the LesiOnTime paper: it is the complete text of \"HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning\" (arXiv:2508.00491). Consequently, the submission contains no description of the LesiOnTime method, no dataset details, no baseline definitions, no experimental tables, and no ablation studies. The only available evidence for the central claim is the abstract itself, which reports a single percentage gain without error bars, significance tests, or external validation.","tokens_in":4580,"tokens_out":2866,"duration_ms":28428,"significance":"If the claimed results were properly documented and validated, the proposed idea would be clinically relevant: incorporating longitudinal imaging and BI-RADS clinical scores into lesion segmentation addresses a real gap in screening workflows, and the two proposed components (TPA and BCR) are plausible and complementary. The promise of public code is also a strength. However, because the submitted manuscript does not contain the LesiOnTime text at all, none of these contributions can be inspected, checked, or placed in context. The significance of the work is therefore entirely conditional on the existence of a correct full text that is not present in this submission.","major_comments":[{"comment":"The body of arXiv:2508.00496 is the paper \"HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning\", not the LesiOnTime manuscript. None of the central elements of the claimed contribution appear in the submission: the Temporal Prior Attention block, the BI-RADS Consistency Regularization loss, the longitudinal DCE-MRI dataset, the baseline definitions, the evaluation protocol, or the ablation studies. The abstract's claim of a 5% Dice improvement is therefore completely unverifiable from the submitted text, and the manuscript as it stands cannot be assessed for correctness or reproducibility.","section":"Full text"},{"comment":"The abstract does not state whether BI-RADS scores are required at inference time or are used only during training. The BCR loss enforces latent-space alignment for scans with similar radiological assessments, so the reliability of the BI-RADS labels is load-bearing for the claimed gains. The authors should report inter-rater agreement or otherwise validate label consistency; without this, the regularization could encode label noise and the reported 5% improvement could be an artifact. This concern can only be properly examined once the actual method description is present.","section":"Abstract"},{"comment":"The reported improvement is given as a single percentage point (5% Dice) with no error bars, no number of patients or scans, no cross-validation scheme, and no statistical significance test. Even with the correct manuscript, a single aggregate number without variance or significance reporting would be insufficient to support the claim that LesiOnTime outperforms state-of-the-art baselines, especially for small lesion segmentation where Dice scores can vary substantially across cases.","section":"Abstract"},{"comment":"The reference list and related-work discussion belong to the HannesImitation paper on prosthetic-hand imitation learning, so the submission provides no context for the claimed longitudinal DCE-MRI segmentation comparison. No prior single-timepoint or longitudinal breast lesion segmentation methods are cited, and no BI-RADS-based modeling approaches are discussed; thus the state-of-the-art comparison mentioned in the abstract cannot be placed in the literature.","section":"Full text"}],"minor_comments":[{"comment":"The abstract mentions a public repository (https://github.com/cirmuw/LesiOnTime), but the submission contains no code, no reproducibility checklist, and no verification that the repository corresponds to the described method; please confirm the repository is accessible and properly linked.","section":"Abstract"},{"comment":"The arXiv submission's abstract and full text describe two completely different papers with different titles, authors, and subject areas; the submission should be corrected or withdrawn so that the abstract and body refer to the same work.","section":"Metadata"},{"comment":"The phrasing \"outperforms state-of-the-art ... by 5% in terms of Dice\" is ambiguous; it should specify the exact Dice variant (e.g., whole-volume vs. per-lesion, lesion-wise vs. scan-wise) and enumerate the baselines included in the comparison.","section":"Abstract"}],"recommendation":"reject","confidential_remarks":"This appears to be a submission error in which the full text of a different arXiv paper was attached. As submitted, the manuscript does not contain the research it claims to present, so no substantive technical review is possible. The appropriate action is rejection; the authors may resubmit a corrected version with the proper LesiOnTime text, at which point the methodological concerns about the BCR loss and evaluation rigor can be evaluated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about arXiv:2508.00496. The abstract describes a sensible idea: a 3D segmentation network for small breast lesions in longitudinal DCE-MRI, with a Temporal Prior Attention block and a BI-RADS Consistency Regularization loss, reporting a 5% Dice gain over single-timepoint and longitudinal baselines on an in-house high-risk screening dataset. The full text of the submission, however, is not that paper. It is a completely unrelated manuscript about imitation learning for the Hannes prosthetic hand. Nothing in the supplied body describes LesiOnTime, its architecture, training, dataset, or evaluation. So the central claim hangs on an abstract alone. What is genuinely new here, taking the abstract at face value, is the combination of temporal attention with a clinical-score consistency loss for longitudinal breast MRI. Each component exists elsewhere, but the pairing is not standard, and small-lesion segmentation in screening is a clinically relevant gap. If the 5% gain is real and reproducible, it matters to the medical imaging community. The soft spot is not subtle. As submitted, there is no evidence to inspect. The comparison, the ablations, the error bars, the handling of BI-RADS noise—none of it exists in this manuscript. Even after fixing the body, I would want to know two things: whether BI-RADS scores are available at inference or only during training, and what the inter-rater agreement is on the scores used to supervise the consistency loss. The abstract does not say. That is a legitimate methodological concern, but it is secondary to the fact that the paper as submitted is not the paper. This paper is for a reader interested in longitudinal medical imaging and clinically conditioned segmentation, but only in a corrected version. The current submission would not survive desk screening, and it shouldn't. My recommendation: send it back to the authors to fix the manuscript mismatch. Once the correct body is available, ask for a statistical significance test and an external validation before taking the 5% claim seriously.","headline":"The abstract describes a sensible longitudinal DCE-MRI segmentation method, but the submitted full text is a different paper on prosthetic hand grasping, so the claimed 5% Dice improvement is completely unverifiable as submitted.","tokens_in":605,"tokens_out":1406,"would_cite":false,"duration_ms":30056,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Prior scans and BI-RADS scores lift lesion segmentation by 5%","keywords":["breast cancer screening","DCE-MRI","lesion segmentation","temporal prior attention","BI-RADS score","longitudinal imaging","3D segmentation","clinical context"],"falsifier":"Re-run the same training and evaluation protocol on a public longitudinal DCE-MRI dataset, but with the BI-RADS labels randomly shuffled during training; if the shuffled-label model still beats the single-timepoint baseline by roughly 5% Dice, the reported gain is not caused by clinical meaning. A second check is to re-score the validation scans with an independent radiologist and confirm the gain persists when only cases with inter-reader agreement are kept.","tokens_in":4222,"feed_emoji":"🧲","tokens_out":6079,"duration_ms":51671,"temperature":0.7,"pith_summary":"LesiOnTime claims that the accuracy ceiling for segmenting small breast lesions in DCE-MRI is set not by image resolution but by the model's ability to use the two cues radiologists rely on: the patient's previous scans and the BI-RADS assessment recorded at each timepoint. The paper introduces a 3D segmentation network with a Temporal Prior Attention block that lets current-scan features query prior-scan features, and a BI-RADS Consistency Regularization loss that pulls latent representations of scans sharing the same BI-RADS category closer together. On an in-house longitudinal dataset of high-risk patients, this combination is reported to outperform single-timepoint and other longitudinal baselines by 5% Dice, with ablations showing both components contribute. The intended upshot is that routine clinical metadata, which is already collected in screening, can be converted into a training signal that improves early lesion detection.","feed_headline":"Prior scans and BI-RADS scores lift lesion segmentation by 5%","feed_subtitle":"Adding earlier scans and BI-RADS alignment to a 3D network improves early breast lesion detection.","key_machinery":"The two load-bearing mechanisms are the Temporal Prior Attention (TPA) block and the BI-RADS Consistency Regularization (BCR) loss. TPA is an attention module that takes features from the current DCE-MRI scan and from earlier scans and computes a dynamic integration of the two, so the network can decide what prior information matters for the current lesion boundary. BCR is an auxiliary loss that pushes latent feature vectors of scans with the same BI-RADS score toward each other during training, injecting the clinical scoring system as a geometric constraint on the representation. The paper's argument is that these two mechanisms, not a larger backbone or more training data, are what close the gap for small lesions.","core_discovery":"The central discovery is that temporal and clinical context are not just useful heuristics but can be made into differentiable network components that directly improve small-lesion segmentation. The Temporal Prior Attention block dynamically weights and integrates features from earlier scans when segmenting the current one, mimicking the radiologist's side-by-side comparison. The BI-RADS Consistency Regularization loss treats the radiological assessment as a structured label: scans with the same BI-RADS category should occupy nearby regions of the learned latent space. Together, the two mechanisms are credited with a 5% Dice gain over state-of-the-art single-timepoint and longitudinal baselines on the curated in-house dataset, and the ablations attribute complementary gains to each term.","pith_inferences":["A natural test would be to shuffle the BI-RADS labels during training; if the BCR loss with random labels gives a similar Dice gain, the reported improvement is coming from a generic regularization effect, not from clinical meaning.","Because the dataset is in-house, the 5% figure is not yet benchmarked against public longitudinal breast MRI datasets; reproducing it there would separate a method-level advantage from dataset-specific cues.","If inter-reader variability of BI-RADS is a concern, weighting the consistency loss by radiologist confidence or adjudicated labels might make the gain more robust; the paper does not explore this.","The TPA block's dynamic integration could also be read as a learned attention that highlights changes between timepoints; a visualization study of the attention maps might reveal whether the network is detecting growth or new enhancement, which radiologists could audit."],"forward_implications":["If the 5% Dice gain reproduces, screening systems can improve segmentation of subtle lesions without any new annotation burden, because prior scans and BI-RADS scores are already recorded in routine follow-up.","The TPA block suggests a general recipe: temporal fusion can be posed as attention between timepoints and trained end-to-end, which could extend to other longitudinal imaging tasks such as tumor growth tracking or multiple-sclerosis lesion monitoring.","The BCR loss offers a template for injecting structured clinical labels into segmentation training; any ordinal or categorical scoring system with known clinical meaning could be used the same way.","Ablations showing complementary gains imply that a clinician-informed training signal and a temporal-fusion mechanism address different sources of error; combining them may be better than either alone."],"supporting_citations":[],"fun_headline_variants":["Temporal and BI-RADS context boost small lesion segmentation by 5%","Adding prior scans and BI-RADS scores lifts lesion Dice by 5%","Radiologist-like context improves breast lesion segmentation 5%","Joint temporal + clinical cues yield 5% Dice gain in lesion MRI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's gain depends on BI-RADS scores being reliable and consistently assigned, because the BCR loss treats them as trustworthy grouping labels; if the scores are noisy or radiologists disagree, the alignment could teach the network spurious structure and the 5% Dice improvement could vanish.","fun_headline_variants_meta":{"raw":{"variants":["Temporal and BI-RADS context boost small lesion segmentation by 5%","Adding prior scans and BI-RADS scores lifts lesion Dice by 5%","Radiologist-like context improves breast lesion segmentation 5%","Joint temporal + clinical cues yield 5% Dice gain in lesion MRI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1514,"prompt_tokens":949,"completion_tokens":565,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":486}},"tokens_in":565,"tokens_out":565,"duration_ms":4831,"temperature":1.0,"reasoning_tokens":486,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:06:25.417332+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same training and evaluation protocol on a public longitudinal DCE-MRI dataset, but with the BI-RADS labels randomly shuffled during training; if the shuffled-label model still beats the single-timepoint baseline by roughly 5% Dice, the reported gain is not caused by clinical meaning. A second check is to re-score the validation scans with an independent radiologist and confirm the gain persists when only cases with inter-reader agreement are kept.","supporting_citations":[],"review_version":1}