{"id":"40196db9-55b5-4bba-8a3a-24a72ffa8a1f","arxiv_id":"2507.15292","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A training-free Lagrangian motion magnification framework with periodic reference resetting and tissue-aware dual-mask control improves vascular pulsation visibility in endoscopic surgery videos.","lead":"This paper presents EndoControlMag, a training-free system that magnifies subtle vascular motion in endoscopic surgery videos by tracking vessels and selectively amplifying their pulsation. The authors report that it outperforms existing motion magnification methods on a new 24-video dataset spanning four surgical procedures.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Accuracy metrics use the same RAFT flow estimator the method is built on; the claimed magnification-accuracy advantage may be a self-consistency artifact.","rationale":"The paper is a competent engineering contribution with a new dataset, thorough ablations, and an acknowledged tracker limitation in Sec. 6.2. The reader's chosen weakest assumption (MFT tracker loss) is real but disclosed and scoped to extreme occlusions/motions; it does not invalidate the reported comparisons if those extremes are absent or handled by reinitialization. The more fundamental problem is that the accuracy metrics (Emotion, Emag) use RAFT, the same optical flow network that is embedded in the synthesis objective (via FlowMag) and in the Motion variant's weighting. Since both EndoControlMag and FlowMag are optimized against RAFT, the accuracy tables measure self-consistency with RAFT, not independently verified motion fidelity. The visual-quality metrics (SSIM, PSNR, MUSIQ) and the surgeon study are more independent, but they do not support the 'magnification accuracy' half of the abstract claim. The proposed test---recomputing accuracy with an independent flow estimator or synthetic ground truth---is feasible given the promised code/data and would settle the issue. Until then, the CONDITIONAL verdict remains appropriate; no change to the reader's verdict is needed.","tokens_in":25036,"tokens_out":8998,"duration_ms":98993,"concrete_test":"Recompute Emotion and Emag for Tables 3 and the Hard Set comparisons using a different optical flow estimator not used by any method (e.g., GMFlow), keeping all magnification outputs fixed. If EndoControlMag's margin over FlowMag shrinks or reverses, the accuracy advantage is an artifact of using RAFT for both synthesis and evaluation. Additionally, build synthetic videos with known ground-truth vessel motion and evaluate both methods against the ground truth; if the ranking changes, the RAFT-based metrics do not support the headline accuracy claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The quantitative claim of superior magnification accuracy rests on Emotion and Emag (Eqs. 11-12), which are computed with RAFT. However, RAFT is not an independent arbiter: EndoControlMag's synthesis path is flow-based, using FlowMag as fVMM, and FlowMag optimizes the magnified frame so that RAFT flow from the PRR reference to the magnified frame matches alpha times the original RAFT flow inside the mask. The Motion variant also derives outer-mask weights from RAFT flow (Eqs. 7-8). Thus the evaluation metric is largely the objective the method optimizes, so low Emotion/Emag can reflect self-consistency rather than true motion fidelity. The reported gains (e.g., Emotion 1.151->0.999, Emag 1.732->1.033 at x2) may come from PRR's shorter reference intervals making the flow-optimization problem easier. Because 'significantly outperforms in magnification accuracy' is half the central claim and this is the only quantitative support for it, the accuracy claim is not independently verified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes EndoControlMag, a training-free Lagrangian framework for mask-conditioned vascular motion magnification in endoscopic surgery. It introduces Periodic Reference Resetting (PRR) to bound optical-flow error accumulation by slicing videos into short overlapping clips, and a Hierarchical Tissue-aware Magnification (HTM) module that recursively tracks a vessel mask and applies either motion-based or distance-based softening to the surrounding tissue. The method is evaluated on a new EndoVMM24 dataset spanning four surgery types and four challenge categories, using image-quality metrics (SSIM, PSNR, MUSIQ), magnification-accuracy metrics (Emotion, Emag), a surgeon scoring study, and ablations.","tokens_in":25259,"tokens_out":3743,"duration_ms":41632,"significance":"If validated, the framework addresses a real clinical need: making subtle vascular pulsations visible during endoscopic procedures while preserving the surrounding tissue. The manuscript is clearly written, the algorithm is fully specified in pseudocode, and the authors commit to releasing code, data, and video results, which is commendable and will aid reproducibility. The use of off-the-shelf tracking and flow models without fine-tuning is a practical strength, and the ablation study on mask dilation strategies is informative. The main scientific value lies in demonstrating that reference resetting and adaptive mask softening improve robustness in dynamic surgical scenes; however, the central quantitative claim of superior magnification accuracy currently rests on metrics that are computed with the same optical flow estimator used inside the method, which creates a self-consistency risk that must be resolved before the accuracy claim can be accepted as independent evidence.","major_comments":[{"comment":"The magnification-accuracy metrics Emotion and Emag are computed with RAFT, which is the same optical flow estimator used by the base FlowMag model inside fVMM and also by the motion-based softening in Eq. (7). Since the magnified frame is synthesized so that RAFT flow from the reference to the magnified frame approximates alpha times the reference-to-current RAFT flow inside the mask, low Emotion/Emag partly reflect that the method directly optimizes what the metric measures. The reported reductions (e.g., Emotion 1.151→0.999, Emag 1.732→1.033 at ×2 in Table 3) may therefore be self-consistency artifacts rather than evidence of true motion fidelity. Please re-evaluate with an independent flow estimator (e.g., a different architecture such as GMFlow or a learned dense tracking method) or on synthetic sequences with known ground-truth motion, so that the accuracy claim is not circular.","section":"§4.2.2, Eqs. (11)-(12), Tables 3 and 5"},{"comment":"The Hard Set evaluation compares EndoControlMag only against FlowMag, not against the other baselines (EVM, DMM, MDL-VMM, STB-VMM, Axial-VMM) that were included on the Easy Set. The abstract and introduction claim that the method 'significantly outperforms existing methods' and maintains 'robustness across challenging surgical conditions'; however, the evidence for robustness in occlusion, view change, vessel deformation, and tool disturbance is based solely on comparison with a single baseline. Please include at least the strongest learning-based baselines on the Hard Set, or explicitly justify their exclusion beyond stating that global methods would produce 'intolerable visual distortions.'","section":"§5.2, Fig. 5"},{"comment":"The phrase 'significantly outperforms' is used repeatedly, but the quantitative tables (Table 2, Table 3) and Fig. 5 report only means and standard deviations with no statistical tests. With only 4 videos in the Easy Set and 3-8 clips per Hard Set category, the observed differences (e.g., SSIM 0.97 vs 0.96) may not be statistically reliable. Provide paired significance tests (e.g., paired t-test or Wilcoxon signed-rank test) across videos, or temper the language to 'consistently outperforms' without the statistical connotation.","section":"§5.1, §5.2, Abstract"},{"comment":"The surgeon evaluation is based on only 3 surgeons, 8 videos, and a binary per-item rubric. The reported score gaps (9.75 vs 8.19 on Easy Set; 9.19/9.22 vs 7.61 on Hard Set) are presented without any measure of inter-rater agreement or significance. Given the small number of raters and items, the standard error of these averages is large, and the claim of 'strong agreement' is not supported by the data shown. Please report per-item agreement statistics (e.g., Fleiss' kappa or a simple percentage agreement) and a statistical comparison (e.g., a permutation test) or explicitly state the pilot nature of the study.","section":"§5.3, Fig. 6"}],"minor_comments":[{"comment":"The PSNR column is labeled 'PNSR' in the header; please correct the typo.","section":"Table 2 header"},{"comment":"The Emotion value at ×16 is printed as '15.152±8.394' and at ×32 as '31.203±17793', which appears to be missing a decimal point (likely 31.203±17.793). Please check all entries for formatting consistency.","section":"Table 5, Vessel-adaptive row"},{"comment":"The link between N=4 and cardiac pulsation frequency is a post-hoc justification; it is not tested independently of the ablation. Consider framing it as a hypothesis rather than a confirmed physiological explanation.","section":"§5.4.1, 'Physiological Relevance' paragraph"},{"comment":"The limitations section honestly acknowledges that the MFT tracker can lose the vessel during occlusions longer than about 2 seconds or very fast camera movements, requiring manual re-initialization. This is a useful caveat, but it should be reflected in the abstract and conclusion, where the claim of 'robustness across challenging surgical conditions' is currently stated without this qualification.","section":"§6.2"},{"comment":"The term 'training-free' may be misleading because the framework uses pretrained models (RAFT, MFT, and the FlowMag base model). The paper clarifies in §6.2 that no fine-tuning on surgical data is performed; consider defining 'training-free' in the introduction to avoid confusion.","section":"§3.1.2"},{"comment":"The description of the scoring scheme is confusing: a 10-item binary (Yes/No) rubric is used, but the text says each item is scored per video and per magnification factor, and the final score is an average. Please clarify how the 0-10 scale is constructed so the reader can interpret the scores (e.g., is it the average number of 'Yes' answers across all 30 item-video-factor combinations?).","section":"Fig. 6(a)"}],"recommendation":"major_revision","confidential_remarks":"The central methodological idea (periodic reference resetting plus dual-mask control) is coherent and the paper is well written, but the evaluation strategy has a load-bearing circularity issue: the accuracy metrics use the same RAFT flow estimator that the method itself relies on. In my view this is not a fatal flaw—it can be addressed by re-evaluating with an independent flow estimator or synthetic ground truth—but it does mean the current quantitative evidence does not support the strong 'significantly outperforms' claim. I also recommend that the authors expand the Hard Set comparison beyond FlowMag and add basic statistical significance testing. The paper fits the scope of eess.IV and, once the evaluation concerns are addressed, could be a solid contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the take: this is a solid, well-organized engineering paper that extends FlowMag with two sensible modules—periodic reference resetting (PRR) and hierarchical tissue-aware masking (HTM)—plus a new 24-clip endoscopic dataset. If you work on surgical vision or motion magnification, it's a useful reference. The writing is clear, the method is described precisely, and the limitations section is honest about tracker failures, 2s/frame processing, and manual mode selection.\n\nThe main soft spot is exactly what the stress-test note flags: the magnification-accuracy metrics (Emotion, Emag) are computed with RAFT, the same flow estimator used in the synthesis objective. FlowMag, which EndoControlMag builds on, optimizes the magnified frame so that RAFT flow from the reference to the magnified frame matches alpha times the original RAFT flow inside the mask. So low Emotion/Emag largely measure how well the method satisfies its own optimization, not true motion fidelity against ground truth. The reported gains over FlowMag may simply reflect that PRR makes the flow-optimization problem easier. This is genuine circularity that weakens the strongest quantitative claim.\n\nThat said, the circularity doesn't sink the paper. The image-quality metrics (SSIM, PSNR, MUSIQ) are independent of flow, and EndoControlMag beats FlowMag and other baselines with modest but consistent margins. The surgeon evaluation, though small (8 clips, 3 raters, binary rubric), also favors it strongly. So the visual-quality claim is credible; only the magnification-accuracy claim is unverified.\n\nMinor issues: the Hard Set comparison is only against FlowMag, so the robustness advantage over other methods is asserted but not shown. The clip-length ablation (N=4) shows tiny differences. The dataset release is promised but not yet verified.\n\nBottom line: this deserves serious peer review. Ask the authors to evaluate accuracy with an independent flow estimator or synthetic ground-truth motion, report confidence intervals for the surgeon study, and compare more baselines on the Hard Set. For a reading group, it's a good example of how evaluation metrics can be circular in video-motion-magnification research.","headline":"A competent, well-written engineering extension of FlowMag for endoscopic vascular motion magnification, but its central accuracy claim rests on a circular metric and needs independent verification.","tokens_in":25766,"tokens_out":3999,"would_cite":true,"duration_ms":41273,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Periodic reference resets and recursive mask tracking make surgical pulse magnification robust.","keywords":["vascular motion magnification","endoscopic surgery","Lagrangian motion magnification","mask-conditioned video editing","periodic reference resetting","hierarchical tissue-aware magnification","video object tracking","EndoVMM24"],"falsifier":"Take a clip in which a vessel is fully hidden behind an instrument for about three seconds and then reappears; annotate its true mask in the frames before and after. If EndoControlMag's output after reappearance shows the magnified pulsation centered more than one vessel radius away from the annotated vessel, or if the magnification map is visibly frozen at the pre-occlusion position, then the tracking assumption fails in precisely the occlusion regime the paper claims to handle.","tokens_in":24865,"feed_emoji":"🩺","tokens_out":6228,"duration_ms":64438,"temperature":0.7,"pith_summary":"EndoControlMag aims to make the faint pulsing of blood vessels visible in endoscopic surgical video without training a new network. The paper claims that two additions to mask-conditioned Lagrangian motion magnification achieve this: periodically resetting the reference frame so optical-flow errors cannot accumulate, and recursively tracking the vessel mask while softening the magnification strength into surrounding tissue. This matters because surgeons must infer vascular location and perfusion from subtle motion cues amid smoke, tool occlusion, camera shifts, and tissue deformation, conditions where existing global or static-mask methods drift or produce boundary artifacts. On the authors' four-surgery EndoVMM24 benchmark, both quantitative metrics and expert surgeon ratings favor EndoControlMag over prior methods.","feed_headline":"Training-free method tracks vessel pulses through smoke and occlusion","feed_subtitle":"Periodic reference resets and a tracked mask keep tiny vascular pulsations visible in endoscopic video.","key_machinery":"The load-bearing mechanism is the combination of two modules. Periodic Reference Resetting (PRR) splits a video into overlapping clips of length $N=4$, using the first frame of each clip as the reference for Lagrangian magnification, so the temporal distance from any frame to its reference is bounded and cumulative optical-flow error stays $O(N)$ instead of $O(T)$. Hierarchical Tissue-aware Magnification (HTM) builds the magnification map $M_t$: 1 inside the tracked vessel mask, a softened weight $W_t$ in the dilated outer region, and 0 outside; the outer weight is either normalized flow magnitude or $\\exp(-\\beta \\cdot d)$ from the vessel boundary. The recursive inner-mask tracking keeps the magnified region aligned to the moving vessel, while the softened outer ring prevents hard boundaries between magnified and unmagnified tissue.","core_discovery":"The central claim is that EndoControlMag robustly outperforms existing video motion magnification methods for endoscopic vascular visualization. It is a training-free Lagrangian pipeline: optical flow from a periodically reset reference frame is scaled by a factor $\\alpha$ and backward-warped, with the scale modulated by a spatially varying magnification mask. The mask comes from a hierarchy: an inner binary mask tracks the target vessel through a pretrained long-term point tracker, and an outer ring around the vessel is softened either by normalized optical-flow magnitude (motion-based) or by exponential decay with distance (distance-based). The paper reports that this design yields lower motion and magnification errors, higher SSIM/PSNR/MUSIQ scores, and better surgeon ratings than FlowMag and other baselines, especially under occlusions, view changes, vessel deformation, and tool disturbance.","pith_inferences":["One step the authors leave implicit is automating the choice between motion-based and distance-based softening by monitoring optical-flow confidence or smoke indicators, which would remove the manual pre-selection and make the system closer to plug-and-play in the operating room.","A testable extension is to read physiological parameters such as pulse rate and relative pulse amplitude from the magnified pulsation signal; if the amplification preserves phase, the enhanced video could serve as a non-contact hemodynamic estimate rather than only a visual aid.","The same PRR-plus-tracked-mask recipe should transfer to other deformable, occluded scenes, such as laparoscopic assessment of bowel motility or fetal ultrasound, where subtle periodic motion matters and static references fail.","If tracker reliability is the bottleneck, replacing the current tracking model with one adapted to surgical appearance would likely improve robustness more than any other component of the pipeline."],"forward_implications":["If the claim holds, a surgeon can select a vessel once and get a pulsing, magnified view of it that stays aligned as the camera moves and instruments pass in front of it, without needing retraining on surgical data.","The PRR clip length $N=4$ implies that error growth is bounded within a short window, so the method should degrade gracefully on long procedures where a fixed-reference baseline would drift.","The dual softening modes give users a choice: motion-based for deformation-heavy scenes, and distance-based for smoke or instrument occlusion where optical flow is unreliable.","The EndoVMM24 dataset provides a reusable benchmark with four surgery types and four challenge categories, so future methods can be compared on the same hard conditions.","Because the method is training-free, it can be combined with improved off-the-shelf optical flow or tracking models as those improve, without changing the framework."],"supporting_citations":[{"why":"Provides the base mask-conditioned Lagrangian magnification model (FlowMag) that EndoControlMag extends and the primary baseline it must beat.","marker":"[31]"},{"why":"Supplies the pretrained long-term point tracker used to propagate the inner vessel mask through occlusions and view changes.","marker":"[29]"},{"why":"Supplies the RAFT optical flow used for motion estimation, motion-based softening, and the two accuracy metrics.","marker":"[40]"},{"why":"Source of the laparoscopic cholecystectomy clips in the EndoVMM24 dataset.","marker":"[33]"},{"why":"Source of the robot-assisted radical prostatectomy clips in the EndoVMM24 dataset.","marker":"[2]"},{"why":"Source of the gastric bypass clips in the EndoVMM24 dataset.","marker":"[32]"},{"why":"Classic Eulerian video magnification baseline compared in the Easy Set evaluation.","marker":"[49]"},{"why":"Learning-based motion magnification baseline compared in the Easy Set evaluation.","marker":"[30]"}],"fun_headline_variants":["Training-free method amplifies vessel pulses despite occlusion","Periodic resets keep endoscopic vessel motion visible","Tracker-guided mask magnifies vascular motion in surgery","Robust vessel pulse magnification without training or tuning","EndoControlMag exposes subtle vessel motion in messy scenes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The design depends on the vessel mask tracker staying locked onto the target while instruments, smoke, and camera motion pass through the scene; if the tracker loses the vessel for longer than about two seconds or during a very fast camera move, the magnified region no longer points at the right structure and a surgeon must reinitialize the mask by hand.","fun_headline_variants_meta":{"raw":{"variants":["Training-free method amplifies vessel pulses despite occlusion","Periodic resets keep endoscopic vessel motion visible","Tracker-guided mask magnifies vascular motion in surgery","Robust vessel pulse magnification without training or tuning","EndoControlMag exposes subtle vessel motion in messy scenes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000171,"raw_usage":{"total_tokens":1299,"prompt_tokens":999,"completion_tokens":300,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":227}},"tokens_in":615,"tokens_out":300,"duration_ms":4746,"temperature":1.0,"reasoning_tokens":227,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:34:50.970938+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a clip in which a vessel is fully hidden behind an instrument for about three seconds and then reappears; annotate its true mask in the frames before and after. If EndoControlMag's output after reappearance shows the magnified pulsation centered more than one vessel radius away from the annotated vessel, or if the magnification map is visibly frozen at the pre-occlusion position, then the tracking assumption fails in precisely the occlusion regime the paper claims to handle.","supporting_citations":[{"cited_title":"Self-supervised motion magnifica- tion by backpropagating through optical flow","cited_arxiv_id":null,"evidence_quote":"Provides the base mask-conditioned Lagrangian magnification model (FlowMag) that EndoControlMag extends and the primary baseline it must beat."},{"cited_title":"Mft: Long-term tracking of every pixel, in: Proceedings of the IEEE /CVF Winter Conference on Applica- tions of Computer Vision, pp","cited_arxiv_id":null,"evidence_quote":"Supplies the pretrained long-term point tracker used to propagate the inner vessel mask through occlusions and view changes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the RAFT optical flow used for motion estimation, motion-based softening, and the two accuracy metrics."},{"cited_title":"Learning-based video motion magnification, in: Pro- ceedings of the European Conference on Computer Vision (ECCV), pp","cited_arxiv_id":null,"evidence_quote":"Source of the laparoscopic cholecystectomy clips in the EndoVMM24 dataset."},{"cited_title":"Weakly supervised tempo- ral convolutional networks for fine-grained surgical activity recognition","cited_arxiv_id":null,"evidence_quote":"Source of the gastric bypass clips in the EndoVMM24 dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Classic Eulerian video magnification baseline compared in the Easy Set evaluation."},{"cited_title":"Motion magnification for endoscopic surgery, in: Medical Imaging 2014: Image- Guided Procedures, Robotic Interventions, and Modeling, SPIE","cited_arxiv_id":null,"evidence_quote":"Learning-based motion magnification baseline compared in the Easy Set evaluation."}],"review_version":1}