{"id":"25a6b17a-2453-4560-ba8d-76fdd98ed058","arxiv_id":"2508.05262","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"A cyclic-consistency particle filter tracks 117 cardiac landmarks in real time with ~5 px error, beating deep learning and conventional trackers in fluorescent imaging.","lead":"Researchers built a tracking algorithm that follows 117 points on live heart videos during surgery, achieving about 5 pixels of error at real-time speed. It is designed to give surgeons immediate feedback on whether a bypass graft is perfusing the heart properly.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed 5.00±0.22 px tracking error cannot be verified because the evaluation protocol is absent and the full text is undecodable; a check of ground-truth independence and baseline fairness should precede acceptance.","rationale":"The reader's weakest assumption identifies precisely the load-bearing point: the tracking error must be measured against independent ground truth and compared against fair baselines. My stress-test agrees and makes the concern more specific by linking it to the cyclic-consistency mechanism: if the error is computed only on accepted tracks, the reported 5.00 ± 0.22 px could be biased. Because the full text is undecodable, this cannot be checked from the manuscript as submitted. The correct verdict remains UNVERDICTED, matching the reader's assessment. The proposed check is concrete and would settle whether the concern lands.","tokens_in":28658,"tokens_out":3284,"duration_ms":34949,"concrete_test":"Obtain an uncorrupted arXiv source or PDF. In the experiments section, verify: (a) ground-truth landmarks were obtained by independent manual annotation or an independent method in the same coordinate frame; (b) the reported 5.00 ± 0.22 px error is computed over all 117 initialized targets across all frames, with the number of targets rejected by cyclic-consistency checks disclosed; and (c) each deep/conventional baseline was run with recommended hyperparameters on identical test sequences. If any of these is absent, or if the error is computed only on cyclic-consistency-accepted tracks, the headline comparison should be re-evaluated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the empirical superiority reported in the abstract: 5.00 ± 0.22 px tracking error versus 22.3 ± 1.1 px for deep trackers and 58.1 ± 27.1 px for conventional trackers, at 25.4 fps over 117 targets. For this claim to be valid, three conditions must hold: (i) tracking error is measured against independent ground-truth landmark positions, not positions derived from the tracker's own cyclic-consistency output; (ii) the error is averaged over all tracked targets, not only the subset that passes cyclic-consistency checks; and (iii) the deep and conventional baselines are standard, appropriately tuned implementations evaluated on identical sequences. The available text is an encoding-corrupted version of the paper, so none of these conditions can be checked. If (ii) fails, the 5.00 px figure could reflect a self-selected subset of easy targets; if (i) fails, the error metric is partly circular. The claim is therefore unverified rather than refuted, but it is the load-bearing empirical assertion of the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a particle-filter-based tracker with cyclic-consistency checks for measuring cardiac perfusion in intraoperative fluorescent cardiac imaging. The abstract claims the method tracks 117 targets simultaneously at 25.4 fps with a tracking error of 5.00 ± 0.22 px, outperforming deep learning trackers (22.3 ± 1.1 px) and conventional trackers (58.1 ± 27.1 px). The body text supplied for review is almost entirely non-decodable replacement characters; no readable methods, mathematical formulation, dataset description, experimental protocol, or baseline implementation details are available. As a consequence, the assessment can be based on the abstract and document structure only.","tokens_in":28964,"tokens_out":2615,"duration_ms":32638,"significance":"If the reported results were fully supported, the paper would offer a practically valuable real-time multi-target tracking solution for intraoperative imaging: 117 targets at 25.4 fps with roughly 5 px error is a strong operational claim, and a non-learned particle-filter alternative that beats deep trackers by a wide margin would be notable. However, the manuscript as submitted does not permit verification of any of these claims. There are no machine-checked proofs, no reproducible code reference, no dataset description, no ground-truth specification, and no identifiable baseline implementations. The paper therefore currently contributes an abstract-level claim rather than a checkable scientific contribution.","major_comments":[{"comment":"The body of the manuscript is not readable: it consists almost entirely of replacement characters, with only fragments such as a repeated 'The ...' and an unrelated arXiv header 'arXiv:2508.05233v3 [astro-ph.SR]' appearing amid garbled text. No algorithm, no cyclic-consistency definition, no particle-filter equations, no dataset description, and no experimental setup can be examined. Because the central claim is an empirical performance comparison, the absence of readable methods and results is a load-bearing defect that cannot be repaired by inference from the abstract.","section":"Full text / all body sections"},{"comment":"The abstract reports precise errors (5.00 ± 0.22 px, 22.3 ± 1.1 px, 58.1 ± 27.1 px) but provides no dataset size, number of sequences, number of frames, ground-truth generation procedure, or list of compared trackers. It is also unclear whether the 5.00 px error is averaged over all 117 targets or only over targets that passed the cyclic-consistency check; if the latter, the error could reflect an easy self-selected subset. To support the central claim, the authors must specify the evaluation protocol, including the target-selection rule and per-target success/failure rates.","section":"Abstract, quantitative claims"},{"comment":"The claimed superiority over 'deep learning trackers' and 'conventional trackers' cannot be interpreted without knowing which specific baselines were used, whether they were standard implementations, and whether they were run on the identical sequences with tuned hyperparameters. The large standard deviation in the conventional-tracker error (27.1 px) further suggests heterogeneous conditions that need explanation. In addition, the paper must show that the ground-truth landmark positions are independent of the tracker's own cyclic-consistency outputs; otherwise the error metric is at risk of circularity. No such evidence is present in the readable portion.","section":"Abstract and missing Methods"}],"minor_comments":[{"comment":"Typographical issue: 'cyclicconsistency checks' should be 'cyclic-consistency checks.'","section":"Abstract"},{"comment":"The body contains an unrelated arXiv identifier 'arXiv:2508.05233v3 [astro-ph.SR]' and repeated garbled boilerplate. If this is a submission artifact, the authors should ensure the submitted PDF text layer is intact.","section":"Document structure"},{"comment":"No references to datasets, prior fluorescent-imaging tracking work, or baseline implementations can be identified from the readable text; a complete reference list is needed.","section":"References"}],"recommendation":"reject","confidential_remarks":"The key problem is that the submitted file is undecodable beyond the abstract. If the unreadable body is the result of an encoding or PDF-extraction failure, the authors could resubmit a properly rendered version, and I would be willing to review that version. Based on the current submission, no technical content can be evaluated, and the empirical claims in the abstract are unaudited. I see no evidence of misconduct, but the manuscript in its present form cannot be accepted or meaningfully revised within its own scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: as submitted, this one is unverifiable. The full text came through as mojibake, so the only content I can engage with is the abstract. The reader's UNVERDICTED verdict is the right call, and the stress-test note correctly identifies what would have to be checked.\n\nWhat the paper does well on the face of it: the clinical task is real. After bypass grafting, surgeons want an immediate read on perfusion, and tracking fluorescent landmarks under heart motion is a legitimate engineering target. Claiming 117 simultaneous targets at 25.4 fps with a particle filter plus cyclic-consistency checks is concrete and testable. The numbers are specific, which is more than many abstracts offer.\n\nWhere it gets soft: I cannot check the three things that matter. (1) How was ground truth obtained? If it is independent of the tracker's cyclic-consistency output, fine; if not, the 5.00 px error is partly circular. (2) Is the error averaged over all 117 targets, or over the subset that passes consistency? The latter could flatter the result. (3) Were the deep and conventional baselines tuned fairly? 22.3 and 58.1 px are big gaps, and I'd want to see identical sequences, same target definitions, and some indication that the baselines weren't starved. None of these are accusations; they are the standard questions the abstract cannot answer. The stress-test's three conditions are exactly right.\n\nOne thing I'd push back on: the discrepancy with deep trackers is large, but not implausible for a domain-specific tracker with strong motion priors. So I wouldn't call the abstract's claim suspect on its face; it's just not testable from what we have.\n\nBottom line: this deserves a serious referee, but only if the authors provide a clean, readable version. I would not cite it yet, and I wouldn't put it in a reading group as-is. If the clean version confirms the protocol, this becomes a modest but useful application paper.","headline":"As submitted, the full text is unreadable, so the verdict hinges on the abstract; the claim is concrete and worth refereeing once a clean manuscript is available.","tokens_in":29371,"tokens_out":3000,"would_cite":false,"duration_ms":32885,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A particle-filter tracker with cyclic-consistency checks keeps 117 fluorescent cardiac landmarks in view at 25.4 fps with a reported mean tracking error of about 5 px, well below the deep-learning and conventional baselines it compares agai","keywords":["particle filter","cyclic consistency","fluorescent cardiac imaging","landmark tracking","intraoperative imaging","perfusion estimation","multi-target tracking","coronary bypass"],"falsifier":"Have an independent observer annotate landmark positions frame by frame on the same fluorescent cardiac sequences, then run this tracker alongside the deep-learning and conventional baselines with each method tuned on a validation split. If the mean per-target error on held-out sequences is far from 5 px, or if any baseline matches or beats the particle filter under equal tuning, the central superiority claim is contradicted.","tokens_in":28621,"feed_emoji":"🫀","tokens_out":5679,"duration_ms":60965,"temperature":0.7,"pith_summary":"This paper tries to establish that a classical particle-filter tracker, extended with a forward-backward cyclic-consistency check, can outperform both deep-learning and conventional trackers in intraoperative fluorescent cardiac imaging, where the heart moves and image brightness and vessel patterns fluctuate. It reports tracking 117 landmarks simultaneously at 25.4 fps with a tracking error of (5.00 +/- 0.22) px, compared with (22.3 +/- 1.1) px for deep-learning trackers and (58.1 +/- 27.1) px for conventional trackers. The proposed mechanism uses consistency checks to discard candidate matches that do not survive a forward-then-backward pass, which stabilizes multi-target tracking during surgery. If the reported numbers hold up, a lightweight, non-learned tracker is sufficient for real-time local perfusion estimates in this setting.","feed_headline":"Particle filter tracks 117 heart landmarks at 25 fps","feed_subtitle":"Cyclic-consistency checks cut tracking error to 5 px, beating deep-learning trackers in fluorescent cardiac imaging.","key_machinery":"The central object is a particle filter with cyclic-consistency checks. For every target landmark, a set of candidate positions (particles) is sampled; each candidate is propagated to the next frame and then matched back to its origin, and candidates whose forward-backward displacement is inconsistent are discarded before the surviving particles are resampled. This forward-backward validation is what keeps a large number of simultaneously tracked landmarks stable in the presence of heart motion and changing image appearance.","core_discovery":"The central claim is that a particle filter which samples candidate positions for each target landmark and rejects those that fail a cyclic-consistency check achieves a mean tracking error of (5.00 +/- 0.22) px while tracking 117 targets at 25.4 fps. On the fluorescent cardiac image sequences evaluated, this beats deep-learning trackers, which report (22.3 +/- 1.1) px, and conventional trackers, which report (58.1 +/- 27.1) px. The authors attribute the gain to the consistency check: a correct landmark track should look the same when followed forward and then backward, so inconsistent candidates are filtered out before they corrupt the state estimate. This makes the tracker stable under the","pith_inferences":["Editorial inference: the same forward-backward consistency idea could transfer to other intraoperative imaging modalities with periodic motion, such as laparoscopy or OCT, though the paper does not claim this.","Editorial inference: the reported dominance of a non-learned method suggests that combining a learned appearance score with the cyclic-consistency check might reduce drift even further, at some computational cost.","Editorial inference: the speed and error figures should be re-tested on a public or independently annotated dataset before being used as a general benchmark, since the abstract does not describe the ground-truth protocol or the baseline hyperparameters."],"forward_implications":["Real-time local perfusion estimates during bypass surgery become feasible, because 117 landmarks can be tracked at 25.4 frames per second.","The cyclic-consistency check provides drift rejection without a learned appearance model, so the approach stays lightweight and interpretable.","If the comparison holds, classical trackers augmented with consistency checks can compete with deep-learning trackers in this specific imaging regime.","Simultaneous tracking of many landmarks supports dense local quantitative maps rather than single-point perfusion measurements."],"supporting_citations":[],"fun_headline_variants":["Particle filter tracks 117 heart landmarks at 25 fps","Cyclic-consistency check cuts heart tracking error to 5 px","Real-time cardiac tracking: particle filter beats deep learning","117 heart landmarks tracked at 25 fps with 5 px error","Heart tracker: 5 px error, outshines deep learning"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The reported advantage depends on ground-truth landmark positions being accurate and independent, and on the deep-learning and conventional baselines being fairly tuned; if either condition fails, the 5-pixel error claim may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Particle filter tracks 117 heart landmarks at 25 fps","Cyclic-consistency check cuts heart tracking error to 5 px","Real-time cardiac tracking: particle filter beats deep learning","117 heart landmarks tracked at 25 fps with 5 px error","Heart tracker: 5 px error, outshines deep learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000647,"raw_usage":{"total_tokens":2764,"prompt_tokens":656,"completion_tokens":2108,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":400,"completion_tokens_details":{"reasoning_tokens":2020}},"tokens_in":400,"tokens_out":2108,"duration_ms":17322,"temperature":1.0,"reasoning_tokens":2020,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:26:02.679930+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an independent observer annotate landmark positions frame by frame on the same fluorescent cardiac sequences, then run this tracker alongside the deep-learning and conventional baselines with each method tuned on a validation split. If the mean per-target error on held-out sequences is far from 5 px, or if any baseline matches or beats the particle filter under equal tuning, the central superiority claim is contradicted.","supporting_citations":[],"review_version":1}