{"id":"4d6498fe-b696-45af-a42d-1cb020e83757","arxiv_id":"2607.12939","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"STIRC2025 benchmarks seven surgical point-tracking algorithms on the STIR infrared-tattoo dataset for accuracy (in vivo/ex vivo) and inference latency.","lead":"This paper reports the 2025 Surgical Tattoos in Infrared Challenge (STIRC2025), a MICCAI EndoVis benchmark that scored seven teams on surgical point-tracking accuracy and inference latency using the STIR dataset. It matters because reliable tissue point tracking is a shared bottleneck for surgical segmentation, 3D reconstruction, landmarking, and autonomy.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the abstract-only incompleteness already noted by the reader.","rationale":"The reader correctly treats this as an abstract-only challenge summary whose contribution is community infrastructure rather than a novel scientific result, and correctly flags that method details, tables, and statistics cannot be verified. The weakest_assumption about clinical proxy validity is real as framing but is not required for the strongest claim (that a dual accuracy/latency evaluation of seven teams was run and artifacts released). Because the full text is unavailable, manufacturing a deeper technical objection would violate the good-faith rule. The verdict therefore stays UNVERDICTED with low confidence; the concrete test simply operationalizes the verification that becomes possible once the paper and ranking table are public.","tokens_in":2045,"tokens_out":371,"duration_ms":3981,"concrete_test":"When the full paper appears, verify that the accuracy metric (e.g., endpoint error or survival rate on STIR ground-truth tattoos) and latency protocol are defined identically for all seven teams and that the public STIRMetrics repo reproduces the published ranking table within floating-point tolerance; any non-reproducible entry would undermine the infrastructure claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper is a challenge report whose central claim is infrastructural: seven algorithms were evaluated for accuracy (in vivo/ex vivo STIR sequences) and inference latency, with public dataset and metrics code. That claim is self-contained and does not require the tattoos or metrics to be a perfect clinical proxy; the proxy question is framing, not a load-bearing premise of the reported bake-off. With only the abstract available, no internal inconsistency, ranking artifact, or metric definition error can be checked, so no new load-bearing concern is identifiable.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This manuscript reports the 2025 Surgical Tattoos in Infrared Challenge (STIRC2025), held within MICCAI EndoVis 2025. Participants submit point-tracking algorithms that are scored on the public STIR (surgical tattoos in infrared) dataset along two quantitative axes: tracking accuracy on in vivo and ex vivo sequences, and inference latency. Seven teams participated. The paper summarizes challenge results and participant methods, and releases the challenge dataset and baseline/metrics code.","tokens_in":2097,"tokens_out":484,"duration_ms":13823,"significance":"A public, dual-track (accuracy and efficiency) benchmark for surgical point tracking is a useful community resource, especially given the downstream tasks named in the abstract (segmentation, 3D reconstruction, virtual landmarking, autonomous probe scanning, subtask autonomy). Explicit release of the STIR challenge data and metrics/baseline code is a concrete reproducibility strength and supports follow-on comparison. The significance of any ranking conclusions depends on protocol, metric, and statistical details that are not assessable from the abstract alone.","major_comments":[],"minor_comments":[{"comment":"Abstract-only review: the full methods, evaluation protocol, result tables, and participant-method descriptions are not available, so ranking validity, metric definitions, and statistical treatment cannot be checked.","section":null},{"comment":"The abstract frames STIR tattoos and the accuracy/latency metrics as enabling several clinical downstream tasks; a short limitations discussion of how well infrared tattoos and the chosen metrics transfer to those tasks would strengthen the framing (not required for the bake-off claim itself).","section":null},{"comment":"Clarify whether this is a continuation of a prior STIR challenge iteration and, if so, what changed in 2025 (data splits, metrics, tracks) so readers can place the results.","section":null}],"recommendation":"uncertain","confidential_remarks":"Assessment is abstract-only; soundness of the reported bake-off cannot be verified without the full text (protocol, tables, statistics, organizer-vs-participant handling). On the abstract alone the design looks standard and the public data/metrics release is a plus. I would re-review if the full manuscript is supplied; until then I cannot responsibly choose accept/minor/major/reject."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a MICCAI EndoVis 2025 challenge summary on surgical point tracking with the STIR infrared-tattoo dataset. The one thing worth knowing is that it ships a clean community result: seven teams ranked on accuracy (in vivo and ex vivo) and inference latency, with the dataset and metrics code public. That is real infrastructure for endoscopic CV and surgical robotics, not a new clinical capability or a deep scientific claim.\n\nWhat is new is the 2025 bake-off itself—the rankings and the short method summaries—not the problem framing. Challenge reports of this kind are an established genre. The design looks standard and defensible: external labeled STIR data, separate accuracy and efficiency tracks, released baseline/metrics code. Circularity risk is low; evaluation is against a fixed public set, not author-fitted tautologies. Credit where due: public Zenodo data and GitHub metrics are the right artifacts, and dual scoring (accuracy plus latency) matches what people actually care about for deployment.\n\nSoft spots are mostly the usual ones for abstract-only challenge papers. We cannot check ranking tables, statistical tests, failure modes, or whether any organizer submissions or labeling choices tilt the field. The abstract’s framing that STIR tattoos and these metrics are a sufficient proxy for downstream tasks (segmentation, 3D recon, landmarking, autonomy) is motivational, not load-bearing for the reported numbers; if the proxy is imperfect, the rankings still stand as rankings on STIR. That is a minor framing issue, not a hole in the result.\n\nWho it is for: people building or benchmarking surgical trackers, and anyone who needs a current public reference on STIR. It is not for readers hunting a new algorithm or a theoretical advance. It deserves a serious referee—challenge reports with public data and multi-team evaluation are exactly what EndoVis exists to publish—even if the scientific novelty is modest. I would skim the methods section when the full text appears and keep the links; I would not reorganize a reading group around it unless someone is actively working on surgical tracking this quarter.","headline":"Standard EndoVis challenge report: useful STIR bake-off with public data and dual accuracy/latency scoring, not a foundational result.","tokens_in":2907,"tokens_out":515,"would_cite":false,"duration_ms":5301,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"STIRC2025 ranks seven surgical point trackers on infrared-tattoo accuracy and latency.","keywords":["point tracking","surgical computer vision","infrared tattoos","STIR dataset","challenge benchmark","MICCAI EndoVis","inference latency"],"falsifier":"Run the top-ranked STIRC2025 trackers on the same sequences without tattoos or with clinical landmarks, and check whether the accuracy ranking and latency numbers still hold under those conditions.","tokens_in":2899,"feed_emoji":"🔬","tokens_out":524,"duration_ms":3916,"temperature":0.7,"pith_summary":"This paper reports the 2025 Surgical Tattoos in Infrared Challenge (STIRC2025), a public benchmark for point tracking in surgical video. Point tracking is framed as a building block for downstream surgical applications such as segmentation, 3D reconstruction, virtual tissue landmarking, autonomous probe scanning, and subtask autonomy. Participants submit tracking algorithms that are scored on the STIR infrared-tattoo dataset for both accuracy (on in vivo and ex vivo sequences) and efficiency (inference latency). Seven teams competed under MICCAI EndoVis 2025. The paper summarizes the rankings, the methods those teams used, and releases the dataset plus metric code so others can reproduce or extend the evaluation. A sympathetic reader cares because a shared, quantitative ranking on real surgical sequences makes it possible to compare trackers fairly rather than relying on private datasets or qualitative demos.","feed_headline":"Seven trackers ranked on surgical infrared tattoos","feed_subtitle":"STIRC2025 scores accuracy on in vivo and ex vivo video plus inference latency, with data and code public.","key_machinery":"The STIR infrared-tattoo evaluation: ground-truth points marked by infrared tattoos on tissue, scored for tracking accuracy on in vivo and ex vivo video plus wall-clock inference latency.","core_discovery":"STIRC2025 establishes a dual accuracy-and-efficiency ranking of seven submitted point-tracking algorithms on the STIR infrared-tattoo dataset, covering both in vivo and ex vivo surgical sequences, and releases the dataset and metrics so the ranking can be audited and extended.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["STIRC2025 ranks seven point trackers on accuracy and latency","Seven algorithms ranked for surgical infrared-tattoo tracking","Dual accuracy-efficiency ranking of seven STIR trackers released","Challenge scores seven trackers on in vivo and ex vivo STIR video","Seven teams ranked on surgical point tracking via STIRC2025"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That infrared tattoo markers and the challenge’s accuracy-plus-latency scores are a good enough stand-in for the tracking quality that real surgical applications actually need.","fun_headline_variants_meta":{"raw":{"variants":["STIRC2025 ranks seven point trackers on accuracy and latency","Seven algorithms ranked for surgical infrared-tattoo tracking","Dual accuracy-efficiency ranking of seven STIR trackers released","Challenge scores seven trackers on in vivo and ex vivo STIR video","Seven teams ranked on surgical point tracking via STIRC2025"]},"model":"grok-4.5","effort":"low","cost_usd":0.004364,"raw_usage":{"total_tokens":1242,"prompt_tokens":724,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":43640000,"prompt_tokens_details":{"text_tokens":724,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":449,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":724,"tokens_out":69,"duration_ms":4195,"temperature":1.0,"reasoning_tokens":449,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T02:10:11.047182+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the top-ranked STIRC2025 trackers on the same sequences without tattoos or with clinical landmarks, and check whether the accuracy ranking and latency numbers still hold under those conditions.","supporting_citations":[],"review_version":1}