{"id":"e7c59db1-74da-4d35-ab1d-435204652862","arxiv_id":"2506.05175","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"TAO pipelines object-centric anomaly scores into SAM2 prompts with a temporal consistency filter to obtain pixel-level anomaly segmentation and tracking.","lead":"The paper presents TAO, a pipeline that combines object-centric video anomaly detection with the SAM2 segmentation model to produce pixel-level anomaly masks and track anomalous objects across frames. It reports state-of-the-art scores on UCSD Ped2 and ShanghaiTech, but the evaluation depends on per-dataset thresholds and no code is released.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SOTA claim depends on per-dataset threshold tuning and a baseline protocol that may distort pixel-F1; abstract's 'no threshold tuning' contradicts Eq. (3) and Sec. 4.2, so the headline numbers are not established.","rationale":"I read the paper in good faith. The core idea—linking object-centric anomaly scores to SAM2 prompts through a temporal filter—is coherent, and the ablation in Table 4 does show that both video tracking mode and robust filtering contribute. However, the central SOTA claim is not yet secure because the method contradicts its own abstract: Eq. (3) is a thresholding step, and Sec. 4.2 sets τ per dataset after the sensitivity analysis in Fig. 5 selects the optimal value on ShanghaiTech. Since all headline metrics are computed from boxes that survive this threshold, the reported margins may be inflated by test-set hyperparameter selection. The baseline comparison also lacks a specified, shared binarization/thresholding protocol, which matters because pixel-F1 is threshold-dependent. The reader's weakest assumption about temporal consistency in Algorithm 1 is plausible and should be tested, but it is more about generalization to fast-moving or occluded anomalies than about whether the reported SOTA numbers are trustworthy; the protocol dependence is the more load-bearing issue. These concerns do not justify rejection, but they do justify keeping the verdict CONDITIONAL until the evaluation is fixed and reproduced.","tokens_in":13501,"tokens_out":4591,"duration_ms":56294,"concrete_test":"Re-run Table 1 under a fixed protocol: choose τ and l on the training split only, apply identical binarization and box-merging code to TAO and all baselines, and report mean±std over at least three runs. If TAO's Pixel-F1 drops by more than 5 points from 64.12, or DRAEM/SimpleNet gain more than 5 points, the SOTA claim is not robust to evaluation choices.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—SOTA Pixel-F1 of 64.12 on Ped2 and RBDC/TBDC of 83.6/93.2—rests on hyperparameters selected per test set. Equation (3) retains boxes only when s_i > τ, yet the abstract says the method 'removes the need for threshold tuning'; Sec. 4.2 fixes τ=1.5 for Ped2 and τ=1.6 for ShanghaiTech, with save intervals l=5 and l=15, and Fig. 5 is a sensitivity sweep whose stated optimum (τ=1.6) is on ShanghaiTech. These choices are made on the evaluation set, and Pixel-F1, RBDC, and TBDC are all computed after this thresholding. The baselines in Table 1 are image-level anomaly detectors applied frame-by-frame with unspecified heatmap thresholds and no temporal filtering, so the large margins (e.g., 64.12 vs 10.89) may reflect threshold/protocol differences rather than the proposed tracking-plus-SAM2 pipeline. No error bars or repeated runs are reported. The Algorithm 1 temporal-consistency assumption is a plausible secondary robustness concern, but the more immediate threat to the SOTA claim is that the reported numbers are protocol-dependent.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Track Any Anomalous Object (TAO), a video anomaly detection pipeline that combines object-centric bounding-box anomaly scoring with a temporal consistency filter (Algorithm 1) and SAM2-based promptable segmentation. The claimed contributions are a new dual-level (pixel and object) evaluation benchmark and state-of-the-art results on UCSD Ped2 and ShanghaiTech Campus, including Pixel-F1 of 64.12 and RBDC/TBDC of 83.6/93.2 on Ped2. The pipeline is described in Sections 3.1-3.3 and evaluated in Section 4 with ablations, sensitivity analyses, and qualitative comparisons.","tokens_in":13810,"tokens_out":7733,"duration_ms":80550,"significance":"If validated, TAO would be a useful integration of object-level anomaly scoring with a frozen video segmentation foundation model, and the proposed dual-level benchmark would address a real gap in VAD evaluation: existing frame-level and object-level metrics do not measure pixel-accurate localization of multiple anomalous objects. The paper is clearly written and Algorithm 1 is well specified, and the idea of linking anomaly scores to SAM2 prompts is simple and falsifiable. However, the headline SOTA claim is not currently established because the evaluation has load-bearing weaknesses: per-dataset threshold selection on the test sets, an underspecified baseline protocol, no error bars, and an asserted removal of threshold tuning that is contradicted by the method itself.","major_comments":[{"comment":"The abstract's claim that TAO 'removes the need for threshold tuning' is contradicted by the method: Eq. (3) retains boxes only when s_i > tau, and Sec. 4.2 assigns dataset-specific thresholds (tau=1.5 for Ped2, tau=1.6 for ShanghaiTech) and save intervals (l=5 vs. l=15). Figure 5 is a sensitivity sweep on the ShanghaiTech test set whose optimum (tau=1.6) is then used in the reported results. Since tau and l affect Pixel-F1, RBDC, and TBDC, the reported SOTA numbers are partly fitted to the test data. Please select all hyperparameters on held-out validation data, fix them across datasets, or report full operating curves; otherwise the central empirical claim is not an independent prediction.","section":"Sec. 4.2, Eq. (3), Fig. 5"},{"comment":"The comparison with image anomaly detectors is not controlled. The baselines are applied frame-by-frame and 'pixel-level segmentation is achieved by applying a threshold to these heatmaps,' but the threshold is never specified, and the baselines are not given temporal filtering. The large Pixel-F1 margin (64.12 vs. 10.89 for DRAEM) could therefore reflect an unfavorable binarization threshold for the baselines rather than a property of the TAO pipeline. Report the heatmap thresholds used for each baseline (or best thresholds under the same protocol), and ideally evaluate baselines with a simple temporal smoothing step to isolate the contribution of temporal consistency.","section":"Sec. 4.3, Table 1"},{"comment":"The proposed 'new comprehensive VAD benchmark' is not reproducible as described. The paper defines pixel-level metrics on UCSD Ped2 and ShanghaiTech Campus but never states how the pixel-level ground-truth masks were obtained, whether they are part of the original datasets, or how any box-level annotations were converted to masks. Without this annotation protocol, the Pixel-AUROC, Pixel-AP, Pixel-AUPRO, and Pixel-F1 numbers cannot be independently verified. Specify the source of pixel labels and any preprocessing, and release the annotation code.","section":"Sec. 4.1, Benchmark Datasets"},{"comment":"The Boxes Robustness Filtering relies on the temporal-consistency premise that true anomalous boxes reappear with IoU > h in at least m of k neighboring frames, while false positives appear sporadically. This premise is asserted without evidence, and no experiments cover fast-moving, occluded, or mutually overlapping anomalies, exactly the cases where the premise can fail. The paper should include failure-case analysis or synthetic stress tests; otherwise the claimed robustness of the pipeline is unsupported for object classes beyond those in the two test sets.","section":"Sec. 3.2, Algorithm 1"},{"comment":"No variance or repeated-run statistics are reported. On ShanghaiTech, the TBDC improvement over STPT is only 0.8 points (85.4 vs. 84.6), and with the threshold-fitted protocol described above it is unclear whether this difference is significant. Report means and standard deviations over at least three runs, or justify why the pipeline is deterministic and the single-run numbers are exact.","section":"Tables 1-3"}],"minor_comments":[{"comment":"The heading 'Comparsion to Conventional Video Anomaly Detection Models' contains a typo; it should be 'Comparison'.","section":"Sec. 4.4 heading"},{"comment":"The sentence 'segmentation probability maps are uniformly binarized' omits the binarization threshold; since Pixel-F1 depends directly on this threshold, its value should be stated.","section":"Sec. 4.2"},{"comment":"The set notation writes I={(c_j, b_j, f_i) | ...}, which mixes frame indices and frame contents; clarify that the prompts contain the actual frame tensors for the saved indices.","section":"Eq. (4)"},{"comment":"The caption says 'Quantitative Results' but does not identify which metric is plotted on the vertical axis; please specify.","section":"Fig. 6"},{"comment":"Several references lack venues or identifiers (e.g., [7], [8], [36], [45]), and some cited papers appear unrelated to video anomaly detection; please verify the relevance and completeness of the reference list.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The reference list contains a number of entries that do not appear to support the related-work claims (e.g., references on image editing and pan-sharpening such as [2], [3], [4], [49], [50]); the editor may wish to verify that these citations are necessary. The paper also makes a 'first' claim about integrating pixel-level tracking into VAD without discussing concurrent or prior detection-plus-SAM2 pipelines; this should be checked during revision. The central issue, however, is empirical: the SOTA claim is tied to test-set threshold selection and an uncontrolled baseline protocol, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a system-level integration that deserves a serious look, but the headline numbers should not be trusted yet. The TAO pipeline — object-centric anomaly scoring, a temporal consistency filter, and SAM2 for mask propagation — is new to VAD, and the idea of evaluating pixel and object metrics together is useful. The ablation is honest: both the filtering and the video tracking mode meaningfully improve pixel-F1 on Ped2.\n\nThe soft spots are in the evaluation. The abstract says the method 'removes the need for threshold tuning,' but Eq. (3) and Sec. 4.2 explicitly set τ=1.5 on Ped2 and τ=1.6 on ShanghaiTech, and Fig. 5 is a sensitivity sweep on ShanghaiTech that picks the optimal value on the test set. The save interval l is also dataset-specific. So the 'no tuning' claim is simply wrong, and the reported SOTA is partially fitted to the benchmark. The baselines are image-level anomaly detectors applied frame-by-frame with unspecified heatmap thresholds and no temporal filtering, so a 64.12 vs 10.89 pixel-F1 gap likely reflects protocol differences more than the pipeline's advantages. No error bars, no code, and the object-centric scorer is not specified precisely enough to reproduce.\n\nThe temporal consistency assumption in Algorithm 1 is a plausible secondary concern — fast-moving or occluded anomalies could be filtered out — but that is not the main issue here.\n\nBottom line: this is a constructive system paper with a useful new evaluation protocol, but the empirical claims are currently unsubstantiated. I'd send it to review and ask for code, detailed baseline thresholds, per-dataset parameter selection explained as such, and error bars. The idea is worth refining, not discarding.","headline":"A useful system integration and evaluation protocol, but the SOTA claims rest on test-set threshold tuning and an apples-to-oranges baseline comparison.","tokens_in":14334,"tokens_out":2603,"would_cite":false,"duration_ms":30582,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TAO is the first pipeline to turn object-centric anomaly scores into pixel-level tracks of multiple anomalies, reporting state-of-the-art Pixel-F1 of 64.12 on UCSD Ped2 and RBDC/TBDC of 83.6/93.2.","keywords":["video anomaly detection","pixel-level anomaly localization","object-centric anomaly detection","SAM2 prompting","temporal consistency filtering","anomaly segmentation","RBDC/TBDC metrics"],"falsifier":"Take a test clip in which a true anomalous object is occluded or moves so quickly that its bounding boxes overlap less than the h=0.2 threshold for at least two frames within any five-frame window, run TAO with the paper's fixed parameters, and measure Pixel-F1 against the ground-truth masks; if the score collapses relative to supplying ground-truth boxes directly to SAM2, the filter, not the segmenter, is the failure point.","tokens_in":13343,"feed_emoji":"🎯","tokens_out":5549,"duration_ms":60583,"temperature":0.7,"pith_summary":"This paper argues that video anomaly detection should be treated as a pixel-level tracking problem, not just a frame-level or object-level scoring problem. It proposes TAO, a pipeline that takes bounding boxes from an object-centric anomaly scorer, filters them for temporal consistency, and feeds the surviving boxes and their center points into SAM2, a prompt-based video segmentation model that propagates masks across the whole clip. On UCSD Ped2, TAO reports Pixel-F1 of 64.12 versus 10.89 for the best image-anomaly baseline, with RBDC/TBDC of 83.6/93.2; on ShanghaiTech Campus it reports RBDC/TBDC of 62.1/85.4. If these numbers hold, TAO is the first framework to localize multiple anomalous objects at pixel level in long videos while still tracking each anomaly as a coherent object.","feed_headline":"Pixel-level anomaly tracking via TAO tops video detection baselines","feed_subtitle":"Object anomaly scores become SAM2 prompts; Pixel-F1 jumps from 10.89 to 64.12 on Ped2.","key_machinery":"The load-bearing mechanism is the coupling of three components: an object-centric anomaly scorer that assigns anomaly scores to detected bounding boxes, the Boxes Robustness Filtering algorithm that enforces temporal consistency by requiring each candidate box to overlap boxes in at least m of the previous or next k frames to inherit or receive a label, and SAM2, a prompt-based video segmentation model that takes the surviving boxes and their center points, saved every l frames, and propagates pixel masks across the entire clip. The filter converts noisy box-level scores into reliable prompts; in the ablation, enabling video tracking and filtering together raises Pixel-AUROC from 41.72 to 75.11 on UCSD Ped2.","core_discovery":"The central claim is that pixel-level anomaly segmentation can be obtained without dense per-pixel training by converting an object-centric anomaly detector into a prompt source for a video segmentation foundation model. The paper introduces a Boxes Robustness Filtering step that keeps only anomaly boxes reappearing with sufficient overlap in at least three of five neighboring frames, preventing redundant boxes from corrupting SAM2's temporal propagation; with this filter, TAO achieves the reported state-of-the-art results on UCSD Ped2 and ShanghaiTech under both pixel-level and object-level metrics. The authors also propose a dual-level benchmark that evaluates pixel metrics (AUROC, AP, AUPRO, F1) and object metrics (RBDC, TBDC) together, arguing that this combined view is better for complex scenes with overlapping or multiple anomalies.","pith_inferences":["A natural extension is to make the filter's window parameters scene-adaptive; fast-moving or intermittently occluded anomalies would likely need a smaller overlap threshold or a shorter window than the fixed k=5, m=3, h=0.2 used here.","The main reported comparisons are against frame-by-frame image anomaly detectors; a video-native baseline with its own temporal smoothing could be a stronger test of TAO's tracking contribution.","Because the pipeline still relies on a score threshold to nominate candidate boxes, the paper's 'threshold-free' claim is better read as 'less sensitive to the threshold' rather than 'no threshold at all'."],"forward_implications":["Any object-centric video anomaly detector can be upgraded to a pixel-level multi-object tracker by appending the filtering and prompt-propagation stages, with no retraining of the segmentation model.","Future video anomaly detection models will be judged on two dimensions at once, spatial precision from pixel metrics and temporal coherence from RBDC and TBDC, so a method that is strong on only one axis is no longer sufficient.","Sparse prompting works: saving filtered boxes every l frames, 5 on Ped2 and 15 on ShanghaiTech, still propagates masks to all frames, reducing storage and compute for long surveillance clips.","The framework stays stable across SAM2 backbone sizes from Tiny to Large and across anomaly thresholds in the range 1.0 to 1.9, indicating the gains are not tied to a particular prompt threshold or model capacity."],"supporting_citations":[{"why":"Supplies SAM2, the prompt-based video segmentation model that propagates masks from filtered boxes to all frames.","marker":"[43]"},{"why":"Supplies the object-level anomaly scoring algorithm that outputs anomalous bounding boxes from pose and depth features used as prompts.","marker":"[44]"},{"why":"DRAEM is one of the strongest image anomaly detection baselines; its binarized heatmaps are compared on pixel and object metrics and beaten on Pixel-F1.","marker":"[61]"},{"why":"SimpleNet is a fully trained image anomaly detection baseline providing pixel-level comparison.","marker":"[34]"},{"why":"Defines the object-level metrics RBDC and TBDC and the evaluation protocol used to compare object-centric video anomaly detection methods.","marker":"[41]"},{"why":"AED-SSMTL is a state-of-the-art object-centric video anomaly detection baseline against which TAO reports higher RBDC and TBDC on both datasets.","marker":"[11]"},{"why":"Provides the ShanghaiTech Campus dataset, one of the two evaluation benchmarks.","marker":"[31]"},{"why":"Provides the UCSD Ped2 dataset, the other evaluation benchmark.","marker":"[48]"}],"fun_headline_variants":["TAO: anomaly boxes prompt SAM2 for pixel-level detection","Object tracking meets pixel segmentation in TAO's VAD pipeline","From boxes to pixels: TAO rethinks anomaly detection as tracking","Pixel-F1 jumps from 10.89 to 64.12 with TAO's prompt filter","No per-pixel training: TAO leverages SAM2 for granular anomaly masks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole tracking stage assumes that a genuine anomaly reappears in at least three of the neighboring five frames with enough spatial overlap, while false alarms appear only sporadically; if an anomaly moves too fast, is occluded, or overlaps a normal object, the filter can drop it and SAM2 then propagates the wrong mask.","fun_headline_variants_meta":{"raw":{"variants":["TAO: anomaly boxes prompt SAM2 for pixel-level detection","Object tracking meets pixel segmentation in TAO's VAD pipeline","From boxes to pixels: TAO rethinks anomaly detection as tracking","Pixel-F1 jumps from 10.89 to 64.12 with TAO's prompt filter","No per-pixel training: TAO leverages SAM2 for granular anomaly masks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000324,"raw_usage":{"total_tokens":1784,"prompt_tokens":879,"completion_tokens":905,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":806}},"tokens_in":495,"tokens_out":905,"duration_ms":11637,"temperature":1.0,"reasoning_tokens":806,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:22:35.450612+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a test clip in which a true anomalous object is occluded or moves so quickly that its bounding boxes overlap less than the h=0.2 threshold for at least two frames within any five-frame window, run TAO with the paper's fixed parameters, and measure Pixel-F1 against the ground-truth masks; if the score collapses relative to supplying ground-truth boxes directly to SAM2, the filter, not the segmenter, is the failure point.","supporting_citations":[{"cited_title":"Attribute-based representa- tions for accurate and interpretable video anomaly detection","cited_arxiv_id":null,"evidence_quote":"Supplies the object-level anomaly scoring algorithm that outputs anomalous bounding boxes from pose and depth features used as prompts."},{"cited_title":"Draem- a discriminatively trained reconstruction embedding for sur- face anomaly detection","cited_arxiv_id":null,"evidence_quote":"DRAEM is one of the strongest image anomaly detection baselines; its binarized heatmaps are compared on pixel and object metrics and beaten on Pixel-F1."},{"cited_title":"Simplenet: A simple network for image anomaly detection and localization","cited_arxiv_id":null,"evidence_quote":"SimpleNet is a fully trained image anomaly detection baseline providing pixel-level comparison."},{"cited_title":"Street scene: A new dataset and evaluation protocol for video anomaly detection","cited_arxiv_id":null,"evidence_quote":"Defines the object-level metrics RBDC and TBDC and the evaluation protocol used to compare object-centric video anomaly detection methods."},{"cited_title":"Anomaly detection in video via self- supervised and multi-task learning","cited_arxiv_id":null,"evidence_quote":"AED-SSMTL is a state-of-the-art object-centric video anomaly detection baseline against which TAO reports higher RBDC and TBDC on both datasets."},{"cited_title":"Fu- ture frame prediction for anomaly detection–a new baseline","cited_arxiv_id":null,"evidence_quote":"Provides the ShanghaiTech Campus dataset, one of the two evaluation benchmarks."},{"cited_title":"Anomaly detection in crowd scene","cited_arxiv_id":null,"evidence_quote":"Provides the UCSD Ped2 dataset, the other evaluation benchmark."}],"review_version":1}