{"id":"47ba4635-2e61-46a1-baa6-3a75d4515702","arxiv_id":"2607.10427","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"BOCCHI’s gradient-overlapping real blur masks force frequency-aware learning, and MSDCT-UNet plus BOCCHI training beat prior sources on in-domain and cross-dataset local blur detection.","lead":"The paper releases BOCCHI, a 633-image real multi-camera benchmark for pixel-level local motion blur, designed so sharp regions share gradient statistics with blur and break simple low-gradient shortcuts. It also introduces MSDCT-UNet, a U-Net that injects multi-scale DCT blur cues, and shows BOCCHI training transfers better across datasets than larger prior sources.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Cross-dataset transfer gains may be driven by Inference Dataset composition and annotation/PSF confounds rather than the claimed gradient-overlap hardness.","rationale":"The reader correctly flags that the transfer superiority is the load-bearing claim and that gradient-overlap causality is the weakest assumption. The manuscript supplies clear numerical support (Tables 2–3, architecture-agnostic 13/13 pattern) and a well-motivated frequency architecture, but never experimentally disentangles PR25/μ_blur from the other factors that co-vary with BOCCHI (multi-camera capture, inclusive LabelMe polygons, scene diversity, and the mixed Inference construction itself). That is the single most important soft spot for the central “more realistic and challenging benchmark” story. The recommended concrete test directly probes both the test-set composition confound and the gradient-overlap causal claim; a positive result would secure the claim, a negative one would relegate BOCCHI to a useful but not uniquely hard data source. Because the empirical tables are still informative and the architecture contribution is independent, the appropriate stance remains CONDITIONAL rather than REJECT—exactly as the reader concluded. No stronger internal inconsistency or calculation error is present.","tokens_in":31737,"tokens_out":674,"duration_ms":8619,"concrete_test":"Re-evaluate the Table 3 Average row after (i) removing the 164 BOCCHI-protocol images from the Inference Dataset and (ii) re-training all 13 models on a gradient-matched synthetic or re-annotated subset of ReLoBlur/OMoBlur that deliberately raises sharp PR25/μ_blur to ≈0.68 while keeping original cameras and annotation protocol. If BOCCHI’s mIoU lead disappears or reverses under either change, the causal hardness claim does not hold.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper’s strongest claim is that BOCCHI-trained models beat every other source on the 572-image Inference Dataset (Table 3 averages: mIoU 0.563 vs OMoBlur 0.541 / ReLoBlur 0.425 / CUHK 0.214) because BOCCHI’s sharp–blur gradient overlap (PR25/μ_blur ≈ 0.68, Sec. 3.3 / Fig. 3) defeats gradient shortcuts. That causal link is not isolated. The Inference set is author-constructed and mixed (28.7 % BOCCHI-protocol, 28.3 % ReLoBlur, 26.2 % OMoBlur, 16.8 % CUHK; Sec. 3.3), so a non-trivial fraction of the test distribution matches BOCCHI’s capture protocol, cameras, and inclusive polygon style. Multi-camera PSF diversity, scene mix, and annotation inclusiveness (all motion artifacts inside the boundary) are confounded with the gradient statistic. Without a controlled ablation that holds annotation style / camera population / scene mix fixed while varying only the PR25/μ_blur ratio, the “defeats shortcuts” story for why BOCCHI is the superior training source remains correlational. Single-run tables without multi-seed variance further leave the modest +2.2 pp edge over OMoBlur fragile.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces BOCCHI, a 633-image real multi-camera benchmark for pixel-level local motion blur detection, arguing that its sharp-region gradient distribution overlaps the blur distribution (high PR25/μ_blur) and thereby defeats low-gradient=blur shortcuts present in prior datasets. It also proposes MSDCT-UNet, a U-Net-style encoder–decoder that injects multi-scale DCT features via multi-head DCT Attention, FiLM fusion, and an AFASPP bottleneck. Empirically, MSDCT-UNet ranks first in-domain on BOCCHI in mIoU and BdF1 among 13 models, and models trained on BOCCHI achieve the highest average cross-dataset mIoU/Dice/Recall on a 572-image mixed Inference Dataset despite fewer training images than ReLoBlur or OMoBlur.","tokens_in":32201,"tokens_out":1477,"duration_ms":23780,"significance":"If the empirical transfer advantage and hardness characterization hold under tighter controls, BOCCHI would be a useful community resource for local blur detection, where existing sets are small (CUHKmotion), scene-biased (ReLoBlur), or soft-mask synthetic (OMoBlur). The multi-camera real capture, human polygon annotations, and explicit gradient-overlap statistics are concrete contributions. The architecture is a solid engineering synthesis of HiFST-style DCT priors with modern fusion (FiLM, SE, ASPP) rather than a foundational theoretical advance, but the frequency-transform ablation and architecture ablations give a clear inductive-bias story. The architecture-agnostic transfer pattern (13/13 models) is a strength of the experimental design if the evaluation set is not biased toward BOCCHI’s capture protocol.","major_comments":[{"comment":"Abstract / Sec. 3.3 / Table 3: The central causal claim—that BOCCHI’s superior cross-dataset transfer is because sharp–blur gradient overlap (PR25/μ_blur ≈ 0.68) defeats gradient shortcuts—is not isolated from confounds. BOCCHI also differs in multi-camera PSF diversity, inclusive motion-artifact annotation style, and scene mix. The Inference Dataset is author-constructed with 28.7% BOCCHI-protocol images (Sec. 3.3). Without per-subset transfer tables (BOCCHI-protocol vs ReLoBlur vs OMoBlur vs CUHK held-out) and/or a controlled comparison that varies only the gradient-overlap statistic while holding annotation protocol and camera population fixed, the “defeats shortcuts” explanation remains correlational. Soften the causal language or add the subset breakdowns and a confound discussion as primary results.","section":null},{"comment":"Table 2–3 and Sec. 5.1–5.2: All headline numbers are single-run (seed 531). The in-domain margin over Cellpose3 is +1.2 pp mIoU / +4.0 pp BdF1, and the average cross-dataset edge over OMoBlur is only +2.2 pp mIoU. The paired 38/63 win rate with +0.7 pp median mIoU is suggestive but not a substitute for multi-seed means and standard deviations (or bootstrap CIs). For claims that BOCCHI “outperforms every other training source” and that MSDCT-UNet “ranks first,” report at least 3 seeds for the main BOCCHI and cross-dataset averages, or qualify rankings as single-run.","section":null},{"comment":"Sec. 5.3 / Table 2: The hardness argument rests heavily on DDRNet-23 trailing MSDCT-UNet by 14.9 pp on BOCCHI vs 2.0 pp on ReLoBlur. That comparison mixes architecture capacity and frequency bias. To support “BOCCHI more strongly penalizes gradient shortcuts,” also report (i) a pure spatial ablation of MSDCT-UNet (already in Table 4 as w/o Freq Branch) trained and tested on each source, and (ii) the same spatial-only backbone’s cross-dataset transfer when trained on BOCCHI vs ReLoBlur. Currently the 7.5× amplification is informative but incomplete as a benchmark-hardness proof.","section":null},{"comment":"Sec. 3.2 Annotation: Three candidate polygons with majority selection is good practice, but no inter-annotator agreement (e.g., mean pairwise IoU or boundary disagreement) is reported. Because annotation inclusiveness (“all visible motion artifacts”) is itself a plausible driver of transfer differences versus ReLoBlur/OMoBlur soft masks, IAA statistics and a short protocol comparison are load-bearing for interpreting Table 3.","section":null}],"minor_comments":[{"comment":"Supplementary Table 7: MSDCT-UNet at 1.0 FPS (1080×720) is a severe practical limitation relative to STDC/BiSeNet-class baselines. Mention this constraint in the main-text conclusion or limitations, not only the supplement.","section":null},{"comment":"Eq. (4): The 0.1·tanh(γ) residual scale is a free hyperparameter; briefly justify the 0.1 choice or note it was not tuned extensively.","section":null},{"comment":"Fig. 3 / Table 1: Mask convention “black→blur, white→sharp” is stated in the table caption but inverted in some qualitative figure descriptions (Fig. 8 says black = predicted blur). Harmonize mask polarity across all figures.","section":null},{"comment":"Sec. 4.2: Channel count 57 is inherited from HiFST; a short sensitivity check (e.g., 32 vs 57 vs 96) in the supplement would strengthen the design choice beyond “design constant.”","section":null},{"comment":"Related Work: OMoBlur is cited as contemporaneous CVPR 2026; ensure the camera-ready citation and any concurrent-work disclaimer match the final venue timeline.","section":null},{"comment":"Typos / polish: “CUHK motion” vs “CUHKmotion” inconsistency; “AF ASPP” spacing in Fig. 1; occasional missing spaces after periods in the abstract PDF text.","section":null}],"recommendation":"major_revision","confidential_remarks":"The empirical package is above average for a dataset+architecture CV paper, but the abstract’s causal framing (“defeat these shortcuts”) overreaches the evidence. If the authors add per-subset Inference results and multi-seed variance and tone down causality, this is a reasonable accept for a solid systems/dataset contribution. If they refuse to address the Inference composition issue, the transfer claim is hard to trust at face value. Scope fit is fine for a vision venue with a dataset track; not a theory paper."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is a new real-captured local-motion-blur benchmark (BOCCHI, 633 multi-camera images with human polygons) built so sharp-region gradients overlap blur ones, plus MSDCT-UNet that injects multi-scale DCT priors via attention and FiLM and tops the in-domain numbers. The cross-dataset result is the part that sticks: models trained only on BOCCHI beat ReLoBlur, OMoBlur and CUHKmotion averages on the mixed 572-image Inference set, even though BOCCHI is the smallest source.\n\nWhat is actually new is the capture protocol and the explicit gradient-overlap design (Table 1 / Fig. 3: PR25/μ_blur ≈ 0.68, highest of the five sets). The architecture recombines known pieces—HiFST-style multi-scale DCT on Sobel, NeXt blocks, FiLM, ASPP, SE—into a task-specific encoder-decoder with per-stage DCT Attention and deep supervision. That is fine; the empirical package is thorough. Thirteen baselines, architecture/loss/frequency-transform ablations, paired per-image check vs Cellpose3, and full cross-metric tables in the supplement. Spatial-only models take a much larger hit on BOCCHI than on ReLoBlur, which supports the hardness claim even if it does not isolate it.\n\nSoft spots are real but proportionate. The Inference set is author-built and 28.7 % BOCCHI-protocol, so part of the transfer edge is distribution match rather than pure gradient-overlap magic; multi-camera PSFs and inclusive annotation style are confounded with the PR25 statistic. No multi-seed runs or error bars, single fixed seed 531, and the model is 1 FPS at full res because of the CPU-side DCT. In-domain mIoU looks high partly because blur ratios are low. None of that sinks the contribution; it just means the causal story is correlational and the absolute numbers need a grain of salt.\n\nMath and citations look solid—standard DCT, ordinary BCE+Dice, appropriate prior work on blur detection and frequency attention. This is for people who care about partial blur masks, selective deblurring, or frequency priors in segmentation. It deserves a serious referee; release the data and code and the paper is accept-shaped after the usual tightening. I would bring it to reading group if we are talking datasets or blur this month, and I would cite the benchmark if I work on the problem.","headline":"Real multi-camera blur-mask set that transfers with 633 images, plus a DCT-UNet that wins in-domain; hardness story is correlational but the tables are useful.","tokens_in":32833,"tokens_out":623,"would_cite":true,"duration_ms":18328,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A real-captured blur benchmark that kills the low-gradient shortcut lets models detect local motion blur better with only 633 images.","keywords":["motion blur detection","local motion blur dataset","DCT attention","frequency-domain learning","semantic segmentation","BOCCHI","MSDCT-UNet"],"falsifier":"Retrain the same models on a version of BOCCHI whose sharp-region gradients have been filtered to remove the low-gradient tail (so PR25 no longer overlaps blur) and check whether the cross-dataset mIoU advantage over ReLoBlur and OMoBlur disappears.","tokens_in":32613,"feed_emoji":"📷","tokens_out":619,"duration_ms":7375,"temperature":0.7,"pith_summary":"Local motion blur detection means marking, at every pixel, which parts of a photo are smeared by moving objects while the rest stays sharp. Prior benchmarks often let models cheat by treating low-gradient patches as blur; those shortcuts collapse when the same model faces new cameras or scenes. This paper builds BOCCHI, 633 real photographs from five consumer cameras whose sharp regions deliberately share the same gradient range as the blurred ones, forcing models to use frequency evidence instead. It also introduces MSDCT-UNet, which feeds multi-scale DCT features into every stage of a U-Net via attention and FiLM fusion. On BOCCHI the new network leads in mask accuracy and boundary quality; more importantly, any model trained only on BOCCHI transfers better to a held-out mixed test set than models trained on larger existing sources.","feed_headline":"633 real photos beat larger blur datasets at transfer","feed_subtitle":"Sharp–blur gradient overlap kills the low-gradient shortcut; a DCT U-Net leads in-domain","key_machinery":"MSDCT-UNet: a U-Net that, at every encoder and decoder stage, extracts 57-channel multi-scale high-frequency DCT features from the Sobel gradient map, reweights them with multi-head DCT Attention, and fuses them into spatial features via FiLM modulation plus an attentive frequency ASPP bottleneck.","core_discovery":"BOCCHI’s sharp-region gradients overlap the blur-region gradient distribution more than any prior local-motion-blur dataset, defeating the low-gradient-equals-blur shortcut; models trained on its 633 images therefore achieve the strongest cross-dataset transfer, and the frequency-aware MSDCT-UNet that injects multi-scale DCT priors ranks first in-domain on BOCCHI.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["633 real shots beat larger blur sets at transfer","BOCCHI gradient overlap defeats low-gradient shortcuts","MSDCT-UNet tops BOCCHI mIoU with multi-scale DCT","Real BOCCHI images win cross-dataset blur transfer","Frequency-aware UNet leads local motion blur on BOCCHI"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The claim that BOCCHI’s transfer gains come mainly from its sharp–blur gradient overlap, rather than from multi-camera diversity, inclusive annotation style, or the makeup of the mixed test set.","fun_headline_variants_meta":{"raw":{"variants":["633 real shots beat larger blur sets at transfer","BOCCHI gradient overlap defeats low-gradient shortcuts","MSDCT-UNet tops BOCCHI mIoU with multi-scale DCT","Real BOCCHI images win cross-dataset blur transfer","Frequency-aware UNet leads local motion blur on BOCCHI"]},"model":"grok-4.5","effort":"low","cost_usd":0.00437,"raw_usage":{"total_tokens":1217,"prompt_tokens":685,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":43700000,"prompt_tokens_details":{"text_tokens":685,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":462,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":685,"tokens_out":70,"duration_ms":4626,"temperature":1.0,"reasoning_tokens":462,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T11:48:38.821564+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Retrain the same models on a version of BOCCHI whose sharp-region gradients have been filtered to remove the low-gradient tail (so PR25 no longer overlaps blur) and check whether the cross-dataset mIoU advantage over ReLoBlur and OMoBlur disappears.","supporting_citations":[],"review_version":1}