{"id":"2977d5af-e7b9-4c05-9446-3a566f3bf244","arxiv_id":"2412.10224","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A sequence-aware transformer for interactive image segmentation that uses previous images and clicks as prompts, claiming state-of-the-art results without reporting them.","lead":"This paper proposes a transformer that uses earlier images, clicks, and masks as prompts to segment the same object in later images. It claims to outperform prior interactive segmentation tools, but the manuscript contains no numerical results to support that claim.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim 'surpasses state-of-the-art' is unverifiable: the main comparison table and the ablation table are both unresolved '??' placeholders, leaving no reported NoC/IoU numbers.","rationale":"The reader's identified weakest assumption (DINOv2 similarity as a proxy for prompt usefulness) is plausible but second-order; the more fundamental gap is that the paper reports no quantitative results at all. The ablation that would test the TPS assumption is itself missing, as is the main comparison table. Therefore the reader's REJECT verdict stands, but for a slightly different reason: the evidence base is absent. No amount of methodological plausibility can support a universal 'outperforms all baselines' claim without numbers, and the paper provides none.","tokens_in":8712,"tokens_out":3770,"duration_ms":33864,"concrete_test":"Render all table environments in the submitted PDF; if Table 1 (Section IV-B) and Table 2 (Section IV-C) both display '??' instead of numeric results, the claim is vacuous. If the authors provide the missing tables, verify that the SPT rows list NoC85/NoC90 values lower than every baseline on all five datasets.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section IV-B claims 'Our SPT framework consistently outperforms the baselines across all datasets and evaluation metrics' but refers only to 'Table ??', which is missing from the manuscript. Section IV-C's ablation study also cites 'Table ??' for ExpID #1-#9, so the claimed benefits of SPT and TPS have no reported data. No quantitative results (NoC85/NoC90, IoU, or MIoU) appear anywhere in the paper; Fig. 4 is referenced but its values are not described, and Fig. 3's qualitative comparison shows inconsistent labels. Because the central contribution is an empirical performance claim, the absence of every comparison table means the claim is not merely weakly supported but wholly unsupported. The ADE20K-Seq benchmark is also not released, so independent evaluation is impossible. This is the single load-bearing concern: if the missing tables were supplied and showed SPT behind SimpleClick, the paper's main conclusion would collapse.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SPT (Sequence Prompt Transformer), a method for interactive image segmentation that, unlike prior single-image methods, exploits a sequence of images depicting the same object category by using previous images, clicks, and predicted masks as prompts. A Top-k Prompt Selection (TPS) module, based on DINOv2 feature similarity, chooses the most relevant prompts from the sequence. The authors also introduce a new benchmark, ADE20K-Seq, constructed from ADE20K by grouping images into seven categories. The empirical section claims state-of-the-art results on GrabCut, Berkeley, COCO-MVal, DAVIS, and ADE20K-Seq, and ablations are said to validate the contributions of SPT and TPS. However, the manuscript as submitted contains no quantitative results: the main comparison table and the ablation table are both unresolved 'Table ??' placeholders, and the only numerical figure (Fig. 4) is not accompanied by aggregate metrics or a statistical description. The central claim of surpassing prior methods is therefore unverifiable from the provided text.","tokens_in":8952,"tokens_out":3652,"duration_ms":33189,"significance":"If the claimed results were supplied and verified, the paper would address a genuinely new and practically relevant variant of interactive segmentation—segmenting the same object category across a sequence of images—and the proposed architecture (causal-mask Transformer over a sequence of image-click-mask features, plus similarity-based prompt selection) is a plausible design. The introduction of a dedicated sequential interactive-segmentation benchmark is also a useful contribution, provided the dataset is described precisely and released. However, at present the paper contains no quantitative evidence: no NoC85/NoC90, IoU, or MIoU numbers appear anywhere, and the ablation analysis that would isolate the effect of the two core modules is missing. The significance of the work cannot be assessed until these results are actually reported. The idea itself is interesting, but the manuscript in its current form is an incomplete research report rather than a verifiable technical contribution.","major_comments":[{"comment":"The main experimental results are entirely absent. Section IV-B states that 'Our SPT framework consistently outperforms the baselines across all datasets and evaluation metrics' but refers only to an unresolved 'Table ??'. No NoC85, NoC90, IoU, or MIoU values are provided for any dataset or baseline in the entire manuscript. Since the paper's central claim is an empirical superiority claim, the absence of the main comparison table means that claim is wholly unsupported. This is a load-bearing deficiency that must be fixed by reporting the actual numbers, not just a promise of a table.","section":"IV-B, Table ??"},{"comment":"The ablation study is also missing. Section IV-C claims that the Sequence Prompt Transformer improves performance (ExpID #1 vs #5) and that Top-k Prompt Selection is effective (ExpID #6 vs #9), but all of these assertions refer to 'Table ??', which does not appear in the manuscript. Without the ablation table, the attribution of performance gains to SPT and TPS is unfounded, and the paper's two named contributions cannot be independently evaluated. The authors must provide the full ablation results, including the settings for different prompt lengths and different selection methods, before the claims can be taken seriously.","section":"IV-C, Table ??"},{"comment":"The newly introduced ADE20K-Seq benchmark is described in only two sentences: it 'extend[s] ADE20K dataset into 7 category-specific benchmarks, with each category containing more than 100 images', and it is said to contain random tasks. No details are given on how images are selected, how objects/instances are paired across images, how segmentation masks are obtained, or how the evaluation protocol is defined. Moreover, the dataset is not released. As a result, the reported evaluation on ADE20K-Seq (if any) would be impossible to reproduce, and the benchmark itself is not a usable contribution. This is a major reproducibility gap that should be addressed, for example by describing the construction procedure precisely and providing a public release link or a clear statement of availability.","section":"IV-A.2, ADE20K-Seq"}],"minor_comments":[{"comment":"The abstract says the method segments 'a series of images featuring the same target object', while Section III-A and the ADE20K-Seq description refer to 'same category'. These are different notions: same object identity versus same object class. The paper should clarify which setting is actually addressed and be consistent throughout.","section":"Abstract and Section I"},{"comment":"The dataset list in Section IV-A.1 misspells 'LVIS' as 'LIVIS' and writes 'DA VIS' instead of 'DAVIS'. Please correct these typos for the camera-ready version.","section":"IV-A, Dataset list"},{"comment":"Section IV-B refers to 'ADE20K-Sep' (a typo for ADE20K-Seq) and states that Fig. 4 shows MIoU versus number of clicks. However, the figure is not described in any quantitative way: no exact MIoU values, no error bars, and no statistical significance test are reported. As a qualitative plot, it cannot substitute for a numerical comparison table.","section":"IV-B, Figure 4"},{"comment":"The definition of the causal mask function in Eq. (4) is notational unclear: it says 'for an element at position i = (x, y) in the sequence' and then defines mask(x, y), but the roles of x and y (sequence positions versus spatial positions) are not explained. Clarify that x indexes the current sequence element and y indexes the attended element, or rewrite the equation to avoid confusion.","section":"III-D, Equation (4)"},{"comment":"The caption of Fig. 3 lists '(e) Results of RITM' and then '(e) Results of SPT (ours)', with the letter (e) repeated. Also, the figure's per-panel IoU values are not accompanied by an aggregate measure over the dataset, and the qualitative selection is not described as representative. Please fix the caption and provide aggregate numbers in the text or table.","section":"Fig. 3 caption"},{"comment":"References [3] and [46] are duplicated; both are the CLIP paper (Radford et al., 2021). Consolidate them into a single entry.","section":"References"},{"comment":"Related work mentions video-based interactive segmentation methods [47], [48] as inapplicable because of timing/scene inconsistency, but it does not discuss context-aware or memory-based segmentation models such as SegGPT [42] or other 'segment everything in context' approaches. A brief comparison would help position the contribution more accurately.","section":"Related Work"}],"recommendation":"major_revision","confidential_remarks":"To the editor: This manuscript is unusually incomplete even for a preprint: both the main results table and the ablation table are unresolved LaTeX placeholders, and no numerical evaluation appears anywhere. The paper's only claim of state-of-the-art performance is therefore not just weakly supported but wholly unsupported. I have recommended major revision rather than outright rejection because the missing content is in principle fixable if the authors actually have the experimental results. However, I want to flag that if the missing tables do not exist or the results do not show a consistent advantage over SimpleClick and similar baselines, the paper's central conclusion collapses. The novelty of the task formulation and the architectural idea is reasonable, but the evaluation infrastructure (including dataset release) must be completed before this can be considered a serious contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Steven,\n\nThe one genuinely new thing here is the task formulation: feeding a sequence of image/click/mask tuples as prompts for interactive segmentation. The components — causal attention over the sequence, DINOv2 similarity for top-k retrieval — are sensible and not in the cited prior work. The idea is worth exploring.\n\nBut the paper as submitted is an empty shell. The main results table and the entire ablation table are unresolved \"Table ??\" placeholders. There are no NoC85/NoC90 or IoU numbers anywhere in the text. The abstract's \"surpasses state-of-the-art\" claim is therefore not just weakly supported; it is unsupported.\n\nWhat is good: the method description is clear enough to reimplement in principle. The ADE20K-Seq benchmark, though not released, addresses a real evaluation gap. The causal masking is standard and correctly specified. The TPS assumption (DINOv2 similarity is a good proxy for prompt usefulness) is plausible and worth testing. So the intellectual core is not nonsense — it just has no empirical evidence attached.\n\nSoft spots beyond the missing tables: Fig. 3 has a mislabeled duplicate '(e)' and inconsistent IoU numbers. The reference list contains duplicates (e.g., [3] and [46]) and irrelevant entries (Eason 1955, BraTS, knee segmentation). That is sloppy and does not inspire confidence. ADE20K-Seq is not released, so even the self-created benchmark cannot be independently evaluated. The claim that \"existing methods ignore sequential information\" overlooks the video-based interactive segmentation work cited later, but that is a minor overstatement.\n\nOverall: this is a draft, not a paper. A serious editor would desk-reject because there is nothing to referee: no numbers, no code, no data. The idea might be worth a workshop submission once the authors actually run the experiments and report them. I would not cite it in its current state. If you are short on time, skip it. If you are curious about the formulation, it is a five-minute skim.\n\nRecommendation: reject/desk-reject as submitted; invite resubmission with real results.","headline":"A plausible new sequence-prompt formulation for interactive segmentation, but zero reported numbers make the central claim unverifiable.","tokens_in":9435,"tokens_out":2830,"would_cite":false,"duration_ms":24823,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Feeding earlier images, clicks, and predicted masks into a causally masked transformer improves interactive segmentation, the paper claims, beating state-of-the-art click methods on five benchmarks.","keywords":["interactive image segmentation","sequence prompting","click-based segmentation","causal self-attention","prompt selection","DINOv2 features","ADE20K-Seq","transformer"],"falsifier":"Re-run SPT on ADE20K-Seq with the TPS module removed (e.g., using the k most recent images as prompts instead of the DINOv2-selected ones) and compare NoC85/NoC90 and M-IoU; if the DINOv2 selection margin disappears, reverses, or shrinks to statistical noise, the claim that Top-k Prompt Selection carries the improvement is false.","tokens_in":8541,"feed_emoji":"🖱️","tokens_out":6839,"duration_ms":62205,"temperature":0.7,"pith_summary":"Interactive image segmentation usually treats each image as an isolated problem: a user clicks, the model segments, repeat per image. This paper claims that when the same object category appears across many images, the history itself is a prompt—the model should read earlier images, clicks, and predicted masks as a sequence and let the current image borrow from similar earlier frames. To make that idea work, the authors build a Sequence Prompt Transformer whose concealed (causal) self-attention hides future frames, and a Top-k Prompt Selection module that uses DINOv2 image features to choose the most similar earlier images as prompts. They report that this setup beats state-of-the-art click-based methods on GrabCut, Berkeley, COCO-MVal, DAVIS, and their new ADE20K-Seq benchmark, with the largest gains at few clicks. If the claim holds, annotating a batch of same-category images would need many fewer clicks, directly cutting the cost of pixel-level data labeling.","feed_headline":"Sequence prompting slashes clicks for interactive segmentation","feed_subtitle":"SPT feeds earlier images, clicks, and masks into a causal transformer, reporting state-of-the-art click counts on five benchmarks.","key_machinery":"The load-bearing mechanism is the Multi-head Concealed Self-Attention inside the Sequence Prompt Transformer: for a feature sequence $F$, position $i=(x,y)$ in the attention is visible only when $x \\ge y$, encoded by a mask function $mask(x,y)=1$ if $x\\ge y$ and $0$ otherwise, so each token attends to preceding positions and itself but never to future frames. Each input frame is formed by concatenating the click map and mask, embedding that concatenation, and adding it to the embedded image before the ViT computes $F_i = \\mathrm{ViT}(\\mathrm{Embed}(C_i \\oplus M_i) + \\mathrm{Embed}(I_i))$. The Top-k Prompt Selection module supplies the prompt subset by ranking DINOv2 features for similarity to the test image; the SPT output then goes through a Feature Pyramid Module and an MLP Segmentation Head, trained with focal loss.","core_discovery":"On its own terms, the paper's discovery is that the sequence of user interactions plus the model's own earlier masks is usable signal, not noise: a causally masked transformer layer placed after a ViT backbone lets the feature of the current image attend to features of earlier images, and the attention is guided by a position-wise mask so no future information leaks. The paper further claims that which earlier images you show matters, and that selecting the top-k most DINOv2-similar images as prompts outperforms using all or the most recent ones. The reported consequence is lower NoC85 and NoC90 on five benchmarks, including the new ADE20K-Seq dataset built from ADE20K sequences of seven categories, and better M-IoU at every click count, especially with very few clicks.","pith_inferences":["Editorial inference: the success of DINOv2-based TPS points to a testable refinement—train a small retrieval network with the segmentation loss so the prompt selector and segmenter are optimized jointly; the paper does not attempt this.","Editorial inference: because DAVIS is a video benchmark, a natural stress test is to compare SPT against video object segmentation methods that also propagate masks over time; the paper only compares against click-based single-image methods, so it has not yet isolated the contribution of sequence prompting from generic mask propagation.","Editorial inference: ADE20K-Seq groups static ADE20K images by category rather than by true temporal continuity; a benchmark built from consecutive video frames would test whether the method's gains persist when appearance changes are large and motion blur occurs."],"forward_implications":["On the paper's reported numbers, a user labeling a series of same-category images should need fewer clicks per image: NoC85 and NoC90 drop relative to RITM, FocalClick, SimpleClick, SAM, HQ-SAM, and the other baselines.","Because the gain is largest at one or two clicks (the M-IoU curves in Figure 4), the method is most valuable in the low-interaction regime where single-image models fail hardest.","Longer prompt sequences improve accuracy up to the tested length of ten prompts, so practitioners can trade memory for precision by increasing sequence length.","The TPS module transfers across datasets: it is trained on COCO and LVIS and evaluated on GrabCut, Berkeley, COCO-MVal, DAVIS, and ADE20K-Seq, so similarity-based prompt retrieval does not need dataset-specific retraining.","The new ADE20K-Seq benchmark gives the community a fixed seven-category sequence test set on which future sequence-aware interactive segmentation methods can be compared."],"supporting_citations":[{"why":"Supplies the DINOv2 feature extractor whose similarity scores drive the Top-k Prompt Selection module.","marker":"[40]"},{"why":"Provides the click-simulation interaction strategy used to mimic iterative user clicks and is a primary baseline the paper must beat.","marker":"[1]"},{"why":"Provides the ViT-based backbone design for interactive segmentation and is a state-of-the-art baseline in the evaluation.","marker":"[8]"},{"why":"FocalClick is an iterative click baseline that SPT claims to surpass on the benchmark datasets.","marker":"[9]"},{"why":"SAM is the zero-shot segmentation baseline whose quality SPT claims to exceed on sequence segmentation tasks.","marker":"[11]"},{"why":"The Vision Transformer architecture used as the backbone for feature extraction from each image, click, and mask input.","marker":"[15]"},{"why":"COCO is one of the two training datasets (118k images and 1.2M instances) on which SPT is trained.","marker":"[35]"},{"why":"LVIS is the second training dataset (100k images and 1.2M instances) used to train the model.","marker":"[36]"},{"why":"ADE20K is the source dataset from which the new ADE20K-Seq benchmark of seven same-category sequences is derived.","marker":"[14]"}],"fun_headline_variants":["Sequence-aware transformer cuts clicks for interactive segmentation","Using past images as prompts improves interactive segmentation","Sequence Prompt Transformer: fewer clicks via sequential context","Causal transformer uses past images to cut interactive clicks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's load-bearing premise is that DINOv2 feature similarity between images is a reliable proxy for how useful an earlier image, click, and mask will be as a prompt for the current image; the paper's ablation for this premise is referenced as ExpID #6–#9 but the table itself is missing from the manuscript.","fun_headline_variants_meta":{"raw":{"variants":["Sequence-aware transformer cuts clicks for interactive segmentation","Using past images as prompts improves interactive segmentation","Sequence Prompt Transformer: fewer clicks via sequential context","Causal transformer uses past images to cut interactive clicks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000582,"raw_usage":{"total_tokens":2697,"prompt_tokens":860,"completion_tokens":1837,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":1779}},"tokens_in":476,"tokens_out":1837,"duration_ms":12939,"temperature":1.0,"reasoning_tokens":1779,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:11:42.998718+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run SPT on ADE20K-Seq with the TPS module removed (e.g., using the k most recent images as prompts instead of the DINOv2-selected ones) and compare NoC85/NoC90 and M-IoU; if the DINOv2 selection margin disappears, reverses, or shrinks to statistical noise, the claim that Top-k Prompt Selection carries the improvement is false.","supporting_citations":[{"cited_title":"Reviving iterative training with mask guidance for interactive segmentation,","cited_arxiv_id":null,"evidence_quote":"Provides the click-simulation interaction strategy used to mimic iterative user clicks and is a primary baseline the paper must beat."},{"cited_title":"Simpleclick: Interactive image segmentation with simple vision transformers,","cited_arxiv_id":null,"evidence_quote":"Provides the ViT-based backbone design for interactive segmentation and is a state-of-the-art baseline in the evaluation."},{"cited_title":"Focalclick: Towards practical interactive image segmentation,","cited_arxiv_id":null,"evidence_quote":"FocalClick is an iterative click baseline that SPT claims to surpass on the benchmark datasets."}],"review_version":1}