{"id":"4745dbc1-a62e-4370-9cf5-f6cbc1f84082","arxiv_id":"2508.15439","paper_version":2,"verdict":"REJECT","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The abstract claims a new video-to-video moment retrieval model with state-of-the-art gains, but the manuscript body is an unrelated pulsar observation paper, so the claimed results are absent from the document.","lead":"The abstract describes MATR, a transformer model for matching moments in videos by example, with large claimed accuracy gains and a new benchmark. But the manuscript body is an astronomy paper on searching for radio pulses from a pulsar, so none of the video retrieval claims appear in the paper itself.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract-only SOTA claims unsupported: full text is an unrelated astronomy manuscript, so there is no method or experiment to verify.","rationale":"The paper's stated goal is to introduce MATR and demonstrate SOTA on ActivityNet-VRL and SportsMoments. For the central claim to hold, the body would need to specify the dual-stage sequence alignment mechanism, the self-supervised pretraining objective, the datasets and evaluation protocols, and report experimental numbers. None of these are present. The full text is the astronomy paper 'Radio Observations of a Candidate Redback Millisecond Pulsar: 1FGL J0523.5-2529', with its own abstract and footer 'arXiv:2508.15435v2 [astro-ph.HE]'. Thus the submission is not a case of a marginally supported architecture; it is a document-level mismatch where the abstract and body share no content. The strongest claim is an existential claim about performance improvements; in the available evidence, the numbers are ungrounded. Reviewing in good faith, I did not find a plausible reading on which the current text supports MATR's results. The reader's weakest assumption was the same: that the manuscript body is the paper described by the abstract. I agree. Because the empirical claims are entirely in the abstract, the correct action remains REJECT; if a corrected submission pairs this abstract with the actual MATR manuscript, it should be reviewed fresh. No independent support (code, proofs, data) appears in the provided text.","tokens_in":10881,"tokens_out":3659,"duration_ms":39511,"concrete_test":"Fetch the arXiv e-print source for 2508.15439 (e.g., via export.arxiv.org/e-print/2508.15439) and grep the manuscript source for the abstract's key terms: 'MATR', 'Moment Alignment TRansformer', 'dual-stage', 'ActivityNet-VRL', 'SportsMoments', 'self-supervised' or 'pretraining', and 'video-to-video'. If none of these tokens occur in the body—or if the body's arXiv footer is 2508.15435v2—then the empirical claims in the abstract have no supporting artifact and the submission cannot be evaluated as a CV paper. If the tokens do appear in a separately supplied correct body, the rejection should be revisited.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing concern is that the central claim—that MATR achieves 13.1% R@1 and 8.1% mIoU over SOTA on ActivityNet-VRL, plus 14.7%/14.4% gains on SportsMoments—appears only in the abstract. The complete body is an astronomy preprint ('Radio Observations of a Candidate Redback Millisecond Pulsar: 1FGL J0523.5-2529', footer arXiv:2508.15435v2 [astro-ph.HE]). It contains no moment-retrieval architecture, no dual-stage sequence alignment, no pretraining objective, no dataset description, no results tables, and no evaluation protocol. Under the review rule that all manuscript text is in-scope evidence, the document's own footer identifies the body as a different paper. The claimed quantitative results therefore have no verifiable derivation; the submission reduces to claim-without-support at the document level. This is not an internal technical flaw in a proposed model; it is the absence of the proposed model itself. A secondary design assumption—that self-supervised clip-localization pretraining transfers to video-to-video moment retrieval—remains untested, but it is secondary because the body never reaches it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission claims to introduce MATR, a transformer-based model for video-to-video moment retrieval, with abstract-level results of 13.1% R@1 and 8.1% mIoU absolute improvement over state-of-the-art on ActivityNet-VRL and 14.7% R@1 / 14.4% mIoU on a new SportsMoments dataset. The full text, however, is an unrelated radio-astronomy manuscript, 'Radio Observations of a Candidate Redback Millisecond Pulsar: 1FGL J0523.5-2529', whose author list, abstract, Introduction, Methods, Results, and references concern pulsar timing observations. No MATR architecture, no dual-stage sequence alignment, no pre-training objective, no dataset description, no experimental protocol, and no results tables appear anywhere in the body.","tokens_in":11029,"tokens_out":1793,"duration_ms":21164,"significance":"If the abstract claims were accompanied by the described model, experiments, and dataset, MATR would represent a substantial advance in video-to-video moment retrieval. The claimed absolute gains of 13.1% R@1 and 8.1% mIoU over prior methods are large, and a new SportsMoments dataset could be a useful community resource. However, none of these contributions are present in the submitted manuscript. The body contains no method to audit, no code or machine-checked artifacts, and no quantitative evaluation relating to moment retrieval. At the document level, the paper reduces to an abstract-only claim with an unrelated supporting text. The significance cannot be assessed because the contribution itself is absent.","major_comments":[{"comment":"The central claims of the paper, namely the MATR architecture, dual-stage sequence alignment, self-supervised pre-training, and the ActivityNet-VRL and SportsMoments results, appear only in the abstract. The full text is an astronomy paper about radio observations of the candidate redback pulsar 1FGL J0523.5-2529, as confirmed by the title, author list, abstract, Sections 1-5, references, and the footer 'arXiv:2508.15435v2 [astro-ph.HE]'. There is no moment-retrieval algorithm, no model diagram, no equations describing attention or alignment, and no dataset. The submitted manuscript is therefore not a paper containing the claimed contribution.","section":"Abstract vs. full text"},{"comment":"The only quantitative results in the body are flux-density upper limits for a pulsar non-detection (Table 3) and related detection limits in Eqs. (3)-(4). These are unrelated to the abstract's R@1 and mIoU claims. The 13.1% R@1, 8.1% mIoU, 14.7% R@1, and 14.4% mIoU figures are not supported by any experimental protocol, baseline description, error bars, or comparison table in the manuscript. There is no way to verify, reproduce, or even locate the source of these numbers.","section":"§4 and Table 3"},{"comment":"The abstract announces 'our newly proposed dataset, SportsMoments' and a self-supervised clip-localization pre-training technique. Neither is described in the manuscript. The dataset name does not occur in the body, and no annotation statistics, collection procedure, evaluation splits, or baseline results are given. The claim that self-supervised clip localization transfers to video-to-video moment retrieval is thus a bare assertion; no experiments address it.","section":"Dataset and pre-training"},{"comment":"The manuscript's own footer identifies the body as arXiv:2508.15435v2 [astro-ph.HE], an arXiv ID different from the submission ID 2508.15439. Internal cross-references (e.g., Section 2 observations, Figure 1) all refer to telescope observations, not to video retrieval. This is not a presentation issue or a missing appendix; it is an absence of the paper being reviewed.","section":"Document integrity"}],"minor_comments":[{"comment":"The arXiv ID and subject class on the footer do not match the submission's stated ID and category; this should be corrected by resubmitting the intended paper.","section":"Metadata"},{"comment":"The astronomy body has several typographical issues (e.g., 'Trinty' for 'Trinity', 'spliting', 'P b≤10' formatting), but these are irrelevant to the video-retrieval manuscript and are listed only for completeness.","section":"Throughout"}],"recommendation":"reject","confidential_remarks":"This appears to be a submission error: the uploaded PDF contains a different paper than the one described in the abstract. I recommend rejecting this version and inviting the authors to resubmit the correct manuscript. If this is instead a deliberate substitution, that is a much more serious matter that the editor should treat accordingly. Either way, the current submission cannot be sent for technical review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — you should know that arXiv:2508.15439 as submitted pairs a computer vision abstract with a full text that is a radio pulsar paper. The footer identifies arXiv:2508.15435v2 [astro-ph.HE]. There is no MATR architecture, no dual-stage alignment, no pretraining, no experiments, no dataset. Every quantitative claim lives in the abstract alone.\n\nWhat is new in the abstract — the Vid2VidMR framing, the transformer, the self-supervised objective, SportsMoments — is plausible-sounding but nothing in the manuscript supports it. As it stands, there is no method and no evidence. The astronomy body is a real paper, and maybe a good one, but it is not the paper described. That is a load-bearing flaw at the document level, not a missing appendix.\n\nThe reader's weakest assumption — that the provided body is the paper described by the abstract — fails on inspection. The reader's verdict of reject with low confidence is right, though I would say the confidence should be higher: the mismatch is verifiable by reading the first page.\n\nOne secondary point: even if a correct body were supplied, the abstract's assumption that clip-localization pretraining transfers to video-to-video retrieval is unargued. That is a genuine design question, but it is minor compared to the missing manuscript.\n\nThere is no formal verification, no shipped code or data, no reproducible experiment. The citation pattern is moot because the body is unrelated. So there is nothing to credit in this submission.\n\nWho gets value from this? Nobody, as it stands. It should be desk rejected, not sent to referees. If the authors later upload the intended full text, that version deserves a normal review. But this artifact does not.","headline":"Abstract claims SOTA video moment retrieval; the body is an unrelated radio pulsar paper — nothing in the submission supports the abstract.","tokens_in":11641,"tokens_out":2036,"would_cite":false,"duration_ms":20576,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A video-query transformer claims state-of-the-art moment retrieval, but the manuscript body is an unrelated pulsar study.","keywords":["video-to-video moment retrieval","Moment Alignment Transformer","MATR","dual-stage sequence alignment","self-supervised pretraining","ActivityNet-VRL","SportsMoments"],"falsifier":"Inspect the full text of the submission: it is titled 'Radio Observations of a Candidate Redback Millisecond Pulsar: 1FGL J0523.5−2529', contains telescope observation tables, and never mentions MATR, dual-stage sequence alignment, ActivityNet-VRL, or SportsMoments. No further computation is needed to confirm that the abstract's claims are unsupported by the provided document.","tokens_in":10684,"feed_emoji":"🎥","tokens_out":5606,"duration_ms":52366,"temperature":0.7,"pith_summary":"The abstract claims that a new architecture called MATR (Moment Alignment Transformer) solves video-to-video moment retrieval: given a query video clip, it finds the semantically matching segment in a target video, and it reports large gains over prior methods on two benchmarks. If true, this would let users search video by example rather than by text. The accompanying manuscript, however, is a radio-astronomy paper titled 'Radio Observations of a Candidate Redback Millisecond Pulsar: 1FGL J0523.5−2529'; it contains no architecture, experiments, or dataset related to MATR. Read in good faith, the abstract's claim is unverifiable from the provided text, and the load-bearing premise — that the body belongs to the abstract — fails on inspection.","feed_headline":"Video-query model claims record gains; full text is a different paper","feed_subtitle":"The abstract reports 13.1% R@1 gains over prior art, but no architecture, experiments, or dataset appear in the text.","key_machinery":"MATR — the Moment Alignment Transformer. The load-bearing components are dual-stage sequence alignment, which conditions target-video features on query-video features to capture cross-video dependencies, and a self-supervised pretraining task where the model localizes random clips inside videos, followed by foreground/background classification and boundary-prediction heads. In the supplied full text none of these components appear; they exist only in the abstract's description.","core_discovery":"The paper's stated discovery is that MATR, a transformer that conditions target-video representations on query-video features through dual-stage sequence alignment, improves moment localization by 13.1% in R@1 and 8.1% in mIoU absolute over state-of-the-art on ActivityNet-VRL, and 14.7% in R@1 and 14.4% in mIoU on the new SportsMoments dataset, using a self-supervised clip-localization pretraining. The full-text document provided for this submission, however, is an unrelated observational study of a candidate redback millisecond pulsar; it does not describe MATR, the alignment mechanism, the pretraining, either dataset, or any of the reported numbers. The author's intended assertion is there","pith_inferences":["The abstract and the body appear to come from different submissions; if so, the abstract's experimental numbers can only be taken as unsupported until the matching manuscript is located and audited.","Independently of this mismatched document, the claim that self-supervised clip localization transfers to video-to-video moment retrieval is a testable hypothesis: one could pretrain a transformer on clip localization, fine-tune on a labeled video-retrieval benchmark, and measure transfer on held-out domains.","A reader seeking the method would need to find the actual MATR paper separately; the currently supplied full text provides no implementation details, architecture diagram, or hyperparameter settings to reproduce.","The reported performance gaps over prior art are large enough that, if verified, they would likely drive adoption of dual-stage cross-video conditioning in other video-understanding tasks such as action localization and dense video captioning."],"forward_implications":["Video-to-video moment retrieval would become practical: users could query by example clip rather than by text, capturing actions words cannot describe.","Self-supervised clip-localization pretraining would transfer to a harder cross-video localization task, reducing the need for labeled video-moment data.","The reported gains — 13.1% R@1 and 8.1% mIoU on ActivityNet-VRL; 14.7% R@1 and 14.4% mIoU on SportsMoments — would set a new state of the art on both benchmarks.","The new SportsMoments dataset would provide a sports-specific benchmark for video-query retrieval, complementing ActivityNet-VRL."],"supporting_citations":[],"fun_headline_variants":["MATR's moment gains don't appear in paper text","Video moment retrieval claims 13% but text is pulsar study","Abstract reports new SOTA; full text is unrelated","MATR transformer results missing from provided text","Paper claims big gains, but body is about a pulsar"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The central claim depends on the manuscript body being the paper described by the abstract, but the body is an unrelated radio-astronomy study titled 'Radio Observations of a Candidate Redback Millisecond Pulsar: 1FGL J0523.5−2529' — that premise fails on inspection.","fun_headline_variants_meta":{"raw":{"variants":["MATR's moment gains don't appear in paper text","Video moment retrieval claims 13% but text is pulsar study","Abstract reports new SOTA; full text is unrelated","MATR transformer results missing from provided text","Paper claims big gains, but body is about a pulsar"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1384,"prompt_tokens":811,"completion_tokens":573,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":493}},"tokens_in":555,"tokens_out":573,"duration_ms":5922,"temperature":1.0,"reasoning_tokens":493,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:53:52.310148+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the full text of the submission: it is titled 'Radio Observations of a Candidate Redback Millisecond Pulsar: 1FGL J0523.5−2529', contains telescope observation tables, and never mentions MATR, dual-stage sequence alignment, ActivityNet-VRL, or SportsMoments. No further computation is needed to confirm that the abstract's claims are unsupported by the provided document.","supporting_citations":[],"review_version":1}