REVIEW 4 major objections 2 minor 39 references
Aligning Moments in Time using Video Queries
T0 review · 4 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A video-query transformer claims state-of-the-art moment retrieval, but the manuscript body is an unrelated pulsar study.
desk verdict Abstract claims SOTA video moment retrieval; the body is an unrelated radio pulsar paper — nothing in the submission supports the abstract. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
MATR — the Moment Alignment Transformer. The load-bearing components are dual-stage sequence alignment, which conditions target-video features on query-video features to capture cross-video dependencies, and a self-supervised pretraining task where the model localizes random clips inside videos, followed by foreground/background classification and boundary-prediction heads. In the supplied full text none of these components appear; they exist only in the abstract's description.
What would settle it
Inspect the full text of the submission: it is titled 'Radio Observations of a Candidate Redback Millisecond Pulsar: 1FGL J0523.5−2529', contains telescope observation tables, and never mentions MATR, dual-stage sequence alignment, ActivityNet-VRL, or SportsMoments. No further computation is needed to confirm that the abstract's claims are unsupported by the provided document.
Extended reading notes
Core claim
The paper's stated discovery is that MATR, a transformer that conditions target-video representations on query-video features through dual-stage sequence alignment, improves moment localization by 13.1% in R@1 and 8.1% in mIoU absolute over state-of-the-art on ActivityNet-VRL, and 14.7% in R@1 and 14.4% in mIoU on the new SportsMoments dataset, using a self-supervised clip-localization pretraining. The full-text document provided for this submission, however, is an unrelated observational study of a candidate redback millisecond pulsar; it does not describe MATR, the alignment mechanism, the pretraining, either dataset, or any of the reported numbers. The author's intended assertion is there
Load-bearing premise
The central claim depends on the manuscript body being the paper described by the abstract, but the body is an unrelated radio-astronomy study titled 'Radio Observations of a Candidate Redback Millisecond Pulsar: 1FGL J0523.5−2529' — that premise fails on inspection.
Editorial extensions
If this is right
- Video-to-video moment retrieval would become practical: users could query by example clip rather than by text, capturing actions words cannot describe.
- Self-supervised clip-localization pretraining would transfer to a harder cross-video localization task, reducing the need for labeled video-moment data.
- The reported gains — 13.1% R@1 and 8.1% mIoU on ActivityNet-VRL; 14.7% R@1 and 14.4% mIoU on SportsMoments — would set a new state of the art on both benchmarks.
- The new SportsMoments dataset would provide a sports-specific benchmark for video-query retrieval, complementing ActivityNet-VRL.
Reading between the lines
- The abstract and the body appear to come from different submissions; if so, the abstract's experimental numbers can only be taken as unsupported until the matching manuscript is located and audited.
- Independently of this mismatched document, the claim that self-supervised clip localization transfers to video-to-video moment retrieval is a testable hypothesis: one could pretrain a transformer on clip localization, fine-tune on a labeled video-retrieval benchmark, and measure transfer on held-out domains.
- A reader seeking the method would need to find the actual MATR paper separately; the currently supplied full text provides no implementation details, architecture diagram, or hyperparameter settings to reproduce.
- The reported performance gaps over prior art are large enough that, if verified, they would likely drive adoption of dual-stage cross-video conditioning in other video-understanding tasks such as action localization and dense video captioning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission claims to introduce MATR, a transformer-based model for video-to-video moment retrieval, with abstract-level results of 13.1% R@1 and 8.1% mIoU absolute improvement over state-of-the-art on ActivityNet-VRL and 14.7% R@1 / 14.4% mIoU on a new SportsMoments dataset. The full text, however, is an unrelated radio-astronomy manuscript, 'Radio Observations of a Candidate Redback Millisecond Pulsar: 1FGL J0523.5-2529', whose author list, abstract, Introduction, Methods, Results, and references concern pulsar timing observations. No MATR architecture, no dual-stage sequence alignment, no pre-training objective, no dataset description, no experimental protocol, and no results tables appear anywhere in the body.
Significance. If the abstract claims were accompanied by the described model, experiments, and dataset, MATR would represent a substantial advance in video-to-video moment retrieval. The claimed absolute gains of 13.1% R@1 and 8.1% mIoU over prior methods are large, and a new SportsMoments dataset could be a useful community resource. However, none of these contributions are present in the submitted manuscript. The body contains no method to audit, no code or machine-checked artifacts, and no quantitative evaluation relating to moment retrieval. At the document level, the paper reduces to an abstract-only claim with an unrelated supporting text. The significance cannot be assessed because the contribution itself is absent.
major comments (4)
- [Abstract vs. full text] The central claims of the paper, namely the MATR architecture, dual-stage sequence alignment, self-supervised pre-training, and the ActivityNet-VRL and SportsMoments results, appear only in the abstract. The full text is an astronomy paper about radio observations of the candidate redback pulsar 1FGL J0523.5-2529, as confirmed by the title, author list, abstract, Sections 1-5, references, and the footer 'arXiv:2508.15435v2 [astro-ph.HE]'. There is no moment-retrieval algorithm, no model diagram, no equations describing attention or alignment, and no dataset. The submitted manuscript is therefore not a paper containing the claimed contribution.
- [§4 and Table 3] The only quantitative results in the body are flux-density upper limits for a pulsar non-detection (Table 3) and related detection limits in Eqs. (3)-(4). These are unrelated to the abstract's R@1 and mIoU claims. The 13.1% R@1, 8.1% mIoU, 14.7% R@1, and 14.4% mIoU figures are not supported by any experimental protocol, baseline description, error bars, or comparison table in the manuscript. There is no way to verify, reproduce, or even locate the source of these numbers.
- [Dataset and pre-training] The abstract announces 'our newly proposed dataset, SportsMoments' and a self-supervised clip-localization pre-training technique. Neither is described in the manuscript. The dataset name does not occur in the body, and no annotation statistics, collection procedure, evaluation splits, or baseline results are given. The claim that self-supervised clip localization transfers to video-to-video moment retrieval is thus a bare assertion; no experiments address it.
- [Document integrity] The manuscript's own footer identifies the body as arXiv:2508.15435v2 [astro-ph.HE], an arXiv ID different from the submission ID 2508.15439. Internal cross-references (e.g., Section 2 observations, Figure 1) all refer to telescope observations, not to video retrieval. This is not a presentation issue or a missing appendix; it is an absence of the paper being reviewed.
minor comments (2)
- [Metadata] The arXiv ID and subject class on the footer do not match the submission's stated ID and category; this should be corrected by resubmitting the intended paper.
- [Throughout] The astronomy body has several typographical issues (e.g., 'Trinty' for 'Trinity', 'spliting', 'P b≤10' formatting), but these are irrelevant to the video-retrieval manuscript and are listed only for completeness.
Circularity Check
No circular derivation identified: the supplied body is an unrelated astronomy paper, so there is no method chain to reduce.
full rationale
The abstract describes a video-to-video moment retrieval model (MATR) with transformer architecture, dual-stage sequence alignment, self-supervised pre-training, and quantitative SOTA gains on ActivityNet-VRL and a new SportsMoments dataset. The full text supplied, however, is an unrelated astrophysics manuscript titled 'RADIO OBSERVATIONS OF A CANDIDATE REDBACK MILLISECOND PULSAR: 1FGL J0523.5−2529' (footer arXiv:2508.15435v2 [astro-ph.HE]). It contains no equations for MATR, no alignment module, no pre-training objective, no dataset description, and no results tables. There is therefore no claimed derivation chain that can be walked, and I cannot exhibit any specific reduction in which a prediction is equivalent to its input by construction, a fitted parameter is renamed as a prediction, or a load-bearing conclusion rests solely on a self-citation. The document-level mismatch is a severe absence of supporting content and a correctness/integrity problem, but it is not circularity under the definitions used here. The secondary assumption that self-supervised clip-localization pre-training transfers to video-to-video moment retrieval is untested, but an untested transfer assumption is not a circular step. Score 0 reflects the absence of any identifiable circular step; it should not be read as validating the abstract's claims.
Assumptions & free parameters
assumptions (3)
- domain assumption Semantic frame-level alignment between query and target video frames is a sufficient and learnable signal for precise moment localization.
- domain assumption Self-supervised pretraining, localizing random clips within videos, transfers to the video-to-video moment retrieval task.
- domain assumption The ActivityNet-VRL split and the constructed SportsMoments dataset correctly instantiate the video-to-video retrieval task.
invented entities (1)
-
SportsMoments dataset
Cite this review
Pith. "Pith review of Aligning Moments in Time using Video Queries." pith.science (2026). https://pith.science/paper/ALHHKS3L
@misc{pith2026250815439,
author = {Pith},
title = {Pith review of: Aligning Moments in Time using Video Queries},
year = {2026},
howpublished = {\url{https://pith.science/paper/ALHHKS3L}},
note = {Machine review of arXiv:2508.15439}
}
read the original abstract
Video-to-video moment retrieval (Vid2VidMR) is the task of localizing unseen events or moments in a target video using a query video. This task poses several challenges, such as the need for semantic frame-level alignment and modeling complex dependencies between query and target videos. To tackle this challenging problem, we introduce MATR (Moment Alignment TRansformer), a transformer-based model designed to capture semantic context as well as the temporal details necessary for precise moment localization. MATR conditions target video representations on query video features using dual-stage sequence alignment that encodes the required correlations and dependencies. These representations are then used to guide foreground/background classification and boundary prediction heads, enabling the model to accurately identify moments in the target video that semantically match with the query video. Additionally, to provide a strong task-specific initialization for MATR, we propose a self-supervised pre-training technique that involves training the model to localize random clips within videos. Extensive experiments demonstrate that MATR achieves notable performance improvements of 13.1% in R@1 and 8.1% in mIoU on an absolute scale compared to state-of-the-art methods on the popular ActivityNet-VRL dataset. Additionally, on our newly proposed dataset, SportsMoments, MATR shows a 14.7% gain in R@1 and a 14.4% gain in mIoU on an absolute scale over strong baselines.
Reference graph
Works this paper leans on
-
[1]
A., Ajello, M., Allafort, A., et al
Abdo, A. A., Ajello, M., Allafort, A., et al. 2013, The Astrophysical Journal Supplement Series, 208, 17, doi:10.1088/0067-0049/208/2/17
-
[2]
2012, The Astrophysical Journal, 753, 83, doi:10.1088/0004-637X/753/1/83
Ackermann, M., Ajello, M., Allafort, A., et al. 2012, The Astrophysical Journal, 753, 83, doi:10.1088/0004-637X/753/1/83
-
[3]
D., Thornton, D., Bailes, M., et al
Bates, S. D., Thornton, D., Bailes, M., et al. 2015, MNRAS, 446, 4019, doi:10.1093/mnras/stu2350
-
[4]
Bhattacharya, D., & van den Heuvel, E. P. J. 1991, Phys. Rep., 203, 1, doi:10.1016/0370-1573(91)90064-S
-
[5]
Broderick, J. W., Fender, R. P., Breton, R. P., et al. 2016, Monthly Notices of the Royal Astronomical Society, 459, 2681, doi:10.1093/mnras/stw794
-
[6]
Chen, H.-L., Chen, X., Tauris, T. M., & Han, Z. 2013, ApJ, 775, 27, doi:10.1088/0004-637X/775/1/27
-
[7]
Cordes, J. M., & Lazio, T. J. W. 2003, NE2001.I. A New Model for the Galactic Distribution of Free Electrons and its Fluctuations, arXiv, doi:10.48550/arXiv.astro-ph/0207156 De Vito, M. A., Benvenuto, O. G., & Horvath, J. E. 2020, MNRAS, 493, 2171, doi:10.1093/mnras/staa395
-
[8]
Deneva, J. S., Ray, P. S., Camilo, F., et al. 2016, ApJ, 823, 105, doi:10.3847/0004-637X/823/2/105
Show all 39 references
- [9]
-
[10]
P., Perez, K
Halpern, J. P., Perez, K. I., & Bogdanov, S. 2022, The Astrophysical Journal, 935, 151, doi:10.3847/1538-4357/ac8161
2022 doi
-
[11]
N., Dunning, A., et al
Hobbs, G., Manchester, R. N., Dunning, A., et al. 2020, Publications of the Astronomical Society of Australia, 37, e012, doi:10.1017/pasa.2020.2
2020 doi
-
[12]
F., et al
Jankowski, F., van Straten, W., Keane, E. F., et al. 2018, MNRAS, 473, 4436, doi:10.1093/mnras/stx2476
2018 doi
-
[13]
F., Barr, E
Keane, E. F., Barr, E. D., Jameson, A., et al. 2018, Monthly Notices of the Royal Astronomical Society, 473, 116, doi:10.1093/mnras/stx2126
2018 doi
-
[14]
Y ., Zharikov, S
Kirichenko, A. Y ., Zharikov, S. V ., Karpova, A. V ., et al. 2024, Monthly Notices of the Royal Astronomical Society, 527, 4563, doi:10.1093/mnras/stad3391
2024 doi
-
[15]
Koljonen, K. I. I., & Linares, M. 2025, SpiderCat: A Catalog of Compact Binary Millisecond Pulsars, arXiv, doi:10.48550/arXiv.2505.11691
2025 doi
-
[16]
H., Manchester, R
Kramer, M., Stairs, I. H., Manchester, R. N., et al. 2021, Physical Review X, 11, 041050, doi:10.1103/PhysRevX.11.041050
2021 doi
-
[17]
R., & Kramer, M
Lorimer, D. R., & Kramer, M. 2005, Handbook of Pulsar Astronomy, Cambridge Observing Handbooks for Research Astronomers (Cambridge University Press)
2005
-
[18]
G., Manchester, R
Lyne, A. G., Manchester, R. N., Lorimer, D. R., et al. 1998, MNRAS, 295, 743, doi:10.1046/j.1365-8711.1998.01144.x
1998
-
[19]
N., Hobbs, G
Manchester, R. N., Hobbs, G. B., Teoh, A., & Hobbs, M. 2005, The Astronomical Journal, 129, 1993, doi:10.1086/428488
2005 doi
-
[20]
2024, Astronomy and Astrophysics, 683, A183, doi:10.1051/0004-6361/202348247
Men, Y ., & Barr, E. 2024, Astronomy and Astrophysics, 683, A183, doi:10.1051/0004-6361/202348247
2024 doi
-
[21]
2022, in Astrophysics and Space Science
Papitto, A., & de Martino, D. 2022, in Astrophysics and Space Science
2022
-
[22]
465, Physics and Astrophysics of Neutron Stars (Springer Nature), 157–200, doi:10.1007/978-3-030-85198-9_6
Library, V ol. 465, Physics and Astrophysics of Neutron Stars (Springer Nature), 157–200, doi:10.1007/978-3-030-85198-9_6
-
[23]
2013, Nature, 501, 517, doi:10.1038/nature12470
Papitto, A., Ferrigno, C., Bozzo, E., et al. 2013, Nature, 501, 517, doi:10.1038/nature12470
2013 doi
-
[24]
2011, Astrophysics Source Code Library, ascl:1107.017
Ransom, S. 2011, Astrophysics Source Code Library, ascl:1107.017. https://ui.adsabs.harvard.edu/abs/2011ascl.soft07017R
2011
-
[25]
M., Greenhill, L
Ransom, S. M., Greenhill, L. J., Herrnstein, J. R., et al. 2001, ApJ, 546, L25, doi:10.1086/318062
2001 doi
-
[26]
Roberts, M. S. E., Noori, H. A., Torres, R. A., et al. 2017, Proceedings of the International Astronomical Union, 13, 43, doi:10.1017/S1743921318000480
2017 doi
-
[27]
S., Bhattacharyya, B., et al
Roy, J., Ray, P. S., Bhattacharyya, B., et al. 2015, The Astrophysical Journal Letters, 800, L12, doi:10.1088/2041-8205/800/1/L12
2015 doi
-
[28]
2016, PhD thesis, University of Virginia
Sanpa-Arsa, S. 2016, PhD thesis, University of Virginia
2016
-
[29]
L., Tout, C
Smedley, S. L., Tout, C. A., Ferrario, L., & Wickramasinghe, D. T. 2015, Monthly Notices of the Royal Astronomical Society, 446, 2540, doi:10.1093/mnras/stu2252
2015 doi
-
[30]
W., Archibald, A
Stappers, B. W., Archibald, A. M., Hessels, J. W. T., et al. 2014, The Astrophysical Journal, 790, 39, doi:10.1088/0004-637X/790/1/39
2014 doi
-
[31]
2016, ApJ, 833, 192, doi:10.3847/1538-4357/833/2/192
Stovall, K., Allen, B., Bogdanov, S., et al. 2016, ApJ, 833, 192, doi:10.3847/1538-4357/833/2/192
2016 doi
-
[32]
2014, The Astrophysical Journal, 788, L27, doi:10.1088/2041-8205/788/2/L27
Strader, J., Chomiuk, L., Sonbas, E., et al. 2014, The Astrophysical Journal, 788, L27, doi:10.1088/2041-8205/788/2/L27
2014 doi
- [33]
-
[34]
J., Breton, R
Thongmeearkom, T., Clark, C. J., Breton, R. P., et al. 2024, Monthly Notices of the Royal Astronomical Society, 530, 4676, doi:10.1093/mnras/stae787
2024 doi
-
[35]
M., Manchester, R
Yao, J. M., Manchester, R. N., & Wang, N. 2017, The Astrophysical Journal, 835, 29, doi:10.3847/1538-4357/835/1/29
2017 doi
-
[36]
M., D’Avanzo, P., Ridolfi, A., et al
Zanon, A. M., D’Avanzo, P., Ridolfi, A., et al. 2021, Astronomy & Astrophysics, 649, A120, doi:10.1051/0004-6361/202040071
2021 doi
-
[37]
S., et al
Zheng, H., Tegmark, M., Dillon, J. S., et al. 2017, Monthly Notices of the Royal Astronomical Society, 464, 3486, doi:10.1093/mnras/stw2525
2017 doi
-
[38]
2022, Research in Astronomy and Astrophysics, 22, 085001, doi:10.1088/1674-4527/ac712b
Zhou, Z.-R., Wang, J.-B., Wang, N., Hobbs, G., & Wang, S.-Q. 2022, Research in Astronomy and Astrophysics, 22, 085001, doi:10.1088/1674-4527/ac712b
2022 doi
-
[39]
Zic, A., Wang, Z., Lenc, E., et al. 2024, Monthly Notices of the Royal Astronomical Society, 528, 5730, doi:10.1093/mnras/stae033 6.APPENDIX 6.1.Detection of J0520-2553 As stated in §4 there was a serendipitous detection of J0520-2553 (see Fig. 6), using both the acceleration ...
2024 doi
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.