Pith. sign in

REVIEW 4 major objections 2 minor 78 references

BSN-II: The First Light Curve Study of Eight Total Eclipsing Contact Binary Stars with Shallow Fillout Factors

T0 review · 4 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper claims that eight total-eclipsing W UMa contact binaries have newly measured ephemerides and period changes, four increasing and four decreasing, with shallow fillout factors.

desk verdict The abstract promises a useful eight-target contact-binary catalog, but the supplied full text is a different paper, so the actual claims are currently unverifiable. read the letter →

arxiv 2508.11901 v1 pith:DL4KL2WP submitted 2025-08-16 astro-ph.SR

classification astro-ph.SR
keywords WUrsaeMajorisbinariescontacteclipsingO-Cdiagramorbitalperiodchangelight-curvemodelingGaiaDR3parallaxfilloutfactor
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to give the first full empirical account of eight total-eclipsing W Ursae Majoris contact binaries, combining fresh ground-based eclipse timings with orbital-period analysis, light-curve modeling, and Gaia DR3 distances. Its central result is that four of the eight binaries have orbits that are slowly growing while four are shrinking, and that all eight have a shallow fillout factor, meaning their shared envelopes barely overflow the Roche lobe. If the result holds, these systems gain measured ephemerides, period-change rates, geometric parameters, and absolute masses and radii, turning them into test points for how contact binaries exchange mass and lose angular momentum as they evolve.

What carries the argument

The O-C diagram is the central instrument: it plots measured times of minimum light against times predicted by a constant-period ephemeris, and a parabolic trend in the residuals yields the period-change rate $dP/dt$. The fillout factor, the fractional amount by which the common envelope overflows the inner Roche lobe, is the geometric quantity used to establish the shallow-contact configuration; PHOEBE light-curve modeling supplies the geometry and the required cold starspots, and the Gaia DR3 parallax turns the dimensionless light-curve solution into absolute masses and radii.

What would settle it

Take one reported period-increasing and one period-decreasing system, gather all archived eclipse minima plus five to ten years of new ones, and compare three fits: a pure parabola, a parabola plus a sinusoidal light-time term, and a linear ephemeris. If the linear fit is statistically adequate, or the sinusoidal term removes the parabolic curvature, the claimed secular $dP/dt$ values and derived mass-loss totals are not established.

Watch

Extended reading notes

Core claim

The paper claims that, for the first time, these eight total-eclipsing W UMa contact binaries have been studied as a set: new times of minimum light were measured from multiband photometry, and the O-C residuals relative to a linear ephemeris were fitted so that four systems show a long-term period increase and four a long-term decrease. Light curves were modeled with the PHOEBE code, two systems needing a cold starspot for an adequate fit, and all systems were found to have a shallow fillout factor. Combining the light-curve solutions with Gaia DR3 parallaxes gives absolute masses, radii, and luminosities, which place the systems on mass-radius and mass-luminosity diagrams and allow A- vers

Load-bearing premise

The four-up, four-down period-change result rests on the assumption that the eclipse-timing residuals are genuine steady orbital changes, not scatter from sparse sampling, starspot-induced timing shifts, or the light-time effect of an unseen third star.

Editorial extensions

If this is right

  • The eight systems now have published ephemerides and period-change rates usable as baselines for future eclipse-timing monitoring.
  • The four increasing and four decreasing periods imply the binaries are not static: material is being exchanged or lost, with the derived total mass lost quantifying the amount.
  • A shallow fillout factor places these systems near the boundary between contact and broken-contact configurations, making them relevant tests of thermal relaxation oscillation models.
  • With Gaia DR3 distances, the measured masses, radii, and luminosities can be checked against stellar evolution tracks on mass-radius and mass-luminosity diagrams.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next test is to watch whether the sign and magnitude of each $dP/dt$ persist over another decade of eclipse timings; if residuals flatten, the reported secular changes would instead be a longer-timescale oscillation or a spot effect.
  • Combining radial-velocity monitoring with the O-C period changes would let observers compare the spectroscopically measured mass-transfer rate with the rate implied by $dP/dt$, directly testing the mass-loss totals.
  • Extending the same O-C plus light-curve pipeline to a larger sample of shallow-contact binaries could reveal whether period-change sign correlates with fillout factor or mass ratio, separating thermal relaxation cycles from genuine angular-momentum loss.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 2 minor

Summary. The manuscript, arXiv:2508.11901, is advertised as an observational study of eight total-eclipsing W UMa contact binaries, reporting new photometry, light-curve modeling with PHOEBE, O-C period-change analysis, Gaia DR3-based absolute parameters, starspot modeling for two systems, shallow fillout factors, A-/W-subtype classification, and mass-loss estimates. However, the supplied full text is not this paper: it is arXiv:2508.11903, a computer-vision paper on online video grounding with hybrid-modal queries, with a different title, abstract, author list, and references. None of the astronomy content described in the abstract — target list, observations, times of minima, PHOEBE parameters, O-C residuals, dP/dt values, spot parameters, parallaxes, or mass-radius/mass-luminosity diagrams — appears anywhere in the submitted text. The central claims are therefore unverifiable from the provided manuscript.

Significance. If the abstract's claims are correct, the paper would provide useful additions to the empirical database on shallow-contact W UMa binaries: new ephemerides, period-change rates, absolute masses and radii, and mass-loss estimates for eight systems. Such studies are of interest to the close-binary community, particularly for testing thermal-relaxation oscillation and angular-momentum-loss scenarios. However, because the submitted full text is a different, unrelated paper, there is no checkable derivation, no tabulated data, no modeling detail, and no comparison with previous work. No machine-checked proofs, reproducible code, or falsifiable quantitative predictions can be assessed. The significance of the astronomy results therefore cannot currently be evaluated; what can be evaluated is only the abstract, which is insufficient for a journal review.

major comments (4)
  1. [Full text (supplied manuscript)] The supplied full text is 'OVG-HQ: Online Video Grounding with Hybrid-modal Queries' (arXiv:2508.11903), a cs.CV paper with no relation to the BSN-II abstract. None of the eight target binaries, observed times of minima, PHOEBE/BSN light-curve fits, starspot parameters, O-C residuals, dP/dt values, Gaia DR3 parallaxes, or absolute parameter tables appear. The abstract is the only astronomy content, and it contains no quantitative support. The central claim of this submission cannot be checked from the manuscript as provided.
  2. [Abstract: O-C analysis] The headline result — four systems with long-term period increases and four with decreases — is asserted without any O-C table, eclipse-timing residuals, fitted ephemeris coefficients, uncertainties, or number of minima. It is therefore impossible to assess whether the timing residuals represent genuine secular period changes or could be affected by sparse sampling, spot-induced eclipse-timing variations, or light-travel-time effects. This is load-bearing for the derived mass-transfer/mass-loss conclusions and is unsupported in the supplied text.
  3. [Abstract: light-curve modeling and fillout] The claim that the targets have shallow fillout factors and that two systems require a cold starspot is not accompanied by any model parameters, spot properties, inclination, mass ratio, temperature ratio, or goodness-of-fit statistics. No spot-degeneracy or parameter-correlation discussion is present. The reader cannot verify the fit quality, the adopted effective temperatures, or the A-/W-subtype classification, all of which are central to the absolute parameter derivation.
  4. [Abstract: absolute parameters and mass loss] The absolute parameters are said to be derived from Gaia DR3 parallax and 'astrophysics equations,' and the initial masses and total mass lost are claimed for each system. No equations, input quantities, evolutionary assumptions, or uncertainty propagation are provided. In particular, the 'total mass lost' is a model-dependent quantity that presupposes initial masses and a mass-loss scenario; without the actual equations and adopted values, the quoted results cannot be reproduced or checked.
minor comments (2)
  1. [General] The abstract does not name the eight target systems or list their coordinates/identifiers, making even the observational scope unreviewable.
  2. [Full text references] The reference list belongs to the video-grounding paper and contains no astronomy citations; all in-text references are to computer-vision literature. The manuscript therefore lacks the required literature context for contact-binary period changes, light-curve modeling conventions, and Gaia DR3 parallax usage.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity demonstrable; the supplied full text is an unrelated cs.CV paper, so the BSN-II derivation chain is absent rather than circular.

full rationale

No circular step can be exhibited from the available material. The BSN-II abstract describes standard, independent observational inputs (ground-based photometry, new times of minima, Gaia DR3 parallaxes) and standard analyses (O-C ephemeris fitting, PHOEBE light-curve modeling, astrophysical parameter equations). The O-C period-change rates are fitted quantities reported from timing residuals; they are not predictions forced by construction, by self-citation, or by definition. The 'initial masses' and 'total mass lost' are model-dependent derived quantities, but model dependence is not circularity. The supplied full text is actually arXiv:2508.11903, an unrelated computer-vision paper on online video grounding, so none of the BSN-II claims can be checked against a derivation in the provided text. That is a missing-evidence problem, not a circularity problem. No self-citation chain, imported uniqueness theorem, or renamed known result is visible in the abstract. Under the rule that circularity requires quoting a specific reduction, the honest finding is no circularity (score 0).

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

Ledger compiled from the abstract only, because the supplied full text belongs to a different paper. All fitted values are per-system numbers reported in the (unavailable) tables, so entries list the model degrees of freedom rather than numeric fits. No new physical entities are introduced; starspots are standard modeling elements.

free parameters (4)
  • fillout factor f (per system)
    Abstract reports all eight systems have a shallow fillout factor; in PHOEBE this is a fitted shape parameter of the Roche common-envelope model, not an independent measurement.
  • PHOEBE light-curve parameters (inclination, mass ratio, temperature ratio) per system
    Standard fitted degrees of freedom implied by 'light curves were analyzed using the PHOEBE Python code'; correlations between these and the fillout/spot parameters are a known degeneracy in contact binaries.
  • cold starspot parameters (temperature factor, angular radius, position) for two systems
    Abstract states two systems required inclusion of a cold starspot to achieve an adequate fit; spot parameters are added free degrees of freedom that trade against temperature ratio and geometry.
  • O-C ephemeris coefficients, including the quadratic (period-change) term per system
    Period increases/decreases are obtained by fitting polynomials to eclipse timing residuals; the dP/dt values are fitted, so the claimed mass transfer/loss rates inherit their uncertainty.
assumptions (5)
  • domain assumption Adopted effective temperatures for each system (from photometric calibrations in prior literature) are accurate.
    Absolute parameters from Gaia DR3 parallax and the astrophysics equations scale with the adopted temperatures; entered before the absolute parameter step in the abstract.
  • domain assumption The PHOEBE/BSN model family (Roche geometry, atmosphere intensities, spot parameterization) adequately and uniquely represents each light curve.
    The shallow-fillout and spot conclusions depend on fit adequacy and on the absence of strong parameter degeneracies; degeneracy among mass ratio, inclination, and fillout is the classic hazard in W UMa modeling.
  • domain assumption O-C timing residuals are dominated by genuine secular period changes, not by sampling artifacts, spot migration, or a light-time effect from a third body.
    Underlies the four-increase/four-decrease result; a short or uneven baseline of ground-based minima can mimic quadratic trends.
  • domain assumption Gaia DR3 parallaxes are reliable and convertible to distances for these binaries.
    Absolute calibration of radii and masses is anchored to 'the Gaia DR3 parallax' per the abstract.
  • domain assumption The adopted stellar evolutionary scenario (tracks and mass-transfer assumptions) is correct for inferring initial masses and total mass lost.
    The abstract's initial masses and mass-loss totals are model-dependent inferences, not measured quantities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of BSN-II: The First Light Curve Study of Eight Total Eclipsing Contact Binary Stars with Shallow Fillout Factors." pith.science (2026). https://pith.science/paper/DL4KL2WP

@misc{pith2026250811901,
  author       = {Pith},
  title        = {Pith review of: BSN-II: The First Light Curve Study of Eight Total Eclipsing Contact Binary Stars with Shallow Fillout Factors},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DL4KL2WP}},
  note         = {Machine review of arXiv:2508.11901}
}
read the original abstract

This study provides the first comprehensive analysis of eight total-eclipse W Ursae Majoris-type contact binary systems. Ground-based photometric multiband observations were conducted at a Mexican observatory, and new times of minima were extracted. The O-C analysis reveals that four of our target binaries exhibit a long-term increase in their orbital periods, while the others show a long-term decrease in their orbital periods. We analyzed the light curves using the PHOEBE Python code and BSN application. Among the target systems, two required the inclusion of a cold starspot on one of the components to achieve an adequate fit. The light curve analysis revealed that the target systems exhibit a shallow fillout factor. Absolute parameters were estimated using the Gaia DR3 parallax and astrophysics equations. Considering the effective temperatures and component masses, each system was classified as either the A- or W-subtype. The stellar evolution of the systems was represented through the mass-radius and mass-luminosity diagrams. Additionally, we calculated the initial masses of the companion stars and the total mass lost for each target system.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

78 extracted references · 72 canonical work pages

  1. [1]

    Activitynet: A large-scale video benchmark for human activity understanding

    Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceed- ings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015. 7

  2. [2]

    On pursuit of designing multi-modal trans- former for video grounding

    Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, and Yuexian Zou. On pursuit of designing multi-modal trans- former for video grounding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Pro- cessing (EMNLP), pages 9810–9823, 2021. 2

  3. [3]

    Temporally grounding natural sentence in video

    Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat- Seng Chua. Temporally grounding natural sentence in video. In Proceedings of the 2018 conference on empirical methods in natural language processing (EMNLP) , pages 162–171,

  4. [4]

    Gatehub: Gated history unit with background sup- pression for online action detection

    Junwen Chen, Gaurav Mittal, Ye Yu, Yu Kong, and Mei Chen. Gatehub: Gated history unit with background sup- pression for online action detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19925–19934, 2022. 7

  5. [5]

    Rethinking the bottom-up frame- work for query-based video localization

    Long Chen, Chujie Lu, Siliang Tang, Jun Xiao, Dong Zhang, Chilie Tan, and Xiaolin Li. Rethinking the bottom-up frame- work for query-based video localization. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 10551– 10558, 2020. 2

  6. [6]

    Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 24185–24198, 2024. 6, 3

  7. [7]

    Moment detection in long tutorial videos

    Ioana Croitoru, Simion-Vlad Bogolin, Samuel Albanie, Yang Liu, Zhaowen Wang, Seunghyun Yoon, Franck Der- noncourt, Hailin Jin, and Trung Bui. Moment detection in long tutorial videos. In Proceedings of the IEEE/CVF in- ternational conference on computer vision (ICCV) , pages 2594–2604, 2023. 2

  8. [8]

    Rethinking weakly-supervised video temporal grounding from a game perspective

    Xiang Fang, Zeyu Xiong, Wanlong Fang, Xiaoye Qu, Chen Chen, Jianfeng Dong, Keke Tang, Pan Zhou, Yu Cheng, and Daizong Liu. Rethinking weakly-supervised video temporal grounding from a game perspective. InEuropean Conference on Computer Vision. Springer, 2024. 2

Show all 78 references
  1. [9]

    Building a digital twin for network optimization using graph neural networks

    Miquel Ferriol-Galm ´es, Jos ´e Su ´arez-Varela, Jordi Pailliss ´e, Xiang Shi, Shihan Xiao, Xiangle Cheng, Pere Barlet-Ros, and Albert Cabellos-Aparicio. Building a digital twin for network optimization using graph neural networks. Com- puter Networks, 217:109329, 2022. 7

  2. [10]

    Temporal sentence grounding in streaming videos

    Tian Gan, Xiao Wang, Yan Sun, Jianlong Wu, Qingpei Guo, and Liqiang Nie. Temporal sentence grounding in streaming videos. In Proceedings of the 31st ACM International Con- ference on Multimedia (ACM MM), pages 4637–4646, 2023. 3, 7

  3. [11]

    Tall: Temporal activity localization via language query

    Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. Tall: Temporal activity localization via language query. In Proceedings of the IEEE international conference on com- puter vision (ICCV), pages 5267–5275, 2017. 1, 2

  4. [12]

    Mac: Mining activity concepts for language-based temporal local- ization

    Runzhou Ge, Jiyang Gao, Kan Chen, and Ram Nevatia. Mac: Mining activity concepts for language-based temporal local- ization. In 2019 IEEE winter conference on applications of computer vision (WACV), pages 245–253. IEEE, 2019. 2

  5. [13]

    Excl: Extractive clip localization us- ing natural language descriptions

    Soham Ghosh, Anuva Agarwal, Zarana Parekh, and Alexan- der G Hauptmann. Excl: Extractive clip localization us- ing natural language descriptions. In Proceedings of the 2019 Conference of the North American Chapter of the As- sociation for Computational Linguistics (NAACL): Hum...

  6. [14]

    Understanding the diffi- culty of training deep feedforward neural networks

    Xavier Glorot and Yoshua Bengio. Understanding the diffi- culty of training deep feedforward neural networks. In Pro- ceedings of the thirteenth international conference on arti- ficial intelligence and statistics (AISTATS) , pages 249–256. JMLR Workshop and Conference Proceed...

  7. [15]

    Minotaur: Multi-task video grounding from multimodal queries

    Raghav Goyal, Effrosyni Mavroudi, Xitong Yang, Sainbayar Sukhbaatar, Leonid Sigal, Matt Feiszli, Lorenzo Torresani, and Du Tran. Minotaur: Multi-task video grounding from multimodal queries. arXiv preprint arXiv:2302.08063, 2023. 3

  8. [16]

    Video activity localisation with uncertainties in temporal boundary

    Jiabo Huang, Hailin Jin, Shaogang Gong, and Yang Liu. Video activity localisation with uncertainties in temporal boundary. In European Conference on Computer Vision (ECCV), pages 724–740. Springer, 2022. 2

  9. [17]

    Knowing where to focus: Event-aware transformer for video grounding

    Jinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon, and Kwanghoon Sohn. Knowing where to focus: Event-aware transformer for video grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 13846–13856, 2023. 2

  10. [18]

    A sliding window scheme for online temporal action localization

    Young Hwi Kim, Hyolim Kang, and Seon Joo Kim. A sliding window scheme for online temporal action localization. In European Conference on Computer Vision, pages 653–669. Springer, 2022. 5

  11. [19]

    Bam-detr: Boundary- aligned moment detection transformer for temporal sentence grounding in videos

    Pilhyeon Lee and Hyeran Byun. Bam-detr: Boundary- aligned moment detection transformer for temporal sentence grounding in videos. In European Conference on Computer Vision (ECCV), pages 220–238. Springer, 2025. 2

  12. [20]

    Detecting moments and highlights in videos via natural language queries

    Jie Lei, Tamara L Berg, and Mohit Bansal. Detecting moments and highlights in videos via natural language queries. Advances in Neural Information Processing Systems (NeurIPS), 34:11846–11858, 2021. 2, 6

  13. [21]

    G2l: Semantically aligned and uniform video grounding via geodesic and game theory

    Hongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li, Zhi- hong Zhu, and Yuexian Zou. G2l: Semantically aligned and uniform video grounding via geodesic and game theory. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 12032–12042, 2023. 2

  14. [22]

    Mo- mentdiff: Generative video moment retrieval from random to real

    Pandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao, Lei Zhang, Yun Zheng, Deli Zhao, and Yongdong Zhang. Mo- mentdiff: Generative video moment retrieval from random to real. Advances in neural information processing systems (NeurIPS), 36, 2024. 2

  15. [23]

    Univtg: Towards unified video- language temporal grounding

    Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shra- man Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan, and Mike Zheng Shou. Univtg: Towards unified video- language temporal grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages...

  16. [24]

    Context-aware biaffine localizing network for temporal sentence ground- ing

    Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, and Yulai Xie. Context-aware biaffine localizing network for temporal sentence ground- ing. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1123...

  17. [25]

    Memory-guided semantic learning net- work for temporal sentence grounding

    Daizong Liu, Xiaoye Qu, Xing Di, Yu Cheng, Zichuan Xu, and Pan Zhou. Memory-guided semantic learning net- work for temporal sentence grounding. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 1665– 1673, 2022. 2, 3

  18. [26]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36:34892–34916, 2023. 2

  19. [27]

    Improved baselines with visual instruction tuning

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024. 2

  20. [28]

    Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

    Ye Liu, Siyuan Li, Yang Wu, Chang-Wen Chen, Ying Shan, and Xiaohu Qie. Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3042–3051, 2022. 2

  21. [29]

    R � -Tuning: Effi- cient Image-to-Video Transfer Learning for Video Temporal Grounding

    Ye Liu, Jixuan He, Wanhua Li, Junsik Kim, Donglai Wei, Hanspeter Pfister, and Chang Wen Chen. R � -Tuning: Effi- cient Image-to-Video Transfer Learning for Video Temporal Grounding. In Proceedings of the European Conference on Computer Vision (ECCV), pages 421–438. Springer, 2024. 7

  22. [30]

    Fixing weight decay regularization in adam

    Ilya Loshchilov, Frank Hutter, et al. Fixing weight decay regularization in adam. arXiv preprint arXiv:1711.05101, 5,

  23. [31]

    Debug: A dense bottom-up grounding approach for natural language video localization

    Chujie Lu, Long Chen, Chilie Tan, Xiaolin Li, and Jun Xiao. Debug: A dense bottom-up grounding approach for natural language video localization. InProceedings of the 2019 Con- ference on Empirical Methods in Natural Language Process- ing and the 9th International Joint Confere...

  24. [32]

    Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training

    Dezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin, and Yang Liu. Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 23045–23055...

  25. [33]

    Zero-shot video moment retrieval from frozen vision-language models

    Dezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin, and Yang Liu. Zero-shot video moment retrieval from frozen vision-language models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 5464–5473, 2024. 2, 3

  26. [34]

    Correlation-guided query-dependency calibration in video representation learning for temporal grounding

    WonJun Moon, Sangeek Hyun, SuBeen Lee, and Jae-Pil Heo. Correlation-guided query-dependency calibration in video representation learning for temporal grounding. arXiv preprint arXiv:2311.08835, 2023. 2

  27. [35]

    Query-dependent video representa- tion for moment retrieval and highlight detection

    WonJun Moon, Sangeek Hyun, SangUk Park, Dongchan Park, and Jae-Pil Heo. Query-dependent video representa- tion for moment retrieval and highlight detection. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 23023–23033, 2023. 2, 1

  28. [36]

    Snag: Scalable and accurate video grounding

    Fangzhou Mu, Sicheng Mo, and Yin Li. Snag: Scalable and accurate video grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18930–18940, 2024. 2

  29. [37]

    Local- global video-text interactions for temporal grounding

    Jonghwan Mun, Minsu Cho, and Bohyung Han. Local- global video-text interactions for temporal grounding. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pages 10810–10819,

  30. [38]

    Chatvtg: Video temporal grounding via chat with video dialogue large language models

    Mengxue Qu, Xiaodong Chen, Wu Liu, Alicia Li, and Yao Zhao. Chatvtg: Video temporal grounding via chat with video dialogue large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1847–1856, 2024. 2

  31. [39]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...

  32. [40]

    Proposal-free temporal moment localization of a natural-language query in video using guided attention

    Cristian Rodriguez, Edison Marrese-Taylor, Fatemeh Sadat Saleh, Hongdong Li, and Stephen Gould. Proposal-free temporal moment localization of a natural-language query in video using guided attention. In Proceedings of the IEEE/CVF winter conference on applications of computer ...

  33. [41]

    Coherent multi-sentence video description with variable level of detail

    Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele. Coherent multi-sentence video description with variable level of detail. In Pattern Recognition: 36th German Conference, GCPR 2014, M¨unster, Germany, September 2-5, 2014, Proceedi...

  34. [42]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 6, 3

  35. [43]

    Focal loss for dense object detection

    T-YLPG Ross and GKHP Doll ´ar. Focal loss for dense object detection. In proceedings of the IEEE conference on com- puter vision and pattern recognition (CVPR) , pages 2980– 2988, 2017. 5, 1

  36. [44]

    Mad: A scalable dataset for language grounding in videos from movie audio descriptions

    Mattia Soldan, Alejandro Pardo, Juan Le ´on Alc´azar, Fabian Caba, Chen Zhao, Silvio Giancola, and Bernard Ghanem. Mad: A scalable dataset for language grounding in videos from movie audio descriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R...

  37. [45]

    Tr- detr: Task-reciprocal transformer for joint moment retrieval and highlight detection

    Hao Sun, Mingyao Zhou, Wenjing Chen, and Wei Xie. Tr- detr: Task-reciprocal transformer for joint moment retrieval and highlight detection. In Proceedings of the AAAI Confer- ence on Artificial Intelligence , pages 4998–5007, 2024. 2, 7

  38. [46]

    Learning to (learn at test time): Rnns with expressive hidden states

    Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, et al. Learning to (learn at test time): Rnns with expressive hidden states. arXiv preprint arXiv:2407.04620, 2024. 2, 3, 4

  39. [47]

    Structured multi-level interaction network for video moment localization via language query

    Hao Wang, Zheng-Jun Zha, Liang Li, Dong Liu, and Jiebo Luo. Structured multi-level interaction network for video moment localization via language query. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7026–7035, 2021. 7

  40. [48]

    Oadtr: On- line action detection with transformers

    Xiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao, Zhengrong Zuo, Changxin Gao, and Nong Sang. Oadtr: On- line action detection with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 7565–7575, 2021. 7

  41. [49]

    Negative sample matters: A renaissance of met- ric learning for temporal grounding

    Zhenzhi Wang, Limin Wang, Tao Wu, Tianhao Li, and Gang- shan Wu. Negative sample matters: A renaissance of met- ric learning for temporal grounding. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2613– 2623, 2022. 2, 3

  42. [50]

    Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

    Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...

  43. [51]

    Boundary proposal network for two-stage natural language video localization

    Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, and Jun Xiao. Boundary proposal network for two-stage natural language video localization. In Proceed- ings of the AAAI Conference on Artificial Intelligence, pages 2986–2994, 2021. 2

  44. [52]

    Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection

    Yicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma, Heng- wei Bian, Yatai Ji, Yujiu Yang, and Xiu Li. Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...

  45. [53]

    Multilevel language and vi- sion integration for text-to-clip retrieval

    Huijuan Xu, Kun He, Bryan A Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko. Multilevel language and vi- sion integration for text-to-clip retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 9062– 9069, 2019. 2

  46. [54]

    Long short-term trans- former for online action detection

    Mingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li, Wei Xia, Zhuowen Tu, and Stefano Soatto. Long short-term trans- former for online action detection. Advances in Neural In- formation Processing Systems, 34:1086–1099, 2021. 7

  47. [55]

    Vtg-gpt: Tuning-free zero-shot video temporal grounding with gpt

    Yifang Xu, Yunzhuo Sun, Zien Xie, Benxiang Zhai, and Sidan Du. Vtg-gpt: Tuning-free zero-shot video temporal grounding with gpt. Applied Sciences, 14(5):1894, 2024. 2, 3

  48. [56]

    Mh-detr: Video moment and highlight detection with cross-modal transformer

    Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Youyao Jia, and Sidan Du. Mh-detr: Video moment and highlight detection with cross-modal transformer. In 2024 International Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE,

  49. [57]

    Unloc: A unified framework for video localization tasks

    Shen Yan, Xuehan Xiong, Arsha Nagrani, Anurag Arnab, Zhonghao Wang, Weina Ge, David Ross, and Cordelia Schmid. Unloc: A unified framework for video localization tasks. In Proceedings of the IEEE/CVF International Confer- ence on Computer Vision (ICCV), pages 13623–13633, 2023. 2

  50. [58]

    Task-driven exploration: Decoupling and inter-task feedback for joint moment retrieval and highlight detection

    Jin Yang, Ping Wei, Huan Li, and Ziyang Ren. Task-driven exploration: Decoupling and inter-task feedback for joint moment retrieval and highlight detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18308–18318, 2024. 2, 7

  51. [59]

    Deconfounded video moment retrieval with causal intervention

    Xun Yang, Fuli Feng, Wei Ji, Meng Wang, and Tat-Seng Chua. Deconfounded video moment retrieval with causal intervention. In Proceedings of the 44th international ACM SIGIR conference on research and development in informa- tion retrieval, pages 1–10, 2021. 2, 3

  52. [60]

    Cogvideox: Text-to-video diffusion models with an expert transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiao- han Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 6, 3

  53. [61]

    Semantic conditioned dynamic modulation for tempo- ral sentence grounding in videos

    Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu. Semantic conditioned dynamic modulation for tempo- ral sentence grounding in videos. Advances in Neural Infor- mation Processing Systems (NeurIPS), 32, 2019. 2

  54. [62]

    To find where you talk: Temporal sentence localization in video with attention based location regression

    Yitian Yuan, Tao Mei, and Wenwu Zhu. To find where you talk: Temporal sentence localization in video with attention based location regression. In Proceedings of the AAAI Con- ference on Artificial Intelligence, pages 9159–9166, 2019. 2

  55. [63]

    A closer look at temporal sentence grounding in videos: Dataset and metric

    Yitian Yuan, Xiaohan Lan, Xin Wang, Long Chen, Zhi Wang, and Wenwu Zhu. A closer look at temporal sentence grounding in videos: Dataset and metric. In Proceedings of the 2nd international workshop on human-centric multime- dia analysis, pages 13–21, 2021. 6

  56. [64]

    Dense regression network for video grounding

    Runhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen, Mingkui Tan, and Chuang Gan. Dense regression network for video grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10287–10296, 2020. 2

  57. [65]

    Unimd: Towards unifying moment retrieval and temporal ac- tion detection

    Yingsen Zeng, Yujie Zhong, Chengjian Feng, and Lin Ma. Unimd: Towards unifying moment retrieval and temporal ac- tion detection. In European Conference on Computer Vision (ECCV), pages 286–304. Springer, 2025. 1, 2

  58. [66]

    Man: Moment alignment network for natu- ral language moment retrieval via iterative graph adjustment

    Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S Davis. Man: Moment alignment network for natu- ral language moment retrieval via iterative graph adjustment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 1247–1257,

  59. [67]

    Localizing events in videos with multimodal queries

    Gengyuan Zhang, Mang Ling Ada Fok, Yan Xia, Yansong Tang, Daniel Cremers, Philip Torr, V olker Tresp, and Jin- dong Gu. Localizing events in videos with multimodal queries. arXiv preprint arXiv:2406.10079, 2024. 3, 6, 2

  60. [68]

    Span-based localizing network for natural language video lo- calization

    Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou. Span-based localizing network for natural language video lo- calization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 6543–6554, 2020. 2, 7

  61. [69]

    Parallel attention net- work with sequence matching for video grounding

    Hao Zhang, Aixin Sun, Wei Jing, Liangli Zhen, Joey Tianyi Zhou, and Siow Mong Rick Goh. Parallel attention net- work with sequence matching for video grounding. In Find- ings of the Association for Computational Linguistics: ACL- IJCNLP 2021, pages 776–790, 2021. 7

  62. [70]

    Exploiting temporal relationships in video moment localization with natural language

    Songyang Zhang, Jinsong Su, and Jiebo Luo. Exploiting temporal relationships in video moment localization with natural language. In Proceedings of the 27th ACM Interna- tional Conference on Multimedia (ACM MM) , pages 1230– 1238, 2019. 2

  63. [71]

    Learning 2d temporal adjacent networks for moment local- ization with natural language

    Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo. Learning 2d temporal adjacent networks for moment local- ization with natural language. In Proceedings of the AAAI Conference on Artificial Intelligence , pages 12870–12877,

  64. [72]

    Cross-modal interaction networks for query-based moment retrieval in videos

    Zhu Zhang, Zhijie Lin, Zhou Zhao, and Zhenxin Xiao. Cross-modal interaction networks for query-based moment retrieval in videos. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 655–664, 2019. 2

  65. [73]

    Localizing unseen activities in video via image query

    Zhu Zhang, Zhou Zhao, Zhijie Lin, Jingkuan Song, and Deng Cai. Localizing unseen activities in video via image query. In Proceedings of the 28th International Joint Con- ference on Artificial Intelligence (IJCAI), pages 4390–4396,

  66. [74]

    Training-free video temporal grounding using large-scale pre-trained models

    Minghang Zheng, Xinhao Cai, Qingchao Chen, Yuxin Peng, and Yang Liu. Training-free video temporal grounding using large-scale pre-trained models. In European Conference on Computer Vision (ECCV), pages 20–37. Springer, 2025. 1, 2

  67. [75]

    Intra-and inter-modal curriculum for multi- modal learning

    Yuwei Zhou, Xin Wang, Hong Chen, Xuguang Duan, and Wenwu Zhu. Intra-and inter-modal curriculum for multi- modal learning. In Proceedings of the 31st ACM Interna- tional Conference on Multimedia (ACM MM) , pages 3724– 3735, 2023. 2, 7 OVG-HQ: Online Video Grounding with Hybrid-...

  68. [76]

    Text Query: The CLIP text encoder [39] generates tex- tual features � �

  69. [77]

    Segment Query: Employing the video feature extractor with 2-second sampling intervals to obtain � �

  70. [78]

    Linear � N

    Image Query : The CLIP image encoder [39] extracts features � �, which are duplicated temporally to match the segment query length. A.2. Memory-guided Multi-modal Fusion Module We employ a two-layer Transformer Decoder [35] to fuse the video features and query features, result...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.