Pith. sign in

REVIEW 4 major objections 5 minor 71 references

The Repeated-Stimulus Confound in Electroencephalography

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that reusing the same stimuli in training and testing inflates EEG decoding accuracy by 4.46–7.42%, affecting 16 published studies.

desk verdict A plausible, high-impact methodological critique whose quantitative claims are unverifiable from the supplied materials; it deserves peer review, not desk rejection. read the letter →

arxiv 2508.00531 v1 pith:WZSW2QAO submitted 2025-08-01 q-bio.NC cs.CV

classification q-bio.NCcs.CV
keywords electroencephalographyneuraldecodingrepeated-stimulusconfounddataleakagestimulusidentityclassificationaccuracyreproducibilitydeeplearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that a widespread evaluation practice in EEG decoding—showing the same stimulus to a participant many times and then using responses to that stimulus in both the training and test sets—artificially inflates reported decoding accuracies. The authors name this the repeated-stimulus confound, identify one susceptible dataset and 16 publications affected by it, and run experiments with models from those studies to estimate the size of the inflation. They conclude that the reported accuracies were overestimated by 4.46–7.42%, and that the overestimation grows by 0.26% for every 1% of reported accuracy. A sympathetic reader should care because the confound undermines a body of published findings and shows how easily leakage can manufacture scientifically implausible results, such as evidence for extrasensory perception.

What carries the argument

The key mechanism is the repeated-stimulus confound: when identical stimuli appear in both the training and test partitions, a model can memorize stimulus identity rather than learn a generalizable neural decoding rule. The paper makes the confound quantitative by running models from the affected studies under confounded versus stimulus-disjoint evaluation on a susceptible dataset, and by fitting the relationship between reported accuracy and the size of the inflation.

What would settle it

Retrain the models from any affected study using a strict stimulus-disjoint split—where responses to the same physical stimulus never appear in both training and test—and compare the resulting accuracy to the study's published number. If the gap falls well below 4.46%, the paper's central estimate loses support.

Watch

Extended reading notes

Core claim

The central claim is that decoding models trained and evaluated on EEG responses to the same physical stimuli can appear far more accurate than they really are, because stimulus identity provides a leakage shortcut. Using a susceptible dataset and models drawn from the affected studies, the paper estimates this inflation at 4.46–7.42% absolute accuracy, with a per-point scaling of 0.26% per 1% of reported accuracy. The paper also demonstrates that the same evaluation pipeline yields 'evidence' for extrasensory perception, which serves as a stark proof that the confound can generate absurd conclusions.

Load-bearing premise

The paper's quantitative estimate assumes that the one susceptible dataset it uses, and the models it selects from the affected publications, fairly represent the evaluation procedures across all 16 papers; if those studies use different stimulus schedules, preprocessing, or architectures, the measured inflation could differ.

Editorial extensions

If this is right

  • Reported accuracies in the 16 identified publications are probably inflated by 4.46–7.42% and need re-evaluation.
  • Future EEG decoding studies should split data at the stimulus level, not the trial level, to avoid this leakage.
  • Because inflation grows with reported accuracy, high-scoring models deserve extra scrutiny for stimulus leakage.
  • The same methodology can produce absurd results like ESP, so decoding pipelines should be validated against impossible-condition baselines.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The confound probably generalizes beyond EEG to other neuroimaging and behavioral decoding studies that reuse stimuli across train and test splits.
  • The specific 4.46–7.42% figure may not transfer across datasets; the durable lesson is the existence of the confound and the need for stimulus-disjoint validation.
  • The ESP demonstration suggests a cheap sanity test: run any proposed decoding pipeline on a condition where only chance performance is possible, and see whether the pipeline's accuracy stays at chance.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript, as identified by the abstract, argues that EEG neural-decoding studies that train and evaluate models on repeated presentations of identical stimuli suffer from a "repeated-stimulus confound" that inflates reported accuracies. It claims to identify a susceptible dataset and 16 affected publications, and to show experimentally that decoding accuracies in those publications were overestimated by 4.46–7.42%, with an additional 0.26% overestimation per 1% increase in confounded accuracy. It further claims that the same methodology could support pseudoscientific claims such as extrasensory perception. The body of the submitted manuscript, however, is an unrelated paper on egocentric spatiotemporal video grounding and contains none of the EEG experiments, data, or analyses described in the abstract.

Significance. If substantiated, the central claim would be important: it would imply that a common experimental convenience (repeated stimuli) biases decoding accuracy in a quantifiable way, and that results in up to 16 publications are inflated. The additional demonstration that the confound can manufacture evidence for extrasensory perception would be a striking, cautionary result. The paper's potential value, however, is entirely conditional on the presentation of the EEG methodology and results. The manuscript does not provide code, data, error bars, or statistical tests for the EEG claims, so the significance cannot currently be assessed.

major comments (4)
  1. [Full text] The body of the submitted manuscript is a different paper ("Fine-grained Spatiotemporal Grounding on Egocentric Videos") and contains no Methods, Results, or supporting material for the EEG repeated-stimulus confound. The abstract's central claims—the 4.46–7.42% overestimation range, the 0.26% slope, the identification of 16 affected publications, and the extrasensory-perception demonstration—are therefore unsupported by any presented evidence. This is a load-bearing gap that prevents evaluation of the paper's central contribution.
  2. [Abstract] The quantitative estimate of 4.46–7.42% overestimation is asserted without describing the susceptible dataset, number of participants, trial counts, stimulus schedules, preprocessing, model architectures, train/test splits, error bars, confidence intervals, or statistical tests. Because the claim is stated as a literature-wide correction for 16 publications, the manuscript must demonstrate representativeness and provide sensitivity analyses; none are present.
  3. [Abstract] The linear relationship "per 1% increase in accuracy under the confound, the magnitude of the overestimation increases by 0.26%" is presented without the fitted data, sample size, range of accuracies, or uncertainty. Without this information, extrapolation to the affected publications is not justified.
  4. [Abstract] The claim that the same methodology "could also be used to justify an array of pseudoscientific claims, such as the existence of extrasensory perception" is presented as a finding, but no experiments or analysis supporting it appear in the manuscript.
minor comments (5)
  1. [Title/Abstract] The title and abstract are inconsistent with the body; the document mismatch should be resolved before any resubmission.
  2. [Abstract] Terms such as "overestimated" and "misreported" need definitions relative to a ground-truth generalization accuracy that is not confounded by stimulus identity.
  3. [Abstract] The "susceptible dataset" is not named or described, so readers cannot assess its relevance to the 16 affected publications.
  4. [Abstract] The 16 affected publications are not listed or cited; a full list with the specific evaluation procedures at issue is necessary for verification.
  5. [Abstract] The extrasensory-perception claim should be framed as a precise methodological demonstration with defined statistical criteria rather than an interpretive aside.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the claimed overestimation is measured by direct confounded-versus-unconfounded comparison, not derived from the definition of the confound.

full rationale

The abstract's central claim is that decoding accuracies in 16 affected publications were overestimated by 4.46-7.42%. This is presented as the result of experiments: 'We conducted experiments using models from the affected studies to investigate the likely extent to which results in the literature have been misreported.' That is a measurement comparing confounded and unconfounded evaluations of the same models, not a logical consequence of the definition of the repeated-stimulus confound. The numerical range depends on the data and models used, so it is not self-definitional. The secondary linear relationship ('per 1% increase in accuracy under the confound, the magnitude of the overestimation increases by 0.26%') is an empirical regression-style summary, not a fitted parameter that is then renamed as a prediction of the same quantity. No load-bearing self-citations, imported uniqueness theorems, or ansatz-by-citation steps are visible in the abstract. The supplied full text is a different arXiv paper (2508.00518v2, cs.CV), so the methods section that would allow further checking is unavailable; that is a completeness gap, not evidence of circularity. The generalization from one susceptible dataset to all 16 publications relies on a representativeness assumption, but that is a correctness or external-validity concern, not circular reasoning. Based on the available derivation chain, the finding is self-contained empirical work.

Assumptions & free parameters 1 free parameters · 2 assumptions · 0 invented entities

The paper's quantitative claims rest on the representativeness of one susceptible dataset and selected models. No new physical or mathematical entities are introduced.

free parameters (1)
  • overestimation slope per 1% accuracy under the confound = 0.26%
    Reported in the abstract as a linear relationship between confounded accuracy and overestimation. It appears to be fitted to the experimental data rather than derived from first principles.
assumptions (2)
  • domain assumption The analyzed susceptible dataset is representative of the evaluation procedures used in the 16 affected publications.
    The abstract states that they 'identify a susceptible dataset' and then generalize the overestimation estimate to 16 publications. This generalization requires the dataset to be representative of those publications' evaluation setups.
  • domain assumption The models selected from the affected studies behave similarly on the susceptible dataset as in the original publications.
    The abstract says they 'conducted experiments using models from the affected studies' to estimate the extent of misreporting. This assumes the models reproduce the original conditions and that their performance differences are attributable to the confound.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Repeated-Stimulus Confound in Electroencephalography." pith.science (2026). https://pith.science/paper/WZSW2QAO

@misc{pith2026250800531,
  author       = {Pith},
  title        = {Pith review of: The Repeated-Stimulus Confound in Electroencephalography},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WZSW2QAO}},
  note         = {Machine review of arXiv:2508.00531}
}
read the original abstract

In neural-decoding studies, recordings of participants' responses to stimuli are used to train models. In recent years, there has been an explosion of publications detailing applications of innovations from deep-learning research to neural-decoding studies. The data-hungry models used in these experiments have resulted in a demand for increasingly large datasets. Consequently, in some studies, the same stimuli are presented multiple times to each participant to increase the number of trials available for use in model training. However, when a decoding model is trained and subsequently evaluated on responses to the same stimuli, stimulus identity becomes a confounder for accuracy. We term this the repeated-stimulus confound. We identify a susceptible dataset, and 16 publications which report model performance based on evaluation procedures affected by the confound. We conducted experiments using models from the affected studies to investigate the likely extent to which results in the literature have been misreported. Our findings suggest that the decoding accuracies of these models were overestimated by between 4.46-7.42%. Our analysis also indicates that per 1% increase in accuracy under the confound, the magnitude of the overestimation increases by 0.26%. The confound not only results in optimistic estimates of decoding performance, but undermines the validity of several claims made within the affected publications. We conducted further experiments to investigate the implications of the confound in alternative contexts. We found that the same methodology used within the affected studies could also be used to justify an array of pseudoscientific claims, such as the existence of extrasensory perception.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

71 extracted references · 57 canonical work pages

  1. [1]

    One token to seg them all: Language instructed reasoning segmentation in videos

    Zechen Bai, Tong He, Haiyang Mei, Pichao Wang, Ziteng Gao, Joya Chen, liulei, Zheng Zhang, and Mike Zheng Shou. One token to seg them all: Language instructed reasoning segmentation in videos. InNeurIPS, 2024. 1, 2, 3, 4, 5, 6

  2. [2]

    End-to-end referring video object segmentation with multi- modal transformers

    Adam Botach, Evgenii Zheltonozhskii, and Chaim Baskin. End-to-end referring video object segmentation with multi- modal transformers. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 4985–4995, 2022. 2

  3. [3]

    Holger Caesar, Jasper R. R. Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. InCVPR, pages 1209–1218. Computer Vision Foundation / IEEE Com- puter Society, 2018. 4

  4. [4]

    Hadzic, Taran Kota, Jimming He, Crist´obal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei

    Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota, Jimming He, Crist´obal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei. Hourvideo: 1-hour video-language understanding. InNeurIPS, 2024. 1

  5. [5]

    Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu, Sanja Fidler, Raquel Urtasun, and Alan L. Yuille. Detect what you can: Detecting and representing objects using holistic models and body parts. InCVPR, pages 1979–1986. IEEE Computer Society, 2014. 4

  6. [6]

    Epic-kitchens visor benchmark: Video segmentations and object relations

    Ahmad Darkhalil, Dandan Shan, Bin Zhu, Jian Ma, Amlan Kar, Richard Higgins, Sanja Fidler, David Fouhey, and Dima Damen. Epic-kitchens visor benchmark: Video segmentations and object relations. InProceedings of the Neural Informa- tion Processing Systems (NeurIPS) Track on Datasets and Benchmarks, 2022. 3

  7. [8]

    Grounded question-answering in long egocentric videos

    Shangzhe Di and Weidi Xie. Grounded question-answering in long egocentric videos. InCVPR, pages 12934–12943. IEEE,

  8. [9]

    Mevis: A large-scale benchmark for video segmentation with motion expressions

    Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, and Chen Change Loy. Mevis: A large-scale benchmark for video segmentation with motion expressions. InICCV, pages 2694–

Show all 71 references
  1. [10]

    Actor and action video segmentation from a sentence

    Kirill Gavrilyuk, Amir Ghodrati, Zhenyang Li, and Cees GM Snoek. Actor and action video segmentation from a sentence. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 5958–5966, 2018. 2

  2. [11]

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrish- nan, Fiona Ryan, Jayant Sharma, Michael Wray, Meng...

  3. [12]

    Context-guided spatio-temporal video grounding

    Xin Gu, Heng Fan, Yan Huang, Tiejian Luo, and Libo Zhang. Context-guided spatio-temporal video grounding. InCVPR, pages 18330–18339. IEEE, 2024. 1

  4. [13]

    Gqa: A new dataset for real-world visual reasoning and compositional question answering

    Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. InCVPR, 2019. 3

  5. [14]

    Tgif-qa: Toward spatio-temporal reasoning in visual question answering

    Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 2758–2766, 2017. 2

  6. [15]

    ISAT with Segment Any- thing: An Interactive Semi-Automatic Annotation Tool, 2024

    Shuwei Ji and Hongyuan Zhang. ISAT with Segment Any- thing: An Interactive Semi-Automatic Annotation Tool, 2024. Updated on 2025-02-07. 4

  7. [16]

    Embrac- ing consistency: A one-stage approach for spatio-temporal video grounding

    Yang Jin, Yongzhi Li, Zehuan Yuan, and Yadong Mu. Embrac- ing consistency: A one-stage approach for spatio-temporal video grounding. InNeurIPS, 2022. 1

  8. [17]

    Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara L. Berg. Referitgame: Referring to objects in pho- tographs of natural scenes. InEMNLP, pages 787–798. ACL,

  9. [18]

    Video object segmentation with language referring expressions

    Anna Khoreva, Anna Rohrbach, and Bernt Schiele. Video object segmentation with language referring expressions. In ACCV (4), pages 123–141. Springer, 2018. 2, 4, 13

  10. [19]

    Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross B

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chlo´e Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross B. Girshick. Segment anything. InICCV, pages 3992–

  11. [20]

    Refego: Re- ferring expression comprehension dataset from first-person perception of ego4d

    Shuhei Kurita, Naoki Katsura, and Eri Onami. Refego: Re- ferring expression comprehension dataset from first-person perception of ego4d. InICCV, pages 15168–15178. IEEE,

  12. [21]

    LISA: reasoning segmentation via large language model

    Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. LISA: reasoning segmentation via large language model. InCVPR, pages 9579–9589. IEEE,

  13. [22]

    Berg, and Mohit Bansal

    Jie Lei, Tamara L. Berg, and Mohit Bansal. Qvhighlights: De- tecting moments and highlights in videos via natural language queries.CoRR, abs/2107.09609, 2021. 1

  14. [23]

    Robust referring video object segmentation with cyclic structural consensus

    Xiang Li, Jinglu Wang, Xiaohao Xu, Xiao Li, Bhiksha Raj, and Yan Lu. Robust referring video object segmentation with cyclic structural consensus. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 22236– 22245, 2023. 2

  15. [24]

    Egocentric video-language pretraining

    Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu, and Mike Zheng Shou. Egocentric video-language pretraining. InNeurIPS,

  16. [25]

    Improved baselines with visual instruction tuning

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. InCVPR, pages 26286–26296. IEEE, 2024. 3

  17. [26]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection.arXiv preprint arXiv:2303.05499, 2023

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection.arXiv preprint arXiv:2303.05499, 2023. 5

  18. [27]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. InEuropean Con- ference on Computer Vision, pages 38–55. Springe...

  19. [28]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InICLR (Poster). OpenReview.net, 2019. 6

  20. [29]

    Egoschema: A diagnostic benchmark for very long- form video language understanding

    Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long- form video language understanding. InNeurIPS, 2023. 3

  21. [30]

    Yuille, and Kevin Murphy

    Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Cam- buru, Alan L. Yuille, and Kevin Murphy. Generation and com- prehension of unambiguous object descriptions. InCVPR, pages 11–20. IEEE Computer Society, 2016. 4

  22. [31]

    Spectrum-guided multi-granularity referring video object segmentation

    Bo Miao, Mohammed Bennamoun, Yongsheng Gao, and Ajmal Mian. Spectrum-guided multi-granularity referring video object segmentation. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 920– 930, 2023. 2

  23. [32]

    Embodiedgpt: Vision-language pre-training via embodied chain of thought

    Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang, Mingyu Ding, Jun Jin, Bin Wang, Jifeng Dai, Yu Qiao, and Ping Luo. Embodiedgpt: Vision-language pre-training via embodied chain of thought. InNeurIPS, 2023. 2

  24. [33]

    Hello gpt-4o

    OpenAI. Hello gpt-4o. https://openai.com/index/ hello-gpt-4o/, 2024. Accessed: 2024-07-29. 2, 3

  25. [34]

    Egovideo: Exploring egocentric foundation model and downstream adaptation.CoRR, abs/2406.18070,

    Baoqi Pei, Guo Chen, Jilan Xu, Yuping He, Yicheng Liu, Kanghua Pan, Yifei Huang, Yali Wang, Tong Lu, Limin Wang, and Yu Qiao. Egovideo: Exploring egocentric foundation model and downstream adaptation.CoRR, abs/2406.18070,

  26. [35]

    Gross, and Alexander Sorkine- Hornung

    Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus H. Gross, and Alexander Sorkine- Hornung. A benchmark dataset and evaluation methodology for video object segmentation. InCVPR, pages 724–732. IEEE Computer Society, 2016. 2

  27. [36]

    Hd-epic: A highly-detailed egocentric video dataset

    Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Parida, Kaiting Liu, Prajwal Gatti, Sid- dhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rho- dri Guerrier, Fahd Abdelazim, Bin Zhu, Davide Moltisanti, Michael Wray, Hazel Doughty, and Dima ...

  28. [37]

    An outlook into the future of egocentric vision.Int

    Chiara Plizzari, Gabriele Goletto, Antonino Furnari, Siddhant Bansal, Francesco Ragusa, Giovanni Maria Farinella, Dima Damen, and Tatiana Tommasi. An outlook into the future of egocentric vision.Int. J. Comput. Vis., 132(11):4880–4936,

  29. [38]

    The 2017 DA VIS challenge on video object segmentation.CoRR, abs/1704.00675, 2017

    Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Ar- bel´aez, Alexander Sorkine-Hornung, and Luc Van Gool. The 2017 DA VIS challenge on video object segmentation.CoRR, abs/1704.00675, 2017. 2, 5

  30. [39]

    Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

    Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, and Pengchuan Zhang. Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. InICCV, pages 5262–5274. IEEE, 2023. 2

  31. [40]

    PACO: parts and attributes of common objects

    Vignesh Ramanathan, Anmol Kalia, Vladan Petrovic, Yi Wen, Baixue Zheng, Baishan Guo, Rui Wang, Aaron Marquez, Rama Kovvuri, Abhishek Kadian, Amir Mousavi, Yiwen Song, Abhimanyu Dubey, and Dhruv Mahajan. PACO: parts and attributes of common objects. InCVPR, pages 7141–7151. IEE...

  32. [41]

    Llava++: extending visual capabilities with llama-3 and phi-3 (2024).URL https://github

    Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad S Khan. Llava++: extending visual capabilities with llama-3 and phi-3 (2024).URL https://github. com/mbzuai- oryx/LLaVA-pp, 2024. 5

  33. [42]

    Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

    Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yux- iong He. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In KDD, pages 3505–3506. ACM, 2020. 6

  34. [43]

    Girshick, Piotr Doll´ar, and Christoph Feichtenhofer

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chlo´e Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross B. Girshick, Piotr Doll´ar, and Christoph Fei...

  35. [44]

    Grounded sam: Assembling open-world models for diverse visual tasks.arXiv preprint arXiv:2401.14159,

    Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, et al. Grounded sam: Assembling open-world models for diverse visual tasks.arXiv preprint arXiv:2401.14159,

  36. [45]

    URVOS: unified referring video object segmentation network with a large-scale benchmark

    Seonguk Seo, Joon-Young Lee, and Bohyung Han. URVOS: unified referring video object segmentation network with a large-scale benchmark. InECCV (15), pages 208–223. Springer, 2020. 1, 2, 4, 5, 13

  37. [46]

    Ego4d goal-step: Toward hierarchical understanding of procedural activities

    Yale Song, Eugene Byrne, Tushar Nagarajan, Huiyu Wang, Miguel Martin, and Lorenzo Torresani. Ego4d goal-step: Toward hierarchical understanding of procedural activities. In NeurIPS, 2023. 3

  38. [47]

    Liang, Kristen Grauman, Matt Feiszli, and Weiyao Wang

    Hao Tang, Kevin J. Liang, Kristen Grauman, Matt Feiszli, and Weiyao Wang. Egotracks: A long-term egocentric visual object tracking dataset. InNeurIPS, 2023. 1, 3, 4, 13

  39. [48]

    Human-centric spatio- temporal video grounding with visual transformers.IEEE Trans

    Zongheng Tang, Yue Liao, Si Liu, Guanbin Li, Xiaojie Jin, Hongxu Jiang, Qian Yu, and Dong Xu. Human-centric spatio- temporal video grounding with visual transformers.IEEE Trans. Circuits Syst. Video Technol., 32(12):8238–8249, 2022. 1

  40. [49]

    Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.CoRR, abs/2409.12191, 2024

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision-language model’s pe...

  41. [50]

    Onlinerefer: A simple online baseline for referring video object segmentation

    Dongming Wu, Tiancai Wang, Yuang Zhang, Xiangyu Zhang, and Jianbing Shen. Onlinerefer: A simple online baseline for referring video object segmentation. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 2761–2770, 2023. 2

  42. [51]

    Language as queries for referring video object segmentation

    Jiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan, and Ping Luo. Language as queries for referring video object segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4974–4984, 2022. 2

  43. [52]

    Next-qa: Next phase of question-answering to explaining tem- poral actions

    Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa: Next phase of question-answering to explaining tem- poral actions. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9777–9786,

  44. [53]

    Video question answering via gradually refined attention over appearance and motion

    Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answering via gradually refined attention over appearance and motion. InProceedings of the 25th ACM international conference on Multimedia, pages 1645–1653, 2017. 2

  45. [54]

    Ning Xu, Linjie Yang, Yuchen Fan, Dingcheng Yue, Yuchen Liang, Jianchao Yang, and Thomas S. Huang. Youtube-vos: A large-scale video object segmentation benchmark.CoRR, abs/1809.03327, 2018. 2, 4

  46. [55]

    VISA: reasoning video object segmentation via large language mod- els

    Cilin Yan, Haochen Wang, Shilin Yan, Xiaolong Jiang, Yao Hu, Guoliang Kang, Weidi Xie, and Efstratios Gavves. VISA: reasoning video object segmentation via large language mod- els. InECCV (15), pages 98–115. Springer, 2024. 4

  47. [56]

    Video object segmentation and tracking: A survey

    Rui Yao, Guosheng Lin, Shixiong Xia, Jiaqi Zhao, and Yong Zhou. Video object segmentation and tracking: A survey. ACM Trans. Intell. Syst. Technol., 11(4):36:1–36:47, 2020. 2

  48. [57]

    Daxberger, Lin Chen, Zongyu Lin, Yanghao Li, Bowen Zhang, Haoxuan You, Dan Xu, Zhe Gan, Jiasen Lu, and Yinfei Yang

    Hanrong Ye, Haotian Zhang, Erik A. Daxberger, Lin Chen, Zongyu Lin, Yanghao Li, Bowen Zhang, Haoxuan You, Dan Xu, Zhe Gan, Jiasen Lu, and Yinfei Yang. Mm-ego: Towards building egocentric multimodal llms.CoRR, abs/2410.07177,

  49. [58]

    Mm-vet: Evaluating large multimodal models for integrated capabilities.arXiv preprint arXiv:2308.02490, 2023

    Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities.arXiv preprint arXiv:2308.02490, 2023. 3

  50. [59]

    Activitynet-qa: A dataset for understanding complex web videos via question answering

    Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 9127–9134, 2019. 2

  51. [60]

    Sa2va: Marrying SAM2 with llava for dense grounded understanding of images and videos.CoRR, abs/2501.04001, 2025

    Haobo Yuan, Xiangtai Li, Tao Zhang, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, and Ming-Hsuan Yang. Sa2va: Marrying SAM2 with llava for dense grounded understanding of images and videos.CoRR, abs/2501.04001, 2025. 1, 2, 3, 4, 5

  52. [61]

    Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502, 2023. 3

  53. [62]

    Llava-grounding: Grounded visual chat with large multimodal models

    Hao Zhang, Hongyang Li, Feng Li, Tianhe Ren, Xueyan Zou, Shilong Liu, Shijia Huang, Jianfeng Gao, Leizhang, Chunyuan Li, et al. Llava-grounding: Grounded visual chat with large multimodal models. InEuropean Conference on Computer Vision, pages 19–35. Springer, 2024. 3

  54. [63]

    Where does it exist: Spatio-temporal video grounding for multi-form sentences

    Zhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang, Huasheng Liu, and Lianli Gao. Where does it exist: Spatio-temporal video grounding for multi-form sentences. InCVPR, pages 10665–10674. Computer Vision Foundation / IEEE, 2020. 1

  55. [64]

    Be- yond embeddings: The promise of visual table in visual rea- soning

    Yiwu Zhong, Zi-Yuan Hu, Michael Lyu, and Liwei Wang. Be- yond embeddings: The promise of visual table in visual rea- soning. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 6876–6911. Association for Computational Linguistics, 2024. 4

  56. [65]

    Scene parsing through ADE20K dataset

    Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Bar- riuso, and Antonio Torralba. Scene parsing through ADE20K dataset. InCVPR, pages 5122–5130. IEEE Computer Society,

  57. [68]

    A short expression with no more than 10 words, starting with ”Short expressions: ”

  58. [69]

    Restriction Policies: - The referring expressions should be concise and informative

    A longer expression with more detailed illustrations, starting with ”Long expressions: ”. Restriction Policies: - The referring expressions should be concise and informative. They can be spatial location in the physical world, OCR characters on the object, spatial relations to...

  59. [70]

    A clear object caption with no more than 10 words, starting with ”Object Caption: ”

  60. [71]

    The visual attributes of the object, starting with ”Visual Attributes: ”

  61. [72]

    the pillows stacked on top of bed

    A concrete affordance description of the object, starting with ”Object: Affordance: ”. Restriction Policies: - Use the provided object tag selectively, as it may contain noise. - The object caption should be a noun phrase. - The object caption should clearly identify the objec...

  62. [2017]

    Total Duration(%)

    4 Appendix In the appendix, we provide more details in addition to our main paper: (1) comparison of existing datasets related to spatiotemporal grounding tasks, (2) verification results of EgoMask annotations, (3) additional statistics of our datasets, (4) additional experime...

  63. [2703]

    1, 2, 4, 5, 13

    IEEE, 2023. 1, 2, 4, 5, 13

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.