REVIEW 4 major objections 6 minor 75 references
ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that ExpStar, a 7-billion-parameter open model, generates step-level scientific experiment commentary—procedures, principles, and safety guidelines—that outperforms 14 leading large multimodal models on the new…
desk verdict ExpInstruct is a genuinely useful new resource and ExpStar is a sensible retrieval-augmented recipe, but the headline numbers rest on GPT-4o-generated references with thin human verification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a set of four special tokens added to the model's vocabulary, which turn retrieval into an explicit decision process. <RET> and <NOT RET> let the model choose whether a given experimental step needs external knowledge, while <REL> and <NOT REL> let it judge whether each retrieved Wikipedia passage is relevant before the passage is used. Around this mechanism, the model uses a CLIP-based retriever (default EVA-CLIP-8B) to pull top-5 passages from a Wikipedia knowledge base using the video and experiment title, then supervised fine-tuning with LoRA teaches the generation, and a rule-based direct preference optimization stage rewards outputs that include experiment-specific safety guidelines.
What would settle it
Have a panel of certified science teachers independently audit the 723 test-set commentaries, scoring factual correctness of every chemical equation, physical law, and safety warning without seeing the GPT-4o labels; if the audit finds a nontrivial fraction (for example, more than 10%) of the labels or ExpStar outputs to be wrong, then the reported metric gains would reflect agreement with flawed labels rather than real instructional quality.
Extended reading notes
Core claim
The central claim is that high-quality experiment commentary is not a captioning problem but a knowledge-augmented generation problem. The paper shows that a 7B model fine-tuned on step-level video clips with a retrieval-augmented mechanism reaches 46.49 BLEU-1, 6.37 BLEU-4, 40.96 CIDEr, and 62.80 BERTScore on ExpInstruct, ahead of 14 baselines that include GPT-4o, Gemini-2.0-Flash, GLM-4V-Plus, and open video models up to 8B. Human evaluation on 200 samples gives ExpStar the highest fluency, instructional clarity, and scientific appropriateness scores. The authors attribute the gain to three design choices: adaptive retrieval decision tokens, relevance-judgment tokens, and a safety-aware reward stage.
Load-bearing premise
The load-bearing premise is that the scientific principles and safety guidelines written by GPT-4o and checked by ten students are accurate and complete enough to serve as ground truth for both training and evaluation.
Editorial extensions
If this is right
- A 7B open-source model can match or beat proprietary systems on a benchmark where output must be concise and grounded, contradicting the assumption that scale alone decides instructional quality.
- Adaptive retrieval is load-bearing: forcing retrieval on every step drops CIDEr from 40.96 to 25.56, so commentary generation needs a decision about when external knowledge is actually needed.
- Relevance filtering matters: removing the <REL>/<NOT REL> judgment reduces BLEU-4 by 1.36 and CIDEr by 2.12, showing that retrieved passages must be screened, not just fetched.
- Safety-aware reward optimization substantially increases both the precision (62.30% to 87.45%) and the frequency (20.93% to 34.27%) of correct safety-guideline generation.
- ExpInstruct itself, with 7,714 step-level clips spanning 21 subjects, gives the community a reusable benchmark for future experiment-commentary and instructional-video systems.
Reading between the lines
- Beyond the paper, the same retrieval-control token scheme could be carried over to other instructional-document generation tasks, such as equipment manuals or clinical procedure write-ups, where selective factual grounding matters.
- Beyond the paper, the reported safety-guideline gains come from a rule-based reward built on the dataset's labels; an independent check with a certified lab-safety expert would show whether the model's warnings are genuinely protective, not merely aligned with the label set.
- Beyond the paper, because train and test clips come from distinct videos, the benchmark says little about how ExpStar handles a brand-new subject area; testing on out-of-domain disciplines would reveal whether the retrieval knowledge generalizes or just matches the Wikipedia index.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ExpInstruct, a new dataset of 7,714 step-level scientific experiment commentaries spanning 21 subjects in science, healthcare, and engineering, and proposes ExpStar, a 7B LMM (built on Qwen2.5-VL-7B) that uses retrieval-augmented generation with special control tokens (<RET>, <NOT RET>, <REL>, <NOT REL>) and a rule-based DPO refinement to improve safety-guideline generation. The authors claim ExpStar substantially outperforms 14 proprietary and open-source LMMs on automatic metrics (Table 2) and on a small human evaluation (Table 3).
Significance. If the reported results hold, the paper would make a useful contribution by defining a new task and benchmark for AI-assisted scientific experiment instruction, and by showing that a 7B open-source model with adaptive retrieval can outperform much larger proprietary models on this specialized task. Strengths include the novel multi-discipline dataset, the clearly described retrieval-control-token mechanism, and the inclusion of both automatic and human evaluation. However, the significance is currently tempered by concerns about the quality of the gold labels, the circularity of using GPT-4o both to create the training targets and to serve as a baseline, the lack of statistical reliability in the comparisons, and qualitative evidence of factual scientific errors in the model's outputs.
major comments (4)
- [Sec. 3.2, Sec. 3.3, Fig. 1] The gold scientific principles and safety guidelines are generated by GPT-4o and then 'thoroughly verified' by ten college students, but the paper reports no inter-annotator agreement, no per-sample verification counts, no error-rate statistics, and does not state that every one of the 7,714 clips was checked. This is load-bearing because all training and all automatic metrics in Table 2 depend on these references. Figure 1 itself contains a GPT-4o output that confusingly mixes copper(II) chloride and copper(II) sulfate, indicating that label noise is present. Please report verification coverage, correction rates, and an error analysis on a sample of test references.
- [Sec. 3.2, Sec. 4.3, Table 2] The evaluation is partly circular. GPT-4o generates the gold principles and safety guidelines, ExpStar is supervised on those same labels, and GPT-4o also appears as a baseline. As a result, BLEU, CIDEr, METEOR, and ROUGE measure fidelity to GPT-4o's style rather than external scientific correctness. Moreover, the proprietary baselines in Table 2 are not fine-tuned on ExpInstruct, so the comparison is not matched. Please add a fine-tuned GPT-4o (or other strong fine-tuned LLM) baseline, and report a human evaluation that is source-blinded and uses expert raters to judge scientific correctness independently of the GPT-4o-derived references.
- [Tables 2 and 5] No error bars or significance tests are reported; all numbers appear to come from a single run. For example, the BLEU-4 gap between ExpStar and ExpStar w/o DPO is only 0.56, which may be within run-to-run variance. Please report results over multiple seeds with mean and standard deviation, and perform paired significance tests (e.g., bootstrap or approximate randomization) for at least the main comparisons in Tables 2 and 5.
- [Sec. 5.6, Fig. 6] The qualitative example in Figure 6 shows ExpStar generating 'Insulating material reduces capacitance by altering dielectric constant (ε)' which is physically incorrect: inserting a dielectric increases capacitance. The paper appears to treat this as a correct output, suggesting that the automatic metrics and even the manual qualitative analysis can miss substantive scientific errors. This directly affects the claim of high 'scientific appropriateness' in Table 3. Please provide a systematic error analysis of the scientific content of generated principles and correct this qualitative assessment.
minor comments (6)
- [Sec. 1, Fig. 1] The phrase 'copper(II) chloride' is inconsistent with the chemical species in the equation and the rest of the example (copper(II) sulfate, CuSO4). If this is part of GPT-4o's output, it is an instance of label noise; otherwise, it is a typo.
- [Sec. 3.3 vs. Appendix A.2] The main text calls the verifiers 'expert annotators' while also referring to them as '10 college students.' Please clarify their qualifications and state how many samples each annotator reviewed, and report inter-annotator agreement.
- [Sec. 4.2.2, Eq. (5) vs. Sec. 4.4] The training sequences in Eq. (5) condition on the gold procedure y_pro, whereas inference must use the model-generated procedure. Please discuss this train/inference mismatch and whether it affects the retrieval decision and relevance judgment.
- [Sec. 5.1] The training configuration is under-specified (only learning rate, LoRA rank, and epochs are given). Please provide full hyperparameters including batch size, learning-rate schedule, warmup, and number of training steps in Appendix D.
- [Sec. 5.5, Table 4] The claim that 'EVA-CLIP with video and title (V+T) yields the best overall performance' is not fully supported because for EVA-CLIP, V alone yields a higher CIDEr (39.08) than V+T (38.99). Specify the criterion used for 'overall' or report a rank aggregation across metrics.
- [Typos and references] There are several typos: Appendix A title 'Detials' should be 'Details'; Figure 7 caption 'biography experiment' should be 'biology experiment'; Figure 7 ground truth says '<Prodcedure>' instead of '<Procedure>'. Also, references [12] and [13] appear to be the same paper and should be merged or corrected.
Circularity Check
ExpStar's automatic gains partly reflect fitting GPT-4o-generated references, because the same GPT-4o outputs serve as training targets and evaluation ground truth.
-
fitted input called prediction
[Sec. 3.2 (Principle and Safety Guideline Annotation) with Sec. 4.2.4 (Model Training)]
"we utilize GPT-4o to determine and generate necessary annotations. Specifically, we follow a structured prompt to extract relevant scientific principles and safety guidelines from experimental procedures. ... apply a token-level cross-entropy loss over all ground-truth tokens within the target commentary sequences (i.e., ypro_i, ypri_i, ysafe_i)."
ExpStar is supervised to predict the exact tokens ypri/ysafe that GPT-4o generated, and then every automatic metric in Table 2 evaluates all models against those same GPT-4o-written references. Thus a high BLEU/CIDEr/BERTScore reflects closeness to GPT-4o's annotation style, not independent scientific correctness. GPT-4o, itself a baseline, is scored against its own output distribution, so ExpStar's automatic margin over GPT-4o is partly forced by construction. The manual verification is described qualitatively without inter-annotator agreement, error rates, or per-sample counts, so it does not break this self-referential loop.
-
fitted input called prediction
[Sec. 4.2.3 (Knowledge Relevance Judgment) with Sec. 5.5 (Ablation Study)]
"Due to the absence of passage relevance annotations, we automatically construct binary relevance labels for the top-K retrieved passages using GPT-4o. ... Passages scoring >=3 are considered relevant, while those scoring <=2 are deemed irrelevant."
The <REL>/<NOT REL> tokens are trained on relevance labels produced by GPT-4o, which scores Wikipedia passages against GPT-4o-written ground-truth principles/safety text. The ablation then attributes a BLEU-4 gain of 1.36 and a CIDEr gain of 2.12 to these relevance tokens, but the supervision signal is generated by the same model whose style defines both the references and the training targets. The claimed benefit of knowledge filtering is therefore partly self-referential: no external or human relevance benchmark is used, so the ablation confirms that ExpStar learned GPT-4o's relevance judgments rather than an independently validated selection ability.
full rationale
The core circularity is in the evaluation/training loop for the principle and safety components. GPT-4o creates the scientific principles, safety guidelines, and relevance labels; ExpStar is fine-tuned with token-level cross-entropy on those GPT-4o outputs; and the headline automatic results are measured against those same outputs. Consequently, the reported superiority over 14 LMMs is partially an artifact of training on the reference distribution, especially because GPT-4o itself is one of the baselines. This is not full circularity: the dataset includes 10 student annotators who review and revise samples, the human evaluation in Table 3 is not directly tied to n-gram matching, and the retrieval-augmented architecture is an independent engineering contribution. However, the paper reports no inter-annotator agreement or error-rate statistics, so the manual verification is insufficient to establish that the references are externally accurate. The relevance-token ablation is similarly self-referential because its training labels come from GPT-4o. These issues affect the central claim of 'substantially outperforms 14 leading LMMs' on automatic metrics, warranting a score of 5 rather than a lower score.
Assumptions & free parameters
free parameters (5)
- relevance threshold =
score >= 3
- retrieval top-K =
K = 5
- query fusion weights =
0.7 * video embedding + 0.3 * title embedding
- LoRA rank and learning rates =
rank 8, SFT lr 1e-4, DPO lr 1e-6
- DPO sampling parameters =
top-p = 0.9, L candidate outputs
assumptions (6)
- domain assumption WhisperX transcripts combined with GPT-4o correction and translation preserve the experimental procedure accurately.
- domain assumption GPT-4o can generate scientifically correct principles and safety guidelines from procedures, and the ten student annotators can verify them.
- domain assumption GPT-4o relevance scores between retrieved passages and ground truth identify useful knowledge.
- domain assumption BLEU, ROUGE, METEOR, CIDEr, and BERTScore are meaningful measures of commentary quality.
- domain assumption Wikipedia introductory paragraphs contain the external knowledge needed for all 21 subjects.
- domain assumption The video frame sampling at 1 FPS and CLIP-based retrieval provide enough visual signal to match steps to knowledge.
invented entities (1)
-
Retrieval control tokens <RET>, <NOT RET>, <REL>, <NOT REL>
Cite this review
Pith. "Pith review of ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments." pith.science (2026). https://pith.science/paper/3NVM5OBE
@misc{pith2026250709693,
author = {Pith},
title = {Pith review of: ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments},
year = {2026},
howpublished = {\url{https://pith.science/paper/3NVM5OBE}},
note = {Machine review of arXiv:2507.09693}
}
read the original abstract
Experiment commentary is crucial in describing the experimental procedures, delving into underlying scientific principles, and incorporating content-related safety guidelines. In practice, human teachers rely heavily on subject-specific expertise and invest significant time preparing such commentary. To address this challenge, we introduce the task of automatic commentary generation across multi-discipline scientific experiments. While recent progress in large multimodal models (LMMs) has demonstrated promising capabilities in video understanding and reasoning, their ability to generate fine-grained and insightful experiment commentary remains largely underexplored. In this paper, we make the following contributions: (i) We construct \textit{ExpInstruct}, the first dataset tailored for experiment commentary generation, featuring over 7\textit{K} step-level commentaries across 21 scientific subjects from 3 core disciplines (\ie, science, healthcare and engineering). Each sample includes procedural descriptions along with potential scientific principles (\eg, chemical equations and physical laws) and safety guidelines. (ii) We propose ExpStar, an automatic experiment commentary generation model that leverages a retrieval-augmented mechanism to adaptively access, evaluate, and utilize external knowledge. (iii) Extensive experiments show that our ExpStar substantially outperforms 14 leading LMMs, which highlights the superiority of our dataset and model. We believe that ExpStar holds great potential for advancing AI-assisted scientific experiment instruction.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrit- twieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P. Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael ...
arXiv 2023
-
[2]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In Proc. of ICLR. OpenReview.net
work page 2024
-
[3]
Hammad A. Ayyubi, Tianqi Liu, Arsha Nagrani, Xudong Lin, Mingda Zhang, Anurag Arnab, Feng Han, Yukun Zhu, Xuande Feng, Kevin Zhang, Jialu Liu, and Shih-Fu Chang. 2024. VIEWS: Entity-Aware News Video Captioning. In Proc. of EMNLP, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, 20220–20239
work page 2024
-
[4]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Frontier Large Vision- Language Model with Versatile Abilities. CoRR abs/2308.12966 (2023)
arXiv 2023
-
[5]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. 2025. Qwen2. 5-VL Technical Report. arXiv preprint arXiv:2502.13923 (2025)
arXiv 2025
-
[6]
Max Bain, Jaesung Huh, Tengda Han, and Andrew Zisserman. 2023. WhisperX: Time-Accurate Speech Transcription of Long-Form Audio. In Proc. of INTER- SPEECH, Naomi Harte, Julie Carson-Berndsen, and Gareth Jones (Eds.). ISCA, 4489–4493
work page 2023
-
[7]
Jiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai, and Qing Li. 2023. Deconfounded Visual Question Generation with Causal Inference. In Proc. of ACM MM, Abdul- motaleb El-Saddik, Tao Mei, Rita Cucchiara, Marco Bertini, Diana Patricia Tobon Vallejo, Pradeep K. Atrey, and M. Shamim Hossain (Eds.). ACM, 5132–5142
work page 2023
-
[8]
Jiali Chen, Xusen Hei, Yuqi Xue, Yuancheng Wei, Jiayuan Xie, Yi Cai, and Qing Li. 2024. Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor. In Proc. of ACM MM, Jianfei Cai, Mohan S. Kankanhalli, Balakrishnan Prabhakaran, Susanne Boll, Ramanathan Subramanian, Liang Zheng, Vivek K. Singh, Pablo César, Lexing ...
work page 2024
Show all 75 references
-
[9]
Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Lin Bin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, and Jiaqi Wang. 2024. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions. InProc. of Neu...
2024
-
[10]
Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che. 2024. M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain- of-Thought. In Proc. of ACL, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, 8199–8221
2024
-
[11]
Yuxiao Chen, Kai Li, Wentao Bao, Deep Patel, Yu Kong, Martin Renqiang Min, and Dimitris N. Metaxas. 2024. Learning to Localize Actions in Instructional Videos with LLM-Based Multi-pathway Text-Video Alignment. In Proc. of ECCV, Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Ru...
2024
-
[13]
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, Lixin Gu, Xuehui Wang, Qingyun Li, Yimin Ren, Zixuan Chen, Jiapeng Luo, Jiahao Wang, Tan Jiang, Bo Wang, Conghui He, Botian Shi, Xingcheng Zhang, Han Lv, Yi...
2024 arXiv
-
[14]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. 2023. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks. CoRR a...
2023 arXiv
-
[15]
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing. 2024. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs. CoRR abs/2406.07476 (2024)
2024 arXiv
-
[16]
Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2024. Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering. CoRR abs/2411.16863 (2024)
2024 arXiv
-
[17]
Denkowski and Alon Lavie
Michael J. Denkowski and Alon Lavie. 2014. Meteor Universal: Language Specific Translation Evaluation for Any Target Language. In Proc. of ACL Workshop . 376–380
2014
-
[18]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sra- vankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston...
2024 arXiv
-
[19]
David Gooding, Trevor Pinch, and Simon Schaffer. 1989. The uses of experiment: Studies in the natural sciences
1989
-
[20]
Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, and Andrew Zis- serman. 2023. AutoAD II: The Sequel - Who, When, and What in Movie Audio Description. In Proc. of ICCV. IEEE, 13599–13609
2023
-
[21]
Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, and Andrew Zis- serman. 2023. AutoAD: Movie Description in Context. In Proc. of CVPR. IEEE, 18930–18940
2023
-
[22]
Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, and Andrew Zisser- man. 2024. AutoAD III: The Prequel - Back to the Pixels. In Proc. of CVPR. IEEE, 18164–18174
2024
-
[23]
Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wanrong Zhu, Jiachen Li, Yue Fan, Jianfeng Wang, Linjie Li, Zhengyuan Yang, Kevin Lin, William Yang Wang, Lijuan Wang, and Xin Eric Wang. 2024. MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos. CoRR...
2024 arXiv
-
[24]
Wenyi Hong, Weihan Wang, Ming Ding, Wenmeng Yu, Qingsong Lv, Yan Wang, Yean Cheng, Shiyu Huang, Junhui Ji, Zhao Xue, Lei Zhao, Zhuoyi Yang, Xiaotao Gu, Xiaohan Zhang, Guanyu Feng, Da Yin, Zihan Wang, Ji Qi, Xixuan Song, Peng Zhang, Debing Liu, Bin Xu, Juanzi Li, Yuxiao Dong, a...
2024 arXiv
-
[25]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In Proc. of ICLR. OpenReview.net
2022
-
[26]
Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Yuanhan Zhang, Xiang Yue, Bo Li, and Ziwei Liu. 2025. Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos. CoRR abs/2501.13826 (2025)
2025 arXiv
-
[27]
David Klahr, Anne L Fay, and Kevin Dunbar. 1993. Heuristics for scientific experimentation: A developmental study. Cognitive psychology 25, 1 (1993), 9 111–146
1993
-
[28]
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles
-
[29]
Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. 2024. LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models. CoRR abs/2407.07895 (2024)
2024 arXiv
-
[30]
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Lou, Limin Wang, and Yu Qiao. 2024. MVBench: A Comprehensive Multi-modal Video Understanding Benchmark. In Proc. of CVPR. IEEE, 22195–22206
2024
-
[31]
Yanwei Li, Chengyao Wang, and Jiaya Jia. 2024. LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models. In Proc. of ECCV, Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol (Eds.), Vol. 15104. Springer, 323–340
2024
-
[32]
Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. 2024. Video-LLaVA: Learning United Visual Representation by Alignment Before Pro- jection. In Proc. of EMNLP, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Lingui...
2024
-
[33]
Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Proc. of ACL Workshop. 74–81
2004
-
[34]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruc- tion Tuning. In Proc. of NeurIPS, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
2023
-
[35]
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan. 2024. DeepSeek-VL: Towards Real-World Vision-Language Understanding. CoRR abs/2403.05525 (2024)
2024 arXiv
-
[36]
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. In Proc. of NeurIPS, Sanmi Koyejo, S. Mohamed, A. Agarw...
2022
-
[37]
Thomas Mensink, Jasper R. R. Uijlings, Lluís Castrejón, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araújo, and Vittorio Ferrari. 2023. Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories. In Proc. of ICCV. IEEE, 3090–3101
2023
-
[38]
OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023)
2023 arXiv
-
[39]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proc. of ACL . 311–318
2002
-
[40]
Ji Qi, Jifan Yu, Teng Tu, Kunyu Gao, Yifan Xu, Xinyu Guan, Xiaozhi Wang, Bin Xu, Lei Hou, Juanzi Li, and Jie Tang. 2023. GOAL: A Challenging Knowledge- grounded Video Captioning Benchmark for Real-time Soccer Commentary Gen- eration. In Proc. of CIKM, Ingo Frommholz, Frank Hop...
2023
-
[41]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proc. of IC...
2021
-
[42]
Manning, Stefano Ermon, and Chelsea Finn
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. InProc. of NeurIPS, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hard...
2023
-
[43]
Jiayuan Rao, Haoning Wu, Hao Jiang, Ya Zhang, Yanfeng Wang, and Weidi Xie
-
[44]
Jiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang, and Weidi Xie. 2024. MatchTime: Towards Automatic Soccer Game Commentary Generation. In Proc. of EMNLP, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Associa- tion for Computational Linguistics, 1671–1685
2024
-
[45]
Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. 2022. Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities. In Proc. of CVPR. IEEE, 21064–21074
2022
-
[46]
Yan Shu, Peitian Zhang, Zheng Liu, Minghao Qin, Junjie Zhou, Tie-Jun Huang, and Bo Zhao. 2024. Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding. CoRR abs/2409.14485 (2024)
2024 arXiv
-
[47]
Quan Sun, Jinsheng Wang, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, and Xinlong Wang. 2024. EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters. CoRR abs/2402.04252 (2024)
2024 arXiv
-
[48]
Lawrence Zitnick, and Devi Parikh
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. CIDEr: Consensus-based image description evaluation. In Proc. of CVPR. 4566–4575
2015
-
[49]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. Qwen2-VL: Enhancing Vision-Language Mode...
2024 arXiv
-
[50]
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Yu Qiao. 2024. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation. In Proc. of IC...
2024
-
[51]
Hongchen Wei, Zhihong Tan, Yaosi Hu, Chang Wen Chen, and Zhenzhong Chen
-
[52]
Stephen Wooding, Alison Cullinane, and Sibel Erduran. 2020. Supporting the teaching of scientific methods in practical science. (2020)
2020
-
[53]
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, Zhenda Xie, Yu Wu, Kai Hu, Jiawei Wang, Yaofeng Sun, Yukun Li, Yishi Piao, Kang Guan, Aixin Liu, Xin Xie, Yuxiang You, Kai Dong, Xingkai Yu, Haowei Zhang,...
2024 arXiv
-
[54]
Zihui Xue, Joungbin An, Xitong Yang, and Kristen Grauman. 2024. Progress- Aware Video Frame Captioning. CoRR abs/2412.02071 (2024)
2024 arXiv
-
[55]
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont- Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. 2023. Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning. In Proc. of CVPR. IEEE, 10714–10726
2023
-
[56]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Cheng- peng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jin...
2024 arXiv
-
[57]
Dingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang, Tiezheng Ge, Bo Zheng, and Qin Jin. 2024. Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline. In Proc. of ACL, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computatio...
2024
-
[58]
Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2024. MMM...
2024
-
[59]
Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Mingdao Liu, Minlie Huang, Peng...
2024 arXiv
-
[60]
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sig- moid Loss for Language Image Pre-Training. InProc. of ICCV. IEEE, 11941–11952
2023
-
[61]
Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction- tuned Audio-Visual Language Model for Video Understanding. InProc. of EMNLP, Yansong Feng and Els Lefever (Eds.). Association for Computational Linguistics, 543–553
2023
-
[62]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi
-
[63]
Yilun Zhao, Lujing Xie, Haowei Zhang, Guo Gan, Yitao Long, Zhiyuan Hu, Tongyan Hu, Weiyuan Chen, Chuhan Li, Junyang Song, Zhijian Xu, Chengye Wang, Weifeng Pan, Ziyao Shangguan, Xiangru Tang, Zhenwen Liang, Yixin Liu, Chen Zhao, and Arman Cohan. 2025. MMVU: Measuring Expert-Le...
2025 arXiv
-
[64]
Xingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan, Austin Myers, Xuehan Xiong, Arsha Nagrani, and Cordelia Schmid. 2024. Streaming Dense Video Captioning. In Proc. of CVPR. IEEE, 18243–18252
2024
-
[65]
id": "",
Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Serena Yeung-Levy, and Xide Xia. 2024. Apollo: An Exploration of Video Understanding in Large Multimodal Models. CoRR abs/2412.10360 ...
2024 arXiv
-
[70]
Generate the "summary" field (experiment summary)
-
[71]
Summarize the experiment procedure into the "steps" field, with numbered steps starting from 1
-
[72]
procedure
Generate a new field "procedure" (a summary description of each step)
-
[73]
Each step’s ASR fragments should consist of a continuous sequence of positive integer numbers (e.g., [2,3,4,5]), without any jumps or regressions (e.g., [2,3,7,5] is not allowed)
Generate the "ASR_id" field, which must strictly follow the order of the ASR fragments. Each step’s ASR fragments should consist of a continuous sequence of positive integer numbers (e.g., [2,3,4,5]), without any jumps or regressions (e.g., [2,3,7,5] is not allowed)
-
[74]
ASR_id" only contains continuous positive integer numbers, without any skips or backtracking. Template: {
Each ASR fragment must only appear in one step and cannot be repeated in different steps. Note: - Only return the JSON result, no additional explanation is required. - Ensure that the "ASR_id" only contains continuous positive integer numbers, without any skips or backtracking...
-
[75]
<Principle>: to annotate explanations of scientific principles or theories
-
[76]
Detecting Sugars, Fats and Proteins in Biological Tissues
<Safety>: to annotate safety precautions or warnings. Please generate a detailed procedure for the current step of the biology experiment titled"Detecting Sugars, Fats and Proteins in Biological Tissues". The images provided are equidistant samples taken from a video. Keep you...
-
[2017]
Dense-Captioning Events in Videos. In Proc. of ICCV . IEEE Computer Society, 706–715
-
[2020]
BERTScore: Evaluating Text Generation with BERT. In Proc. of ICLR . OpenReview.net
-
[2024]
CoRR abs/2412.01820 (2024)
Towards Universal Soccer Video Understanding. CoRR abs/2412.01820 (2024)
2024 arXiv
-
[2025]
CoRR abs/2502.15393 (2025)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models. CoRR abs/2502.15393 (2025)
2025 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.