Pith. sign in

REVIEW 4 major objections 6 minor 75 references

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments

T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that ExpStar, a 7-billion-parameter open model, generates step-level scientific experiment commentary—procedures, principles, and safety guidelines—that outperforms 14 leading large multimodal models on the new…

desk verdict ExpInstruct is a genuinely useful new resource and ExpStar is a sensible retrieval-augmented recipe, but the headline numbers rest on GPT-4o-generated references with thin human verification. read the letter →

arxiv 2507.09693 v1 pith:3NVM5OBE submitted 2025-07-13 cs.CV

classification cs.CV
keywords experimentcommentarygenerationscientificvideosretrieval-augmentedlargemultimodalmodelssafetyguidelinesinstructionalvideocaptioningExpInstructdatasetdirectpreferenceoptimization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ExpStar is a system for automatically writing commentary for scientific experiment videos, and ExpInstruct is the dataset built to train and test it. The paper's central claim is that a 7-billion-parameter open model, starting from Qwen2.5-VL, can produce step-by-step commentary—what the experimenter does, why it works, and what to be careful of—that human raters judge better than commentary from larger proprietary systems such as GPT-4o. The reason, the authors argue, is that the model is trained to decide on its own when to consult outside knowledge, to filter which retrieved passages are actually relevant, and to be rewarded specifically for including correct safety guidance. If the claim holds, AI-assisted science instruction becomes feasible with a modest-size open model, lowering the time cost of preparing laboratory explanations across science, healthcare, and engineering.

What carries the argument

The load-bearing mechanism is a set of four special tokens added to the model's vocabulary, which turn retrieval into an explicit decision process. <RET> and <NOT RET> let the model choose whether a given experimental step needs external knowledge, while <REL> and <NOT REL> let it judge whether each retrieved Wikipedia passage is relevant before the passage is used. Around this mechanism, the model uses a CLIP-based retriever (default EVA-CLIP-8B) to pull top-5 passages from a Wikipedia knowledge base using the video and experiment title, then supervised fine-tuning with LoRA teaches the generation, and a rule-based direct preference optimization stage rewards outputs that include experiment-specific safety guidelines.

What would settle it

Have a panel of certified science teachers independently audit the 723 test-set commentaries, scoring factual correctness of every chemical equation, physical law, and safety warning without seeing the GPT-4o labels; if the audit finds a nontrivial fraction (for example, more than 10%) of the labels or ExpStar outputs to be wrong, then the reported metric gains would reflect agreement with flawed labels rather than real instructional quality.

Watch

Extended reading notes

Core claim

The central claim is that high-quality experiment commentary is not a captioning problem but a knowledge-augmented generation problem. The paper shows that a 7B model fine-tuned on step-level video clips with a retrieval-augmented mechanism reaches 46.49 BLEU-1, 6.37 BLEU-4, 40.96 CIDEr, and 62.80 BERTScore on ExpInstruct, ahead of 14 baselines that include GPT-4o, Gemini-2.0-Flash, GLM-4V-Plus, and open video models up to 8B. Human evaluation on 200 samples gives ExpStar the highest fluency, instructional clarity, and scientific appropriateness scores. The authors attribute the gain to three design choices: adaptive retrieval decision tokens, relevance-judgment tokens, and a safety-aware reward stage.

Load-bearing premise

The load-bearing premise is that the scientific principles and safety guidelines written by GPT-4o and checked by ten students are accurate and complete enough to serve as ground truth for both training and evaluation.

Editorial extensions

If this is right

  • A 7B open-source model can match or beat proprietary systems on a benchmark where output must be concise and grounded, contradicting the assumption that scale alone decides instructional quality.
  • Adaptive retrieval is load-bearing: forcing retrieval on every step drops CIDEr from 40.96 to 25.56, so commentary generation needs a decision about when external knowledge is actually needed.
  • Relevance filtering matters: removing the <REL>/<NOT REL> judgment reduces BLEU-4 by 1.36 and CIDEr by 2.12, showing that retrieved passages must be screened, not just fetched.
  • Safety-aware reward optimization substantially increases both the precision (62.30% to 87.45%) and the frequency (20.93% to 34.27%) of correct safety-guideline generation.
  • ExpInstruct itself, with 7,714 step-level clips spanning 21 subjects, gives the community a reusable benchmark for future experiment-commentary and instructional-video systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the same retrieval-control token scheme could be carried over to other instructional-document generation tasks, such as equipment manuals or clinical procedure write-ups, where selective factual grounding matters.
  • Beyond the paper, the reported safety-guideline gains come from a rule-based reward built on the dataset's labels; an independent check with a certified lab-safety expert would show whether the model's warnings are genuinely protective, not merely aligned with the label set.
  • Beyond the paper, because train and test clips come from distinct videos, the benchmark says little about how ExpStar handles a brand-new subject area; testing on out-of-domain disciplines would reveal whether the retrieval knowledge generalizes or just matches the Wikipedia index.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces ExpInstruct, a new dataset of 7,714 step-level scientific experiment commentaries spanning 21 subjects in science, healthcare, and engineering, and proposes ExpStar, a 7B LMM (built on Qwen2.5-VL-7B) that uses retrieval-augmented generation with special control tokens (<RET>, <NOT RET>, <REL>, <NOT REL>) and a rule-based DPO refinement to improve safety-guideline generation. The authors claim ExpStar substantially outperforms 14 proprietary and open-source LMMs on automatic metrics (Table 2) and on a small human evaluation (Table 3).

Significance. If the reported results hold, the paper would make a useful contribution by defining a new task and benchmark for AI-assisted scientific experiment instruction, and by showing that a 7B open-source model with adaptive retrieval can outperform much larger proprietary models on this specialized task. Strengths include the novel multi-discipline dataset, the clearly described retrieval-control-token mechanism, and the inclusion of both automatic and human evaluation. However, the significance is currently tempered by concerns about the quality of the gold labels, the circularity of using GPT-4o both to create the training targets and to serve as a baseline, the lack of statistical reliability in the comparisons, and qualitative evidence of factual scientific errors in the model's outputs.

major comments (4)
  1. [Sec. 3.2, Sec. 3.3, Fig. 1] The gold scientific principles and safety guidelines are generated by GPT-4o and then 'thoroughly verified' by ten college students, but the paper reports no inter-annotator agreement, no per-sample verification counts, no error-rate statistics, and does not state that every one of the 7,714 clips was checked. This is load-bearing because all training and all automatic metrics in Table 2 depend on these references. Figure 1 itself contains a GPT-4o output that confusingly mixes copper(II) chloride and copper(II) sulfate, indicating that label noise is present. Please report verification coverage, correction rates, and an error analysis on a sample of test references.
  2. [Sec. 3.2, Sec. 4.3, Table 2] The evaluation is partly circular. GPT-4o generates the gold principles and safety guidelines, ExpStar is supervised on those same labels, and GPT-4o also appears as a baseline. As a result, BLEU, CIDEr, METEOR, and ROUGE measure fidelity to GPT-4o's style rather than external scientific correctness. Moreover, the proprietary baselines in Table 2 are not fine-tuned on ExpInstruct, so the comparison is not matched. Please add a fine-tuned GPT-4o (or other strong fine-tuned LLM) baseline, and report a human evaluation that is source-blinded and uses expert raters to judge scientific correctness independently of the GPT-4o-derived references.
  3. [Tables 2 and 5] No error bars or significance tests are reported; all numbers appear to come from a single run. For example, the BLEU-4 gap between ExpStar and ExpStar w/o DPO is only 0.56, which may be within run-to-run variance. Please report results over multiple seeds with mean and standard deviation, and perform paired significance tests (e.g., bootstrap or approximate randomization) for at least the main comparisons in Tables 2 and 5.
  4. [Sec. 5.6, Fig. 6] The qualitative example in Figure 6 shows ExpStar generating 'Insulating material reduces capacitance by altering dielectric constant (ε)' which is physically incorrect: inserting a dielectric increases capacitance. The paper appears to treat this as a correct output, suggesting that the automatic metrics and even the manual qualitative analysis can miss substantive scientific errors. This directly affects the claim of high 'scientific appropriateness' in Table 3. Please provide a systematic error analysis of the scientific content of generated principles and correct this qualitative assessment.
minor comments (6)
  1. [Sec. 1, Fig. 1] The phrase 'copper(II) chloride' is inconsistent with the chemical species in the equation and the rest of the example (copper(II) sulfate, CuSO4). If this is part of GPT-4o's output, it is an instance of label noise; otherwise, it is a typo.
  2. [Sec. 3.3 vs. Appendix A.2] The main text calls the verifiers 'expert annotators' while also referring to them as '10 college students.' Please clarify their qualifications and state how many samples each annotator reviewed, and report inter-annotator agreement.
  3. [Sec. 4.2.2, Eq. (5) vs. Sec. 4.4] The training sequences in Eq. (5) condition on the gold procedure y_pro, whereas inference must use the model-generated procedure. Please discuss this train/inference mismatch and whether it affects the retrieval decision and relevance judgment.
  4. [Sec. 5.1] The training configuration is under-specified (only learning rate, LoRA rank, and epochs are given). Please provide full hyperparameters including batch size, learning-rate schedule, warmup, and number of training steps in Appendix D.
  5. [Sec. 5.5, Table 4] The claim that 'EVA-CLIP with video and title (V+T) yields the best overall performance' is not fully supported because for EVA-CLIP, V alone yields a higher CIDEr (39.08) than V+T (38.99). Specify the criterion used for 'overall' or report a rank aggregation across metrics.
  6. [Typos and references] There are several typos: Appendix A title 'Detials' should be 'Details'; Figure 7 caption 'biography experiment' should be 'biology experiment'; Figure 7 ground truth says '<Prodcedure>' instead of '<Procedure>'. Also, references [12] and [13] appear to be the same paper and should be merged or corrected.

Circularity Check

2 steps flagged · score 5.0 of 10

ExpStar's automatic gains partly reflect fitting GPT-4o-generated references, because the same GPT-4o outputs serve as training targets and evaluation ground truth.

  1. fitted input called prediction [Sec. 3.2 (Principle and Safety Guideline Annotation) with Sec. 4.2.4 (Model Training)]
    "we utilize GPT-4o to determine and generate necessary annotations. Specifically, we follow a structured prompt to extract relevant scientific principles and safety guidelines from experimental procedures. ... apply a token-level cross-entropy loss over all ground-truth tokens within the target commentary sequences (i.e., ypro_i, ypri_i, ysafe_i)."

    ExpStar is supervised to predict the exact tokens ypri/ysafe that GPT-4o generated, and then every automatic metric in Table 2 evaluates all models against those same GPT-4o-written references. Thus a high BLEU/CIDEr/BERTScore reflects closeness to GPT-4o's annotation style, not independent scientific correctness. GPT-4o, itself a baseline, is scored against its own output distribution, so ExpStar's automatic margin over GPT-4o is partly forced by construction. The manual verification is described qualitatively without inter-annotator agreement, error rates, or per-sample counts, so it does not break this self-referential loop.

  2. fitted input called prediction [Sec. 4.2.3 (Knowledge Relevance Judgment) with Sec. 5.5 (Ablation Study)]
    "Due to the absence of passage relevance annotations, we automatically construct binary relevance labels for the top-K retrieved passages using GPT-4o. ... Passages scoring >=3 are considered relevant, while those scoring <=2 are deemed irrelevant."

    The <REL>/<NOT REL> tokens are trained on relevance labels produced by GPT-4o, which scores Wikipedia passages against GPT-4o-written ground-truth principles/safety text. The ablation then attributes a BLEU-4 gain of 1.36 and a CIDEr gain of 2.12 to these relevance tokens, but the supervision signal is generated by the same model whose style defines both the references and the training targets. The claimed benefit of knowledge filtering is therefore partly self-referential: no external or human relevance benchmark is used, so the ablation confirms that ExpStar learned GPT-4o's relevance judgments rather than an independently validated selection ability.

full rationale

The core circularity is in the evaluation/training loop for the principle and safety components. GPT-4o creates the scientific principles, safety guidelines, and relevance labels; ExpStar is fine-tuned with token-level cross-entropy on those GPT-4o outputs; and the headline automatic results are measured against those same outputs. Consequently, the reported superiority over 14 LMMs is partially an artifact of training on the reference distribution, especially because GPT-4o itself is one of the baselines. This is not full circularity: the dataset includes 10 student annotators who review and revise samples, the human evaluation in Table 3 is not directly tied to n-gram matching, and the retrieval-augmented architecture is an independent engineering contribution. However, the paper reports no inter-annotator agreement or error-rate statistics, so the manual verification is insufficient to establish that the references are externally accurate. The relevance-token ablation is similarly self-referential because its training labels come from GPT-4o. These issues affect the central claim of 'substantially outperforms 14 leading LMMs' on automatic metrics, warranting a score of 5 rather than a lower score.

Assumptions & free parameters 5 free parameters · 6 assumptions · 1 invented entities

The central claim rests on a chain of human and AI curation choices: GPT-4o produces the gold annotations and relevance labels, students verify a subset, and the model is trained and evaluated on those labels. The main free parameters are the relevance threshold, retrieval width, query fusion weights, and LoRA/DPO hyperparameters. No new physical entities are proposed; the only invented constructs are the four control tokens, which have no independent falsifiable handle outside the in-house dataset.

free parameters (5)
  • relevance threshold = score >= 3
    GPT-4o relevance scores from 1 to 5 are binarized at 3; this changes the training labels for <REL>/<NOT REL> tokens and therefore the model's knowledge selection behavior.
  • retrieval top-K = K = 5
    Default in main experiments; Table 5 compares K = 0, 1, 3, 5, 8. The choice affects how much noise enters the generator.
  • query fusion weights = 0.7 * video embedding + 0.3 * title embedding
    Hand-chosen weighted sum for the multimodal retrieval query in Appendix C.
  • LoRA rank and learning rates = rank 8, SFT lr 1e-4, DPO lr 1e-6
    Training configuration that determines the fine-tuned model behavior; not derived from data.
  • DPO sampling parameters = top-p = 0.9, L candidate outputs
    Sampling configurations for generating preference pairs; the value of L is not specified.
assumptions (6)
  • domain assumption WhisperX transcripts combined with GPT-4o correction and translation preserve the experimental procedure accurately.
    Used in Sec. 3.1-3.2 to build step-level procedures; errors here propagate to all annotations.
  • domain assumption GPT-4o can generate scientifically correct principles and safety guidelines from procedures, and the ten student annotators can verify them.
    The gold standard is the basis of all training and evaluation, per Sec. 3.2-3.3.
  • domain assumption GPT-4o relevance scores between retrieved passages and ground truth identify useful knowledge.
    Sec. 4.2.3 constructs binary relevance labels with GPT-4o; no human validation of relevance labels is reported.
  • domain assumption BLEU, ROUGE, METEOR, CIDEr, and BERTScore are meaningful measures of commentary quality.
    Sec. 5.2.1; n-gram overlap metrics may not capture scientific correctness or safety coverage.
  • domain assumption Wikipedia introductory paragraphs contain the external knowledge needed for all 21 subjects.
    Sec. 4.2.2 and Appendix C; if a subject's key facts are missing, retrieval cannot help.
  • domain assumption The video frame sampling at 1 FPS and CLIP-based retrieval provide enough visual signal to match steps to knowledge.
    Sec. 5.1 and Appendix C; short clips lose temporal detail.
invented entities (1)
  • Retrieval control tokens <RET>, <NOT RET>, <REL>, <NOT REL>
    purpose: Let the model decide whether to retrieve external knowledge and which retrieved passages to use.
    These are internal vocabulary elements introduced for this task; their effectiveness is demonstrated only on the in-house ExpInstruct benchmark, with no external benchmark or formal specification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments." pith.science (2026). https://pith.science/paper/3NVM5OBE

@misc{pith2026250709693,
  author       = {Pith},
  title        = {Pith review of: ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3NVM5OBE}},
  note         = {Machine review of arXiv:2507.09693}
}
read the original abstract

Experiment commentary is crucial in describing the experimental procedures, delving into underlying scientific principles, and incorporating content-related safety guidelines. In practice, human teachers rely heavily on subject-specific expertise and invest significant time preparing such commentary. To address this challenge, we introduce the task of automatic commentary generation across multi-discipline scientific experiments. While recent progress in large multimodal models (LMMs) has demonstrated promising capabilities in video understanding and reasoning, their ability to generate fine-grained and insightful experiment commentary remains largely underexplored. In this paper, we make the following contributions: (i) We construct \textit{ExpInstruct}, the first dataset tailored for experiment commentary generation, featuring over 7\textit{K} step-level commentaries across 21 scientific subjects from 3 core disciplines (\ie, science, healthcare and engineering). Each sample includes procedural descriptions along with potential scientific principles (\eg, chemical equations and physical laws) and safety guidelines. (ii) We propose ExpStar, an automatic experiment commentary generation model that leverages a retrieval-augmented mechanism to adaptively access, evaluate, and utilize external knowledge. (iii) Extensive experiments show that our ExpStar substantially outperforms 14 leading LMMs, which highlights the superiority of our dataset and model. We believe that ExpStar holds great potential for advancing AI-assisted scientific experiment instruction.

Figures

Figures reproduced from arXiv: 2507.09693 by the authors.

Figure 1
Figure 1. Generated commentaries by different models ( [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. (a) Distribution of 21 subjects from 3 core disciplines [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Overview of our proposed ExpStar model. It is built on the Qwen2.5-VL-7B architecture. Special control tokens are [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Qualitative results of Qwen-2.5-VL-7B, GPT-4o and our ExpStar. The [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: A case of the optics experiment. The green and red denote correct and incorrect prediction, respectively. GPT-4o: < Procedure> Insert a dielectric material, such as a glass plate, between the parallel plates while keeping other conditions constant. Observe the deflecti…
Figure 6
Figure 6. Figure 6: A case of the mechanics experiment.The green and red denote correct and incorrect prediction, respectively. Expstar: < Procedure> Place the test tube in a water bath at 60°C for one minute. <Principle> Fehling's reagent requires a specific temperature to form the compl…
Figure 7
Figure 7. Figure 7: A case of biography experiment.The green and red denote correct and incorrect prediction, respectively. 15 [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]
Figure 8
Figure 8. Figure 8: A case of the material engineering experiment. The [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: A failure case. The green and red denote correct and incorrect prediction, respectively. 16 âĂă Qing Li „ [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

75 extracted references · 48 canonical work pages

  1. [1]

    Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrit- twieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P

    Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrit- twieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P. Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael ...

  2. [2]

    Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In Proc. of ICLR. OpenReview.net

  3. [3]

    Ayyubi, Tianqi Liu, Arsha Nagrani, Xudong Lin, Mingda Zhang, Anurag Arnab, Feng Han, Yukun Zhu, Xuande Feng, Kevin Zhang, Jialu Liu, and Shih-Fu Chang

    Hammad A. Ayyubi, Tianqi Liu, Arsha Nagrani, Xudong Lin, Mingda Zhang, Anurag Arnab, Feng Han, Yukun Zhu, Xuande Feng, Kevin Zhang, Jialu Liu, and Shih-Fu Chang. 2024. VIEWS: Entity-Aware News Video Captioning. In Proc. of EMNLP, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, 20220–20239

  4. [4]

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Frontier Large Vision- Language Model with Versatile Abilities. CoRR abs/2308.12966 (2023)

  5. [5]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. 2025. Qwen2. 5-VL Technical Report. arXiv preprint arXiv:2502.13923 (2025)

  6. [6]

    Max Bain, Jaesung Huh, Tengda Han, and Andrew Zisserman. 2023. WhisperX: Time-Accurate Speech Transcription of Long-Form Audio. In Proc. of INTER- SPEECH, Naomi Harte, Julie Carson-Berndsen, and Gareth Jones (Eds.). ISCA, 4489–4493

  7. [7]

    Jiali Chen, Zhenjun Guo, Jiayuan Xie, Yi Cai, and Qing Li. 2023. Deconfounded Visual Question Generation with Causal Inference. In Proc. of ACM MM, Abdul- motaleb El-Saddik, Tao Mei, Rita Cucchiara, Marco Bertini, Diana Patricia Tobon Vallejo, Pradeep K. Atrey, and M. Shamim Hossain (Eds.). ACM, 5132–5142

  8. [8]

    Jiali Chen, Xusen Hei, Yuqi Xue, Yuancheng Wei, Jiayuan Xie, Yi Cai, and Qing Li. 2024. Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor. In Proc. of ACM MM, Jianfei Cai, Mohan S. Kankanhalli, Balakrishnan Prabhakaran, Susanne Boll, Ramanathan Subramanian, Liang Zheng, Vivek K. Singh, Pablo César, Lexing ...

Show all 75 references
  1. [9]

    Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Lin Bin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, and Jiaqi Wang. 2024. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions. InProc. of Neu...

  2. [10]

    Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che. 2024. M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain- of-Thought. In Proc. of ACL, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, 8199–8221

  3. [11]

    Yuxiao Chen, Kai Li, Wentao Bao, Deep Patel, Yu Kong, Martin Renqiang Min, and Dimitris N. Metaxas. 2024. Learning to Localize Actions in Instructional Videos with LLM-Based Multi-pathway Text-Video Alignment. In Proc. of ECCV, Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Ru...

  4. [13]

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, Lixin Gu, Xuehui Wang, Qingyun Li, Yimin Ren, Zixuan Chen, Jiapeng Luo, Jiahao Wang, Tan Jiang, Bo Wang, Conghui He, Botian Shi, Xingcheng Zhang, Han Lv, Yi...

  5. [14]

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. 2023. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks. CoRR a...

  6. [15]

    Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing. 2024. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs. CoRR abs/2406.07476 (2024)

  7. [16]

    Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2024. Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering. CoRR abs/2411.16863 (2024)

  8. [17]

    Denkowski and Alon Lavie

    Michael J. Denkowski and Alon Lavie. 2014. Meteor Universal: Language Specific Translation Evaluation for Any Target Language. In Proc. of ACL Workshop . 376–380

  9. [18]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sra- vankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston...

  10. [19]

    David Gooding, Trevor Pinch, and Simon Schaffer. 1989. The uses of experiment: Studies in the natural sciences

  11. [20]

    Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, and Andrew Zis- serman. 2023. AutoAD II: The Sequel - Who, When, and What in Movie Audio Description. In Proc. of ICCV. IEEE, 13599–13609

  12. [21]

    Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, and Andrew Zis- serman. 2023. AutoAD: Movie Description in Context. In Proc. of CVPR. IEEE, 18930–18940

  13. [22]

    Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, and Andrew Zisser- man. 2024. AutoAD III: The Prequel - Back to the Pixels. In Proc. of CVPR. IEEE, 18164–18174

  14. [23]

    Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wanrong Zhu, Jiachen Li, Yue Fan, Jianfeng Wang, Linjie Li, Zhengyuan Yang, Kevin Lin, William Yang Wang, Lijuan Wang, and Xin Eric Wang. 2024. MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos. CoRR...

  15. [24]

    Wenyi Hong, Weihan Wang, Ming Ding, Wenmeng Yu, Qingsong Lv, Yan Wang, Yean Cheng, Shiyu Huang, Junhui Ji, Zhao Xue, Lei Zhao, Zhuoyi Yang, Xiaotao Gu, Xiaohan Zhang, Guanyu Feng, Da Yin, Zihan Wang, Ji Qi, Xixuan Song, Peng Zhang, Debing Liu, Bin Xu, Juanzi Li, Yuxiao Dong, a...

  16. [25]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In Proc. of ICLR. OpenReview.net

  17. [26]

    Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Yuanhan Zhang, Xiang Yue, Bo Li, and Ziwei Liu. 2025. Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos. CoRR abs/2501.13826 (2025)

  18. [27]

    David Klahr, Anne L Fay, and Kevin Dunbar. 1993. Heuristics for scientific experimentation: A developmental study. Cognitive psychology 25, 1 (1993), 9 111–146

  19. [28]

    Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles

  20. [29]

    Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. 2024. LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models. CoRR abs/2407.07895 (2024)

  21. [30]

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Lou, Limin Wang, and Yu Qiao. 2024. MVBench: A Comprehensive Multi-modal Video Understanding Benchmark. In Proc. of CVPR. IEEE, 22195–22206

  22. [31]

    Yanwei Li, Chengyao Wang, and Jiaya Jia. 2024. LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models. In Proc. of ECCV, Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol (Eds.), Vol. 15104. Springer, 323–340

  23. [32]

    Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. 2024. Video-LLaVA: Learning United Visual Representation by Alignment Before Pro- jection. In Proc. of EMNLP, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Lingui...

  24. [33]

    Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Proc. of ACL Workshop. 74–81

  25. [34]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruc- tion Tuning. In Proc. of NeurIPS, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)

  26. [35]

    Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan. 2024. DeepSeek-VL: Towards Real-World Vision-Language Understanding. CoRR abs/2403.05525 (2024)

  27. [36]

    Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. In Proc. of NeurIPS, Sanmi Koyejo, S. Mohamed, A. Agarw...

  28. [37]

    Thomas Mensink, Jasper R. R. Uijlings, Lluís Castrejón, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araújo, and Vittorio Ferrari. 2023. Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories. In Proc. of ICCV. IEEE, 3090–3101

  29. [38]

    OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023)

  30. [39]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proc. of ACL . 311–318

  31. [40]

    Ji Qi, Jifan Yu, Teng Tu, Kunyu Gao, Yifan Xu, Xinyu Guan, Xiaozhi Wang, Bin Xu, Lei Hou, Juanzi Li, and Jie Tang. 2023. GOAL: A Challenging Knowledge- grounded Video Captioning Benchmark for Real-time Soccer Commentary Gen- eration. In Proc. of CIKM, Ingo Frommholz, Frank Hop...

  32. [41]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proc. of IC...

  33. [42]

    Manning, Stefano Ermon, and Chelsea Finn

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. InProc. of NeurIPS, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hard...

  34. [43]

    Jiayuan Rao, Haoning Wu, Hao Jiang, Ya Zhang, Yanfeng Wang, and Weidi Xie

  35. [44]

    Jiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang, and Weidi Xie. 2024. MatchTime: Towards Automatic Soccer Game Commentary Generation. In Proc. of EMNLP, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Associa- tion for Computational Linguistics, 1671–1685

  36. [45]

    Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. 2022. Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities. In Proc. of CVPR. IEEE, 21064–21074

  37. [46]

    Yan Shu, Peitian Zhang, Zheng Liu, Minghao Qin, Junjie Zhou, Tie-Jun Huang, and Bo Zhao. 2024. Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding. CoRR abs/2409.14485 (2024)

  38. [47]

    Quan Sun, Jinsheng Wang, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, and Xinlong Wang. 2024. EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters. CoRR abs/2402.04252 (2024)

  39. [48]

    Lawrence Zitnick, and Devi Parikh

    Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. CIDEr: Consensus-based image description evaluation. In Proc. of CVPR. 4566–4575

  40. [49]

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. Qwen2-VL: Enhancing Vision-Language Mode...

  41. [50]

    Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Yu Qiao. 2024. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation. In Proc. of IC...

  42. [51]

    Hongchen Wei, Zhihong Tan, Yaosi Hu, Chang Wen Chen, and Zhenzhong Chen

  43. [52]

    Stephen Wooding, Alison Cullinane, and Sibel Erduran. 2020. Supporting the teaching of scientific methods in practical science. (2020)

  44. [53]

    Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, Zhenda Xie, Yu Wu, Kai Hu, Jiawei Wang, Yaofeng Sun, Yukun Li, Yishi Piao, Kang Guan, Aixin Liu, Xin Xie, Yuxiang You, Kai Dong, Xingkai Yu, Haowei Zhang,...

  45. [54]

    Zihui Xue, Joungbin An, Xitong Yang, and Kristen Grauman. 2024. Progress- Aware Video Frame Captioning. CoRR abs/2412.02071 (2024)

  46. [55]

    Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont- Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. 2023. Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning. In Proc. of CVPR. IEEE, 10714–10726

  47. [56]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Cheng- peng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jin...

  48. [57]

    Dingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang, Tiezheng Ge, Bo Zheng, and Qin Jin. 2024. Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline. In Proc. of ACL, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computatio...

  49. [58]

    Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2024. MMM...

  50. [59]

    Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Mingdao Liu, Minlie Huang, Peng...

  51. [60]

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sig- moid Loss for Language Image Pre-Training. InProc. of ICCV. IEEE, 11941–11952

  52. [61]

    Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction- tuned Audio-Visual Language Model for Video Understanding. InProc. of EMNLP, Yansong Feng and Els Lefever (Eds.). Association for Computational Linguistics, 543–553

  53. [62]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi

  54. [63]

    Yilun Zhao, Lujing Xie, Haowei Zhang, Guo Gan, Yitao Long, Zhiyuan Hu, Tongyan Hu, Weiyuan Chen, Chuhan Li, Junyang Song, Zhijian Xu, Chengye Wang, Weifeng Pan, Ziyao Shangguan, Xiangru Tang, Zhenwen Liang, Yixin Liu, Chen Zhao, and Arman Cohan. 2025. MMVU: Measuring Expert-Le...

  55. [64]

    Xingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan, Austin Myers, Xuehan Xiong, Arsha Nagrani, and Cordelia Schmid. 2024. Streaming Dense Video Captioning. In Proc. of CVPR. IEEE, 18243–18252

  56. [65]

    id": "",

    Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Serena Yeung-Levy, and Xide Xia. 2024. Apollo: An Exploration of Video Understanding in Large Multimodal Models. CoRR abs/2412.10360 ...

  57. [70]

    Generate the "summary" field (experiment summary)

  58. [71]

    Summarize the experiment procedure into the "steps" field, with numbered steps starting from 1

  59. [72]

    procedure

    Generate a new field "procedure" (a summary description of each step)

  60. [73]

    Each step’s ASR fragments should consist of a continuous sequence of positive integer numbers (e.g., [2,3,4,5]), without any jumps or regressions (e.g., [2,3,7,5] is not allowed)

    Generate the "ASR_id" field, which must strictly follow the order of the ASR fragments. Each step’s ASR fragments should consist of a continuous sequence of positive integer numbers (e.g., [2,3,4,5]), without any jumps or regressions (e.g., [2,3,7,5] is not allowed)

  61. [74]

    ASR_id" only contains continuous positive integer numbers, without any skips or backtracking. Template: {

    Each ASR fragment must only appear in one step and cannot be repeated in different steps. Note: - Only return the JSON result, no additional explanation is required. - Ensure that the "ASR_id" only contains continuous positive integer numbers, without any skips or backtracking...

  62. [75]

    <Principle>: to annotate explanations of scientific principles or theories

  63. [76]

    Detecting Sugars, Fats and Proteins in Biological Tissues

    <Safety>: to annotate safety precautions or warnings. Please generate a detailed procedure for the current step of the biology experiment titled"Detecting Sugars, Fats and Proteins in Biological Tissues". The images provided are equidistant samples taken from a video. Keep you...

  64. [2017]

    Dense-Captioning Events in Videos. In Proc. of ICCV . IEEE Computer Society, 706–715

  65. [2020]

    BERTScore: Evaluating Text Generation with BERT. In Proc. of ICLR . OpenReview.net

  66. [2024]

    CoRR abs/2412.01820 (2024)

    Towards Universal Soccer Video Understanding. CoRR abs/2412.01820 (2024)

  67. [2025]

    CoRR abs/2502.15393 (2025)

    LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models. CoRR abs/2502.15393 (2025)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.