Pith. sign in

REVIEW 9 cited by

VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.05355 v1 pith:KSYAV4EB submitted 2024-07-07 cs.CV cs.CL

classification cs.CVcs.CL
keywords annotationdatasetshumanvideoactivemllmstoolvideos
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multimodal large language models (MLLMs) are flourishing, but mainly focus on images with less attention than videos, especially in sub-fields such as prompt engineering, video chain-of-thought (CoT), and instruction tuning on videos. Therefore, we try to explore the collection of CoT datasets in videos to lead to video OpenQA and improve the reasoning ability of MLLMs. Unfortunately, making such video CoT datasets is not an easy task. Given that human annotation is too cumbersome and expensive, while machine-generated is not reliable due to the hallucination issue, we develop an automatic annotation tool that combines machine and human experts, under the active learning paradigm. Active learning is an interactive strategy between the model and human experts, in this way, the workload of human labeling can be reduced and the quality of the dataset can be guaranteed. With the help of the automatic annotation tool, we strive to contribute three datasets, namely VideoCoT, TopicQA, TopicCoT. Furthermore, we propose a simple but effective benchmark based on the collected datasets, which exploits CoT to maximize the complex reasoning capabilities of MLLMs. Extensive experiments demonstrate the effectiveness our solution.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flow reorganization and transport enhancement in two-dimensional horizontal convection near a density extremum

    physics.flu-dyn 2025-08 unverdicted novelty 7.0 of 10

    Horizontal convection with a water-like density maximum reorganizes into a single roll with full-depth plumes, and heat transport scales as Nu ~ Ra^{1/4} to Ra^{1/3}, faster than the classical Ra^{1/5} law.

  2. Weak-to-Strong On-Policy Distillation

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A strong LLM is improved by distilling from the logit difference of two weaker models instead of from a stronger teacher.

  3. AdsQA: Towards Advertisement Video Understanding

    cs.CV 2025-09 conditional novelty 6.0 of 10

    AdsQA adds an ad-video question-answering benchmark and ReAd-R, a GRPO-trained model that beats 7B baselines but not larger closed models.

  4. RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    RadarQA introduces a specialized MLLM and a 70,000-example dataset for descriptive weather radar forecast quality analysis, outperforming general-purpose MLLMs on its own benchmark.

  5. DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A video-language model fine-tuned on a new defect-annotated dataset detects AI-generated videos from unseen generators with 76.7% accuracy and gives written explanations, though the test set is small and the dataset i...

  6. MINERVA: Evaluating Complex Video Reasoning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MINERVA provides 1,515 multi-step video QA questions with human reasoning traces; frontier models score far below humans and fail mainly on temporal localization and perception.

  7. VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

    cs.CV 2024-11 conditional novelty 6.0 of 10

    VideoEspresso is a large automatically generated video QA dataset with chain-of-thought reasoning and core frame selection, plus a hybrid LVLM framework that reportedly outperforms baselines on its own benchmark.

  8. Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

    cs.SD 2026-08 conditional novelty 5.0 of 10

    AudioRubrics uses evolving, audio-grounded rubric rewards from a powerful judge model to improve reinforcement learning for audio reasoning, beating baselines on MMAU, MMAR, and MMSU.

  9. PySeizure: A single machine learning classifier framework to detect seizures in diverse datasets

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A unified EEG seizure-detection framework with standardized preprocessing and majority voting reaches within-dataset AUC 0.86-0.90 and cross-dataset AUC 0.615-0.762 across CHB-MIT and TUSZ.

Pith tools