Pith. sign in

REVIEW 13 cited by

MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08407 v3 pith:NN6SBON2 submitted 2024-06-12 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords mmworldvideosmllmsevaluationmodelsmulti-disciplinemulti-facetedunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models" -- interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal medium, as they encapsulate rich representations of real-world dynamics and causalities. To this end, we introduce MMWorld, a new benchmark for multi-discipline, multi-faceted multimodal video understanding. MMWorld distinguishes itself from previous video understanding benchmarks with two unique advantages: (1) multi-discipline, covering various disciplines that often require domain expertise for comprehensive understanding; (2) multi-faceted reasoning, including explanation, counterfactual thinking, future prediction, etc. MMWorld consists of a human-annotated dataset to evaluate MLLMs with questions about the whole videos and a synthetic dataset to analyze MLLMs within a single modality of perception. Together, MMWorld encompasses 1,910 videos across seven broad disciplines and 69 subdisciplines, complete with 6,627 question-answer pairs and associated captions. The evaluation includes 2 proprietary and 10 open-source MLLMs, which struggle on MMWorld (e.g., GPT-4V performs the best with only 52.3\% accuracy), showing large room for improvement. Further ablation studies reveal other interesting findings such as models' different skill sets from humans. We hope MMWorld can serve as an essential step towards world model evaluation in videos.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

    cs.CV 2026-08 conditional novelty 6.0 of 10

    LVLM judges are near chance when asked to pick the correctly ordered version of an image sequence, and this temporal blindness persists after fine-tuning and at larger scale.

  2. HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A hierarchical benchmark for multimodal models on human-centric visual understanding finds frontier models average under 60% and miss question-uncued visual evidence, with test-time scaling helping only marginally.

  3. VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    VLM4D benchmarks spatiotemporal reasoning in VLMs and finds large gaps versus humans, with proposed methods showing partial improvement.

  4. "Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    The authors construct the first Chinese youth-toxicity dataset, show that youth and adult perceptions of toxic language diverge, and report that adding contextual meta information improves detection accuracy.

  5. ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ExpStar, with a new 7,714-sample ExpInstruct dataset, generates step-level scientific experiment commentary including procedures, principles, and safety guidelines.

  6. VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    VRBench is a benchmark of 960 long narrative videos with 8,243 human-written multi-step questions, plus a two-level evaluation of answer accuracy and reasoning quality for 31 large models.

  7. VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new video benchmark asks models to reorder shuffled clips from everyday tasks; most large video language models perform near chance, and a recognize-then-reason prompt decomposition gives large gains.

  8. Multi-Agent System for Comprehensive Soccer Understanding

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The authors present SoccerBench (10K multimodal soccer QA pairs), SoccerWiki (a soccer knowledge base), and SoccerAgent, a multi-agent system that outperforms general multimodal LLMs on the benchmark.

  9. Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

    cs.CV 2025-04 conditional novelty 6.0 of 10

    Video-MMLU: a 1,065-video lecture benchmark where most AI video models score 10-50%, but text-only models answer 40% of quiz questions without video.

  10. EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

    cs.AI 2024-12 conditional novelty 6.0 of 10

    EgoPlan-Bench2 evaluates multimodal LLMs on next-action planning across 24 egocentric real-world scenarios, finding most models near random chance, and a prompt-based method raises GPT-4V's accuracy to 43 percent on a...

  11. VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A new benchmark, VF-Eval, measures how well multimodal LLMs check, detect, and reason about errors in AI-generated videos, and shows frontier models remain far below human performance.

  12. From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A structured survey of recent jailbreak attacks and defenses across LLMs, multimodal LLMs, and agents, with taxonomies for methods, datasets, metrics, and defenses.

  13. VideoLLM Benchmarks and Evaluation: A Survey

    cs.CV 2025-05 unverdicted novelty 1.0 of 10

    A survey of VideoLLM benchmarks and evaluation protocols that organizes known datasets and metrics, with proposed future benchmark designs.

Pith tools