Pith. sign in

REVIEW 22 cited by

EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.06722 v3 pith:4GMVRD56 submitted 2023-12-11 cs.CV cs.CLcs.RO

classification cs.CVcs.CLcs.RO
keywords mllmsplanningegoplan-benchhuman-levelmultimodalbenchmarkcapabilitiesevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The pursuit of artificial general intelligence (AGI) has been accelerated by Multimodal Large Language Models (MLLMs), which exhibit superior reasoning, generalization capabilities, and proficiency in processing multimodal inputs. A crucial milestone in the evolution of AGI is the attainment of human-level planning, a fundamental ability for making informed decisions in complex environments, and solving a wide range of real-world problems. Despite the impressive advancements in MLLMs, a question remains: How far are current MLLMs from achieving human-level planning? To shed light on this question, we introduce EgoPlan-Bench, a comprehensive benchmark to evaluate the planning abilities of MLLMs in real-world scenarios from an egocentric perspective, mirroring human perception. EgoPlan-Bench emphasizes the evaluation of planning capabilities of MLLMs, featuring realistic tasks, diverse action plans, and intricate visual observations. Our rigorous evaluation of a wide range of MLLMs reveals that EgoPlan-Bench poses significant challenges, highlighting a substantial scope for improvement in MLLMs to achieve human-level task planning. To facilitate this advancement, we further present EgoPlan-IT, a specialized instruction-tuning dataset that effectively enhances model performance on EgoPlan-Bench. We have made all codes, data, and a maintained benchmark leaderboard available to advance future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    EgoSAT is the first benchmark unifying retrospective, online, and prospective reasoning tasks in egocentric streaming video to evaluate VLMs, revealing struggles with temporal modeling and mis-calibration.

  2. SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    SVI-Bench provides 35K hours of sports video with 9 tasks across four cognitive levels, revealing models drop from ~74% on action QA to 5% on agentic evidence integration.

  3. EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning

    cs.CV 2025-11 conditional novelty 7.0 of 10

    EgoVITA, a GRPO-based plan-then-verify framework with dense visual-grounding rewards, improves egocentric video reasoning by up to +7.7 points and keeps exocentric video performance intact.

  4. RoBoSR: Structured Scene Representations for Embodied Robotic Reasoning

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    RoBoSR uses structured object-centric scene graphs as an intermediate representation to enable causal reasoning and subtask planning in embodied robotics, outperforming baselines on benchmarks and real demos.

  5. SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    SVI-Bench is a 35K-hour sports video benchmark with 9 tasks across four cognitive pillars that reveals multimodal models drop from ~73% on action QA to 5% on agentic evidence-gathering tasks.

  6. PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    PRTS pretrains VLA models with contrastive goal-conditioned RL to embed goal-reachability probabilities from offline data, yielding SOTA results on robotic benchmarks especially for long-horizon and novel instructions.

  7. SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    SocialGrid benchmark shows even top LLMs achieve below 60% in embodied planning and task completion, with deception detection near random chance regardless of model scale.

  8. LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Frontier LLMs exceed human performance on easy Wikipedia navigation tasks but finish fewer than 25% of hard games, with failures driven by looping and an inability to replan after mistakes.

  9. Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A 6K-question embodied-reasoning benchmark plus a flow-matching action tokenizer let one 3B vision-language model reason and manipulate better than continuous- or discrete-action VLA baselines.

  10. MiMo-Embodied: X-Embodied Foundation Model Technical Report

    cs.RO 2025-11 unverdicted novelty 6.0 of 10

    MiMo-Embodied is a single foundation model that achieves state-of-the-art results on 17 embodied AI benchmarks and 12 autonomous driving benchmarks through multi-stage learning, curated data, and CoT/RL fine-tuning th...

  11. Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    An MLLM trained with auxiliary goal-prediction tasks and multi-token prediction achieves SOTA on COIN and CrossTask visual planning and matches SOTA on Ego4D LTA.

  12. DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DisCo assigns each visual token to a unique concept from the caption and aligns its attention across frames, improving video MLLM accuracy and token efficiency.

  13. RoboBrain 2.0 Technical Report

    cs.RO 2025-07 conditional novelty 6.0 of 10

    RoboBrain 2.0, a 7B/32B embodied vision-language model built on Qwen2.5-VL, reports state-of-the-art or near-top scores on several spatial and temporal reasoning benchmarks for robotics.

  14. GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GRPO-CARE improves answer accuracy and reasoning coherence over standard GRPO on a new video reasoning benchmark, with a 6.7 point gain on the hardest level and a 24.5 point higher consistency rate.

  15. A Survey on Vision-Language-Action Models for Embodied AI

    cs.RO 2024-05 unverdicted novelty 6.0 of 10

    This is the first survey on vision-language-action models, providing a taxonomy across three lines, plus summaries of datasets, simulators, benchmarks, challenges, and future directions in embodied AI.

  16. Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    Embodied-R1.5 is an 8B EFM achieving SOTA on 16 of 24 embodied VLM benchmarks, fine-tunable to outperform leading VLAs, with claimed zero-shot real-robot generalization.

  17. EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    EgoCoT-Bench provides 3,172 verifiable QA pairs across perception, anticipation, and reasoning tasks on egocentric videos, revealing that many MLLMs give answer-correct but evidence-inconsistent explanations.

  18. Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

    cs.RO 2025-08 unverdicted novelty 5.0 of 10

    This survey organizes large VLM-based VLA models for robotic manipulation into monolithic and hierarchical paradigms, reviews their integrations and datasets, and outlines future directions.

  19. ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A 7B multimodal model that fuses audio and visual signals with explicit timestamps achieves strong measured comprehension of real-world short videos on the authors' new ShortVid-Bench benchmark.

  20. ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

    cs.CV 2025-07 unverdicted novelty 5.0 of 10

    ThinkAct introduces reinforced visual latent planning in a dual VLA system to enable better long-horizon reasoning and adaptation for embodied tasks.

  21. Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

    cs.RO 2026-06 unverdicted novelty 4.0 of 10

    An 8B embodied foundation model trained on 15B tokens with multi-task RL and a Planner-Grounder-Corrector loop claims SOTA on 16/24 embodied VLM benchmarks and strong VLA/real-robot transfer.

  22. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

Pith tools