Pith. sign in

Paper Citation Record · LEDGER

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2412.04447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04447 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:29:11.117135Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:49:52.033876Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec888612-2f53-4d8c-93fc-a4f9ef993653 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:11.117135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:11.117135Z digest=sha256:c568ce8a2b44dc1aa5a49890729ea40f69eb678d3617e4c5b5064339b64a5d74

Observation 5900ee82-62fd-4757-bae9-cfd7a231a0fc · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.946680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.946680Z digest=sha256:5bcaa8597076f27925fd6759460f2a154d711530fef9a891cb93d9913eccbaeb

Observation fed62ce0-a034-485a-bf36-63085f2fbbd6 · inbound

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning cites this paper.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.993024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.993024Z digest=sha256:cf1cde1a4ce905a8019d7e84fa50ec96e1dff1b870383729f9c3d0ad6d55bcd6

Observation 0cb8f075-81df-44af-a819-b23870384605 · inbound

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents cites this paper.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.779856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.779856Z digest=sha256:6a7ec700eb78f59fd00c90f2381f0e8eb38271ee8629bb47b6e9fd06400be158

Observation 0a77f47e-7049-499a-a220-3aa2d3e26933 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.033928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:d33c8c66258e10aff10c9b43336e929ed6f6a3da21852711daac1a14541e083a

Observation 9b6c2843-4cbc-4135-a943-43880b395e63 · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.150961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.150961Z digest=sha256:df157c143fdbe941fbd552bc31d69a2cb1a3be26398a8a6b689397dc36dfd91a

Observation f73b685f-a738-490e-a587-fa771a72ec12 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.708154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:b245c6c700ec023fbd08dbd7a8ba28cf85e89e0d1f2715ba718feba46fb1bcd7

Observation aabe2527-e28c-4601-9756-9fb3e63a8cbe · inbound

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models cites this paper.

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:28.347894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:27:18.284698Z digest=sha256:c95cfa9abf441f6aeb5649b701a0ac7986541963309fffc92e1a3009e5f37c8f

Observation 44818236-e821-4b69-a825-9897082ce3de · inbound

Long-Horizon Manipulation via Trace-Conditioned VLA Planning cites this paper.

Long-Horizon Manipulation via Trace-Conditioned VLA Planning EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:39.077637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T21:10:03.554504Z digest=sha256:52f5bebf526632efa13f07b7bcbd60a4cfca52a1d2befc650f5b26027d1f4e9f

Observation 61229f3d-2fe9-42ee-a71a-d73c8c0e3e7b · inbound

RECIPE: Procedural Planning via Grounding in Instructional Video cites this paper.

RECIPE: Procedural Planning via Grounding in Instructional Video EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:28:05.566482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T06:24:50.714859Z digest=sha256:a1e1b486abe0b3875b61decc22a2465fa13e0443f7c2786adef08e4b3aeef955

Observation 2117b87c-cca6-43bf-8b2d-c1ace6e4cad7 · inbound

RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes cites this paper.

RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.452888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T18:56:37.114948Z digest=sha256:631dc1d6bdfdc7962676daab2dc57c308c9b2202d1ddeb7d3d3db4f0cdfbe332

Observation 037bf477-997e-4777-a198-11fab6a70ab1 · inbound

When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA cites this paper.

When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:07:27.135883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:26:47.393632Z digest=sha256:5f580a5aca00caac61a43ca607a61af4263f375e4bee064fa9ad1ecf41e14eca

Observation 0f306271-80f7-4e66-b6de-0a52b681faef · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.230001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:059e1c4e61acd8ac1ce1e6cee0215935c0a94b81e3bfa265c2205285b663553a

Observation adde64d9-d24a-4b7c-a1ed-84ed942ec750 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:b05df13bb0b9b109e0c54ed1bfff91aad436d75a3c716206cb44184429acc999

Observation 05c4a96d-0dd5-4fcd-b5f9-d2b44778827e · inbound

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy cites this paper.

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.035634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T04:52:51.524022Z digest=sha256:fa9c779c14d7eb852ee29d4c136aec75c6320c6b7f33852e74055e28f6712523