Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T20:23:18.667677Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2606.07962.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T20:23:18.667677Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bf0541a9-2202-44c7-b345-c270330acc71 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? pi_0.5: a vision-language-action model with open-world generalization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 910f267a-b047-4ef7-bec9-d3769cbb93ba · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? OpenVLA: An Open-Source Vision-Language-Action Model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c9e4979-3a23-41b0-a78b-2608e2ed50db · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bcfff88e-07d6-46e0-a3b7-4e8cef209b65 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Re-align: Aligning vision language models via retrieval-augmented direct preference optimization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af4b3146-4303-475e-9324-f5f53971d929 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Emerging Properties in Unified Multimodal Pretraining
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ccbfccdb-0628-4274-974c-5bc0f3215510 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Representation alignment for generation: Training diffusion transformers is easier than you think
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ae5149-4b9f-4832-8314-29adaa73d87e · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Weakly-supervised 3d spatial reasoning for text-based visual question answering.IEEE Transactions on Image Processing, 32:3367–3382, 2023
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f7d9f75-087b-4e17-9d97-477d99b2ca05 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? V-fat: Benchmarking visual fidelity against text-bias, 2026
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0862f782-6095-41ad-8478-f4dd03a7d89a · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mitigating hallucination in visual-language models via re-balancing contrastive decoding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 170f310b-dafd-4d5d-841c-4c505427b94b · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ebb366bc-11ca-4c02-9194-c0a6c28de517 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Robust multimodal large language models against modality conflict
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b412b7a-76e0-49e3-a476-e5b9e2255f71 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mitigating modality prior-induced hallucinations in multimodal large language models via deciphering attention causality
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4c6dcbe-4472-44a1-bf04-71ca73eb9610 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Quantum physics intelligent question answering (q&a) system based on retrieval-augmented generation.Concurrency and Computation: Practice and Experience, 38(1):e70379, 2026
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbd94af2-b3eb-4169-8e70-07eacc036cfe · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81065ccc-859e-4c68-bc0f-c883bfe7f79d · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Large vision-language model alignment and misalignment: A survey through the lens of explainability
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b397e0b1-1b2e-4a2a-bbc7-90c5013a1e00 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bd0110c8-f3d1-4d28-82ae-c5374eeae13d · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a4ad08b-812e-4496-a1e4-4a1e60a9d0ca · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mm-vet: Evaluating large multimodal models for integrated capabilities
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56e30cd9-bb06-48ab-a870-2bc5984bc98a · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c791b204-4b76-4e26-be4b-612271297f20 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a07a93fa-9fbd-43d8-81b9-5768c7d6d442 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.Advances in Neural Information Processing Systems, 37:89098–89124, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e7df6d-df4c-4bd7-b012-c1704e6591d5 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa802e29-c999-41bc-a413-89a1db1cb71f · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Lvbench: An extreme long video understanding benchmark
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef33675a-6350-4390-ad5c-53bfea09d88e · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93a41931-5898-48ea-bf7f-a2ce6852876b · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd69f7f-2839-4ff7-9249-b74e8c26b73d · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mlvu: Benchmarking multi-task long video understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 684f383e-b0a0-49ef-8502-2c42132a79c8 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Behavior in habitat 2.0: Simulator-independent logical task description for benchmarking embodied ai agents, 2022
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24df9feb-bf49-44df-9b0a-fa33147d580e · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32794ebc-2422-43b8-8482-139a441032b6 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Affordance Benchmark for MLLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7e053c8f-27e8-4918-a40a-43d004865a00 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Phystoolbench: Benchmarking physical tool understanding for mllms
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c8c42eb3-de11-4ba8-ae44-72d02d4eedc3 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ee49f16-c076-4817-9bae-c03b8296568f · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? QuantiPhy: A quantitative benchmark evaluating physical reasoning abilities of vision-language models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 012b3919-45ed-4e1e-b211-7b779cd4bde7 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958, 2026
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5541e0bd-d6fb-49d9-9bbf-de7e114fb2f1 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 846e9bad-35a8-4e0c-a7be-4c49aa3d42c2 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Intphys 2019: A benchmark for visual intuitive physics understanding.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9):5016–5025, 2021
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21aa4551-0cc1-4dc5-961f-c27edbaf7cab · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29da7131-22a1-4951-8475-d270991ad4c0 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Physion: Evaluating Physical Prediction from Vision in Humans and Machines
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3615da6-f8a3-4fba-9760-37cfad50b70a · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fa09806-9e8c-4453-b0c5-4bdba8143b62 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? ComPhy: Compositional Physical Reasoning of Objects and Events from Videos
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 178f3444-3905-4505-bef9-d75a1c981d15 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? ContPhy: Continuum Physical Concept Learning and Reasoning from Videos
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 847de4d2-6bbe-4c0c-800b-16f2fc34a3f8 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 22bbed5b-1b7d-4d4c-87bf-c95da090f955 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d8f0afaf-4af6-4b32-b691-5135409340bd · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50f1a3af-5596-4128-ad9d-c0b632926066 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 46bb0bf4-c1d4-4ca0-9f42-5f6873fc888d · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environments, 2025
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 247f534e-37e3-49ec-add2-34809bed93d8 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Compositional physical reasoning of objects and events from videos
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6d2c36-e0cc-480d-af49-4f313ed60854 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Craft: A benchmark for causal reasoning about forces and interactions
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d1a870a-19fc-44cb-8043-7c4b6b8886b8 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0677d9f-3162-40e3-93f3-a10432fe406a · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? STAR: A Benchmark for Situated Reasoning in Real-World Videos
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 63c365ef-b8cf-4732-b275-be649d161d0e · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748–42761, 2023
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3921d17e-1f03-4205-b931-c1c531c532e6 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Seeing sound, hearing sight: Uncovering modality bias and conflict of ai models in sound localization
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2449d16-3e95-47f8-8557-8c265af7382f · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Vmbench: A benchmark for perception-aligned video motion generation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9b47191-b58c-42da-af36-52e7a56a3d64 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Unipc: A unified predictor-corrector framework for fast sampling of diffusion models.Advances in Neural Information Processing Systems, 36:49842–49869
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9bff6b47-c3d9-478a-9ff8-1ef17734dc33 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Physunibench: An undergraduate-level physics reasoning benchmark for multimodal models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 581918d4-a989-411e-b8e9-4d7e39d59181 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? GPT-4 Technical Report
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f8f1d2a6-44ba-42b6-8af2-59fa58f21aa8 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Gemini: A Family of Highly Capable Multimodal Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f21d63bb-17f0-49f3-9cdb-6988fa0df3c2 · outbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.