Pith. sign in

Paper Citation Record · LEDGER

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?

As of 5 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2606.07962.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07962 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T20:23:18.667677Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact22
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf0541a9-2202-44c7-b345-c270330acc71 · outbound

This paper cites pi_0.5: a vision-language-action model with open-world generalization.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? pi_0.5: a vision-language-action model with open-world generalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:b30e82073e5ad3b65e67d8483e07a409b81e358a6ebf7016c894220c095eebf7

Observation 910f267a-b047-4ef7-bec9-d3769cbb93ba · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:27:22.640801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:14041513f0e5958f56d71cef0e20646bd5fede0c3682be00f229a14539a2b901

Observation 2c9e4979-3a23-41b0-a78b-2608e2ed50db · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:27:22.596414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:c7d4fb23e46faa2746195e83e5833e30c50df3d9d9287eb00355f52fef7339f4

Observation bcfff88e-07d6-46e0-a3b7-4e8cef209b65 · outbound

This paper cites Re-align: Aligning vision language models via retrieval-augmented direct preference optimization.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Re-align: Aligning vision language models via retrieval-augmented direct preference optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:401f7037ef98ca6e41f4f32dbe13146650d7b6ddc14f016162b593ea550f0337

Observation af4b3146-4303-475e-9324-f5f53971d929 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Emerging Properties in Unified Multimodal Pretraining

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:27:22.632913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:0cebea3b1e80cddac5138173bbd6195769909dace51810919327a2f001db4fcc

Observation ccbfccdb-0628-4274-974c-5bc0f3215510 · outbound

This paper cites Representation alignment for generation: Training diffusion transformers is easier than you think.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Representation alignment for generation: Training diffusion transformers is easier than you think

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:e98edd7adf1b53a746a2ae8f2736d39ab950b9b2d4f63059eff539fe9b6e4020

Observation 03ae5149-4b9f-4832-8314-29adaa73d87e · outbound

This paper cites Weakly-supervised 3d spatial reasoning for text-based visual question answering.IEEE Transactions on Image Processing, 32:3367–3382, 2023.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Weakly-supervised 3d spatial reasoning for text-based visual question answering.IEEE Transactions on Image Processing, 32:3367–3382, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:cd2f2d75dbcd295a1720dce97b20cf566f0ea98a367d2850c25874d6d7ee1409

Observation 7f7d9f75-087b-4e17-9d97-477d99b2ca05 · outbound

This paper cites V-fat: Benchmarking visual fidelity against text-bias, 2026.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? V-fat: Benchmarking visual fidelity against text-bias, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:48fe9efd07845d404d95f91624dcbe1d7312be4a0dd98795735bc0d48aa66e31

Observation 0862f782-6095-41ad-8478-f4dd03a7d89a · outbound

This paper cites Mitigating hallucination in visual-language models via re-balancing contrastive decoding.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mitigating hallucination in visual-language models via re-balancing contrastive decoding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:17c5df1539f6fa585639995a8993c3732cabfa3c50d301b4def5ef21a00d1012

Observation 170f310b-dafd-4d5d-841c-4c505427b94b · outbound

This paper cites DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:27:22.636691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:4d0f7b9d4ea60b396c83df8a9e596e6b2cfbcd0492d689517f9f47cdabee1505

Observation ebb366bc-11ca-4c02-9194-c0a6c28de517 · outbound

This paper cites Robust multimodal large language models against modality conflict.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Robust multimodal large language models against modality conflict

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:541c01ca46bb11a0dde7870c407c25de34146f6c9b38c70ae70577af33834816

Observation 0b412b7a-76e0-49e3-a476-e5b9e2255f71 · outbound

This paper cites Mitigating modality prior-induced hallucinations in multimodal large language models via deciphering attention causality.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mitigating modality prior-induced hallucinations in multimodal large language models via deciphering attention causality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:75eb8bef4b80570f2e5743a4dce10a35d8575024d64c12492f31c257e0c03cbc

Observation e4c6dcbe-4472-44a1-bf04-71ca73eb9610 · outbound

This paper cites Quantum physics intelligent question answering (q&a) system based on retrieval-augmented generation.Concurrency and Computation: Practice and Experience, 38(1):e70379, 2026.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Quantum physics intelligent question answering (q&a) system based on retrieval-augmented generation.Concurrency and Computation: Practice and Experience, 38(1):e70379, 2026

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:9fb878b9e02e033e79a8acb0a5d6e5c3bd745720dc5ba6daaf0b953b18b5a260

Observation bbd94af2-b3eb-4169-8e70-07eacc036cfe · outbound

This paper cites Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.638397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:671c7ab01c12f558a6d9586dcc1cd55ed5cb3c3c57267fc80c538ead1782166c

Observation 81065ccc-859e-4c68-bc0f-c883bfe7f79d · outbound

This paper cites Large vision-language model alignment and misalignment: A survey through the lens of explainability.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Large vision-language model alignment and misalignment: A survey through the lens of explainability

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.587656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:dcf34c432abb9c7fc120ee1f4ee2de3c1e10800a8148f7a43d50177bbde39421

Observation b397e0b1-1b2e-4a2a-bbc7-90c5013a1e00 · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.590378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:1e9b9cd181594632dbb8cb7d3c12b5404c5107ab47bc4aa982380e5592ac0e9a

Observation bd0110c8-f3d1-4d28-82ae-c5374eeae13d · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:27:22.652231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:9a53286ae200c348ffa63e6d260f4b78096c11993e641209f91a50038c5603a8

Observation 4a4ad08b-812e-4496-a1e4-4a1e60a9d0ca · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:65ad4baafd5eec1fd89f42a5c21cb5f5db7a32e9e598bdbb61a67acfb5e55ee7

Observation 56e30cd9-bb06-48ab-a870-2bc5984bc98a · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:f04de70e835c03c7af0361b584246000d79c3875e4df1c699a54f201e8770f45

Observation c791b204-4b76-4e26-be4b-612271297f20 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:654582b2746da3e0a8c443b40c9ec45ce97140f55cea6753c62fce230d5d59ee

Observation a07a93fa-9fbd-43d8-81b9-5768c7d6d442 · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.Advances in Neural Information Processing Systems, 37:89098–89124, 2024.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.Advances in Neural Information Processing Systems, 37:89098–89124, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:dfc766960d546f732d072e8bc68ada7a7f5733d224e9c9ae21d235437b549867

Observation 40e7df6d-df4c-4bd7-b012-c1704e6591d5 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:6bc5a5c8511759929be20a699ce15b3e4be305f9a9e62a044307a5e93f9980d9

Observation fa802e29-c999-41bc-a413-89a1db1cb71f · outbound

This paper cites Lvbench: An extreme long video understanding benchmark.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Lvbench: An extreme long video understanding benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:806e37fd350f1aa0c07f1aaf651585281080b9d1b2e5d65868cc2ea99e516293

Observation ef33675a-6350-4390-ad5c-53bfea09d88e · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:3e7eb9d2b740eba8e225d92cf54640b99c22e8c45b10b3cc5f55c9f5bfe30e1c

Observation 93a41931-5898-48ea-bf7f-a2ce6852876b · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:d37bd524ee9bb8643474ee8561cb884abdac801c83cc48c1ce38fb5e43e6d2c9

Observation 4fd69f7f-2839-4ff7-9249-b74e8c26b73d · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Mlvu: Benchmarking multi-task long video understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:f1c2f77fbb0205d9f1312b1c23eedd40faececdd214d581698b03e7469bf7847

Observation 684f383e-b0a0-49ef-8502-2c42132a79c8 · outbound

This paper cites Behavior in habitat 2.0: Simulator-independent logical task description for benchmarking embodied ai agents, 2022.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Behavior in habitat 2.0: Simulator-independent logical task description for benchmarking embodied ai agents, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:f8c07196a52a23d57cdde53f4feb66de0a84a4db072ed43fa24f635b82cd270a

Observation 24df9feb-bf49-44df-9b0a-fa33147d580e · outbound

This paper cites ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:c6c6d2abae5ffeb474a6d76c82e7c9c927555daad6e1ec8b85c9f4d2751198ae

Observation 32794ebc-2422-43b8-8482-139a441032b6 · outbound

This paper cites Affordance Benchmark for MLLMs.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Affordance Benchmark for MLLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.654711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:6aa3dc36c965c6163d9d0bf9bac44d4bbc20cae4234e90f1812536910c23ad80

Observation 7e053c8f-27e8-4918-a40a-43d004865a00 · outbound

This paper cites Phystoolbench: Benchmarking physical tool understanding for mllms.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Phystoolbench: Benchmarking physical tool understanding for mllms

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.602087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:2de4b1fecf5bdc88a4b47670d8cc0a7141e094980fc059ea7a1c3c4ce1b1af5b

Observation c8c42eb3-de11-4ba8-ae44-72d02d4eedc3 · outbound

This paper cites RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.609663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:c6e358b9655c6a6037342694c7287b33d507c1979a878ebd883b2a97a05ae8af

Observation 3ee49f16-c076-4817-9bae-c03b8296568f · outbound

This paper cites QuantiPhy: A quantitative benchmark evaluating physical reasoning abilities of vision-language models.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? QuantiPhy: A quantitative benchmark evaluating physical reasoning abilities of vision-language models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.609970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:b655aa3050804377b550b4557a70630b2b46bd63dd303f601ca82e180a3b3795

Observation 012b3919-45ed-4e1e-b211-7b779cd4bde7 · outbound

This paper cites Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958, 2026.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958, 2026

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:8879d39cd95045399da8a8e149d5f4660e9b1a037add59580618d5882e61c547

Observation 5541e0bd-d6fb-49d9-9bbf-de7e114fb2f1 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:27:22.654943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:317923f0f9311efb358d45ee7a42bb260468f3710148d19fca35eabd233b47a1

Observation 846e9bad-35a8-4e0c-a7be-4c49aa3d42c2 · outbound

This paper cites Intphys 2019: A benchmark for visual intuitive physics understanding.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9):5016–5025, 2021.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Intphys 2019: A benchmark for visual intuitive physics understanding.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9):5016–5025, 2021

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:a8a2abfada9a4b0e85d0273c61163c4f75b74f40005b19673bd83412c1dd444d

Observation 21aa4551-0cc1-4dc5-961f-c27edbaf7cab · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:27:22.604481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:c8bd68c42540d2fac2c5b552a33c8fbcc7bc65c7e592f9ed2d449ed2cb9b6527

Observation 29da7131-22a1-4951-8475-d270991ad4c0 · outbound

This paper cites Physion: Evaluating Physical Prediction from Vision in Humans and Machines.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Physion: Evaluating Physical Prediction from Vision in Humans and Machines

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:27:22.625194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:fee16d2e5233d08496b422bbc001798365ac33074b10e85679fbad39079786d1

Observation b3615da6-f8a3-4fba-9760-37cfad50b70a · outbound

This paper cites an unresolved cited work.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:beff8a23bb2f996b57b63f1bc7a9f3a22fa0fde38a9a27641f508519529e43bd

Observation 0fa09806-9e8c-4453-b0c5-4bdba8143b62 · outbound

This paper cites ComPhy: Compositional Physical Reasoning of Objects and Events from Videos.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? ComPhy: Compositional Physical Reasoning of Objects and Events from Videos

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.627870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:e077baacc87b4bfea7cd485b2256f6b89ccf904d8a8aecc11806a199d342c7e4

Observation 178f3444-3905-4505-bef9-d75a1c981d15 · outbound

This paper cites ContPhy: Continuum Physical Concept Learning and Reasoning from Videos.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? ContPhy: Continuum Physical Concept Learning and Reasoning from Videos

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:27:22.635819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:b88509eda22b4ed3a2246b5b5f4a8619147c5bf202c02c2339ed9a72e7260116

Observation 847de4d2-6bbe-4c0c-800b-16f2fc34a3f8 · outbound

This paper cites UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.657341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:bc2c654ee30401130a41f039bd2c23b166511535aad238601db465e1c160f383

Observation 22bbed5b-1b7d-4d4c-87bf-c95da090f955 · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.619660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:791406d039e0c8ef0679858063f74aa6ba0773581c37f4221f44ad7f39192c28

Observation d8f0afaf-4af6-4b32-b691-5135409340bd · outbound

This paper cites PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.649811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:739882d5c9e5f00518800fd9c7b703b3e516964272d18f73819635fef5ab4682

Observation 50f1a3af-5596-4128-ad9d-c0b632926066 · outbound

This paper cites CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.646770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:cb5b8e362c0389651135ce0afc30b32a9761c05bc252f681b02e059936309552

Observation 46bb0bf4-c1d4-4ca0-9f42-5f6873fc888d · outbound

This paper cites Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environments, 2025.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environments, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:715ac4e527e6b521159ee628a306905c88c2cccd0c7bc3739acd61ff41d94ec0

Observation 247f534e-37e3-49ec-add2-34809bed93d8 · outbound

This paper cites Compositional physical reasoning of objects and events from videos.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Compositional physical reasoning of objects and events from videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:eac3f00dbb57ba49c32b66d925223df50fd4bb5b62b1a5799d32bb853df0bb15

Observation 5d6d2c36-e0cc-480d-af49-4f313ed60854 · outbound

This paper cites Craft: A benchmark for causal reasoning about forces and interactions.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Craft: A benchmark for causal reasoning about forces and interactions

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:697b92101a55fcf99e1c0f7435eaee951e1ab7b8918588c96a6fea7f9d3d47fe

Observation 3d1a870a-19fc-44cb-8043-7c4b6b8886b8 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:ba2443da19ec9ef43c8a431abbfb958dfc10e5670c273844008f17a22e1450b5

Observation a0677d9f-3162-40e3-93f3-a10432fe406a · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.643440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:3e692a971eb5819161dbef63e469e80b199931e713f2cc3a84685c0e8f9ead5e

Observation 63c365ef-b8cf-4732-b275-be649d161d0e · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748–42761, 2023.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748–42761, 2023

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:4154a63e2ade4fe90fd75c26643ce42bdf238d31a02ae984f0117bcb58f229ed

Observation 3921d17e-1f03-4205-b931-c1c531c532e6 · outbound

This paper cites Seeing sound, hearing sight: Uncovering modality bias and conflict of ai models in sound localization.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Seeing sound, hearing sight: Uncovering modality bias and conflict of ai models in sound localization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:4db07e088919009da83fda1612dd1f05cff385a5fd99d31029f9837f1f31d6cb

Observation f2449d16-3e95-47f8-8557-8c265af7382f · outbound

This paper cites Vmbench: A benchmark for perception-aligned video motion generation.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Vmbench: A benchmark for perception-aligned video motion generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:645098a7c180758dd6177c4f967c48c36d16ea41e4cb5cd54b12e66018cf8a56

Observation e9b47191-b58c-42da-af36-52e7a56a3d64 · outbound

This paper cites Unipc: A unified predictor-corrector framework for fast sampling of diffusion models.Advances in Neural Information Processing Systems, 36:49842–49869.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Unipc: A unified predictor-corrector framework for fast sampling of diffusion models.Advances in Neural Information Processing Systems, 36:49842–49869

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:27:22.630518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:c3cc680c22886d08ca8ba63d0830d2038b856d90db0421956e4cd1cc3fbc2664

Observation 9bff6b47-c3d9-478a-9ff8-1ef17734dc33 · outbound

This paper cites Physunibench: An undergraduate-level physics reasoning benchmark for multimodal models.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Physunibench: An undergraduate-level physics reasoning benchmark for multimodal models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.607461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:f2cab092268fcfcfb627af761f191f3d5735d1a2917eaac7b3032acea57996d6

Observation 581918d4-a989-411e-b8e9-4d7e39d59181 · outbound

This paper cites GPT-4 Technical Report.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? GPT-4 Technical Report

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:27:22.641767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:408a5a2dd674f2c18bf925b221acee866625c487a1e0bf44243ea4a17d41038a

Observation f8f1d2a6-44ba-42b6-8af2-59fa58f21aa8 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Gemini: A Family of Highly Capable Multimodal Models

Reference 56

Resolution
malformed identifier
local_arxiv, observed 2026-07-02T20:27:22.644118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:8cc1d189324822c51a554964ceab8a0abc929711bc428cfbd7063743e14ab26e

Observation f21d63bb-17f0-49f3-9cdb-6988fa0df3c2 · outbound

This paper cites Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T20:23:18.667677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:88bbf0eda86982587f7147d6eb97ac5e0309084d4c34893db66c30f5d0f03373

Pith citing papers

No inbound Pith citation observations are available.