Pith. sign in

Paper Citation Record · LEDGER

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2501.10674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10674 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:41:35.841271Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d3718325-5de1-4fc0-92c6-5c98b9263dea · inbound

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs cites this paper.

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:07:15.464858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:03:59.222849Z digest=sha256:e03ab399624c9ee4917ff810781276db85c403d8e3336c38c2a6cc2976efe6fc

Observation 6039192c-0f0b-445f-9fff-c3c2200c1846 · inbound

Pluri-perspectivism in Human-robot Co-creativity with Older Adults cites this paper.

Pluri-perspectivism in Human-robot Co-creativity with Older Adults Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:41:35.841271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:41:35.841271Z digest=sha256:1ded65e4e9c295414a6182278164899c082d200ce619a4ce88ef328f4de83b6a

Observation 11c15949-1f24-437e-bb22-6a43f81c0ec0 · inbound

Lost in the Vibrations: Vision Language Models Fail the Dynamic Gauges Test cites this paper.

Lost in the Vibrations: Vision Language Models Fail the Dynamic Gauges Test Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:26.657532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:21:17.509796Z digest=sha256:0300fae520ec8f5a26daa08a2631063675b09b0412dd0f94e760778428b73455

Observation d7480da3-8d0c-4b46-9d8b-5d4996d25f45 · inbound

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice cites this paper.

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:05:50.282084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:03:50.055081Z digest=sha256:4a9b3f72f124e1e5787338354a1d7a80a9163d6e7520e342e112497a21f97913

Observation b9359344-7ff2-41b1-a7bf-0e96f2dc3622 · inbound

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models cites this paper.

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:59.623227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:30:36.491137Z digest=sha256:4f90b1f99c1be5236d66427c2c95778c562f9c5e4286bb63ef416d59aa65e223

Observation b98cd653-213b-47b3-b2ea-f2fa067e1798 · inbound

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning cites this paper.

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:13:52.947250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T04:53:25.259840Z digest=sha256:3b83e8dfecba6f7dc31fa25a28d15516d9ce104d042fb5c541ae84baab93065d