Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2410.23266.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:06:46.226297Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T05:57:41.270398Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation b73cabd9-3ec1-4fdf-ac0f-828b168ce800 · inbound
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32543ac6-b29e-42a2-954b-c83eec350e2b · inbound
Seed1.5-VL Technical Report TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ebbce05a-4324-48b2-8795-b111606f7354 · inbound
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51d896f8-66df-4027-9c6e-e48beb66ab3b · inbound
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1084599e-7509-4f0e-a156-9c5c0b60a8f8 · inbound
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd1c37b8-704b-4224-888d-0992d7313db4 · inbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ffdbdba9-ab23-41ab-b009-7575a0a0e9f7 · inbound
ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b5b6e6-29d1-4b52-93ca-c5210d58cc63 · inbound
SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6adaaf95-1685-46ff-8bf3-bcd8f52a8fdb · inbound
Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cd74b951-ddc4-4b90-84ca-f9892bb792f2 · inbound
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c71e0dfd-8391-4f1b-9d1b-0a901beee377 · inbound
VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 43914f2a-3cf6-4734-a493-640c3ec42e7f · inbound
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6bb766fe-00df-468e-b65c-38e645fc268b · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 205
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.