Pith. sign in

Paper Citation Record · LEDGER

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 55 inbound Pith citation observations for arXiv:2410.10818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.10818 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.911916Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:39:03.234328Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d5b6241f-c506-4dde-888a-5ebddaa3d455 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 216

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.170271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:b0634a4d5f1e15de1c3efd5e10e95689c5f8059587d9e82b14ef600e5d0e6982

Observation 1512597d-7a36-4870-8895-35fd4e258513 · inbound

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models cites this paper.

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:09.398031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:07:09.398031Z digest=sha256:5cd73c6ec7b3d2e2c0913dc8d66fbe909b1550743ef503d722fdfc820d18ea8d

Observation a1d00ac4-016f-4473-9f03-864bc1a8fed2 · inbound

VideoOrion: Tokenizing Object Dynamics in Videos cites this paper.

VideoOrion: Tokenizing Object Dynamics in Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.338126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.338126Z digest=sha256:5462760ae12f417c00eb2fd8615d1428790452a295bc3178df0159425aedf276

Observation 3545e38d-7d81-460e-b8f4-3e6589535254 · inbound

Progress-Aware Video Frame Captioning cites this paper.

Progress-Aware Video Frame Captioning TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:58.274135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:58.274135Z digest=sha256:05ee99b962c9fc1aa55a73f8d4f3e0aabad23e4aa7edf380a381b5e433eb59f6

Observation d4e27493-a762-4e80-956f-7a5cfad2ee83 · inbound

Apollo: An Exploration of Video Understanding in Large Multimodal Models cites this paper.

Apollo: An Exploration of Video Understanding in Large Multimodal Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:11:10.414228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:11:10.414228Z digest=sha256:01d05549da827f993705b4c3f0bbfd095a0f1ce274b80661b558065d0d3174bb

Observation bd4773ee-68fa-4ec1-8340-d4322bef8daa · inbound

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models cites this paper.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.575147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.575147Z digest=sha256:cc22ace7aea08d4e19146ad991b5eda835ca38cb7a4dce26bf0a9c2b963c35ef

Observation 28e5d89d-eb04-4f64-9bef-1ee19a008c81 · inbound

Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation cites this paper.

Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:28.275944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:15:28.275944Z digest=sha256:64d0f466dbcaea39496f4966de4dad12b19bd31d572727da4e216860502fcf8c

Observation 43e812f6-e863-40dd-8404-0d52591f11a0 · inbound

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! cites this paper.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.147800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.147800Z digest=sha256:6b6097e92a4c5c76ea6294ff48039c06469ce33f7901d84ad5fdb61c7f0686d0

Observation 12dd80fa-586b-45d7-9b83-9d90bb319641 · inbound

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding cites this paper.

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T17:15:23.810158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:15:23.810158Z digest=sha256:5643445a0f4cd0d715b5765c463748a3d271bd0c706ab635fdf24b14b348d021

Observation 8728683d-c1dd-4cae-8913-cb685f76a829 · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.302744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:c13e462a497e8809dfc2935c412be00ddbc8eb2bf4ea802bf12c5fa1b55c978a

Observation 333f9254-8eb1-45d4-bb86-ce42e95dab06 · inbound

HD-EPIC: A Highly-Detailed Egocentric Video Dataset cites this paper.

HD-EPIC: A Highly-Detailed Egocentric Video Dataset TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T23:26:20.268946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:26:20.268946Z digest=sha256:ef8527bef0ee32598332fe26e16003a2b26192b684e4fa93f2035a9dd2b44a21

Observation 9658c1c3-82cd-45bc-8251-b58c942b3496 · inbound

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding cites this paper.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.911916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.911916Z digest=sha256:1d5771389c7f5229440ac2a3b0383447d2b7c8d7689df60802c459966f8be0ab

Observation a2544e8b-e3ca-4e4f-904a-808be547e498 · inbound

MINERVA: Evaluating Complex Video Reasoning cites this paper.

MINERVA: Evaluating Complex Video Reasoning TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:19.023361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:19.023361Z digest=sha256:2f20420057d9d7375ccaed6aec6a1bbaedbdd42c2fedd417acc412faae46b5ab

Observation 4d7bc3d6-cca5-45e2-9860-084fb8cfc5d9 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.681583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:28070fe5bb9602baaaaeaae5202f0a6e23cd6ee6f625e620886d3f1567e30227

Observation dd700fdb-2a23-43c3-9a61-04b000dd0e50 · inbound

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models cites this paper.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.684509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.684509Z digest=sha256:e90e963dbf0ab7d2ac88ccb28110d85169f04ccd6532da332c5eec6ff9c68d54

Observation a6bc6d7e-b309-427f-aa28-fe817f779cc3 · inbound

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models cites this paper.

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:07.209929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:07.209929Z digest=sha256:93df1fc561eb8327b068403aa408c12708a49a61b19a7b530affb805bacd553d

Observation 0068b704-aecd-46cc-a8c5-175b3944aa05 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.572640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:02:59.572640Z digest=sha256:44e443904730a335200e8edaf00967c12e49aa413ceb198859db97a28172fc42

Observation 358b7850-828e-4e44-905f-3a296da34640 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.545824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:148aaf199a3e011255b802b21785d888476662f0a04b63fcfeaea4890efc7d83

Observation 8c97c53f-09c4-49cc-a2be-eebe6f273efc · inbound

Fostering Video Reasoning via Next-Event Prediction cites this paper.

Fostering Video Reasoning via Next-Event Prediction TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.431117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.431117Z digest=sha256:933f3880c1a24b1d7d1d090421d2931917479c6befe479e68aa1e46107bfd94f

Observation 305f6973-6c73-49c7-a805-d19b14918540 · inbound

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times cites this paper.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.121411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.121411Z digest=sha256:360a0570e8655779eb01e7eb5c386a6a1ea177d41417886b912a770ff2c54832

Observation 0a22fff6-cf20-4c15-89bc-b46ffd20a18e · inbound

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs cites this paper.

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:07:15.380828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T11:03:59.222849Z digest=sha256:ac43611a77cc56078c155f40a59a16ad8d57b25733d34020e6250044f9108ff0

Observation b3e7157c-7b90-4817-989b-f5200612b389 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.948202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:61af3fe668932d2992fe020e5539af599a758736b7c38091e68765085dca21f5

Observation e849daa4-bf69-475c-8e19-353675a6c28c · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:13.972207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:13.972207Z digest=sha256:0d157846f2e8aef424e66a5b02cb4e1e894a1f61ad7fe431b5af4c0852f5c978

Observation bbe643d1-35b2-4d2e-9e1a-e0b64b67fc02 · inbound

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? cites this paper.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:26.999928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:26.999928Z digest=sha256:42b7ef5097b85e175bef99b5ba25d0dcee86c7ebca8c4116992ab729628806c3

Observation f91d00e8-c6f6-491f-b1a5-7bdc970597a8 · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.035101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.035101Z digest=sha256:9185e45fe59c9b00aa0674162ece1cea2735f01e72b6a59c3c13579299089dc4

Observation ad492195-bd08-400d-a93c-a65b26054333 · inbound

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos cites this paper.

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:17.287285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:17.287285Z digest=sha256:06243435ecf939a45a18e6d130c3a398d159e8d50d645dbc5557e84700e38f60

Observation 15f9e0ce-387e-4e56-a998-6d9afa19ef88 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.661579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.661579Z digest=sha256:e379ba59656b956b02c38e624cb8b76d393831b7986da42354531f603a50aba9

Observation af06448f-fba1-4f4c-bcbd-94a2990c6610 · inbound

NeMo: Needle in a Montage for Video-Language Understanding cites this paper.

NeMo: Needle in a Montage for Video-Language Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:11.277795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:11.277795Z digest=sha256:354f1ed177254f4d1ba1b7671c036223bc35c8cdcf6a00bd98a340d21ab0d10a

Observation 6a8aaab0-b1f6-4fde-bec6-dbbb05df58b9 · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:31:32.156486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:d7dfa280342e57eaf0c4064315c34939d8eebb75be95060ab478b6d40d5e08a8

Observation d12b6c69-c134-451c-910b-20cabb957091 · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.103038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:3a59428e6175767e4c6538aaaf420531b35e8cba3958d7cb78ca870b884649cc

Observation c84d5d53-66bb-46aa-8f21-429c8c812f36 · inbound

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models cites this paper.

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T18:28:15.922879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:28:15.922879Z digest=sha256:c721a23e68d7414dddc2e9b078ede5b58d39f9d6e7c356dc337c3a39ad8e6c24

Observation f6290141-24a4-455e-b8f9-c5e6fe3d58e8 · inbound

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models cites this paper.

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:26:01.762020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:38:10.956726Z digest=sha256:55a246f0d54c16103c14df4f05bcf69a555eca5752b03a81ce68aca247d48e50

Observation 6968e55c-af11-43d4-910d-b458eb5ba282 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:04.015830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:61b7643f71ccf39d6c12d7e932bde5564aa8c9362e2865e6fe6fb2457290031c

Observation 4d934730-f911-4785-8ed2-1e50700138de · inbound

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models cites this paper.

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:46:37.546797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T06:41:59.641410Z digest=sha256:82b61b38b665242823c3b99f7567bd3c6ae40e3318e9285be8120afb4a59530a

Observation ebfb5115-25aa-42a9-8026-15d9f12fafa5 · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.264335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:0e6dbeb160f3151b4dbf3a52971f84043bddafb2bcd380d31d8f4543aec67da9

Observation 83c7d2dd-2323-4155-b984-2468ff757f85 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:26.606021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:c1a6c98716ea5283b798fd5d80b744304752cbcf3ba975d9dea38427bfe9e893

Observation 7d3ca994-57b8-40e7-917c-a5a53953eb14 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:57:28.298673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:2792d702737929adfbc5b4b58a0f829e8d878fe0f0ff27625edae0695d62c9c7

Observation 7559d5ad-eb08-48c6-aa60-ead607b633ac · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:52.564646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:52f4381ba31d950f241cb4efb602c950e529f7396f6a08ce1ad64de6d58051ec

Observation f1700d39-3c5f-45fb-8f20-19724b4ed41a · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.044136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:2bf8ed26b30445f7b6136eac112df099fb6e3b33673f2b4b5c3fca6d8bf1bde0

Observation aaa66853-e66b-4ffb-8c41-8f8517e6def5 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:13:05.462877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T06:08:12.887394Z digest=sha256:322f7aa4d9917a838dc89961a96b1e62725179b5b5057820435e2ddd7d8276fe

Observation a057f476-9899-436f-91f4-48608cfcacc4 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:04:02.791065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T07:59:57.488107Z digest=sha256:e59129f3b47efebe731ec7bb8e8e49c20ef567d162d8c82575e8528d048d8d24

Observation 165ec764-c51b-4b13-8e50-4780a9e9f8e9 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:14:59.999305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:09:49.178090Z digest=sha256:56a1ad3c939ece9c6f9a4e29244d56e7bd87d6e780846dee12d43faefff008f3

Observation e6080247-df3f-4d8b-944b-8736e0d26d6a · inbound

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning cites this paper.

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:46:14.880613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T07:45:56.473188Z digest=sha256:74e5b425e3fba39ffcc1ff59159d146b519bd70cfeacfc49f4c2bf3171b86053

Observation 9aca9f84-e2ee-4eb9-a074-cb6fe0b21f5e · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:46:40.010650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T05:40:33.752341Z digest=sha256:9b85fcc592086d4fae528fefe0623f9f7023a120bac039a1ec890b11f86bc47e

Observation 5fa44cfa-c497-42cd-8aa1-84798f581a31 · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-15T11:06:23.089564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:06:23.089564Z digest=sha256:6809342dc60f8d72449f183d428e224e12b44d44b3ee5f20859187123407642a

Observation de64e0c2-82c2-414a-a70c-739353ac54ac · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:27:48.232517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:27:48.232517Z digest=sha256:3a2ae4b76f5a5f2aed190eac0bf1cba8b9ea7edbeb3ca36d1956a881f20a5933

Observation b5eb221b-48ea-48d7-8e50-f892bc12077b · inbound

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams cites this paper.

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:51.073077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T18:24:57.881644Z digest=sha256:41ebaf2a0fc844d5458b7a936ec2b58481a44b0c3b66800eff581ecb74aa1896

Observation 2a848339-a283-4131-80eb-97e8b045757f · inbound

YoCausal: How Far is Video Generation from World Model? A Causality Perspective cites this paper.

YoCausal: How Far is Video Generation from World Model? A Causality Perspective TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:33:15.588037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T08:27:03.674229Z digest=sha256:85d2976d368549afdc343ed4a1ba7e1cd246c222612f3aa32ce7c58da3d4f195

Observation f8049609-c217-4dd0-9b2a-88c25912653e · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.638618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:d253fc4ce1a605f2538180cc5e1173b9a743c9bbe8c2b1253070e85d2872b31f

Observation 8ba65bc2-aaee-4236-9f7a-44708d0b76f4 · inbound

Benchmarking Visual State Tracking in Multimodal Video Understanding cites this paper.

Benchmarking Visual State Tracking in Multimodal Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:27.984113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T10:43:06.811228Z digest=sha256:b3780b03e58f99884b9818386ae8895c5a61525cfb5c5327958803989e9272c0

Observation a21a8d1f-e3e0-4d18-957e-68a4b6df0fbb · inbound

MAOAM: Unified Object and Material Selection with Vision-Language Models cites this paper.

MAOAM: Unified Object and Material Selection with Vision-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.337706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T11:08:59.900161Z digest=sha256:b9556f77eed0ccd7851e5e73f9d35392395c5086f8c07277dd1bdd78f3f6d5b0

Observation 4544884c-2f76-421b-ae9b-7d6aabf2f597 · inbound

APT: Atomic Physical Transitions for Causal Video-Language Understanding cites this paper.

APT: Atomic Physical Transitions for Causal Video-Language Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:39:03.236843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T22:06:00.249819Z digest=sha256:adf36ab4aadf1e0eb8f3238b46051f123c299a77dea95b39d0b7503664f02033

Observation fe62423f-d61e-4782-aac6-fa2f3577de9c · inbound

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation cites this paper.

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:05:49.779006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T00:53:52.724902Z digest=sha256:971ff7e9df1948fb0d1c50e668388df1bcad56a2d83a18297bd3b0a42f6620f9

Observation 30703464-7cc2-487a-b790-330a78c00b24 · inbound

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models cites this paper.

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 247

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:53.254750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:35:53.254750Z digest=sha256:17cc745cc556a84e0f2cd26afbf748b997aff54b06d76207dfc63f09ba1ce763

Observation e3fccc12-aa01-4200-9141-0de88513741f · inbound

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding cites this paper.

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:56:48.428575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:56:48.428575Z digest=sha256:a0669f99820b55ef5291268f994f72597a0396002007c6e42394429ee47869c7