Pith. sign in

Paper Citation Record · LEDGER

On the Consistency of Video Large Language Models in Temporal Comprehension

As of 22 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2411.12951.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12951 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:07:00.966961Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.570984Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved30
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08c990da-1e8e-4903-88b8-653ca79ca3cd · outbound

This paper cites GPT-4 Technical Report.

On the Consistency of Video Large Language Models in Temporal Comprehension GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.665661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.665661Z digest=sha256:2c8b4ee76a906bce33c5ecfa5d3dfc9468445c7109207a7479d7c57be5226e02

Observation b955f145-fd79-42bf-b5b4-2da3da0f60c4 · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval.

On the Consistency of Video Large Language Models in Temporal Comprehension The surprising effectiveness of multimodal large language models for video moment retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.670834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.670834Z digest=sha256:804def709f7f2bb15c35454183aca2f144816f8cad1bf383fce94be5c45d9bd8

Observation b3c56589-7156-44e0-8e30-add47b684379 · outbound

This paper cites End- to-end object detection with transformers.

On the Consistency of Video Large Language Models in Temporal Comprehension End- to-end object detection with transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.257703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.675443Z digest=sha256:8b3cd46b55bd02166f51d3dcf929e6915e2f041588d01135d0730fd62c6a9985

Observation 7f1de3a9-2cdc-4bca-ae3d-de878c1e28fa · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

On the Consistency of Video Large Language Models in Temporal Comprehension VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.680015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.680015Z digest=sha256:faf47a4da52402d3962738840dfd25f6438b24f2767a28d986b91a48b02fb88d

Observation ad1d2f73-aa96-4fe9-a1c5-a424690178f3 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023.

On the Consistency of Video Large Language Models in Temporal Comprehension Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.239855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.685381Z digest=sha256:8b992c760fd406b49992b86cdc986eceae8f8fbcb5f92e451c45f074e395c17b

Observation 6a659422-42fe-42db-b2f5-a86f256dd7a9 · outbound

This paper cites Measuring and improving consistency in pretrained language models.

On the Consistency of Video Large Language Models in Temporal Comprehension Measuring and improving consistency in pretrained language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.222390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.690358Z digest=sha256:8c3ee6a8ed9994a8c65e9df54dd4797cfaa5fa96aee412ccee000b5887c9e746

Observation 628ecb47-2ece-4484-bbbe-89b43659ec8e · outbound

This paper cites Tall: Temporal activity localization via language query.

On the Consistency of Video Large Language Models in Temporal Comprehension Tall: Temporal activity localization via language query

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.695685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.695685Z digest=sha256:32a959803f34933c4cbcef3a25ea9fa70cd7ecfacd8d28c9c1cbbdd07b58d621

Observation d1c10303-a580-4b2c-85bb-671a12b0ab79 · outbound

This paper cites VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding.

On the Consistency of Video Large Language Models in Temporal Comprehension VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.700588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.700588Z digest=sha256:ef25a53bf0a3134e8ca64d55518a816b7ec5d83bf955d5163b3257561ca53990

Observation c50191b1-4e91-419b-94f6-ceb545b407d0 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

On the Consistency of Video Large Language Models in Temporal Comprehension Vtimellm: Empower llm to grasp video moments

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.192540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.705927Z digest=sha256:c35201300999bd48775243d57e54a64ccec484b3b4c66e674f765347bbba9373

Observation 83e26ae3-5819-42ba-a3f9-862648089c34 · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

On the Consistency of Video Large Language Models in Temporal Comprehension LITA: Language Instructed Temporal-Localization Assistant

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.716282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.716282Z digest=sha256:b1d9ece27c565c50f1cc178bd7630c41621b0e2ab4d9a5a1350fe18be2b58eb3

Observation f0cca322-944d-419b-a950-650c1a0b051e · outbound

This paper cites Modal-specific pseudo query genera- tion for video corpus moment retrieval.

On the Consistency of Video Large Language Models in Temporal Comprehension Modal-specific pseudo query genera- tion for video corpus moment retrieval

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.173036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.722101Z digest=sha256:94d5f0f0ef97254781fec5634b9b029bd41298b9be4c481f0f46bf55b9e327d6

Observation 68388e89-0a61-4557-ba7c-0b2692f18aef · outbound

This paper cites Background-aware moment detection for video moment retrieval.

On the Consistency of Video Large Language Models in Temporal Comprehension Background-aware moment detection for video moment retrieval

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.156093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.726946Z digest=sha256:159e42b074f76fe611566da403e20dca2b0ec7250a790515ff44d9f1c4bb8543

Observation 129af2f7-9a6b-4b7e-bc82-8b43005c2458 · outbound

This paper cites Language Repository for Long Video Understanding.

On the Consistency of Video Large Language Models in Temporal Comprehension Language Repository for Long Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.731830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.731830Z digest=sha256:4566cfa319282a523923a5db8324ca752cbfd079a1a5e1a740d6d9320b63b40b

Observation a9a11936-c5ec-4787-9ec8-97e76b1767f8 · outbound

This paper cites Dense-captioning events in videos.

On the Consistency of Video Large Language Models in Temporal Comprehension Dense-captioning events in videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.137308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.737259Z digest=sha256:9b257cec3b99d27503b5d8d1b28df859e9ab1e8ba94e621ec937d50bf22c8093

Observation 6ff56cb3-4f18-41ec-ae05-55a96c0bb826 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

On the Consistency of Video Large Language Models in Temporal Comprehension Detecting mo- ments and highlights in videos via natural language queries

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.116551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.742401Z digest=sha256:0bde4f25c798ac0a9386bd744fdfc54665f96b97cf662359fa2536411b0e61b1

Observation 3b98d44e-4679-4ccc-9437-5f7e00f4ee6b · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

On the Consistency of Video Large Language Models in Temporal Comprehension VideoChat: Chat-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.747917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.747917Z digest=sha256:807f297f41d1b8899de761d8a924f3620eacba8caaf8a39d42084a6438fc680a

Observation daf696bf-3bbd-4c77-a87b-b3a180fd56cd · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

On the Consistency of Video Large Language Models in Temporal Comprehension Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.096278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.753337Z digest=sha256:8443776b0cc993a4ef27a522b8de382e7a3cac42fa60b82c3860c96e00e943a8

Observation 6eb5061f-cf7e-4ea8-b19a-63b639977b23 · outbound

This paper cites VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.758146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.758146Z digest=sha256:17046670970457f370ad0ed8a9985a65f0ee7b17ed55e6b3428ccb610e6f3ba9

Observation cb8b437f-aed0-497b-b36b-f42d083d8a54 · outbound

This paper cites Benchmarking and Improving Generator-Validator Consistency of Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension Benchmarking and Improving Generator-Validator Consistency of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.763059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.763059Z digest=sha256:bcb4a79a41121aae4dc7fc4bc8ac1ae050db157506a98b53f2102035834a238e

Observation a5ef0812-feb0-4a45-9f1d-dbc5e192265f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

On the Consistency of Video Large Language Models in Temporal Comprehension Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.768085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.768085Z digest=sha256:6bef41e1e6b0d2c0465cf34edf0e090fa8c00303aa5da1d41ca9c0b84ce290db

Observation f8fe26f4-9732-427d-ab48-ac4823d85ce5 · outbound

This paper cites Visual instruction tuning, 2023.

On the Consistency of Video Large Language Models in Temporal Comprehension Visual instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.773379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.773379Z digest=sha256:6d2dcdc8006e73a857581d11e10ce4bd4a5ab8041cebae82e8a4bc16e0047076

Observation bcc114b1-bc9b-4935-b563-bf168f02d0ec · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

On the Consistency of Video Large Language Models in Temporal Comprehension TempCompass: Do Video LLMs Really Understand Videos?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.778290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.778290Z digest=sha256:5ee7a6f892b956edcf4e37df804d1d86c553c0ccf40dfee3546f77bcdbd13495

Observation 6af2b771-a107-4a7a-a8d2-f466382770b7 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

On the Consistency of Video Large Language Models in Temporal Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.783631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.783631Z digest=sha256:94fdda3a703676e4621776624d2abec1037bae1f8d878ac5a21f40490828d9f8

Observation fd341e08-3ec7-49fb-a242-d8f604083cb9 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

On the Consistency of Video Large Language Models in Temporal Comprehension Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.056181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.796312Z digest=sha256:322877c2a6821244cc031a6fc48feeb46f42e571c0991e17bd2f690c70d97bf8

Observation cf2e423e-53b8-461c-a37a-a13460241155 · outbound

This paper cites Query-dependent video representation for moment retrieval and highlight detection.

On the Consistency of Video Large Language Models in Temporal Comprehension Query-dependent video representation for moment retrieval and highlight detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.034221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.800616Z digest=sha256:52582db2c32fc2a555cab0b56007fd826f680fb29cfa83e46108034dfa1bc401

Observation f10669a0-866d-4530-b045-cb19b8ef92db · outbound

This paper cites Local-global video-text interactions for temporal grounding.

On the Consistency of Video Large Language Models in Temporal Comprehension Local-global video-text interactions for temporal grounding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:02.013322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.805391Z digest=sha256:56bb98b81eb3d4745f03b09efe393bfb315be1438d6532e584a32673c381b7e2

Observation 081f7dd8-e3a6-4ac6-a1ad-3459c8ad70bb · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.809662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.809662Z digest=sha256:ed9285edbbbb1f847d75b16610e4f75fc392696aafc44041791f5414019602af

Observation f32a4cb9-4c3a-4245-bf2f-05710e46e5d8 · outbound

This paper cites Uncovering Hidden Challenges in Query-Based Video Moment Retrieval.

On the Consistency of Video Large Language Models in Temporal Comprehension Uncovering Hidden Challenges in Query-Based Video Moment Retrieval

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.814437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.814437Z digest=sha256:a1caa640be420b86a38d3eca3b9a72bbc21c5cc8cc621a26d4a81f89dbfff722

Observation 52ba0e91-0b05-483b-a6b9-dbba00470632 · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

On the Consistency of Video Large Language Models in Temporal Comprehension Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.819490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.819490Z digest=sha256:583cb19378a87390dffcef1a8cea7c61d1ddbb97994cd493cff0b6c42bce7ee5

Observation d7fd9c77-a4ce-4769-99bc-27b13b521333 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

On the Consistency of Video Large Language Models in Temporal Comprehension Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.824434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.824434Z digest=sha256:27c31b917e39a6ab7ea380521aa66ce76d9a14cff63047ecc8d908643bd41f17

Observation 7525b355-3d6c-414c-b3e2-48a18b4b3b0f · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

On the Consistency of Video Large Language Models in Temporal Comprehension Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.996307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.829905Z digest=sha256:7fa6eb16cb630f6d46ab858a6f56f67a57ccea23ac366c35038a4768aa6e3a5e

Observation edd3e7d5-3664-4011-9117-43be9ee7776d · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.835798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.835798Z digest=sha256:bf271559f7bcb813373d5ca10a80753ff06880cd08855211b34ee00a2911a555

Observation 8ad9f57c-8738-44b4-9e96-6089277393f9 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

On the Consistency of Video Large Language Models in Temporal Comprehension HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.840663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.840663Z digest=sha256:df50280a619e0360ee2f2fa0b416daf3e37c5a5351cba94d4021204b9ed0e185

Observation 3c426dd8-6ad6-4685-a14d-c9b05c6a5b37 · outbound

This paper cites Negative sample matters: A renaissance of metric learning for temporal grounding.

On the Consistency of Video Large Language Models in Temporal Comprehension Negative sample matters: A renaissance of metric learning for temporal grounding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.978865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.845487Z digest=sha256:2af61269b29e5ac311d1f87b8879eb2e72c0f4087d9b089062274372934f8228

Observation c0706820-e09c-44c9-b3e5-afab73f553b6 · outbound

This paper cites Chain-of- thought prompting elicits reasoning in large language models.

On the Consistency of Video Large Language Models in Temporal Comprehension Chain-of- thought prompting elicits reasoning in large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.961610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.850430Z digest=sha256:d1eec3907836e1e0cfb3cca00ced60c969f4f162e8a8e15143b1842e4be6d1b3

Observation 97b95638-3e78-4d88-8d31-996da5028e42 · outbound

This paper cites VideoQA in the Era of LLMs: An Empirical Study.

On the Consistency of Video Large Language Models in Temporal Comprehension VideoQA in the Era of LLMs: An Empirical Study

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.854915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.854915Z digest=sha256:fde372d1f38cb9598df1d1f5d2ff6ff1c0c16fe1fb546a061cd73837641ed13f

Observation 574da5c1-99cd-4299-8ab0-9072c58b7e7d · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

On the Consistency of Video Large Language Models in Temporal Comprehension Can i trust your answer? visually grounded video question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.943589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.860000Z digest=sha256:2c97546179d00aba61b773071122e39ff017f7b048bb6a15e94e14ceaa96aa03

Observation f9ec1c82-fc35-44e0-b10d-bd132f8e91b3 · outbound

This paper cites A closer look at temporal sentence ground- ing in videos: Dataset and metric.

On the Consistency of Video Large Language Models in Temporal Comprehension A closer look at temporal sentence ground- ing in videos: Dataset and metric

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.924923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.864538Z digest=sha256:839a3b0bf0be490afdcea086ee371b8a5ca0c70810cd002c8efe30bf17e54e2d

Observation 7101815f-3ff4-4f8e-80eb-738370f61abe · outbound

This paper cites Sc-tune: Unleashing self-consistent referential compre- hension in large vision language models.

On the Consistency of Video Large Language Models in Temporal Comprehension Sc-tune: Unleashing self-consistent referential compre- hension in large vision language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.908456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.869110Z digest=sha256:3dded0834950f879bd83d5e4c05de1c0a2e6759673794c70f30ad3ace9184f56

Observation 5bdc71a0-763f-44ab-8f84-70530c3739b9 · outbound

This paper cites Span-based Localizing Network for Natural Language Video Localization.

On the Consistency of Video Large Language Models in Temporal Comprehension Span-based Localizing Network for Natural Language Video Localization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.874181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.874181Z digest=sha256:bf752791b7f982f0e782f16359a9e116836464f198e01ca73f65c6081d8efebe

Observation 27c0a6f0-4963-4354-98a4-d5c6d0850ad0 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

On the Consistency of Video Large Language Models in Temporal Comprehension Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.879008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.879008Z digest=sha256:20b22475a719ee4eb7f6812c52c3ff668cde6114e6be2d8929dc249d92691097

Observation 6964b318-5963-4117-aa1a-4d7bd5816eca · outbound

This paper cites Learning 2d temporal adjacent networks for moment local- ization with natural language.

On the Consistency of Video Large Language Models in Temporal Comprehension Learning 2d temporal adjacent networks for moment local- ization with natural language

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.884824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.884824Z digest=sha256:ceb273b01837284d583f32b1298e96c7be16dfcc7874e4bc421f6df477b1bebf

Observation 5f458d21-c5b7-4dbf-a0ca-533ea3693a1b · outbound

This paper cites Unveiling the Tapestry of Consistency in Large Vision-Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension Unveiling the Tapestry of Consistency in Large Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.889826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.889826Z digest=sha256:46bcdaa39a3ea5f138bc5747f8245008a5228153428d88f84ff7ef2ab028f31f

Observation ab874dcf-3866-4f0b-80a6-0a7d39a40a66 · outbound

This paper cites Prompt Consistency for Zero-Shot Task Generalization.

On the Consistency of Video Large Language Models in Temporal Comprehension Prompt Consistency for Zero-Shot Task Generalization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.895247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.895247Z digest=sha256:6387f8fdf64f02200433a9223deb83ec4a90e10f9b6b2361c802677e81944380

Observation 6af2e0c8-f895-43dd-854b-1d6da6277b2e · outbound

This paper cites Towards auto- matic learning of procedures from web instructional videos.

On the Consistency of Video Large Language Models in Temporal Comprehension Towards auto- matic learning of procedures from web instructional videos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.901267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.901267Z digest=sha256:6f11293278c80870e5ce7ce8956952964f7f5c76b851804eb5bafa03fe508b93

Observation 07b15573-2b1f-4fb0-b96f-1eb1ac93e7a7 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

On the Consistency of Video Large Language Models in Temporal Comprehension MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.905873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.905873Z digest=sha256:581e05113a7e1ab1a23488d2c08f6cf60361f68d51df78f765b581289976d520

Observation fff885c0-7168-462e-bd0c-cd1a399c9171 · outbound

This paper cites It shows a remarkable zero-shot audio understanding capability and also generates responses to the visual and audio information presented in the videos.

On the Consistency of Video Large Language Models in Temporal Comprehension It shows a remarkable zero-shot audio understanding capability and also generates responses to the visual and audio information presented in the videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.865606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.910597Z digest=sha256:d0da86b6e056d209b373b5dc7f28f73fc3f1d33adaa0537125de60da1bde3848

Observation b40b6262-0a19-4173-b722-e609557a1753 · outbound

This paper cites To do this, Video- LLaV A collects both image and video-text datasets and incorporates them in its instruction tuning.

On the Consistency of Video Large Language Models in Temporal Comprehension To do this, Video- LLaV A collects both image and video-text datasets and incorporates them in its instruction tuning

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T17:07:01.843878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.915214Z digest=sha256:0e15b6612cc30a41ec8060c90c3913b2c6fb5c9848ebfb9a0b028e10c92ba66d

Observation a35b7f22-f46b-48f1-ae24-83cb385aaf2d · outbound

This paper cites It introduces a new dataset for video instruction tuning, containing 100,000 high-quality video-instruction pairs.

On the Consistency of Video Large Language Models in Temporal Comprehension It introduces a new dataset for video instruction tuning, containing 100,000 high-quality video-instruction pairs

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.827566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.919502Z digest=sha256:081bd1c34309fe3b39a82008bfa758cb59103eac527e1315014d71c845df6ac0

Observation 17f59f9c-44dc-49ae-827c-ec549a3eddba · outbound

This paper cites an unresolved cited work.

On the Consistency of Video Large Language Models in Temporal Comprehension Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:07:01.807893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.924171Z digest=sha256:9a43e283c7d68b2bf05ad921c216aceefb556d27a63a9a70d59198d14ddd640f

Observation 3996aac8-ead6-49ad-879d-829a2fe2e155 · outbound

This paper cites The format should be: ’start time - end seconds’.

On the Consistency of Video Large Language Models in Temporal Comprehension The format should be: ’start time - end seconds’

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.789263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.928772Z digest=sha256:3d91dca229b67f050409033f2b6bfe857fd99bedd1ac5373a2ad63734680afde

Observation a304e0bb-3c63-451d-a8fa-d2311d2200bc · outbound

This paper cites The output format should be: ’start - end seconds’.

On the Consistency of Video Large Language Models in Temporal Comprehension The output format should be: ’start - end seconds’

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.771741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.932889Z digest=sha256:0f1c067278685a7a4848023bdc9960ab003aab1872f4f7498947861f9f273262

Observation 5cc8dc9c-75cd-4577-bfff-c5c5fa42642f · outbound

This paper cites Specifically, they aim to align vision and text in the first stage and then generate captions from various image-text pairs.

On the Consistency of Video Large Language Models in Temporal Comprehension Specifically, they aim to align vision and text in the first stage and then generate captions from various image-text pairs

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.753056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.937581Z digest=sha256:e0f55b4f8b9d3e18d1df6c78e437dfeca7310569abd5aa08727b45a4dac31fa1

Observation 57a7ffd1-1db5-4dbe-bc4c-4e3989e145c7 · outbound

This paper cites They seamlessly integrate both visual and audio modalities in videos and propose STC connector to understand spatiotemporal video informa- tion.

On the Consistency of Video Large Language Models in Temporal Comprehension They seamlessly integrate both visual and audio modalities in videos and propose STC connector to understand spatiotemporal video informa- tion

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.731635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.941979Z digest=sha256:0cce0c1e077a1645b39d0023ef939a06328bd2e6bb6424b9785883c6d53cf801

Observation 3bb54507-8220-42f6-8272-5a8da55627ff · outbound

This paper cites an unresolved cited work.

On the Consistency of Video Large Language Models in Temporal Comprehension Unresolved cited work

Reference 55

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T17:07:01.711940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.946401Z digest=sha256:d542560b3fc63286fa83b7b4aa35614d2ba4f670b5dedb2bfc38bac9b128da0a

Observation 46f06485-429f-4d83-8e75-e4127754a740 · outbound

This paper cites an unresolved cited work.

On the Consistency of Video Large Language Models in Temporal Comprehension Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:07:01.694046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.951178Z digest=sha256:f0dc148ef300b17ca804308ca5329e16f2312db212e518899da0b86acb5d1e0a

Observation 041bec98-5982-4e8c-8da6-4fb868e61657 · outbound

This paper cites I’m unable to find timestamps in the video.

On the Consistency of Video Large Language Models in Temporal Comprehension I’m unable to find timestamps in the video

Reference 57

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T17:07:01.674919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.955768Z digest=sha256:dda380025900333b28d19c372afbc6f947a73034c738862da3ec8e995cec6273

Observation 83407866-e070-4e37-98b1-af452de9d1ae · outbound

This paper cites Experiments on Charades-CON with TimeChat.

On the Consistency of Video Large Language Models in Temporal Comprehension Experiments on Charades-CON with TimeChat

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.657803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.961927Z digest=sha256:b90d6ac917707baca197e285ac2d51ca50c3842e19b04fb4e7336fcc1489a971

Observation 64614f41-df1e-4010-a7fd-abc914e6d17b · outbound

This paper cites 16 Figure 11.

On the Consistency of Video Large Language Models in Temporal Comprehension 16 Figure 11

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:07:01.639697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T17:07:00.966961Z digest=sha256:313ede1c593e514649c13c87a6bfda2305170c438c176a8f788cd99ffe02e281

Pith citing papers

Observation 594ddd65-eb4d-4f0f-b519-f53b293afe38 · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark On the Consistency of Video Large Language Models in Temporal Comprehension

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.570984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.570984Z digest=sha256:db30e9e01e13cd98c7906b3fbfdd910aee433219c8e3c982fb3701e39e5ab914