Pith. sign in

Paper Citation Record · LEDGER

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times

As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2506.00928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00928 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:08.643991Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact5
  • verified fuzzy6
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 661074a2-ba1d-4c2a-bb75-d26358803388 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.065790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.065790Z digest=sha256:8b13017aaaa6dc3c4db97ecb910c8771d08e1bdcc92e7bfcec5a38ce6c0576d3

Observation 6fbbaa43-276b-4ea5-8110-f20991041e7c · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.753863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.079290Z digest=sha256:22559a4536fbda96b668cf2118fcf13c57329b626863d33d47178b6ef042a04c

Observation f205f64e-d814-4608-b617-3f2535bf5cc7 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.090193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.090193Z digest=sha256:5bd48e260688a8f212c75ec4c6b6482cf43f27e2497d9ed783326efa90b5cc02

Observation 8149cf40-87c9-4658-8393-38ccfb40a981 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.712770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.100925Z digest=sha256:eb581a1c7d8ed02d933a6fb503d88aa41314343a1c666e7d854ebf9624f96a7c

Observation 1f05cdf1-6cb3-4bc6-b8dc-6e29a695e0b9 · outbound

This paper cites Qwen Technical Report.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.107125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.107125Z digest=sha256:86c8ac199e81b7ab8ccebeb5abf27c6721906b10a8bbc29f489e5a5798815b12

Observation 190c08d5-ef08-4c04-bad2-3aaafccb369c · outbound

This paper cites Bosch, Mathilde Chailleux, and Francesca Foppolo.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Bosch, Mathilde Chailleux, and Francesca Foppolo

Reference 6

Resolution
verified exact
doi, observed 2026-08-07T11:58:08.773160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.113594Z digest=sha256:fab988d38386e75ce2db334e27c3f0ef03133fefd14d20b6190780b625b1977b

Observation 305f6973-6c73-49c7-a805-d19b14918540 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.121411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.121411Z digest=sha256:67d41b073ac6c750cfebd96883bb7cdf3a3de7b12ca38f0b8c2aaa2e080d9511

Observation 5174ea98-94dd-4396-b95f-d4ecb38eeee4 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.684219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.137586Z digest=sha256:eeb141a9a3fec2b4f67d376d080706c45fc9f6def6824defb277331132ea39a7

Observation 45f4a0c7-581e-43c9-9e16-0f09708d677b · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.148172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.148172Z digest=sha256:7356009f1480f5f69e156eb4c9de699e24d8ddd1ca8b5e205d6b99ca5ac2439f

Observation 7af4940d-214d-4cf0-b8e1-93b564dfb6be · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.163052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.163052Z digest=sha256:13e0ebb71d25ab976b99f72e056abbffa4affff5c9fd280af5a50b3516944f50

Observation 780dd187-28cd-426c-b1c3-4b8d4be3d655 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Gonzalez, Ion Stoica, and Eric P

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.169555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.169555Z digest=sha256:9ffbde5a74f4674b7cbe7866fc70076ef44e91e34c6267ac60a0dfbd0e7b942c

Observation b3ea757f-cda6-4708-969b-56f7866b53d9 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Scaling Instruction-Finetuned Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.177631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.177631Z digest=sha256:d4cf99b62d5313941aa7af9d069bf7445435111e142cbfcc34cee167b61db138

Observation abe87411-74c7-4815-b49a-6688ee090634 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.613626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.185701Z digest=sha256:eb610fbc2122d1164a84a5e1395608ad4202135099d67c435ec26ec9de380df4

Observation ea391800-8fe8-4cec-ac86-e26b7f1da069 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.589722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.194140Z digest=sha256:95f24594c1f6b169395cefdbab3da398fd0608facbc9a86eb28ea44d9e88bac6

Observation a67c9357-e855-4797-a47e-c00cf4454665 · outbound

This paper cites Bosch, Ciro Greco, Maria Nella Carminati, and Francesca Panzeri.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Bosch, Ciro Greco, Maria Nella Carminati, and Francesca Panzeri

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.548424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.205597Z digest=sha256:d44c436d5eb71303741bd248b69dc56f198c9527a60291d1129e303c6ff0fc40

Observation 3123e7e9-6174-4333-8504-59bf6df14862 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.511873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.222972Z digest=sha256:d73f4b8deeddfca22ec4875a78ca903ac114ae9182d87f3a45970abfc4e65d42

Observation 88f3d769-d1dd-4af2-870a-aa6bc7679af3 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.229550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.229550Z digest=sha256:21c3169406b0c9b1f505a1e6539581bc59370789d22706538d58c1b3b9139c0b

Observation 33d5291c-fc49-4bb6-ac6b-ff35595f70e3 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.243662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.243662Z digest=sha256:e83ec3a1e4ab0bd6b02385596db57749a970d1d1002351b502d2722c9f006e6f

Observation 88dececd-f3ee-419b-8eab-d98ae24b3419 · outbound

This paper cites Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.452676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.251953Z digest=sha256:25b7b2d5bc874713d4b10a1b4eef30a4f17f10a5eb62234364554593d16631fd

Observation 101a6953-0694-4eef-9cb4-f98d0e2002ab · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.409327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.261181Z digest=sha256:bd8d70e867a7e9d38ee9b6ae22122c5f546c3e029a691e26a3a3f64589d17ffe

Observation 7a86d6d0-763f-4431-9bb2-a47548082387 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times The Kinetics Human Action Video Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.269901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.269901Z digest=sha256:1c53b4881e82f320c49daadfd3bca35a542f8f0e970e0c5b3cbae94b83e7406a

Observation 8ab32041-ba94-48b4-9e0e-184aec77ae79 · outbound

This paper cites Richard Landis and Gary G.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Richard Landis and Gary G

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.379407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.279360Z digest=sha256:70122afe0de72fcc919fa24e2c9d865ba2ed53a2a8b4fda31a278c22f7372133

Observation 6ba89717-7e6b-444d-8765-2e136566d15a · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.355305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.294485Z digest=sha256:6e10fcf84719d44160f5044f50497556e4bc87b82187f4ec8897c9f03cdcf03f

Observation 8d77b7a2-50c7-406a-9503-d99203cf17a7 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.328245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.304150Z digest=sha256:672f65584b236bf7310ae587027911607acaa98557cb6a669f9acc9c4bf1ef4e

Observation 84b0703e-e012-4999-b0df-827ebc4bc01f · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.299380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.309019Z digest=sha256:33862fe51b3413dff6e8c7f10e0b2479f89b3a880a2032917cb5b4f7f8e2e948

Observation dbb4b5e4-3a7e-4e76-b17d-325ebe888241 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.318031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.318031Z digest=sha256:01b1745c55ff52c8d3098b0358250d3ae08a4f943dffce5b4195205f067017f9

Observation 003702f3-d3bb-4edb-aeea-0450b3168719 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.270819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.326369Z digest=sha256:b6d39727408124a0cc60456ce294a3612523b7a35165ae2d09a2355f0b1d64ef

Observation eda7a662-bae1-4f30-a5ad-1902a40b912d · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times TempCompass: Do Video LLMs Really Understand Videos?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.340059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.340059Z digest=sha256:7f14060f1557980d838ad7c4da9cd775509c1b9fc50df05de399e3278c58cc68

Observation 57753b8c-62ef-4678-b831-78e636faa163 · outbound

This paper cites Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.355931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.355931Z digest=sha256:b0a247d923bbcbc314897c33cc9dcedaa50c1d7943d00610dc6b32fbf242ca56

Observation 77c7fac6-a0ae-4e54-812f-90d8282ef350 · outbound

This paper cites Agentivit\`a e telicit\`a in GilBERTo: implicazioni cognitive.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Agentivit\`a e telicit\`a in GilBERTo: implicazioni cognitive

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:58:09.253700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.363088Z digest=sha256:50264950b3e71e14c1e6ef3e87efa8e73c7ae7667cdc2188b805a47cc9147702

Observation df6d49b8-2b3a-4ee0-aa00-408b69daaa68 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.371169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.371169Z digest=sha256:a10f5b437540c83f1cbbdec2e7d3438006048bf540ab38e476e88c42604c7ff8

Observation 7f14d5e1-ef21-48b7-91bb-e9ded914a0e7 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 32

Resolution
verified exact
doi, observed 2026-08-07T11:58:08.730314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.380118Z digest=sha256:ba60f638813383e3d3b86a9681d811bbdfa4923488a0a3fff2eb284cf08141b9

Observation 65869796-290d-4a5f-ad21-318311b7bf4b · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.240152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.391886Z digest=sha256:569e59734d6619f1da6981d9fb05f0fc5fcc19eea41251b6ecda01c4dc0d2714

Observation d6e8a864-4a98-46ec-9279-f20853c0ec43 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.203402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.398684Z digest=sha256:8f3ed24b6ca081e625595edd7e773d1f38a3a5b4abfc3f7cfc8eb7ae46300699

Observation 357174b5-9dd7-47cb-931d-64eadb196113 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.163004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.405770Z digest=sha256:f42fba690e09a18a533a5563585107819d84fdf1a12a4cc22aa484dad0dc1325

Observation 5814384c-e745-4a4d-ae37-4a4db5be6698 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.414163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.414163Z digest=sha256:d079084668c737390d83e8ca737303663540c47a4609ff6f18ce887a4884967a

Observation 97b4ab7a-f919-43e0-b666-cb0904b0fbda · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.420869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.420869Z digest=sha256:57027c069e15b523f60bd4553cf4671bcb70138b20a2592ac4a4078f58c5b2ae

Observation 9911d2db-23b7-4b05-af9d-22fde91c33ed · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.067753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.427772Z digest=sha256:b43acc5c8d22cb6d6332be5db027f84ade49f5c8a78ff3b6a19fefc075cbea85

Observation 8a725ab7-a015-451f-beec-04fa9aeb1814 · outbound

This paper cites Anwer, Tim Baldwin, Michael Felsberg, and Fahad S.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Anwer, Tim Baldwin, Michael Felsberg, and Fahad S

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.011609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.437777Z digest=sha256:3d599f2e67e08daea373019ab7734aaa85c61446362b6e1de97b3c3966ac8185

Observation 3b6ff0a6-05be-4876-a83e-04b6bbee94d9 · outbound

This paper cites Video Question Answering with Phrases via Semantic Roles.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video Question Answering with Phrases via Semantic Roles

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:58:09.185982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.449747Z digest=sha256:2b0e587acf3f0032f8bca3dc806e338c7f6ad425433b3b3dfefa3c90b4c57ebc

Observation 7b4c77ae-1998-49f9-96b9-b7dfba8e7405 · outbound

This paper cites Sigurdsson, G \"u l Varol, X.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Sigurdsson, G \"u l Varol, X

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:09.970795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.463803Z digest=sha256:b7f8f3b4ad3701372bc5b5e443d0d99c7b79c84528b3c00b82ef369dab7850ec

Observation 47f487db-cb26-4286-83b1-8ad8e5541a67 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Gemini: A Family of Highly Capable Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.473338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.473338Z digest=sha256:348b8d39f2adb91ef1670831bfac8c279995467dd7280f9c8d4bdb5283b1d4a4

Observation 99e957c6-d17e-40dd-832f-b4d4bd4c4be1 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.940208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.484690Z digest=sha256:c6977cfb8706e0f62e40ccea63f60c13c9709e287e49a6cc4b1092386e648e3f

Observation 6f28e24c-fd11-4442-b9a1-5dc98113b0fe · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.490321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.490321Z digest=sha256:d245b703324543d0b0482f2bb8943dfeb066f960f0cfbf2d65e3e48032ef6aae

Observation f9412774-1016-4c9d-ae3d-b17c4d922938 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.907561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.498573Z digest=sha256:d19d0794187043a5178c2f49d44a52e6d18fb2aa53d957ccfccaab98837e7a88

Observation 142db803-8e86-4a28-bcc5-e8239baec08e · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:09.879043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.509837Z digest=sha256:a038994cb1d0587efa9bfcca2f7a1903eed15f5cef16f968edc04154a80486b4

Observation 71817efc-3217-43ed-ad6e-3f631e95fc82 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.520364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.520364Z digest=sha256:14b6e3756febec83d5e8771ca48b5604b474e0407531fa6f26918ca6613c166b

Observation 3a903595-c68d-4c9e-9d41-4151345d5212 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.844726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.531169Z digest=sha256:3f2d49d039e5fc3905cf176f250259122e474e0948e0954db6829ba8fcc6d4f9

Observation 7de2259f-8e1f-466d-8cfb-4bea5307059a · outbound

This paper cites InstructionBench: An Instructional Video Understanding Benchmark.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times InstructionBench: An Instructional Video Understanding Benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.540817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.540817Z digest=sha256:da7d3b723e8a6342f9489ef1f13a3d4bdf648ca8a7c63d668a6cab15a8360635

Observation 73c04c96-407e-490a-a3fd-f5f0974c377e · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.813619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.551390Z digest=sha256:777984f36aeb11ef08476d963ded7f606ada69491d7436a9c2de5500324480fc

Observation f87a7a24-3391-4e34-ad8c-22d14662f196 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.559129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.559129Z digest=sha256:22da3a6280d2d238c15babbc8fe2f47d291a8094a1241c53dc7ed30a75cf1c05

Observation dfd73e6f-180f-468e-9b48-2121b184fcd2 · outbound

This paper cites Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.572193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.572193Z digest=sha256:5be884ddfdfb7669b44fc4e76d6cb3c7a49c55b1d8cb7249e8fe63fca64f8618

Observation 583eb29f-cf09-4dcb-a04d-1b18f0d70251 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.580265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.580265Z digest=sha256:1a41050d576b2659dc0dae9c207d2a531febda01c7a9844f569eda08dc9ddda1

Observation 0f910687-a0df-43de-a1f3-02378c2032ba · outbound

This paper cites ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.586217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.586217Z digest=sha256:6da2d4307c048d9930d26ab99327e3ba009fe8791fe76571d59445e2850aff54

Observation 4a733a98-b3dc-49d6-b2e2-f4ad12fd0db0 · outbound

This paper cites Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.593276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.593276Z digest=sha256:63e06e97aafc3751115f4e2432b3da1f472949359b549991ef2498f3e07f4aac

Observation a858a6da-6e70-4864-8619-3aaf6e95981d · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Sigmoid Loss for Language Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.598707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.598707Z digest=sha256:78a8384d5e1175f37f34c714bbc871431dddddc40bb1f88a36a85e8c0ba8f48d

Observation 5762f209-5a99-4b11-9c1c-55cd14564bbe · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.604138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.604138Z digest=sha256:fc0ca199ecc9df9ee8fc338d169d337db0f1a9575f9898e32e102355ecdf6b6b

Observation 41957909-cf19-4920-ac01-d848ae17fb67 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 58

Resolution
verified exact
doi, observed 2026-08-07T11:58:08.704124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.611110Z digest=sha256:4881893a0e536aa8769310e877b82dffc864a6ac04dd67a30300ee8684ae2a48

Observation 8896bb95-5376-4b02-a7d7-7c0f7a5e5aa9 · outbound

This paper cites Video Question Answering: Datasets, Algorithms and Challenges.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video Question Answering: Datasets, Algorithms and Challenges

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.626602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.626602Z digest=sha256:a88f789724c8294e11262040fa2c84980bbbbae2d00c06732cd0b2530127a304

Observation f4ff0b5f-be66-46ea-94b5-cd49718bdd99 · outbound

This paper cites online" 'onlinestring :=.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times online" 'onlinestring :=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.634847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.634847Z digest=sha256:38a5db46dab2dde1caa006129ef79b103ddf0f3daa81ae12c28b1e40c32fa2df

Observation 7ee018d2-2192-4ea8-be87-1c20cc788438 · outbound

This paper cites write newline.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.643991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.643991Z digest=sha256:dd6c9eeb3a101b3f28579e63fe9fee36d9bd54362a940ba258e9b5ad924169e5

Pith citing papers

No inbound Pith citation observations are available.