Pith. sign in

Paper Citation Record · LEDGER

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives

As of 16 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2412.10720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10720 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:44:54.926584Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee7622d1-5dfa-46c8-9bb8-a4d93a8bba3e · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.753530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.753530Z digest=sha256:3d5d6417fadcdfedcd5d46de692ffed7f383aacad45ffd6fc1669012c23adcf9

Observation 7a630dca-16d1-4102-86e9-6628a724d1ca · outbound

This paper cites Triple sequence generativ e adversarial nets for unsupervised image captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Triple sequence generativ e adversarial nets for unsupervised image captioning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.766203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.766203Z digest=sha256:1894e34dbf81073e206b19b5e4fc4f34fd47d44ba6d65218d70ddd1079cd4c2b

Observation 1f7f8107-574d-4ca9-9461-74136da45bf9 · outbound

This paper cites Sketch storytelling,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Sketch storytelling,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.778107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.778107Z digest=sha256:5fbd39d72e86c3473fdf06ef6908e5e047d159226649739c1788827c3aa49f68

Observation 4b1fbd5f-e1eb-4d48-860b-c6adca70a371 · outbound

This paper cites NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.791036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.791036Z digest=sha256:4ece1b13c2e8a516c70897e8f7cd1c47e5f1ae6199938060457e83b0e5884510

Observation 1b0f9cdd-c751-4ab1-b6cc-6445be56d32c · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.799390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.799390Z digest=sha256:bc7c1c6d1d4f167625db11cf609ec9181cac12d8164c7de322c331d4b60c51ac

Observation e0fda45c-3af6-47f4-b2b8-6419469532d0 · outbound

This paper cites Visual in-context le arning for large vision-language models,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Visual in-context le arning for large vision-language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.803663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.803663Z digest=sha256:d1c166b9662e81fe4c7add9078ff2edaf1fc7a9cdd6d9dddd94e8fd921f01db8

Observation 421e803c-170a-4b62-bd17-366913b0296d · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.809836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.809836Z digest=sha256:4ddd4a52e1e5905f78ca3f6c85fbbb7d8563fcc33fa81ccc154de4fc1442c838

Observation 4e582af8-2a10-421e-bb8c-385d906994b7 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.831133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.831133Z digest=sha256:1c18e4d45568e6be960c335ae09fc0f66f99b0c2e60bb0de1499a8c31a753498

Observation fc0a0539-7834-4bc9-850e-c2416b0f0f81 · outbound

This paper cites A su rvey on multimodal large language models,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives A su rvey on multimodal large language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:44:56.029034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:44:54.838345Z digest=sha256:59541de0faae24684b9c2295eba9d352075aa31bf148d01862e74d6844132a9c

Observation 34f69918-805c-4277-9c18-291ad00a3b46 · outbound

This paper cites Multimodal event transformer for i mage-guided story ending generation,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Multimodal event transformer for i mage-guided story ending generation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.842939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.842939Z digest=sha256:64efcb863019ddc02d818123af5a7b6863185157b1bd221ced7d07909118e856

Observation 81961021-753c-4c11-b087-7455143f4495 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.848987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.848987Z digest=sha256:e30c7ef3d859de228b6c9927a8fa3c173f4d13677f3af4194d199d220e689a37

Observation c9795afb-6b79-4556-a418-6a729a2bbd57 · outbound

This paper cites Style-aware contrastive learning for multi-style image captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Style-aware contrastive learning for multi-style image captioning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.855522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.855522Z digest=sha256:7fcd4b5a0b59202ccd96617bbe59e5ba4d821b950c927de49359774ad46a7445

Observation 332e282d-e0f1-4489-8562-954c63e89838 · outbound

This paper cites Improving cross-modal alignment for text-guided image inpaint- ing,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Improving cross-modal alignment for text-guided image inpaint- ing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:44:55.981517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:44:54.861556Z digest=sha256:e0d18cf501eb17f89eec71ac565fb90710bb575ba124114fc05bf715da1c999c

Observation 4f0ddb9d-d718-4827-bd03-1fc2844c8e8c · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.869813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.869813Z digest=sha256:24be4128511ceb023a117bd9f8580c2bda054005469ee90b14bd5b41bc09f426

Observation 38fb7a67-f1d4-430b-9953-5a35a52b049c · outbound

This paper cites Thread of Thought Unraveling Chaotic Contexts.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Thread of Thought Unraveling Chaotic Contexts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.875118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.875118Z digest=sha256:2ed57fb218e0f0705a21cd07e3f2c7abb47b69c9b2ba94368e5b48c9c0803bd3

Observation 3681400b-2e88-4dc2-9723-10bb43ceb208 · outbound

This paper cites Streaming dense video captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Streaming dense video captioning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.883452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.883452Z digest=sha256:106e663be3eac7154bb3035163055b3cb4fe11baa9d881aa5ae44938de10d355

Observation 46227d4c-2f77-4c3c-a2c2-594de139b5e4 · outbound

This paper cites Livecap: Live video captioning wit h sequential encoding network,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Livecap: Live video captioning wit h sequential encoding network,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.888620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.888620Z digest=sha256:a57d65dee9e5ee6d2339af71e5815eaeb0fa01de882e5f5bf8457e4049fa9f35

Observation 48aa3127-6575-4f59-8bc3-bbb94c45feb0 · outbound

This paper cites Ret rieval enhanced zero-shot video captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Ret rieval enhanced zero-shot video captioning,

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-11T15:44:55.563935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:44:54.893937Z digest=sha256:f66e2a6a3b059ea5265c707a1adc87ad1e4b6282d89f24eda8a20ebb2107dfaf

Observation d7fb0076-51c6-4b82-8ee5-05e8022df916 · outbound

This paper cites Accura te and fast compressed video captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Accura te and fast compressed video captioning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.911301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.911301Z digest=sha256:4eb71bb95909c723e271f5aed9e77ec56898a488eb0587aef824de93cb1db313

Observation a6c4e297-062b-4f96-8d31-97a676d7646b · outbound

This paper cites Available: https://doi.org/10.48550/ar Xiv.2405.07046.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Available: https://doi.org/10.48550/ar Xiv.2405.07046

Reference 20

Resolution
verified exact
doi, observed 2026-08-11T15:44:55.151058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:44:54.901405Z digest=sha256:15313edae78cd7e8b27c21f31cb5386b7b09be8d84d1b56975d7c9cdae142e09

Observation c9161954-b612-47f5-9f4e-5d78b07a56b0 · outbound

This paper cites Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.916642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.916642Z digest=sha256:5f9eed36d0963f365398ccbbfebec0d7701c362c1e15116d6dd30c534c2e2def

Observation 8d5e2da0-b6e7-48f7-8378-34e76cb4a432 · outbound

This paper cites Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion

Reference 2023

Resolution
malformed identifier
local_arxiv, observed 2026-08-11T15:44:54.992595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:44:54.926584Z digest=sha256:ad4c0d0c51cbb1d915bb0f730097aeae0b3f41652725b3c63e4cc194fe956d42

Observation d1daa596-d7c4-4448-8d5a-9e83036f87b1 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.816312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.816312Z digest=sha256:4350e24b65b968d0cc0ac036baefc4eb3c2655b76be1698303f9ee611dec8c53

Pith citing papers

No inbound Pith citation observations are available.