Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:44:54.926584Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2412.10720.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:44:54.926584Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ee7622d1-5dfa-46c8-9bb8-a4d93a8bba3e · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a630dca-16d1-4102-86e9-6628a724d1ca · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Triple sequence generativ e adversarial nets for unsupervised image captioning,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f7f8107-574d-4ca9-9461-74136da45bf9 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Sketch storytelling,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b1fbd5f-e1eb-4d48-860b-c6adca70a371 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b0f9cdd-c751-4ab1-b6cc-6445be56d32c · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0fda45c-3af6-47f4-b2b8-6419469532d0 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Visual in-context le arning for large vision-language models,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 421e803c-170a-4b62-bd17-366913b0296d · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e582af8-2a10-421e-bb8c-385d906994b7 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc0a0539-7834-4bc9-850e-c2416b0f0f81 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives A su rvey on multimodal large language models,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 34f69918-805c-4277-9c18-291ad00a3b46 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Multimodal event transformer for i mage-guided story ending generation,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81961021-753c-4c11-b087-7455143f4495 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9795afb-6b79-4556-a418-6a729a2bbd57 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Style-aware contrastive learning for multi-style image captioning,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332e282d-e0f1-4489-8562-954c63e89838 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Improving cross-modal alignment for text-guided image inpaint- ing,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4f0ddb9d-d718-4827-bd03-1fc2844c8e8c · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38fb7a67-f1d4-430b-9953-5a35a52b049c · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Thread of Thought Unraveling Chaotic Contexts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3681400b-2e88-4dc2-9723-10bb43ceb208 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Streaming dense video captioning,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46227d4c-2f77-4c3c-a2c2-594de139b5e4 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Livecap: Live video captioning wit h sequential encoding network,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48aa3127-6575-4f59-8bc3-bbb94c45feb0 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Ret rieval enhanced zero-shot video captioning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d7fb0076-51c6-4b82-8ee5-05e8022df916 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Accura te and fast compressed video captioning,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c4e297-062b-4f96-8d31-97a676d7646b · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Available: https://doi.org/10.48550/ar Xiv.2405.07046
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c9161954-b612-47f5-9f4e-5d78b07a56b0 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d5e2da0-b6e7-48f7-8378-34e76cb4a432 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d1daa596-d7c4-4448-8d5a-9e83036f87b1 · outbound
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.