Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:03:10.621431Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2412.11391.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:03:10.621431Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T21:55:17.388570Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T20:06:13.285932Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 530b7ad4-bc9f-49a5-8cfd-dba9fa042326 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Improving cross-modal alignment fo r text- guided image inpainting,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4b571c40-358a-438a-84ad-99db03fbd50b · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Visual semantic role labeling for video understanding,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f4e8022b-8435-4489-814f-c3b5f11a9ff7 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Learning transferable visual models from na tural language supervision,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e91f26dc-2b2c-4e74-bb65-af1488614639 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a03d22b1-0641-482d-b3d2-473d1fd811a5 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Video-llava: Learning united visual representation by al ignment before projection,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 609e8ad4-1268-4d79-a882-7571ce8c09f2 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5165f5d6-26ab-4088-baea-38512f4dd074 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Visual in-context le arning for large vision-language models,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7b85c097-d117-4282-bc78-9a6cf8be69cd · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models An Introduction to Vision-Language Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ba79b8-b211-4aff-9698-fd29df02bdbe · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Triple sequence generativ e adversarial nets for unsupervised image captioning,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 13e4d032-0556-4830-963c-ac3152413ad6 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2845fadf-6512-4b6d-a673-60fa1db1973d · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Style-aware contrastive learning for multi-style image captioning,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a129508-079c-4be8-b8be-12d100853811 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Sketch storytelling,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb20cf11-9596-484d-8e78-98dcd360342c · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6280c8fe-0ba1-4b28-a26b-5b53566389d8 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfec1ed5-b8aa-42b3-97f2-4f04d3fbab7c · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models RelationVLM: Making Large Vision-Language Models Understand Visual Relations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df81e3f0-51c2-4e88-8ccf-aa11e3bbb0b7 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Multimodal event transformer for i mage-guided story ending generation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 028d54c5-aad0-4a89-a7ed-b0f224fa0b34 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Multimod al large language models: A survey,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2955da0a-9076-4f0e-b18b-bdea223b1209 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Thread of Thought Unraveling Chaotic Contexts
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4773ee-68fa-4ec1-8340-d4322bef8daa · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83b130a2-be03-470a-953b-9dcad23c942c · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51f9acce-d756-4f12-b611-3ee2ed4b045a · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TESTA: temporal-spatial token aggregation for long-form video-l anguage understanding,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea147321-d0f4-4ff8-8735-cbc5263c3825 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Electrophysiological responses i n the ventral temporal cortex during reading of numerals and calculation ,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e421dcfb-d0e2-4719-8a6b-74238e67754a · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb7e6470-1192-4b0f-92d1-2a9f66ea954f · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Enhancing video-language representations with structur al spatio- temporal alignment,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6649c21a-592f-445a-b08e-a77b3f96058c · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal Sentence Grounding in Videos: A Survey and Future Directions
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced600e6-8623-4ccf-81d1-c6fb5dabc37c · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Towards Effective Time-Aware Language Representation: Exploring Enhanced Temporal Understanding in Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c32036d5-1ebe-4b68-b961-09da0f4f71f6 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal Reasoning Transfer from Text to Video
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd473bb-d232-4b56-9282-6d7b0b9e27b6 · outbound
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8e505cdf-3815-43b9-b414-1c01695c8cb7 · inbound
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f830223a-3a6f-4c7d-a475-893a5c70b8d8 · inbound
MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.