Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2501.07888.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T21:07:32.473373Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation ad0af566-b5c3-4249-9bc2-44d4c77eb671 · inbound
Goku: Flow Based Video Generative Foundation Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4820ee3-cf4b-4940-b5c7-8710dff26a8d · inbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a745cefa-6f64-4205-8f2c-d09036d15a8b · inbound
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71e9c9c5-ee26-4ec7-af1a-7aa1e7509853 · inbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c96d5b97-b387-450b-a5f1-13dc333d2a9f · inbound
How Important are Videos for Training Video LLMs? Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4a72311-99ba-4367-80b0-3e7fe61effcf · inbound
Seedance 1.0: Exploring the Boundaries of Video Generation Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 506152f2-e699-4975-852f-f94d86af54a0 · inbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a136efc8-662b-450c-90c5-22835e454eb3 · inbound
Kwai Keye-VL Technical Report Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5672ce2c-f641-4552-ba1f-88334ca51faf · inbound
Kwai Keye-VL 1.5 Technical Report Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71939508-e1c3-4376-8bb6-6eb0b7b8a73a · inbound
NeMo: Needle in a Montage for Video-Language Understanding Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8945482-6b3c-4c05-8c01-c345ae36d1ed · inbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 837614d4-94f0-44a2-b888-7a5e2b1dad60 · inbound
Syn4D: A Multiview Synthetic 4D Dataset Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 97f91e0e-d866-4983-9f5d-63cfd6aafa82 · inbound
Syn4D: A Multiview Synthetic 4D Dataset Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 129
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f5c3bc-281a-4040-813c-dd57eb4f54a0 · inbound
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2fee3852-83c5-41a1-9c57-bdca37c063c4 · inbound
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c13f9945-5d2d-4cd3-b539-cdc63e8a7fcd · inbound
GazeBehavior Annotation Toolkit (GBAT): AI-powered toolkit for automatic annotation of egocentric eye-tracking and video data of child-caregiver interaction Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e825296b-f6c7-4712-8e04-fc4d80848624 · inbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e656377-e5d8-4dd9-a5dd-4f8f19bbae3b · inbound
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d08a864-9ca4-44ec-9909-28ed378eedee · inbound
Task-Focused Memorization for Multimodal Agents Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d344e5c1-beb9-4a5f-a9ab-a1c298b9f080 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8cce8127-212d-43fe-8fed-4edf35c38d80 · inbound
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 28359097-2996-405b-9bfd-b14c2bf3938d · inbound
CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d43b88f3-90dd-4d8c-b055-ba896156cbef · inbound
MotionAtlas: Detailed Region Captioning for Motion-Centric Videos Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bba6e8b5-3184-4c67-af0b-ba8051221a73 · inbound
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e23db478-1672-4805-8959-7441d03e82d8 · inbound
Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dc8bdbe5-db5f-4970-bc1b-09f5f5777cae · inbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a5680dd-462b-49e5-a284-ee7adb09ce1f · inbound
LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.