Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2407.00634.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:40:23.734088Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T17:40:00.906584Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 7868b652-9b62-4519-befa-8237bb57d600 · inbound
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0be139c6-460b-4b87-ae5c-c34e807f64df · inbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f6822db5-7cbd-4ebd-8023-6aa5e8df054d · inbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e108ea2e-3b48-4b8c-91e9-463ea9edbf50 · inbound
Open-Sora Plan: Open-Source Large Video Generation Model Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8b4458b4-060b-4a33-bdfa-057b2f6d6052 · inbound
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee4e1f2-cd67-4798-a8a9-e436db706eac · inbound
Progress-Aware Video Frame Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f9c87a1-a224-42c3-87d2-2b5769e407f0 · inbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0db57b7-51b8-4e78-8b50-f50461a269ac · inbound
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 121
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b4c1064-cd9d-4bfc-afc1-2fc19c77a335 · inbound
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29df199b-4fd9-4a7e-9473-d21570d2bad3 · inbound
BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63f8375a-2a2e-47a9-a9bf-2f2200751a76 · inbound
Goku: Flow Based Video Generative Foundation Models Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ba35c61-2595-4c3e-81cc-798898ead23d · inbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d31b4ecc-a26f-4da3-b04f-050dde267969 · inbound
Seed1.5-VL Technical Report Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4165373e-c41e-4f73-bb44-aa18251c6109 · inbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e152143e-f3e1-440b-be7d-f896473bbd0b · inbound
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f47666e-7b6c-4623-bfbb-7765955299fa · inbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c314e2d9-ad47-490c-a7b0-549e292d25e3 · inbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 543f8e17-0ffc-4fd6-aac0-f39935557265 · inbound
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b82aeb2f-ca6e-4fb8-a96a-3af13f1bd4a5 · inbound
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb707a5-6917-4e9f-b00d-1c255f096643 · inbound
How Important are Videos for Training Video LLMs? Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c455dc02-2b25-4543-a95a-d0536125090a · inbound
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b62dbf61-4e24-4aa2-be39-3f4e90809c3f · inbound
Adapting MLLMs for Nuanced Video Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 245512fd-c346-4587-9d79-985335bf2157 · inbound
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation acb55947-e818-425b-b074-60982699ea2b · inbound
SCP: Spatial Causal Prediction in Video Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 31ef9471-cb5c-43dd-924d-6b79638221e1 · inbound
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3b876ced-1ea4-4dad-ac02-ba66176a193c · inbound
Building a Precise Video Language with Human-AI Oversight Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 97a2f1b6-2f1e-4d05-bab6-e6c6324a2baa · inbound
Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 38714149-295f-4abe-9ab5-d23f4216ba24 · inbound
CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c55f7ec3-6104-41fc-870c-a9ab35371b3e · inbound
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 779d8d9e-b6f3-4051-86a8-a7c4d19b3f41 · inbound
UNIVID: Unified Vision-Language Model for Video Moderation Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d58a1299-f53b-4cbb-8f78-1f7c9e333eba · inbound
MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6909c904-fb97-4a95-aee1-5bbac09e37f0 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f129b174-1b38-463b-a386-c40262bf78de · inbound
AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e33d93dd-136a-4769-8026-647da8b3bed4 · inbound
Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9b7429ba-3931-4175-861f-4f270f630a92 · inbound
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2902fe35-5cb6-4ef3-83a3-04289e077e83 · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 297
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1efce4af-dece-45af-8191-c58b9bd5045a · inbound
CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 87c325a0-12db-4fe8-b67c-73b20cc29f7e · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation de2f74fc-0519-4b49-ba64-ce5e88f3af4a · inbound
MotionAtlas: Detailed Region Captioning for Motion-Centric Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a0274c15-76c9-48f4-a8cb-a533aa800a29 · inbound
Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1a17a28f-d18a-4b89-92f5-7f5b45e0daa2 · inbound
MentalThink: Shaping Thoughts in Mental SVG World Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 167
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4661c7-a4c4-4aeb-8791-76af60bd4357 · inbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9643093e-bfed-4251-9083-b15b50c20ec5 · inbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90db9274-bbf3-46ad-8514-45a615392547 · inbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.