Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T17:31:59.030963Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 8 inbound Pith citation observations for arXiv:2411.02327.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T17:31:59.030963Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:59:16.530036Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-25T06:10:24.042274Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6b7a964c-7182-4033-8d99-3d38e54cd72c · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0709323d-2412-477a-83f5-52e1cc6e9a92 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c00ffd46-e694-4b32-ae1e-7efc539b6ac1 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9c0f558-4d9f-4e83-9ba7-53a353b33657 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Instructblip: towards general-purpose vision-language models with instruction tuning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42a13d49-e85f-4e53-b8bd-dd7c2fd2682a · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2794c98-6dde-432a-abc5-45cdba567235 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance VTimeLLM: Empower LLM to Grasp Video Moments
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11e849b7-9432-467b-b9d8-f9d404e0e492 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c674f23e-cc95-4b6d-aa9d-0f395eaddf1e · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance The Kinetics Human Action Video Dataset
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9512c314-7c3b-47c8-b3c5-fdfe3c0b33bc · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7eb635f6-8a01-4dee-9ed7-66b4293a8751 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1642530-20a1-405b-b95c-0c4ab130da09 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Visual Instruction Tuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc8993eb-df65-4539-8b6b-4bd24028e0e2 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f5d6a4b-5ef7-445b-ba7e-d036c0bcd49f · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82efa17d-1292-46ee-ad95-f8bcdd83b5ca · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Disentangled Representation Learning for Text-Video Retrieval
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac24b98d-7588-4608-a5a6-f4da5a267dfc · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ab69d3a-edc3-4c0e-9fa9-4e7915dcf147 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance xgen-mm (blip-3): A family of open large multimodal models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a37cdb24-04e9-4629-941b-eb8447461fd4 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50c3b1a7-8470-40c9-9e15-4a369b372166 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa9c9685-13da-4874-b1b9-c8ad5fc4ef16 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fed152d1-acea-4a91-8536-678bbe328e50 · outbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54f950b8-9c35-4a49-b604-db6aed4d3927 · inbound
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2ad6307-3313-41c9-887d-fdf025e90a0e · inbound
Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a00f510-85b7-4864-9be1-0d8183088207 · inbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a91b56c2-f9f7-4d82-87bb-54d86a864abb · inbound
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba40a180-65e0-4704-a181-399ea7ad7aeb · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3650b38-4333-43a9-a53b-26c46d9f4728 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b1d1496-031a-4bbc-bb17-c191991d2ca0 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 894a8947-ea88-48fa-80a8-fb259b92b349 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.