Pith. sign in

Paper Citation Record · LEDGER

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 8 inbound Pith citation observations for arXiv:2411.02327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.02327 v4

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T17:31:59.030963Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:59:16.530036Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-25T06:10:24.042274Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact19
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b7a964c-7182-4033-8d99-3d38e54cd72c · outbound

This paper cites Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.688660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:597b379e3229d8eeb523c524ce3759948f1fa2228da5b02c79d8741ca97f58d9

Observation 0709323d-2412-477a-83f5-52e1cc6e9a92 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.683478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:9761835416e748ddbb3b0577ca3b9722d8e4d2643db2dde4ceca442223d36acd

Observation c00ffd46-e694-4b32-ae1e-7efc539b6ac1 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.645235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:e6b210b7eff3b934fca3d52adf398b2546aaac9bc86cbbb957710d6611c57a7f

Observation a9c0f558-4d9f-4e83-9ba7-53a353b33657 · outbound

This paper cites Instructblip: towards general-purpose vision-language models with instruction tuning.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Instructblip: towards general-purpose vision-language models with instruction tuning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:33:16.615806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:ee56284de439db315d34d0fcbfb9145dc71aff4174e4d75a9fc84325dc032912

Observation 42a13d49-e85f-4e53-b8bd-dd7c2fd2682a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.678252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:51cb9db2c513665ed986282db7fb098c863f4b242479bfc4823f2ddc2ee7dbbe

Observation e2794c98-6dde-432a-abc5-45cdba567235 · outbound

This paper cites VTimeLLM: Empower LLM to Grasp Video Moments.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance VTimeLLM: Empower LLM to Grasp Video Moments

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.634988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:be25a4cd03a202c2eeac2f0e4d7a6b098c1c57f1450184094ae8590d245738b4

Observation 11e849b7-9432-467b-b9d8-f9d404e0e492 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.629641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:aec30b0cf100255ea2a0476bdcc5a7c285c187192a573ed6b3951874fbe45e1d

Observation c674f23e-cc95-4b6d-aa9d-0f395eaddf1e · outbound

This paper cites The Kinetics Human Action Video Dataset.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance The Kinetics Human Action Video Dataset

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.660977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:e85537d6d78161b8a995fc0cfa47ce696bc6f0ddf8050e2702c110448fa597ef

Observation 9512c314-7c3b-47c8-b3c5-fdfe3c0b33bc · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.639650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:2ba6055e9b5db71a5a4ee48b50ef8a52e1cb0ee1d883b4d5f2431bcc8870b366

Observation 7eb635f6-8a01-4dee-9ed7-66b4293a8751 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.655629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:b96bb0abca906cfdbf704d5c981c9a60e21c40f40a4e7bad134ab451f90c7772

Observation b1642530-20a1-405b-b95c-0c4ab130da09 · outbound

This paper cites Visual Instruction Tuning.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Visual Instruction Tuning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.666100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:556521c261db79773f45532f8137b7cc5687722eb6ae41bf44e5c12ee42d2f28

Observation dc8993eb-df65-4539-8b6b-4bd24028e0e2 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.706077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:ff1e918c7275afa629e183c5765fb523c8f9231877f8fbec7ac276435eef0140

Observation 0f5d6a4b-5ef7-445b-ba7e-d036c0bcd49f · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.711594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:7a3eace162b18c12e377ea159ceff80cad09bef6dd52d738575497c3bff947a2

Observation 82efa17d-1292-46ee-ad95-f8bcdd83b5ca · outbound

This paper cites Disentangled Representation Learning for Text-Video Retrieval.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Disentangled Representation Learning for Text-Video Retrieval

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.650502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:0dc37e818adeada91abc2b37b0241dc8ffbeb8472fa43ee35b0f8cc6098a82f1

Observation ac24b98d-7588-4608-a5a6-f4da5a267dfc · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.699217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:2e38af04ec79908025ac999c7f6684523dde2ea5d47602daa0b5d241b711a4e0

Observation 0ab69d3a-edc3-4c0e-9fa9-4e7915dcf147 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance xgen-mm (blip-3): A family of open large multimodal models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.694290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:9cd10ccd348ef920bc487e5c61424630502739720984484738572cf1a8b2643a

Observation a37cdb24-04e9-4629-941b-eb8447461fd4 · outbound

This paper cites CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.721769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:c551a3bd5c19950a26829fb7711187a190ba07540cc8eb503e239505068f11ef

Observation 50c3b1a7-8470-40c9-9e15-4a369b372166 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.726334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:1ef5e212fac30abbd15c67b383cf2fd6141bb826f026d91e401b35a2ac6fbc23

Observation fa9c9685-13da-4874-b1b9-c8ad5fc4ef16 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:33:15.716201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:73655e887d75025ac0f6c0c0d84ee27225a0de036f0abdd0f8713dd9e22ddb0c

Observation fed152d1-acea-4a91-8536-678bbe328e50 · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.672665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:b271a3d78e1c5519f0edd2163147f61e03eb5e2151dc0bb63fbd5139bcd63cff

Pith citing papers

Observation 54f950b8-9c35-4a49-b604-db6aed4d3927 · inbound

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory cites this paper.

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.873276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T12:54:31.765909Z digest=sha256:5bf61d780762e3ea655127f17c60895d2c8af95a2f53114981f440b39f45ba17

Observation a2ad6307-3313-41c9-887d-fdf025e90a0e · inbound

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding cites this paper.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.530036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.530036Z digest=sha256:0ca50ff6ebcf0c8e805153a1a2d3b2d3cae30d0c2814db58437b5d7554c73668

Observation 0a00f510-85b7-4864-9be1-0d8183088207 · inbound

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding cites this paper.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.444553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.444553Z digest=sha256:6f73bf79fd1f1723c4cff209f72a71db43232f0865e85297c9f19d22877dcde2

Observation a91b56c2-f9f7-4d82-87bb-54d86a864abb · inbound

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning cites this paper.

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:40:42.234279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T06:39:29.010937Z digest=sha256:f982d1bf6108d937736b6367097c44aaca93239a9515a5939d9a3b947b79c80d

Observation ba40a180-65e0-4704-a181-399ea7ad7aeb · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:46:08.634289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T14:06:27.953376Z digest=sha256:af8c07a4758c369379e1f3a9fa9b88b6d590782deceb9490a85724d9d61ed6b9

Observation a3650b38-4333-43a9-a53b-26c46d9f4728 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:50:51.301451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:49:41.654207Z digest=sha256:aac0257de5ca0d22c7d931375aebe8835d6a73141a36ab19454e70131a383e8a

Observation 5b1d1496-031a-4bbc-bb17-c191991d2ca0 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:27.873311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:35:59.553683Z digest=sha256:b5ddeea469bbd45e63f33a2a78b416394f35dab9ecd2b6ccc51ce95f8244fc8a

Observation 894a8947-ea88-48fa-80a8-fb259b92b349 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:10:24.055171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:08:19.956833Z digest=sha256:c2afbcde2f7ce3ad2aa8d45cfdade290c931d223be9cef18019833e1eb17ee66