Pith. sign in

Paper Citation Record · LEDGER

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler

As of 11 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2501.15513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15513 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:18:18.203170Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:39:26.597389Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0594859a-02ad-4274-b5cb-675f791c02be · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.356749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.356749Z digest=sha256:e4d688be69da3552d967846ec8aa57836668e8ee102f3a1831a6cf85a6b24325

Observation 8b723536-4ad9-4cf2-a9a7-dd830a5299c9 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.386325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.386325Z digest=sha256:2b83da37fc9a12e5061fd28ff3838228bbd88141fdf9367092a76950956dc9e8

Observation b8072065-f2ca-4289-aeef-d172b1250d6f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.411410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.411410Z digest=sha256:219913c3fcae4f4efd2102d08f0015de17febc9506167ae428f0dcc5390bf1e3

Observation da7685aa-0d17-435c-8389-f3141c41ff2a · outbound

This paper cites On attention redundancy: A comprehensive study.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler On attention redundancy: A comprehensive study

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.819859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:18:17.446172Z digest=sha256:e6ed556ac4f2f2d30c8f63417a6a1fc41b839ea5042549a31a294172ab71d1ac

Observation d0610ac8-318e-447b-a2a3-8986869d718d · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.464986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.464986Z digest=sha256:4f687b7333e87c601d9d9ab440b6818e3e8c2b5c530d954823592cd81c0d254c

Observation 5fde84b7-b003-4cf6-8e4e-75596a287cc8 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.488069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.488069Z digest=sha256:c4ba19144f2f13e2e71a031e043d87d0bdd6737e4fd33027ad93526559fc379e

Observation c371f93b-dd3b-4003-8e41-3bdd221f361a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.511758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.511758Z digest=sha256:2e64c9939908b2f602e28469ff4538410f196df8945d82997d61fa4a8af3ebce

Observation 94764db4-f09f-4834-9d11-812e5a47907d · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Efficient Multimodal Learning from Data-centric Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.531656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.531656Z digest=sha256:77e1d0b0987d3c22962f5aee0f405316b166694f5a4fdf8c23fd527a8a93a1a7

Observation 1228130b-8da1-477f-8bf7-98b9e7adb483 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.549431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.549431Z digest=sha256:b1c79fb7fca8d63f02ded517a9ecb4ba07c65ecb2996831520fcd086be203f50

Observation 2ea93c03-052e-44e2-a882-852655171b02 · outbound

This paper cites Qwen2.5-Coder Technical Report.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Qwen2.5-Coder Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.558446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.558446Z digest=sha256:9cb4996790e1661b0e38eeed24eed364c32c807eacbb7923ae163086f4ca2d32

Observation ee1e81e6-e1d9-427f-89f6-59ac41686970 · outbound

This paper cites Phi-2: The surprising power of small language models.Microsoft Research Blog, 1(3):3, 2023.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Phi-2: The surprising power of small language models.Microsoft Research Blog, 1(3):3, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.727706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:18:17.570633Z digest=sha256:08efcc10445f955827f690a74a3abb1cbe1c6305b4a1dad381848f430fcbd958

Observation e7f2384a-558d-4916-b860-04b6f227d507 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.580777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.580777Z digest=sha256:df6cdcf8982ec79f5c5fb422a127b202b2648243bac0995fec37c5efa9dfca8a

Observation 5035d900-4a80-48fe-97fb-804f35289731 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.592863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.592863Z digest=sha256:166f047ef1fe20cd6615323dec60ecc1a5a5df94c14d3b84e63f968a1d6fdff7

Observation 39898d34-2355-4ed0-82d8-51ca383af8ed · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler VideoChat: Chat-Centric Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.605226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.605226Z digest=sha256:cd488ad4f83d910ce109fb688039eccf3fd816b08cc4cbb2a5b9149b5e673983

Observation 07ab9965-e9d8-471e-94ff-916d0f37064c · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.618502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.618502Z digest=sha256:dded732670d57c09b31d09656ba5291a3bff7c1a2ae738e8fdbc76b7343e1b23

Observation 584475cf-7c6c-47b5-9226-14e34cbec53e · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Llama-vid: An image is worth 2 tokens in large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.618280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:18:17.645852Z digest=sha256:1130a84cee2d89329bcf5ce6855373542e87e2f48a86514b04a1c29d83f24e97

Observation 6334ebe2-a2e4-40c8-a56f-ac9683e0c18d · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.694223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.694223Z digest=sha256:217c8de20c933f538a765b6758a2535c74391e4db4856112ab90606631e1694c

Observation d0a71bda-cba0-4041-965d-e16e645a9f62 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.735844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.735844Z digest=sha256:169d534d6583d39493e8689f06d0b01831cced77d29763b9e8f39d855c70b925

Observation ee1fd6ce-9700-4652-9bbc-78cbd819c634 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.787631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.787631Z digest=sha256:30f66d216ab24e757990c109c9ff78e058089ac1b4bfe040c9c200beebce57c0

Observation 87cf6f5f-bda8-418f-8086-cc2bfa6f68b9 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.816391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.816391Z digest=sha256:d31d1f11414023510f47f5cba971872f0e7e3a0cc357223858fa43cd4cb7d3dc

Observation 72a5ebb0-6080-4aa6-8c4f-b2863db0fc86 · outbound

This paper cites St-llm: Large language models are effective temporal learners.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler St-llm: Large language models are effective temporal learners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.580148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:18:17.845024Z digest=sha256:607481547291bd124aa657c5594641135d994034bd9f32a2105e03b871841730

Observation 38287e0a-7220-4daf-9cbd-e6fcfdf0004d · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.857839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.857839Z digest=sha256:262108fc2031af1e154bfac659af93e7579666e33d0e8139e698f71e62e0d67b

Observation 1ee46f66-a036-46a2-be2b-a8e42416aa54 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.898432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.898432Z digest=sha256:9f2d3500f3fcf1c9d5c219dd8538bc408f04235bc27dfc0c6abb2ede1dbed607

Observation c07d4a9d-4c35-428a-866d-5cca2a9b9ade · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler DINOv2: Learning Robust Visual Features without Supervision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.916023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.916023Z digest=sha256:e3db6dedf20f824880404a441c13e4e0fe5050178eefada15d53a6fd1b160eda

Observation 94812177-582d-490a-9e94-0bbc74402419 · outbound

This paper cites Learning transferable visual models from natural language supervision.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.929562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.929562Z digest=sha256:0bbf08b1f3b21721f7004b5a69d316bb38d8bc4aca1a9070ded918aee018a5a4

Observation 294d4da3-d181-4240-89ef-097f885f3295 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Moviechat: From dense token to sparse memory for long video understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.934235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.934235Z digest=sha256:f5fb87a70a7d4720c92d3afccb29bfb35c73c8741361c9c0057addbb3904e3fc

Observation 722f96a5-9ef2-4191-b2a7-581d452959de · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Gemma: Open Models Based on Gemini Research and Technology

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.947719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.947719Z digest=sha256:77ebf762691712df6d1a5d4d963d1a32b8da629fd7abac99455f467ea0db4cd8

Observation d49048db-9381-46d5-b43a-c05e9d6b6cf5 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.964511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.964511Z digest=sha256:97cb864b0ec62ee205a718c3582148d0cb7f5e6ec7a9ace3af34aa8c6f5d3573

Observation a0a81bc9-a7e6-4d8f-9d69-ee685f2fb426 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.977599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.977599Z digest=sha256:9337317eface2681e0981eed761aa6dda31ab214f954e4b0fd8d07465f0c4e5c

Observation aa8b79b0-698f-4177-b0f7-7f4bc1b232d7 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.557231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:18:17.988770Z digest=sha256:42dc11977c4a9ca62d93462db140c85c27d1183898c96f947e755e2585dcbf86

Observation 35d0e1c7-193d-47c4-8c84-d15690c79cfa · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.992793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.992793Z digest=sha256:23fafefea5609fc5c01dfff143876aaceb894b5ea83057614c10c7814c76090a

Observation a8d7feb4-d9bf-4e6f-bcd9-228401fa8cf8 · outbound

This paper cites Sigmoid loss for language image pre-training.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Sigmoid loss for language image pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.998128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.998128Z digest=sha256:e22f729ca606329226566271161c74bbd822792bf5fe084d9e25e2f812fd3648

Observation 4b280291-3a34-49c4-b416-3ed501db93fa · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.004242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.004242Z digest=sha256:6ece3bd171ea052b6eaee906d9b4fecaae7249dfe3d3ceacd63ea469ea8eb865

Observation 8ef89105-5643-4fc6-b894-96a13c1bc000 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler TinyLlama: An Open-Source Small Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.008511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.008511Z digest=sha256:6a28882d8a5ef5c5eda53994f7edd7989be93033ad506b1bb699eeec2a2d5f50

Observation 86aeb704-3adc-4170-afc5-5eac48021193 · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.012399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.012399Z digest=sha256:464e162153c2956d6d713fbc1e15b9471446c533d96956a29eb8a44e88712402

Observation fe5d7b0f-5675-4012-8c79-da07e7ebffe9 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.042072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.042072Z digest=sha256:a9c0d4906952ae71a24618e45053b432e9930ebc80c7154ca82e419c849d42b7

Observation 7ce12156-d8cc-4ea6-ba88-ee60aab2a8c3 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.086229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.086229Z digest=sha256:e9395724776b4d5c639f800b7e77038c0e1bc9ede1937e6960e5efebc3df379e

Observation ede505cb-e091-4b28-a534-6a91bffb337e · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.139242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.139242Z digest=sha256:92796f5fb1dc1e01315d081cdb53da386fa4011275bcd0a7dad23fceb8384766

Observation 4eba6fca-bc13-46c2-9fb9-6984553edddb · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler MLVU: Benchmarking Multi-task Long Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.168610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.168610Z digest=sha256:fd09c25fc2ddb756af3c5e98a22a79da0df7e4001e7b79aebf648e6a89be4404

Observation 2216e17b-a605-4ba7-9dc2-b2692de0cbf6 · outbound

This paper cites The best results are indicated byboldface.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler The best results are indicated byboldface

Reference 512

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:18:18.533114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:18:18.203170Z digest=sha256:c97659e1b5781396c5de76f2337f5c4bba56dd8975a30b92199aba8ce44077cf

Pith citing papers

Observation ca6bb9bc-a892-4943-a3f2-fc70019a7e86 · inbound

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs cites this paper.

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T04:39:26.597389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:39:26.597389Z digest=sha256:aff749115cb2006d84799ffee8f6387fffce5c589eaf7fedd580b74230e76920