Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:18:18.203170Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2501.15513.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:18:18.203170Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T04:39:26.597389Z
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0594859a-02ad-4274-b5cb-675f791c02be · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b723536-4ad9-4cf2-a9a7-dd830a5299c9 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8072065-f2ca-4289-aeef-d172b1250d6f · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da7685aa-0d17-435c-8389-f3141c41ff2a · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler On attention redundancy: A comprehensive study
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d0610ac8-318e-447b-a2a3-8986869d718d · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fde84b7-b003-4cf6-8e4e-75596a287cc8 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c371f93b-dd3b-4003-8e41-3bdd221f361a · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94764db4-f09f-4834-9d11-812e5a47907d · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Efficient Multimodal Learning from Data-centric Perspective
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1228130b-8da1-477f-8bf7-98b9e7adb483 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea93c03-052e-44e2-a882-852655171b02 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Qwen2.5-Coder Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee1e81e6-e1d9-427f-89f6-59ac41686970 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Phi-2: The surprising power of small language models.Microsoft Research Blog, 1(3):3, 2023
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e7f2384a-558d-4916-b860-04b6f227d507 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Chat-univi: Unified visual representation empowers large language models with image and video understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5035d900-4a80-48fe-97fb-804f35289731 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39898d34-2355-4ed0-82d8-51ca383af8ed · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler VideoChat: Chat-Centric Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07ab9965-e9d8-471e-94ff-916d0f37064c · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 584475cf-7c6c-47b5-9226-14e34cbec53e · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Llama-vid: An image is worth 2 tokens in large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6334ebe2-a2e4-40c8-a56f-ac9683e0c18d · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a71bda-cba0-4041-965d-e16e645a9f62 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee1fd6ce-9700-4652-9bbc-78cbd819c634 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Llava-next: Improved reasoning, ocr, and world knowledge, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87cf6f5f-bda8-418f-8086-cc2bfa6f68b9 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72a5ebb0-6080-4aa6-8c4f-b2863db0fc86 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler St-llm: Large language models are effective temporal learners
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 38287e0a-7220-4daf-9cbd-e6fcfdf0004d · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee46f66-a036-46a2-be2b-a8e42416aa54 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c07d4a9d-4c35-428a-866d-5cca2a9b9ade · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler DINOv2: Learning Robust Visual Features without Supervision
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94812177-582d-490a-9e94-0bbc74402419 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Learning transferable visual models from natural language supervision
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294d4da3-d181-4240-89ef-097f885f3295 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Moviechat: From dense token to sparse memory for long video understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 722f96a5-9ef2-4191-b2a7-581d452959de · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Gemma: Open Models Based on Gemini Research and Technology
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49048db-9381-46d5-b43a-c05e9d6b6cf5 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0a81bc9-a7e6-4d8f-9d69-ee685f2fb426 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8b79b0-698f-4177-b0f7-7f4bc1b232d7 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 35d0e1c7-193d-47c4-8c84-d15690c79cfa · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d7feb4-d9bf-4e6f-bcd9-228401fa8cf8 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Sigmoid loss for language image pre-training
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b280291-3a34-49c4-b416-3ed501db93fa · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef89105-5643-4fc6-b894-96a13c1bc000 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler TinyLlama: An Open-Source Small Language Model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86aeb704-3adc-4170-afc5-5eac48021193 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5d7b0f-5675-4012-8c79-da07e7ebffe9 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce12156-d8cc-4ea6-ba88-ee60aab2a8c3 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede505cb-e091-4b28-a534-6a91bffb337e · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eba6fca-bc13-46c2-9fb9-6984553edddb · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler MLVU: Benchmarking Multi-task Long Video Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2216e17b-a605-4ba7-9dc2-b2692de0cbf6 · outbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler The best results are indicated byboldface
Reference 512
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca6bb9bc-a892-4943-a3f2-fc70019a7e86 · inbound
GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.