Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:38:23.397829Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2412.09919.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:38:23.397829Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6e269592-11af-445e-aa6e-a0cf055d1045 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Accessed: 2024-09-30
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5cce60a5-a4a9-4ed6-90ef-b513cafce434 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Flamingo: A Visual Language Model For Few-shot Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5b1a2745-d8f5-40b8-8072-e37ceeccd44d · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Deepspeed-Inference: Enabling Efficient Inference of Trans- former Models at Unprecedented Scale
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0ed46894-0977-4b78-b1a0-fcc333a4411a · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24e6e467-6e9d-48d9-84df-282eb0896330 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 66863b66-26b1-4de9-999c-5059196ed2d0 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Token Merging for Fast Stable Diffusion
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 165db5d1-38ec-480f-bb5e-65fb6fdd29d3 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Token Merging: Your ViT But Faster
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fb4e4d0b-64f3-44c5-bbd6-066c0d878c26 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Dif- fusiondet: Diffusion model for object detection
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 97cfd238-8e33-4990-bfe2-38f7dded1efd · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a20ac0e-7c12-4615-828d-320b0218b5bb · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c5aca54-ace4-42a8-8118-b1de65e80b0e · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens The Llama 3 Herd of Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2330e1df-0570-4baf-9e50-13fe6758d348 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6c8bf51d-4b97-4f92-8a27-f03d3374c844 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8781f6cb-a499-4491-b7e3-42c24537ec2a · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3f706293-107d-4993-bbd4-e83e824ee93d · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8f1a4dc-1609-4fb8-9d2b-03606732a319 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Making the V in VQA Matter: Ele- vating the Eole of Image Understanding in Visual Question Answering
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f1228d8d-5a4c-4164-9e72-a3f607853442 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Vizwiz Grand Challenge: Answering Visual Questions from Blind People
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6a00835a-f980-4756-a9df-70f5f278827b · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LoRA: Low-Rank Adaptation of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 858c5d64-c8a9-495f-be26-980af22e3251 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Vision-based freezing of gait detection with anatomic patch based representation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5b56b643-4a71-4cfc-9c8a-74f0e0f86360 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens GQA: A New Dataset for Real-World Visual Eeasoning and Compositional Question Answering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7794ffa5-db67-469b-b98c-b57fb734b362 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3431e70-d808-426f-8882-589bf490cab8 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Chat-univi: Unified Visual Representation Em- powers Large Language Models with Image and Video Un- derstanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e95e7ffa-4f76-4850-ac58-5a023780b521 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens ReferItGame: Referring to objects in pho- tographs of natural scenes
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 050cfcae-719a-4d2b-aa14-f9261a73546f · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Shamma, Michael S
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 68eec932-30cf-47ef-bfd9-ffb94fc0e74a · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988303e2-7d51-4828-bbcf-77e5927944d9 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 06f1b6c1-cfe4-4fcd-b12b-9167853b30d2 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MVBench: A Comprehensive Multi-Modal Video Under- standing Benchmark
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 518762f7-23b3-4235-92d4-4c1344c7aecb · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VidToMe: Video Token Merging for Zero-Shot Video Edit- ing
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51dd0415-e64e-4783-9a9e-fda1710687c5 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Evaluating Object Hallucination in Large Vision-Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 92a16398-b545-44e4-8979-57ae413ca2be · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2e3f976d-c981-44d5-a858-17164d855aee · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5c0778-e792-4e59-8a65-efc0971f955e · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VILA: On Pre-training for Vi- sual Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32ebfc29-6c55-4996-aae1-a76ca8583fbc · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Visual Instruction Tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 07d0d442-44ae-4980-a08c-383de4b60c55 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MMBench: Is Your Multi-Modal Model an All-Around Player? In ECCV, pages 216–233
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7a9a20a6-08bb-4b17-9cb3-637fd6a280f8 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Decoupled Weight Decay Regularization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e65f279-9a63-403a-824c-3ac0e044ba28 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 548e78ee-8e54-4b6a-8702-cc6a77bc6639 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Autoregressive omni-aware outpainting for open- vocabulary 360-degree image generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e4ff90b6-9572-4e9e-bb04-3ebbee2a1dcc · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f691b681-a29b-4662-b668-5adb143c233a · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1380e099-43ab-480b-8e23-5aa7f9838e04 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Egoschema: A Diagnostic Benchmark for Bery Long-Form Video Language Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ddd3961b-6767-4666-8164-aac1912eaec5 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Generation and comprehension of unambiguous object descriptions, 2016
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 185ae4d6-ef84-462f-a676-5b1346e15f11 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Ocr-vqa: Visual question answering by reading text in images
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0ae5b099-655e-46a9-b95f-99c036a5e7cf · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Perception Test: A Diagnostic Benchmark for Multimodal Video Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f9c5f32d-f17c-4c41-a43f-f69bcfcbb382 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Learning Transferable Visual Models from Natural Language Supervi- sion
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f605098e-75ae-43b5-9fdf-74ae1e73dc5b · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge, 2022
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f567e276-c483-4e2b-ac6b-118259650460 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Llava-prumerge: Adaptive Token Reduction for Efficient Large Multimodal Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b534e526-2aa0-4464-8f86-c197afa5c564 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0203478c-b5fc-41e7-a016-a25cda20fc82 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens TextCaps: a Dataset for Image Captioning with Reading Comprehension, 2020
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be74fcc7-6dc7-4e10-a015-78bf9100fdce · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Towards VQA Models That Can Read
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4c438198-f11a-4cd1-a214-35405fb7ca7a · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Moviechat: From Dense To- ken to Sparse Memory for Long Video Understanding
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2d20aec9-ae63-4483-a3d9-132da3fc07ce · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 927d1f5a-eb39-4d25-a0d6-5dd9d7e7c996 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58606c6-b27b-4f42-bc8c-f950dcfea81b · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VisionLLM: Large Language Model is Also An Open-ended Decoder for Vision-Centric Tasks
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 40d9be43-8b8c-4ab3-b5f0-119fd11e7d43 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LongVLM: Efficient Long Video Under- standing Via Large Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e7dd3e38-dc97-4301-ab75-fa784b31980b · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video Question Answer- ing via Gradually Refined Attention over Appearance and Motion
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ad4deac9-1809-4d49-8119-ad02323728de · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Qwen2 Technical Report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1440bf94-3c60-4b0a-a6ba-06e2acd78b34 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Segmentation
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bf437b6-282a-4b7d-9fd5-9bfd9b010e8e · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2b71cb-7949-42ce-9489-ce89004ba5e7 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLaV A-NeXT: A Strong Zero-shot Video Understanding Model, 2024
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2e7fbe65-0514-494c-b5a9-b5f9956e60c7 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Needle In A Video Haystack: A Scalable Syn- thetic Framework for Benchmarking Video MLLMs
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1d57da18-8074-49fb-bc71-4fbfdcd2012d · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Clip in medical imaging: A survey
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4e9f2669-99f0-4716-809f-fdfd699558a7 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f8d223f-0972-4550-9f92-406505997ab6 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens A closer look at the cls token for cross-domain few-shot learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1a170a45-5858-417e-ad21-d9e763a9e8ce · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Training Details We adopt a two-stage training strategy [9, 30, 33], dividing training into pretraining for modality alignment and fine- tuning for instruction tuning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be9a685a-eca2-48a4-a5dc-bca19ff2bb72 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Additional Discussion on Different Frame Se- lection Features
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f95bd0b0-8aba-4b16-b736-e932672b0924 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Game Science
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3039831d-a05b-48b2-935a-8e8423126b96 · outbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Unresolved cited work
Reference 470
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.