Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-14T19:15:13.205594Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2605.13375.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-14T19:15:13.205594Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c517aff6-f6d2-4768-929b-ed4bf980772d · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8cff7443-81f2-4ae9-a985-b13d792f7200 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d83dc8ba-3a99-492a-9398-c1c3c92df709 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Improved baselines with visual instruction tuning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e23d07a7-28ae-4173-a89c-86cf1e426026 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Visual instruction tuning.Advances in neural information processing systems
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27be793a-9cc5-4b91-88c6-9813ead50fc1 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7e907cc3-cd78-41d9-acc3-44c064633476 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 14720828-cb05-44e2-a3d0-8bf9473d259c · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Qwen Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ba6c2b71-a99f-41a4-a12d-708ff2c80317 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Qwen2.5 Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 62663dba-f79a-442d-b3fa-bbb6a665b7db · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f37c86b6-fbc2-4b30-a1a2-112474976abd · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Internvideo2: Scaling video foundation models for multimodal video understanding.Arxiv e-prints, pages arXiv–2403
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d98ab346-77fb-4b13-9544-93c49b75b348 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97efd2c1-4cd3-4cf7-98d8-60f29608211d · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 942b7986-1252-47f9-8b46-db1d21578ef7 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Qwen2.5-vl technical report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e041039-bc6a-4e0d-a90d-44ec21f78f2f · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 42f656de-bbed-48e0-af1e-38644ace1ebc · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a55e99f-20c1-41ae-bdf9-fa694754708a · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cf81fa2d-4e0e-4663-8e9d-6d135465d057 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2876ae61-2a2e-4ddd-8fc3-e6edf48d55dc · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Sparsevlm: Visual token sparsification for efficient vision-language model inference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fafd32d3-7b06-4cb4-8b3d-f098384652b7 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Token merging: Your vit but faster
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 001e77ed-87ee-41e7-a706-7678085aa48c · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Framefusion: Combining similarity and importance for video token reduction on large vision language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 406b3d0f-0746-4229-ac86-6b0547434ea0 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Dynamicvit: Efficient vision transformers with dynamic token sparsification
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b294bdee-7fdc-4773-9b5b-15ac212bfd8e · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Smarttrim: Adaptive tokens and attention pruning for efficient vision-language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f8081c2-751a-483a-bc72-76116d8f089b · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Visionselector: End-to-end learnable visual token compression for efficient multimodal llms
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 16e4cb1c-cbaf-48ed-9bd7-0fa66784b4e2 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Efficient multi-modal large language models via progressive consistency distillation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2977d013-112d-4498-b2e5-6dfc73ffaa3d · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 232a81b5-dff8-4c7f-adb9-9021a270e528 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4e1e9b15-5537-4968-aa6f-ab794d2ab01d · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6234653-aa1e-496f-9d94-a5d7c58ec71b · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 204b7347-2110-4eb7-bbdc-7a15872115fd · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2894c32c-03d0-45ff-8cf1-628b324c22d6 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Qwen2.5-VL Technical Report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e7a0828-266f-44f0-981e-5f06a324a12e · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Attention is all you need.Advances in neural information processing systems, 30
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01acc121-5f1e-419a-bd97-6bb4d4ed9dca · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 06aa54cb-1a7a-4416-aa69-5b7e6356ad6d · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Film: Visual reasoning with a general conditioning layer
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63ef317b-ac85-4a25-ac83-29ad5f54fa58 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Sft or rl? an early investigation into training r1-like reasoning large vision-language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01623c75-8ef5-4c04-a3f4-4c736f1406e4 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8929df7f-efb7-4e82-98ee-24b299109575 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cf90a508-84b8-4496-83ce-544465813af4 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cf64293f-0f2b-4f7b-b18f-67484eccb2ca · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Towards vqa models that can read
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf289a30-9e9a-4147-a3fe-367855a28437 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models A diagram is worth a dozen images
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2835d0ef-7f8f-4f45-b5f1-9a2361fad672 · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cf4eb7f7-0569-4c19-88f2-f3548d869dde · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models MMBench: Is Your Multi-modal Model an All-around Player?
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8b473c9e-256e-4ab2-ad61-0ae7543b54bb · outbound
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.