Pith. sign in

Paper Citation Record · LEDGER

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models

As of 6 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2605.13375.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.13375 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T19:15:13.205594Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact16
  • verified fuzzy25
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c517aff6-f6d2-4768-929b-ed4bf980772d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.117329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:b268e23aac5bd5623b25b20ca58435d2f9346ac90f8b82a7e7120b19a77eb91c

Observation 8cff7443-81f2-4ae9-a985-b13d792f7200 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.756941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:e803830a515c5dc11959999041fc2864b6ee645e8515aff546b313cb1787d512

Observation d83dc8ba-3a99-492a-9398-c1c3c92df709 · outbound

This paper cites Improved baselines with visual instruction tuning.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Improved baselines with visual instruction tuning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.121890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:d13a74204885227917ee667adc713de8dfc81b039f0c204ca3028424f7bd4be1

Observation e23d07a7-28ae-4173-a89c-86cf1e426026 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Visual instruction tuning.Advances in neural information processing systems

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.126711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:e4aeff0aa2dca47196a4949d2a913b3109a5559f1dcf32dacd685a6ea9a1a248

Observation 27be793a-9cc5-4b91-88c6-9813ead50fc1 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.749786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:56668d78c669c01ef8f0f78c177cf7b112c91216131764e43d7c8e82f91df48e

Observation 7e907cc3-cd78-41d9-acc3-44c064633476 · outbound

This paper cites The Llama 3 Herd of Models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models The Llama 3 Herd of Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.762418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:4c78f8b208221e78d8f5e7ceaf44f8b42ee947c96167074da0ccb4f9adba3bac

Observation 14720828-cb05-44e2-a3d0-8bf9473d259c · outbound

This paper cites Qwen Technical Report.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Qwen Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.768044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:42168eb1a0ca80b8821109999647aad82f2d00f32f4d871401cdabe457c8f75c

Observation ba6c2b71-a99f-41a4-a12d-708ff2c80317 · outbound

This paper cites Qwen2.5 Technical Report.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Qwen2.5 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.785305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:064cbbeaade7ca2d40e3b1e922c4225b58c2e3e9e6f4580287c81427794dc1ae

Observation 62663dba-f79a-442d-b3fa-bbb6a665b7db · outbound

This paper cites Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:50.819402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:c717edebdde8a07fcc79a6876f2b13d22d11ef36b711ba00b5db206e10ba922f

Observation f37c86b6-fbc2-4b30-a1a2-112474976abd · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.Arxiv e-prints, pages arXiv–2403.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Internvideo2: Scaling video foundation models for multimodal video understanding.Arxiv e-prints, pages arXiv–2403

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:31:33.511181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:c7c00b310a59ba54292e411436621e128b1e681ff8354c03cca10baf4311d0d4

Observation d98ab346-77fb-4b13-9544-93c49b75b348 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:50.811951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:a9ba4cb4e0a7ab72a4bc1ac8d8360f0ae1994414e1bd433354c52e272ee29d4b

Observation 97efd2c1-4cd3-4cf7-98d8-60f29608211d · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.017253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:ff5f4597e195c28b5a72da79a084f0484cf4b63b155417013dc41262b9a9ffee

Observation 942b7986-1252-47f9-8b46-db1d21578ef7 · outbound

This paper cites Qwen2.5-vl technical report.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Qwen2.5-vl technical report

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.045910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:7cf9cdb9926db7a93deccba018c9b4ef4583f844ded19cf6912c21a86665ec9d

Observation 9e041039-bc6a-4e0d-a90d-44ec21f78f2f · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:44:47.683285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:12e6c5a56f91fbad18df87c5fc5b9ccd0974666f39042bdb579428ca281d8291

Observation 42f656de-bbed-48e0-af1e-38644ace1ebc · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:17:50.798831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:3759f4e4848bb1f832a9eb932bf13531c9768861c3d15916a2e67c451b3f9058

Observation 8a55e99f-20c1-41ae-bdf9-fa694754708a · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:50.792438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:817ddc436c52add1f2dbcce87806b0efdd327a4c8b15b6f1240fbd92b8d8d849

Observation cf81fa2d-4e0e-4663-8e9d-6d135465d057 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.093017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:5cd23660576401423691d830b59c101e7c17d5473fea135172386173c2e6db8f

Observation 2876ae61-2a2e-4ddd-8fc3-e6edf48d55dc · outbound

This paper cites Sparsevlm: Visual token sparsification for efficient vision-language model inference.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Sparsevlm: Visual token sparsification for efficient vision-language model inference

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.087782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:7c6d56e45c00a1ae6cd6c8e0f1841ea8351635253105578d1d93577b3cdf45d4

Observation fafd32d3-7b06-4cb4-8b3d-f098384652b7 · outbound

This paper cites Token merging: Your vit but faster.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Token merging: Your vit but faster

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.041341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:2512fe73b24913854058265007145f85d39fd8c9c6accbb02c2f8830212d0e84

Observation 001e77ed-87ee-41e7-a706-7678085aa48c · outbound

This paper cites Framefusion: Combining similarity and importance for video token reduction on large vision language models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Framefusion: Combining similarity and importance for video token reduction on large vision language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.032076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:2ed9208bc9f18744b8517b2a30d3dd85f2315b52042708734cbc876525fca9f8

Observation 406b3d0f-0746-4229-ac86-6b0547434ea0 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:31:33.520559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:336fc9eaa2ef362ad632f1b584365bc2d2173f4679d8b5df9ae6ec0e0fcb0846

Observation b294bdee-7fdc-4773-9b5b-15ac212bfd8e · outbound

This paper cites Smarttrim: Adaptive tokens and attention pruning for efficient vision-language models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Smarttrim: Adaptive tokens and attention pruning for efficient vision-language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:31:33.524662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:bf73315f6a4bcd415c483c14d31cb2536866582104063cfbe3ca6acbf6bec5ff

Observation 1f8081c2-751a-483a-bc72-76116d8f089b · outbound

This paper cites Visionselector: End-to-end learnable visual token compression for efficient multimodal llms.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Visionselector: End-to-end learnable visual token compression for efficient multimodal llms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.082831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:c438249db6abe609162bc57e75136a12acca53386d657f611b6617bee0887dfd

Observation 16e4cb1c-cbaf-48ed-9bd7-0fa66784b4e2 · outbound

This paper cites Efficient multi-modal large language models via progressive consistency distillation.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Efficient multi-modal large language models via progressive consistency distillation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.107841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:a99d46be929b742d9c85cf021778e68855187ec5055d0330c49f540b82b10bd2

Observation 2977d013-112d-4498-b2e5-6dfc73ffaa3d · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.073062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:c77bdd3641138c446c5a8a9c3cc296a3432491f4939f6c8a188978a45599a80f

Observation 232a81b5-dff8-4c7f-adb9-9021a270e528 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.036573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:9fbfb08285da007cdce7428a9ae4383b11faed7972399ae1912816c7ecc82f88

Observation 4e1e9b15-5537-4968-aa6f-ab794d2ab01d · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.022689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:0046622ae0a90a8b1df6a6c39759dccb1293236ea894a056b207aaf72bb3b49f

Observation d6234653-aa1e-496f-9d94-a5d7c58ec71b · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.780093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:93473ce7cc0be5299cfcb7b5d3384071c1e80c785470dbe6f752d60c51707f2c

Observation 204b7347-2110-4eb7-bbdc-7a15872115fd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.824171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:c26b933c081ff653cfd4a8bbc225c43188adb62ec36916ebf61d2ef98e8d9f7c

Observation 2894c32c-03d0-45ff-8cf1-628b324c22d6 · outbound

This paper cites Qwen2.5-VL Technical Report.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Qwen2.5-VL Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.774101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:47c9a733894085347e60fc01c0c573751f3835a3776c27128241d90a128fbbda

Observation 5e7a0828-266f-44f0-981e-5f06a324a12e · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Attention is all you need.Advances in neural information processing systems, 30

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.027652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:82716a4504dd8e0917a95ef56a805f2b955045e5ba21ba018442723b4d52a04d

Observation 01acc121-5f1e-419a-bd97-6bb4d4ed9dca · outbound

This paper cites Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:50.839649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:97b3f0c6ac353393b956e1da0f158e959f084bb2eaf5384a3965b0d02c37fb4c

Observation 06aa54cb-1a7a-4416-aa69-5b7e6356ad6d · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Film: Visual reasoning with a general conditioning layer

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.112409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:3089412eb65dc124b78f8c766b6dec01308dd49a19a195afb16a8b7138334662

Observation 63ef317b-ac85-4a25-ac83-29ad5f54fa58 · outbound

This paper cites Sft or rl? an early investigation into training r1-like reasoning large vision-language models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Sft or rl? an early investigation into training r1-like reasoning large vision-language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:31:33.528814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:596ad4f14b5ca1bea3f7ecf9eeefb41b06c6e4f4760dfe6e0668c1b67043a1cf

Observation 01623c75-8ef5-4c04-a3f4-4c736f1406e4 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.844426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:5bb87940f266065b4f6401c802f2fe530f34820ee14a6856dd8eeb7f854139b2

Observation 8929df7f-efb7-4e82-98ee-24b299109575 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.804755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:a72c1596c0bf0e237fa76445eb4814e80783decec89e9772a4099d3894e6b4f2

Observation cf90a508-84b8-4496-83ce-544465813af4 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.098008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:24ec5fab635ae1e0f09d25bcf4e93352fb781b57beadf391e3239aa848a9fe2b

Observation cf64293f-0f2b-4f7b-b18f-67484eccb2ca · outbound

This paper cites Towards vqa models that can read.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Towards vqa models that can read

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.102959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:6cd57843423076f40e5665d218bfa8d4f63beb6d9f749e399eb3bea633186926

Observation bf289a30-9e9a-4147-a3fe-367855a28437 · outbound

This paper cites A diagram is worth a dozen images.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models A diagram is worth a dozen images

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.065991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:75dee91712b608f886a8cdc48ddc245dcceca2d14f6ec9e27257b9c88316e304

Observation 2835d0ef-7f8f-4f45-b5f1-9a2361fad672 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:51:34.077728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:4c600a5d1581c156d07a334b610661d59e840f7ae80e0a7b5301de95b477466a

Observation cf4eb7f7-0569-4c19-88f2-f3548d869dde · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:17:50.833919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:7d5ab1d700b501ae798ca87ed473786079917fa5c8cfe9c68cb54ce54d0e317a

Observation 8b473c9e-256e-4ab2-ad61-0ae7543b54bb · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T19:31:33.515784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:bb42bb8b3928bdea686cb12b703992ce30b378fcfbe4d7f957517beaf97a81cb

Pith citing papers

No inbound Pith citation observations are available.