Pith. sign in

Paper Citation Record · LEDGER

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

As of 7 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 27 inbound Pith citation observations for arXiv:2506.10967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10967 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:20:33.522423Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:11:49.179772Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:59:57.457886Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9580dc0-9e63-41c9-860c-842923ee40f6 · outbound

This paper cites DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.500266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.500266Z digest=sha256:0137ce1e6f9d6824c7defc20c78307805f31904213492005973a48e3eff632b1

Observation e25074c2-8e0b-46be-a3d5-4b0e7c0947ca · outbound

This paper cites Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:20:35.249313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:20:30.843591Z digest=sha256:4415f49d037a0665e760527daa26bdbc552c45a6049ce660b1854ad5b8809bba

Observation bda22afa-70e1-4b24-b976-6ab55fef2e88 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.095679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.095679Z digest=sha256:27a53031e4256224d72d211d4b1abde6a76c19f2e367b70967a870215086130b

Observation 75faae6a-b724-4b8b-abd7-d58a1dbf889b · outbound

This paper cites Mistral 7B.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Mistral 7B

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.290255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.290255Z digest=sha256:39d550e9a1bc79bfc3a367bd7c14ccc7ba8ec2de1af993013f9828b9b9bb6edd

Observation cb6f6bb3-cb43-4549-877f-ce54049d6d93 · outbound

This paper cites A diagram is worth a dozen images.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs A diagram is worth a dozen images

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.400104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.400104Z digest=sha256:d1cc92f61db7784b4417829156e3605d51b483f3ca1e2f2ee2585729c06be4cd

Observation 0b5a60a5-a5be-480d-9837-88bf333bfa83 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.581672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.581672Z digest=sha256:7a276f3ae8adb5c9569961423839a9e5171d28f90b055018fcaf3fac2da28957

Observation 880c77ef-a514-40ae-8dac-e2ee90fdcceb · outbound

This paper cites Microsoft coco: Common objects in context.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Microsoft coco: Common objects in context

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:20:34.975202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:20:31.664718Z digest=sha256:ce5cb9b4aba7baa5ef7e7aabe8e53e440981362bf7cfb15f3e33db465a361a75

Observation 731567cb-7de2-446a-b222-3b662f87a2da · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.753984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.753984Z digest=sha256:7dd26a07f36247adbfc04c09ff7d7b6e1bd4ed3f8ec007be5f4da681da439edf

Observation f46ba898-242d-4dcc-a67e-660f1cd1d0a9 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.837013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.837013Z digest=sha256:14402d21a7654f96b0c43dce9690dabed4e18d95b0dc25bba51c1f3fe086e3b6

Observation c8783900-4263-4e24-8bf3-dfdd7c03543e · outbound

This paper cites Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.068161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.068161Z digest=sha256:cd72ed41756aa266c946fae6e8e0659cc2f6b06e7f9c83f16783320fa8ba9a58

Observation 5d13ddfd-0112-46ce-bfb0-8a3a9baa8099 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Gemma: Open Models Based on Gemini Research and Technology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.245624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.245624Z digest=sha256:00b081ac694c4960c9788b9f8eb672ee89fc23edc5804d4f430fdebef412d1bc

Observation b416dba4-3258-4a0e-ae97-07e441622bd0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.325718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.325718Z digest=sha256:1b867615ff993ae46f774dd9599c2bcb700554d02d25c1313a308e7e10b077b1

Observation e2fd1dc9-808b-4d11-be3a-44643dea61ac · outbound

This paper cites Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.526592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.526592Z digest=sha256:ec6dfc888bc820a65d9991032c59e314f251589a8d83b1b3c129fc94a9cbb78e

Observation 5a32adc7-b150-4ae5-b10a-6b9661ed0910 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.613673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.613673Z digest=sha256:e2c509cdbc9a7870900730d8b0fe21e565111683b8bd05f431f718754dc63d0e

Observation c735d4a5-be86-430e-9d94-702ab8595c14 · outbound

This paper cites Qwen2.5 Technical Report.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.701129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.701129Z digest=sha256:7b9d6f21b241ba29fd02f1c7c27a17141a2f1fb517e6e590f8588bbd02f6f3d2

Observation be676b47-8fa3-4528-acab-1b7fa509bc28 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.927138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.927138Z digest=sha256:420cd0a5129ffccbe0be1efbe247ed2b6c0601f7c17dc5b10cc745bd57e1e965

Observation fc8edb99-3740-495f-8468-1737e9cc18aa · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:33.004390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:33.004390Z digest=sha256:6c76b76dd1f6b1a10fa7a5740bdefdb83faddaeb5f1645e2c8a701fd0719bbd3

Observation 26b7f326-3691-4767-a429-eddfeaa91aef · outbound

This paper cites Long Context Transfer from Language to Vision.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Long Context Transfer from Language to Vision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:33.079691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:33.079691Z digest=sha256:0ed73da7fd7cf8f6808c205b59311687ef16f89531dbdb1d157defdffee34be6

Observation fe9d77f7-1109-4755-be87-518e9ba954a8 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:33.148022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:33.148022Z digest=sha256:f4f9bda01b7a15ae3aea4069e52d2753fceb21e11a95c42d4412f9814b23acb3

Observation a8344eeb-4409-476e-8514-35f8472891fe · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:33.229755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:33.229755Z digest=sha256:a32900afe120c7411aff01b89d0fa6f16763153d1d82957278f43fa97eaf77af

Observation c2123e3a-1bd5-40af-9385-258fcf6886da · outbound

This paper cites Appendix B provides some details of the experimental setup, including information about model architectures, evaluation benchmarks, comparison methods and implemen- tation.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Appendix B provides some details of the experimental setup, including information about model architectures, evaluation benchmarks, comparison methods and implemen- tation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:20:34.643823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:20:33.317956Z digest=sha256:5e993a9ec9183ecd27ab9f3dcd383de64523a8864d04ae7d9f87e5100b0e81fc

Observation ac6f7c56-d168-4c80-b242-6bc6aaf2d9a3 · outbound

This paper cites Therefore, the greedy algorithm runs in O(nm2) time.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Therefore, the greedy algorithm runs in O(nm2) time

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:20:34.442063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:20:33.433526Z digest=sha256:a183c117afa7ba9ea05bf4946dd8065457ef9061e15f983f4ec04b132d9f10ba

Observation 82d897ab-868b-4802-9c48-46acc1b2dfde · outbound

This paper cites an unresolved cited work.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:20:34.277604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:20:33.522423Z digest=sha256:aadb623aeb46ba0ed8b28ac3e1a43fbd5cec84b5182ff6a956015c9b9e3e035a

Observation 80118ccd-ae30-4541-b8f9-194c876cf8b5 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 1975

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.906806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.906806Z digest=sha256:266e1ca896a0f999ad0ddf10ecedaff5bce0deb418a21d00eff2b2ea1e97c2aa

Observation 6e0c639a-dbd6-4e51-b060-7be707b167b8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.510710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.510710Z digest=sha256:79c5069d30665df350c8051ac5c246805564341cb42d0888853a4b65f431068f

Observation 9b2a4f8a-0bb3-4df5-b1a9-9275074cf3d4 · outbound

This paper cites Qwen Technical Report.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Qwen Technical Report

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.562632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.562632Z digest=sha256:01432abe16284f8016c3ba381a2020a8cefc1948d12c04214989b2a18331027b

Observation 9fccbff3-c235-473c-827c-647ccd7a1c5a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.413526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.413526Z digest=sha256:fb5ab089df5f806af1b669fa638f86df3f787a71624f3f0acdba477455aa442b

Observation 38da4126-715b-4717-a507-f8506b4447bf · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.765313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.765313Z digest=sha256:dba2fc548318cf4d850fc65269e6e2a5a054c0beb950ce8b9a281d4e6d0aab7b

Observation 8e4ee5fe-b3d4-4fb5-864d-64b5088904fe · outbound

This paper cites Similarity-Aware Token Pruning: Your VLM but Faster.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Similarity-Aware Token Pruning: Your VLM but Faster

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.175344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.175344Z digest=sha256:a23c6ae265586e62fdc67e1f0887332994aeb3459e11dbe765e57bda8ad8bae0

Observation 2895cc3f-6db3-47f7-999b-8ece88d84bc0 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.982439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.982439Z digest=sha256:29702fda0d29196d8b8b8d4bbbefeca2a15e8610206ad31ba842ec3a2b994040

Observation b7ffd2ae-9a9a-4e15-bd0c-5f1fa2a75316 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.949286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.949286Z digest=sha256:e2779dd3afe597649ff13110c0b63146ac3689d2139a654b830bbe7a1811e9ff

Observation 480d693e-9f13-40e3-96a6-4f0a64979a92 · outbound

This paper cites Qwen2.5-VL Technical Report.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Qwen2.5-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.635629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.635629Z digest=sha256:bf972e433901eb95ef5fd38cc037811d6df9104abe10a0dc64d858825cd7bf5b

Observation 8c99589e-8e7b-40ab-8eb9-0c9b25a35eea · outbound

This paper cites Mdp3: A training-free approach for list-wise frame selection in video-llms.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Mdp3: A training-free approach for list-wise frame selection in video-llms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.151078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.151078Z digest=sha256:1d974634d04c6b0097fbd0e47a3c0cad61eb955d158368e42e5d043c6d4777bd

Observation fd123cd7-c6c9-4391-93fb-7314899df7ac · outbound

This paper cites InternLM2 Technical Report.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs InternLM2 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.698008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.698008Z digest=sha256:d3b1f9c9d71309ac07e75cacb2fa6ee6d146c5a919401e21428ef8738c5fafc7

Pith citing papers

Observation 00c809f9-844e-4f80-8722-f6bc21ce394c · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:45.462835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:45.462835Z digest=sha256:59ee9213b35e1331fa4c422fa962890bc92514b7b0c3653e85257a038c3760b2

Observation d6e1b3bf-d94e-4c4b-959d-e59484535d8d · inbound

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models cites this paper.

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:37:35.450623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:37:35.450623Z digest=sha256:1f9e8d2af2bd29b556ad51c40643078f8dfc86c0e607848190eb368c3e0ff2ae

Observation fb5cc1b2-a8b6-494d-ace8-c05441757435 · inbound

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics cites this paper.

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:34.271065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:34.271065Z digest=sha256:85d6d0203a52d5726a98fe261239df96bc32e5957c5309c22a271f30cfa861b6

Observation 2867c7e5-692b-4aee-9976-eee272cbf749 · inbound

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models cites this paper.

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:05.867186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:20:19.653806Z digest=sha256:00b33e8c05d081f15b1345d9a9f36a91debdfe9dff6ea7030b851f8cb79a0941

Observation b952a8a0-3d83-4f70-b36c-5b1ea2f0bffe · inbound

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs cites this paper.

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:44:48.417048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:43:44.921189Z digest=sha256:453762de1fba19547b74d427d629886845a7f371092e6a7107fe7baaf379c8b1

Observation 5e597c98-ab8b-4547-8774-cd44bd08c998 · inbound

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference cites this paper.

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:08.455532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:49:28.591419Z digest=sha256:e13fca1e023e8c24387ac516cc574f18a0ff59179d233232d6768c9bbd759e5d

Observation 23d1383c-d06e-4817-b8b6-88c678f3ea48 · inbound

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference cites this paper.

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:21:19.376132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:16:31.918225Z digest=sha256:11d7b94c473796568eb7052878e969db6e78e4e5f70a0eb2ed0eaa235e15f401

Observation 3490a9f9-7ee5-4a16-8cbe-0a9c697dee77 · inbound

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference cites this paper.

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.772099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T10:28:15.293381Z digest=sha256:47c004a490e74058b14c6914fec5f8858437aa138a01f4a151b85db2c47208ad

Observation c50d254f-c09e-4cd1-872d-b44ea1beec21 · inbound

Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models cites this paper.

Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:23:58.587932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T05:20:55.448430Z digest=sha256:cf4fd43d7cf1fddf000c43051a1972f746cae09a72002f6fc5fba736ddf57b24

Observation 70eb3121-3269-4247-8ead-97d6e6805ba0 · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:43:14.998270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:40:10.152344Z digest=sha256:d6b7a57b020c73f7de3a4395154068f607d09c42c352cd0ec751245cb9b505bc

Observation bb008883-066e-4558-af78-5629603e4af6 · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T12:54:43.228543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:54:43.228543Z digest=sha256:26d1265349358c3d8a7778fa6969b85a216cb252cce5358ea75fc4ed18575298

Observation fda3f518-4f09-46cc-95a7-8aa849d3982e · inbound

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding cites this paper.

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.963802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T15:06:22.102725Z digest=sha256:dddc676fc85da08d318fa7badcb124ad547091c977d00774408120cce2fed3c2

Observation edbc6b09-c160-4a09-9a4a-b6b3f11b6fc7 · inbound

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding cites this paper.

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:44:36.407579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:42:37.401221Z digest=sha256:a83e1d03829dc33eb8944997d4a47d9439cdf48b350e8ea405097c00d3179b8e

Observation 722a29f6-c5a9-4c48-847a-f1da057ebcbd · inbound

TGV-KV: Text-Grounded KV Eviction for Vision-Language Models cites this paper.

TGV-KV: Text-Grounded KV Eviction for Vision-Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.715894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:04:54.171291Z digest=sha256:52d66e60970f4fb78f6a5b85fcaeb48c73e40e4b807f7dc3c368b17d47a1689e

Observation e6a61e6e-7b9c-428d-a4cc-5f379bf31ec5 · inbound

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models cites this paper.

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:59:57.459369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:04:20.802531Z digest=sha256:80960388938e1f1c478202bc8d04d9255b71f01f6ca656462e49e5726ff2e0de

Observation 3d777fc8-3027-41a4-9912-24a5f4ce079e · inbound

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference cites this paper.

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:09:53.798585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T04:24:38.917137Z digest=sha256:76d0baf842fcd2aba1d9dd2e44c0ed0fc786eb1e27a5321888b91e9020d36f8e

Observation e7538e8f-8d90-4758-9d8c-e24379874ebc · inbound

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning cites this paper.

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.406972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T14:47:35.377391Z digest=sha256:1f68e6342e58ae2e4d0f698835b4861457a08142f67336bd87e1b8b519ebc882

Observation af6ec779-02cc-4cd9-9823-eff642eca4cb · inbound

RADIO1D: Elastic Representations for Condensed Vision Modeling cites this paper.

RADIO1D: Elastic Representations for Condensed Vision Modeling Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:8b045ea078bb86345993681599af507852a0fbcf04af1b363518558dda0db468

Observation f15b7aa2-54f0-49dc-a303-7d5f893e19aa · inbound

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models cites this paper.

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T22:32:25.110039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:32:25.110039Z digest=sha256:e1e32fc8e19c63eb82d416226f25d67e9076b7d38e7a21722d9b135bf809388c

Observation ab72c83f-49fb-468a-906a-74a8aca1be14 · inbound

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models cites this paper.

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T04:21:09.449182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:21:09.449182Z digest=sha256:60183e072b21b33c37b33935eea5e94af1d014ad3021ea03eaec217bcce38fe5

Observation 1d712359-c0ab-48d8-8081-4f12ab56b5ef · inbound

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs cites this paper.

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T03:50:32.506198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:50:32.506198Z digest=sha256:c769edd35183d5a0254a551c8dcf284f27da662e54659cd74e184a4265aa9a05

Observation 787a3619-c726-44eb-b279-9170356085bd · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T23:59:18.339215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:59:18.339215Z digest=sha256:dadc990dc83181fa24032b4d6d5ab6cb6752570dcc10880a13f06ecff009c29e

Observation c940e4a4-e27a-4e34-b139-3498d8d14cc3 · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.622068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.622068Z digest=sha256:b3165297515ae27b5c8e3dd5513b03ef44a02fb82056384fc465de46051f2b5f

Observation d99230ee-52b9-41a9-911f-f0300818985a · inbound

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models cites this paper.

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T01:26:50.833164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:26:50.833164Z digest=sha256:19f61a4d6eae9de961a86bce12e2993e94c7a4b0845569316c5826427c43d2b4

Observation 814f8c48-6d41-4d5d-8183-f3ad966214e2 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:57.677655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:57.677655Z digest=sha256:384a9f12459e08b181f6efd56e34a264241e76c90003799a2c9e5c5f496d2891

Observation b32330f0-3ef6-43b0-a933-40627f1db851 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T00:11:49.179772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:11:49.179772Z digest=sha256:f3e89c07b1ba92db0bf0edcbe67b82748ea3925e115619e29a0d48ecc1a1d8ba

Observation 9355748a-d097-48c7-a5bc-667d093bcbcd · inbound

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding cites this paper.

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T18:36:14.731592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:36:14.731592Z digest=sha256:425b4e31f6ead859e5b36a5e2676eb9e164158d7390e5e23431f0195181e2ec5