Pith. sign in

Paper Citation Record · LEDGER

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM

As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.15816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15816 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:12.577103Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T06:33:36.846345Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T06:34:40.990815Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8fc6bbe7-7958-4e92-bd2d-706296eb87c7 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.474248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.474248Z digest=sha256:a18c2119b772b8f1d2889f63f3bc6419cd3237109d6a81703a48e2d114125f9a

Observation 8d8a034f-3e83-4553-b13a-e6c2bf8d0a3d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.862879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.862879Z digest=sha256:27971b5fe48b317b29faca7091e890cbc5174e915e688c990b0b172ed49c7e43

Observation a33f1fad-b67d-4fe2-9011-183d95200ed5 · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.933843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.933843Z digest=sha256:4de582f85d4404d932562f01a3939a96bcfe65f844f6f60f600957872d6fc3ea

Observation f6891c77-ab35-42fa-b24e-ef3c1e9f02e6 · outbound

This paper cites The Llama 3 Herd of Models.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.001417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.001417Z digest=sha256:c0b813e574bd46d07a28e3476a9f6cd9fe4c88004d4922bc9655922a7821d2a1

Observation 94df1105-1a71-4884-8784-4d314371a3d8 · outbound

This paper cites Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.049280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.049280Z digest=sha256:4482973c87e2d8ccdf3011339aecf062ca7343ffc19a731937e1e5701f24bfa9

Observation d2b73828-db75-412c-864a-12046d6b3fda · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.166214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.166214Z digest=sha256:129d041859fb60d4cd4eb6c3c6810e87158d94366b602d2d195a70d1f9516056

Observation 1ec9bd7f-f068-49cd-898b-b626a761f3dc · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.230877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.230877Z digest=sha256:6588cea6200bea41527dfde30ebfab66f100b1da46b49ea39d2de2d16ecf7070

Observation 94b725be-4ed1-45d0-8108-8c76e00512fb · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.315553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.315553Z digest=sha256:500a61786f11acc7ac46b70a482f75c68f737e89d1e3cfe09509cfae8abec27a

Observation 196ae3f1-9380-4d33-b451-9513513ea221 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.437146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.437146Z digest=sha256:f3a0d87b29be9680f403dc4285d59ca8961c66e40862f0f27a81beef31303afd

Observation 4a86901b-42a9-4789-a6b3-f2f91fcd9c1f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.683559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.683559Z digest=sha256:37c62dd82d89785fcf2b5e46780e749042a289d246038ca7130811d0f872aea5

Observation 0ab86f51-7054-440a-9b1d-2c94dd4b4332 · outbound

This paper cites EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.721014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.721014Z digest=sha256:645e3d7ecee2b06a8de2fbc2732c6be2b397010c5a1e9bea8ee936f90a54f576

Observation e0eeda1f-9f52-4ae6-9519-2ecb36638b25 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.922497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.922497Z digest=sha256:718d091ad4c7da909c8af96bc437a91d5a0bb772944638a1f3f41c91270b4c7f

Observation b2fbaa53-77e2-40d8-81f3-e9f1d92dd550 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.997946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.997946Z digest=sha256:592a6eb4c27394c15a20c8f83a1f85a5f0e250df16467a2131af5500ec0732e7

Observation 176cf2c9-49f7-4d31-8b89-614b50ad2a88 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.085010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.085010Z digest=sha256:ec870ca2fbcda2b69177ff7686c70c98bb2316cb9c1e5e4a5b7164e05ab56baf

Observation cccd9eeb-d8cc-44b5-a363-03d7c376bcdd · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.159249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.159249Z digest=sha256:c1ad188981ae058d4d6edc5eb662a18b697aba5adbff34d0d758f861d7c0f7d6

Observation 4a2e253e-b2ae-4f1c-89e2-10eceb1764c8 · outbound

This paper cites Qwen2 Technical Report.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Qwen2 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.191635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.191635Z digest=sha256:cf93bf6ea537c613d1650b533cdbda2f3a306718717c4615991a319641c6e1e3

Observation c45523ae-bfe6-4356-b6b4-3f7ce6e92d7d · outbound

This paper cites Long Context Transfer from Language to Vision.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Long Context Transfer from Language to Vision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.258703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.258703Z digest=sha256:eab4628ac2f6fdefa97ba975844c144f0c3a6f7ec2036cc7badef61c7f78f67d

Observation 709f18f3-12eb-4ec6-8bff-aff98e463c23 · outbound

This paper cites Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.354668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.354668Z digest=sha256:a7f087a766976e67c1137399a9ebd0a3e4baf46ad7a2243f55717de632512de9

Observation f61d0b3a-33cc-4d09-9490-858863c934ec · outbound

This paper cites AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.434364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.434364Z digest=sha256:0856cd40f55dea9f1b07cb8d65fd3d027170f8f22399df49cecaf79a49cd81cd

Observation bfd9234e-3dbf-4d97-8a2c-b8b439605a8a · outbound

This paper cites an unresolved cited work.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:13.582959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:12.501855Z digest=sha256:47a204f3d46acdcacee10d52b6893b01596555e2df8eac9786dccef7f6f77598

Observation 5871d1eb-3091-4db2-ba61-b4acaf0125b4 · outbound

This paper cites Table 7.Evaluation on a comprehensive set of general multimodal benchmarks.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Table 7.Evaluation on a comprehensive set of general multimodal benchmarks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:13.371875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:12.577103Z digest=sha256:8f0e71b3bc75323b6d9ca50767613dcd54e54602dad4c862d5ae59748639d1b9

Observation 4251acff-7482-4f44-af19-847fd10a068c · outbound

This paper cites Cross-Self KV Cache Pruning for Efficient Vision-Language Inference.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Cross-Self KV Cache Pruning for Efficient Vision-Language Inference

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.783480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.783480Z digest=sha256:ec4b3decc34de519a607420ead37967ee9afa5a62ebd75f979aa39e74d516c45

Observation c1cb8a37-0ca7-474e-885f-539cee6ae2e8 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.612694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.612694Z digest=sha256:7cc88dbf6a686ca1f1030105146389f2fb69db71a95d461ec3a2d1c7a79f0299

Observation 4daf7dd9-aab9-4b1e-9937-3369bc9fb644 · outbound

This paper cites J., and Yan, Y.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM J., and Yan, Y

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.852764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.852764Z digest=sha256:374e2d7f062c5ed94475d3706284a2bd0ded8dbbf98937d00750086c5f2b1373

Observation 75c46191-e6b4-40e8-a942-e0c02d96293f · outbound

This paper cites InternLM2 Technical Report.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM InternLM2 Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.553356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.553356Z digest=sha256:05646dc20b514be980cdc095ae3d175a98fb4ea3f5c338a97a268d915038eeb8

Observation 8824895c-1718-43fe-88e5-984ba981523e · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM NVLM: Open Frontier-Class Multimodal LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.736485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.736485Z digest=sha256:6a6b8bc00b81cdc4bfc22e79f751ca3d7cfc803a4ce175ee3d541d89d1af2fe5

Observation 519b9643-f2d2-4432-b1a8-638dcc103ef5 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.643408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.643408Z digest=sha256:6cac024cf3b41e7ce83fc656df1a71a999b14a3cdab324bf49df6c7191afcb41

Observation d6970a0a-68eb-4b8b-8841-49d7ff231925 · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.527741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.527741Z digest=sha256:a7ece58b2a6b9a1b1a2497456d8dc132bd065992f4952d1a06c7036a5f867745

Pith citing papers

Observation f2b7f28a-61ff-4f35-b0ed-ca024f3c5c1b · inbound

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning cites this paper.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:34:40.992784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:33:36.846345Z digest=sha256:80ea74c43be116c405a77107cc1687851fb760a2692f3626250046e60ce9506e