Pith. sign in

Paper Citation Record · LEDGER

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.15816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15816 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:12.577103Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T06:33:36.846345Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T06:34:40.990815Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8fc6bbe7-7958-4e92-bd2d-706296eb87c7 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.474248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.474248Z digest=sha256:08f914d318ec9104aaa8ef3833328fcec50fb170ab18cc0cfb98c88158ff73c2

Observation 8d8a034f-3e83-4553-b13a-e6c2bf8d0a3d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.862879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.862879Z digest=sha256:245b6406e368e33973964f4db64659a0bdfa55201659387fee5b117f24e623a6

Observation a33f1fad-b67d-4fe2-9011-183d95200ed5 · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.933843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.933843Z digest=sha256:60afac6b7c47d177af465ba7dd1c471239c8e1c76eebe68b87fd4b5f7e09ef98

Observation f6891c77-ab35-42fa-b24e-ef3c1e9f02e6 · outbound

This paper cites The Llama 3 Herd of Models.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.001417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.001417Z digest=sha256:9d4fe86708661251f29710a90283fc055944163f3fdaea11c07e1efbd6793aca

Observation 94df1105-1a71-4884-8784-4d314371a3d8 · outbound

This paper cites Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.049280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.049280Z digest=sha256:f367e5e8cf9e5ba6a2351588cc8577ef583d1185acedc1e56e3d1035b713bf8a

Observation d2b73828-db75-412c-864a-12046d6b3fda · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.166214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.166214Z digest=sha256:9375a3550c431890cee31990eae6d483107ea0ceaccc2572a01cde60f90566b8

Observation 1ec9bd7f-f068-49cd-898b-b626a761f3dc · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.230877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.230877Z digest=sha256:6294ebc7bd62c14266dad0c23d1089fa947f93d0925d2314202f61aafdb6b25e

Observation 94b725be-4ed1-45d0-8108-8c76e00512fb · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.315553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.315553Z digest=sha256:86353f3ce365f554946f2137a2f373acac57f9914ef884ebab23e4c9ab1e04b8

Observation 196ae3f1-9380-4d33-b451-9513513ea221 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.437146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.437146Z digest=sha256:a200e3523fa4287ad438a865c3c0d48f6bd4291fe3e8790b18925f695e9efda6

Observation 4a86901b-42a9-4789-a6b3-f2f91fcd9c1f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.683559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.683559Z digest=sha256:8dbaa3a1e26cb446066f0be7982e5769f20f2c5c91131790c39f6757cd3c4df8

Observation 0ab86f51-7054-440a-9b1d-2c94dd4b4332 · outbound

This paper cites EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.721014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.721014Z digest=sha256:e3b9e6fe1b1db6be02fcc8857ac17d842ef3b1d2b9d57f0bf571bf834db7a8c0

Observation e0eeda1f-9f52-4ae6-9519-2ecb36638b25 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.922497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.922497Z digest=sha256:b42684b4fd7c1dbe852c88cb6152c406338b2a47f254af0eff5dedaae539e419

Observation b2fbaa53-77e2-40d8-81f3-e9f1d92dd550 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.997946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.997946Z digest=sha256:25cc01f56d27fea47eb1fe0af8273b243b36991e15903579c665ba6a6658b7dc

Observation 176cf2c9-49f7-4d31-8b89-614b50ad2a88 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.085010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.085010Z digest=sha256:bf020138c79fafc45a3a15857ab9167dd29a8b0e952a52b7a970f5040b2118fb

Observation cccd9eeb-d8cc-44b5-a363-03d7c376bcdd · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.159249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.159249Z digest=sha256:83f9090575dcc369e7ad4d1f045f5252343458201a2935aeed02758dc1f74cf4

Observation 4a2e253e-b2ae-4f1c-89e2-10eceb1764c8 · outbound

This paper cites Qwen2 Technical Report.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Qwen2 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.191635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.191635Z digest=sha256:08c3844931bf8f1f8246a2212dc82658f47d89c72be3434c57b25b96d55b7dff

Observation c45523ae-bfe6-4356-b6b4-3f7ce6e92d7d · outbound

This paper cites Long Context Transfer from Language to Vision.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Long Context Transfer from Language to Vision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.258703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.258703Z digest=sha256:d4f0be2d98ad61dfea66843e97210db4eb299b021fe9af966643ee4a755f8a8a

Observation 709f18f3-12eb-4ec6-8bff-aff98e463c23 · outbound

This paper cites Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.354668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.354668Z digest=sha256:5f401b57f073857d073b3e620c3167c00336b9d77d405799f01dd9509cb9f81c

Observation f61d0b3a-33cc-4d09-9490-858863c934ec · outbound

This paper cites AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.434364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.434364Z digest=sha256:4acfb20b15b9588ae61a6a84e52b1be67f09e2cd1c6f4351da352ea4df3ef527

Observation bfd9234e-3dbf-4d97-8a2c-b8b439605a8a · outbound

This paper cites an unresolved cited work.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:13.582959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:12.501855Z digest=sha256:9432bac5d0944528235154845c0714fc0656804c1fddd95abb3eccfabdcd8912

Observation 5871d1eb-3091-4db2-ba61-b4acaf0125b4 · outbound

This paper cites Table 7.Evaluation on a comprehensive set of general multimodal benchmarks.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Table 7.Evaluation on a comprehensive set of general multimodal benchmarks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:13.371875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:12.577103Z digest=sha256:d2f7b3614c8e3363c3ed8c1c28353322791a3bf0f1dd001364451db3965e31f0

Observation 4251acff-7482-4f44-af19-847fd10a068c · outbound

This paper cites Cross-Self KV Cache Pruning for Efficient Vision-Language Inference.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Cross-Self KV Cache Pruning for Efficient Vision-Language Inference

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.783480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.783480Z digest=sha256:58e3fca3cddc00edac716153929ed8fec53139f5e7118a4fb988e44e1e253d33

Observation c1cb8a37-0ca7-474e-885f-539cee6ae2e8 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.612694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.612694Z digest=sha256:11faf21093f1cf42d4563cb542c2a1c760c535c52148b8d2ae50ac9b844a680d

Observation 4daf7dd9-aab9-4b1e-9937-3369bc9fb644 · outbound

This paper cites J., and Yan, Y.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM J., and Yan, Y

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.852764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.852764Z digest=sha256:a9f4e3db8893617bfc66c390082354a8165cad7745807b3431c40e0d25e19819

Observation 75c46191-e6b4-40e8-a942-e0c02d96293f · outbound

This paper cites InternLM2 Technical Report.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM InternLM2 Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.553356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.553356Z digest=sha256:bdd42ef01ee85aa7f1fb930ae49ef0a3c3e1311b450a23a8ff5140481770bfe3

Observation 8824895c-1718-43fe-88e5-984ba981523e · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM NVLM: Open Frontier-Class Multimodal LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.736485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.736485Z digest=sha256:402c8b48466d91df72d6698cf830ffb364caa474ba771889b674653228787ddc

Observation 519b9643-f2d2-4432-b1a8-638dcc103ef5 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.643408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.643408Z digest=sha256:4ec27ee3389d88fc6beb5ffcbdf4ea294ca4010510fcebe3fe62bd32b6ad9f70

Observation d6970a0a-68eb-4b8b-8841-49d7ff231925 · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.527741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.527741Z digest=sha256:b66d6b4d43377be682547f56d1ce5576dfe21174bfd03f2fb8021e3fca1c268f

Pith citing papers

Observation f2b7f28a-61ff-4f35-b0ed-ca024f3c5c1b · inbound

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning cites this paper.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:34:40.992784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:33:36.846345Z digest=sha256:34ff7e2fef8b3b6cf2cc41854f50d76694cbcf99fff7d2670ff808459a8a65c3