Pith. sign in

Paper Citation Record · LEDGER

Token Sequence Compression for Efficient Multimodal Computing

As of 19 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2504.17892.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17892 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:34:16.703536Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:31:44.517145Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T21:31:45.066251Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a78ded4-8a92-4aaa-a465-8944dcc5a272 · outbound

This paper cites A multimodal architecture for ai agents, 2023.

Token Sequence Compression for Efficient Multimodal Computing A multimodal architecture for ai agents, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:17.061184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.602109Z digest=sha256:6833ad6ebb0e8c5f6d16b66e535038c01e5dc28ca06be65ca2a4c2475d546ea0

Observation a5d2f209-181e-496d-8490-615ba2603449 · outbound

This paper cites Token merging: Your vit but faster.

Token Sequence Compression for Efficient Multimodal Computing Token merging: Your vit but faster

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:17.047608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.608026Z digest=sha256:9b79df47cd5c6dd8b0301718b4e6adf486547317acc60b75acea23b382a7ce83

Observation 693a7eb5-bfbe-4cb5-81e3-e487fa364f70 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Token Sequence Compression for Efficient Multimodal Computing An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:17.033981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.612332Z digest=sha256:f1f7e346ac80432f8f4ee97bd09cf39aa4203cc7b246272b37e1a3ee1fd528cb

Observation c43e1943-2d31-402e-8503-8685a083bf33 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Token Sequence Compression for Efficient Multimodal Computing MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:34:16.616677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:34:16.616677Z digest=sha256:2c0b90ebdfac873390e0a9a1f95e0b1ce98fa0cf564f763f0dbadbffb151799b

Observation ea89dfbb-b98e-4bb9-913a-73c111eb7851 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Token Sequence Compression for Efficient Multimodal Computing Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:17.020084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.621133Z digest=sha256:5b5175d9157e6c510b58adb40752b4efc6ba9f6b1d7b4f259f4f4f52a4698c14

Observation c0443a45-2164-455c-b96d-6d119e3cdb53 · outbound

This paper cites Perceiver: General Perception with Iterative Attention.

Token Sequence Compression for Efficient Multimodal Computing Perceiver: General Perception with Iterative Attention

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:34:16.625879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:34:16.625879Z digest=sha256:1eb33ddc01f83cd62a7f914be05fd031e7a2ad90cf35304756a0a48a9966b251

Observation 3092b459-a3eb-40c7-8e2b-4c055d4cefc8 · outbound

This paper cites Towards efficient visual-language alignment of the q-former for visual reason- ing tasks.

Token Sequence Compression for Efficient Multimodal Computing Towards efficient visual-language alignment of the q-former for visual reason- ing tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:17.006773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.630399Z digest=sha256:70e4d5245fd84f366822360b61d0c703b6967b0ab9680bc866aee9a15333f84a

Observation 769df055-957c-4c70-9a3c-61711c810f86 · outbound

This paper cites Llava-onevision: Easy visual task transfer.

Token Sequence Compression for Efficient Multimodal Computing Llava-onevision: Easy visual task transfer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.993732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.634727Z digest=sha256:afa261b0dd01dba0bf9f5dcc2127039c2ca363ffbc42126856e1d234640ed404

Observation 9cc093b6-24cf-42f2-8872-ac28aa6964cf · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Token Sequence Compression for Efficient Multimodal Computing Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.980827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.638677Z digest=sha256:0c64b55f0fe22d3ebf8c3da6cd4c6bbfdf36d2afaffd82d9bf3142592c7a050f

Observation b6c77503-6edc-415d-b963-a5a1caea6dfb · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Token Sequence Compression for Efficient Multimodal Computing Evaluating Object Hallucination in Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:34:16.642522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:34:16.642522Z digest=sha256:2f3bd3a8ad171fb8a3e77b2e3db15cdc4cb408191385e24c5421d6795f7dc5ad

Observation ef17429c-fcbc-4651-a657-d8d1282c74f3 · outbound

This paper cites Vila: On pre-training for visual language models.

Token Sequence Compression for Efficient Multimodal Computing Vila: On pre-training for visual language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.967529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.646783Z digest=sha256:33dbf6e7de4b62fdef207487e4a760d5824185a3f7b40a4e8395b0920a5ced5e

Observation 3d034a8a-b008-4a08-9316-0f1a5407658d · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Token Sequence Compression for Efficient Multimodal Computing Improved Baselines with Visual Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:34:16.650651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:34:16.650651Z digest=sha256:3ceedad3fdc41a069f96382549c891908b4e49fa25884abb5a961ff7e8cd6e6c

Observation 8b10226c-38a9-49f4-8ad8-231b41d9079d · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge.

Token Sequence Compression for Efficient Multimodal Computing Llava-next: Im- proved reasoning, ocr, and world knowledge

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.954232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.655013Z digest=sha256:34425074481b199d0edd776c1cc9fec8b046ae88970451508a1ba985990dc149

Observation 50708892-7149-437b-9553-899a08765bc6 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Token Sequence Compression for Efficient Multimodal Computing MMBench: Is Your Multi-modal Model an All-around Player?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:34:16.659149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:34:16.659149Z digest=sha256:912023bdf0dbf776c82f5f55adb7eb68474e1164a278d0c4db1eba765bb72998

Observation ade9eed7-f91b-4eab-8e3e-677dacce5f1b · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Token Sequence Compression for Efficient Multimodal Computing Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.941028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.663960Z digest=sha256:7338a20791be79bd0ad375a0cce4c9086a8da102fa3c8268642a7970b520f168

Observation 84364a5f-d43c-4bcb-95ec-5e1b3323eb15 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Token Sequence Compression for Efficient Multimodal Computing Learning transferable visual models from natural language supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.927347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.667608Z digest=sha256:0c1750634b58bbd064d6762db627c3220c3ede263e984e7eaafa8d3f78e03b74

Observation 169b392e-d44b-4f91-a208-eb8275377bae · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Token Sequence Compression for Efficient Multimodal Computing Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.914243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.671649Z digest=sha256:f9a7c6e83e49239f6a4e9b0d643a1895b96abbb72f7644d40685db89f5ed5ac1

Observation 2255e0cc-cfa5-4979-9417-87a98d688b15 · outbound

This paper cites Towards vqa models that can read.

Token Sequence Compression for Efficient Multimodal Computing Towards vqa models that can read

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.900023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.675605Z digest=sha256:df08a81b59330d7b623f5368f29f4b59de6195793045f8c2d30607f672f5dc5b

Observation ed5fc4ea-c91a-443b-b807-3a5321940efe · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

Token Sequence Compression for Efficient Multimodal Computing Next-gpt: Any-to-any multimodal llm

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.885111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.679365Z digest=sha256:40907f786afe493fb8bdf344ac883abc703b7ea02eb5a01f451fa06ff6ca8318

Observation 72e1df33-5ca7-4d9e-8236-2a4417ee3bf8 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Token Sequence Compression for Efficient Multimodal Computing Visionzip: Longer is better but not necessary in vision language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.871734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.683267Z digest=sha256:358ddecc986b27a57887cf0023da63af5dfae62a23419729c3f30e6655a7c13f

Observation 581969a4-9e02-4872-bedd-83577b5e6cc4 · outbound

This paper cites X-vila: Cross-modality align- ment for large language model.

Token Sequence Compression for Efficient Multimodal Computing X-vila: Cross-modality align- ment for large language model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.858512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.687215Z digest=sha256:1dc25f554bfa5bba2f49a32d078dd4b06e60635148fcb9ad05c98e3afd9a1c1d

Observation 736433ff-84ff-4c46-b937-5ef4e7766373 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

Token Sequence Compression for Efficient Multimodal Computing Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.844846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.691295Z digest=sha256:3025f7d6325c88f674de1d0f87f002b95ef35eafb32a286db8cce3e5c7449d58

Observation 1976f3d4-e1ff-4750-82be-1a0e20bd777b · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Token Sequence Compression for Efficient Multimodal Computing LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:34:16.695161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:34:16.695161Z digest=sha256:9c10f9311c6c6e159d16c03a79006e6180fccde86550b26cf58a478bccf7a9ba

Observation 980db933-c9f3-4f20-b082-57dd7fdb5880 · outbound

This paper cites Sigmoid loss for language image pre-training.

Token Sequence Compression for Efficient Multimodal Computing Sigmoid loss for language image pre-training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:34:16.831138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T10:34:16.699603Z digest=sha256:787f0ea8edee90d4a9d3eabd0afab940524362ea0ae4163480b574eedd26d650

Observation c34eabbf-2f40-45e8-8169-338c768cda8b · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Token Sequence Compression for Efficient Multimodal Computing SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:34:16.703536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:34:16.703536Z digest=sha256:4e59b0751cb78268f85ee02c552b383a3e9b75bbd5751c9e19484df328c4cd20

Pith citing papers

Observation a809c974-a1d4-4aa0-a0aa-b8b498adb9a9 · inbound

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay cites this paper.

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay Token Sequence Compression for Efficient Multimodal Computing

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:31:45.074721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:31:44.517145Z digest=sha256:f3c24049458bd914e1fa027add1c10b86b959acf1dadb8236937eb84699fb650