Pith. sign in

Paper Citation Record · LEDGER

TokenPacker: Efficient Visual Projector for Multimodal LLM

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2407.02392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.02392 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:55:03.126780Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:15:44.617819Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5ecbc9d4-5d8e-43c3-a3a6-0753f84636a3 · inbound

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step cites this paper.

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:35:25.953654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T11:35:25.894465Z digest=sha256:c51ae6b99af506f9696642fe3f46ec013c468425922bb7b04a3f90cec83e5ee3

Observation b9df33b0-369c-4531-afcb-3e94984d212a · inbound

FoPru: Focal Pruning for Efficient Large Vision-Language Models cites this paper.

FoPru: Focal Pruning for Efficient Large Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:31:59.552458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:31:59.552458Z digest=sha256:4ddda2cf67cc9e14dc9bb0c87351cdaf7b66199811180193374b17cb0d16daf3

Observation f0a986b0-1320-418f-be09-e018279d5e4d · inbound

freePruner: A Training-free Approach for Large Multimodal Model Acceleration cites this paper.

freePruner: A Training-free Approach for Large Multimodal Model Acceleration TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:07.386710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:07.386710Z digest=sha256:faa4e9d30489488cde4f12aafc423f08f899808d0c4cc18ac6c1372abdb3c9f5

Observation 0f87dbc8-cfd3-48d1-934e-08ca0e0c21ee · inbound

Importance-Based Token Merging for Efficient Image and Video Generation cites this paper.

Importance-Based Token Merging for Efficient Image and Video Generation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:43.479995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:43.479995Z digest=sha256:8c0a1ee6f3f9b0d95775f7fac52d832593e4a32870d650799fcad6f39e8f91fd

Observation d5a7df4a-714f-4097-893a-1e71631a0529 · inbound

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings cites this paper.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.920297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.920297Z digest=sha256:d7a653774414e1807ccb879daacc9394dff54833cfad6b1cff1a5863995438e2

Observation eb1ee894-a87f-497c-acc4-0e3fbf1ef691 · inbound

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction cites this paper.

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:21:10.833574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:21:10.833574Z digest=sha256:7708bbe5bfadb92b0c20313124afd79eae3ed53475925264caa80835c15b7768

Observation 715ec7c7-15a8-4847-9b3e-4a3cdc18d2b3 · inbound

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs cites this paper.

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T00:57:32.455630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:57:32.455630Z digest=sha256:0e7de147f4569d6d79dbc224b70e000153155090aecda0d3fc3236347c099dab

Observation 64483dff-5527-4c2b-ba48-8dfb8a2a7b10 · inbound

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression cites this paper.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.306619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.306619Z digest=sha256:917235e9d33062453e03b085577a134da5a43274752a8a27a0f20d16a4518eca

Observation a0873528-fedb-49f4-88f7-8c9689af5dca · inbound

DocVLM: Make Your VLM an Efficient Reader cites this paper.

DocVLM: Make Your VLM an Efficient Reader TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.772001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.772001Z digest=sha256:175a6ba43275c21e70f3f6abb6d2eb84b2787da8aa90c610967bfb54fb7d8536

Observation 59e7046c-d4fb-4a3a-b6fb-0d8f77c9a4cd · inbound

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer cites this paper.

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:13.286584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:13.286584Z digest=sha256:0b622368a23185929cd03ee2d884bbabf3a214f3f150c38c44603984e9e6ac14

Observation fb05bf86-a2c6-460a-8300-fb344eaa035b · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.651832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.651832Z digest=sha256:af520ef966e53032ad8cc802dcb5ac5997a18a2877b1a7cc7deb35a4677ade47

Observation 3d591470-788b-498c-9f25-6218ac1d9be2 · inbound

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming cites this paper.

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:39:18.105337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:39:18.105337Z digest=sha256:ee38f8d82f93becdff67a47c8dc648fc57bfdc35d18215663ff957511270cb02

Observation 726c841c-55fc-46e1-8a32-8eb056635bd1 · inbound

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM cites this paper.

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:30.221093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:30.221093Z digest=sha256:d3edef61ad4733bd7f7f39499c2e8f65dc008f8a7c174cd07b3cdc1f6c962835

Observation 71909ff6-8976-435a-8fde-f3e3fed2aac0 · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.769861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.769861Z digest=sha256:d869eb5d444ac05589b395e2d112151b8e50f34ddfd55ef47a840349db1d1db0

Observation 996775d2-0132-4d8e-9dc0-582210f9d4f1 · inbound

DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs cites this paper.

DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:55:03.126780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:55:03.126780Z digest=sha256:abf1ef92200e33769ca718fda626a9487a44d142e4df5bc72a6bb2bdb348b936

Observation 8b7a0790-66f6-455b-a4f4-4dedf1e16902 · inbound

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning cites this paper.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.856988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.856988Z digest=sha256:c83066ef7d7975520cb68c3783fb1c1619a87d0d0868986b1fad66da8021fa66

Observation b57cb93c-d8e1-4482-b0b8-4d92a2d3f68b · inbound

STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference cites this paper.

STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:41.755234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:41.755234Z digest=sha256:0da1180368bc4577fa6ac7553f901d773fc11236824e29f842d155c14435aa8a

Observation 2ad7117d-d750-496d-bf95-cb4432bb11c7 · inbound

PixelThink: Towards Efficient Chain-of-Pixel Reasoning cites this paper.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.521580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.521580Z digest=sha256:9975a395ffab56c37f86a2c05cb787c24c3eb1cca185b0aeeec7aa8905e8cf09

Observation 9f5a24b9-6c20-4983-ad8f-d9e2bcdde671 · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:04.930789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:04.930789Z digest=sha256:395aca4ab63a016d457863da3664fbb9cc5e462aaea45399bd935e52934cf1ed

Observation 81818563-8738-439b-a096-f401ce88c5b9 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.333184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.333184Z digest=sha256:83226f68f9fbf1ec4f9ab72d73801e6156650b569029a73eded5eec4fb31bde8

Observation f1413e60-66b2-4474-8cc5-4e90d4930ee2 · inbound

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings cites this paper.

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:48.866029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:48.866029Z digest=sha256:55ac3e3cce88a3f5578bfce7afa228da5831ba11d5854304578055e303b1513a

Observation e8315937-1c4c-4eed-a4d5-020faa7412bf · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:09.538863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:09.538863Z digest=sha256:9fdc235de5f26b749255893b6e95c4352923968f72dc8d68ce19d32d5ee2caf0

Observation 4adaa931-5207-495e-bbd1-7f0ca55cb3db · inbound

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding cites this paper.

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:07.697677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:47:07.697677Z digest=sha256:6f9ae7d2a215937cc28d05e6327d46a6f0ddaddb272afd35f479908b098017a3

Observation c62664f3-9f31-4c76-a0de-2a28e1bfcca7 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:46.824918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:46.824918Z digest=sha256:7f65e7f9449c68ff244fb9ec8b44221ebbc0fdddcb372985ad25ef0c8faf04d8

Observation 4294a87b-37ca-4dd7-8684-3793d217ceef · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.416053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.416053Z digest=sha256:2441b4689911afb586bc874d3f7f38aea21bc58fd7d735058de6b59a19de037d

Observation ede358af-656a-4c9c-a204-21158acab1c5 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:25.023913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:25.023913Z digest=sha256:288f68dbdc0ece44cbf9e973c54f7230e673b50678403f2027dd9ce482388c45

Observation c934db64-3e50-42a6-b795-2596f8f744e4 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.543299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.543299Z digest=sha256:ae38069db0192cf2c68ee45943fe3e5a55ae17bccda1d5c9748a6175dd671254

Observation bb2fa175-544f-4b44-8abd-a3127afaa569 · inbound

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval cites this paper.

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:00.242206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:00.242206Z digest=sha256:96f8389af258f75333c30278c49cf76dd21a9f640e27af378c2eac015d36635d

Observation 0c424190-ddff-486c-a6ef-16aaf8b4e4ef · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.298449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.298449Z digest=sha256:25e6479fc8762919d8d4b1df19916ae2c9afcd04ef9e3d06cd7be7e0247426b1

Observation 3243f840-cfe4-4a19-8f2d-3f79e8bfccb5 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.589632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.589632Z digest=sha256:690798d4ca2c48473c38a7a9fbfe077432df30b82e9655ad96ff48df7ae07797

Observation d3e74dd7-8293-4ec6-9204-f21aaa904a8f · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:51.265156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:51.265156Z digest=sha256:88459d40f137b4703066adb7fedf1d55754d048ebf7bd4d83f3050dd684f6427

Observation b04a3492-ffc0-4989-92c4-271b99e7bb8a · inbound

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces cites this paper.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.074772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.074772Z digest=sha256:5af9dfd7d701d745f8853cdf9ba6aeae9dcb544a4d0d7575e22bc2b8b76fecad

Observation 1a3c9763-f343-4d6b-95d7-997527b62566 · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:15.952781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:15.952781Z digest=sha256:7d9d4082b02d2395caf677d4c4df9f51910febe355bc9f93e1e4fedb9473336a

Observation 9edf9c59-0678-425a-9243-243392fd084d · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.675460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.675460Z digest=sha256:bb3258829c7922eda68d1f6e1792e8874c34d3477afaf58bcca0c895139cdcac

Observation 4b131883-97a4-4f14-97af-e2abeb8c3d79 · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:54.422800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:54.422800Z digest=sha256:19e703c27e993a1316eb380f47823cd9ba4679816d6a2d07a9735c7b955a3ccb

Observation 8d937da5-6184-48f7-91e4-6eaf389f73b2 · inbound

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation cites this paper.

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:25:33.013951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T00:22:41.611893Z digest=sha256:770e28f061a95900611112c1392af4ef3560ac2527f674ac4992f958e822abcb

Observation 730fc652-2479-4f49-9878-86f56d5b5466 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.210581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:a8af1e802bbf3f8ca98bf24459f27a4854b8f0487e1373f74b2a8cd936249adf

Observation de47aec6-9bef-40d4-9786-bfe0e2b88f70 · inbound

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding cites this paper.

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.864498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T08:13:42.526597Z digest=sha256:6c68db65ff2487c7eb4f01e53cd005c124f8a3fa2c835e87d2f374be8b7494b3

Observation 2500fc84-b94d-418c-a669-5923dcef4800 · inbound

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs cites this paper.

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.619104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:41:04.184461Z digest=sha256:15170ba424472cefbd21436a19ed9911a736327f46dd7e10ffccb7408ee6ff80

Observation 8f953d7f-c32b-4170-910c-9dc353963321 · inbound

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning cites this paper.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.532712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.532712Z digest=sha256:c554946b5d5e936ccfc9c3fa99be8b8ee55a2e6da2dd437cf18a33a199797b70

Observation 029a34b8-1856-4299-8a32-a5378955ea5a · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.568021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.568021Z digest=sha256:749f867fba060c7658acedfe23289340fcdb5267838a576c9b5f4a5605e8698b

Observation a842a616-3f4c-4e81-9abb-82e1759158ba · inbound

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware cites this paper.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.805044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.805044Z digest=sha256:ec3ecd3ea1d8dc66fd994c77b8f7cc92cada74857a0eac5ed35b388d0cf10464

Observation 3c74d0f9-56b4-442e-91c8-51e2e8a3ec73 · inbound

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin cites this paper.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.255625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.255625Z digest=sha256:68753a8f0032fdbf842878f21af47dcd2f5ef1f725fdd9d704c22b25260d35b2