Pith. sign in

Paper Citation Record · LEDGER

TokenPacker: Efficient Visual Projector for Multimodal LLM

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2407.02392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.02392 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:47.521580Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:15:44.617819Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5ecbc9d4-5d8e-43c3-a3a6-0753f84636a3 · inbound

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step cites this paper.

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:35:25.953654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T11:35:25.894465Z digest=sha256:7862f988e48e6e8f07137899ddda58f1b2975aa23098d0043d7b588ac3755081

Observation 2ad7117d-d750-496d-bf95-cb4432bb11c7 · inbound

PixelThink: Towards Efficient Chain-of-Pixel Reasoning cites this paper.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.521580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.521580Z digest=sha256:a3d402d6192696e465a588f22260565a5b821d730f4547868cbc618f4a506c14

Observation 9f5a24b9-6c20-4983-ad8f-d9e2bcdde671 · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:04.930789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:04.930789Z digest=sha256:6a7bdf9121a7ddd67dcf01630f9decf44d096e07c3618431d2b250d85d77b099

Observation 81818563-8738-439b-a096-f401ce88c5b9 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.333184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.333184Z digest=sha256:0a5d684ceb88c8067c2bb13056b492d27fe4de5c2b3f4285983c37428c8fe716

Observation f1413e60-66b2-4474-8cc5-4e90d4930ee2 · inbound

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings cites this paper.

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:48.866029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:48.866029Z digest=sha256:46218405a5fd73ac2ae8da464708acf69c9e0fc985dd827ccb8dfc0e800ee5ee

Observation e8315937-1c4c-4eed-a4d5-020faa7412bf · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:09.538863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:09.538863Z digest=sha256:6d7fc37ddf79bd3ac865d3c49bfd2e68f918cfad4096d1c36bf7804cf12d075d

Observation 4adaa931-5207-495e-bbd1-7f0ca55cb3db · inbound

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding cites this paper.

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:07.697677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:47:07.697677Z digest=sha256:36a0776f71187efcd95872270ed6fbf82e867670462b4dcc9ff50c3b057de29a

Observation c62664f3-9f31-4c76-a0de-2a28e1bfcca7 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:46.824918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:46.824918Z digest=sha256:7dc6b9d10f5f74c8afe26c49f8440d4393b9a6045dfcde8db84c738a5594d9d6

Observation 4294a87b-37ca-4dd7-8684-3793d217ceef · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.416053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.416053Z digest=sha256:ce2d387b58ca1f8bbdd3543cdaacb9ba95f53211c75f4418070b8b4529cac7be

Observation ede358af-656a-4c9c-a204-21158acab1c5 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:25.023913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:25.023913Z digest=sha256:93b843588cb3c50c80f403eaf2bab687ba81864bdb78bf3675a3fa3684379ce9

Observation c934db64-3e50-42a6-b795-2596f8f744e4 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.543299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.543299Z digest=sha256:67e75680a22feece7025c808dd724b0584de598a487b0a139f6f45a719559ecb

Observation bb2fa175-544f-4b44-8abd-a3127afaa569 · inbound

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval cites this paper.

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:00.242206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:00.242206Z digest=sha256:67c8269256f89b17565c9060a6e529472069731ec84760542a1ae5d554d42920

Observation 0c424190-ddff-486c-a6ef-16aaf8b4e4ef · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.298449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.298449Z digest=sha256:4b12491123ffcdedec6ae79361500bae520fc5274dee593e0ba515b38d7e4b2a

Observation 3243f840-cfe4-4a19-8f2d-3f79e8bfccb5 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.589632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.589632Z digest=sha256:ea95fb6fe080dcf812606ecdf07053613a06bf961460a76bd1acde580d3968dd

Observation d3e74dd7-8293-4ec6-9204-f21aaa904a8f · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:51.265156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:51.265156Z digest=sha256:951acd422ee26e82aa71ae81ba2bf2f35700847c1f1823e9fcd28596067c4027

Observation b04a3492-ffc0-4989-92c4-271b99e7bb8a · inbound

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces cites this paper.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.074772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.074772Z digest=sha256:d511fbd0b4d920dfdac6b08ff9c63beb937fc99773e0bf682ec950a294de22f8

Observation 1a3c9763-f343-4d6b-95d7-997527b62566 · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:15.952781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:15.952781Z digest=sha256:330b512dddf2303d479d6ee2a63cb8c9fb774868a80dff6ea4f14d7b04557e15

Observation 9edf9c59-0678-425a-9243-243392fd084d · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.675460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.675460Z digest=sha256:03f3f6f57ffb9552530eb3f99df9d556557abc38a43faee7bf39dbe0843fdb5d

Observation 4b131883-97a4-4f14-97af-e2abeb8c3d79 · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:54.422800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:54.422800Z digest=sha256:29177fa094cfe02e0d65bf035c89dd5d7c51ed9c8b4f996fbcef609d4929efbe

Observation 8d937da5-6184-48f7-91e4-6eaf389f73b2 · inbound

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation cites this paper.

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:25:33.013951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T00:22:41.611893Z digest=sha256:34fbefdf60a6118aaf42c49cd29fa569fe5b0bd359495f90d538fbe5ab21caff

Observation 730fc652-2479-4f49-9878-86f56d5b5466 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.210581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:444875937c149b206a7aad04fd6797223bd31f7edbca3de9d8b6e9ba7662c305

Observation de47aec6-9bef-40d4-9786-bfe0e2b88f70 · inbound

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding cites this paper.

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.864498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:13:42.526597Z digest=sha256:9a7334c53d0cde3a0b1351777b68cb63fa0e9285366b6c7db96d6a7cb2be07f7

Observation 2500fc84-b94d-418c-a669-5923dcef4800 · inbound

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs cites this paper.

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.619104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:41:04.184461Z digest=sha256:cfcbef2a6a9bc17709ceac26328a503fbafbcd4ee7897328e3e6b38d6ac2b444

Observation 8f953d7f-c32b-4170-910c-9dc353963321 · inbound

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning cites this paper.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.532712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.532712Z digest=sha256:d8ce17af8c9714e1623cfe4b66ed6a69f0815b215182fbd373defb44054fc888

Observation 029a34b8-1856-4299-8a32-a5378955ea5a · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.568021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.568021Z digest=sha256:318aadc9bd8432f9a4f259a0fc237ef0d32d6cd958298829d64b29280a89ee12

Observation a842a616-3f4c-4e81-9abb-82e1759158ba · inbound

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware cites this paper.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.805044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.805044Z digest=sha256:75b62e163fa5b6c25b051a48afbedf14f6c747ad67aac8af7a8af6347a2a80a7