Pith. sign in

Paper Citation Record · LEDGER

Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2410.14072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.14072 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:15:26.802931Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T22:35:40.607778Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e4cab989-0a6d-4dfa-9ae4-4b98aa9cc0e1 · inbound

NanoVLMs: How small can we go and still make coherent Vision Language Models? cites this paper.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.167918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.167918Z digest=sha256:e1a710efe12a6b65a7d4176d2c8473ed33e12bd59f37154a8b8a4578c3f83588

Observation 3ab00f2c-d575-4688-9b0d-11d75f682727 · inbound

MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention cites this paper.

MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:26.802931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:15:26.802931Z digest=sha256:506d89d70bdeb11a744ec40f3b3413bae3578475e8d34871e5f0a56ef714a36e

Observation cc86dd3a-0cd2-4c62-9ca1-94c6bc54c6a6 · inbound

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning cites this paper.

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:39:06.590851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:39:06.590851Z digest=sha256:23b3bd42074a346ad2dd8b8aad7fa65b297ca2418cf3b3d9dd26ce92b440e247

Observation 00402166-99a3-4c38-b206-feb971afe9d9 · inbound

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning cites this paper.

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:58.959189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:58.959189Z digest=sha256:1627210b4099a42683062fdd818b88cc765202b6b4049f5e5f475981857b1db5

Observation b869b41c-3b6d-4852-a8d3-bb4d66106308 · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:06:20.575068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:e14b992ff17a9919874ed3f4944614772a5dbc95280784880538ada0e1838bc5

Observation ad6581d7-6d82-468c-bc17-d4b3f4c91d32 · inbound

Where and How to Prune: An Empirical Study of Visual Token Pruning for GUI Agent Navigation cites this paper.

Where and How to Prune: An Empirical Study of Visual Token Pruning for GUI Agent Navigation Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:03:20.148788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T23:59:49.017251Z digest=sha256:90d7e9d6038d9e9c15c4562b51d605db0b5fef2b8db691d013eab6dcac396dad

Observation 963c0819-29ba-4857-a82c-9b74e7c00c16 · inbound

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens cites this paper.

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:03:36.946492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T18:00:18.315737Z digest=sha256:3e7171a59ddb1f180e7e051f405c276814ea248964a40a11dac9c4131f15de22

Observation 845780f1-5d67-4024-8896-568df97ed835 · inbound

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring cites this paper.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:35:40.609039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:06f3de4d66fc04583e2bdf5bd3ba9ef62fbfed1a8d08b0be04f8ea4b5b5b8a54

Observation 73dadc04-5cae-49e7-adf0-cd500339a263 · inbound

TDSal: Task-Based Top-Down Saliency Prediction Model cites this paper.

TDSal: Task-Based Top-Down Saliency Prediction Model Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T15:15:45.858365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:15:45.858365Z digest=sha256:20aae7f0b14bf55472a82b81da0e5cc18dff929fcfb8c61b75477f4b5c9ba525

Observation 6fc7e5d8-6bd2-47c6-9e9b-6d743ae275d5 · inbound

RegisterBridgeMM: A Register-Centric Framework for RGB-Infrared Object Detection cites this paper.

RegisterBridgeMM: A Register-Centric Framework for RGB-Infrared Object Detection Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:59.237877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:59.237877Z digest=sha256:99e20604d156a3be96c7e5ac32e9dcf85d54fc040ff97577f6625f888ae6601f

Observation f7e6a82c-d1ad-4b30-9c0c-272f3de478b9 · inbound

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin cites this paper.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.299362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.299362Z digest=sha256:509192301e3fcfd1ca466efd8b5cd2cb7df8ab996f21343790ab016e6693a0a6