Pith. sign in

Paper Citation Record · LEDGER

InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2503.21307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21307 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:18:06.939468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:09:53.831006Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e08e65ee-056b-403a-9f2a-8b6b957b00c7 · inbound

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO cites this paper.

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:33.396784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:33.396784Z digest=sha256:dc9c162e19073aa307da096a47c23da30c36cc4b23f1b9b26171597b8a2254e4

Observation a11f7e75-7d53-4f32-8fc8-0fb0710c46c6 · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.681525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.681525Z digest=sha256:8c916a6721453ec9fd2ba58681670d6916fd4ea63f3f1ccaffc6af4f2ec12348

Observation 48ae9154-6455-4740-9750-a32c2ad19818 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.064975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.064975Z digest=sha256:a94aa665a7ada6353c07a7bb93a5863518ba723547f9a879473042899d2004f2

Observation 65c1534e-1b89-400c-b96a-1bbc0500d2d5 · inbound

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? cites this paper.

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T05:06:45.171417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:06:45.171417Z digest=sha256:ff96635b85fb079d6df3cb5a195d9570d1cea46e343a7d40d6603975c9d2c4ca

Observation b8713d11-63f4-4866-8f3a-31267e02be27 · inbound

An Efficient Token Compression Framework for Visual Object Tracking cites this paper.

An Efficient Token Compression Framework for Visual Object Tracking InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:35.749787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:27:18.535624Z digest=sha256:f1de215954df230f04ae18162d929bc138707fb072d3a625ee3fce1e9b7d9716

Observation 9caa1331-a672-4e77-995d-2ea219c237d4 · inbound

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? cites this paper.

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:16.340053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T02:06:45.858231Z digest=sha256:109b66519bf34d781deb5c22371799750c22a3854019618cf95fc31e03368aa2

Observation 0e34dc88-bf54-4c21-bf7f-74d5a6de833e · inbound

CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models cites this paper.

CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.442979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T12:12:51.867760Z digest=sha256:fdf441440ec4c441b8e39535970aa9db1f0714fb0fc8db81d22984b427ead808

Observation 8bd093a2-dbd9-4379-a643-7dbba2c8a816 · inbound

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference cites this paper.

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:09:53.832588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T04:24:38.917137Z digest=sha256:5c2a4b0ea29c81a51421dab6ac96fa5b17f281bc989c93d69ed9be9b5061f975

Observation 9ba06075-77b6-45c8-97a1-15ca8e42cf37 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.576340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.576340Z digest=sha256:2472a825cf3ece14d645553819b81f1396119861df29c3caa3719629798484f8

Observation 28c532bd-5bee-4187-a525-2142fc9a8462 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T04:28:23.606938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:28:23.606938Z digest=sha256:9d388a8ed2090fdb9f9cc5c58787977c04b8c7bf684e37a8028a983ca49ba1a7

Observation 6921e4da-165f-4615-a4bd-f248cc995554 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:40.925614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:30:40.925614Z digest=sha256:0298009d6188ed8b114940f5f8222cc2544ceae50305a13b9631a5e17fc12948

Observation a8788006-6be9-4c2b-a03c-b18ed1b5c7ef · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:06.939468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:06.939468Z digest=sha256:907fcb0d66652734cc48b2845f07755cc3389fcb34cf4fee92a04b3b43dae5fb

Observation 3e278458-c443-4484-aac3-85b23c675074 · inbound

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin cites this paper.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.475211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.475211Z digest=sha256:7689a34dbba4378b631a3fe90e40163968ace8f7c90e36c028f3aa6ce989f990