Pith. sign in

Paper Citation Record · LEDGER

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models

As of 16 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 2 inbound Pith citation observations for arXiv:2605.13178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.13178 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:05:05.400221Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T16:20:19.874172Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact5
  • verified fuzzy5
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 51ea143e-6bff-494f-961b-860080fa0f8e · outbound

This paper cites GPT-4 Technical Report.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models GPT-4 Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T14:15:47.684821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:3f253defa603bff12dfbfb625dc714e67aa0f99a3779c931471be4eacf275be7

Observation 34b6cc85-005d-4c4f-ae20-c644ba93b500 · outbound

This paper cites Qwen Technical Report.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models Qwen Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.671143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:0511fc622f3b8202d0ca40b8c557a7589654bbbbe8e536dacd9081486c3b20bb

Observation 508c962d-33c4-473a-aaea-368a9c79a308 · outbound

This paper cites Qwen3-VL Technical Report.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models Qwen3-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.687077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:93bb86bafd0052fc26f81e2ec278b37a776071279dde33482bb03312688d18c1

Observation 103c0351-9261-4bfa-bc3c-2eae9b239183 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.682740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:bcc66127f59869f1ba59e6911b6c7e12c4d966ee2e5d3757fccc161d0849ac84

Observation 613bfecb-4ba3-4020-9bf0-e870ac82da89 · outbound

This paper cites an unresolved cited work.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-07T15:13:53.992380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:84aaf9980e07529458070c68e22bae52ebcfce50b17897192d063146c1e1fbe9

Observation 88d195e1-dfef-4328-843b-89d66ac1aac2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T14:15:47.678231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:154b04f070db67c43717f058b18921ad8553917c2f7f96e124bd585e32bf124a

Observation 4e26ab1f-f3a4-4cbb-a923-b9356b8d9393 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.680439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:c6a5cf943086af7f3bd94a1e4e2459012fa8f7aa75b07c0c1ed918085ee44a0d

Observation 437c3db6-78de-4789-b8b0-00e141a90be5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.675961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:e91e8c76b9acc0ef6746c46c4b3ca245f641d428dcda05326f087471c2787ff4

Observation 3beae1b5-cac5-4c86-878c-1eb4a8fe52ae · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T14:15:47.673691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:e1d9ef28e548d45444a22a9a88eb7454ed7302653eab09f756a8f6e6a9e84277

Observation d06879aa-db81-4f8c-be4b-a9562f98743a · outbound

This paper cites Dataset We conduct experiments with LiteLVLM on 6 widely used benchmarks, including 3 referring expression segmentation datasets and3referring video object segmentation datasets.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models Dataset We conduct experiments with LiteLVLM on 6 widely used benchmarks, including 3 referring expression segmentation datasets and3referring video object segmentation datasets

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T15:13:53.988727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:9336769fbf5857847361140742b5fd3fbab09248dbca25109094ec2dc2ee28dc

Observation bab79e1b-9888-40af-a4b7-4b6abf133f79 · outbound

This paper cites left/right.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models left/right

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T15:13:53.986798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:35b417f3484635999b4516e8e5f260cd9c499d2b65848204f156e19dac40d237

Observation 7390316f-ad6f-4c4a-a40c-04ab86994878 · outbound

This paper cites As shown in Table 7, LiteLVLM maintains its performance with only a 0.2% drop while pruning 65.9% of the total visual tokens (192 tokens).

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models As shown in Table 7, LiteLVLM maintains its performance with only a 0.2% drop while pruning 65.9% of the total visual tokens (192 tokens)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T15:13:53.983467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:75317b1a4a05777a0f6b4087ca601d90a4b9209e5b311b5cd796010be56b23b1

Observation 25c55b4a-dba2-46c2-8911-135d1022088d · outbound

This paper cites an unresolved cited work.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-07-07T15:13:53.980027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:1955bde1b2801bcb445fa8897e79e00246054e3bcb718873d699e685f3042aab

Observation 736e1e67-25a1-430e-9007-d67df00e89d2 · outbound

This paper cites MetaCLIP builds upon CLIP by scaling up the pretraining data and improving data quality.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models MetaCLIP builds upon CLIP by scaling up the pretraining data and improving data quality

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T15:13:53.990598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:6bf7d97caa80246a2d8e315e853762a03556297b51964cc49d406898e6de7591

Observation bed0136a-5d33-4174-bb79-2e7902bda128 · outbound

This paper cites an unresolved cited work.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models Unresolved cited work

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-07-07T15:13:53.981576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:db49be58373e71a995cc996a8a46b970b8eb44691b88fbd0e92985164c60b61c

Observation 71b7b7ea-b0ff-4968-bff8-2698f41b8574 · outbound

This paper cites Here, we employ LiteLVLM‡, an enhanced version that primarily uses context-aware tokens.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models Here, we employ LiteLVLM‡, an enhanced version that primarily uses context-aware tokens

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T15:13:53.985129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:5c95373dc9a85c4a87eb4c6292215ceebc4bcb6b38c5eab6356ba0edf91510ff

Pith citing papers

Observation 3be9d4dd-d3a8-486b-baca-99b17db2de85 · inbound

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval cites this paper.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T16:32:55.757864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:32:55.757864Z digest=sha256:1d68b63c291ab6060cf45ca0602328cabfd3aaee9e6f95d935e7a17c81581efb

Observation ba733c62-eea7-47ec-ae6c-25dc6e4a0a15 · inbound

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval cites this paper.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:a7d5a476eb3a2bdce5bb07dba05a28bac691dcacd27caaf75d0857d6915dda4e