Pith. sign in
Pith Number

pith:QUSHR6WH

pith:2026:QUSHR6WH544A2PINBG53GSRP6Q
not attested not anchored not stored refs resolved

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models

Sangin Lee, Yukyung Choi

Reversing CLIP visual-text similarity retains the tokens needed for accurate pixel grounding without training.

arxiv:2605.13178 v1 · 2026-05-13 · cs.CV · cs.AI

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{QUSHR6WH544A2PINBG53GSRP6Q}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

LiteLVLM significantly outperforms existing methods by over 5% across diverse token budgets. Without any training or fine-tuning, LiteLVLM maintains 90% of the original performance with a 22% speedup and a 2.3x memory reduction.

C2weakest assumption

The observation that referent-region visual tokens exhibit low similarity to text in CLIP analysis generalizes directly to the large vision-language models used for pixel grounding, and that reversing the similarity ranking will reliably retain the necessary tokens across inputs and models.

C3one line summary

LiteLVLM prunes visual tokens for pixel grounding by reversing CLIP visual-text similarity to retain referent region tokens, outperforming prior methods by over 5% with 22% speedup and 2.3x memory reduction without any training.

References

15 extracted · 15 resolved · 8 Pith anchors

[1] GPT-4 Technical Report · arXiv:2303.08774
[2] Qwen Technical Report · arXiv:2309.16609
[3] Qwen3-VL Technical Report · arXiv:2511.21631
[4] VideoPoet: A Large Language Model for Zero-Shot Video Generation · arXiv:2312.14125
[5] Liu, H., Li, C., Li, Y ., and Lee, Y . J. Improved baselines with visual instruction tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024a. Liu, H., Li, C. 2024
Receipt and verification
First computed 2026-05-18T03:08:56.450163Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

852478fac7ef380d3d0d09bbb34a2ff4131e3b3a4dee98b0e4e64acf152cb334

Aliases

arxiv: 2605.13178 · arxiv_version: 2605.13178v1 · doi: 10.48550/arxiv.2605.13178 · pith_short_12: QUSHR6WH544A · pith_short_16: QUSHR6WH544A2PIN · pith_short_8: QUSHR6WH
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/QUSHR6WH544A2PINBG53GSRP6Q \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 852478fac7ef380d3d0d09bbb34a2ff4131e3b3a4dee98b0e4e64acf152cb334
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "9c82ac07718480edb4405202e8eea1d3e1767846b96faaf68464ebc7e84c35a9",
    "cross_cats_sorted": [
      "cs.AI"
    ],
    "license": "http://creativecommons.org/licenses/by/4.0/",
    "primary_cat": "cs.CV",
    "submitted_at": "2026-05-13T08:40:40Z",
    "title_canon_sha256": "3723d4e9b36ff76cf8d891b08b3348d7cf26c77e5522988302b02f7ec78f7e66"
  },
  "schema_version": "1.0",
  "source": {
    "id": "2605.13178",
    "kind": "arxiv",
    "version": 1
  }
}