Pith. sign in

An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

citation-role summary

baseline 1

citation-polarity summary

fields

cs.CV 3

years

2026 3

roles

baseline 1

polarities

baseline 1

representative citing papers

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

cs.CV · 2026-05-19 · unverdicted · novelty 5.0

SEATS adaptively selects and removes non-text tokens before and inside the LLM layers of omni-modal models, yielding 9.3x FLOPs reduction and 4.8x prefill speedup at 10% token retention while keeping 96.3% performance.

citing papers explorer

Showing 3 of 3 citing papers.