Pith. sign in

Multi-stage vision token dropping: Towards efficient multimodal large language model

12 Pith papers cite this work. Polarity classification is still indexing.

12 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 11 2025 1

roles

background 1

polarities

background 1

representative citing papers

EarlyTom: Early Token Compression Completes Fast Video Understanding

cs.CV · 2026-05-28 · unverdicted · novelty 6.0

EarlyTom is a training-free early token compression method inside the vision encoder with decoupled spatial selection that reduces TTFT up to 2.65x and FLOPs 61% on LLaVA-OneVision-7B while keeping accuracy comparable to full tokens.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs

cs.CV · 2026-04-13 · unverdicted · novelty 6.0

POINTS-Long is a dual-mode multimodal large language model that uses dynamic visual token scaling to retain 97.7-99.7% accuracy on long-form tasks with 1/40 to 1/10th the tokens and supports streaming via detachable KV-cache.

citing papers explorer

Showing 12 of 12 citing papers.