FastVGGT achieves 4x speedup on VGGT for 1000-image inputs using training-free token merging tailored to 3D architectures while reducing error accumulation.
Learn- ing to merge tokens in vision transformers
6 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
REDI combines supervised TF-IDF corpus scores over DINOv3 visual words with attention maps to rank patches, reducing sequence length 46.8% while raising Top-1 accuracy from 83.514% to 84.706% on ImageNet-1K.
ConsisFormer reduces WFM Transformer complexity by over 83% via adaptive token aggregation and feature interpolation while preserving performance on channel tasks.
ST-Merge is a plug-and-play spatio-temporal token merging method that delivers 2x speedup on VLMs and 8.3x on a VLA at high resolution with minimal accuracy loss via 3D coordinate matching and positional correction.
ViCoStream is a new coordinated pipeline framework for streaming VideoLLMs that achieves 134 FPS video throughput and less than 50 ms TTFT on A100 while keeping accuracy near full-history baselines.
SEPatch3D accelerates ViT-based 3D object detectors up to 57% faster than StreamPETR via dynamic patch sizing and cross-granularity enhancement while keeping comparable accuracy on nuScenes and Argoverse 2.
citing papers explorer
-
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
FastVGGT achieves 4x speedup on VGGT for 1000-image inputs using training-free token merging tailored to 3D architectures while reducing error accumulation.
-
REDI: Corpus Aware Patch Ranking for DINOv3 Token Reduction
REDI combines supervised TF-IDF corpus scores over DINOv3 visual words with attention maps to rank patches, reducing sequence length 46.8% while raising Top-1 accuracy from 83.514% to 84.706% on ImageNet-1K.
-
ConsisFormer: Compute-Efficient Transformer for Wireless Foundation Models Based on Channel Consistency
ConsisFormer reduces WFM Transformer complexity by over 83% via adaptive token aggregation and feature interpolation while preserving performance on channel tasks.
-
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs
ST-Merge is a plug-and-play spatio-temporal token merging method that delivers 2x speedup on VLMs and 8.3x on a VLA at high resolution with minimal accuracy loss via 3D coordinate matching and positional correction.
-
ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference
ViCoStream is a new coordinated pipeline framework for streaming VideoLLMs that achieves 134 FPS video throughput and less than 50 ms TTFT on A100 while keeping accuracy near full-history baselines.
-
Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors
SEPatch3D accelerates ViT-based 3D object detectors up to 57% faster than StreamPETR via dynamic patch sizing and cross-granularity enhancement while keeping comparable accuracy on nuScenes and Argoverse 2.