Pith. sign in

REVIEW 3 cited by

Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.02891 v3 pith:QT5EOHB6 submitted 2025-02-26 cs.CV cs.AR

classification cs.CVcs.AR
keywords edgedeploymenttechniquesvitsaccelerationcompressionhardwaremodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In recent years, vision transformers (ViTs) have emerged as powerful and promising techniques for computer vision tasks such as image classification, object detection, and segmentation. Unlike convolutional neural networks (CNNs), which rely on hierarchical feature extraction, ViTs treat images as sequences of patches and leverage self-attention mechanisms. However, their high computational complexity and memory demands pose significant challenges for deployment on resource-constrained edge devices. To address these limitations, extensive research has focused on model compression techniques and hardware-aware acceleration strategies. Nonetheless, a comprehensive review that systematically categorizes these techniques and their trade-offs in accuracy, efficiency, and hardware adaptability for edge deployment remains lacking. This survey bridges this gap by providing a structured analysis of model compression techniques, software tools for inference on edge, and hardware acceleration strategies for ViTs. We discuss their impact on accuracy, efficiency, and hardware adaptability, highlighting key challenges and emerging research directions to advance ViT deployment on edge platforms, including graphics processing units (GPUs), application-specific integrated circuit (ASICs), and field-programmable gate arrays (FPGAs). The goal is to inspire further research with a contemporary guide on optimizing ViTs for efficient deployment on edge devices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions

    cs.HC 2025-05 conditional novelty 6.0 of 10

    EdgeWisePersona is a new synthetic dataset and benchmark for reconstructing structured smart-home user routines from multi-session dialogues, on which large LLMs clearly outperform small on-device models.

  2. Recursive transformers for semiconductor thermo-mechanical reliability

    cs.LG 2026-07 reject novelty 4.0 of 10

    Depth Recursive transformer, which injects the depth index as a state and uses per-step losses, gives the best accuracy-per-FLOP trade-off among three recursive weight-sharing designs on small engineering surrogate be...

  3. Evaluating Deep Learning Models for African Wildlife Image Classification: From DenseNet to Vision Transformers

    cs.CV 2025-07 reject novelty 2.0 of 10

    On a four-species African wildlife dataset, ViT-H/14 reaches 99% accuracy versus 67% for the best CNN (DenseNet-201), but at far greater computational cost.

Pith tools