Pith. sign in

REVIEW 1 cited by

Cross-Axis Transformer with 3D Rotary Positional Embeddings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.07184 v3 pith:77PHNYB4 submitted 2023-11-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords transformersvisioncross-axisimagemodelingrequiredtransformeraccurately
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite lagging behind their modal cousins in many respects, Vision Transformers have provided an interesting opportunity to bridge the gap between sequence modeling and image modeling. Up until now however, vision transformers have largely been held back, due to both computational inefficiency, and lack of proper handling of spatial dimensions. In this paper, we introduce the Cross-Axis Transformer. CAT is a model inspired by both Axial Transformers, and Microsoft's recent Retentive Network, that drastically reduces the required number of floating point operations required to process an image, while simultaneously converging faster and more accurately than the Vision Transformers it replaces.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Survey of Retentive Network

    cs.CL 2025-06 conditional novelty 2.0 of 10

    A review that describes the RetNet architecture and enumerates its applications across many domains, without presenting new experimental results.

Pith tools