Pith. sign in

REVIEW 4 cited by

White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.13110 v4 pith:HIFXUNEY submitted 2023-11-22 cs.LG cs.CLcs.CV

classification cs.LGcs.CLcs.CV
keywords architecturesrepresentationcompressioncratedeepratewhite-boxcompress
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we contend that a natural objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a low-dimensional Gaussian mixture supported on incoherent subspaces. The goodness of such a representation can be evaluated by a principled measure, called sparse rate reduction, that simultaneously maximizes the intrinsic information gain and extrinsic sparsity of the learned representation. From this perspective, popular deep network architectures, including transformers, can be viewed as realizing iterative schemes to optimize this measure. Particularly, we derive a transformer block from alternating optimization on parts of this objective: the multi-head self-attention operator compresses the representation by implementing an approximate gradient descent step on the coding rate of the features, and the subsequent multi-layer perceptron sparsifies the features. This leads to a family of white-box transformer-like deep network architectures, named CRATE, which are mathematically fully interpretable. We show, by way of a novel connection between denoising and compression, that the inverse to the aforementioned compressive encoding can be realized by the same class of CRATE architectures. Thus, the so-derived white-box architectures are universal to both encoders and decoders. Experiments show that these networks, despite their simplicity, indeed learn to compress and sparsify representations of large-scale real-world image and text datasets, and achieve performance very close to highly engineered transformer-based models: ViT, MAE, DINO, BERT, and GPT2. We believe the proposed computational framework demonstrates great potential in bridging the gap between theory and practice of deep learning, from a unified perspective of data compression. Code is available at: https://ma-lab-berkeley.github.io/CRATE .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Intermediate transformer states sit far from the output axis on purpose: that position insulates attention's cross-token mixing, and the frame can be prescribed in advance without loss.

  2. Towards White-Box Deep Wireless Sensing

    cs.LG 2025-07 conditional novelty 6.0 of 10

    RF-CRATE derives a fully complex-valued white-box transformer for RF sensing from the sparse rate reduction principle and shows it matches black-box baselines across five datasets.

  3. Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A variational reformulation of the MCR2 objective yields a linear-complexity attention operator, ToST, that matches transformer performance without computing pairwise token similarities.

  4. ESS-ReduNet: Enhancing Subspace Separability of ReduNet via Dynamic Expansion with Bayesian Inference

    cs.LG 2024-11 conditional novelty 4.0 of 10

    ESS-ReduNet speeds ReduNet training by dynamically boosting the expansion operator and correcting membership estimates with label-derived Bayesian posteriors, reporting more than 10x faster convergence on several datasets.

Pith tools