Hard-attention Transformers with W parameters and depth L have VC dimension Θ(WL log(TW)); teacher forcing is sample-optimal for chain-of-thought learning.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
representative citing papers
KIVI applies asymmetric 2-bit quantization to KV cache with per-channel keys and per-token values, reducing memory 2.6x and boosting throughput up to 3.47x with near-identical quality on Llama, Falcon, and Mistral.
citing papers explorer
-
Tight Sample Complexity of Transformers
Hard-attention Transformers with W parameters and depth L have VC dimension Θ(WL log(TW)); teacher forcing is sample-optimal for chain-of-thought learning.
-
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
KIVI applies asymmetric 2-bit quantization to KV cache with per-channel keys and per-token values, reducing memory 2.6x and boosting throughput up to 3.47x with near-identical quality on Llama, Falcon, and Mistral.