Pith. sign in

Average-Hard Attention Transformers are Constant-Depth Uniform Threshold Circuits

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Transformers have emerged as a widely used neural network model for various natural language processing tasks. Previous research explored their relationship with constant-depth threshold circuits, making two assumptions: average-hard attention and logarithmic precision for internal computations relative to input length. Merrill et al. (2022) prove that average-hard attention transformers recognize languages that fall within the complexity class TC0, denoting the set of languages that can be recognized by constant-depth polynomial-size threshold circuits. Likewise, Merrill and Sabharwal (2023) show that log-precision transformers recognize languages within the class of uniform TC0. This shows that both transformer models can be simulated by constant-depth threshold circuits, with the latter being more robust due to generating a uniform circuit family. Our paper shows that the first result can be extended to yield uniform circuits as well.

fields

cs.AI 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

A First-Principles Theory of Slow Thinking and Active Perception

cs.AI · 2026-07-09 · conditional · novelty 7.5

Active lifting of data distributions via latent-sequence sampling and max-rate uncertainty reduction formally derives slow-thinking LLMs and places them on representation and sampler hierarchies that can be climbed.

citing papers explorer

Showing 1 of 1 citing paper.

  • A First-Principles Theory of Slow Thinking and Active Perception cs.AI · 2026-07-09 · conditional · none · ref 148 · internal anchor

    Active lifting of data distributions via latent-sequence sampling and max-rate uncertainty reduction formally derives slow-thinking LLMs and places them on representation and sampler hierarchies that can be climbed.