Pith. sign in

arXiv preprint arXiv:2402.09268 , year=

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

years

2026 6 2025 1

representative citing papers

Higher-Order Token Interactions via Quantum Attention

quant-ph · 2026-06-10 · unverdicted · novelty 7.0

QHA represents order-k token interactions in O(log k) quantum circuit depth, with an expressivity separation from classical self-attention and empirical gains on high-order parity and application tasks at reduced parameter count.

Transformer Approximations from ReLUs

cs.LG · 2026-04-27 · unverdicted · novelty 7.0

A recipe translates ReLU approximations to softmax attention with target-specific economic bounds for multiplication, reciprocal computation, and min/max primitives.

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory

cs.LG · 2026-03-27 · unverdicted · novelty 7.0

Muon achieves higher storage capacity than SGD and matches Newton's method in one-step recovery rates for associative memory under power-law distributions, while saturating at larger critical batch sizes and showing faster initial multi-step dynamics.

Scaling Latent Reasoning via Looped Language Models

cs.CL · 2025-10-29 · conditional · novelty 5.0

A 1.4B and a 2.6B looped (weight-tied, recurrent-depth) language model trained on 7.7T tokens match or exceed several 4B–8B transformer baselines on selected reasoning benchmarks.

There Will Be a Scientific Theory of Deep Learning

stat.ML · 2026-04-23 · unverdicted · novelty 2.0

A mechanics of the learning process is emerging in deep learning theory, characterized by dynamics, coarse statistics, and falsifiable predictions across idealized settings, limits, laws, hyperparameters, and universal behaviors.

citing papers explorer

Showing 7 of 7 citing papers.

  • Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers cs.LG · 2026-05-21 · accept · none · ref 34

    Graph tokenizations for Transformers induce distinct depth regimes with proven separations and impossibility results for converting between them at limited depth.

  • Higher-Order Token Interactions via Quantum Attention quant-ph · 2026-06-10 · unverdicted · none · ref 7

    QHA represents order-k token interactions in O(log k) quantum circuit depth, with an expressivity separation from classical self-attention and empirical gains on high-order parity and application tasks at reduced parameter count.

  • Transformer Approximations from ReLUs cs.LG · 2026-04-27 · unverdicted · none · ref 4

    A recipe translates ReLU approximations to softmax attention with target-specific economic bounds for multiplication, reciprocal computation, and min/max primitives.

  • How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers cs.LG · 2026-04-20 · unverdicted · none · ref 15

    Transformers need depth scaling as the product of ceil(k/s) and log n terms for k-hop pointer chasing under cache size s, with a conjectured lower bound, proved upper bound via windowed pointer doubling, and an adaptive-oblivious error separation.

  • Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory cs.LG · 2026-03-27 · unverdicted · none · ref 47

    Muon achieves higher storage capacity than SGD and matches Newton's method in one-step recovery rates for associative memory under power-law distributions, while saturating at larger critical batch sizes and showing faster initial multi-step dynamics.

  • Scaling Latent Reasoning via Looped Language Models cs.CL · 2025-10-29 · conditional · none · ref 79

    A 1.4B and a 2.6B looped (weight-tied, recurrent-depth) language model trained on 7.7T tokens match or exceed several 4B–8B transformer baselines on selected reasoning benchmarks.

  • There Will Be a Scientific Theory of Deep Learning stat.ML · 2026-04-23 · unverdicted · none · ref 245

    A mechanics of the learning process is emerging in deep learning theory, characterized by dynamics, coarse statistics, and falsifiable predictions across idealized settings, limits, laws, hyperparameters, and universal behaviors.