Pith. sign in

Sliding window attention training for efficient large language models

9 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

9 Pith papers citing it
1 external citations · external index

citation-role summary

background 3

citation-polarity summary

years

2026 6 2025 3

roles

background 3

polarities

background 3

representative citing papers

IAFormer: Interaction-Aware Transformer network for collider data analysis

hep-ph · 2025-05-06 · unverdicted · novelty 7.0

IAFormer uses boost-invariant pairwise quantities and differential attention to create a sparse Transformer that achieves state-of-the-art classification on top-quark and quark-gluon jet datasets while using over an order of magnitude fewer parameters than prior Particle Transformer models.

A Survey of Scaling in Large Language Model Reasoning

cs.AI · 2025-04-02 · unverdicted · novelty 3.0

A survey categorizing scaling in LLM reasoning across input size, steps, rounds, training, and future directions, noting that scaling can negatively affect performance.

citing papers explorer

Showing 9 of 9 citing papers.