Pith. sign in

Attention is all you need

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

representative citing papers

HRM-Text: Efficient Pretraining Beyond Scaling

cs.CL · 2026-05-20 · unverdicted · novelty 6.0

A 1B-parameter hierarchical recurrent model pretrained on 40B instruction-response tokens achieves 60.7% MMLU and strong results on ARC-C, DROP, GSM8K, and MATH while using 100-900x fewer tokens than standard baselines.

Graph Star Net for Generalized Multi-Task Learning

cs.SI · 2019-06-21 · unverdicted · novelty 6.0

GraphStar is a new GNN that adds star nodes and relay attention to achieve non-local representations for node, graph, and link tasks, claiming 2-5% gains over prior SOTA on benchmarks.

citing papers explorer

Showing 6 of 6 citing papers.