Pith. sign in

REVIEW 3 cited by

ENTP: Encoder-only Next Token Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.01600 v3 pith:YGJ3GC43 submitted 2024-10-02 cs.LG cs.CL

classification cs.LGcs.CL
keywords entpdecoder-onlypredictiontransformersencoder-onlyintroducenextnext-token
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Next-token prediction is conventionally done using decoder-only Transformers with causal attention, as this approach allows for efficient reuse of keys and values. What if we were not compute-limited, should we still use decoder-only Transformers? In this work, we introduce Encoder-only Next Token Prediction (ENTP). We explore the differences between ENTP and decoder-only Transformers in expressive power and complexity, highlighting potential advantages of ENTP in settings with unbounded compute. We introduce the $\operatorname{Count3}$ task and show, both theoretically and experimentally, that while ENTP can perform this task easily, a decoder-only Transformer cannot. Finally, we empirically demonstrate the superior performance of ENTP across representative tasks where next-token prediction based Transformers can be evaluated, including addition, in-context learning, and language modeling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Theoretical limitations of multi-layer Transformer

    cs.LG 2024-12 conditional novelty 8.0 of 10

    An L-layer decoder-only Transformer requires polynomial model dimension to compute L-step sequential function composition, and this is proven without any unproven complexity conjecture.

  2. Hierarchical Domain Generalization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Over infinite domains, hierarchy-uniform domain generalization is impossible for every nontrivial hypothesis class; a length-generalization bound is a property of the length hierarchy, not a hierarchy-free guarantee.

  3. Small Languages, Big Models: A Study of Continual Training on Languages of Norway

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A three-stage continual training recipe (tokenizer change, embedding alignment, full retraining) produces NorMistral-11B, an open Norwegian and Northern Sámi language model that improves on most Norwegian benchmarks a...

Pith tools