Pith. sign in

REVIEW 2 cited by

What Do Position Embeddings Learn? An Empirical Study of Pre-Trained Language Model Positional Encoding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.04903 v1 pith:KMADJ5XT submitted 2020-10-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords positionembeddingspre-trainedtransformerstasksempiricaldifferentencoding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, pre-trained Transformers have dominated the majority of NLP benchmark tasks. Many variants of pre-trained Transformers have kept breaking out, and most focus on designing different pre-training objectives or variants of self-attention. Embedding the position information in the self-attention mechanism is also an indispensable factor in Transformers however is often discussed at will. Therefore, this paper carries out an empirical study on position embeddings of mainstream pre-trained Transformers, which mainly focuses on two questions: 1) Do position embeddings really learn the meaning of positions? 2) How do these different learned position embeddings affect Transformers for NLP tasks? This paper focuses on providing a new insight of pre-trained position embeddings through feature-level analysis and empirical experiments on most of iconic NLP tasks. It is believed that our experimental results can guide the future work to choose the suitable positional encoding function for specific tasks given the application property.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AttentionSmithy: A Modular Framework for Rapid Transformer Development and Customization

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A modular transformer framework is introduced and validated by replicating the original transformer, running neural architecture search on WMT14 translation, and fine-tuning a BERT-style model for cell type classification.

  2. HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models

    cs.CL 2025-09 reject novelty 4.0 of 10

    HoPE replaces RoPE's sine/cosine rotations with hyperbolic functions plus an exponential damping term to enforce monotonic attention decay, but the claimed consistent superiority and the 'RoPE as special case' theorem...

Pith tools