Pith. sign in

REVIEW 5 cited by

A Survey on Long Text Modeling with Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.14502 v2 pith:DMPIUPXN submitted 2023-02-28 cs.CL

classification cs.CL
keywords longmodelingtexttextsmodelstransformercharacteristicsdiscuss
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modeling long texts has been an essential technique in the field of natural language processing (NLP). With the ever-growing number of long documents, it is important to develop effective modeling methods that can process and analyze such texts. However, long texts pose important research challenges for existing text models, with more complex semantics and special characteristics. In this paper, we provide an overview of the recent advances on long texts modeling based on Transformer models. Firstly, we introduce the formal definition of long text modeling. Then, as the core content, we discuss how to process long input to satisfy the length limitation and design improved Transformer architectures to effectively extend the maximum context length. Following this, we discuss how to adapt Transformer models to capture the special characteristics of long texts. Finally, we describe four typical applications involving long text modeling and conclude this paper with a discussion of future directions. Our survey intends to provide researchers with a synthesis and pointer to related work on long text modeling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    LongReD reduces short-text performance loss after long-context extension by training the extended model to match the original model's hidden states on short texts and using skipped position indices to bridge short and...

  2. Generative Retrieval for Book search

    cs.IR 2025-01 conditional novelty 5.0 of 10

    GBS applies generative retrieval to book search by augmenting training data with hierarchical book identifiers and pseudo-queries, and encoding books with outline-based bi-level positions and retentive attention, repo...

  3. A Study on Context Length and Efficient Transformers for Biomedical Image Analysis

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Smaller image patches improve transformer accuracy on biomedical tasks, attention-window size matters less, and Hyena/MambaVision operators match attention with up to 80% faster training.

  4. Boosting Long-Context Management via Query-Guided Activation Refilling

    cs.CL 2024-12 conditional novelty 4.0 of 10

    ACRE uses a bi-layer KV cache with query-guided refilling to answer long-context questions beyond a model's native window, reporting gains over RAG and compression baselines.

  5. Words of War: Exploring the Presidential Rhetorical Arsenal with Deep Learning

    cs.LG 2024-12 reject novelty 3.0 of 10

    Deep learning classifiers appear to distinguish pre-war presidential speeches from others, but the reported high performance is inflated by resampling the full dataset before splitting into train and test sets.

Pith tools