Pith. sign in

REVIEW 9 cited by

Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.12351 v2 pith:5HFLHMFR submitted 2023-11-21 cs.CL cs.LG

classification cs.CLcs.LG
keywords llmslong-contextmodelstransformer-basedarchitecturecomprehensivecurrentlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer-based Large Language Models (LLMs) have been applied in diverse areas such as knowledge bases, human interfaces, and dynamic agents, and marking a stride towards achieving Artificial General Intelligence (AGI). However, current LLMs are predominantly pretrained on short text snippets, which compromises their effectiveness in processing the long-context prompts that are frequently encountered in practical scenarios. This article offers a comprehensive survey of the recent advancement in Transformer-based LLM architectures aimed at enhancing the long-context capabilities of LLMs throughout the entire model lifecycle, from pre-training through to inference. We first delineate and analyze the problems of handling long-context input and output with the current Transformer-based models. We then provide a taxonomy and the landscape of upgrades on Transformer architecture to solve these problems. Afterwards, we provide an investigation on wildly used evaluation necessities tailored for long-context LLMs, including datasets, metrics, and baseline models, as well as optimization toolkits such as libraries, frameworks, and compilers to boost the efficacy of LLMs across different stages in runtime. Finally, we discuss the challenges and potential avenues for future research. A curated repository of relevant literature, continuously updated, is available at https://github.com/Strivin0311/long-llms-learning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PARTREP: Learning What to Repeat for Decoder-only LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    PartRep selects high-NLL tokens via a lightweight early-exit gate for partial prompt repetition, retaining most full-repetition gains at 59.4% KV cache and 79% prefill FLOPs on eight benchmarks.

  2. VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory

    cs.RO 2026-03 conditional novelty 6.0 of 10

    A transformer-based compressor that turns old observations into fixed-size memory tokens improves non-Markovian imitation-learning robot policies, with large gains on memory-intensive simulated tasks.

  3. CHEM: Estimating and Understanding Hallucinations in Deep Learning for Image Processing

    cs.CV 2025-12 unverdicted novelty 6.0 of 10

    The paper defines the Conformal Hallucination Estimation Metric (CHEM) that localizes hallucination-prone regions in image reconstruction models via multiscale representations and distribution-free conformal regression.

  4. Customizing the Inductive Biases of Softmax Attention using Structured Matrices

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Structured-matrix scoring functions, BTT and MLR, let attention escape the low-rank bottleneck and add a distance-dependent compute bias, improving accuracy for fixed compute on regression, language modeling, and forecasting.

  5. Multi-User SLNR-Based Precoding With Gold Nanoparticles in Vehicular VLC Systems

    cs.IT 2025-08 unverdicted novelty 6.0 of 10

    Gold-nanoparticle decorrelation of LED channels plus optimized RGB ratios improves multi-user vehicular visible light communication rate and secrecy.

  6. When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A new audit framework, SODA, measures demographic bias in objects generated by text-to-image models and finds strong default-to-majority and stereotype-collapse patterns across five models.

  7. Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial

    cs.NI 2025-09 conditional novelty 4.0 of 10

    A survey and tutorial that organizes LLM-enabled wireless network optimization into formulation, solution, and verification stages, with case studies drawn from the authors' own prior papers.

  8. Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior

    cs.CR 2025-08 conditional novelty 4.0 of 10

    Embedding a short 'system instruction' in a .docx file causes several commercial LLMs to refuse, substitute, redirect, or bias their output during summarization tasks.

  9. A Survey on the Memory Mechanism of Large Language Model based Agents

    cs.AI 2024-04 accept novelty 3.0 of 10

    A systematic review of memory designs, evaluation methods, applications, limitations, and future directions for LLM-based agents.

Pith tools