Pith. sign in

REVIEW 4 cited by

System 2 Attention (is something you might need too)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.11829 v1 pith:MB4NWU57 submitted 2023-11-20 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords attentioncontextllmsinformationirrelevantlanguagesystemability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Soft attention in Transformer-based Large Language Models (LLMs) is susceptible to incorporating irrelevant information from the context into its latent representations, which adversely affects next token generations. To help rectify these issues, we introduce System 2 Attention (S2A), which leverages the ability of LLMs to reason in natural language and follow instructions in order to decide what to attend to. S2A regenerates the input context to only include the relevant portions, before attending to the regenerated context to elicit the final response. In experiments, S2A outperforms standard attention-based LLMs on three tasks containing opinion or irrelevant information, QA, math word problems and longform generation, where S2A increases factuality and objectivity, and decreases sycophancy.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Concretized Proposition Prompting Resolves Composition-Knowledge Dichotomy in Large Language Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    CPP improves LLM QA by generating TP/TN/FP/FN propositions then answering with CoT, outperforming many prompting baselines on medical and commonsense benchmarks.

  2. Leveraging Language Prior for Infrared Small Target Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Language priors from GPT-4V, converted by CLIP into text embeddings, improve infrared small target detection when used only during training.

  3. Reasoning in machine vision by learning fast and slow thinking

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A dual-process vision system improves segmentation accuracy by spending more inference-time compute, using a fast predictor and a slow self-play refiner, reporting gains on cancer localisation with only 8-16 labels.

  4. Survey of Specialized Large Language Model

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A survey of 24 specialized LLMs (2022-2025) claims a shift from domain fine-tuning to native architectures, but the synthesis is undermined by citation errors and selection bias.

Pith tools