Pith. sign in

REVIEW 5 cited by

Needle in the Haystack for Memory Based Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.01437 v2 pith:DW6GUDZL submitted 2024-07-01 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords memorycontextsexternallanguagelarimartaskstraininglarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current large language models (LLMs) often perform poorly on simple fact retrieval tasks. Here we investigate if coupling a dynamically adaptable external memory to a LLM can alleviate this problem. For this purpose, we test Larimar, a recently proposed language model architecture which uses an external associative memory, on long-context recall tasks including passkey and needle-in-the-haystack tests. We demonstrate that the external memory of Larimar, which allows fast write and read of an episode of text samples, can be used at test time to handle contexts much longer than those seen during training. We further show that the latent readouts from the memory (to which long contexts are written) control the decoder towards generating correct outputs, with the memory stored off of the GPU. Compared to existing transformer-based LLM architectures for long-context recall tasks that use larger parameter counts or modified attention mechanisms, a relatively smaller size Larimar is able to maintain strong performance without any task-specific training or training on longer contexts.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

    cs.AI 2025-10 unverdicted novelty 6.0 of 10

    TRACE uses an evidence bank to score tool-augmented LLM agents on efficiency, hallucination, and adaptivity without ground-truth trajectories.

  2. Can an Actor-Critic Optimization Framework Improve Analog Design?

    cs.LG 2026-03 conditional novelty 5.0 of 10

    An actor-critic framework with two LLM agents—one proposing and one auditing search regions—improves analog sizing by 38.9% in top-10 FoM and 24.7% in regret over a single-LLM baseline.

  3. Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Identical video questions get different accuracy when placed at the start, middle, or end of a long context, and the new benchmark maps this bias across 27 video-language models.

  4. Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A new maze-navigation benchmark for LLMs reports that reasoning models outperform standard ones, but the link from performance gaps to a lack of persistent self-awareness is an overreach.

  5. SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models

    cs.CL 2025-07 reject novelty 4.0 of 10

    SCOPE estimates a model's position bias with nonsense prompts, puts correct answers in disliked slots, and spreads similar distractors apart to cap lucky guessing.

Pith tools