Pith. sign in

REVIEW 1 cited by

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.19293 v2 pith:NLMF3BJF submitted 2025-05-25 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords long-contextbenchmarksabilitybaselineevaluatingllmslongbenchmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Long-context capability is considered one of the most important abilities of LLMs, as a truly long context-capable LLM enables users to effortlessly process many originally exhausting tasks -- e.g., digesting a long-form document to find answers vs. directly asking an LLM about it. However, existing real-task-based long-context evaluation benchmarks have two major shortcomings. First, benchmarks like LongBench often do not provide proper metrics to separate long-context performance from the model's baseline ability, making cross-model comparison unclear. Second, such benchmarks are usually constructed with fixed input lengths, which limits their applicability across different models and fails to reveal when a model begins to break down. To address these issues, we introduce a length-controllable long-context benchmark and a novel metric that disentangles baseline knowledge from true long-context capabilities. Experiments demonstrate the superiority of our approach in effectively evaluating LLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning

    cs.CL 2026-07 conditional novelty 8.0 of 10

    WILDTRACE evaluates long-context models on 481 natural multi-hop evidence trails from 214 real documents, with top systems at 75.3% and geometry-specific weaknesses.

Pith tools