Pith. sign in

REVIEW 8 cited by

LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.02076 v7 pith:XSMADUCB submitted 2024-09-03 cs.CL

classification cs.CL
keywords textgenerationllmslonggenbenchlonglong-formmodelsability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current benchmarks like Needle-in-a-Haystack (NIAH), Ruler, and Needlebench focus on models' ability to understand long-context input sequences but fail to capture a critical dimension: the generation of high-quality long-form text. Applications such as design proposals, technical documentation, and creative writing rely on coherent, instruction-following outputs over extended sequences - a challenge that existing benchmarks do not adequately address. To fill this gap, we introduce LongGenBench, a novel benchmark designed to rigorously evaluate large language models' (LLMs) ability to generate long text while adhering to complex instructions. Through tasks requiring specific events or constraints within generated text, LongGenBench evaluates model performance across four distinct scenarios, three instruction types, and two generation-lengths (16K and 32K tokens). Our evaluation of ten state-of-the-art LLMs reveals that, despite strong results on Ruler, all models struggled with long text generation on LongGenBench, particularly as text length increased. This suggests that current LLMs are not yet equipped to meet the demands of real-world, long-form text generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    RedKnot reports up to 3.5× lower time-to-first-token and 4.7–7.8× more concurrent sessions by reusing KV cache at the granularity of attention heads rather than tokens, while preserving QA accuracy.

  2. ReportLogic: Evaluating Logical Quality in Deep Research Reports

    cs.CL 2026-01 conditional novelty 6.0 of 10

    An auditability-based benchmark with three logic layers and eight dimensions shows a distilled judge agrees with human experts ~74-75% versus ~62-74% for frontier LLM judges.

  3. Power Law Guided Dynamic Sifting for Efficient Attention

    cs.LG 2025-06 conditional novelty 6.0 of 10

    SiftAttention skips top-k sorting in sparse attention by thresholding attention weights with a threshold predicted from a power-law fit of score quantiles over generation steps.

  4. SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 7B writing model trained on plan-write-refine thinking data with multi-stage preference optimization matches or beats several larger models on long-form generation benchmarks.

  5. GEIS: A Generation-Evaluation-Improvement Loop of Agent Skills for Long-Form Article Generation

    cs.CL 2026-07 conditional novelty 5.0 of 10

    A skill-based generate–evaluate–improve loop raises Wikipedia-style long-form article quality over fixed multi-agent pipelines and self-improves via permanent writing-rule patches.

  6. Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

    cs.CL 2026-01 conditional novelty 5.0 of 10

    On a new extended needle-in-a-haystack benchmark, explicit anti-hallucination prompts and dispersed fact placement cause some long-context LLMs to over-refuse or collapse in accuracy, while others remain robust.

  7. MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    MiniLongBench, a 237-sample compression of LongBench, is claimed to reproduce model rankings with a 0.97 Spearman correlation at 4.5% of the evaluation cost.

  8. An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3

    cs.CL 2025-05 conditional novelty 4.0 of 10

    LLMs can produce fluent movie reviews that readers often mistake for human-written ones, but the models differ in emotional balance and depth.

Pith tools