Pith. sign in

REVIEW 1 cited by

Evaluating, Understanding, and Improving Constrained Text Generation for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16343 v2 pith:OAIWUGEW submitted 2023-10-25 cs.CL

classification cs.CL
keywords generationllmstextconstrainedconstraintslanguageevaluatingimproving
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Advancements in natural language generation (NLG) and large language models (LLMs) have led to proficient text generation in various tasks. However, integrating intricate constraints into neural text generation, due to LLMs' opacity, remains challenging. This study investigates constrained text generation for LLMs, where predefined constraints are applied during LLM's generation process. Our research mainly focuses on mainstream open-source LLMs, categorizing constraints into lexical, structural, and relation-based types. We also present various benchmarks to facilitate fair evaluation. The study addresses some key research questions, including evaluating, understanding and improving constrained text generation for LLMs. Results illuminate LLMs' capacity and deficiency to incorporate constraints and provide insights for future developments in constrained text generation. Codes and datasets will be released upon acceptance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Instruction-Tuning Imparts Length Control: A Cross-Lingual Mechanistic Analysis

    cs.CL 2025-09 reject novelty 5.0 of 10

    Instruction-tuned Llama 3.1 controls word count far better than the base model, and attribution scores point to later layers, but the scoring rule mishandles outputs that are too short.

Pith tools