REVIEW 3 cited by
COLLIE: Systematic Construction of Constrained Text Generation Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text generation under constraints have seen increasing interests in natural language processing, especially with the rapidly improving capabilities of large language models. However, existing benchmarks for constrained generation usually focus on fixed constraint types (e.g.,generate a sentence containing certain words) that have proved to be easy for state-of-the-art models like GPT-4. We present COLLIE, a grammar-based framework that allows the specification of rich, compositional constraints with diverse generation levels (word, sentence, paragraph, passage) and modeling challenges (e.g.,language understanding, logical reasoning, counting, semantic planning). We also develop tools for automatic extraction of task instances given a constraint structure and a raw text corpus. Using COLLIE, we compile the COLLIE-v1 dataset with 2080 instances comprising 13 constraint structures. We perform systematic experiments across five state-of-the-art instruction-tuned language models and analyze their performances to reveal shortcomings. COLLIE is designed to be extensible and lightweight, and we hope the community finds it useful to develop more complex constraints and evaluations in the future.
Forward citations
Cited by 3 Pith papers
-
HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
HSS-Synth generates 230k instruction-tuning samples for 14 humanities/social-science fields and reports state-of-the-art fine-tuning results on 16 benchmarks.
-
JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
JSONSchemaBench is a new 10K-schema benchmark showing that constrained decoding frameworks differ widely in efficiency, coverage, and quality, with the best tool supporting roughly twice as many schemas as the worst.
-
LCTG Bench: LLM Controlled Text Generation Benchmark
LCTG Bench is a new Japanese benchmark that scores LLM controllability on format, character count, keyword, and prohibited-word constraints across summarization, ad text, and pros/cons generation.
Discussion (0). Continue with ORCID to comment.