Pith. sign in

REVIEW 2 cited by

Can Large Language Models Understand Real-World Complex Instructions?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.09150 v2 pith:V2CDNRAZ submitted 2023-09-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords complexinstructionsllmscelloconstraintsmodelsunderstandability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) can understand human instructions, showing their potential for pragmatic applications beyond traditional NLP tasks. However, they still struggle with complex instructions, which can be either complex task descriptions that require multiple tasks and constraints, or complex input that contains long context, noise, heterogeneous information and multi-turn format. Due to these features, LLMs often ignore semantic constraints from task descriptions, generate incorrect formats, violate length or sample count constraints, and be unfaithful to the input text. Existing benchmarks are insufficient to assess LLMs' ability to understand complex instructions, as they are close-ended and simple. To bridge this gap, we propose CELLO, a benchmark for evaluating LLMs' ability to follow complex instructions systematically. We design eight features for complex instructions and construct a comprehensive evaluation dataset from real-world scenarios. We also establish four criteria and develop corresponding metrics, as current ones are inadequate, biased or too strict and coarse-grained. We compare the performance of representative Chinese-oriented and English-oriented models in following complex instructions through extensive experiments. Resources of CELLO are publicly available at https://github.com/Abbey4799/CELLO.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios

    cs.CL 2024-12 conditional novelty 6.0 of 10

    RuleArena evaluates LLMs on realistic rule-guided reasoning and finds that even o1-preview solves only about half of the easiest problems and near zero of the hardest.

  2. Med-Banana: Learning Quality-Controlled Medical Image Editing from Success-and-Failure Trajectories

    cs.CV 2025-11 reject novelty 5.0 of 10

    Med-Banana-50K is a dataset of ~88K AI-generated medical image edits (accepts and rejects) across 23 diseases, labeled by a single commercial LLM judge with minimal expert validation.

Pith tools