Pith. sign in

REVIEW 4 cited by

WildIFEval: Instruction Following in the Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.06573 v3 pith:AKNJI62I submitted 2025-03-09 cs.CL cs.AI

classification cs.CLcs.AI
keywords constraintsinstructionswildifevaluserconditionsdatasetfollowinginstruction-following
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent LLMs have shown remarkable success in following user instructions, yet handling instructions with multiple constraints remains a significant challenge. In this work, we introduce WildIFEval - a large-scale dataset of 7K real user instructions with diverse, multi-constraint conditions. Unlike prior datasets, our collection spans a broad lexical and topical spectrum of constraints, extracted from natural user instructions. We categorize these constraints into eight high-level classes to capture their distribution and dynamics in real-world scenarios. Leveraging WildIFEval, we conduct extensive experiments to benchmark the instruction-following capabilities of leading LLMs. WildIFEval clearly differentiates between small and large models, and demonstrates that all models have a large room for improvement on such tasks. We analyze the effects of the number and type of constraints on performance, revealing interesting patterns of model constraint-following behavior. We release our dataset to promote further research on instruction-following under complex, realistic conditions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prompts in the Wild: A Large Analyzed Collection of Transactional Prompts in Code

    cs.CL 2026-08 conditional novelty 7.0 of 10

    A large dataset of 57.5K GitHub prompts is annotated with a new ontology, enabling quantitative analysis of how transactional prompts are used in real code.

  2. Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Agentic coding prompts grow because the reasoning behind old instructions decays, and comments that preserve that reasoning halt the growth and recover instruction-following.

  3. UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following

    cs.CL 2026-08 conditional novelty 6.0 of 10

    UNSPECIFIC generates constraints common to two similar articles, hardens only too-easy constraints, and scores satisfaction after summarization, yielding a harder and more natural instruction-following benchmark.

  4. Revisiting the Reliability of Language Models in Instruction-Following

    cs.SE 2025-12 conditional novelty 6.0 of 10

    LLMs exhibit up to 61.8% performance drops on nuanced rephrasings of instruction-following tasks, revealing insufficient nuance-oriented reliability across 46 tested models.

Pith tools