Pith. sign in

REVIEW 7 cited by

Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.02823 v1 pith:IU55N3OB submitted 2024-04-03 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords conifermodelscomplexinstructionsconstraintsdatasetinstruction-followingllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability of large language models (LLMs) to follow instructions is crucial to real-world applications. Despite recent advances, several studies have highlighted that LLMs struggle when faced with challenging instructions, especially those that include complex constraints, hindering their effectiveness in various tasks. To address this challenge, we introduce Conifer, a novel instruction tuning dataset, designed to enhance LLMs to follow multi-level instructions with complex constraints. Utilizing GPT-4, we curate the dataset by a series of LLM-driven refinement processes to ensure high quality. We also propose a progressive learning scheme that emphasizes an easy-to-hard progression, and learning from process feedback. Models trained with Conifer exhibit remarkable improvements in instruction-following abilities, especially for instructions with complex constraints. On several instruction-following benchmarks, our 7B model outperforms the state-of-the-art open-source 7B models, even exceeds the performance of models 10 times larger on certain metrics. All the code and Conifer dataset are available at https://www.github.com/ConiferLM/Conifer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

    cs.AI 2025-05 conditional novelty 7.0 of 10

    AgentIF introduces a realistic, long-form instruction-following benchmark for agentic scenarios and shows that current LLMs follow fewer than 30% of such instructions perfectly.

  2. CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

    cs.AI 2026-07 conditional novelty 6.0 of 10

    CoRT uses per-token likelihood contrasts between rubric-conditioned and criteria-free prompts to redistribute GRPO advantages, improving rubric instruction-following RL without an auxiliary token scorer.

  3. VerIF: Verification Engineering for Reinforcement Learning in Instruction Following

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A hybrid verifier that combines rule-based code checks and a reasoning-LLM judge enables reinforcement learning to improve LLM instruction following on several benchmarks.

  4. STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    STEER-BENCH is a Reddit-derived benchmark of 5,552 multiple-choice questions on which the best of 13 large language models scores near 65 percent, versus human experts near 81 percent.

  5. From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    An explainable model using BLEURT, CometKiwi, pause features, and Chinese phraseological diversity predicts human-rated quality dimensions in English-Chinese consecutive interpreting, with SHAP identifying the stronge...

  6. DSMentor: Enhancing Data Science Agents with Curriculum Learning and Online Knowledge Accumulation

    cs.AI 2025-05 conditional novelty 5.0 of 10

    Ordering data science problems easy-to-hard and accumulating their solutions in a memory buffer improves LLM agent pass rates on DSEval and QRData by up to 5.2%.

  7. Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A hierarchical direct preference optimization with four alignment levels plus automated data selection improves physical plausibility of text-to-video models.

Pith tools