Pith. sign in

REVIEW 1 cited by

StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14494 v2 pith:EAKTCZXM submitted 2025-02-20 cs.CL

classification cs.CL
keywords multi-turnstructuralevaluationfollowinginstructionbenchmarkdialogueflow
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-turn instruction following capability constitutes a core competency of large language models (LLMs) in real-world applications. Existing evaluation benchmarks predominantly focus on fine-grained constraint satisfaction and domain-specific capability assessment, yet overlook the crucial structural dependencies between dialogue turns that distinguish multi-turn from single-turn interactions. These structural dependencies not only reflect user intent but also establish an essential second dimension for the instruction following evaluation beyond constraint satisfaction. To address this gap, we propose StructFlowBench, a multi-turn instruction following benchmark with structural flow modeling. The benchmark defines an innovative structural flow framework with six fundamental inter-turn relationships. These relationships introduce novel structural constraints for model evaluation and also serve as generation parameters for creating customized dialogue flows tailored to specific scenarios. Adopting established LLM-based automatic evaluation methodologies, we conduct systematic evaluations of 13 leading open-source and closed-source LLMs. Experimental results reveal significant deficiencies in current models' comprehension of multi-turn dialogue structures. The code is available at https://github.com/MLGroupJLU/StructFlowBench.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Across 23 reasoning models on 420 constrained math problems, stronger reasoning-oriented training and longer chains of thought are associated with worse adherence to user-specified constraints.

Pith tools