Pith. sign in

REVIEW 21 cited by

MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.14762 v3 pith:AIRPX2A7 submitted 2024-02-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords multi-turndialoguesllmsabilitiesdialoguemt-bench-101fine-grainedtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The advent of Large Language Models (LLMs) has drastically enhanced dialogue systems. However, comprehensively evaluating the dialogue abilities of LLMs remains a challenge. Previous benchmarks have primarily focused on single-turn dialogues or provided coarse-grained and incomplete assessments of multi-turn dialogues, overlooking the complexity and fine-grained nuances of real-life dialogues. To address this issue, we introduce MT-Bench-101, specifically designed to evaluate the fine-grained abilities of LLMs in multi-turn dialogues. By conducting a detailed analysis of real multi-turn dialogue data, we construct a three-tier hierarchical ability taxonomy comprising 4208 turns across 1388 multi-turn dialogues in 13 distinct tasks. We then evaluate 21 popular LLMs based on MT-Bench-101, conducting comprehensive analyses from both ability and task perspectives and observing differing trends in LLMs performance across dialogue turns within various tasks. Further analysis indicates that neither utilizing common alignment techniques nor chat-specific designs has led to obvious enhancements in the multi-turn abilities of LLMs. Extensive case studies suggest that our designed tasks accurately assess the corresponding multi-turn abilities. The data and code are available at \url{https://github.com/mtbench101/mt-bench-101}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    DecisionBench supplies a fixed task suite, model pool, delegation interface, and multi-axis metrics to evaluate emergent delegation, showing similar quality across awareness conditions but 15-31 point headroom under p...

  2. EditPropBench: Measuring Factual Edit Propagation in Scientific Manuscripts

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    EditPropBench evaluates LLM editors on propagating factual edits to dependent claims in synthetic scientific manuscripts, showing that even the strongest systems miss roughly 30% of required updates on hard cases.

  3. EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving

    cs.AI 2025-09 unverdicted novelty 7.0 of 10

    EngiBench shows LLMs accuracy drops with task complexity, degrades under perturbations, and stays below human performance on open-ended engineering problems.

  4. Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A 17-model benchmark with a decoupled simulator/judge design finds frontier chatbots indistinguishable on subjective warmth but sharply separated on long-horizon intent tracking, with a reasoning-mode gain that appear...

  5. Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A three-party-decoupled multi-turn chat benchmark finds closed and open models nearly tied on subjective empathy/persona scores but separated by up to 9× on objective intent tracking, with reasoning and persona format...

  6. SOMA: Efficient Multi-turn LLM Serving via Small Language Model

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    SOMA estimates a local response manifold from early turns and adapts a small surrogate model via divergence-maximizing prompts and localized LoRA fine-tuning for efficient multi-turn serving.

  7. Mechanistic Analysis of Alignment Algorithms in Language Models

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Mechanistic analysis of six preference optimization methods reveals distinct geometric shifts in model representations, with KTO/GRPO enhancing separability while DPO/ORPO degrade it.

  8. Data-dependent Exploration for Online Reinforcement Learning from Human Feedback

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    DEPO uses historical data to build a data-dependent uncertainty bonus for exploration in online RLHF, yielding an adaptive regret bound and stronger empirical performance than baselines.

  9. CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks

    cs.CR 2026-04 unverdicted novelty 6.0 of 10

    CoopGuard deploys cooperative agents to track conversation history and counter evolving multi-round attacks on LLMs, achieving a 78.9% reduction in attack success rate on a new 5,200-sample benchmark.

  10. SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding

    cs.DC 2026-02 unverdicted novelty 6.0 of 10

    SPEED-Bench is a new standardized benchmark for speculative decoding that supplies semantically diverse qualitative data and throughput-oriented splits across concurrency levels, integrated with vLLM and TensorRT-LLM.

  11. SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding

    cs.DC 2026-02 conditional novelty 6.0 of 10

    A new benchmark for speculative decoding that maximizes semantic diversity and supports throughput evaluation across input lengths, exposing biases in synthetic benchmarks.

  12. TRINITY: An Evolved LLM Coordinator

    cs.LG 2025-12 unverdicted novelty 6.0 of 10

    A compact 0.6B-parameter coordinator with a 10K-parameter head uses evolutionary strategy to dynamically delegate roles to LLMs, achieving SOTA results such as 86.2% on LiveCodeBench.

  13. VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

    cs.SD 2025-10 conditional novelty 6.0 of 10

    A Chinese benchmark built on real human speech evaluates large audio language models across instruction following, knowledge, and robustness, revealing large performance gaps.

  14. On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Paraphrasing benchmark questions keeps LLM rankings stable but reduces their accuracy, suggesting static benchmarks overestimate model robustness.

  15. Data-dependent Exploration for Online Reinforcement Learning from Human Feedback

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    DEPO constructs uncertainty bonuses from historical data for exploration in online RLHF and provides a data-dependent regret bound that adapts to task hardness.

  16. Computational Hermeneutics: Evaluating generative AI as a cultural technology

    cs.AI 2026-03 unverdicted novelty 5.0 of 10

    Generative AI should be evaluated through computational hermeneutics using iterative, human-inclusive benchmarks that measure cultural context rather than isolated model outputs.

  17. ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing

    cs.LG 2025-07 unverdicted novelty 5.0 of 10

    ReasonCache reuses similar KV cache states across reasoning steps in LRMs via collaborative filtering to boost serving throughput by up to 89.2% while preserving accuracy.

  18. Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks

    cs.CR 2026-07 reject novelty 4.0 of 10

    CoopGuard's defer-tempt-analyze-coordinate agents cut reported jailbreak success and raise attacker token costs on the new EMRA benchmark, but the deceptive-rate metric is partly defined by the paper's own scoring rubric.

  19. HEAL: A Hypothesis-Based Preference-Aware Analysis Framework

    cs.CL 2025-08 conditional novelty 4.0 of 10

    HEAL evaluates preference optimization by measuring ranking accuracy and strength correlation between model likelihoods and proxy reward scores over multi-response hypothesis spaces.

  20. HAEPO: History-Aggregated Exploratory Policy Optimization

    cs.LG 2025-08 conditional novelty 4.0 of 10

    HAEPO weights each trajectory by its softmax-normalized cumulative log-likelihood, adds entropy and KL penalties, and matches or slightly surpasses PPO, GRPO, and DPO on small RL and summarization tasks.

  21. Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector

    cs.CL 2025-09 unverdicted novelty 3.0 of 10

    Fine-tuned LLaMA 3.1-8B variants for the energy sector outperform the base model on domain QA benchmarks, with LoRA delivering similar gains at lower training cost.

Pith tools