SCICONVBENCH is a new benchmark evaluating LLMs on multi-turn disambiguation and inconsistency resolution for task formulation in computational science, with frontier models reaching only 52.7% success on fluid mechanics disambiguation cases.
arXiv preprint arXiv:2309.13233 , year=
7 Pith papers cite this work, alongside 3 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 7representative citing papers
PERCEIVE is the first bilingual benchmark integrating author content, reader emotions from comments, communication behavior, user attributes, and social graphs for personalized social media emotion understanding.
DITTO uses RL with verbal feedback to train LLMs for human behavior simulation, reporting 36% average gains over base models and outperforming GPT-5.4 on 6 of 10 SOUL benchmark tasks.
A retail user-simulator benchmark and GRPO training recipe claim improved persona adherence, but the paper's abstract and body disagree on core numbers.
MUSE generates realistic, persona-consistent Chinese user responses across domains via self-evolving profiles, role-reversal fine-tuning, and rubric-guided multi-turn RL, outperforming baselines in utterance and session evaluations.
DiPS uses Implicit Q-Learning over dialogue history embeddings to select persuasion policies turn-by-turn, raising evacuation success above zero-shot LLM and RAG baselines in simulation and human studies.
Fine-tuned simulators grounded in real human data produce LLM assistants that win more often against real users than those trained against role-playing simulators.
citing papers explorer
-
SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science
SCICONVBENCH is a new benchmark evaluating LLMs on multi-turn disambiguation and inconsistency resolution for task formulation in computational science, with frontier models reaching only 52.7% success on fluid mechanics disambiguation cases.
-
PERCEIVE: A Benchmark for Personalized Emotion and Communication Behavior Understanding on Social Media
PERCEIVE is the first bilingual benchmark integrating author content, reader emotions from comments, communication behavior, user attributes, and social graphs for personalized social media emotion understanding.
-
Reinforcing Human Behavior Simulation via Verbal Feedback
DITTO uses RL with verbal feedback to train LLMs for human behavior simulation, reporting 36% average gains over base models and outperforming GPT-5.4 on 6 of 10 SOUL benchmark tasks.
-
CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
A retail user-simulator benchmark and GRPO training recipe claim improved persona adherence, but the paper's abstract and body disagree on core numbers.
-
MUSE: Multi-Domain Chinese User Simulation via Self-Evolving Profiles and Rubric-Guided Alignment
MUSE generates realistic, persona-consistent Chinese user responses across domains via self-evolving profiles, role-reversal fine-tuning, and rubric-guided multi-turn RL, outperforming baselines in utterance and session evaluations.
-
DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents
DiPS uses Implicit Q-Learning over dialogue history embeddings to select persuasion policies turn-by-turn, raising evacuation success above zero-shot LLM and RAG baselines in simulation and human studies.
-
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants
Fine-tuned simulators grounded in real human data produce LLM assistants that win more often against real users than those trained against role-playing simulators.