REVIEW 4 major objections 6 minor 1 cited by
Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Slowing dialogue with questions and revealed assumptions lifts task success
desk verdict A genuinely useful ontology of positive friction in dialogue, but the central task-success claim leans on an LLM-as-judge loop that may reward verbosity; the paper deserves review, not blind acceptance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the positive-friction ontology: five movement classes—assumption reveal, reflective pause, reinforcement, overspecification, and probing—each with subcategories, built from discourse and cognitive theories. It does two jobs: it is a coding scheme that lets human annotators and GPT-4o label turns, and it is a prompt-level intervention whose definitions and in-context examples are added to LLM dialogue agents (AutoTOD for MultiWOZ; a dialogue-enabled ReAct agent for ALFWorld). The ontology is what turns 'slow down' from a vague value into a testable independent variable.
What would settle it
Run the same MultiWOZ and ALFWorld comparisons with human users or with a human-validated success label, and include a control condition that injects extra but non-frictive turns; the central claim would collapse if the success gains disappear or if arbitrary slowdowns reproduce them.
Extended reading notes
Core claim
The paper claims that deliberately slowing a goal-oriented conversation—through questions that probe the user's context, utterances that reveal assumptions, and overspecification of constraints—improves, rather than harms, the collaboration. On MultiWOZ, injecting these friction categories into the AutoTOD agent's prompts raises task success from 56.4% without friction to 62.8% with all three categories; on ALFWorld, probing raises success from 51.5% to 59.0% and cuts the average number of physical actions from 19.9 to 6.1. The paper also reports that conversations containing certain friction movements produce lower mean-squared error when a model infers user satisfaction, and that humans deploy friction at strategic dialogue positions. The intended consequence is that dialogue policies and evaluation metrics should treat utterance valence—whether a turn slows or speeds the interaction—as a first-class signal rather than optimizing only for shortness and superficial preference.
Load-bearing premise
The load-bearing assumption is that the automated evaluation loop—GPT-4o-mini acting as user and as success judge—measures real human goals and task completion faithfully; if simulated users simply reward longer, information-rich dialogues, the reported success gains may come from the evaluation setup rather than from positive friction itself.
Editorial extensions
If this is right
- In multi-domain booking, combining assumption reveal, probing, and overspecification raises success from 56.4% to 62.8%, while in embodied ALFWorld, probing raises success from 51.5% to 59.0% and reduces physical actions from about 19.9 to 6.1.
- Friction turns reduce model error when inferring user satisfaction from dialogue history in MultiWOZ, which suggests that slowed exchanges reveal more about the user's mental state.
- Friction categories cut across traditional dialogue acts: most acts can be performed with or without friction, and request-like acts are inherently frictive because they probe for information.
- Because friction lengthens dialogues, evaluation metrics that penalize every extra turn are misaligned with long-term task success and should be rethought.
- Using all friction categories at once can hurt embodied performance when the environment imposes a step limit, so friction must be timed and selected rather than applied uniformly.
Reading between the lines
- Editorial inference: the ontology could be turned into reward-shaping features for RLHF-style training, since preference data collected over whole interactions—rather than single turns—would let models learn when friction pays off and when it does not.
- Editorial inference: the timing results (probing early, pauses later) suggest a testable policy: inject friction only at uncertainty-critical decision points, and leave fluent stretches of dialogue untouched.
- Editorial inference: applying the same annotation scheme to other high-stakes domains, such as medical or financial advice, could test whether the task-success gains generalize beyond booking and embodied household tasks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces "positive friction" as a design principle for goal-oriented dialogue systems, formalizes it as an ontology of five categories (assumption reveal, reflective pause, reinforcement, overspecification, probing) with subcategories, and collects human annotations on MultiWOZ and TEACh as well as GPT-4o proxy annotations on three corpora. It then reports two utility studies: a correlational analysis in §4.1 linking automatically detected friction to lower error in inferring user satisfaction, and task-success experiments in §4.2 that prompt AutoTOD and a ReAct-based ALFWorld agent with friction definitions and evaluate against GPT-4o-mini user simulation. The paper concludes that positive friction improves mental-state modeling and task success.
Significance. The taxonomy and annotation effort are a useful conceptual contribution, and the attempt to measure downstream task success with environment-based ALFWorld success is commendable. If the causal claims were established, the paper would challenge the common efficiency-first assumption in dialogue policy and provide a practical vocabulary for designing reflective interactions. The strengths are the clear ontology, the human annotation protocol, and the planned public release of code and data. The main weakness is that the headline empirical results currently rest on a simulated-user evaluation loop and a correlational analysis, so the empirical support is not yet commensurate with the strength of the conclusions.
major comments (4)
- [§4.2, Table 2, Appendix E] The MultiWOZ success improvements (e.g., 56.4% no-friction vs. 62.8% all-friction) do not yet establish that the friction ontology causes better task success. The online Success metric is computed by GPT-4o-mini from the dialogue transcript, and the same model also serves as the user simulator. Because the friction prompt explicitly instructs the assistant to probe and overspecify, frictive dialogues are longer and restate constraints, which an LLM judge answering a QA-style 'were all attributes provided?' question may reward regardless of whether the booking outcome is genuinely better. A length-matched control (e.g., a non-frictive system that repeats constraints without the friction framing) and human verification of a sample of successful dialogues are needed before the performance gap can be attributed to friction.
- [§4.2, ALFWorld results] The claim of a 'significant improvement' for probing (58.96% vs. 51.49%) is not supported by a reported statistical test, and the setup gives the GPT-4o-mini user simulator the ability to answer the agent's probing questions with task-relevant information that may exceed what a real user would know. The authors should report a significance test over the 134 evaluation games and add a control that constrains the simulated user's answers to information available from the task instructions alone, or replace or supplement the simulator with human users. The step-limit explanation for the 'All three' drop (46.06%) should also be tested by ablating the step limit.
- [§4.1, Figure 4] The claim that friction 'improves user modeling' is not supported by the experimental design. Each dialogue contributes one randomly sampled turn annotated with a friction category, with no control for dialogue length, turn position, or topic, and the labels come from the GPT-4o proxy rather than from human annotations. The Kruskal-Wallis test only establishes that error distributions differ across categories; it does not establish that the friction causes the lower errors or that the category is not a proxy for turn position or dialogue length. The authors should either present a regression or matching analysis that adjusts for turn index and dialogue length, or soften the causal language to a reported association.
- [§3.3 and subsequent analyses] The paper relies on GPT-4o automatic friction labels for all downstream quantitative claims, but the agreement with the human majority vote is only moderate (Cohen's kappa 0.50) and lower against individual annotators (0.34), while the human annotators themselves agree only fairly (0.42 at category level). The paper should include robustness checks, for example by repeating key analyses on the subset of turns with high human agreement, by reporting per-category agreement, or by showing that the §4.1 and §3.4 findings are stable across alternative label sources.
minor comments (6)
- [Author affiliations] The author affiliation line contains a typo: 'Northeastearn University' should be 'Northeastern University'.
- [Table 1] The reinforcement example uses 'Turnt' and 'Turnt + 1' where 'Turn t' and 'Turn t+1' seem intended.
- [Figure 5 caption] The caption reports p = 0.1; the text says 'strategically use friction at different time points (p < 0.01) and friction often slows down conversations (p = 0.1)'; the latter is not significant at conventional levels and should be described as marginal rather than as a confirmatory result.
- [§3.2] Calling undergraduate annotators 'expert annotators' after a short lecture is potentially misleading; consider 'trained annotators' instead.
- [Table 2] The heading 'Fric. (%)' should be defined in the caption as the percentage of turns containing at least one friction movement.
- [Appendix E] The paper states temperature 0 for generations and averages over three runs; please clarify how variance arises (e.g., API nondeterminism) and report seeds if applicable.
Circularity Check
No significant circularity: the friction ontology is an empirical taxonomy, and the utility claims are tested on external corpora rather than derived from the ontology's definitions.
full rationale
The paper's central chain is: define positive friction; collect human annotations; detect friction automatically; then measure whether friction improves user-satisfaction inference (§4.1) and task success (§4.2). None of these steps reduces to an input by construction. The ontology's categories are defined descriptively (Definition 1, Table 1) from prior discourse and cognitive work (Stalnaker, Tannen, Wilkes-Gibbs and Clark, Zellner), not from the outcome measures. The §4.1 result compares squared errors in predicting external MultiWOZ satisfaction annotations (Sun et al., 2021) across GPT-4o, LLaMA, and Mixtral; the friction label at a random turn is not a fitted parameter and does not determine the regression target. The §4.2 MultiWOZ result uses AutoTOD's online Success metric, which checks whether all requested attributes are provided; although the judge is GPT-4o-mini and the friction prompts ask for overspecification, success is still anchored to MultiWOZ goals and the comparison is empirical, not a definitional identity. ALFWorld success is environment-based (object positions and states), so the main MultiWOZ judge concern does not transfer. The paper does cite prior work by overlapping authors (Sicilia and Alikhani 2024 for the satisfaction-inference method; Dongre et al. 2024 for the ReAct dialogue extension), but these are tools or baselines, not uniqueness theorems or assumptions that already contain the friction conclusion. The most serious caveat is construct validity, not circularity: the MultiWOZ success gap (56.4% vs. 62.8%) may partly reflect an LLM judge rewarding longer, constraint-restating dialogues, and no length-matched control or human verification is reported; that is a confound for the claim that friction per se causes the gain. It is not, however, a case of a prediction being equivalent to its inputs by construction. Score 1 reflects the minor self-citations and the evaluator-overlap caveat, not equation-level circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption GPT-4o can serve as a reliable automatic annotator of friction movements.
- domain assumption The dual-process (System 1/System 2) framework from cognitive science applies to conversational AI users as described.
- domain assumption AutoTOD's online Success metric, computed by GPT-4o-mini via a question-answering task, is a valid measure of task completion.
invented entities (5)
-
Assumption Reveal
-
Reflective Pause
-
Reinforcement
-
Overspecification
-
Probing
Cite this review
Pith. "Pith review of Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems." pith.science (2026). https://pith.science/paper/7BQCCSYK
@misc{pith2026250117348,
author = {Pith},
title = {Pith review of: Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/7BQCCSYK}},
note = {Machine review of arXiv:2501.17348}
}
read the original abstract
While theories of discourse and cognitive science have long recognized the value of unhurried pacing, recent dialogue research tends to minimize friction in conversational systems. Yet, frictionless dialogue risks fostering uncritical reliance on AI outputs, which can obscure implicit assumptions and lead to unintended consequences. To meet this challenge, we propose integrating positive friction into conversational AI, which promotes user reflection on goals, critical thinking on system response, and subsequent re-conditioning of AI systems. We hypothesize systems can improve goal alignment, modeling of user mental states, and task success by deliberately slowing down conversations in strategic moments to ask questions, reveal assumptions, or pause. We present an ontology of positive friction and collect expert human annotations on multi-domain and embodied goal-oriented corpora. Experiments on these corpora, along with simulated interactions using state-of-the-art systems, suggest incorporating friction not only fosters accountable decision-making, but also enhances machine understanding of user beliefs and goals, and increases task success rates.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 1 Pith paper
-
Dynamic Epistemic Friction in Dialogue
A vector-based belief-update model, grounded in dynamic epistemic logic, predicts final block-weight beliefs in a collaborative task from dialogue friction.
Reference graph
Works this paper leans on
-
[1]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Ian A. Anderson and Wendy Wood. 2021. https://doi.org/10.1002/arcp.1063 Habits and the electronic herd: The psychology behind social media ' s successes and failures . Consum. Psychol. Rev., 4(1):83--99
-
[4]
Nicholas Asher, Julie Hunter, Mathieu Morey, Benamara Farah, and Stergos Afantenos. 2016. https://aclanthology.org/L16-1432 Discourse structure and dialogue acts in multiparty dialogue: the STAC corpus . In Proceedings of the Tenth International Conference on Language Resources and Evaluation ( LREC '16) , pages 2721--2727, Portoro z , Slovenia. European ...
2016
-
[5]
Katherine Atwell, Mert Inan, Anthony B Sicilia, and Malihe Alikhani. 2024. Combining discourse coherence with large language models for more inclusive, equitable, and robust task-oriented dialogue. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 3538--3552
work page 2024
-
[6]
Jonathan Back, Duncan P. Brumby, and Anna L. Cox. 2010. https://doi.org/10.1145/1753846.1754054 Locked-out: investigating the effectiveness of system lockouts to reduce errors in routine tasks . In CHI '10 Extended Abstracts on Human Factors in Computing Systems
-
[7]
Anouck Braggaar, Christine Liebrecht, Emiel van Miltenburg, and Emiel Krahmer. 2024. http://arxiv.org/abs/2312.13871 Evaluating task-oriented dialogue systems: A systematic review of measures, constructs and their operationalisations
arXiv 2024
-
[8]
Duncan P Brumby, Anna L Cox, Jonathan Back, and Sandy JJ Gould. 2013. Recovering from an interruption: Investigating speed- accuracy trade-offs in task resumption behavior. Journal of Experimental Psychology: Applied, 19(2):95
work page 2013
Show all 75 references
-
[9]
Zana Bu c inca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021. To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making. Proceedings of the ACM on Human-computer Interaction
2021
-
[10]
Pawe Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, I \ n igo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Ga s i \'c . 2018. https://doi.org/10.18653/v1/D18-1547 M ulti WOZ - a large-scale multi-domain W izard-of- O z dataset for task-oriented dialogue modelling . In P...
2018 doi
-
[11]
Paul Cairns, Anna Cox, and A Imran Nordin. 2014. Immersion in digital games: review of gaming experience research. Handbook of digital games
2014
-
[12]
Ana Caraban, Evangelos Karapanos, Daniel Gon c alves, and Pedro Campos. 2019. 23 ways to nudge: A review of technology-mediated nudging in human-computer interaction. In Proceedings of the 2019 CHI conference on human factors in computing systems
2019
-
[13]
Birte Carlmeyer, Simon Betz, Petra Wagner, Britta Wrede, and David Schlangen. 2018. https://doi.org/10.1145/3173386.3176992 The Hesitating Robot - Implementation and First Impressions . In ACM Conferences , pages 77--78. Association for Computing Machinery, New York, NY, USA
2018
-
[14]
Cecchinato, Anna L
Marta E. Cecchinato, Anna L. Cox, and Jon Bird. 2015. https://doi.org/10.1145/2702123.2702537 Working 9-5? professional differences in email and boundary management practices . In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems
2015
-
[15]
Chai, Rui Fang, Changsong Liu, and Lanbo She
Joyce Y. Chai, Rui Fang, Changsong Liu, and Lanbo She. 2016. https://doi.org/10.1609/AIMAG.V37I4.2684 Collaborative language grounding toward situated human-robot dialogue . AI Mag. , 37(4):32--45
2016 doi
-
[16]
positive friction
Zeya Chen and Ruth Schmidt. 2024. http://arxiv.org/abs/2402.09683 Exploring a behavioral model of "positive friction" in human-ai interaction
2024 arXiv
-
[17]
Katherine M Collins, Valerie Chen, Ilia Sucholutsky, Hannah Rose Kirk, Malak Sadek, Holli Sargeant, Ameet Talwalkar, Adrian Weller, and Umang Bhatt. 2024. Modulating language model experiences through frictions. CoRR
2024
-
[18]
Marc-Alexandre C \^o t \'e , Akos K \'a d \'a r, Xingdi Yuan, Ben Kybartas, Tavian Barnes, Emery Fine, James Moore, Matthew Hausknecht, Layla El Asri, Mahmoud Adada, et al. 2019. Textworld: A learning environment for text-based games. In Computer Games: 7th Workshop, CGW 2018,...
2019
-
[19]
Cox, Sandy J
Anna L. Cox, Sandy J. J. Gould, Marta E. Cecchinato, Ioanna Iacovides, and Ian Renfree. 2016. https://doi.org/10.1145/2851581.2892410 Design Frictions for Mindful Interactions: The Case for Microboundaries . In ACM Conferences , pages 1389--1397. Association for Computing Mach...
2016
-
[20]
Claudio De Stefano, Carlo Sansone, and Mario Vento. 2000. To reject or not to reject: that is the question-an answer in case of neural classifiers. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 30(1):84--94
2000
-
[21]
Diekhof and Oliver Gruber
Esther K. Diekhof and Oliver Gruber. 2010. https://doi.org/10.1523/JNEUROSCI.4690-09.2010 When Desire Collides with Reason: Functional Interactions between Anteroventral Prefrontal Cortex and Nucleus Accumbens Underlie the Human Ability to Resist Impulsive Desires . J. Neurosc...
2010 doi
-
[22]
Vardhan Dongre, Xiaocheng Yang, Emre Can Acikgoz, Suvodip Dey, Gokhan Tur, and Dilek Hakkani-Tür. 2024. http://arxiv.org/abs/2411.00927 Respact: Harmonizing reasoning, speaking, and acting towards building large language model-based conversational ai agents
2024 arXiv
-
[23]
Engelhardt, Karl G
Paul E. Engelhardt, Karl G. D. Bailey, and Fernanda Ferreira. 2006. https://doi.org/10.1016/j.jml.2005.12.009 Do speakers and listeners observe the Gricean Maxim of Quantity? Journal of Memory and Language, 54(4):554--573
2006 doi
-
[24]
Mihail Eric, Rahul Goel, Shachi Paul, Abhishek Sethi, Sanchit Agarwal, Shuyang Gao, Adarsh Kumar, Anuj Goyal, Peter Ku, and Dilek Hakkani-Tur. 2020. Multiwoz 2.1: A consolidated multi-domain dialogue dataset with state corrections and state tracking baselines. In Proceedings o...
2020
-
[25]
Jonathan Ericson. 2022. Reimagining the role of friction in experience design. Journal of User Experience, 17(4)
2022
-
[26]
Jonathan St . B. T. Evans. 2003. https://doi.org/10.1016/j.tics.2003.08.012 In two minds: dual-process accounts of reasoning . Trends in Cognitive Sciences, 7(10):454--459
2003 doi
-
[27]
Martin Fishbein and Icek Ajzen. 2011. Predicting and changing behavior: The reasoned action approach. Psychology press
2011
-
[28]
Kristina Lundholm Fors. 2015. Production and perception of pauses in speech. Ph.D. thesis, Department of Philosophy, Linguistics, and Theory of Science, University of …
2015
-
[29]
Christian Geishauser, Carel van Niekerk, Nurul Lubis, Hsien-chin Lin, Michael Heck, and Shutong Feng. 2024. https://doi.org/10.1109/TASLP.2024.3385289 Learning With an Open Horizon in Ever-Changing Dialogue Circumstances . IEEE/ACM Trans. Audio Speech Lang. Process., 32:2352--2366
2024
-
[30]
a s and Johan Redstr o \
Lars Halln a \" a s and Johan Redstr o \" o m. 2001. https://doi.org/10.1007/PL00000019 Slow Technology Designing for Reflection . Personal Ub. Comp., 5(3):201--212
2001 doi
-
[31]
Julian Hough and David Schlangen. 2017. https://doi.org/10.1145/2909824.3020214 It's Not What You Do, It's How You Do It: Grounding Uncertainty for a Simple Robot . In ACM Conferences , pages 274--282. Association for Computing Machinery, New York, NY, USA
2017
-
[32]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Info...
2023
-
[33]
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088
2024 arXiv
-
[34]
i'm not sure, but
Sunnie SY Kim, Q Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, and Jennifer Wortman Vaughan. 2024. " i'm not sure, but...": Examining the impact of large language models' uncertainty expression on user reliance and trust. In The 2024 ACM Conference on Fairness, Accountabil...
2024
-
[35]
Chei Sian Lee, Dion Hoe-Lian Goh, Alton YK Chua, and Rebecca P Ang. 2010. Indagator: Investigating perceived gratifications of an application that blends mobile content sharing with gameplay. Journal of the American Society for Information Science and Technology, 61(6):1244--1257
2010
-
[36]
Anna Lembke. 2023. Dopamine nation: Finding balance in the age of indulgence. Dutton, an imprint of Penguin Random House LLC
2023
-
[37]
Stephan J Lemmer, Anhong Guo, and Jason J Corso. 2023. Human-centered deferred inference: Measuring user interactions and setting deferral criteria for human-ai teams. In Proceedings of the 28th International Conference on Intelligent User Interfaces, pages 681--694
2023
-
[38]
Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016. https://doi.org/10.18653/v1/D16-1127 Deep reinforcement learning for dialogue generation . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages ...
2016 doi
-
[39]
Xinmeng Li, Wansen Wu, Long Qin, and Quanjun Yin. 2021. http://arxiv.org/abs/2108.01369 How to evaluate your dialogue models: A review of approaches
2021 arXiv
-
[40]
Xiujun Li, Yun-Nung Chen, Lihong Li, Jianfeng Gao, and Asli Celikyilmaz. 2017. https://aclanthology.org/I17-1074/ End-to-end task-completion neural dialogue systems . In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Pap...
2017
-
[41]
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. https://doi.org/10.18653/v1/D16-1230 How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation . In Proceed...
2016 doi
-
[42]
are you really sure?
Shuai Ma, Xinru Wang, Ying Lei, Chuhan Shi, Ming Yin, and Xiaojuan Ma. 2024. “are you really sure?” understanding the effects of human self-confidence calibration in ai-assisted decision making. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1--20
2024
-
[43]
Lars Malmqvist. 2024. Sycophancy in large language models: Causes and mitigations. arXiv preprint arXiv:2411.15287
2024 arXiv
-
[44]
Roland Mangold and Rupert Pobel. 1988. https://doi.org/10.1177/0261927X8800700403 Informativeness and Instrumentality in Referential Communication . Journal of Language and Social Psychology, 7(3-4):181--191
1988 doi
-
[45]
Rudnicky
Matthew Marge and Alexander I. Rudnicky. https://doi.org/10.1109/ROMAN.2013.6628486 Towards evaluating recovery strategies for situated grounding problems in human-robot dialogue . In 2013 IEEE RO-MAN , pages 26--29. IEEE
2013
-
[46]
McClure, Keith M
Samuel M. McClure, Keith M. Ericson, David I. Laibson, George Loewenstein, and Jonathan D. Cohen. 2007. https://doi.org/10.1523/JNEUROSCI.4246-06.2007 Time discounting for primary rewards . J. Neurosci., 27(21):5796--5804
2007 doi
-
[47]
Hussein Mozannar, Hunter Lang, Dennis Wei, Prasanna Sattigeri, Subhro Das, and David Sontag. 2023. Who should predict? exact algorithms for learning to defer to humans. In International conference on artificial intelligence and statistics. PMLR
2023
-
[48]
Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gokhan Tur, and Dilek Hakkani-Tur. 2021. http://arxiv.org/abs/2110.00534 Teach: Task-driven embodied agents that chat
2021 arXiv
-
[49]
Joon Sung Park, Rick Barber, Alex Kirlik, and Karrie Karahalios. 2019. A slow algorithm improves users' assessments of the algorithm's accuracy. Proceedings of the ACM on Human-Computer Interaction, 3(CSCW):1--15
2019
-
[50]
Piantadosi
Steven T. Piantadosi. 2014. https://doi.org/10.3758/s13423-014-0585-6 Zipf ' s word frequency law in natural language: A critical review and future directions . Psychonomic bulletin & review , 21(5):1112
2014 doi
-
[51]
James Pierce. 2014. Undesigning interaction. Interactions, 21(4):36--39
2014
-
[52]
Beatrice Szczepek Reed. 2017. Analysing conversation: An introduction to prosody. Bloomsbury Publishing
2017
-
[53]
Craige Roberts. 2012. https://doi.org/10.3765/sp.5.6 Information Structure: Towards an integrated formal theory of pragmatics . S & P , 5:6:1--69
2012 doi
-
[54]
Nicolas Ruiz, Gabriela Molina Le \'o n, and Hendrik Heuer. 2024. Design frictions on social media: Balancing reduced mindless scrolling and user satisfaction. In Proceedings of Mensch und Computer 2024, pages 442--447
2024
-
[55]
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al. 2023. Towards understanding sycophancy in language models. arXiv preprint arXiv:2310.13548
2023 arXiv
-
[56]
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020 a . Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF conference on computer vision and...
2020
-
[57]
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre C \^o t \'e , Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2020 b . Alfworld: Aligning text and embodied environments for interactive learning. arXiv preprint arXiv:2010.03768
2020 arXiv
-
[58]
Anthony Sicilia and Malihe Alikhani. 2024. Evaluating theory of (an uncertain) mind: Predicting the uncertain beliefs of others in conversation forecasting. arXiv preprint arXiv:2409.14986
2024 arXiv
-
[59]
Anthony Sicilia, Yuya Asano, Katherine Atwell, Qi Cheng, Dipunj Gupta, Sabit Hassan, Mert Inan, Jennifer Nwogu, Paras Sharma, and Malihe Alikhani. 2023. Isabel: An inclusive and collaborative task-oriented dialogue system
2023
-
[60]
Anthony Sicilia, Mert Inan, and Malihe Alikhani. 2024. Accounting for sycophancy in language model uncertainty estimation. arXiv preprint arXiv:2410.14746
2024 arXiv
-
[61]
Frank Soboczenski, Paul Cairns, and Anna L Cox. 2013. Increasing accuracy by decreasing presentation quality in transcription tasks. In Human-Computer Interaction--INTERACT
2013
-
[62]
Stalnaker
Robert C. Stalnaker. 1978. https://doi.org/10.1163/9789004368873_013 Assertion . In Pragmatics , pages 315--332. Brill, Leiden, The Netherlands
1978 doi
-
[63]
Manfred Stede and Arne Neumann. 2014. http://www.lrec-conf.org/proceedings/lrec2014/pdf/579_Paper.pdf P otsdam commentary corpus 2.0: Annotation for discourse research . In Proceedings of the Ninth International Conference on Language Resources and Evaluation ( LREC '14) , pag...
2014
-
[64]
Weiwei Sun, Shuo Zhang, Krisztian Balog, Zhaochun Ren, Pengjie Ren, Zhumin Chen, and Maarten de Rijke. 2021. Simulating user satisfaction for the evaluation of task-oriented dialogue systems. In Proceedings of the 44th International ACM SIGIR Conference on Research and Develop...
2021
-
[65]
D. Tannen. 1989. https://books.google.com/books?id=MuaMgeJ4FF8C Talking Voices: Repetition, Dialogue and Imagery in Conversational Discourse . Studies in Interactional Sociolinguistics. Cambridge University Press
1989
-
[66]
Tourtouri, Francesca Delogu, and Matthew W
Elli N. Tourtouri, Francesca Delogu, and Matthew W. Crocker. 2021. https://doi.org/10.1111/cogs.13071 Rational Redundancy in Referring Expressions: Evidence from Event-related Potentials . Cognitive Science, 45(12):e13071
2021 doi
-
[67]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971
2023 arXiv
-
[68]
Xuewei Wang, Weiyan Shi, Richard Kim, Yoojung Oh, Sijia Yang, Jingwen Zhang, and Zhou Yu. 2019. https://doi.org/10.18653/v1/P19-1566 Persuasion for good: Towards a personalized persuasive dialogue system for social good . In Proceedings of the 57th Annual Meeting of the Associ...
2019 doi
-
[69]
Adam Waytz, Kurt Gray, Nicholas Epley, and Daniel M Wegner. 2010. Causes and consequences of mind perception. Trends in cognitive sciences, 14(8):383--388
2010
-
[70]
Deanna Wilkes-Gibbs and Herbert H. Clark. 1992. https://doi.org/10.1016/0749-596X(92)90010-U Coordinating beliefs in conversation . Journal of Memory and Language, 31(2):183--194
1992 doi
-
[71]
Heng-Da Xu, Xian-Ling Mao, Puhai Yang, Fanshu Sun, and Heyan Huang. 2024. https://doi.org/10.18653/v1/2024.acl-long.152 Rethinking task-oriented dialogue systems: From complex modularity to zero-shot autonomous agent . In Proceedings of the 62nd Annual Meeting of the Associati...
2024 doi
-
[72]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. https://openreview.net/forum?id=WE_vluYUL-X React: Synergizing reasoning and acting in language models . In The Eleventh International Conference on Learning Representations
2023
-
[73]
Brigitte Zellner. 1994. Pauses and the temporal structure of speech. In Zellner, B.(1994). Pauses and the temporal structure of speech, in E. Keller (Ed.) Fundamentals of speech synthesis and speech recognition.(pp. 41-62). Chichester: John Wiley., pages 41--62. John Wiley
1994
-
[74]
Tong Zhang, Chen Huang, Yang Deng, Hongru Liang, Jia Liu, Zujie Wen, Wenqiang Lei, and Tat-Seng Chua. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.26 Strength lies in differences! improving strategy planning for non-collaborative dialogues via diversified user simulation ...
2024 doi
-
[75]
Kaitlyn Zhou, Jena Hwang, Xiang Ren, and Maarten Sap. 2024. https://doi.org/10.18653/v1/2024.acl-long.198 Relying on the unreliable: The impact of language models' reluctance to express uncertainty . In Proceedings of the 62nd Annual Meeting of the Association for Computationa...
2024 doi
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.