A new multimodal dataset of simulated operating-room team dialogues shows that existing LLMs and fine-tuned models reach only about 51% macro F1 on team reflection behavior classification, indicating substantial room for improvement.
Annals of surgery, 247(4):699–706
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation
A new multimodal dataset of simulated operating-room team dialogues shows that existing LLMs and fine-tuned models reach only about 51% macro F1 on team reflection behavior classification, indicating substantial room for improvement.