{"id":"72bd6b46-1330-4fbf-ae3a-3d11227b3603","arxiv_id":"2605.26747","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"MeDial-Speech provides 111+ hours of spoken medical dialogues from robot-patient and doctor-patient interactions across four conditions, with a 20-option sentence selection benchmark where Claude Sonnet 4 reaches 71-75% accuracy.","lead":"The paper releases MeDial-Speech, a 111-hour speech dataset of robot-patient and doctor-patient dialogues covering four medical conditions, plus a sentence-selection benchmark on three LLMs. A smart generalist might read it to understand what data is now available for training medical dialogue systems.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Representativeness of robot/doctor-patient dialogues for real medical consultations is unverified","rationale":"The reader's weakest_assumption directly identifies the load-bearing condition for the dataset's claimed utility. No stronger internal inconsistency (e.g., in the reported accuracies or overconfidence finding) appears from the abstract; the realism gap is the point that must be true for the headline value proposition to hold. Full-text details on collection would be the natural next check, but the concern remains load-bearing regardless.","tokens_in":1769,"tokens_out":344,"duration_ms":22355,"concrete_test":"In the methods or data-collection section, locate the protocol for dialogue elicitation (e.g., participant recruitment, use of scripts, recording setup). If it shows heavy scripting or non-clinical participants without a realism check (e.g., comparison to real consultation transcripts), re-run the sentence-selection benchmark on a matched real-consultation subset; a >15-point accuracy drop would confirm the representativeness gap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that MeDial-Speech enables training/evaluating Med-AIs for consultations because it contains 111+ hours of speech from robot-patient and doctor-patient dialogues across four conditions, collected in realistic environments. This requires the dialogues to be sufficiently representative of actual clinical interactions. The abstract provides no evidence on participant type (real patients vs. actors), elicitation method (scripted vs. free-form), presence of time pressure or physical exams, or any validation against real consultation corpora. If the data are primarily scripted role-play, downstream utility for Med-AI training collapses even if the benchmark numbers hold.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces MeDial-Speech, a speech dataset of 111+ hours collected from robot-patient and doctor-patient dialogues in realistic environments, covering four conditions (Lewy body dementia, heart failure, shoulder pain, angina). It proposes a sentence-selection benchmark (20 options) to evaluate LLMs on medical dialogues, reporting that Claude Sonnet 4 achieves the highest accuracy (71.1% on manual transcriptions, 74.7% on automatic transcriptions) while all tested models (including GPT-5 mini and DeepSeek-V3) exhibit high overconfidence in their predictions. The dataset is released freely for non-commercial use via Hugging Face.","tokens_in":1881,"tokens_out":548,"duration_ms":33198,"significance":"A large-scale spoken medical dialogue dataset would address a clear resource gap for training and evaluating spoken Med-AIs. The sentence-selection benchmark supplies a reproducible task and initial LLM comparison, and the public release itself is a concrete contribution. However, the significance is conditional on the dialogues being sufficiently representative of actual clinical interactions; without supporting details this remains an open question.","major_comments":[{"comment":"Abstract: The claim that the dataset 'can carry out consultations with patients' and supports training/evaluating Med-AIs rests on the unverified assumption that the collected dialogues are representative of real medical consultations. No information is supplied on participant type (real patients vs. actors), elicitation method (scripted vs. free-form), presence of time pressure or physical exams, or any quantitative validation against existing real consultation corpora.","section":"Abstract"},{"comment":"Abstract (benchmark results): The reported accuracies of 71.1% and 74.7% and the statement that 'all LLMs are highly overconfident' are given without error bars, confidence intervals, or statistical significance tests comparing models or conditions. This prevents assessment of whether the observed differences and overconfidence pattern are reliable.","section":"Abstract"},{"comment":"Abstract: The benchmark relies on both manual and automatic transcriptions, yet no details are provided on transcription accuracy, inter-annotator agreement, or word-error-rate of the automatic system. These factors directly affect the validity of the 71.1% and 74.7% figures.","section":"Abstract"}],"minor_comments":[{"comment":"Abstract: The model identifier 'GPT-5 mini' is non-standard; please specify the exact model versions and checkpoints used.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. We address each major comment below and will revise the manuscript to provide the requested details.","responses":[{"response":"The abstract summarizes the collection process but omits granular methodological details due to space limits. The full manuscript states that dialogues were collected in realistic environments from robot-patient and doctor-patient interactions, but we agree that explicit information on participant types, elicitation procedures, time pressure, physical exams, and validation against real corpora would better support claims of representativeness. We will expand the Methods section in the revision to include these details.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The claim that the dataset 'can carry out consultations with patients' and supports training/evaluating Med-AIs rests on the unverified assumption that the collected dialogues are representative of real medical consultations. No information is supplied on participant type (real patients vs. actors), elicitation method (scripted vs. free-form), presence of time pressure or physical exams, or any quantitative validation against existing real consultation corpora."},{"response":"The abstract reports point accuracies and the overconfidence observation without accompanying statistical measures. We agree that error bars, confidence intervals, and significance tests are needed to assess reliability. We will add these analyses to the results section and update the abstract in the revised manuscript.","revision_made":"yes","referee_comment":"[Abstract] Abstract (benchmark results): The reported accuracies of 71.1% and 74.7% and the statement that 'all LLMs are highly overconfident' are given without error bars, confidence intervals, or statistical significance tests comparing models or conditions. This prevents assessment of whether the observed differences and overconfidence pattern are reliable."},{"response":"Transcription details were not included in the abstract. We will add information on manual transcription inter-annotator agreement and the word-error-rate of the automatic system to the revised manuscript to support the reported benchmark figures.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The benchmark relies on both manual and automatic transcriptions, yet no details are provided on transcription accuracy, inter-annotator agreement, or word-error-rate of the automatic system. These factors directly affect the validity of the 71.1% and 74.7% figures."}],"tokens_in":1510,"tokens_out":515,"duration_ms":37679,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core contribution is the public release of MeDial-Speech: 111+ hours of spoken dialogues covering robot-patient and doctor-patient interactions for Lewy body dementia, heart failure, shoulder pain, and angina. It also includes a sentence-selection benchmark where Claude Sonnet 4 reaches 71.1% on manual transcripts and 74.7% on automatic ones, with a side note that all three tested LLMs are overconfident.\n\nThe data release itself is the part that adds something concrete. Prior work on medical dialogue corpora exists, but this specific mix of robot and human interactions in spoken form for these four conditions, made available on Hugging Face, is new. The benchmark numbers are reported clearly enough to be reproducible at the level described.\n\nThe main limitation is the lack of detail on how representative the dialogues actually are. The abstract mentions collection in realistic environments but gives no information on whether participants were real patients or actors, whether the exchanges were scripted or free-form, or how the data compares to actual clinical recordings. If the material is mostly role-play without time pressure or physical exams, its value for training systems that handle real consultations drops sharply. The sentence-selection task is also quite constrained; it does not test full dialogue management or error recovery.\n\nNo error bars or statistical tests accompany the LLM results, and the paper does not address transcription accuracy or consent procedures in the provided summary. These are standard expectations for a data paper.\n\nThis work is mainly for groups already building or evaluating spoken medical dialogue systems who need additional speech data. It is coherent on its own terms as a data contribution and shows honest engagement with the task of releasing the corpus. It deserves peer review so that referees can examine the collection protocol and assess whether the realism claim holds up. I would send it out rather than desk reject.","headline":"A straightforward dataset release of spoken medical dialogues with a narrow LLM benchmark, but the key claim of usefulness for real consultations rests on unverified representativeness.","tokens_in":2360,"tokens_out":448,"would_cite":false,"duration_ms":27496,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A new dataset of 111+ hours of robot-patient and doctor-patient medical dialogues enables benchmarking of LLMs on consultation sentence selection.","keywords":["medical dialogues","spoken language processing","robot-patient dialogues","doctor-patient dialogues","large language models","sentence selection benchmark","health conditions","MeDial-Speech"],"falsifier":"A controlled test in which models trained or fine-tuned on MeDial-Speech perform no better than general-purpose models when evaluated on actual patient consultations with human doctors would indicate the dataset does not deliver the claimed training or evaluation value.","tokens_in":2647,"feed_emoji":"🩺","tokens_out":682,"duration_ms":27403,"temperature":0.7,"pith_summary":"This paper introduces MeDial-Speech, a speech dataset gathered from realistic robot-patient and doctor-patient dialogues covering four health conditions: Lewy body dementia, heart failure, shoulder pain, and angina. The collection totals over 111 hours without augmentation and is positioned for training and evaluating AI systems that perform medical consultations. The authors define a sentence-selection benchmark with 20 options per turn and test three large language models, showing Claude Sonnet 4 reaches 71.1% accuracy on manual transcriptions and 74.7% on automatic transcriptions while all models remain highly overconfident in their predictions. The work targets the open problem of applying LLMs effectively to spoken medical interactions by supplying a dedicated, publicly available resource.","feed_headline":"111+ hours of medical dialogues dataset released for AI","feed_subtitle":"MeDial-Speech covers four conditions and benchmarks LLMs on sentence selection with top model at 74.7 percent accuracy.","key_machinery":"The MeDial-Speech dataset of spoken medical dialogues together with its sentence-selection benchmark using 20 options per turn.","core_discovery":"The paper establishes MeDial-Speech as a resource of 111+ hours of speech from robot-patient and doctor-patient dialogues across four health conditions, paired with a sentence-selection benchmark that identifies Claude Sonnet 4 as the strongest of three tested LLMs at 71.1% accuracy on manual transcriptions and 74.7% on automatic transcriptions, while documenting that all evaluated models exhibit high overconfidence irrespective of whether their selected sentence is correct.","pith_inferences":["The resource could support development of AI assistants that handle spoken medical exchanges more reliably than current general models.","The documented overconfidence pattern suggests a broader need for uncertainty-aware methods when applying LLMs to high-stakes conversational domains.","Robot-patient dialogues within the collection may enable separate study of human-robot medical communication patterns."],"forward_implications":["The dataset supplies training material specifically for spoken medical consultation tasks.","Automatic speech recognition transcriptions can serve as a viable alternative to manual ones for model evaluation in this domain.","Large language models require additional calibration techniques before deployment in medical dialogue settings due to consistent overconfidence.","Both robot-patient and doctor-patient interaction data become available for spoken language processing research."],"fun_headline_variants":["MeDial-Speech offers 111+ hours of robot and doctor patient dialogues","Benchmark shows Claude Sonnet 4 at 74.7% on MeDial-Speech transcriptions","MeDial-Speech dataset spans four medical conditions for AI training","All LLMs overconfident regardless of correctness in MeDial-Speech test","Robot-patient dialogues in MeDial-Speech for spoken language processing"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The collected dialogues are representative enough of real medical consultations to be useful for training and evaluating medical AI systems.","fun_headline_variants_meta":{"raw":{"variants":["MeDial-Speech offers 111+ hours of robot and doctor patient dialogues","Benchmark shows Claude Sonnet 4 at 74.7% on MeDial-Speech transcriptions","MeDial-Speech dataset spans four medical conditions for AI training","All LLMs overconfident regardless of correctness in MeDial-Speech test","Robot-patient dialogues in MeDial-Speech for spoken language processing"]},"model":"grok-4.3","cost_usd":0.005993,"raw_usage":{"total_tokens":2774,"prompt_tokens":700,"num_sources_used":0,"completion_tokens":98,"cost_in_usd_ticks":59928000,"prompt_tokens_details":{"text_tokens":700,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1976,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":700,"tokens_out":98,"duration_ms":18072,"temperature":1.0,"reasoning_tokens":1976,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T17:40:51.944051+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled test in which models trained or fine-tuned on MeDial-Speech perform no better than general-purpose models when evaluated on actual patient consultations with human doctors would indicate the dataset does not deliver the claimed training or evaluation value.","supporting_citations":[],"review_version":1}