{"id":"6eb1a650-7132-4de3-8101-d2ff8538b995","arxiv_id":"1908.10023","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new hierarchical multi-label annotation scheme for human-machine spoken dialog, with a 24K-utterance corpus and a dialog act classifier reaching 79% micro-F1.","lead":"The paper introduces MIDAS, a new set of labels for tagging what people are trying to do when they talk to a voice assistant or social bot, and releases a 24,000-utterance annotated dataset built from real conversations. It also trains a model that tags these utterances with 79% F1, showing that the scheme can be learned automatically.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Train/test split may mix turns from the same conversation, which could inflate the reported 0.79 F1 and undermine the usability claim.","rationale":"The paper's central claim is that MIDAS is usable in practice: labels can be assigned reliably (kappa=0.94) and predicted automatically (F1=0.79). The reader's weakest assumption focuses on ASR segmentation, which the paper itself acknowledges can hurt prediction. I agree that is a real limitation, but the more load-bearing issue is the lack of a stated conversation-level train/test split. If segments from the same dialogue appear in both training and testing, the context features (previous system unit, previous user unit) let the model exploit conversation-level regularities that would not be available for new conversations. This directly affects the 0.79 F1, which is the primary evidence for automatic predictability. The concern is concrete and testable: rerunning with a conversation-level split would settle it. I do not see a fatal flaw in the annotation scheme itself; the taxonomy is described in detail, the data is shared, and the kappa pilot is a reasonable first step. The verdict should remain CONDITIONAL, but the condition should include re-evaluating with a conversation-level split in addition to the segmentation and statistical-significance issues already noted.","tokens_in":11820,"tokens_out":2575,"duration_ms":28350,"concrete_test":"Re-run the dialog act prediction experiments with a conversation-level split: partition the 468 conversations so that no conversation appears in both training and test sets, preserving the 10.3K/2.6K user-segment ratio as closely as possible, and report micro-F1 for BERT-text, BERT F-text, and BERT F-DA+text. If the F1 drops materially (for example, below 0.75) or the context gain shrinks, the reported 0.79 overstates the practical predictability of MIDAS labels.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 6 reports 10.3K user segments for training and 2.6K for testing, with the remaining 11.1K machine segments used as context, but it does not state that the split is conversation-level. Since the context representation in Section 5.2 includes the previous system unit and previous user unit from the same dialogue, a random segment-level split can place a test utterance in the same conversation as training utterances. The model can then exploit conversation-specific lexical patterns or near-duplicate adjacent turns, inflating the central F1=0.79 result. The paper's own error analysis notes that context helps (adding text context improves F1 from 71.30 to 79.11), but if the context is drawn from training conversations at test time, part of this gain may be leakage rather than generalizable dialog act understanding. This is more directly load-bearing than the ASR-segmentation concern: even with perfect segmentation, an overlapping split would make the headline number optimistic. The paper also reports that several transfer-learning improvements are not statistically significant, yet the base model comparison itself may be affected by the same leakage. The kappa=0.94 agreement, while measured only on a 1,185-utterance pilot, is a separate reliability claim and does not resolve the evaluation concern.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MIDAS, a hierarchical multi-label dialog act annotation scheme designed for open-domain human-machine spoken conversations. The authors describe the scheme's structure, collect and annotate 24K utterances from real Gunrock (Alexa Prize) conversations, report an inter-annotator agreement of kappa = 0.94 on a pilot subset, and train multi-label dialog act prediction models with transfer learning, achieving a micro-F1 score of 0.79. The paper also releases the annotated data and trained models.","tokens_in":12081,"tokens_out":3634,"duration_ms":33786,"significance":"If the results are reliable, MIDAS would be a useful resource for dialog system development and discourse analysis. The paper has notable strengths: it uses real user-system interactions rather than scripted data, reports results averaged over six random seeds, provides per-tag counts and annotated examples in the appendix, and publicly releases the data and code. However, two evaluation concerns—the potential for training/test leakage through the context representation and the reliance on imperfect automatic segmentation—mean that the headline numbers need further validation before the practical applicability of the scheme is fully established.","major_comments":[{"comment":"The paper reports 10.3K training and 2.6K test user segments but does not state that the split is at the conversation level. Because the context representation in Section 5.2 includes the previous system unit and previous user unit from the same dialogue, a random segment-level split can place a test utterance in the same conversation as training utterances. The model can then exploit conversation-specific lexical patterns or near-duplicate adjacent turns, inflating the central F1=0.79 result. The authors must either specify a conversation-level split and re-run the experiments, or provide a controlled experiment showing that overlapping contexts do not affect the reported results.","section":"Section 6, Setting"},{"comment":"All dialog act annotation and predictions are performed on automatic segmentation results from a model with 84.43% micro-F1, and the paper itself notes that some incorrectly segmented units led to inaccurate dialog act prediction. Consequently, the reported kappa and F1 characterize agreement and performance on machine-produced units rather than on true utterance boundaries. The authors should quantify how segmentation errors affect tag distribution and model performance, or at least delimit the claims to the automatic-segmentation setting.","section":"Section 4 and Section 7"},{"comment":"The inter-annotator agreement kappa = 0.94 is computed on only 1,185 pilot utterances, after which the two annotators annotated the remaining 24K utterances separately. Thus, the reliability of the final corpus is not directly measured. The paper should provide at least a small random sample of double-annotated utterances from the main annotation phase, or explicitly acknowledge this as a limitation.","section":"Section 4"}],"minor_comments":[{"comment":"The phrase \"inter-annotated agreement\" should be \"inter-annotator agreement\" (also appears in Section 4).","section":"Introduction"},{"comment":"The text says \"two annotators reached 0.94 in Kapa\"; this should be \"kappa\".","section":"Section 2"},{"comment":"The example contains \"the great gastby\", which should be \"the great gatsby\".","section":"Table 1"},{"comment":"The word \"syntatic\" should be \"syntactic\", and \"ﬁne-turning\" should be \"fine-tuning\".","section":"Section 8"},{"comment":"The phrase \"complimentary information\" should be \"complementary information\".","section":"Section 7"},{"comment":"The maximum of two tags per utterance is a somewhat arbitrary constraint; a more detailed discussion of its trade-off would be helpful.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The training/test split concern is the most consequential issue. If the authors can demonstrate a conversation-level split or show that the results are unchanged under such a split, the paper would likely be suitable for publication. Given the public release of the data and models, this should be verifiable by the reviewers. Please ensure the authors address this directly, as it bears on the validity of the headline F1 score."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid resource paper. MIDAS gives the community a dialog act scheme built for machine-directed speech, a 24K-utterance corpus from real Gunrock conversations, and a working classifier at roughly 79% F1. Worth engaging, but the evaluation section has a train/test split ambiguity that could make the reported number optimistic, and the reliability evidence is thinner than the headline suggests.\n\nWhat's actually new: the MIDAS taxonomy itself. It adapts SWBD-DAMSL and ISO/DIT++ to the human-machine setting, with a two-tree hierarchy (semantic and functional request), 23 leaf tags, and multi-label annotation capped at two. The invalid-command and nonsense tags are genuinely useful for ASR-heavy user input. The corpus is real and the appendix gives per-tag counts and examples, which makes the scheme fairly easy to inspect. The classifier experiments are standard: Bi-LSTM and BERT variants, context from previous turns, six-seed averaging. The paper also honestly says several transfer-learning gains are not statistically significant, which is more than many papers do.\n\nSoft spots, in order of importance. First, Section 6 reports 10.3K user segments for training and 2.6K for testing, but never says the split is conversation-level. Since the context representation uses the previous system unit and previous user unit from the same dialogue, a segment-level split could leak conversation-specific patterns into test. The context gain (71.30 to 79.11 F1) could partly be an artifact of that leakage. This is the most load-bearing issue and is easy to fix: state the split procedure, or rerun with a conversation-level split. Second, the kappa=0.94 is computed on just 1,185 pilot utterances; the rest of the 24K was annotated separately by the two annotators. That is a thin reliability base. Third, all annotation and prediction happen on automatic ASR segmentation with 84.43% micro-F1, and the paper itself notes that incorrect segmentation leads to dialog act errors. So the tag distribution and F1 are partly contingent on the segmenter.\n\nNone of this kills the paper. The scheme and corpus are useful even if the classifier number moves under a proper split. The authors need to document the split, ideally report conversation-level results, and be clearer about the scope of the reliability claim.\n\nWho this is for: dialog system builders, especially social-bot and chitchat people, and anyone working on annotation schemes for machine-directed speech. The citation pattern looks fine. I'd send it to peer review with a request for revision, not a desk reject.","headline":"MIDAS is a legitimate resource paper for human-machine dialog act annotation, but the evaluation needs a documented conversation-level split before the headline F1 is fully trustable.","tokens_in":12624,"tokens_out":2711,"would_cite":true,"duration_ms":27468,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A dialog act annotation scheme designed for open-domain human-machine spoken conversation reaches 94% annotator agreement and 0.79 F1 on automatic prediction.","keywords":["dialog act annotation","human-machine conversation","multi-label classification","open-domain dialogue","hierarchical annotation scheme","spoken dialogue system","BERT transfer learning","utterance segmentation"],"falsifier":"Annotate a held-out sample of user turns twice: once after automatic ASR segmentation and once after manual human segmentation, then measure MIDAS label agreement on the two versions. If classifier F1 or annotator kappa on manually segmented units falls clearly below the reported 0.79 and $\\kappa = 0.94$, the usability claim is undermined.","tokens_in":11595,"feed_emoji":"💬","tokens_out":9743,"duration_ms":83248,"temperature":0.7,"pith_summary":"Existing dialog act schemes were built for human-human dialogue, in which both sides understand language perfectly. This paper argues that machines, which understand less, need an annotation vocabulary matched to their limitations, and proposes MIDAS, a hierarchical multi-label dialog act scheme for open-domain human-machine spoken conversation. The authors annotate 24K utterances from real social-bot conversations: two annotators reach $\\kappa = 0.94$, and a BERT-based transfer-learning classifier reaches 0.79 F1. The claim is that machine-directed speech differs enough from human-human talk to deserve its own tags, and that those tags can be defined, reliably assigned, and automatically predicted.","feed_headline":"Dialog act scheme for human-machine chat hits 94% agreement, 0.79 F1","feed_subtitle":"Old schemes trained on human-human talk fail on bots; MIDAS labels real assistant conversations and predicts intent.","key_machinery":"The load-bearing object is the MIDAS scheme itself: 23 dialog act tags arranged as leaf nodes under semantic request and functional request subtrees, with multi-label support capped at two tags. The hierarchy guides annotators to the right tag; the multi-label rule captures utterances that do more than one thing, such as a negative answer that is also a task command; context completion lets annotators resolve ellipsis before tagging; and the priority rule keeps only the two tags most useful for dialog planning. On the prediction side, the mechanism is a BERT encoder fine-tuned on in-domain unlabeled conversation before supervised multi-label training, with context formed by concatenating the previous system segment, previous user segment, and current user segment.","core_discovery":"MIDAS organizes dialog acts into two trees, a semantic request type and a functional request type, with classes, categories, and 23 leaf-node tags; utterances receive up to two tags, selected by a priority order that favors answer, command, opinion, statement non-opinion, and question. Annotators reconstruct elliptical meaning from context before tagging, though the original text stays unchanged. Applied to 24K segments from Gunrock conversations, the scheme is reliable enough for two annotators to agree at $\\kappa = 0.94$, and a multi-label classifier trained with in-domain fine-tuned BERT and textual context from previous system and user turns predicts the tags at 0.79 F1.","pith_inferences":["Because the segmentation model is the main identified error source, jointly training utterance segmentation and dialog act prediction, as the paper lists for future work, would likely raise the 0.79 F1; the paper does not demonstrate this.","The priority order answer > command > opinion > statement > question encodes a dialog-management reflex: detect compliance first, then topic shifts; that ordering could be evaluated downstream by measuring task success under alternative label-selection rules.","Nothing in the tag definitions ties MIDAS to Gunrock, so the scheme should transfer to other voice assistants; a cross-bot annotation study would test that portability.","Automating the annotators' ellipsis completion step, rather than leaving it to humans, could improve prediction on short answers such as 'the great gatsby'; the paper leaves this to future work."],"forward_implications":["A dialog system can use MIDAS tags directly for policy decisions: follow a proposed topic, answer a question, execute a command, or recognize a complaint.","The 24K annotated segments provide a training resource for open-domain dialog act prediction that does not rely on the poorly transferring human-human Switchboard labels.","The 0.79 F1 result indicates automatic dialog act prediction on raw ASR output is feasible for open-domain social bots, not just task-oriented systems.","Separating semantic request tags from functional request tags gives a system two orthogonal signals: what topic the user wants and what discourse move the user is making.","The reported 47.38% accuracy of a BERT model trained on SWBD-DAMSL and tested on human-machine conversation motivates replacing human-human schemes with machine-directed ones."],"supporting_citations":[{"why":"Supplies the SWBD-DAMSL scheme and Switchboard corpus that the paper argues do not transfer to human-machine dialogue; also the source of the 42-to-23 tag mapping.","marker":"(Jurafsky et al., 1997)"},{"why":"Describes Gunrock, the Alexa Prize social bot whose 380K conversations provide the collected dataset.","marker":"(Chen et al., 2018)"},{"why":"Provides the DIT++ taxonomy and multi-dimension/multi-function design that MIDAS's hierarchical structure follows.","marker":"(Bunt, 2009)"},{"why":"Defines the ISO standard dialogue act annotation framework that motivates MIDAS and its context-completion requirement.","marker":"(Bunt et al., 2010)"},{"why":"Supplies the BERT pretrained model that the transfer-learning classifier fine-tunes.","marker":"(Devlin et al., 2018)"},{"why":"Motivates unsupervised in-domain fine-tuning of BERT on 50M unlabeled conversational utterances.","marker":"(Siddhant et al., 2018)"},{"why":"Provides the Cornell Movie-Quotes Corpus used to train the sentence segmentation model.","marker":"(Danescu-Niculescu-Mizil and Lee, 2011)"},{"why":"Supplies the boundary-only sentence segmentation approach that avoids predicting punctuation.","marker":"(Favre et al., 2008)"},{"why":"Establishes the earlier dialog act modeling and segmentation pipeline that MIDAS's predictor setting extends.","marker":"(Stolcke et al., 2000)"},{"why":"Introduces an alternative human-machine dialog act scheme with 14 tags, the comparison point for why MIDAS provides richer intents.","marker":"(Khatri et al., 2018)"}],"fun_headline_variants":["Dialog act scheme for bots: 94% agreement, 0.79 F1","MIDAS: tagging human-bot chats with 94% reliability","New dialog act scheme targets imperfect bot understanding","Machine-friendly dialog acts: multi-label, hierarchical, 0.94 kappa","Forget human-human dialog tags: MIDAS models bot-human chat"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the automatic ASR segmentation model, at 84.43% micro-F1, produces usable utterance units; if its boundaries are systematically wrong, the $\\kappa = 0.94$ agreement and the 0.79 F1 describe units that do not match the user's intended utterance boundaries.","fun_headline_variants_meta":{"raw":{"variants":["Dialog act scheme for bots: 94% agreement, 0.79 F1","MIDAS: tagging human-bot chats with 94% reliability","New dialog act scheme targets imperfect bot understanding","Machine-friendly dialog acts: multi-label, hierarchical, 0.94 kappa","Forget human-human dialog tags: MIDAS models bot-human chat"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000582,"raw_usage":{"total_tokens":2681,"prompt_tokens":825,"completion_tokens":1856,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":1763}},"tokens_in":441,"tokens_out":1856,"duration_ms":13956,"temperature":1.0,"reasoning_tokens":1763,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:55:12.869405+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Annotate a held-out sample of user turns twice: once after automatic ASR segmentation and once after manual human segmentation, then measure MIDAS label agreement on the two versions. If classifier F1 or annotator kappa on manually segmented units falls clearly below the reported 0.79 and $\\kappa = 0.94$, the usability claim is undermined.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the ISO standard dialogue act annotation framework that motivates MIDAS and its context-completion requirement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SWBD-DAMSL scheme and Switchboard corpus that the paper argues do not transfer to human-machine dialogue; also the source of the 42-to-23 tag mapping."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes Gunrock, the Alexa Prize social bot whose 380K conversations provide the collected dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the DIT++ taxonomy and multi-dimension/multi-function design that MIDAS's hierarchical structure follows."},{"cited_title":"Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent Agents","cited_arxiv_id":"1811.05370","evidence_quote":"Motivates unsupervised in-domain fine-tuning of BERT on 50M unlabeled conversational utterances."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces an alternative human-machine dialog act scheme with 14 tags, the comparison point for why MIDAS provides richer intents."}],"review_version":1}