{"id":"7a7c94cd-5ce6-40f8-91c6-3628e7f59231","arxiv_id":"2607.14310","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Dialogs is a new 20.6-hour studio-quality Russian conversational speech corpus with emotion labels and a VITS2 proof-of-concept.","lead":"The authors release Dialogs, a 20.6-hour studio-recorded Russian speech dataset of acted conversations by three actors, with 12 emotion and style labels per utterance. It targets a gap in Russian TTS data: no open corpus combines studio audio, conversational style, and per-utterance emotion labels at this scale.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dialogs' expressiveness/conversational advantage may be a sampling artifact: emotion-stratified 188 clips vs random 100 baselines.","rationale":"The reader's weakest assumption is exactly the most load-bearing vulnerability. The corpus itself has real independent support: public release, open commercial license, concrete recordings, and training code; there is no internal inconsistency in the construction. But the headline claim about expressiveness and conversational naturalness rests entirely on a between-corpus MOS comparison whose two arms are selected under different principles. The paper's assertion that CIs make subset differences irrelevant is incorrect for selection bias. A representative-sample rerun is straightforward. Therefore I agree with the CONDITIONAL verdict; no further adjustment is needed.","tokens_in":5119,"tokens_out":4136,"duration_ms":42927,"concrete_test":"Draw 100 clips from Dialogs with probability proportional to style duration (or as a simple random sample from the full corpus), run the identical MOS protocol with the same rater pool and outlier rule, and compare expressiveness and conversational-naturalness means/95% CIs to Ruslan/Natasha in Table 3. If the gap collapses or reverses, the current expressiveness advantage is an artifact of emotion-stratified sampling; if it persists, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is Table 3's +0.23–0.30 expressiveness/conversational edge for Dialogs over Ruslan/Natasha. The Dialogs MOS subset is intentionally emotion-stratified (5 per speaker per emotion, 188 clips), while baselines are random 100-clip draws. This makes the comparison asymmetric in exactly the dimension being measured. Corpus style durations (Table 4) are heavily skewed: neutral 643.5 min, happy 349.6 min, surprise 78.6 min, then rare styles such as whispering 6.1 min, laughing 19.9 min, angry 19.2 min. Uniform 5-per-emotion selection massively oversamples rare, highly expressive styles relative to their corpus share, so raters hear a curated expressive set from Dialogs but typical read speech from baselines. CIs quantify within-condition sampling error, not selection bias; subset-size differences are irrelevant only if subsets are representative. The additional removal of outlier annotators per corpus (41/26/23 raters, no sensitivity analysis) could shift condition means asymmetrically. The dataset's own limitation admits dialogs are scripted/acted, so even the naturalness gap may reflect listeners responding to emotional variety in the selected clips rather than turn-taking behavior. Unless the same representative sampling is applied to all corpora, Table 3 does not establish corpus-level expressiveness superiority.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Dialogs, a 20.6-hour studio-quality Russian conversational speech corpus with 11,796 utterances from 3 actors, recorded at 44.1 kHz and released under OpenRAIL. The corpus is claimed to fill a gap for expressive, conversational Russian TTS by combining studio audio, dialog recordings, and per-utterance style/emotion labels across 12 categories. The authors validate the corpus with crowd-sourced MOS tests, reporting comparable audio quality and intelligibility to the Russian read-speech baselines Ruslan and Natasha, and higher expressiveness and conversational naturalness. They also train VITS2 on the corpus as a proof of concept, reporting MOS and UTMOS scores for synthesized speech.","tokens_in":5486,"tokens_out":3828,"duration_ms":40119,"significance":"If the claims hold, Dialogs is a useful public resource for Russian expressive and dialog TTS, with commercial-friendly licensing, per-utterance emotion labels, and a reproducible training demonstration. The paper is transparent about the data release, and the availability of code and corpus is a concrete contribution. However, the headline expressiveness and conversational-naturalness advantages rest on a MOS comparison with asymmetric sampling across conditions, and the annotation aggregation procedure is not validated; these issues currently prevent the paper from fully establishing its central claims.","major_comments":[{"comment":"The headline expressiveness (+0.23–0.25) and conversational (+0.24–0.30) advantages are not established by the reported comparison because the evaluation subsets differ by design. Dialogs' 188 clips are stratified by 12 emotion/style categories (5 per speaker per style), while Ruslan and Natasha are random 100-clip draws from unlabelled read corpora. Table 4 shows extreme duration skew (neutral 643.5 min, happy 349.6 min, whisper 6.1 min, laughing 19.9 min, angry 19.2 min), so Dialogs raters hear a curated set that oversamples rare, highly expressive styles, whereas baseline raters hear typical read speech. Confidence intervals quantify sampling variability within each subset, not selection bias across subsets. Please either evaluate all corpora under the same sampling protocol, or supply a sensitivity analysis using a random Dialogs subset to show the advantage persists.","section":"§3.5, Table 3"},{"comment":"The annotation aggregation is not validated and may bias labels. With three annotators and majority vote, all-disagree ties are resolved by selecting the globally least-frequent category. This can assign a rare label to an utterance for which no annotator chose that category, artificially inflating rare-style durations in Table 4 and injecting label noise into training. No inter-annotator agreement, distribution of tie cases, or comparison to alternative tie-breaking is reported. Since per-utterance style/emotion labels are a central new contribution, please quantify the disagreement rate and justify or change the tie-breaking rule.","section":"§3.4"},{"comment":"Outlier annotator removal lacks transparency. The paper reports 41/26/23 retained raters per corpus after 'response pattern analysis' but does not state the removal criterion, the number removed per corpus, or results without removal. If removal was more aggressive for Ruslan/Natasha than for Dialogs, the apparent expressiveness gap in Table 3 could be an artifact. Please provide the criterion, the counts removed, and a sensitivity analysis (e.g., all raters included, or an identical outlier rule across conditions).","section":"§3.5"}],"minor_comments":[{"comment":"No inter-annotator agreement measure (e.g., Fleiss' kappa) is reported. Also specify annotator instructions and whether annotators heard full dialog context or isolated utterances.","section":"§3.4"},{"comment":"The statement that expressiveness (2.56) and conversational (2.59) 'notably exceed' intelligibility (2.28) is not supported by the reported 95% CIs: the intervals overlap substantially. Use a paired test or soften the claim.","section":"§5, Table 5"},{"comment":"The Natasha row lacks a citation. Add a reference or state the source.","section":"Table 1"},{"comment":"Clarify the recording setup: were two Behringer XM8500 microphones used, one per actor? How was stereo captured? Place this information in the metadata as well.","section":"§3.2"},{"comment":"The corpus UTMOS score (3.17±0.07) and the TTS UTMOS score (3.36±0.06) are close; since UTMOS is not calibrated across different audio conditions, avoid direct numerical comparison or state that they are not comparable.","section":"§3.1, Table 2"}],"recommendation":"major_revision","confidential_remarks":"The corpus release is valuable and likely to be useful to the Russian TTS and speech-resource community. The principal risk is overclaiming from an asymmetric MOS evaluation; if the authors cannot re-run the comparison with matched sampling, they should reframe the expressiveness/conversational advantages as properties of the corpus's curated content rather than as corpus-level superiority. No concerns about novelty disclosure or authorship."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: Dialogs is a useful new resource — 20.6 hours of studio-recorded Russian dialog across three speakers, per-utterance emotion labels in 12 categories, OpenRAIL license, and public data and code. If you work on Russian expressive TTS, it fills a real gap. The paper's gap analysis is honest, and the audio-quality and intelligibility parity with Ruslan/Natasha is consistent with the recording setup. But the headline claim — that Dialogs is more expressive and conversational than those baselines — is not yet established. The stress-test note has it right. The MOS comparison draws Dialogs clips stratified at 5 per speaker per emotion (188 clips), while Ruslan and Natasha get random 100-clip draws. Rare styles like whisper, laughing, and angry make up only a few minutes each in the corpus, so that stratified selection massively oversamples the very styles most likely to be rated expressive. Confidence intervals cover within-condition sampling error, not selection bias. The paper's assertion that subset-size differences don't affect comparability is not an argument. Outlier annotator removal (41/26/23 raters) without sensitivity analysis is a second, smaller worry. Neither issue is fatal, but both need to be addressed before the +0.23-0.30 expressiveness/conversational edges are taken at face value. What's solid: the corpus itself, the clear documentation of recording and annotation, and the limitation note that dialogs are acted, so the authors do not overclaim spontaneity. The VITS2 proof-of-concept is just that: it shows you can train a TTS on this data, and it does not compare against a baseline trained on another corpus. That's acceptable for a data paper. The tie-breaking rule for three-way annotation disagreements (pick the globally least-frequent category) is a slightly odd choice, but it only affects rare labels and is unlikely to change the main outcome. My recommendation: send it to peer review. It's a data resource with public artifacts, and the community will use it regardless of whether Table 3 survives contact with matched sampling. Reviewers should ask the authors to redo the comparison with the same sampling scheme, or at least show a sensitivity analysis. If the expressiveness edge shrinks, the paper still stands as a corpus release. If you're working on Russian TTS, cite it for the dataset, not for the expressiveness claim.","headline":"A genuinely useful Russian expressive dialog corpus, but the headline expressiveness advantage over baselines rests on an asymmetric MOS comparison that should be fixed before the claim is cited.","tokens_in":758,"tokens_out":724,"would_cite":true,"duration_ms":28532,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Dialogs is a new open Russian conversational speech corpus that matches studio read-speech quality while scoring substantially higher on expressiveness and conversational naturalness, and it supports training an expressive dialog TTS.","keywords":["Russian speech corpus","conversational TTS","expressive speech synthesis","emotion annotation","dialog assistants","studio-quality audio","VITS2","MOS evaluation"],"falsifier":"A listening test that scores randomly sampled Dialogs clips (not stratified by emotion) against Ruslan and Natasha using the same rater pool would show whether the expressiveness and conversational naturalness advantages shrink or vanish; if the gap persists on random samples, the claim is robust.","tokens_in":5024,"feed_emoji":"🎙️","tokens_out":5634,"duration_ms":57145,"temperature":0.7,"pith_summary":"The paper introduces Dialogs, a 20.6-hour Russian conversational speech corpus recorded in a professional studio by three actors facing each other and speaking from script prompts they were encouraged to improvise around. It argues that this fills a gap in Russian speech resources: existing studio-quality corpora are single-speaker read speech without emotion labels, while web-mined corpora are uncontrolled. Crowd-sourced listening tests show Dialogs matches the read-speech baselines Ruslan and Natasha in audio quality, intelligibility, and prosody, while scoring 0.23–0.30 points higher (on a 5-point scale) for expressiveness and conversational naturalness. As a proof of concept, the authors train a VITS2 text-to-speech model on Dialogs; the resulting speech scores higher on expressiveness and conversational dimensions than on intelligibility, indicating the model absorbed the corpus's conversational prosody. If these ratings hold, Dialogs provides Russian TTS developers with a publicly and commercially usable studio-quality conversational corpus with per-utterance style labels.","feed_headline":"Dialog speech corpus rates higher on expressiveness than read speech","feed_subtitle":"Studio-recorded, emotion-labeled audio scores 0.23–0.30 points above read-speech baselines in crowd tests.","key_machinery":"The central object is the Dialogs corpus itself: 11,796 stereo utterances, 21,262 unique words, and 12 style/emotion labels per utterance (neutral, happy, surprise, sad, disgust, angry, tongue-twister, poem, whisper, arrogance, laughing, fear). The load-bearing design choices are the face-to-face studio recording protocol that encourages natural turn-taking and improvisation, the crowd-based triple annotation with majority-vote (ties broken toward rare styles), and a stratified test set of 188 utterances (5 per speaker per emotion) used for MOS evaluation. This structure is what lets the paper claim that the corpus, not just the recording hardware, drives the conversational and expressive ad","core_discovery":"Dialogs is a studio-quality, openly licensed Russian corpus of acted conversational speech with per-utterance labels across 12 style/emotion categories. Its recording protocol—actors seated face-to-face, improvising around script prompts—captures turn-taking rhythm and expressive prosody that read-speech resources lack. In crowd MOS evaluation, Dialogs is rated comparable to strong single-speaker baselines on overall quality, audio quality, and intelligibility, while receiving substantially higher ratings for expressiveness and conversational naturalness. Training a VITS2 model on the corpus yields speech whose expressiveness and conversational scores exceed its intelligibility score, eviden","pith_inferences":["The evaluation asymmetry—stratified emotion-rich excerpts for Dialogs versus random clips for baselines—means the expressiveness gap could be partly an artifact of sampling; a matched random-sample listening test would settle it.","The corpus's acted, script-improvised nature means it does not contain true spontaneous speech, overlapping turns, or disfluencies; extending the recording protocol to unscripted interaction could test whether the conversational advantage generalizes beyond acted dialogs.","The per-utterance multi-label annotation opens avenues beyond TTS, such as expressive resynthesis or emotion-transfer research, where a controlled studio corpus with rare styles (whisper, tongue-twisters) is currently scarce for Russian.","If the expressiveness ratings reproduce in independent listening studies, the face-to-face recording protocol itself—actors reacting to each other rather than to a microphone—could become a standard recipe for constructing conversational TTS corpora in other languages."],"forward_implications":["Russian TTS systems can be trained or fine-tuned on Dialogs to produce expressive, conversational output without resorting to uncontrolled web-mined data.","The 12 style/emotion labels enable emotion-conditioned synthesis, letting developers choose a delivery style per utterance in dialog assistants.","Because the corpus is released under an open commercial license, unlike most existing Russian studio corpora, production teams can legally use it.","The stratified test set provides a reproducible benchmark for evaluating expressive Russian TTS, with per-speaker and per-emotion coverage.","Mixing Dialogs with larger read-speech corpora is expected to lift synthesis quality, since the low per-speaker hours cap raw naturalness (a limitation the authors note)."],"fun_headline_variants":["Russian dialog corpus outshines read speech on expressiveness","Acted dialog corpus boosts expressive TTS training","New Russian speech data scores higher on naturalness than read","Studio-recorded dialogs with emotion labels for expressive TTS","Corpus captures turn-taking and prosody for better dialog TTS"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim that Dialogs is more expressive and conversational than the baselines rests on the assumption that its emotion-stratified 188-clip evaluation subset is comparable to the baselines' random 100-clip subsets, despite the different sampling and the removal of outlying raters.","fun_headline_variants_meta":{"raw":{"variants":["Russian dialog corpus outshines read speech on expressiveness","Acted dialog corpus boosts expressive TTS training","New Russian speech data scores higher on naturalness than read","Studio-recorded dialogs with emotion labels for expressive TTS","Corpus captures turn-taking and prosody for better dialog TTS"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000172,"raw_usage":{"total_tokens":1064,"prompt_tokens":647,"completion_tokens":417,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":391,"completion_tokens_details":{"reasoning_tokens":335}},"tokens_in":391,"tokens_out":417,"duration_ms":4808,"temperature":1.0,"reasoning_tokens":335,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:27:26.182058+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A listening test that scores randomly sampled Dialogs clips (not stratified by emotion) against Ruslan and Natasha using the same rater pool would show whether the expressiveness and conversational naturalness advantages shrink or vanish; if the gap persists on random samples, the claim is robust.","supporting_citations":[],"review_version":1}