{"id":"7fcedc5f-a2f0-49e7-ada1-462a6ef404fa","arxiv_id":"2607.03744","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Dyadic turn-pair timing alone is the strongest single modality on DAIC-WOZ development data, and convex late fusion with RoBERTa reaches 0.804/0.669 macro-F1 while assigning zero weight to WavLM acoustics.","lead":"A 24-feature timing model of interviewer-participant turn pairs beats or matches large frozen speech and text encoders for depression screening on DAIC-WOZ, and late fusion with text reaches 0.804/0.669 macro-F1 while dropping acoustics. Timing of the exchange may be a cheap, interpretable signal for dyadic mental-health screening.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The CTD gain may be inflated by wizard-controlled ask-side timing rather than participant depression, so the strongest single-modality and fusion claims rest on a non-transferable artifact.","rationale":"The Reader correctly isolates the wizard-controlled Ask-side confound as the weakest assumption and already conditions the verdict on response-side ablations and external validation. That concern is load-bearing: without it, the strongest empirical pattern (CTD best unimodal on dev + fusion gain that zeros acoustics) cannot be interpreted as evidence for participant depression dynamics. No stronger internal inconsistency appears; protocol hygiene, public code, and transparent CIs are real strengths, and the small-n overlapping intervals are already acknowledged. The concrete Res-only re-run is the single decisive check that would either vindicate or deflate the claim. Therefore the CONDITIONAL verdict stands; no upgrade or downgrade is warranted until that ablation is reported.","tokens_in":13263,"tokens_out":575,"duration_ms":5190,"concrete_test":"Re-train the exact CTD logistic regression (C=0.3, train-fit standardization, session-mean) using only the Res-side subset of Table I (res_d, res voiced/silence ratios, res_bt/st, res_h, and any pure Res ratios) plus, separately, an Ask-ablated set that zeros all Ask-involving features. Report dev/test macro-F1 and the re-tuned convex weights with RoBERTa. If Res-only CTD falls below ~0.65 dev or loses the fusion lift over RoBERTa alone, the headline claim that dyadic timing is a robust complement is substantially weakened.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim treats the 24-d CTD bank (Table I) as a first-class depression modality that beats frozen WavLM/RoBERTa on dev and drives fusion gains (0.804/0.669). Half the features are Ask-side (ask_d, ask_ud/du/sd/ds/su/us, ask_bt/st, duration ratios/diffs involving ask). Sec. VIII itself notes Ellie is wizard-controlled, so ask-side durations, silences, and response latency (res_h) can encode prompt selection, scripted pacing, or operator reaction to perceived depression rather than participant-intrinsic psychomotor retardation. If those Ask features carry most of the signal, the reported CTD superiority and the zero acoustic weight are DAIC-specific artifacts, not evidence that conversational timing is a transferable complement for real clinician interviews. The paper never reports a response-side-only or Ask-ablated ablation, so the load-bearing causal link remains untested.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper asks whether dyadic conversational temporal dynamics (CTD)—a fixed 24-dimensional bank of Ask/Res turn-pair timing descriptors (durations, voiced/silence ratios, response latency, backchannel/silence counts)—can serve as a first-class modality for depression screening on DAIC-WOZ. Under a subject-independent, leakage-safe protocol with documented session exclusions, a compact L2-logistic CTD detector is compared to frozen WavLM-large (utterance-level acoustic probe) and frozen RoBERTa-large (turn-level semantic probe). On the development set CTD is the strongest unimodal system (macro-F1 0.746 vs. 0.690 RoBERTa and 0.673 WavLM). Convex-weighted late fusion of session probabilities reaches 0.804 / 0.669 macro-F1 on dev/test and drives the acoustic weight to exactly zero, so the deployed system reduces to RoBERTa+CTD. The authors frame the study as preliminary and discuss statistical and transfer caveats in Sec. VIII.","tokens_in":13530,"tokens_out":996,"duration_ms":12624,"significance":"If the result holds beyond this cohort, treating transcript-level turn timing as an equal-footing modality is a useful, lightweight, and interpretable complement to large self-supervised encoders for dyadic clinical screening. Strengths that support taking the work seriously include: (i) reuse of a previously published, task-agnostic CTD feature bank with no DAIC-specific feature engineering; (ii) a carefully documented, subject-independent protocol with train-only statistics, published integrity exclusions, and both parameter-free mean-prob and dev-tuned convex fusion; (iii) explicit reporting of bootstrap CIs, paired prediction changes, and seed variability; and (iv) public code with seed/hyperparameter/deployed-model details. These practices raise the bar relative to much of the DAIC-WOZ literature even if the absolute gains remain small-sample and indicative.","major_comments":[{"comment":"Sec. IV-C / Table I and Sec. VIII: Roughly half of the 24 CTD features are Ask-side (ask_d, the six ask voiced/silence ratios, ask_bt/st, and all duration diffs/ratios involving Ask), and res_h is the Ask-end→Res-start gap. Because Ellie is wizard-controlled, these quantities can encode prompt selection, scripted pacing, or operator reaction rather than participant psychomotor retardation. The manuscript correctly flags this in Sec. VIII but never reports a response-side-only or Ask-ablated detector. Without that ablation (or an equivalent prompt-controlled check), the claim that CTD is a transferable ‘first-class modality for dyadic depression screening’ rests on a DAIC-specific artifact risk. This is load-bearing for the abstract and conclusion; a response-side-only row in Table II (or a clear statement that the result is DAIC-Ellie-specific) is needed before the interpretation can sta","section":null},{"comment":"Sec. VII-D and Table II: Dev n=33 and test n=45 yield wide, overlapping bootstrap CIs (e.g., proposed system test macro-F1 [0.509, 0.806] vs. CTD alone [0.472, 0.771]), and the paired intervals for RoBERTa+CTD vs. each unimodal arm include or touch zero. The paper already notes this, yet the abstract and bolded claims still present 0.804 / 0.669 as an unambiguous improvement and CTD as ‘the highest single-modality performance.’ The central ordering should be stated as indicative on this split, with the more robust observations (consistent dev lift, parameter-free mean_prob[T+CTD] test 0.650, learned zero acoustic weight) foregrounded over point-estimate ranking.","section":null},{"comment":"Sec. IV-A and Table II/III: The acoustic stream is weak on test (0.545) and receives weight 0.0 in the three-way convex solution. The authors attribute this partly to utterance- vs. turn-level pooling and to the frozen SUPERB-style probe, but they do not show that a stronger acoustic baseline (fine-tuned WavLM, different pooling, or a published competitive acoustic system under the same split) would still be down-weighted. The zero-weight conclusion is therefore specific to this probe/cohort; either strengthen the acoustic baseline or qualify the claim that ‘acoustics contributes little’ as probe-dependent rather than modality-general.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: a fixed 24-d Ask/Res timing bank, taken unchanged from Chou et al.’s deception work, is the strongest single modality on DAIC-WOZ development data (0.746 macro-F1) and, when late-fused with frozen RoBERTa, lifts both splits to 0.804/0.669 while the learned convex weights drive WavLM to exactly zero. That is a clean empirical pattern, not a re-derivation of the features.\n\nWhat the paper does well is protocol hygiene and honesty. They exclude the ten integrity-problem sessions from a published catalog, keep all statistics train-only, freeze the SSL encoders, report both parameter-free mean fusion and dev-tuned convex weights, ship code, and flag the wizard-controlled Ellie problem themselves in Sec. VIII. The acoustic-versus-turn granularity ablation is also useful: utterance-level WavLM beats turn-level pooling, so they correctly leave turn structure to CTD. Citation pattern is tight and the math is just logistic regression plus score-level convex combination—nothing load-bearing is hidden.\n\nSoft spots are real but proportionate. Dev n=33 and test n=45 produce wide, overlapping bootstrap CIs; paired intervals for the fusion gain touch zero, so the test-side win is indicative only. Half the CTD features are Ask-side; because Ellie is wizard-controlled, those durations, silences, and res_h can encode prompt selection or operator reaction rather than participant psychomotor retardation. The paper never reports a response-side-only or Ask-ablated run, so the transfer claim to real clinician interviews is untested. The acoustic probe is also weak (test 0.545), which may exaggerate the “acoustics get zero weight” story. None of this is circularity or fabrication; it is just the usual DAIC-WOZ small-sample and artifact risk, which they mostly own.\n\nThis is for people who already work on multimodal depression screening or dyadic timing and want a lightweight, interpretable stream that can sit beside text. It is not a clinical-deployment paper and does not reorganize the field. I would send it to peer review: the experiment is carefully run, the result is new on this protocol, and the limitations are stated clearly enough that referees can demand the missing ablations. Worth reading and citing if you work in this niche; not something to build a grant around until the ask-side check is done.","headline":"Solid small-scale transfer of fixed dyadic timing features to DAIC-WOZ depression screening; CTD beats frozen SSL unimodals on dev and fuses cleanly with text, but wizard-controlled ask-side timing and tiny n leave the transfer claim provisional.","tokens_in":14184,"tokens_out":617,"would_cite":true,"duration_ms":5377,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A 24-feature map of how people take turns can match or beat large speech and text models for depression screening in clinical interviews.","keywords":["depression detection","conversational temporal dynamics","turn-taking","multimodal late fusion","DAIC-WOZ","WavLM","RoBERTa"],"falsifier":"Recompute the same CTD features using only participant-side response timing (or on interviews with a fully scripted, non-adaptive interviewer) and check whether the unimodal CTD advantage and the fusion gains over text alone disappear.","tokens_in":14115,"feed_emoji":"⏱️","tokens_out":644,"duration_ms":5001,"temperature":0.7,"pith_summary":"Most automatic depression screening from clinical interviews focuses on what the participant says or how their voice sounds. This paper asks whether the timing of the exchange itself — how long each person holds the floor, how long they hesitate before answering, how much silence sits inside a turn — carries a depression signal of its own. The authors build a fixed, 24-number summary of dyadic turn-pair timing, treat it as a full modality next to frozen large acoustic and language models, and show that on the standard DAIC-WOZ interview corpus this tiny timing module is the strongest single stream on the development set. When its session scores are fused with the text model by a simple convex weight search, performance rises further, while the acoustic stream is driven to zero weight. The practical claim is that conversational timing is a lightweight, readable complement that should be treated as a first-class input for dyadic depression screening.","feed_headline":"Turn-taking timing beats large models for depression screening","feed_subtitle":"A 24-number map of interview pauses and response latency lifts dyadic detection while acoustics get zero weight","key_machinery":"Conversational temporal dynamics (CTD): a fixed, task-agnostic 24-feature bank of dyadic Ask/Res turn-pair timing (turn durations and ratios, voiced-versus-silence ratios, response hesitation, and silence/backchannel counts) averaged per session and classified by regularized logistic regression, then fused at the probability level with other modality scores.","core_discovery":"On DAIC-WOZ, a compact 24-dimensional conversational temporal dynamics module built from Ask/Res turn-pair durations, silences, response latency, and backchannel counts is the strongest single modality on the development set (macro-F1 0.746), and convex-weighted late fusion with a frozen RoBERTa-large semantic detector reaches 0.804 development and 0.669 test macro-F1 while assigning zero weight to the acoustic stream.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["24D turn timing tops frozen large models for depression detection","Dyadic pause and latency alone beat WavLM and RoBERTa on DAIC-WOZ","Compact conversational timing is strongest single modality on dev set","Late fusion zeros acoustics yet lifts depression screening F1","Turn-pair dynamics complement semantics while ignoring audio stream"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The timing features, including interviewer-side durations and silences, mainly reflect the participant's depression rather than the virtual interviewer's wizard-controlled prompt choices or other session artifacts.","fun_headline_variants_meta":{"raw":{"variants":["24D turn timing tops frozen large models for depression detection","Dyadic pause and latency alone beat WavLM and RoBERTa on DAIC-WOZ","Compact conversational timing is strongest single modality on dev set","Late fusion zeros acoustics yet lifts depression screening F1","Turn-pair dynamics complement semantics while ignoring audio stream"]},"model":"grok-4.5","effort":"low","cost_usd":0.004454,"raw_usage":{"total_tokens":1293,"prompt_tokens":730,"num_sources_used":0,"completion_tokens":89,"cost_in_usd_ticks":44540000,"prompt_tokens_details":{"text_tokens":730,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":474,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":730,"tokens_out":89,"duration_ms":4270,"temperature":1.0,"reasoning_tokens":474,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T00:14:44.671866+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Recompute the same CTD features using only participant-side response timing (or on interviews with a fully scripted, non-adaptive interviewer) and check whether the unimodal CTD advantage and the fusion gains over text alone disappear.","supporting_citations":[],"review_version":1}