{"id":"b7f56a25-f3ea-477a-82f3-a2e55f4f8c64","arxiv_id":"2507.15871","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"During an emotional working memory task, socially anxious and non-anxious participants showed similar electrodermal responses, but resting-state EDA classified the groups with moderate accuracy.","lead":"Researchers measured skin conductance in 50 people with and without social anxiety during rest and during a memory game with emotional faces. The task raised arousal equally in both groups, while resting-state signals alone separated the groups moderately well.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 2-minute 'resting' baseline was recorded after consent and task briefing, immediately before practice and the main task (Section III-D), so the baseline ML signal (AUC 0.73) and baseline-to-task contrasts are plausibly contaminated by anticipatory anxiety in the SA group; a within-baseline…","rationale":"Good-faith reading: the paper is careful with its null result, releases data and code, and transparently reports effect sizes and caveats. The central claim, however, includes a contrast between a discriminative 'resting-state' baseline and a non-discriminative task. The procedure makes the baseline a poor operationalization of rest: Section III-D places it after consent and briefing and before practice and the main task, so SA participants are likely in an anticipatory state. Because the paper offers no manipulation check or state-anxiety measure, the baseline classification (AUC 0.73, RF 0.85) could reflect group differences in anticipatory reactivity instead of trait EDA. This also contaminates the repeated-measures baseline-to-task comparison, so 'similar increases' cannot be cleanly interpreted. I agree with the reader's weakest-assumption identification. The task-phase null result itself is less vulnerable, but the central 'context-dependent' interpretation and the resting-state component of the claim would need revision if the proposed check shows a within-baseline decline or an unstable AUC. Other concerns identified by the reader, such as uncorrected multiple testing, missing CI/permutation tests for AUCs, and possible data overlap with reference [34], are real but secondary; they do not change the conditional recommendation.","tokens_in":16016,"tokens_out":8145,"duration_ms":100774,"concrete_test":"Using the released EDA data, recompute the baseline classification while using only the final 60 seconds of each participant's 2-minute baseline segment, and compare the resulting AUC to the reported full-baseline AUC with a permutation-based confidence interval. If the last-60-second AUC drops toward chance or is significantly lower, the resting-state signal is an anticipatory artifact; if the AUC is stable, the concern is refuted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the 2-minute baseline is a neutral resting state. Per Section III-D and Fig. 3, it is recorded after informed consent and the task briefing, immediately before a 25-trial practice and the 13-minute main task. For socially anxious participants this window is likely dominated by anticipatory anxiety about the upcoming task and by the experimenter's presence; the paper includes no manipulation check, no subjective state-anxiety rating, and no post-session appraisal. If this concern lands, the central claim's 'resting-state EDA' component is not established: the average baseline AUC of 0.73 and the RF AUC of 0.85 could reflect group differences in anticipatory state rather than a stable trait signature. The baseline-to-task contrast is also contaminated: if SA participants begin from elevated anticipatory arousal, 'both groups showed similar increases' is not a clean measure of task reactivity, and the early-phase interactions (SCR Amplitude Mean, EDASymp) cannot be unambiguously interpreted as task-onset effects. The task-phase null result is less affected, but the paper's context-dependence conclusion rests on comparing a non-neutral baseline with the task, so the interpretation would shift from 'resting-state trait differences are masked by the task' to 'pre-task anticipatory differences are masked by the task.' The released data and code make this concern directly testable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a wearable-EDA study of 50 participants (25 socially anxious, 25 non-socially anxious) who performed an emotional 2-back working-memory task with facial expressions. Ten EDA features spanning tonic, phasic, sympathetic, spectral, and nonlinear domains were extracted from a 2-minute baseline and from three 2-minute task phases, and analyzed with mixed ANOVAs plus complementary machine-learning classifiers (LR, SVM, RF, GB, 1D-CNN). The main results are: the task reliably increased autonomic arousal in both groups; no group main effects were found in any phase; two brief Condition-by-Group interactions appeared only in the early phase; and classification could distinguish groups from baseline EDA (average AUC 0.73) but not from task EDA (average AUC 0.60, 0.54, 0.55 across phases). The conclusion is that cognitive-emotional load without explicit social-evaluative threat does not differentiate socially anxious from non-anxious individuals at the autonomic level, and that task engagement masks resting-state discriminative information.","tokens_in":16264,"tokens_out":5800,"duration_ms":73575,"significance":"If the result holds, it is a useful contribution to the debate about context-dependent physiological biomarkers for social anxiety: it provides a clear null result for non-evaluative cognitive-emotional stress and a moderate positive signal for resting-state EDA. The paper has several strengths that should be acknowledged: the data and code are publicly released; the 1D-CNN evaluation uses participant-level grouped cross-validation to prevent window leakage; the authors are explicit about the small sample size and the exploratory nature of the ML results; and the convergence across five very different classifiers supports the temporal pattern even if individual AUC estimates are uncertain. The main caveats are that the 'resting' baseline is recorded immediately before the task and may be contaminated by anticipatory anxiety, that the statistical tables involve many uncorrected comparisons, and that the ML performance estimates lack confidence intervals and nested validation.","major_comments":[{"comment":"The 2-minute 'baseline' is recorded after informed consent and task briefing and immediately before the practice session and the main task. For socially anxious participants this window is likely dominated by anticipatory anxiety about the upcoming task and by the presence of the experimenter. This is load-bearing because the paper's central 'resting-state EDA' classification result (average AUC 0.73, RF 0.85) and the baseline-to-task contrast ('both groups showed similar increases') both assume that the baseline is a neutral resting state. If the baseline reflects anticipatory arousal, the resting-state discriminability could be a state effect rather than a trait signature, and the claim that task engagement 'masks' resting differences would need to be reframed. The released data make this testable: I would like to see a within-baseline temporal analysis (first vs. second minute), a manipulation check or post-session state-anxiety rating, and/or a comparison with a true rest condition. Without such evidence, the 'resting-state' component of the central claim is not established.","section":"Section III-D, Fig. 3"},{"comment":"The ANOVA results are presented without correction for multiple comparisons. Across ten features and three task phases, each table contains Condition, Group, and Interaction tests, totaling 90 hypothesis tests. Several effects that drive the narrative have p-values near 0.02-0.05, and the EDASymp interaction (p = .045) would not survive even a simple Bonferroni correction; the authors themselves call this finding weak and mention regression to the mean. The central null (no group main effects) is less affected, but the claims of 'broad initial activation' (six of ten features) and 'only transient interactions' depend on uncorrected p-values. Please report corrected p-values or false-discovery-rate q-values, and specify which effects remain significant after correction. This is particularly important for the early-phase interactions and for the late-phase TVSymp and SCR-amplitude effects.","section":"Section IV, Tables II-IV"},{"comment":"The ML classification results are central to the conclusion that 'machine learning confirmed' the statistical findings, but the reported AUC values have no confidence intervals, no permutation-based significance tests, and no nested cross-validation for hyperparameter tuning. With N = 50 and 5-fold cross-validation, each test fold contains only 10 participants, so the standard errors on AUC are large; the baseline RF AUC of 0.85 could be substantially optimism-biased without an inner tuning loop. The authors are appropriately cautious in their caveats, but the caveats mean that the ML evidence is currently insufficiently quantified. Please report bootstrap or DeLong confidence intervals for each AUC, permutation p-values, and ideally a nested CV scheme, so the reader can judge whether baseline AUC 0.73 is statistically distinguishable from task-phase AUCs around 0.55-0.60.","section":"Section V-B, V-B.3"},{"comment":"The conclusion states that SA and NSA individuals 'did not show different EDA responses' and that the task 'does not produce different physiological responses' in socially anxious individuals. This is a claim of equivalence or null effect, but no equivalence testing (e.g., TOST) or effect-size confidence intervals are provided. With N = 25 per group, the ANOVAs may simply be underpowered to detect small-to-moderate group differences, especially for Group main effects where several F values are near zero but a few approach significance (e.g., late-phase SCR Amplitude Mean Group effect p = .083). Please temper the language to 'no statistically significant differences were detected' and, if the authors wish to claim comparability, add equivalence bounds or report the smallest detectable effect size given the sample.","section":"Section VI-D"}],"minor_comments":[{"comment":"The text contains typographical artifacts such as 'ANOV As' and 'ANOV A' instead of 'ANOVAs'; these should be cleaned throughout.","section":"Abstract and Section II"},{"comment":"The abstract says task-phase classification performance is 'average AUC <= 0.57', but the reported Early-phase average is 0.60 (LR 0.56, SVM 0.57, RF 0.69, GB 0.65, CNN 0.53). Please use a description that is accurate for all phases, such as 'near-chance' or '0.54-0.60'.","section":"Abstract and Section V-B.2"},{"comment":"The figure caption says the orange line corresponds to the EDA data-acquisition timeline, but in the manuscript text the orange line is not visible; please ensure the figure rendering includes the line or clarify in the caption what the orange line represents.","section":"Figure 3"},{"comment":"The text repeatedly refers to Supplementary Tables A.1-A.4 (feature descriptions and simple effects), but these tables are not included in the manuscript; please ensure they are uploaded with the submission.","section":"Supplementary tables"},{"comment":"The sentence 'The Group effect approached significance (p = .083, d = -0.501)' for the late phase is potentially misleading without a clear statement that this is a main effect in the absence of a significant interaction; please clarify the interpretation.","section":"Section IV-C.1"},{"comment":"The paper relies on the companion paper [34] for detailed task and behavioral-performance descriptions. Since behavioral performance is relevant to interpreting the EDA results, please include at least the key performance statistics (accuracy, response time, group comparison) in the main text or supplement.","section":"Section III-D"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the authors have been transparent about limitations, which is commendable. The main concern is that the 'resting' baseline may be contaminated by anticipatory anxiety; because the data are released, this concern is directly testable, and I would recommend requesting that analysis. The multiple-comparison issue and the lack of ML uncertainty quantification are also fixable within the manuscript's scope. I do not see a need for rejection, but the paper should not be accepted in its current form because the central 'context-dependent' conclusion rests on a comparison whose baseline validity is not yet demonstrated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the quick read. This is a transparently reported, small-sample wearable EDA study. Twenty-five socially anxious and twenty-five non-socially anxious participants did an emotional 2-back face task while wearing a Shimmer sensor. The task clearly raised arousal in everyone, but there were no group main effects on any of the ten EDA features in any of the three task phases—only two transient early interactions that disappeared within two minutes. Machine learning, including a 1D-CNN, separated the groups at 'rest' (average AUC 0.73) but fell to near chance during the task. That pattern is coherent, and the authors deserve credit for reporting null results so fully and for releasing code and data.\n\nCredit where it's due: the ANOVA tables are complete, the ML methods are diverse, and the listed limitations (N=50, no nested cross-validation, no confidence intervals) are honest. The specific resting-versus-task ML contrast on an emotional 2-back is new, although the underlying null in non-evaluative settings has been reported before (their refs 49-52).\n\nThe biggest soft spot is the baseline. The 'resting' two minutes were recorded after the consent/information sheet and before the task overview and practice. For an SA sample, that window is plausibly dominated by anticipatory anxiety. There is no manipulation check, no state-anxiety rating, and no post-session appraisal. So the baseline ML signal (AUC 0.73, and 0.85 for random forest) could be picking up anticipatory arousal rather than a stable trait. That also contaminates the baseline-to-task reactivity contrast. The task-phase null result is less affected, but the paper's central 'context-dependent biomarker' conclusion rests on comparing a non-neutral baseline with the task. The interpretation shifts from 'resting traits are masked by the task' to 'pre-task anticipatory states are masked by the task.' They should relabel the baseline or collect a proper neutral rest.\n\nOther issues, in order of importance: thirty ANOVAs without multiple-comparison correction; the two early interactions are exactly the kind of thing that can be chance. The ML AUCs have no confidence intervals or permutation tests. The 'necessary ingredient' claim about social-evaluative threat overreaches—one non-evaluative task cannot establish necessity. And they should disclose whether the EDA data overlap with their companion paper [34].\n\nOverall, the core null result is probably real, and the paper is useful for the digital mental health audience. Send it to peer review with major revision on the baseline confound, multiple comparisons, ML validation, and a tempered conclusion.","headline":"A transparently reported small-sample EDA study whose null task result looks solid, but the 'resting' baseline is likely anticipatory anxiety, so the context-dependence conclusion needs reframing.","tokens_in":16823,"tokens_out":4216,"would_cite":false,"duration_ms":45269,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that emotionally salient cognitive load without social-evaluative threat produces equivalent electrodermal responses in socially anxious and non-socially anxious people, whereas resting-state skin conductance can…","keywords":["social anxiety disorder","electrodermal activity","wearable sensing","2-back working memory","skin conductance","machine learning classification","digital mental health","autonomic arousal"],"falsifier":"Record a true resting baseline before the participant is briefed about the task, or after a long habituation period, and rerun the classification. If the resting-state AUC drops to chance when anticipatory arousal is removed, then the reported baseline signal is an artifact of the experimental procedure rather than a stable social-anxiety signature. A second test: add an explicit social-evaluative condition to the same 2-back task; if group discrimination still fails to reach the resting-state level, the paper's claim that evaluation is necessary would be weakened.","tokens_in":15789,"feed_emoji":"🧠","tokens_out":7256,"duration_ms":78213,"temperature":0.7,"pith_summary":"This study asks whether social anxiety shows up in the body during a demanding emotional task that does not involve being judged by others. Fifty participants (25 socially anxious, 25 not) wore a skin-conductance sensor during a resting baseline and a 2-back working-memory task with facial expressions. The task reliably raised autonomic arousal, but the two groups responded equally: only two brief group differences appeared in the first two minutes and then vanished. Machine-learning classifiers could tell the groups apart from resting-state skin conductance (average AUC about 0.73) but not from task data (average AUC at or below 0.57). The paper concludes that social-evaluative threat, not cognitive-emotional load alone, is what produces anxiety-related autonomic differences.","feed_headline":"Social anxiety leaves no skin-conductance trace under cognitive stress","feed_subtitle":"Resting-state skin conductance separates the groups; a demanding emotional task erases the gap.","key_machinery":"The carrying machinery is a standardised interval-based EDA pipeline that splits the first six minutes of the task into three two-minute phases (Early, Middle, Late) and compares each with a duration-matched two-minute baseline. Ten features spanning tonic, phasic, sympathetic, spectral, and nonlinear domains are extracted after a convex-optimisation decomposition of the signal into tonic and phasic components, then analysed with mixed ANOVAs and five classifiers (logistic regression, SVM, random forest, gradient boosting, and a 1D-CNN on the raw signal). The consistent temporal pattern across both statistical and machine-learning analyses—most discriminability at baseline, none during the task—is what carries the claim that context, not trait anxiety alone, governs the signal.","core_discovery":"On the paper's own terms, the central discovery is that emotionally salient cognitive load without explicit social evaluation is not enough to differentiate socially anxious from non-socially anxious individuals at the level of electrodermal activity. Both groups showed a substantial and statistically indistinguishable rise in arousal from baseline to task, with only transient interaction effects in the first two minutes (SCR amplitude and EDASymp) that did not persist. Multivariate classification confirmed the asymmetry: resting-state EDA carried a moderate group signal, while task-phase EDA fell to near chance across all five model types. The paper reads this as evidence that anxiety-related autonomic signatures are context-dependent, and that wearable EDA biomarkers for social anxiety may be informative mainly in socially evaluative or resting contexts.","pith_inferences":["Editorial extension: the resting-state signal that classifies SA versus NSA may be partly anticipatory arousal, because the baseline was recorded after consent and task briefing and immediately before practice; a pre-briefing rest recording would test whether the AUC 0.73 reflects stable trait physiology or imminent-task anxiety.","Editorial extension: the same participants and task could be run in two versions, one with neutral stimuli and one with explicit evaluative feedback, to isolate whether cognitive load interacts with social-evaluative threat—a design the current data do not cover.","Editorial extension: the transient crossover pattern in SCR amplitude (NSA increasing, SA slightly decreasing at task onset) suggests testable hypotheses about orienting responses and task engagement in social anxiety, but with N=50 and no correction across many tests it should be treated as hypothesis-generating.","Editorial extension: for wearable anxiety monitoring, a practical implication is to acquire a short rest segment before demanding tasks or to embed evaluative elements; otherwise the very signal of interest may be masked by a common arousal response."],"forward_implications":["Because the task masks group differences, wearable EDA screening for social anxiety should sample resting states or socially evaluative situations rather than generic cognitive load.","Averaging EDA over the whole task would have missed the transient early interactions; phase-level analysis is needed to see the brief group differences that do exist.","Attentional Control Theory receives only partial support from these data: cognitive load raises arousal, but anxiety-linked EDA dysregulation appears to require an evaluative component.","The decline in classification performance from baseline (average AUC 0.73) to task phases (at most 0.57) across all five model types makes the context-dependence pattern the robust finding, not any single classifier's accuracy.","A missing recovery phase means the design cannot rule out delayed group differences in return to baseline after the task ends."],"supporting_citations":[{"why":"Provides the EDA spectral measures and the task-inspired procedure for capturing sympathetic function, which the study adapts for its baseline/task design.","marker":"[30]"},{"why":"Companion paper that supplies the emotional 2-back task paradigm and the behavioural results showing no group differences in accuracy or response time.","marker":"[34]"},{"why":"Supplies the cvxEDA convex optimisation algorithm used to decompose skin conductance into tonic and phasic components.","marker":"[36]"},{"why":"Supplies the open-source toolbox used for preprocessing, decomposition, and feature extraction.","marker":"[37]"},{"why":"Validates the Social Interaction Anxiety Scale used to classify participants into SA and NSA groups.","marker":"[31]"},{"why":"Prior evidence on autonomic recovery and habituation in social anxiety that frames why baseline versus task contrasts matter.","marker":"[9]"},{"why":"Systematic review documenting inconsistent skin-conductance findings in social anxiety, motivating the context-dependence claim.","marker":"[17]"},{"why":"Attentional Control Theory, the framework the results only partially support.","marker":"[53]"}],"fun_headline_variants":["Resting skin conductance flags social anxiety; task stress does not","Social anxiety's EDA signature vanishes under cognitive load","Cognitive-emotional stress masks resting EDA differences in social anxiety","Wearable EDA: resting signal separates social anxiety, task signal doesn't"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The resting-state baseline is treated as neutral, but it was recorded after informed consent and a task briefing, immediately before practice, so it may capture anticipatory anxiety rather than rest.","fun_headline_variants_meta":{"raw":{"variants":["Resting skin conductance flags social anxiety; task stress does not","Social anxiety's EDA signature vanishes under cognitive load","Cognitive-emotional stress masks resting EDA differences in social anxiety","Wearable EDA: resting signal separates social anxiety, task signal doesn't"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000176,"raw_usage":{"total_tokens":1304,"prompt_tokens":974,"completion_tokens":330,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":257}},"tokens_in":590,"tokens_out":330,"duration_ms":4106,"temperature":1.0,"reasoning_tokens":257,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:59:26.982256+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record a true resting baseline before the participant is briefed about the task, or after a long habituation period, and rerun the classification. If the resting-state AUC drops to chance when anticipatory arousal is removed, then the reported baseline signal is an artifact of the experimental procedure rather than a stable social-anxiety signature. A second test: add an explicit social-evaluative condition to the same 2-back task; if group discrimination still fails to reach the resting-state level, the paper's claim that evaluation is necessary would be weakened.","supporting_citations":[{"cited_title":"Power spectral density analysis of electrodermal activity for sympathetic function assessment,","cited_arxiv_id":null,"evidence_quote":"Provides the EDA spectral measures and the task-inspired procedure for capturing sympathetic function, which the study adapts for its baseline/task design."},{"cited_title":"The impact of acute social stress on working memory updating in social anxiety disorder,","cited_arxiv_id":null,"evidence_quote":"Companion paper that supplies the emotional 2-back task paradigm and the behavioural results showing no group differences in accuracy or response time."},{"cited_title":"cvxeda: A convex optimization approach to electrodermal activity processing,","cited_arxiv_id":null,"evidence_quote":"Supplies the cvxEDA convex optimisation algorithm used to decompose skin conductance into tonic and phasic components."},{"cited_title":"Neurokit2: A python toolbox for neurophysiological signal processing,","cited_arxiv_id":null,"evidence_quote":"Supplies the open-source toolbox used for preprocessing, decomposition, and feature extraction."},{"cited_title":"Validation of the Social Interaction Anxiety Scale and the Social Phobia Scale across the anxiety disorders","cited_arxiv_id":null,"evidence_quote":"Validates the Social Interaction Anxiety Scale used to classify participants into SA and NSA groups."},{"cited_title":"Autonomic recovery and habituation in social anxiety,","cited_arxiv_id":null,"evidence_quote":"Prior evidence on autonomic recovery and habituation in social anxiety that frames why baseline versus task contrasts matter."},{"cited_title":"A systematic review of thermosensa- tion and thermoregulation in anxiety disorders,","cited_arxiv_id":null,"evidence_quote":"Systematic review documenting inconsistent skin-conductance findings in social anxiety, motivating the context-dependence claim."},{"cited_title":"Anxiety and cognitive performance: attentional control theory","cited_arxiv_id":null,"evidence_quote":"Attentional Control Theory, the framework the results only partially support."}],"review_version":1}