{"id":"63186f24-9b89-4053-92b6-4aa3700ef501","arxiv_id":"2607.04350","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A dense mixture-of-experts model guided by training-only weak evidence-layout priors outperforms single-detector baselines on Chinese and English user-level depression detection.","lead":"WPG-MoE routes social-media users to specialized depression detectors using soft evidence-layout priors learned only at training time. It improves detection over strong baselines while keeping inference to PHQ-9 templates and a shared LLM backbone.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Gains may be driven by Path-A privileged evidence rather than by soft routing on stable layouts.","rationale":"The reader correctly flags Path-A stability and SWDD label correction as the softest assumptions, and CONDITIONAL is the right overall call: matched-backbone gains (Table 5), human layout diagnostics (Fig. 3), controlled mixing (Fig. 4), and seed spreads (A.12) are unusually careful for this area and do not contradict the claim that dense specialization helps. I only partially agree on the weakest assumption. The audits already show substantial A/B agreement and non-trivial weak-prior\to reference alignment; the sharper internal risk is causal attribution. Table 6 shows MoE and Path A dominate the drop, while weak priors and route loss are small. That leaves open that LUPI-style privileged evidence construction, not soft layout routing, is doing most of the work. The concrete test isolates that mechanism without requiring new data or clinical labels. If the test fails (gains collapse without real π/L_route), the reader’s concern is upgraded; if it holds, the paper’s MoE+LUPI claim is stronger than the layout narrative, and CONDITIONAL remains appropriate mainly for missing code/release and non-clinical evaluation rather than for a broken central comparison.","tokens_in":36388,"tokens_out":701,"duration_ms":7229,"concrete_test":"Re-train the full SWDD and Twitter models with Path A retained for candidate/block construction and L_evi, but replace π with either (i) a constant prior or (ii) a random 3-vector renormalized per user, and set α=0 (no L_route). Compare to Table 6’s w/o Weak Priors / w/o Route Loss and to full Ours on the same seeds. If AUPRC stays within ~0.01 of full Ours, the layout-prior mechanism is not load-bearing for the headline gains.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that post-screening heterogeneity is a real failure mode and that weak-prior-guided dense MoE under LUPI fixes it while remaining deployable with only Path B. The paper’s own ablations make the soft-prior story less secure than the MoE story: on SWDD, w/o MoE drops AUPRC 0.830\to0.693, w/o Path A drops to 0.758, while w/o Weak Priors and w/o Route Loss stay high (0.800 / 0.818; Table 6). That ordering is consistent across datasets. So the load-bearing risk is not that layouts are meaningless, but that Path A’s rich structured fields (and the evidence blocks built from them) mainly act as stronger training-time supervision/features, while the three soft layouts and L_route are secondary. If so, the LUPI framing still holds, but the paper’s distinctive claim—that clinically grounded weak priors for self-disclosure / episode / sparse layouts are what make routing work—overstates what the ablations isolate. The human audits (A.1/A.3) support that Path-A fields are recoverable, not that those three tendencies are the causal mechanism of the Table 4 gains.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that after risk-post screening, depressed users still express risk through heterogeneous evidence layouts (self-disclosure, episode-supported, sparse high-risk, mixed), so a single detector averages away localized cues. It proposes WPG-MoE: a dense five-expert MoE on a shared LLM backbone, softly routed by user-level weak priors π derived from training-only LLM-structured Path-A fields, cast as LUPI so inference uses only PHQ-9 template screening (Path B), the shared backbone, history segments, and lightweight statistics. Contributions are (i) diagnosing post-screening heterogeneity, (ii) weak-prior-guided dense routing without hard subtypes, and (iii) a deployable LUPI split. Support includes three datasets under a unified holdout, seven baselines, matched-backbone replacements, ablations, human layout diagnostics (κ≈0.71), Path-A audits, controlled mixing, gate analyses, seed-wise spreads, and case studies.","tokens_in":36711,"tokens_out":994,"duration_ms":16940,"significance":"If the result holds, this is a useful systems contribution for clinical NLP: it reframes user-level depression detection as a post-screening specialization problem and shows a practical LUPI recipe that keeps costly LLM annotation out of deployment. Strengths that should be credited include multi-dataset controlled comparison (Table 4), matched-backbone isolation of architecture vs encoder (Table 5 / A.9), human evidence-layout diagnostics with substantial agreement (Table 3; A.8), Path-A field audits (A.1/A.3), full ablations and seed-wise reliability (Tables 6, 23–24), and explicit train–deploy alignment mechanisms (Table 1). The work is significant for mental-health screening pipelines even if the three named layouts are better treated as soft inductive biases than as validated clinical subtypes.","major_comments":[{"comment":"Table 6 (§4.4) and the full matrix in Table 23 show a consistent component ordering that undercuts the paper’s distinctive mechanism claim. On SWDD, removing dense MoE drops AUPRC 0.830→0.693 and removing Path A drops to 0.758, while w/o Weak Priors (0.800) and w/o Route Loss (0.818) remain close to full model. The same pattern holds on Twitter and eRisk25. The Abstract, §1 contributions, and §2.3 present clinically grounded weak priors for self-disclosure / episode / sparse layouts as what makes routing work; the ablations instead isolate multi-view MoE capacity plus privileged Path-A supervision as the load-bearing pieces, with L_route and π as secondary shapers. The central claim remains defensible under a broader LUPI+dense-MoE reading, but the manuscript should restate the mechanism to match what Table 6 isolates, or add a control that keeps Path-A features while randomizing/shuffli","section":null},{"comment":"§2.3 Eq. (5) and Appendix A.3: the three soft layouts are induced from the same Path-A fields used to build evidence blocks and silver evidence labels. Human audits show Path-A fields and user-layout assignment are recoverable (κ≈0.65–0.71; assignment Acc. 0.772), which rules out pure noise, but does not establish that these three tendencies are the causal routing structure rather than convenient summaries of privileged score mass. Controlled mixing (Fig. 4) and gate plots (Figs. 5–6) are compatible with overlapping multi-view capacity; they do not falsify a “richer training supervision” alternative. A load-bearing revision is either (i) an experiment that freezes Path-A candidate quality while ablating layout-specific prior structure, or (ii) a clearer claim that π is a soft training regularizer, not a validated clinical typology.","section":null},{"comment":"§3.1 / Appendix A.7: SWDD results use a manually corrected self_reported split (399 raw 0→1, 16 raw 1→0). This is responsible, but main Table 4 and transfer cells that train on SWDD are not accompanied by a sensitivity check on the raw vs corrected labels for WPG-MoE and the strongest baselines. Because Path-A priors and self-disclosure routing lean on self-report structure, the SWDD gains (and SWDD→* transfer) need a short raw-label or leave-correction-out control so the improvement is not partly an artifact of the audit that also defines the self-disclosure slice.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: after risk-post screening, a single user vector still averages away sparse and episode-like evidence, and a dense five-expert MoE on a shared backbone fixes a lot of that under a deployable PHQ-9 interface. That is a real, practical contribution for computational mental-health NLP, not a rehash of Santos-style MoE or another screening paper.\n\nWhat is new is the LUPI packaging. Path A (offline Qwen structured fields) builds training-only weak priors and evidence blocks; inference keeps only Path B templates, the shared encoder, and history segments. They test this carefully: three datasets under a unified holdout, seven baselines, matched-backbone replacements (mDeBERTa, Qwen, Ministral), full ablations, human layout audits (κ≈0.71), controlled mixing, gate plots, seed spreads, and anonymized cases. Gains are consistent; removing MoE hurts most. The human diagnostics and mixing curves make the heterogeneity claim more than rhetoric.\n\nThe soft spot is proportional, not fatal. Ablations show Path A and dense MoE carry the lift; weak priors and route loss are smaller (SWDD AUPRC: full 0.830, w/o MoE 0.693, w/o Path A 0.758, w/o priors 0.800, w/o route loss 0.818). So the paper’s distinctive story—that clinically grounded self-disclosure / episode / sparse priors are what make routing work—overstates what the numbers isolate. The three layouts are recoverable and useful soft hints, not proven causal subtypes. SWDD’s author-corrected labels and the proprietary Path-A scorer also mean the training signal is not fully transparent, and there is no public code/data claim in the manuscript. Evaluation is controlled holdout, not clinical outcomes or official eRisk early-risk protocol—which they themselves flag.\n\nMath and citation pattern look fine: standard MoE + LUPI, honest related work, no load-bearing contradiction. Free parameters (loss weights, drop rates, K(n)) are ordinary for the area.\n\nThis is for people building deployable user-level detectors who care about post-screening specialization. I would bring it to reading group, cite the LUPI/MoE framing if I work in this space, and send it to peer review. Expect referees to push on mechanism isolation and release, not desk-reject.","headline":"Solid empirical methods paper: dense MoE under LUPI is the real win; the three soft layouts are secondary scaffolding, not the load-bearing mechanism.","tokens_in":37377,"tokens_out":589,"would_cite":true,"duration_ms":8512,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"After screening, depressed users express risk in different layouts; one flat detector averages them away, so weak-prior routing to dense experts improves detection while staying deployable.","keywords":["depression detection","social media","mixture of experts","weak priors","learning using privileged information","PHQ-9 screening","user-level classification","evidence heterogeneity"],"falsifier":"On a held-out set with human-adjudicated evidence layouts, replace weak-prior-guided dense experts by a single shared classifier (or by unguided MoE) while keeping the same backbone and PHQ-9 screening; if the AUPRC/F1 gap on sparse and episode-supported slices disappears or reverses, the central claim fails.","tokens_in":37221,"feed_emoji":"🧠","tokens_out":720,"duration_ms":6755,"temperature":0.7,"pith_summary":"Social-media depression detectors usually improve the screening step—picking risk posts, grounding symptoms, building clinical features—then hand everything to a single classifier. This paper argues that the real failure is what happens next: users still express risk differently (explicit self-disclosure, recurring episode-like blocks, or a few sparse high-intensity posts), and a monolithic model averages those patterns into one representation and one decision boundary. That averaging dilutes localized cues and especially hurts non-self-disclosing users. WPG-MoE answers with a dense mixture-of-experts on a shared language-model backbone. During training, offline LLM-extracted structure supplies soft user-level weak priors that gently route each user toward experts matched to those evidence layouts; at inference the system keeps only PHQ-9 template screening and the shared backbone, casting the gap as learning using privileged information. On Chinese and English user-level datasets the method beats strong screening-based baselines and shows readable gate behavior, while ablations and case studies tie the gains to preserving sparse and episode-supported cues that flat detectors blur.","feed_headline":"One flat detector averages away how users show depression risk","feed_subtitle":"Weak-prior MoE routes sparse and episode cues to specialists, yet inference keeps only PHQ-9 screening.","key_machinery":"WPG-MoE: dual-path evidence construction (privileged Path-A LLM scores for training; deployable Path-B PHQ-9 template similarity at inference) plus a dense five-expert mixture whose gates are softly guided by a three-component weak prior (self-disclosure, episode-supported, sparse high-risk) under LUPI, so privileged structure shapes routing without being required at deployment.","core_discovery":"Post-screening evidence heterogeneity is a distinct failure mode for user-level social media depression detection: after risk posts are selected, depressed users still present overlapping but different evidence layouts, and a single detector that averages them dilutes localized signals. Softly routing users to dense experts with training-only weak priors induced from those layouts, under a LUPI train–deploy split that leaves only PHQ-9 screening at inference, recovers performance that flat detectors lose.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Flat detectors dilute heterogeneous depression evidence layouts","One classifier averages away sparse and episode depression cues","Weak-prior MoE stops single models from washing out risk signals","Post-screening layouts diverge; flat detectors lose localized cues","Dense experts recover what averaging erases in user depression detection"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The three soft evidence layouts pulled from the offline LLM scorer are stable, meaningful routing tendencies rather than artifacts of the scorer or residual label noise in the training data.","fun_headline_variants_meta":{"raw":{"variants":["Flat detectors dilute heterogeneous depression evidence layouts","One classifier averages away sparse and episode depression cues","Weak-prior MoE stops single models from washing out risk signals","Post-screening layouts diverge; flat detectors lose localized cues","Dense experts recover what averaging erases in user depression detection"]},"model":"grok-4.5","effort":"low","cost_usd":0.00789,"raw_usage":{"total_tokens":1897,"prompt_tokens":777,"num_sources_used":0,"completion_tokens":79,"cost_in_usd_ticks":78900000,"prompt_tokens_details":{"text_tokens":777,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1041,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":777,"tokens_out":79,"duration_ms":9561,"temperature":1.0,"reasoning_tokens":1041,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T19:52:35.033866+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a held-out set with human-adjudicated evidence layouts, replace weak-prior-guided dense experts by a single shared classifier (or by unguided MoE) while keeping the same backbone and PHQ-9 screening; if the AUPRC/F1 gap on sparse and episode-supported slices disappears or reverses, the central claim fails.","supporting_citations":[],"review_version":1}