{"id":"e7070466-c708-4f44-b437-3d65a277fbea","arxiv_id":"2605.04873","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Semantic projection of Sentence-BERT embeddings onto axes from validated clinical scales yields continuous scores for depression, anxiety, and worry that correlate with standard measures, especially in structured text formats.","lead":"This paper introduces a theory-driven unsupervised method that projects text embeddings onto semantic axes derived from clinical scale items to score psychological states like depression and anxiety. A smart generalist might read it for a potentially more interpretable and scalable way to assess mental health from language without supervised training.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Semantic axes from clinical scale items risk capturing lexical overlap rather than independent psychological constructs","rationale":"The reader's weakest assumption is precisely the load-bearing point. The structural dependence between axis construction and validation criterion makes the evidence for an 'interpretable and scalable alternative' weaker than presented, independent of any specific numerical results in the full text.","tokens_in":1751,"tokens_out":287,"duration_ms":45875,"concrete_test":"Re-run all reported correlations after ablating participant responses to remove any tokens overlapping with the anchor lexicon; if the association with clinical measures drops by >0.15 or loses significance while split-half reliability remains high, the headline claim that the method measures the target states (rather than lexical match) does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that projection onto axes defined by lexical anchors and scale items yields scores that validly measure depression/anxiety/worry. Because anchors are drawn directly from the same validated scales used for criterion validation, correlations may arise from surface-level semantic similarity to scale language rather than from the embedding space encoding the underlying constructs. This is especially plausible given stronger results for constrained response formats (selected/written words, phrases) versus free text, and the paper does not report controls such as orthogonal axes, expert-rated construct validity independent of the scales, or ablation of anchor-specific terms.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces a theory-driven unsupervised framework for assessing psychological states (depression, anxiety, worry) from text. It defines semantic axes in Sentence-BERT space using lexical anchors and items drawn from validated clinical scales, projects participant responses (across selected words, generated words, phrases, and free text) onto these axes to produce continuous scores, and evaluates them via correlations with the same clinical measures, split-half reliability, attenuation corrections, Wasserstein distributional comparisons, and benchmarks against VADER sentiment analysis. Stronger associations are reported for structured response formats, with improvements from sentence-level aggregation in free text; the work positions semantic projection as an interpretable, scalable alternative to supervised language models.","tokens_in":1858,"tokens_out":557,"duration_ms":55087,"significance":"If the validity issues can be addressed, the approach would represent a meaningful contribution by offering a fully unsupervised, theory-grounded method that avoids the need for large labeled training sets and provides direct interpretability through explicit semantic axes. The attention to response format and aggregation strategies highlights practical considerations for language-based assessment that are often overlooked.","major_comments":[{"comment":"Abstract and implied Methods (axis construction): The semantic axes are defined using items from the same validated clinical scales later employed for criterion validation. This creates a partial circularity risk, as reported correlations may largely reflect lexical/semantic overlap between the input text and scale language rather than independent capture of the underlying constructs. The pattern of stronger results for constrained formats (selected/written words, phrases) versus free text is consistent with surface-level similarity rather than deep construct measurement. No controls (orthogonal axes, independent expert construct-validity ratings, or ablation of anchor-specific terms) are described.","section":"Abstract and Methods"},{"comment":"Abstract and Results: No sample size, demographic details, exact mathematical definition of the projection operation (e.g., how anchors are combined into an axis vector), effect sizes with confidence intervals, or statistical controls for confounds (text length, lexical diversity) are provided. These omissions make it impossible to assess the magnitude, robustness, or generalizability of the claimed associations and reliability coefficients.","section":"Abstract and Results"}],"minor_comments":[{"comment":"The abstract references attenuation corrections and Wasserstein distance but does not report the numerical outcomes or implementation details for these analyses.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript sits at the intersection of computational linguistics and psychometrics; while the topic fits cs.CL, the absence of open data, code, or precise formulas limits independent verification and may warrant a request for supplementary materials if revised."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their thoughtful and constructive review. We address each major comment point by point below, providing the strongest honest defense of the manuscript while indicating where revisions will be incorporated.","responses":[{"response":"We appreciate the referee's identification of this methodological consideration. The axes are constructed from validated scale items precisely to ground the projections in established psychological theory, with participant responses generated independently and embedded via Sentence-BERT to capture semantic rather than purely lexical similarity. The observed pattern across response formats is consistent with our interpretation that more structured prompts better elicit construct-relevant language. However, we acknowledge that lexical overlap remains a plausible alternative explanation. In the revision we will add ablation analyses that remove scale-specific anchor terms, introduce comparisons to orthogonal axes derived from unrelated constructs, and expand the discussion of construct validity versus surface similarity. Expert ratings would require new data collection and are noted as a direction for future work.","revision_made":"partial","referee_comment":"[Abstract and Methods] Abstract and implied Methods (axis construction): The semantic axes are defined using items from the same validated clinical scales later employed for criterion validation. This creates a partial circularity risk, as reported correlations may largely reflect lexical/semantic overlap between the input text and scale language rather than independent capture of the underlying constructs. The pattern of stronger results for constrained formats (selected/written words, phrases) versus free text is consistent with surface-level similarity rather than deep construct measurement. No controls (orthogonal axes, independent expert construct-validity ratings, or ablation of anchor-specific terms) are described."},{"response":"We regret that these elements were not foregrounded in the abstract. The full Methods section defines the projection as the cosine similarity between a response embedding and the axis vector (the mean Sentence-BERT embedding of the lexical anchors and scale items). Sample size and demographics appear in the Participants subsection. In the revised manuscript we will update the abstract to report sample size, key demographics, and a concise statement of the projection operation. We will also add effect sizes with 95% confidence intervals throughout the Results and include regression models that control for text length and lexical diversity (type-token ratio) as potential confounds. These changes will directly address the concerns about transparency and robustness.","revision_made":"yes","referee_comment":"[Abstract and Results] Abstract and Results: No sample size, demographic details, exact mathematical definition of the projection operation (e.g., how anchors are combined into an axis vector), effect sizes with confidence intervals, or statistical controls for confounds (text length, lexical diversity) are provided. These omissions make it impossible to assess the magnitude, robustness, or generalizability of the claimed associations and reliability coefficients."}],"tokens_in":1475,"tokens_out":575,"duration_ms":63590,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this paper builds semantic axes directly from items on standard depression, anxiety, and worry scales, embeds text with Sentence-BERT, and projects responses onto those axes to produce scores. It tests the approach on several response formats and reports correlations with the original clinical measures plus split-half reliability and a VADER comparison. That combination of theory-driven axes and projection across formats is the actual novelty; most prior work trains supervised models on questionnaire targets instead. The results look stronger for constrained formats like selected words or phrases than for raw free text, and sentence-level aggregation helps the latter. Those patterns are worth noting for anyone thinking about response design in language-based assessment. The approach is at least trying to stay interpretable and avoid fitting to specific training distributions. The soft spots are more substantial. Because the axes are defined from the same scale items later used for criterion validation, the reported associations could largely reflect surface lexical overlap rather than independent measurement of the constructs. The fact that performance drops in free text is consistent with that possibility. The abstract supplies no effect sizes, no sample details, no exact projection formula, and no controls such as orthogonal axes or external construct-validity checks. Without those pieces it is hard to judge how much new signal is actually being captured. This is the sort of paper that might interest computational psychologists or NLP groups working on mental-health measurement who want to move beyond black-box prediction. A reader could extract the basic projection idea and try it on their own data, but the current evidence is too preliminary to treat as a settled method. I would send it to peer review because the core framing is coherent and the authors are engaging with real limitations of supervised alternatives, even if the execution needs tighter validation and more transparent reporting.","headline":"Semantic projection onto clinical-scale axes gives an unsupervised scoring method but the validation looks partly circular and under-specified.","tokens_in":2330,"tokens_out":415,"would_cite":false,"duration_ms":49654,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Semantic projection onto axes from clinical scale items produces psychological scores from text that align with standard clinical measures.","keywords":["semantic projection","language-based assessment","psychological states","depression","anxiety","worry","unsupervised measurement","clinical scales"],"falsifier":"A replication study in which projection scores show no reliable correlation with participants' scores on the same clinical scales used to construct the axes, or fail split-half reliability checks.","tokens_in":2635,"feed_emoji":"🧠","tokens_out":661,"duration_ms":53275,"temperature":0.7,"pith_summary":"The paper develops an unsupervised method that turns natural language responses into scores for depression, anxiety, and worry by projecting sentence embeddings onto semantic axes. Each axis is built directly from lexical anchors and items drawn from existing validated questionnaires, so the resulting scores stay tied to established clinical definitions without any task-specific training data. Strong correlations emerge between these projection scores and participants' questionnaire results, especially when responses come in structured formats such as word selections or short phrases. The approach therefore offers a way to assess psychological states that remains interpretable and does not require retraining models for every new population or context.","feed_headline":"Text projection on clinical axes matches questionnaire scores","feed_subtitle":"Embedding responses onto axes built from scale items yields interpretable scores for depression and anxiety that align with standard tests, ","key_machinery":"semantic projection of Sentence-BERT embeddings onto axes defined by lexical anchors and items from clinical scales","core_discovery":"Psychological constructs are operationalized as interpretable semantic axes derived from lexical anchors and items taken from validated clinical scales for depression, anxiety, and worry. Participant responses in several formats are embedded with Sentence-BERT and projected onto these axes to yield continuous scores. These scores exhibit strong associations with the original clinical measures, particularly for structured formats, while free-text responses improve when processed at the sentence level rather than as whole documents. The results position semantic projection as a theory-driven, fully unsupervised alternative to supervised language models for psychological assessment.","pith_inferences":["The same axis-construction process could be applied to additional psychological constructs simply by substituting the relevant scale items.","Because the axes are built from public clinical scales, the approach might support measurement in settings where administering full questionnaires is impractical.","Multilingual embeddings could allow the same clinical-scale anchors to generate comparable scores across languages without new validation data."],"forward_implications":["Structured response formats such as selected words, written words, and phrases produce stronger alignment with clinical measures than whole free-text responses.","Sentence-level aggregation markedly improves results on free-text input compared with document-level projection.","The method supplies continuous, interpretable scores without supervised training on labeled questionnaire data.","Direct comparisons to lexicon-based sentiment tools and distributional checks support its reliability as an assessment technique."],"fun_headline_variants":["Projecting text onto clinical scale axes matches questionnaire results","Semantic projection of responses yields scores aligned with clinical measures","Language embeddings on anxiety depression axes correlate with standard tests","Axis projection from scale items tracks depression anxiety questionnaire scores"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Lexical anchors and items from validated clinical scales, when placed in Sentence-BERT space, accurately and fully represent the target constructs of depression, anxiety, and worry.","fun_headline_variants_meta":{"raw":{"variants":["Projecting text onto clinical scale axes matches questionnaire results","Semantic projection of responses yields scores aligned with clinical measures","Language embeddings on anxiety depression axes correlate with standard tests","Axis projection from scale items tracks depression anxiety questionnaire scores"]},"model":"grok-4.3","cost_usd":0.004574,"raw_usage":{"total_tokens":2210,"prompt_tokens":707,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":45740500,"prompt_tokens_details":{"text_tokens":707,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1442,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":707,"tokens_out":61,"duration_ms":23540,"temperature":1.0,"reasoning_tokens":1442,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-08T17:14:46.462678+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A replication study in which projection scores show no reliable correlation with participants' scores on the same clinical scales used to construct the axes, or fail split-half reliability checks.","supporting_citations":[],"review_version":1}