{"id":"f05944d1-f1da-463a-8417-ebb6ab4f345d","arxiv_id":"2501.03064","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"Introduces MENTAL-TRUST, a seven-level expert-annotated trust dataset for counseling dialogues, and benchmarks 14 models, reporting that fine-tuned smaller models outperform zero-shot LLMs.","lead":"This paper builds a counseling dataset, MENTAL-TRUST, where experts score each patient utterance on a seven-level trust scale. It benchmarks 14 language models and reports that fine-tuned smaller models predict these trust scores better than large or closed-source models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5.1's 5-class label set contradicts the 7-level ordinal dataset; all TrustBench accuracy numbers are suspect until resolved.","rationale":"I read the paper as a resource paper: the value is the dataset and the benchmark. The dataset's validity is a matter of annotation design, and the paper reports a reasonable Cohen's Kappa (0.77). The most immediately falsifiable component is the benchmark evaluation. The Section 5.1 label-space inconsistency is a concrete internal contradiction that affects every number in Table 3. Unlike the philosophical question of whether trust is observable in text, this can be checked by inspecting code and re-running experiments. The session-count mismatch and wrong model category in Figure 2 reinforce the need for a corrected revision. The reader's weakest_assumption (text-only trust observability and the 2.5 initialization) is legitimate but less diagnostic; the label mismatch is a smoking gun that should be resolved before the benchmark is trusted. Hence I agree with the CONDITIONAL verdict but identify a different primary concern.","tokens_in":13957,"tokens_out":4574,"duration_ms":44728,"concrete_test":"Release the evaluation code and exact label mapping; then re-run the full TrustBench with the seven ordinal labels from Section 3.2. Compare Mental-BART's accuracy and the model ranking against Table 3. Also include a persistence baseline (predict the previous trust level) and a majority-class baseline; if the persistence baseline approaches 89%, the benchmark's discriminative power is not demonstrated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central benchmark claim (Mental-BART at 89.03 accuracy) depends on a consistent label vocabulary between the dataset and the evaluation. Section 5.1 defines the task as predicting trust score ti ∈ {1, 2, 3, 4, 5}, while Section 3.2 defines seven ordinal levels (1, 1.5, 2, 2.5, 3, 3.5, 4), and Table 1 and Table 5 show half-integer labels. If the models were trained or evaluated on integer-only labels, the gold standard must have been collapsed or the model output space excluded valid classes; either way, the reported accuracies in Table 3 do not describe performance on the stated 7-level task. This is not a stylistic typo: it changes the number of classes, the loss function, and the chance accuracy. A second indicator of fragility is the session count (212 in the abstract, 167 in Table 2) and the mislabeling of Mental-BART as decoder-only in Figure 2. These inconsistencies, while individually correctable, collectively mean the experimental pipeline is not currently reproducible, so the headline 'smaller models outperform larger ones' is not yet established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MENTAL-TRUST, a dataset of counseling sessions in which patient utterances are annotated with seven ordinal trust levels, and TRUST-BENCH, a benchmark evaluating 14 language models on predicting these levels as an ordinal classification task. The authors report that a fine-tuned Mental-BART model achieves the highest accuracy (89.03%) and that smaller encoder-based models outperform larger decoder-only and closed-source models. They also present trust trajectory and topic analyses, and commit to releasing the dataset.","tokens_in":14182,"tokens_out":3995,"duration_ms":34315,"significance":"If the dataset and benchmark are reliable, the work addresses a meaningful gap in counseling dialogue research: quantifying the therapeutic bond as a dynamic trust trajectory. Strengths include the iterative annotation design with expert involvement (Cohen's Kappa 0.77), the inclusion of half-level trust labels to capture ambiguity, and the systematic comparison across model families. However, several internal inconsistencies currently prevent the results from being taken at face value.","major_comments":[{"comment":"Section 5.1 defines the label set as {1,2,3,4,5}, while Section 3.2 and Tables 1 and 5 use seven levels with half-integers (1, 1.5, 2, 2.5, 3, 3.5, 4). This discrepancy changes the number of classes, the chance accuracy, and the loss function. The reported accuracies in Table 3 are therefore not clearly attributable to the described 7-level task. The authors must specify the label set actually used for training and evaluation; if the 5-class set was used, they must explain how the half-level labels were collapsed and recompute all metrics accordingly.","section":"Section 5.1"},{"comment":"The dataset size is inconsistently reported. The abstract and Section 3 state 212 counseling sessions and 12.9K utterances, while Section 4 and Table 2 sum to 167 sessions and 10,172 utterances (116+17+34 sessions; 6,902+949+2,321 utterances). These mismatches make the resource description ambiguous and undermine the reproducibility of the benchmark split.","section":"Abstract and Section 4 / Table 2"},{"comment":"Section 3.3 reports only a single overall Cohen's Kappa (0.77) for the final annotation iteration. It does not state how many annotators labeled each session, how disagreements were adjudicated, or what per-label or per-level agreement was. The trajectory-level analyses in Table 4 and Figure 4 assume that the labels are reliable at every level; without per-level reliability evidence, the validity of these analyses is not established.","section":"Section 3.3"},{"comment":"The decision to initialize every conversation at trust level 2.5 (Section 3.2) is a modeling assumption that feeds directly into the trajectory statistics in Table 4 and the outcome classification in Section 6.2, where positive and negative outcomes are defined relative to the initial trust level. Because all trajectories are anchored at the same neutral point, the reported asymmetry between positive and negative jumps may be an artifact of this initialization rather than a property of trust dynamics. The authors should either justify this anchor with prior data or analyze sensitivity to it.","section":"Section 3.2 and Section 6.2"}],"minor_comments":[{"comment":"The caption refers to Mental-BART as a decoder-only model, but Table 3 and Section 5.2 classify it as an encoder-decoder model; this makes the trajectory comparison figure misleading.","section":"Figure 2 caption"},{"comment":"The references for Llama 3.1 and Phi-3.5 are incomplete: they appear as 'et al., 2024a' and 'et al., 2024b' without author names.","section":"References"},{"comment":"The subsection heading contains a typo: 'Enoder-only Methods' should be 'Encoder-only Methods'.","section":"Section 5.2"},{"comment":"The text states that validation utterances are slightly longer (65.21 tokens) than training (56.56) and test (52.71), but the table also shows that the average utterance length per speaker is higher for validation; the current phrasing is acceptable but should be checked for clarity.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The paper has a promising direction, but the label-set inconsistency and dataset size mismatch are load-bearing and prevent acceptance in the current form. It may be that the 5-class label set in Section 5.1 is a typo, but the authors must clarify the exact experimental setup and, if needed, rerun or re-report the benchmark. The annotation reliability details also need strengthening to support the trajectory-level claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real new resource—expert-annotated dynamic trust labels for counseling dialogues—but the paper as posted has enough internal inconsistencies that its headline accuracy numbers shouldn't be trusted yet.\n\nThe new thing here is MENTAL-TRUST, a dataset of counseling sessions annotated with seven ordinal trust levels plus topic-shift markers. That fills a real gap: prior trust measures in therapy are static outcome scales, not per-utterance trajectories. The annotation process is also a credit: three iterative rounds, kappa improving from 0.22 to 0.77, with clear guidelines and examples. Benchmarking 14 models gives a reasonable first map: fine-tuned BART-family models do well, zero-shot LLMs do poorly.\n\nThe problems are mostly in the reporting. The most serious: Section 5.1 defines the task as predicting trust scores from {1,2,3,4,5}, while Section 3.2 and the data examples define seven levels (1, 1.5, 2, 2.5, 3, 3.5, 4). If the models were trained on the 7-level scheme, the 5-level definition is wrong; if they were trained on a collapsed 5-level version, the paper should say so and the accuracy numbers apply to that collapsed task. Either way, the results in Table 3 don't currently describe the stated task. This isn't a cosmetic typo—it changes the number of classes, the loss, and the chance level.\n\nThere are smaller but compounding inconsistencies: the abstract says 212 sessions, Table 2 sums to 167; the utterance counts don't obviously match (12.9K vs. 10,172); Mental-BART is called decoder-only in Figure 2's caption but is listed as encoder-decoder in Table 3, and the reference points to Mental-LLaMA, not Mental-BART. Also, the \"smaller models beat larger ones\" claim is confounded: the small models are fine-tuned, the closed-source LLMs are zero-shot. That's a fine-tuning effect, not evidence about model size.\n\nThe fixed initial trust of 2.5 and the lack of per-label agreement or variance reporting are minor by comparison. Data and code aren't released yet, which is a limitation for a benchmark paper but not fatal if they appear on acceptance.\n\nOverall: the resource and the analysis are worth engaging. A serious reviewer could set this straight with a few targeted experiments or just a faithful description of what was done. I'd send it to review, but the authors need to fix the task definition and the consistency issues before the numbers carry any weight.","headline":"A genuinely useful new dataset for trust dynamics in counseling, but the paper's internal inconsistencies—especially the 5-vs-7-class definition—mean the benchmark numbers aren't interpretable as written.","tokens_in":14734,"tokens_out":3225,"would_cite":false,"duration_ms":27913,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Mental-BART, a compact model fine-tuned on mental-health text, best predicts patient trust on a seven-level ordinal scale in counseling dialogues.","keywords":["trust modeling","therapeutic bond","ordinal classification","counseling conversations","mental health NLP","MENTAL-TRUST dataset","TrustBench benchmark"],"falsifier":"Have two independent expert teams annotate the same held-out set of counseling sessions using the paper's guidelines and check whether Cohen's kappa stays near the reported 0.77, and additionally compare utterance-level trust labels against a post-session patient-reported trust questionnaire to see whether the ordinal trajectory predicts the questionnaire outcome.","tokens_in":13738,"feed_emoji":"🧠","tokens_out":4626,"duration_ms":41534,"temperature":0.7,"pith_summary":"The paper argues that the quality of a counseling session can be tracked by measuring a patient's trust in the therapist, defined as the willingness and openness to disclose sensitive personal material. To make this measurable, it introduces MENTAL-TRUST, a dataset of 212 counseling sessions in which every patient utterance carries one of seven expert-verified ordinal trust levels, together with topic-shift annotations. The authors frame trust prediction as an ordinal classification task and benchmark 14 models under the name TrustBench. Their central finding is that a relatively small, mental-health-tuned model, Mental-BART, reaches 89.03% accuracy and outperforms larger and closed-source models across most metrics. If correct, this gives therapists and AI-assisted counseling tools a concrete way to monitor whether the therapeutic bond is strengthening or weakening during a session.","feed_headline":"Fine-tuned mental-health model best at scoring patient trust","feed_subtitle":"A seven-level trust scale for counseling sessions lets smaller models beat GPT-4o and Gemini.","key_machinery":"The central object is the seven-level ordinal trust scale and its annotation protocol. Trust is defined by three observable textual behaviors: sharing personal, detailed, or sensitive information; opening up about relevant concerns; and staying aligned with the topic of concern. The four anchor levels are least trust (1), low trust (2), building trust (3), and achieved trust (4), with intermediate levels 1.5, 2.5, and 3.5 to absorb ambiguous cases. Conversations are seeded at the neutral midpoint 2.5, and annotators track upward and downward jumps from that point, additionally marking topic shifts. This scale is the mechanism that turns the therapeutic bond from a vague qualitative idea into a prediction target for ordinal classification models.","core_discovery":"The central claim is that trust in counseling can be operationalized as a dynamic, moment-by-moment ordinal trajectory visible in patient utterances, and that language models can learn to track this trajectory well enough to assist therapists. MENTAL-TRUST is the paper's evidence for this idea: seven ordinal levels, anchored by least trust, low trust, building trust, and achieved trust, plus intermediate half-levels, annotated through an iterative expert calibration that raised inter-annotator agreement from a Cohen's kappa of 0.22 to 0.77. TrustBench then evaluates 14 models, with the paper reporting that fine-tuned encoder-decoder models, especially Mental-BART, follow trust trajectories most closely while large and closed-source models lag substantially. The authors interpret this as evidence that specialized smaller models capture local, trust-specific textual indicators better than models relying on broad world knowledge.","pith_inferences":["A natural next step the paper does not take is to convert the ordinal scale into a real-time therapist alert, flagging drops of one or more levels as moments to recalibrate the session; the asymmetry between gradual building and sharp drops suggests such alerts would be rare but clinically meaningful.","Because the scale is defined purely from text, the same annotation protocol could be tested on multilingual or multimodal sessions where vocal tone and facial expression are available; the paper itself notes the text-only limitation, so this is an extension the authors imply rather than test.","The convention of initializing every conversation at trust level 2.5 is a testable assumption: re-annotating a sample with different starting priors, or letting annotators infer an initial level from the first exchanges, would show how much of the reported trajectory statistics depend on that choice.","The benchmark uses standard cross-entropy-style losses; an ordinal-aware loss that penalizes larger trust jumps more heavily could plausibly improve models, since the paper's own analysis shows that trust decreases are typically larger than trust increases."],"forward_implications":["Mental-BART, a mental-health fine-tuned BART, achieves 89.03% accuracy and the best scores on seven of nine metrics, showing that the ordinal trust task is learnable with moderate-size models.","Closed-source large models such as GPT-4o and Gemini 1.5 score around 23% accuracy, below even the smallest fine-tuned encoders, suggesting that general-purpose instruction-following does not transfer to fine-grained trust rating.","Trust trajectories have a positive bias: positive level changes are about twice as common as negative ones, with average upward steps near +0.5 and downward steps near -1.0, meaning trust builds gradually but can drop sharply.","Stable plateaus of about 8 to 9 consecutive utterances are common, indicating that trust levels are not noisy moment-to-moment states but settle into temporary equilibria.","Sessions with positive therapeutic outcomes tend to stay on coherent core topics, while negative-outcome sessions feature scattered topics and frequent digressions, linking topic alignment to trust maintenance."],"supporting_citations":[{"why":"Provides the HOPE counseling dataset, the source of the 212 sessions and 12.9K utterances that MENTAL-TRUST extends with trust annotations.","marker":"Malhotra et al., 2022"},{"why":"Supplies Mental-BART, the mental-health fine-tuned model that the paper reports as the best-performing system on seven of nine metrics.","marker":"Yang et al., 2023"},{"why":"Defines BERT, a standard encoder-only baseline whose fine-tuned performance helps establish the benchmark comparison.","marker":"Devlin et al., 2019"},{"why":"Defines RoBERTa, an encoder-only baseline used to test the effect of pretraining strategy on trust prediction.","marker":"Liu et al., 2019"},{"why":"Defines DeBERTa, the best encoder-only competitor in the benchmark, providing the main within-category comparison for Mental-BART.","marker":"He et al., 2021"},{"why":"Defines BART, the base encoder-decoder architecture that Mental-BART fine-tunes, making it a load-bearing baseline for the top result.","marker":"Lewis et al., 2019"},{"why":"Supplies Mental-BERT, another mental-health specialized model whose performance is compared against the best system.","marker":"Ji et al., 2021"},{"why":"Presents an earlier static psychometric scale for patient trust and respect, which the paper explicitly contrasts with its dynamic trajectory conceptualization.","marker":"Crits-Christoph et al., 2019b"}],"fun_headline_variants":["Smaller models beat GPT-4o and Gemini at therapy trust scoring","Seven-level trust scale helps small models outscore big LLMs in counseling","Mental-BART tops GPT-4o on therapy trust benchmark","Fine-tuned encoder-decoder best at tracking therapy trust"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset assumes that a patient's trust can be read from the text of their utterances and topic alignment alone, without any nonverbal cues, and that expert annotators can rate it consistently; if that assumption fails, the labels and every model comparison are not actually measuring trust.","fun_headline_variants_meta":{"raw":{"variants":["Smaller models beat GPT-4o and Gemini at therapy trust scoring","Seven-level trust scale helps small models outscore big LLMs in counseling","Mental-BART tops GPT-4o on therapy trust benchmark","Fine-tuned encoder-decoder best at tracking therapy trust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000341,"raw_usage":{"total_tokens":1866,"prompt_tokens":922,"completion_tokens":944,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":879}},"tokens_in":538,"tokens_out":944,"duration_ms":94141,"temperature":1.0,"reasoning_tokens":879,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:57:07.576825+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent expert teams annotate the same held-out set of counseling sessions using the paper's guidelines and check whether Cohen's kappa stays near the reported 0.77, and additionally compare utterance-level trust labels against a post-session patient-reported trust questionnaire to see whether the ordinal trajectory predicts the questionnaire outcome.","supporting_citations":[],"review_version":1}