{"id":"8f594863-9e4b-48ae-bea0-f6f395bb1b91","arxiv_id":"2607.24754","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"CARE-MH is a unified evaluation framework showing that mental-health LLM benchmark results are strongly affected by evaluator-model stability and metric definitions.","lead":"A team built CARE-MH, a standardized way to evaluate mental-health chatbots that makes different benchmarks comparable. Re-running three major benchmarks showed results shift when evaluator models change, and that most disagreement comes from how each benchmark defines its scoring rubrics.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim that metric definitions are the primary driver of cross-benchmark disagreement is not isolated: RQ4 vs RQ5 confounds construct choice and evaluator family with metric definition.","rationale":"The reader's weakest assumption (LLM-as-judge validity without human calibration) is real, but the more load-bearing problem is that the paper's central causal attribution is not actually identified in the reported experiments. RQ4 and RQ5 are designed as a natural experiment, but they vary more than one factor at a time: RQ4 keeps the construct (Empathy) roughly constant while varying evaluator family and prompt text; RQ5 varies construct, metric definition, and evaluator family simultaneously. Therefore the larger inconsistency in RQ5 cannot be specifically attributed to 'metric definitions' as opposed to the different constructs being measured or the different evaluator families used to measure them. This directly undermines Takeaway 2 and the abstract's causality claim. The concrete test would settle whether definitional wording alone changes rankings. The paper's framework and reproducibility data are valuable, and the more conservative claims (evaluator-model changes shift scores; stability matters) are supported. The verdict remains CONDITIONAL because the central claim needs the proposed controlled comparison, but no harsher verdict is warranted given the framework's other contributions.","tokens_in":46535,"tokens_out":4929,"duration_ms":49692,"concrete_test":"Recompute Table 4 using a single evaluator model (e.g., Gemini-2.5-Pro) to evaluate the same SUT response sets under (a) the three existing metric prompts (CB Specificity, MB Relevance, MC Active Listening) and (b) several reworded variants of the same construct (e.g., three different phrasings of 'Empathy' with different scales/instructions). Compute rank correlations (e.g., Spearman) across conditions. If rankings remain stable across Empathy rewordings but diverge across the three different constructs, the Table 4 inconsistency is driven by construct choice, not metric definitions; if rewordings alone change rankings, the paper's attribution is supported. This directly isolates metric definition from construct and evaluator family.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim (abstract, Takeaway 2) is that cross-benchmark disagreement arises primarily from differences in metric definitions. The evidence is the contrast between RQ4 (Table 3) and RQ5 (Table 4): same-named Empathy yields consistent rankings across benchmark evaluator families, while Specificity/Relevance/Active Listening yield less consistent rankings. However, this contrast does not isolate metric definition. In RQ4, 'Empathy' is not defined identically across benchmarks — scales differ (1–5 vs 1–10) and wording differs (CounselBench vs MentalBench vs MentalChat16K prompts in Figures 7/11/16) — so consistency there shows robustness to wording/scale for a single construct. In RQ5, the three compared metrics are different constructs (tailoring vs on-topicness vs reflective listening), not the same construct with different definitions. The design also varies evaluator model family alongside the metric (CB/MB/MC evaluator sets), so model-family or prompt-template differences are not controlled. Hence the larger inconsistency in RQ5 could reflect genuine differences in what the benchmarks measure (construct mismatch) or evaluator-family/prompt confounds, rather than 'metric definitions' as the primary driver. This is load-bearing because the headline recommendation — adopt shared metric definitions — depends on the claim that definitional variance, not construct selection or evaluator variance, is the main cause of cross-benchmark disagreement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CARE-MH, a configurable evaluation framework for non-clinical mental-health LLM benchmarks. CARE-MH explicitly parameterizes prompt templates, evaluator instructions, generation settings, and metric definitions, and organizes benchmark-specific metrics into a unified taxonomy. Using this framework, the authors reproduce three benchmarks—CounselBench, MentalBench, and MentalChat16K—and report three main findings: (1) benchmark reproducibility depends strongly on evaluator and SUT model stability; (2) cross-benchmark disagreement arises primarily from differences in metric definitions; and (3) a unified evaluation design with shared metrics improves comparability and reproducibility. The evidence includes reproduction tables, cross-evaluation experiments (RQ3–RQ5), and a unified 16-metric evaluation across four evaluator families.","tokens_in":46808,"tokens_out":5031,"duration_ms":58734,"significance":"The paper makes a useful practical contribution: CARE-MH is a concrete, modular pipeline with explicit configuration tracking, and the reproduction and cross-evaluation artifacts are valuable for the mental-health LLM benchmarking community. The demonstration that deprecated evaluator models can substantially shift scores on identical responses (Figure 3) is a concrete, falsifiable observation that supports the reproducibility part of the message. The unified taxonomy and the release of prompts/configurations are also strengths. If the central causal claim about metric definitions were properly isolated, the paper would be an important call for standardized evaluation configurations. As it stands, the headline conclusion is not yet supported by the experimental design, though the underlying recommendations remain plausible.","major_comments":[{"comment":"The central claim that cross-benchmark disagreement 'primarily arises from differences in metric definitions' is not isolated by the RQ4/RQ5 comparison. In RQ5, the three columns differ simultaneously in metric construct (tailored advice vs. on-topicness vs. reflective listening), evaluator-model family (CB vs. MB vs. MC evaluator sets), prompt template, rating scale (1–5 vs. 1–10), and output format. Thus the larger inconsistency in Table 4 relative to Table 3 could be caused by construct mismatch, evaluator-family differences, prompt-format differences, or scoring-scale differences, not specifically by 'metric definitions.' To support the causal attribution, the authors would need to hold evaluator family and prompt template fixed while varying only the operational definition of a metric, or to vary one factor at a time. Without this, Takeaway 2 and the abstract overstate what the expe","section":"§4.2, Table 4, Takeaway 2"},{"comment":"RQ4 is not a clean control for RQ5. The 'same metric' Empathy is implemented with different wordings, rating scales, and evaluator prompts across the three benchmarks (e.g., CounselBench 1–5, MentalChat16K 1–10), and different evaluator models are used. The observed consistency under Empathy is informative—it shows robustness across those variations for one construct—but it does not by itself prove that metric definition is the cause of the RQ5 disagreement. Moreover, the paper does not provide a quantitative measure of consistency (e.g., rank correlation, Kendall's tau, or score variance) for RQ3–RQ5. Table 4 actually shows many SUTs with identical or near-identical rank orders across columns, so the claim that RQ5 exhibits 'larger inconsistency' needs a formal comparison rather than visual inspection.","section":"§4.2, Table 3 and Table 4"},{"comment":"The evaluation pipeline treats LLM-as-a-judge outputs as valid measurements of response quality without any human-validation or calibration evidence. The paper defines an evaluator as 'an LLM used to assess SUT responses' and then uses those scores as ground truth throughout Section 4. If LLM judges are biased by prompt wording, scale format, or evaluator family, then the observed 'disagreement' in RQ5 could reflect judge-model behavior rather than properties of the SUT responses or the metric definitions. At minimum, the authors should report judge self-consistency (e.g., repeated scoring), agreement with human annotations on a subsample, or a sensitivity analysis showing that the RQ5 pattern is not driven by evaluator-model artifacts. This is a correctness-risk concern, not a demand for a particular philosophical stance on LLM evaluation.","section":"§3.2, §4, Limitations"},{"comment":"Takeaway 3 recommends 'shared evaluation standards with standardized metric definitions, evaluator prompts, generation settings, and structured evaluation schemas.' This is a reasonable recommendation, but it is broader than what the experiments test. The unified evaluation in Section 4.3 changes many components at once (new rubric, new prompts, new output format, fixed temperature), so the observed improvement in comparability cannot be attributed to standardized metric definitions alone. The paper should either present this as a framework demonstration rather than a causal claim, or include a decomposition experiment showing which components drive the improvement.","section":"§4.3, Table 5, Takeaway 3"}],"minor_comments":[{"comment":"The normalized scores in parentheses are not explained. Please state the normalization formula (e.g., min–max scaling per column or per benchmark) and whether normalization is done before or after averaging across evaluators.","section":"Table 4"},{"comment":"In the GPT-4.1 panel, the Qwen-2.5 and Qwen-3 rows report identical values across all metrics. This looks like a copy-paste error or a genuine data anomaly; please check whether Qwen-3 was actually evaluated under this condition.","section":"Appendix Table 20"},{"comment":"The text says MentalChat16K is 'not fully reproducible' because the original evaluators are deprecated, but then proceeds to use replacement evaluators. Please clarify that the reported 'reproduction' is an adapted reproduction, not an exact replication, and that the original benchmark results cannot be directly compared without caveats.","section":"Section 4.1"},{"comment":"Several headings and paragraph markers contain stray non-text characters (e.g., '♂redoReproducibility', '♂searchRQ 1', '/balance-scaleCross-Benchmark Consistency'). These appear to be LaTeX artifacts and should be cleaned before publication.","section":"Throughout"},{"comment":"The Limitations section states that 'we report all data sources, prompts, generation parameters, models, and our code,' but the main text does not give a repository URL or artifact DOI. Please include the exact link or indicate that code will be released upon publication.","section":"Limitations"},{"comment":"The radar plots use different scales per axis but the axis labels are not visible in the main text. A shared scale or explicit axis range would help readers compare the magnitude of changes across metrics.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The paper's practical framework and reproduction data are valuable, but the headline causal claim ('primarily from differences in metric definitions') is not supported by the current experimental design. The authors should either add a controlled experiment that isolates metric definitions from evaluator family/prompt/scale, or substantially soften the causal language in the abstract and takeaways. The LLM-judge validity issue is a further correctness risk that should be addressed with at least a small human-agreement or self-consistency analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful piece of evaluation infrastructure work — a factorized framework plus the first systematic reproduction of three mental-health LLM benchmarks — and the evidence that model/evaluator stability affects scores is solid. The stronger claim, that cross-benchmark disagreement primarily arises from metric definitions, is not established by the reported design. I'd send it to review, and reviewers should push on exactly that point.\n\nWhat's new: CARE-MH explicitly parameterizes prompts, generation settings, metric definitions, and evaluators, and maps the three benchmarks' metrics onto a common taxonomy. The reproduction results are a real contribution, especially the MentalChat16K evaluator-replacement experiment, which shows large score shifts on identical responses when newer GPT/Gemini models replace deprecated ones. The paper also documents substitutions, token-limit adjustments, and failure rates, and it honestly flags several of its own limitations.\n\nWhere the soft spots are: the abstract and Takeaway 2 claim that metric definitions are the primary driver of cross-benchmark disagreement. The evidence is the contrast between same-named Empathy (consistent across benchmarks) and Specificity/Relevance/Active Listening (less consistent). But those are different constructs, not the same construct with different operational definitions. The comparison also changes evaluator family alongside the metric, and the \"same\" Empathy is not actually identical across benchmarks — different scales (1–5 vs 1–10) and different prompt wording appear in the appendix. So the causal attribution is confounded. That's the main structural issue.\n\nBeyond that: the study assumes LLM-as-a-judge outputs are valid measurements without any human calibration or agreement check; the sample sizes are modest (100–200 prompts for two of the three benchmarks); there are no error bars; and several deprecated models are substitutes for the originals. Code and exact dataset versions aren't public in this version. All of these are fixable or at least disclosable, and several are already acknowledged.\n\nWho this is for: anyone building or using mental-health LLM benchmarks, and broadly anyone doing LLM-as-a-judge evaluation. It deserves a serious referee. The framework is sensible, the reproduction data are valuable, and the confound is identifiable and addressable rather than fatal. I'd invite revision with a request to either weaken the causal claim or run an experiment that holds construct and evaluator family fixed while varying only metric definition. A small human-validation sanity check on the evaluator scores would also strengthen it considerably.","headline":"Useful framework and reproduction study, but the headline claim that metric definitions are the primary driver of disagreement is not actually isolated by the experiments.","tokens_in":47291,"tokens_out":1694,"would_cite":true,"duration_ms":20955,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Mental-health LLM benchmark disagreement is driven by metric definitions, not model quality.","keywords":["mental health LLM evaluation","benchmark reproducibility","metric definitions","LLM-as-a-judge","evaluation framework","cross-benchmark comparison","model stability","evaluation taxonomy"],"falsifier":"Recruit a panel of licensed clinicians to score the same set of system responses under the three benchmarks' original rubrics and under the unified rubric. If the clinicians' rankings show the same cross-rubric shifts the LLM judges showed, the metric-definition claim is supported; if human rankings stay stable across rubrics while LLM rankings flip, the disagreement is a judge-model artifact rather than a property of the metric definitions.","tokens_in":46391,"feed_emoji":"🧠","tokens_out":5710,"duration_ms":57823,"temperature":0.7,"pith_summary":"The paper sets out to explain why mental-health LLM benchmarks often disagree about which models are best, and to make such evaluation reproducible at all. It claims that the main culprit is not the models being tested but how each benchmark defines its evaluation metrics: semantically similar criteria such as specificity, relevance, and active listening are worded differently and produce different rankings. It also claims that benchmark outcomes depend heavily on the stability of the models involved—both the model being evaluated and the LLM judge that scores it—since replacing a deprecated judge with a newer one shifts scores even when all responses are identical. To establish this, it introduces CARE-MH, a framework that turns every evaluation choice (prompts, decoding settings, metric definitions, judge and evaluated models) into an explicit, swappable parameter, and it reproduces three leading benchmarks inside that framework. It concludes that future benchmarks should release complete evaluation configurations and standardized metric definitions, and it demonstrates that a unified 17-metric rubric yields more consistent cross-benchmark rankings.","feed_headline":"Cross-benchmark gaps trace to metric definitions, not models","feed_subtitle":"Evaluator and rubric choices, not model quality, move mental-health benchmark scores.","key_machinery":"The CARE-MH framework is the load-bearing mechanism: a factorized evaluation pipeline that separates prompt formatting, SUT response generation, evaluator judgment, and metric aggregation into four explicit stages, with a unified taxonomy that groups semantically related metrics from different benchmarks into six categories (therapeutic communication, content quality and problem fit, actionability, safety/ethics/scope, trustworthiness, and evaluation artifacts). Its role is to convert every hidden evaluation choice into a controlled variable, so the paper can swap one component at a time and attribute disagreements to a specific source—most importantly, to the wording and scoring scale of th","core_discovery":"On the paper's own terms, the central discovery is that cross-benchmark inconsistency in mental-health LLM evaluation is attributable primarily to metric definition differences rather than model behavior. Using CARE-MH, the authors reproduced three established benchmarks and found that when the same metric is defined identically, rankings are stable across benchmarks; when semantically related metrics are defined differently, rankings shift even though the underlying responses are unchanged. They further found that benchmark results are reproducible only when the original evaluator and system-under-test model versions are preserved, and that substituting newer LLM judges systematically lower","pith_inferences":["The same confounding likely applies beyond mental health: any LLM evaluation that relies on rubric-based LLM judges (e.g., medical advice, legal advice) may be sensitive to metric wording, and the CARE-MH parameterization could expose that.","A low-cost test of the paper's central attribution would be to have the same system responses scored by the same judge model under multiple paraphrased rubrics; if scores move with paraphrases, metric definition is confirmed as the causal driver.","Because CARE-MH uses LLM judges as ground truth without human calibration, its unified rubric may well be measuring a stable artifact of judge preference; requiring human-validated reference scores would strengthen the framework."],"forward_implications":["Benchmark scores are only interpretable relative to a full evaluation configuration; publishing a dataset without evaluator versions, prompts, and metric definitions makes results non-reproducible by design.","Replacing a deprecated evaluator model can change scores substantially even when the responses under test are identical, so longitudinal comparisons require frozen judges.","Leaderboards that compare models across differently worded rubrics may be ranking the rubric, not the models.","A shared metric taxonomy with explicit definitions can reduce cross-benchmark disagreement and improve comparability."],"fun_headline_variants":["Mental-health LLM rankings hinge on metric definitions","Same LLM, different scores? Check the metric definitions","Benchmark gaps in mental-health LLMs traced to metrics","It's not the LLM, it's the metric: mental-health eval"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper assumes that LLM-as-a-judge scores are valid, stable measurements of mental-health response quality, without calibrating those judgments against human expert ratings.","fun_headline_variants_meta":{"raw":{"variants":["Mental-health LLM rankings hinge on metric definitions","Same LLM, different scores? Check the metric definitions","Benchmark gaps in mental-health LLMs traced to metrics","It's not the LLM, it's the metric: mental-health eval"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1079,"prompt_tokens":602,"completion_tokens":477,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":346,"completion_tokens_details":{"reasoning_tokens":408}},"tokens_in":346,"tokens_out":477,"duration_ms":5502,"temperature":1.0,"reasoning_tokens":408,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T13:03:13.495599+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recruit a panel of licensed clinicians to score the same set of system responses under the three benchmarks' original rubrics and under the unified rubric. If the clinicians' rankings show the same cross-rubric shifts the LLM judges showed, the metric-definition claim is supported; if human rankings stay stable across rubrics while LLM rankings flip, the disagreement is a judge-model artifact rather than a property of the metric definitions.","supporting_citations":[],"review_version":1}