{"id":"2fa7720e-d5fa-477c-986c-709b7736fc30","arxiv_id":"2608.02046","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"CompanionBench, a real-data-grounded bilingual benchmark with a hidden disclosure gate and IRT-corrected judging, ranks 28 AI companions and finds most fail to earn deeper disclosure, often substituting warmth for substance.","lead":"This paper introduces CompanionBench, a bilingual benchmark that evaluates AI emotional companions through a theory-built rubric and a simulated user whose hidden disclosure gate changes the conversation depending on the AI's behavior. It ranks 28 chatbots and finds that most substitute warm, encouraging language for substantive relational support, with role-play chatbots ranking lowest.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The disclosure gate and simulator have no external validation; the 2% deep-trust and warmth-over-substance findings could be artifacts of hand-set gate thresholds and single-judge labels.","rationale":"The reader's weakest assumption names exactly the load-bearing point: the gate/simulator mapping to real user trust is hand-authored and unvalidated. My pass agrees and sharpens it. The paper's strongest empirical claims — rankings reproducible, role-play agents at the bottom, no deep trust in 20 turns, warmth over substance — all depend on the environment (disclosure gate plus simulator) rewarding the right behaviors and punishing the wrong ones. If the gate's thresholds or the LLM's ai_move labels diverge from real user disclosure dynamics, every downstream conclusion inherits that divergence. The paper is unusually transparent about this: §9 admits no human or counselor baseline and no external-validity studies. The proposed test uses data the authors already possess (the real corpus behind the simulator) to falsify or support the gate as a transition model; it does not require collecting new human data, though a human/counselor study would be the strongest follow-up. I find no internal inconsistency or artifact in the release plan, and the internal evidence (test-retest, permutation responsiveness, IRT agreement) is credible and honestly reported. The condition on external validation is therefore sufficient and appropriate; no verdict change is needed.","tokens_in":38152,"tokens_out":6575,"duration_ms":69043,"concrete_test":"Use the authors' held-out real corpus to validate the gate: sample real AI-turn/user-next-turn pairs with matching persona/scenario fields, classify each AI turn with the gate policy as advancing/neutral/violating, and compare the predicted next-turn disclosure depth to the actual user depth transition (ordinal agreement or transition log-loss). Also run the same check on the trained simulator's outputs to separate gate error from simulator error. If the gate's transitions do not significantly predict real user depth transitions (e.g., weighted kappa < 0.4 or no improvement over a depth-prior baseline), then axis-2 and the 2% deep-trust result lack external validity, and the ranking's claim to measure relational competence would require human validation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that CompanionBench separates substantive relational support from surface warmth depends on the disclosure gate and trained simulator (§4.1–§4.2) being a valid model of how real users disclose, retreat, and rebuild trust. The gate's advance/hold/retreat transitions and 'earned depth' thresholds are hand-authored from reciprocity, Social Penetration, and rupture-repair theory, and the anti_goal trips are triggered by LLM-labeled ai_move categories (§5.2). The synthetic counterfactual training branch exists precisely because real good-AI turns are scarce (§4.2), so those transitions encode the authors' theoretical priors rather than observed user reactions. The paper explicitly reports that expert inter-rater reliability and external-validity studies are undone (§9). Without evidence that the gate's transition probabilities match real user disclosure behavior, the headline results could be properties of the gate design rather than of the SUTs. In particular, 'no SUT earns deep trust in 20 turns' (§6.3a) is consistent with an arbitrarily strict hand-set gate, and the dominant warmth-over-substance failure counts come from a single judge's reverse-event labels (§6.3c). Internal reliability, permutation checks, and IRT de-biasing are strong, but they establish that the instrument is stable, not that it measures the intended construct.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents CompanionBench, an interactive bilingual benchmark for evaluating AI emotional companionship. It constructs 500 Chinese–English parallel personas grounded in de-identified real conversations, maps ten capabilities from 25 theories, trains a user simulator with a hidden disclosure gate that advances, holds, or retreats based on the system-under-test's behavior, and scores trajectories on two axes: a subjective ten-capability rubric (axis-1) and a deterministic gate replay measuring earned disclosure depth (axis-2). Twenty-eight agents are evaluated in 20-turn role-flipped rollouts, and rankings are de-biased with a many-facet Rasch model across a three-family judge panel. The reported findings are that rankings are highly reproducible (Spearman rho = 0.996 Chinese / 0.953 English), the two axes diverge on max_depth, no agent earns deep trust in 20 turns, and the dominant failure mode is substituting surface warmth for substantive relational support, with role-play-optimized models ranking near the bottom.","tokens_in":38441,"tokens_out":4588,"duration_ms":76035,"significance":"If the instrument measures what it claims, CompanionBench is a significant methodological contribution: it operationalizes relational capabilities that prior benchmarks collapse into a single warmth score, grounds scenarios and a user simulator in real-world data, and provides a reproducible bilingual evaluation pipeline. The internal reliability work is genuinely strong: test-retest stability (rho 0.996/0.953), permutation-based environmental responsiveness (p < 0.001), bootstrap tiering, sample-size saturation analysis, and a worked IRT example showing how naive cross-family means can invert a ranking. The paper is also unusually explicit about its limitations. However, construct validity remains the central open question. The disclosure gate and the reverse-pattern counts are hand-authored or single-judge, and the paper reports no human or counselor validation, so the headline findings are not yet established as properties of the SUTs rather than properties of the instrument.","major_comments":[{"comment":"The central claim that CompanionBench separates substantive relational support from surface warmth rests on the disclosure gate being a valid model of how real users disclose, retreat, and rebuild trust. The gate's advance/hold/retreat transitions and earned-depth thresholds are hand-authored from reciprocity, Social Penetration, and rupture-repair theory, and the training branch deliberately uses synthetic counterfactual turns because real good-AI turns are scarce (only 2.28% of slices). The paper itself states that expert inter-rater reliability and external-validity studies are undone (§9). Without an external criterion, the finding that 'no SUT earns deep trust in 20 turns' (§6.3a) is consistent with an arbitrarily strict hand-set gate, and the warmth-over-substance counts could be a property of the gate design. I ask for a validation study in which human or counselor annotators judge a subset of rollouts for expected disclosure depth and trust responses, compared against the gate transitions; if thresholds are adjusted, sensitivity analyses should show that the rankings and the deep-trust ceiling are robust across a plausible threshold range.","section":"§4.1–§4.2, §9"},{"comment":"Axis-2 labels (ai_move categories and disclosure depth) and all reverse-event counts come from a single judge (deepseek-v4-pro), and inter-judge agreement is reported only for axis-1 rankings (0.845–0.951). The dominant failure-mode claim — that empty_encouragement and forced_positivity together account for 83.6% of tier-2 events (Table L.2) — is therefore based on labels whose reliability is unmeasured. Please report inter-judge agreement on axis-2 and on reverse-item labels, or run a multi-judge axis-2 pass on a subset, and show that the tier-2 dominance and its rank correlation survive.","section":"§5.2, §6.3c, Appendix L"},{"comment":"The statement in §6.3a that max_depth clustering at mid is 'both a deliberate gate constraint and the single-session capability ceiling' acknowledges that the deep-trust result is partly by construction. To make this finding informative, the paper should provide a threshold-sensitivity analysis: vary the hand-set gate thresholds (for example, the number of earned-depth signals required per layer, or the asymmetry of retreat-and-repair) and show that no-SUT-earns-core persists across a plausible range, or state explicitly which part of the ceiling is an instrument property.","section":"§6.3a, §4.1"},{"comment":"The IRT model is additive in judge severity, and the limitations section notes that a residual judge×family×language interaction survives, with opus being stricter on Chinese open-weight models. An additive β can only dilute, not remove, such interactions. Please quantify the effect of the residual interaction on the final ranking — for example, by comparing θ estimates from judge subsets or from a model with an interaction term — and show that the reported tier structure is not driven by that residual interaction.","section":"§5.3, §9"}],"minor_comments":[{"comment":"The phrase 'Per-cellnships with the code' is a typographical fragment; it should read 'Per-cell n ships with the code' or be rewritten as a complete sentence.","section":"Appendix H"},{"comment":"The 'overall ρ = 0.951' should state explicitly that it is Spearman correlation on IRT θ and should clarify whether it is computed over the 28-SUT ordering or a different set.","section":"§6.1"},{"comment":"Abbreviations such as 'dep' and short model names (e.g., 'db-character-251128') are defined only in captions or appendix; please define all abbreviations at first use in the main text.","section":"Table 3, Figure 4"},{"comment":"The synthetic counterfactual branch is described in one sentence; a brief description of the three-agent layer and the coverage grid, or a pointer to an appendix with implementation details, would aid reproducibility.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well within scope for this venue, and the internal reliability work is unusually careful. My main concern is the missing external validation of the disclosure gate and the single-judge axis-2 labels; without it, the headline findings remain ambiguous between instrument properties and SUT properties. I would support publication after the authors add a human/counselor validation study or substantially recalibrate the claims to match the currently internal-only evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is the most careful benchmark paper I've seen for AI emotional companionship. The genuinely new thing is the hidden disclosure gate: a deterministic state machine that forks the user-simulator trajectory on the agent's own behavior, so the environment reacts to the system under test instead of following a fixed script. That, plus grading four capabilities prior work ignores (holding ambiguity, selfobject responsiveness, positive resonance, calibrated challenge), and an IRT treatment of judge severity that exposes the scale-incommensurability artifact under diagonal exclusion, makes it a real contribution.\n\nThe reliability evidence is unusually solid. Test-retest Spearman 0.996/0.953 across languages, inter-judge agreement 0.845-0.951, permutation tests for environmental responsiveness, and a worked example showing naive cross-family means invert the opus/gpt-5.5 ranking. The paper also has a proper section 9 that lists undone validity studies instead of hiding them.\n\nWhere it is soft: the disclosure gate and the simulator are hand-authored from theory and trained with synthetic counterfactuals, and there is no evidence yet that gate transitions match how real users disclose, retreat, or repair. The 2% deep-trust and warmth-over-substance findings could be partly a property of the instrument. The paper says this explicitly, but that does not make the finding any less instrument-relative. I would want human or counselor baselines before betting on the substantive claims. Also the code and data are promised, not released, and the persona population remains skewed toward young women and preoccupied attachment despite sampling. Sparse capability samples (C8 as low as n=32) mean some per-capability conclusions rest on thin ice.\n\nNet: this deserves a serious referee. The central object is novel, the internal evidence is strong, and the limitations are front-loaded. I would condition acceptance on shipping the code and adding at least one external-validity check. It is a solid piece of work.","headline":"Careful, novel benchmark with strong internal reliability; treat the headline findings as instrument-relative until external validation lands.","tokens_in":38953,"tokens_out":2502,"would_cite":true,"duration_ms":26174,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A bilingual benchmark with a hidden disclosure gate shows that no AI companion earns deep trust in a 20-turn session, and that the dominant failure across 28 agents is substituting surface warmth for substantive relational support.","keywords":["emotional companionship","LLM evaluation","benchmark","user simulator","disclosure gate","Item Response Theory","bilingual","relational competence"],"falsifier":"Record licensed counselors' or real users' moment-by-moment willingness to disclose during the same 20-turn dialogues and compare it to the gate's state; if the correlation between gate-earned depth and human-rated earned depth is weak, the deterministic axis would not measure trust.","tokens_in":37967,"feed_emoji":"💬","tokens_out":6853,"duration_ms":59673,"temperature":0.7,"pith_summary":"This paper builds a bilingual benchmark that separates two things often fused in AI companionship scores: the warmth of what an agent says, and whether it actually earns a user's deeper trust. The benchmark's central mechanism is a hidden disclosure gate: a simulated user opens up only when the agent's behavior earns it, and retreats when the agent judges, lectures, or rushes. Running 28 models through 20-turn sessions in Chinese and English, the paper finds that no model earns deep trust within a session, and that the dominant failure is replacing substantive relational support with surface warmth. Because the same trajectories are scored both by a ten-capability rubric and by the deterministic gate, the benchmark can show where tone and substance diverge. Rankings are reproducible across languages, and the authors argue existing single-score empathy benchmarks would miss these distinctions.","feed_headline":"AI companions substitute warmth for real support, 28-model test finds","feed_subtitle":"A hidden disclosure gate shows no model earns deep trust in 20 turns; rankings replicate across languages.","key_machinery":"The key mechanism is the hidden disclosure gate: a deterministic state machine over three disclosure depths (surface < mid < core) with five ordered gates, each opened only by behavior that earns it and re-locked by behavior that violates the persona's anti-goal (judging, lecturing, rushing), with retreat quick and re-building slow. It turns the trained user simulator into a transition function that forks each persona's trajectory on the agent's own behavior, so the environment itself is the measurement: how far disclosure gets is the second evaluation axis. This gate is what makes warmth and substance empirically separable from aggregate scores.","core_discovery":"On the paper's own terms, the discovery is that relational competence in AI companions is measurable as two separable axes: a subjective ten-capability rubric (ten capabilities anchored in 25 psychology and counseling theories, four of them newly graded) and a deterministic disclosure-gate replay that records how far and how often the user's deeper disclosure is earned rather than taken. Across 28 agents in both languages, emotion regulation and calibrated challenge are the weakest capabilities even at the top, holding ambiguity discriminates agents most, and no agent reaches the deepest (core) disclosure layer in a 20-turn first encounter. Role-play-optimized agents rank near the bottom, which the paper reads as evidence that immersion is not relational competence. The dominant failure across agents is not overt harm but substituting warmth for substance: high verbal-warmth scores with large drops once reverse anti-patterns (empty encouragement, forced positivity, overpromising) are deducted.","pith_inferences":["If the disclosure gate were validated against human-counselor or real-user ratings of the same dialogues, the deterministic axis could become a training signal; until then, the gate's mapping from behavior to disclosure is an assumption, not an observed law.","The 20-turn ceiling on deep trust might be partly a design artifact of a deliberately conservative gate; longer or repeated sessions could show whether deeper trust is reachable by any model, or whether single-session benchmarks structurally cap it.","The cross-family IRT correction could transfer to other LLM-as-judge settings where per-agent exclusion of same-family judges creates scale drift; the paper's β-span diagnostic gives a cheap way to decide when de-biasing is needed.","The newly graded capabilities (holding ambiguity, selfobject responsiveness, positive resonance, calibrated challenge) suggest concrete training targets; using them as rewards might reduce the warmth-for-substance substitution the benchmark exposes."],"forward_implications":["Aggregate warmth/empathy scores hide capability-level differences; reporting primary, deduction, and final scores separately is necessary to see whether an agent's warmth is substantive.","Role-play immersion does not imply relational competence; role-play-optimized models ranked near the bottom, so training for immersion may be orthogonal to earning trust.","No agent earns core disclosure in 20 turns, so cross-session, longitudinal evaluation is the natural next test of earned deep trust.","The near-redundancy of the gate advance rate with the rubric ranking (ρ≈0.94) but the divergence of max depth (ρ≈0.25) shows that earned depth is the non-redundant signal; benchmarks should measure depth, not just advance rate.","Because rankings saturate at about 200 personas, future evaluations can use smaller sets when only ranking fidelity is needed, reserving larger sets for sparse capabilities and gate peaks."],"supporting_citations":[{"why":"Supplies the disclosure-depth ladder that the gate's earned-depth progression is built on.","marker":"Altman and Taylor 1973"},{"why":"Grounds the asymmetric retreat-and-repair dynamics the gate uses after an agent violates the persona's anti-goal.","marker":"Safran and Muran 2000"},{"why":"Provides the reciprocity principle by which disclosure advances only when the agent's behavior earns it.","marker":"Jourard 1971"},{"why":"Supplies the many-facet Rasch model used to estimate judge severity and produce de-biased rankings.","marker":"Linacre 1989"},{"why":"Documents LLM judges favoring their own family, motivating the cross-family panel and the self-preference diagnostic.","marker":"Panickssery, Bowman, and Feng 2024"},{"why":"Establishes the agenda-based user-simulation paradigm that the trained user simulator extends into a disclosure-gate transition function.","marker":"Schatzmann et al. 2007"},{"why":"Provides ESConv, an earlier multi-turn emotional-support benchmark that aggregates empathy into a warm-score, the design the paper contrasts with.","marker":"Liu et al. 2021"},{"why":"Defines social sycophancy, the failure mode the benchmark's reverse layer and calibrated-challenge capability target.","marker":"Cheng et al. 2026"},{"why":"Prior counselling benchmark that scores calibrated challenge only within counselling competence, which the paper extends into a standalone capability.","marker":"Wang et al. 2026a"}],"fun_headline_variants":["AI companions substitute warmth for support, 28-agent test finds","No AI companion earns deep trust in 20-turn test","Role-play AI ranks lowest in relational competence test","Warmth without substance: AI companions' top failure","Emotion regulation and calibrated challenge trip up AI companions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the disclosure gate's behavior-to-disclosure mapping—when a real user would open up or withdraw—matches how actual people respond; the paper does not yet have human or counselor validation for this mapping.","fun_headline_variants_meta":{"raw":{"variants":["AI companions substitute warmth for support, 28-agent test finds","No AI companion earns deep trust in 20-turn test","Role-play AI ranks lowest in relational competence test","Warmth without substance: AI companions' top failure","Emotion regulation and calibrated challenge trip up AI companions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001414,"raw_usage":{"total_tokens":5751,"prompt_tokens":1028,"completion_tokens":4723,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":4644}},"tokens_in":644,"tokens_out":4723,"duration_ms":33471,"temperature":1.0,"reasoning_tokens":4644,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:09:49.754206+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record licensed counselors' or real users' moment-by-moment willingness to disclose during the same 20-turn dialogues and compare it to the gate's state; if the correlation between gate-earned depth and human-rated earned depth is weak, the deterministic axis would not measure trust.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the disclosure-depth ladder that the gate's earned-depth progression is built on."},{"cited_title":"D.; and Muran, J","cited_arxiv_id":null,"evidence_quote":"Grounds the asymmetric retreat-and-repair dynamics the gate uses after an agent violates the persona's anti-goal."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the reciprocity principle by which disclosure advances only when the agent's behavior earns it."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the many-facet Rasch model used to estimate judge severity and produce de-biased rankings."},{"cited_title":"R.; and Feng, S","cited_arxiv_id":null,"evidence_quote":"Documents LLM judges favoring their own family, motivating the cross-family panel and the self-preference diagnostic."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the agenda-based user-simulation paradigm that the trained user simulator extends into a disclosure-gate transition function."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides ESConv, an earlier multi-turn emotional-support benchmark that aggregates empathy into a warm-score, the design the paper contrasts with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines social sycophancy, the failure mode the benchmark's reverse layer and calibrated-challenge capability target."}],"review_version":2}