{"id":"20421b38-ee26-41ec-80ee-2a7b7af39cc4","arxiv_id":"2506.01357","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A large Japanese counseling dialogue dataset collected via role-play by trained counselors, with per-dialogue client feedback, improves LLM counseling response generation and evaluation.","lead":"KokoroChat, a Japanese psychological counseling dialogue dataset of 6,589 role-played sessions between trained counselors, is introduced. Each dialogue includes 20-item client feedback, and fine-tuning open-source LLMs on it improves both response generation and automated dialogue evaluation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Client feedback scores from role-playing counselors are unvalidated as ground truth for counseling quality; the automatic-evaluation claim in §5.2 rests on this assumption.","rationale":"The reader's weakest-assumption analysis already identified the core issue: role-play authenticity and client-feedback validity are asserted but not externally validated. My stress-test converges on the same point, sharpened to the specific claim that the score labels are the sole ground truth for the automatic-evaluation experiments. This is the single most load-bearing concern because the paper's second and third contributions (improved generation via quality filtering, improved automatic evaluation) both rely on the client-feedback scores being meaningful measures of counseling quality. If those scores are biased—for example, because role-playing counselors rate each other leniently, or because their professional perspective diverges from real clients—then the score-prediction results in Table 5 and the Kokoro-Low/Kokoro-High comparison in §5.1 become difficult to interpret. Other limitations noted by the reader (test set restricted to scores 99–100, post-hoc exclusion of all-3 dialogues, small human evaluation, no released code) are real but secondary: they affect the strength or generality of the empirical demonstration, not the foundational validity of the dataset's labels. The proposed concrete test directly measures rater agreement between the original role-play clients and independent external raters using the same instrument, which would settle whether the labels can serve as ground truth. Since the reader already returned a CONDITIONAL verdict, my read does not move the verdict; the paper should be accepted only if the validity of the client-feedback scores is demonstrated or the claims are appropriately scaled back.","tokens_in":35705,"tokens_out":4489,"duration_ms":52855,"concrete_test":"Randomly sample about 200 KokoroChat dialogues. Have a separate panel of real clients or standardized patients (without counselor training), plus licensed counseling supervisors, rate the counselor-side utterances using the same 20 items from Table 2. Compute inter-rater agreement (ICC or weighted kappa) between the original client-role scores and the external ratings, overall and per item. If ICC is below 0.6 or weighted kappa below 0.4, the score labels are not sufficiently reliable to support the automatic-evaluation claims in Table 5 and the quality-based data partitions in §5.1.1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The automatic-evaluation contribution depends entirely on the 20-item client feedback scores (Table 2, §3.3) being reliable and valid measures of counseling quality. The raters are trained counselors playing the client role, and their scores are immediately shared with the counselor-role player and monitored by the platform administrator, which may introduce social-desirability or peer-evaluation bias. No external validation is provided: the paper asserts authenticity in §1 and §3.1, but never compares the proxy-client scores against real-client perceptions, independent expert ratings, or an established counseling-quality instrument. Consequently, the improved ACC/ACCsoft/MAE in Table 5 demonstrates only that a fine-tuned model predicts the role-play clients' scores better than GPT-4o does; it does not demonstrate better evaluation of counseling quality. The same scores also drive the high/low quality partitions used in the response-generation experiments (§5.1.1), so if the scores are biased or noisy, the conclusion that Kokoro-High outperforms Kokoro-Low because of data quality is likewise unsupported. The generation claim itself has partial independent support from the human pairwise evaluation (§5.1.3), but the evaluation and data-quality claims hinge on the unvalidated score labels.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces KokoroChat, a Japanese psychological counseling dialogue dataset of 6,589 role-played dialogues between trained professional and trainee counselors, each accompanied by a 20-item client feedback score (0-100). The authors claim KokoroChat is the largest human-collected counseling dialogue dataset to date. They fine-tune Llama-3.1-Swallow on the dataset and report that it improves both response generation quality (automatic metrics and human pairwise evaluation) and automatic dialogue evaluation (score prediction accuracy, soft accuracy, and MAE compared with Llama-3.1 and GPT-4o). The dataset is released publicly.","tokens_in":35918,"tokens_out":5612,"duration_ms":61336,"significance":"If the underlying data and scores are valid, KokoroChat fills a clear gap as a large Japanese-language counseling resource with long, realistic-length dialogues and fine-grained evaluation labels. The paper's strengths include the scale of human collection (6,589 dialogues, 480 participants), the expert-designed 20-item feedback instrument, the use of trained counselors as role players, a public release, and careful documentation of fine-tuning and inference settings (QLoRA, seeds, deterministic decoding). The human evaluation by five professional counselors is a useful independent check. However, the significance of the automatic-evaluation contribution rests entirely on the validity of the role-play client feedback scores, which are not externally validated, and the response-generation experiments are evaluated on a narrow high-score subset of the data.","major_comments":[{"comment":"The automatic-evaluation claim in Section 5.2 is load-bearing and depends on the 20-item client feedback scores being valid measures of counseling quality, but the paper provides no external validation of these scores. The raters are trained counselors acting as simulated clients, the scores are immediately shared with the counselor-role player, and the platform administrator monitors the process; these conditions may introduce social-desirability or peer-evaluation bias. The paper asserts authenticity in Sections 1 and 3.1 but never compares the proxy-client scores against real-client perceptions, independent expert ratings, or an established counseling-quality instrument. The paper's own Ethical Considerations state that the dialogues are not real counseling sessions. Consequently, Table 5 demonstrates only that a fine-tuned model can predict the role-play clients' scores better than GPT-4o; it does not establish better evaluation of counseling quality. The authors should either add an external validation study or substantially soften the abstract's claim about 'automatic evaluation of counseling dialogues.'","section":"3.3 and 5.2"},{"comment":"The response-generation test set consists exclusively of 118 dialogues with client feedback scores of 99 or 100, while all training variants use dialogues with scores at or below 98. This restricts evaluation to the extreme upper tail of the score distribution and does not reflect performance on typical dialogues (the mean score is 63.58, median 64.00). Moreover, the Kokoro-High vs Kokoro-Low comparison is confounded by the mismatch between the score distribution of each training set and the test set: neither training set contains dialogues from the 99-100 score range. To support the general claim that fine-tuning on KokoroChat improves response quality, the evaluation should include a test set sampled from the full score range, or the conclusions should be limited to the high-score regime.","section":"5.1.1"},{"comment":"The human pairwise evaluation reports only raw win/lose/tie percentages without significance tests, bootstrap confidence intervals, or inter-annotator agreement. With 100 responses per model and only 10 dialogues, margins of a few percentage points (e.g., in the first comparison in Figure 5, the win and lose rates differ by only a few percentage points) are within sampling noise, yet the text concludes that 'even when using only the lower-scoring portions of KokoroChat, it still enhances open-source LLMs.' The authors should report statistical significance or confidence intervals for each pairwise comparison, or refrain from strong conclusions where the margin is not significant.","section":"5.1.3 and Figure 5"}],"minor_comments":[{"comment":"The dataset name 'C ACTUS' has an extra space and should be written as 'CACTUS' (also in Table 1).","section":"2.2"},{"comment":"The justification for the score threshold of 70 ('to ensure balanced data segmentation') is not substantiated; the authors should report the number of dialogues and utterances in each partition to show balance.","section":"5.1.1 (footnote 3)"},{"comment":"Standard deviations are reported only for the fine-tuned model (averaged over five seeds), while the baseline Llama-3.1 and GPT-4o results appear to be single runs, making the stability comparison asymmetric; reporting multiple runs for the baselines or clearly stating the limitation would be fairer.","section":"Table 5 and Appendix D"},{"comment":"The dialogue-topic distribution is produced by GPT-4o-mini without any human validation or agreement measure; the diversity claims based on Figure 2 should be treated as preliminary unless a sample of the topic labels is verified.","section":"4.1"}],"recommendation":"major_revision","confidential_remarks":"The dataset itself is a valuable community resource and the authors are transparent about its simulated nature. The main risk is overclaiming in the abstract: the automatic-evaluation contribution currently lacks external validation of the role-play scores, and the response-generation evaluation is restricted to a high-score subset. If the authors add validation or substantially temper these claims, the paper could be acceptable for publication. No concerns about novelty or citation patterns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"KokoroChat is a substantial resource: 6,589 Japanese counseling role-play dialogues, averaging 91 utterances, each with 20-item client feedback. That fills a real gap—there is no comparable Japanese counseling dataset at this scale—and the dataset is the contribution, not the experiments. The role-play approach with trained counselors is a reasonable ethical compromise, and the paper is honest that these are simulated sessions, not real therapy. The fine-tuning experiments are a credible demonstration. The human evaluation gives independent support that Kokoro-High beats Kokoro-Low and Kokoro-Full, so the data-quality signal is not just a BLEU artifact. The score-prediction result shows the fine-tuned model learns the role-play clients' scoring patterns better than GPT-4o zero-shot, which is a legitimate supervised learning result. The soft spots are real but not fatal. The biggest one is that the 20-item scores are unvalidated. The raters are trained counselors acting as clients; their scores are shared with the counselor-role player and monitored by the platform, and there is no comparison against real-client perceptions or an established counseling-quality instrument. So Table 5 shows prediction of role-play scores, not counseling quality. The paper overclaims when it calls this 'automatic evaluation of counseling dialogues.' The generation partition also leans on these scores: the test set is only the 99–100 dialogues, and the high/low split uses the same labels. That said, the value of KokoroChat as a training resource does not depend on the scores being externally valid; researchers can use the dialogues for generation and treat the feedback as a role-play signal. A small validation study—independent expert ratings on a sample—would substantially strengthen the evaluation claim. Minor issues: no code for the fine-tuning experiments, only the dataset; the all-3 exclusion is post-hoc and could be described as such; the 70% threshold is arbitrary but clearly stated. The limitations section is candid about the lack of direct comparisons and about demographic skew. Who benefits: anyone working on Japanese counseling language models, cross-cultural counseling NLP, or dialogue evaluation datasets. The paper deserves a serious referee; it will need revision, mainly to narrow the evaluation claims and add validation, but the dataset is a solid contribution. Send it to review.","headline":"KokoroChat is a genuinely useful new dataset and the paper deserves a serious referee; the main caveat is that the client-feedback scores are role-play labels whose validity as counseling-quality measures is never tested.","tokens_in":740,"tokens_out":1824,"would_cite":true,"duration_ms":46957,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A role-play dataset of 6,589 Japanese counseling dialogues, built by trained counselors, improves both response generation and automated evaluation.","keywords":["psychological counseling dialogue dataset","role-playing","Japanese NLP","client feedback","dialogue evaluation","fine-tuning LLM","emotional support conversation"],"falsifier":"Run an inter-rater reliability study where several trained counselors each play the client role for the same recorded counselor session and independently complete the 20-item feedback form; if the scores disagree substantially across raters, the feedback signal that powers both fine-tuning and score prediction is not a stable measure of counseling quality.","tokens_in":35501,"feed_emoji":"🧑‍⚕️","tokens_out":4204,"duration_ms":42696,"temperature":0.7,"pith_summary":"The paper claims that a dataset of 6,589 Japanese psychological counseling dialogues, collected by trained counselors playing both the counselor and client roles, is the largest human-collected counseling dialogue dataset to date. It argues that this role-playing approach yields longer, more authentic dialogues than LLM-generated alternatives while avoiding the privacy risks of using real counseling records. The paper then reports that fine-tuning an open-source large language model on this data improves the quality of generated counseling responses, and that a model trained on the dataset's client feedback scores predicts human ratings more accurately than GPT-4o. A sympathetic reader would care because better dialogue data and automated evaluation could make counseling support more accessible to people who cannot see a professional.","feed_headline":"6,589 role-played therapy chats help train counseling AI","feed_subtitle":"Japanese dataset with per-session client scores lifts response quality and automated evaluation.","key_machinery":"The role-playing collection protocol combined with the 20-item client feedback instrument. Trained professional and trainee counselors, 480 in total and all with 10 hours of online text-counseling training, alternate as counselor and client in hour-long text sessions; after each session the client-role player scores the counselor on 10 overall-impression items and 10 counseling-skill items, with three screening flags that can zero or halve the total. This design supplies both the dataset's authenticity claim and the supervision signal used for fine-tuning response generation on high-versus-low score partitions and for training the score-prediction evaluator.","core_discovery":"The paper's central claim is that a large-scale role-play protocol staffed by trained counselors can produce counseling dialogues that carry detailed per-session client feedback, and that this data improves both generation and evaluation of counseling responses. Each KokoroChat dialogue is a roughly one-hour text exchange, with client-role players rating counselor-role players on 20 items covering overall impression and professional skills on a 0-to-5 scale. The paper reports that fine-tuning Llama-3.1-Swallow-8B on the high-scoring subset (Kokoro-High) outperforms fine-tunes on low-scoring and full datasets in human pairwise evaluation, and that a score-prediction model trained on the feedback exceeds GPT-4o in accuracy (35.35 vs 30.92) and mean absolute error (0.828 vs 1.015). The claim is that the dataset's size, dialogue length (averaging 91.2 utterances), topic coverage, and itemized scores make it a stronger resource for Japanese counseling dialogue systems than existing LLM-augmented or smaller human-collected datasets.","pith_inferences":["If the score-prediction model proves reliable, it could serve as an automated coaching tool that gives novice counselors itemized feedback without requiring a human supervisor.","The authenticity claim rests on an unvalidated premise; comparing role-play transcripts with de-identified real counseling sessions would be a direct test of whether the dialogues faithfully represent real client-counselor interactions.","The strong correlation among feedback items such as gaining new insights, feeling hopeful, and perceiving value suggests that a model optimized only for empathy may miss the insight-oriented components that clients associate with a valuable session.","Translating or adapting KokoroChat into other languages would allow direct comparison with English and Chinese counseling datasets and would test whether the score distributions and item correlations hold across cultures."],"forward_implications":["Fine-tuning on KokoroChat, especially the high-scoring subset, improves the quality of counseling responses generated by an 8-billion-parameter open-source Japanese LLM.","A dialogue evaluation model trained on KokoroChat's client feedback predicts the 20 per-dimension scores with higher accuracy and lower error than GPT-4o in zero-shot mode.","Client word count shows the strongest positive correlation with feedback scores (rho = 0.42), suggesting that encouraging client expression matters more for positive evaluation than counselor verbosity.","Counselor response time correlates negatively with feedback scores, implying that faster replies may contribute to a better counseling experience.","The dataset's size and per-item scores enable future work such as dialogue-act annotation, cross-lingual comparison, and fine-grained analysis of questioning strategies."],"supporting_citations":[{"why":"Supplies ESConv, the main human-collected emotional support dialogue baseline and defines the emotional support conversation task that KokoroChat compares against.","marker":"Liu et al. (2021)"},{"why":"Provides Client-Reactions, the closest prior human-collected dataset with client ratings, used as a comparison baseline for scale and feedback coverage.","marker":"Li et al. (2023)"},{"why":"Offers Anno-MI, an expert-annotated counseling dialogue dataset, as another human-collected comparison point for dialogue length.","marker":"Wu et al. (2022)"},{"why":"Motivates the role-play approach by showing GPT-4-generated counseling responses can approach professional counselors in role-play settings.","marker":"Inaba et al. (2024)"},{"why":"Provides GPT-4o, the baseline model used in response generation and score prediction comparisons.","marker":"OpenAI (2024)"},{"why":"Provides Llama-3.1-Swallow, the continuously pre-trained Japanese LLM that is the base model for fine-tuning experiments.","marker":"Fujii et al. (2024)"}],"fun_headline_variants":["6.5K role-played therapy chats hone counseling AI","Japanese role-play dataset trains counseling AI on 6,589 chats","Trained counselors role-play to build better therapy AI in Japanese","Role-played Japanese therapy chats improve AI counseling and scoring","Counselor role-play yields 6,589-session dataset for better AI therapy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's value rests on the premise that role-playing by trained counselors produces dialogues and client feedback scores that faithfully represent real counseling interactions, a premise the paper asserts but does not validate against real counseling data or an external benchmark.","fun_headline_variants_meta":{"raw":{"variants":["6.5K role-played therapy chats hone counseling AI","Japanese role-play dataset trains counseling AI on 6,589 chats","Trained counselors role-play to build better therapy AI in Japanese","Role-played Japanese therapy chats improve AI counseling and scoring","Counselor role-play yields 6,589-session dataset for better AI therapy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000651,"raw_usage":{"total_tokens":2975,"prompt_tokens":924,"completion_tokens":2051,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":1960}},"tokens_in":540,"tokens_out":2051,"duration_ms":16148,"temperature":1.0,"reasoning_tokens":1960,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:43:07.236719+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an inter-rater reliability study where several trained counselors each play the client role for the same recorded counselor session and independently complete the 20-item feedback form; if the scores disagree substantially across raters, the feedback signal that powers both fine-tuning and score prediction is not a stable measure of counseling quality.","supporting_citations":[],"review_version":1}