{"id":"ac259531-7bd2-44f8-9474-b8f069e71949","arxiv_id":"2412.14971","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Most Chinese and Western LLMs pass the Chinese social work exam's knowledge sections, with Chinese models ahead on jurisprudence, and both groups showing cultural bias in practice scenarios.","lead":"Researchers tested eight Chinese and Western AI chatbots on the Chinese national social work exam and found that most passed the knowledge sections. Chinese models were stronger on laws and policies, but both sides struggled with culturally specific practice scenarios and showed cultural biases.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The core assumption that the 2023 CNSWE items were unseen by the eight models is not secured; the paper's own limitation concedes this, and the published 2024 self-study guides make contamination plausible.","rationale":"The reader's verdict is already CONDITIONAL with moderate confidence, and the training-contamination assumption is the right load-bearing point. I agree with the reader's identification: the paper's own limitation section concedes the uncertainty, and the use of 2024-published self-study guides containing 2023 exam items makes the 'unseen items' assumption particularly insecure. The proposed paraphrasing test is feasible with the existing API infrastructure and the released repository: if paraphrased scores match the original scores, the foundational-knowledge interpretation is substantially strengthened; if they drop, the paper must be reframed as measuring memorization or retrieval rather than understanding. I do not see a separate concern that would move the verdict. The cultural-bias claims are more illustrative than systematic, but they are not the central quantitative claim, and the expert-review protocol with PABAK values provides some support for the qualitative findings. Thus the appropriate recommendation is to keep the reader's CONDITIONAL verdict unchanged, with the concrete contamination check as the condition.","tokens_in":19475,"tokens_out":4917,"duration_ms":34748,"concrete_test":"Re-administer all 160 items to the same eight models in a paraphrased form: rewrite each question stem and each option in Chinese, preserving meaning, official answer, and question structure but changing surface wording, and use identical prompts and temperature settings. Then compare per-question accuracy and answer agreement with the original administration. If paraphrased accuracy drops by more than about five points, or if answer flips cluster on items with distinctive legal or factual phrasing, the original scores are inflated by memorization; if accuracy is stable, the unseen-items assumption is supported. A supplementary check is to release the exact item-answer mapping and per-question response logs so the comparison can be independently audited.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the 2023 CNSWE questions used as the test set were not in the models' training data. The Methods section ('Chinese National Social Work Examination') argues that the 2023 version was chosen because it 'coincides with our selected models' approximate training cutoff date,' but the same section reports that the items were drawn from 2024-published self-study guides (National Social Worker Professional Exam Question Compilation Group, 2024a, 2024b). Those guides are exactly the kind of Chinese-language exam-preparation material that is widely scraped into pretraining corpora, and unofficial copies of the 2023 exam circulated online immediately after administration. The paper's Strengths and Limitations section concedes: 'we cannot definitively determine if performance reflects true knowledge or pattern matching from training data.' This is not a peripheral caveat: the headline finding that seven of eight models pass both sections, and especially the Chinese-model advantage on jurisprudence, would be reinterpreted as retrieval or memorization if the items are present in training. The cultural-bias examples may remain informative because they come from explanations, but the central 'foundational knowledge' claim and the comparison between Chinese and Western models rest on the unseen-items assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper evaluates eight cloud-based large language models (four Chinese, four Western) on the jurisprudence and applied knowledge sections of the 2023 Chinese National Social Work Examination (160 questions). The authors administer three testing conditions (required response, option to skip, answer-options-only), score the responses using official rules, and subject explanation texts to expert review by bilingual social work professionals. They report that seven of eight models met the official 60-point pass threshold on both sections; Chinese models scored higher on jurisprudence (median 77.0 vs. 70.3) and lower on applied knowledge (65.5 vs. 67.0); and both groups displayed cultural biases, particularly around gender and family issues, despite strong command of professional terminology. The study is positioned as the first systematic cross-cultural benchmark of LLMs in a non-Western professional context.","tokens_in":19681,"tokens_out":8411,"duration_ms":60950,"significance":"If the findings hold, this study provides a novel and useful benchmark for evaluating LLMs in a non-Western professional domain. Its principal strengths are the use of an official national licensure examination with transparent, official scoring rules; a bilingual expert review of model explanations with reported inter-rater reliability; and a publicly available code and data repository. The paper also goes beyond simple pass/fail metrics by analyzing reasoning validity and by including a condition designed to detect construct-irrelevant variance. However, the central findings depend on an assumption about training-data contamination and on single-run measurements, and the cultural-bias claim lacks systematic coding. These issues materially affect the strength of the conclusions, as elaborated in the major comments.","major_comments":[{"comment":"The headline claim that seven models possess foundational knowledge sufficient to pass the CNSWE depends on the items being unseen during pretraining. The authors justify the 2023 version by its proximity to model training cutoffs, yet the items were drawn from 2024-published guide materials, and related exam content has circulated online; the paper's own limitation section concedes 'we cannot definitively determine if performance reflects true knowledge or pattern matching from training data.' Memorization would inflate pass rates and could explain the Chinese-model advantage on jurisprudence (e.g., if Chinese-language exam guides are overrepresented in Chinese-model training corpora). To support the 'foundational knowledge' and cross-region comparison claims, please add a contamination check (e.g., membership-inference probes, testing on a post-cutoff exam, or a control set of newly written items), or revise the abstract and conclusions to present the results as conditional on the unseen-items assumption.","section":"Methods: Chinese National Social Work Examination (p. 7-8); Strengths and Limitations (p. 27-28)"},{"comment":"All scores are based on a single API call per condition with temperature set to 0, and the analysis is restricted to descriptive statistics with no uncertainty quantification. Given only n=4 models per region, the reported median differences (jurisprudence 77.0 vs. 70.3; applied knowledge 65.5 vs. 67.0) could be within run-to-run or item-level variance. The authors should report item-level bootstrap confidence intervals around the medians, or a permutation-based comparison, and discuss the limitations of the n=8 sample for drawing any regional inferences. Without this, statements such as 'Chinese models demonstrate advantages in regulatory content' (Discussion) are not statistically supported.","section":"Methods: Data Management (p. 11-12); Analytic Plan (p. 12)"},{"comment":"The claim that both Chinese and Western models exhibit cultural biases, especially around gender equality and family dynamics, is supported only by a few qualitative examples (e.g., the inheritance-rights scenarios in which models favored sons, and the 'loving putting self in the spotlight' answer). The expert-review protocol in the Methods section (p. 12-13) does not include a pre-specified coding scheme for bias themes, so no prevalence counts, per-model breakdowns, or reliability figures are available for this assertion. Please either present the bias finding as an exploratory observation from the qualitative review, or add a systematic content analysis of all 160 items across the eight models with inter-rater reliability for the bias categories.","section":"Discussion: Cultural Competency and Language Processing (p. 24-25)"},{"comment":"The finding that 16.4-45.0% of incorrect answers contain 'valid alternative reasoning' is used to argue that binary scoring understates model understanding. However, the inter-rater reliability for incorrect-answer classifications is only moderate (PABAK = .64), and the paper does not describe how disputes were resolved or how reviewers distinguished 'valid alternative' logic from an unacceptable rationale. Please document the decision rule, report the number of cases needing adjudication, and provide representative examples of 'valid alternative reasoning' from incorrect answers.","section":"Methods: Expert Review of Explanations (p. 12-13); Table 5"}],"minor_comments":[{"comment":"The abstract contains the typo 'STAT questions'; this should read 'SATA questions' (select-all-that-apply).","section":"Abstract"},{"comment":"The Chinese text contains what appears to be a typo in the law name: '中华人民共和未国未成年人保护法' should likely be '中华人民共和国未成年人保护法'; please verify against the original guide.","section":"Appendix A, first SATA example"},{"comment":"Please clarify whether the 2024-published guide materials reproduce the 2023 exam verbatim or are a compilation of similar items; this affects how the contamination risk should be interpreted.","section":"Methods: CNSWE"},{"comment":"Please report the exact API access dates and model snapshot identifiers, since cloud models are updated without notice and the authors themselves note that the results are a point-in-time snapshot.","section":"Methods: Data Management"},{"comment":"In the extracted text the y-axis labels and legends are not legible; please ensure the final figures have clearly visible labels for normalized vs. raw scores and for question types.","section":"Figures 2 and 3"},{"comment":"The reference to 'advanced reasoning models like DeepThink' lacks a citation or a specific model identifier; please add a reference or remove the sentence.","section":"Discussion (p. 28)"}],"recommendation":"major_revision","confidential_remarks":"For the editor: The manuscript addresses a timely topic, and the empirical design has notable strengths (official exam, expert review, open data). The main risk is that the central claims may be overinterpreted relative to the training-data contamination and uncertainty limitations. If the authors can supply robustness checks (e.g., contamination probes, uncertainty intervals) or explicitly temper the conclusions, the paper could be a valuable contribution to the social work and AI literatures. I would also encourage the editor to verify that the 'first systematic cross-cultural benchmark' claim is consistent with the existing literature, as several related studies on cultural alignment of LLMs are cited but not compared directly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read if you care about cross-cultural LLM evaluation. This is the first systematic comparison of four Chinese and four Western models on the Chinese National Social Work Examination, using the official 2023 intermediate-level items. The authors score under official rules, run three conditions (forced response, optional skip, options-only guessing), and add an expert review of model explanations with inter-rater reliability. The code is on GitHub. That is real work.\n\nThe headline numbers are credible: seven of eight models clear the 60-point bar on both sections; Chinese models lead on jurisprudence (median 77 vs 70.3), not on applied knowledge (65.5 vs 67). The options-only condition is a nice check: models land at 37-38%, well above chance but far below the 73% ChatGPT hit on the ASWB, suggesting this exam has less construct-irrelevant variance.\n\nThe main soft spot is exactly the one the stress test flags: the 2023 items came from a 2024-published self-study guide, and the models' training cutoffs are close to that date. The paper concedes it cannot rule out that the questions were in training data. That clouds the pass-rate and the Chinese-vs-Western score comparison. I'd call this a legitimate threat, not a lethal one, because the reasoning analyses partially survive. The expert review found that 16-45% of incorrect answers still had valid professional reasoning, and the cultural-bias examples (e.g., sons favored over daughters despite equal legal standing) come from generated explanations, not just answer recall. So the 'technical language does not equal cultural competence' point is on firmer ground than the 'foundational knowledge' claim.\n\nOther soft spots: one API run per condition with no temperature variance or error bars; n=8 models, so the medians are descriptive; the Chinese vs Western binary is crude; and the cultural-bias finding is more illustrative than systematic—there is no formal content-analysis protocol with counts or reliability for the bias categories. Still, the paper is honest about most of these.\n\nBottom line: a useful, transparent benchmark paper. The contamination risk means the results should be read as a snapshot with a caveat, not as a definitive proof of cultural incompetence. It deserves a serious referee and probably a revision that either tests newer models with truly held-out questions or explicitly frames the scores as upper bounds. If I were in the editor's chair, I'd send it out.","headline":"Solid, honest benchmark of LLMs on the Chinese social work exam; the training-contamination risk is real but the reasoning analyses keep the paper useful.","tokens_in":20207,"tokens_out":3801,"would_cite":true,"duration_ms":31076,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Seven of eight tested large language models pass both written sections of the Chinese National Social Work Examination, while both Chinese- and Western-built models show cultural blind spots on gender and family scenarios.","keywords":["large language models","Chinese National Social Work Examination","cultural competence","cross-cultural assessment","professional licensure","cultural bias","artificial intelligence in social work","select-all-that-apply scoring"],"falsifier":"Retest the same eight models on a freshly written 160-question exam built to the same content blueprint and official scoring rules; if scores fall to the 13.8%–26.3% guessing range or the gender and family bias patterns disappear, the reported results would be explained by memorization of the published 2023 items rather than by cultural-professional knowledge.","tokens_in":1940,"feed_emoji":"🤖","tokens_out":6208,"duration_ms":131482,"temperature":0.7,"pith_summary":"The paper sets out to establish whether large language models understand Chinese social work as a professional and cultural domain, not just as Chinese-language text. Using the 160 written questions from the 2023 intermediate-level Chinese National Social Work Examination, the authors scored eight cloud models, four Chinese and four Western, under official rules. Seven of the eight cleared the 60-point passing bar on both the jurisprudence and applied-knowledge sections; Chinese models led on jurisprudence (median 77.0 vs. 70.3) but not on applied knowledge (65.5 vs. 67.0). Expert review found that models frequently produced professionally valid reasoning even when their chosen answers were marked wrong, and that both groups exhibited cultural biases, especially in inheritance and family scenarios where sons were favored despite equal-rights law. The paper's conclusion is that strong command of professional terminology does not guarantee cultural competence, which matters for any attempt to deploy AI in cross-cultural social work.","feed_headline":"Seven of eight AI models pass China's social work exam","feed_subtitle":"Chinese models lead on law questions; both sides stumble on gender and family scenarios.","key_machinery":"The central instrument is the 2023 intermediate-level Chinese National Social Work Examination (CNSWE), a standardized 160-question test whose jurisprudence and applied-knowledge sections are scored with official rules, including partial credit for select-all-that-apply items. It serves as a culturally grounded proxy for foundational social work knowledge: jurisprudence questions test command of Chinese law and policy, and applied-knowledge questions test practice reasoning. The study wraps this instrument in a three-condition testing protocol, required answers, optional skipping with confidence ratings, and options-only presentation to detect test-pattern artifacts, and then adds bilingual expert review of the models' explanations to separate genuine reasoning from pattern matching.","core_discovery":"The central claim, stated on the paper's own terms, is that current large language models have enough embedded Chinese social work knowledge to pass a national licensing examination, but this knowledge is uneven and culturally shallow in specific ways. Seven of eight models exceeded the official 60-point threshold in both sections, with only DeepSeek-2.5 falling slightly below (59.5) on applied knowledge. Chinese models outperformed Western models on jurisprudence (median 77.0 vs. 70.3) but not on applied knowledge (65.5 vs. 67.0), and the Chinese advantage disappeared on select-all-that-apply items. Explanations attached to wrong answers showed valid professional reasoning 16.4% to 45.0% of the time, and models on both sides exhibited biased judgments in scenarios about gender equality and family dynamics despite clear legal provisions. The paper concludes that technical language ability and formal regulatory knowledge do not ensure culturally competent practice, and that licensing-style tests alone overstate practical cultural knowledge.","pith_inferences":["Beyond the paper: a direct deployment audit could ask the same models to produce open-ended advice for an inheritance or custody scenario; if the patriarchal bias appears there too, it is a client-safety issue rather than a test artifact.","Beyond the paper: the same three-condition design could be run on a Chinese professional exam from another field, such as law or nursing, to test whether the jurisprudence-versus-application gap is specific to social work or a general feature of how models handle formal versus scenario-based Chinese content.","Beyond the paper: the authors' retrieval-augmented-generation suggestion implies a testable prediction, namely that giving models a curated Chinese social work practice manual at inference time should improve applied-knowledge scores more than jurisprudence scores, because the missing resource is practice knowledge rather than policy text.","Beyond the paper: because 'valid alternative reasoning' was judged by bilingual social work professionals against an official guide, the 16.4% to 45.0% range may partly reflect reviewers' professional norms; replicating the review with practitioners from different Chinese regions would test how stable that estimate is."],"forward_implications":["If the pass-rate results hold, cloud LLMs already carry enough Chinese social work content to act as knowledge-delivery aids, which shifts the design problem from basic capability to safeguards, source verification, and professional oversight.","Because Chinese models beat Western models on jurisprudence but not applied knowledge, local training data appears to confer an advantage in formal policy and legal text rather than in culturally specific practice scenarios.","The finding that 16.4% to 45.0% of wrong answers contained valid reasoning implies that binary pass/fail scoring understates models' professional understanding and may misclassify legitimate alternative approaches.","The biased inheritance reasoning, despite explicit equal-rights law, implies that training-data stereotypes can override legal and professional frameworks in model outputs, so cultural-bias auditing is a prerequisite for deployment.","The options-only condition staying above chance but far below the prior U.S. licensing-exam result (73.3%) suggests the Chinese exam contains less construct-irrelevant variance, making it a relatively cleaner test of knowledge."],"supporting_citations":[{"why":"The 2024a and 2024b guide materials provide the 160 intermediate-level Chinese National Social Work Exam questions used as the test set.","marker":"National Social Worker Professional Exam Question Compilation Group"},{"why":"Demonstrated that ChatGPT can reason about the ASWB licensing exam, establishing licensure tests as a probe for social work knowledge.","marker":"Victor, Kubiak, Angell, and Perron (2023)"},{"why":"Supplies the options-only construct-irrelevant-variance methodology and the 73.3% guessing-rate comparison used in Condition 3.","marker":"Victor et al. (2024)"},{"why":"Documents Western cultural alignment in popular LLMs, framing the cross-cultural bias question that motivates the study.","marker":"Tao et al. (2024)"},{"why":"Provides evidence that LLMs misrepresent local cultural values even when prompted in different languages, supporting the language-versus-culture distinction.","marker":"Cao et al. (2023)"},{"why":"Establishes the construct-irrelevant variance testing approach for licensing exams that Condition 3 adapts.","marker":"Albright & Thyer (2010)"},{"why":"Establishes the official 60-point passing threshold that defines successful performance in both exam sections.","marker":"Ministry of Human Resources and Social Security (2022)"},{"why":"Describes the CNSWE's intermediate-level structure and historical pass rates, grounding the choice of exam level.","marker":"Zeng, Li, & Chen (2019)"}],"fun_headline_variants":["Seven of eight AI models pass China's social work exam, but lose on culture","AI passes China's social work exam, flunks cultural nuance","Chinese AI wins law, but all AI stumbles on gender scenarios","AI models pass social work exam, yet show cultural blind spots","Seven AIs pass China's exam; all show gender bias"],"cache_read_input_tokens":22400,"weakest_assumption_plain":"The results depend on the 2023 exam questions not having appeared in the models' training data, a point the paper's limitations section concedes cannot be confirmed, because if the models memorized the questions the pass rates and bias findings would reflect recall rather than knowledge or reasoning.","fun_headline_variants_meta":{"raw":{"variants":["Seven of eight AI models pass China's social work exam, but lose on culture","AI passes China's social work exam, flunks cultural nuance","Chinese AI wins law, but all AI stumbles on gender scenarios","AI models pass social work exam, yet show cultural blind spots","Seven AIs pass China's exam; all show gender bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000379,"raw_usage":{"total_tokens":2039,"prompt_tokens":993,"completion_tokens":1046,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":955}},"tokens_in":609,"tokens_out":1046,"duration_ms":7287,"temperature":1.0,"reasoning_tokens":955,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:44:14.074499+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retest the same eight models on a freshly written 160-question exam built to the same content blueprint and official scoring rules; if scores fall to the 13.8%–26.3% guessing range or the gender and family bias patterns disappear, the reported results would be explained by memorization of the published 2023 items rather than by cultural-professional knowledge.","supporting_citations":[],"review_version":1}