{"id":"a264119a-99a4-4e3c-85e6-d877219ae829","arxiv_id":"2608.01629","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Native Japanese raters and six LLMs both penalize L2-written Japanese emails on fluency, status, and solidarity, with AIs reproducing the bias in attenuated form and over-differentiating learner backgrounds.","lead":"Researchers compared how native Japanese speakers and six AI chatbots rate emails written by non-native Japanese learners. Both humans and the AIs marked the non-native emails lower, showing the AIs copy a human bias that could affect automated hiring and school assessments.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human-LLM comparison uses mismatched item sets: human baseline from 347 emails vs. LLM gaps from 145 pairs (290); unaddressed and could shift the reported divergences.","rationale":"The paper's core empirical contribution—a direct human-LLM comparison using parallel L1/L2 emails—rests on comparing human gaps (Experiment 1) with LLM gaps (Experiment 2). The 'Email Stimuli' section states humans rated 174 L1 + 173 L2 emails, while the LLM experiment used only 145 complete L1-L2 pairs. If the human baseline is computed on the larger set, the divergence claims ('all understated the solidarity gap', 'humans did not differentiate L1 backgrounds') are not strictly apples-to-apples. This is the most load-bearing concern because it is unacknowledged, directly testable, and could alter the headline divergences. The register confound flagged by the reader is a valid limitation and is explicitly conceded in the manuscript; the item-set mismatch is not conceded and is at least as threatening to the central comparison: if the 57 extra emails differ systematically from the paired ones, the human reference line itself moves. The human results are otherwise credible (large sample, factor analysis consistent with the three dimensions), and the LLM pattern is consistent across six models; the free-text justifications are supporting rather than central. My proposed reanalysis would settle whether the headline divergences survive on a common item set. Since the issue is fixable by reanalysis, the conditional verdict stays appropriate.","tokens_in":7701,"tokens_out":17499,"duration_ms":166672,"concrete_test":"Recompute the three human linear mixed-effects models (and the learner-L1 likelihood-ratio tests) using only the 145 complete L1-L2 pairs that were presented to the LLMs, i.e., the same 290 emails. Compare the writer-type β coefficients for fluency, status, and solidarity, the writer-type × subscale interaction, and the L1-background LR-test p-values to the reported values (fluency β = −1.11; status β = −0.57; solidarity β = −0.49; status vs solidarity β = 0.07, p = .23; learner-L1 ps ≥ .055). If any β shifts by more than ~0.1 or a p-value crosses 0.05, the claims that 'all LLMs understated the solidarity gap' and 'humans did not differentiate L1 backgrounds' need to be revisited. Ideally, also report the exact overlap between the 347 human-rated emails and the 290 LLM-rated emails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 'Email Stimuli' reports human ratings on 174 L1 and 173 L2 emails (347 total), while the LLM experiment uses only the 145 complete L1-L2 pairs (290 emails). All human β coefficients (fluency −1.11, status −0.57, solidarity −0.49) and the learner-L1 likelihood-ratio tests (ps ≥ .055) come from the full 347-email set, whereas every LLM gap is computed on the 290-email pair set. The paper never reports a human analysis on the common 290-email subset, nor the overlap between the two item sets. If the 57 unpaired emails (e.g., missing a counterpart in the corpus or dropped by the >5-ratings screening rule) differ from the paired emails in speech-act composition, L1 background, or perceived difficulty, the human baseline is not the correct reference for the LLM gaps. The claimed divergences—'all understated the solidarity gap' and 'humans did not differentiate among learner L1 backgrounds'—depend directly on this comparison. This is an unacknowledged methodological mismatch, separate from the conceded register limitation, and is checkable with the existing data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares human and LLM evaluations of Japanese emails written by L1 speakers and L2 learners, using matched L1-L2 email pairs from the I-JAS FOLAS corpus. Human raters (1,536 native Japanese speakers, each rating one email) evaluated 347 emails on three Likert subscales—fluency, status, and solidarity—and penalized L2 emails on all three dimensions, with the fluency gap (β = −1.11) roughly twice the status (β = −0.57) and solidarity (β = −0.49) gaps; the status-over-solidarity ordering was not statistically significant. Six LLMs, prompted in Japanese with no writer identity, also penalized L2 emails on all three dimensions; five of six reproduced the human dimension ordering. All LLMs understated the human solidarity gap, and all LLMs showed significant learner-L1 differentiation that human raters did not. The paper argues that the language-attitudes framework provides a ready-made audit yardstick for LLM evaluators beyond English.","tokens_in":7880,"tokens_out":7457,"duration_ms":85646,"significance":"If the results hold, the paper makes a valuable contribution: it extends the fluency principle from speech to writing, provides the first quantitative human-LLM comparison of language attitudes in Japanese, and uses a content-controlled corpus with factor-validated Japanese subscales. The design has notable strengths: a fully between-subjects human rating procedure, an EFA-confirmed three-factor structure, deterministic LLM calls with verified reproducibility, and direct comparisons on identical scales. The main divergences reported—attenuated solidarity gaps and learner-L1 differentiation in LLMs—are provocative and practically important for LLM-as-judge deployment. However, two load-bearing points need additional work before the headline claims are fully supported: the human and LLM analyses are computed on different item sets, and the native-rewrite manipulation leaves an acknowledged register confound unbounded.","major_comments":[{"comment":"The human baseline and the LLM results are computed on different stimulus sets. The human βs (−1.11 fluency, −0.57 status, −0.49 solidarity) and the human learner-L1 likelihood-ratio tests come from all 347 rated emails (174 L1, 173 L2), while the LLM gaps and LLM learner-L1 tests use only the 145 complete L1-L2 pairs (290 emails). The 57 unpaired emails are never described. If they differ from paired emails in speech-act mix, learner-L1 composition, or difficulty, the headline divergences ('all understated the solidarity gap'; 'humans did not differentiate' vs 'all models differentiated') are not comparisons on equal footing. This is checkable with the existing data and should be reported.","section":"§Email Stimuli; §Analysis; Results"},{"comment":"The causal interpretation that the L1/L2 gaps reflect attitudes to non-native writing presupposes that the native rewrites differ from the L2 originals only in nativeness. The paper acknowledges that the rewrites 'may not eliminate subtle register differences,' but the acknowledgment does not bound the confound. Because status and solidarity ratings are sensitive to politeness and formality, a systematic register difference is an alternative explanation for the size and ordering of the gaps. Please add a manipulation check (e.g., native-speaker ratings of formality or appropriateness) or a robustness analysis restricted to pairs matched on register.","section":"§Email Stimuli; §Discussion (Limitations)"},{"comment":"The learner-L1 divergence rests on 18 unadjusted tests (6 models × 3 dimensions). Human null results (p ≥ .055, three tests) are interpreted as 'did not differentiate,' yet no equivalence test is provided; failing to reject a null is not strong evidence of absence. For LLMs, 17/18 tests significant at p < .05 are reported without multiple-testing control, and the p-values are not given, so the reader cannot judge whether the claim survives. Please report adjusted p-values (FDR or Bonferroni) and effect sizes, and consider a rater-type × learner-L1 interaction test for the human-LLM contrast.","section":"§Results (Human Raters; LLMs)"}],"minor_comments":[{"comment":"The statement 'the fidelity of reproduction scaled with model capability' is asserted without a formal statistical test. With only six models, this claim should be softened or supported by a quantitative association.","section":"§Discussion, first paragraph"},{"comment":"Human distributions are over individual ratings (n ≈ 770 per condition) while LLM distributions are over item-level scores (n = 145). Comparing the widths of these distributions is misleading; please show LLM distributions at a comparable aggregation level or clearly label the different units.","section":"Figure 3"},{"comment":"The screening step that removed 433 submissions 'belonging to emails that had received more than the five planned ratings' needs clarification: if each rater saw one randomly assigned email, how could an email exceed a planned quota? Please describe the assignment/quota mechanism.","section":"§Method, Human Raters"},{"comment":"The automated free-text coding is non-validated and non-exclusive, as the authors note. The classification of 丁寧 'polite' as a solidarity keyword is debatable; please provide the full keyword list and a sensitivity analysis excluding ambiguous keywords.","section":"§Method, Analysis; Figure 4"},{"comment":"H2 is stated as a three-way ordering (fluency > status > solidarity), but the status-over-solidarity comparison was not significant. Consider splitting H2 into two component predictions so that the partial support is not masked by the unsupported ordering.","section":"Introduction, H2"}],"recommendation":"major_revision","confidential_remarks":"The central finding is plausible and the paper is well positioned for a sociolinguistics or CL venue. The item-set mismatch is the key obstacle: it is a load-bearing issue for the headline human-LLM comparison, and it is fixable with the existing data. I recommend major revision rather than rejection because the required re-analyses are feasible without new data collection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's what I'd tell you about this paper. It does something genuinely new: it takes the matched-guise approach from English dialect work and applies it to non-native Japanese writing, comparing native raters to six LLMs on content-controlled L1/L2 emails. The human experiment is well designed—parallel emails from I-JAS, one rating per rater, mixed-effects models, and the authors honestly report that the predicted status-over-solidarity gap is not significant. The core result is robust: L2 emails are penalized on fluency, status, and solidarity, with the fluency gap roughly twice the others. And five of six LLMs reproduce the ordering, though all understate the solidarity gap and all differentiate among learner L1s where humans don't. That's a meaningful contribution to LLM-as-judge fairness outside English, and the paper deserves a serious referee.\n\nThe soft spots are real. The biggest is the mismatched item sets: human gaps are estimated on all 347 emails, while LLM gaps come from the 145 complete pairs (290 emails). The paper never reports the human coefficients on that common subset, so the baseline you're comparing LLMs against isn't exactly what the LLMs saw. This is checkable and should be reported; it might shift the quantitative divergences, though I doubt it flips the qualitative pattern. Second, the screening rule that dropped 433 submissions for emails with more than five ratings is under-explained and arguably post hoc — it could bias the stimulus set. Third, the concession that native rewrites may not fully control register is important, because the fluency subscale overlaps with the L1/L2 contrast by construction. The lack of data/code deposit makes these harder to probe.\n\nNone of this breaks the central claim. The direction of the bias is consistent across all models, and the key divergences are large enough that a rerun on the paired subset would likely hold. But the paper needs a revision that addresses the item-set mismatch and the screening rule before I'd call it sound. I'd bring it to reading group and would cite it if I were working on non-English LLM evaluation. Send it out.","headline":"A solid, novel human-LLM comparison of language attitudes in Japanese, with an addressable item-set mismatch that should be fixed before publication.","tokens_in":8437,"tokens_out":2634,"would_cite":true,"duration_ms":32004,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Native Japanese raters penalize L2 emails on fluency, status, and solidarity; six LLMs reproduce the penalty without being told the writer is non-native.","keywords":["language attitudes","non-native Japanese","LLM-as-a-judge","fluency principle","status and solidarity","Japanese email evaluation","bias in large language models","I-JAS corpus"],"falsifier":"Present the six LLMs and, separately, a new panel of Japanese raters with one native-written email under two conditions: identical text with no writer information, and identical text with a metadata line naming the writer's L1. If the metadata line alone shifts the solidarity or status gap, category-based attitudes are at work independent of form; if ratings are unchanged, the bias is carried entirely by surface non-native form.","tokens_in":7524,"feed_emoji":"🤖","tokens_out":6448,"duration_ms":73594,"temperature":0.7,"pith_summary":"This paper asks whether large language models share native speakers' negative attitudes toward non-native writing, using Japanese as the test case. Across content-matched emails written by native speakers and learners, 1,536 Japanese raters scored learner emails lower on all three attitude dimensions—fluency, status, and solidarity—with the fluency gap roughly twice the status and solidarity gaps. Six LLM judges, given the same emails with no writer identity, reproduced the direction of that bias, and five reproduced the ordering across dimensions. The models differed from humans in a consistent way: every model understated the solidarity gap, and every model distinguished among learner L1 backgrounds where human raters did not. The paper concludes that human language attitudes are encoded, in attenuated form, in LLM evaluators, and that the language-attitudes framework can serve as a non-English yardstick for auditing them.","feed_headline":"Six LLMs reproduce Japan's bias against non-native writing","feed_subtitle":"Machine judges copy the fluency-first penalty but soften the warmth gap and split learner groups humans treat alike.","key_machinery":"The fluency principle—borrowed from speech-accent research and extended here to writing—is the load-bearing mechanism: non-native form is harder to process, and processing difficulty lowers evaluations most on fluency, about half as much on status, and least on solidarity. The study's engine is the I-JAS FOLAS corpus, which provides learner-written Japanese emails and native rewrites of the same content, so that content is held constant and only nativeness of form varies. The LLM-as-judge protocol converts the same nine Likert items into a prompt battery administered to six models, making human and machine evaluations directly comparable.","core_discovery":"In a content-controlled comparison of Japanese emails written by native speakers and learners, native raters scored L2 emails lower on all three attitude dimensions, with the fluency gap (−1.11) roughly twice the status gap (−0.57) and solidarity gap (−0.49). Six LLMs, given no writer identity, reproduced the direction of this bias, and five of six reproduced the ordering across dimensions. The models diverged from humans in two ways: no model reached the human solidarity gap, and all six differentiated among learner L1 backgrounds on at least two dimensions, where human raters showed no such differentiation. The paper interprets this as language attitudes encoded in LLMs, with fluency being","pith_inferences":["If the solidarity gap is driven by outgroup categorization rather than text, a matched-guise variant of these emails—identical text but an implied or labeled writer origin—should reproduce a human solidarity gap; this would cleanly separate category-based from form-based bias.","The open-weight 8B models showed the weakest L1/L2 gaps, suggesting bias magnitude may scale with model capability; if so, larger future models will track human language attitudes more closely, for better and worse.","The authors' closing observation about AI-polished non-native writing implies a testable prediction: when learners revise with LLMs and surface errors vanish, the fluency penalty should shrink, leaving any residual bias to register or content cues.","The models' differentiation among learner L1 backgrounds suggests their training data carries frequency signals of particular learner varieties; auditing those distinctions could reveal new bias axes absent from human attitude surveys."],"forward_implications":["In automated screening of Japanese writing for hiring or assessment, non-native surface form alone can trigger lower competence and warmth ratings even when writer identity is hidden.","The three-scale language-attitudes battery offers a ready-made, non-English audit tool for LLM evaluators.","Because all models understated the solidarity gap, LLM judges reproduce less of the interpersonal warmth penalty than humans do, while still approximating human-level fluency and status penalties.","Because models distinguished among learner L1 backgrounds where humans did not, LLM judges may introduce new, text-driven hierarchies among non-native groups.","Extending the fluency principle to writing means content-matched written corpora can test processing-based bias without acoustic confounds, opening a new empirical route for sociolinguistics."],"supporting_citations":[{"why":"Supplies the status/solidarity dimensions and the state of the field that the study extends to Japanese writing.","marker":"Dragojevic et al., 2021"},{"why":"Meta-analysis establishing that non-native or accented speakers are penalized more strongly on status, the prediction H2 extends.","marker":"Fuertes et al., 2012"},{"why":"Defines the fluency principle that the paper extends from speech to writing.","marker":"Dragojevic, 2020"},{"why":"Shows that surface errors in email alone shape person perception, grounding the written-text design.","marker":"Vignovic & Thompson, 2010"},{"why":"The I-JAS FOLAS corpus is the source of the content-matched L1/L2 email stimuli.","marker":"Sakoda, 2020"},{"why":"Adapts matched-guise testing to LLMs, the direct precedent for probing LLM language attitudes.","marker":"Hofmann et al., 2024"},{"why":"Frames the LLM-as-a-judge deployment that gives the bias real-world stakes.","marker":"Gu et al., 2026"}],"fun_headline_variants":["LLMs mirror Japan's bias against non-native writing","AI judges copy human bias on non-native Japanese","LLMs show same fluency bias as humans on Japanese","Machine raters inherit Japan's language prejudice","Six LLMs reproduce native bias on L2 Japanese emails"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The entire L1/L2 comparison rests on the assumption that the native rewrites of the learner emails differ only in nativeness; the paper concedes they may retain subtle register differences, so part of the gap could be politeness or formality rather than non-native form.","fun_headline_variants_meta":{"raw":{"variants":["LLMs mirror Japan's bias against non-native writing","AI judges copy human bias on non-native Japanese","LLMs show same fluency bias as humans on Japanese","Machine raters inherit Japan's language prejudice","Six LLMs reproduce native bias on L2 Japanese emails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000195,"raw_usage":{"total_tokens":1160,"prompt_tokens":679,"completion_tokens":481,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":423,"completion_tokens_details":{"reasoning_tokens":407}},"tokens_in":423,"tokens_out":481,"duration_ms":5090,"temperature":1.0,"reasoning_tokens":407,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T23:50:58.764733+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Present the six LLMs and, separately, a new panel of Japanese raters with one native-written email under two conditions: identical text with no writer information, and identical text with a metadata line naming the writer's L1. If the metadata line alone shifts the solidarity or status gap, category-based attitudes are at work independent of form; if ratings are unchanged, the bias is carried entirely by surface non-native form.","supporting_citations":[],"review_version":1}