{"id":"13c6565d-787f-4872-b55e-3aa499cae06a","arxiv_id":"2508.02502","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Prompted language identity changes both the outputs and the internal layer representations of Llama-3.3-70B and Qwen2.5-72B on sound symbolism and word valence tasks.","lead":"LLMs shift their judgments about word sounds and emotional tone depending on whether they are prompted as Dutch, Chinese, or bilingual speakers. The paper shows this through two psycholinguistic tasks and by decoding the models' internal layers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Language identity is confounded with prompt and stimulus language: no condition varies the persona while holding task text and words fixed, so the central claim is underdetermined.","rationale":"The reader identifies the English-only ground truth as the weakest assumption. That is a real limitation for interpreting accuracy as psycholinguistic alignment, but it does not directly threaten the more basic claim that outputs and internal representations differ under different language conditions. The stronger threat is causal: the paper's manipulation of 'language identity' is confounded with instruction language in every condition and with stimulus language in Task 2, so the observed differences cannot be attributed uniquely to the assigned identity. This is a distinct concern, hence partial agreement. The concern is not fatal to all of the paper: Task 1 uses identical pseudowords across personas, and even Task 2 demonstrates that prompts differ in their effects; what is undermined is the specifically cognitive 'language identity' interpretation and the probing comparisons between Dutch and Chinese. Because the authors could address this with an identity-only control or factorial design, conditional acceptance remains the right verdict; I would not reject outright. The proposed test would settle whether the central claim survives the confound.","tokens_in":13740,"tokens_out":10511,"duration_ms":134452,"concrete_test":"Run a controlled experiment on Qwen2.5-72B-Instruct: keep all user-message task text and all stimulus words in English (or a neutral language), and vary only the system-prompt persona sentence (e.g., 'You are a native Dutch speaker' vs 'You are a native Mandarin Chinese speaker'). Re-measure behavior and layer-wise probing accuracy for both tasks. If the Dutch/Chinese differences in Tables 2-4 and Figures 3-4 disappear or shrink to noise, the reported effects are due to prompt/stimulus language, not language identity. As a second check, run Task 2 as a 2x2 factorial crossing persona language (Dutch/Chinese) with stimulus word language (Dutch/pinyin Chinese) to estimate the identity effect holding stimuli fixed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'language identity conditions both output behavior and internal representations' requires comparisons in which only the identity (persona) changes. The design never provides such comparisons. In Section 4.2, monolingual conditions write both system and user prompts in the target language, while bilingual conditions keep the system prompt in English and put the target language in the user message; thus 'language identity' is always co-varied with the language of the instruction text. The valence task is worse: Appendix A shows that the Dutch persona is asked to judge a romanized Chinese word (e.g., niao), while the Chinese persona is asked to judge a Dutch word (e.g., vogel). So in Task 2, persona language and stimulus word language are perfectly confounded: Dutch prompts always see pinyin words and Chinese prompts always see Dutch words. Tables 2-4 and Figures 3-4 therefore cannot separate an effect of the assigned linguistic identity from an effect of the surface tokens (or their training-data associations). The reader's English-ground-truth concern is genuine, but it is not the only load-bearing gap: even with correct Dutch and Chinese human norms, these comparisons would still be ambiguous. The paper needs an identity-only manipulation or a full crossing of persona language by stimulus language before the abstract's cognitive 'language identity' reading is supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether LLMs exhibit human-like psycholinguistic responses under different linguistic identities by evaluating Llama-3.3-70B-Instruct and Qwen2.5-72B-Instruct on two tasks: sound symbolism (round/spiky judgments on pseudowords) and word valence (positive/negative judgments on translated ANEW words). The authors compare monolingual (Dutch, Chinese, English) and bilingual (Dutch-English, Chinese-English) prompting conditions, reporting behavioral accuracy relative to English-derived ground truth and layer-wise probing accuracy on the hidden states of Llama (and Qwen in the appendix). The central claim is that 'language identity conditions both output behavior and internal representations in LLMs,' with Qwen showing stronger language sensitivity and Chinese prompts yielding more robust valence representations than Dutch.","tokens_in":14010,"tokens_out":5514,"duration_ms":60512,"significance":"If the central claim were established, the paper would contribute useful evidence about how LLMs simulate cross-linguistic psycholinguistic behavior and whether prompt-based persona conditioning affects not just outputs but also internal encodings. Strengths include deterministic generation at temperature 0, the use of two model families with contrasting multilingual training profiles, a layer-wise probing design with hyperparameters in Appendix C, and an explicit acknowledgment in the Table 2 note that English-based labels are a reference rather than an absolute standard. However, the design confounds language identity with the surface language of the prompt and the stimulus, so the observed differences cannot currently be attributed to the assigned linguistic identity. The significance is accordingly contingent on additional control conditions or a substantially weakened interpretation.","major_comments":[{"comment":"'Language identity' is never manipulated independently of the surface language of the prompt and the stimulus. In the monolingual condition both system and user prompts are written in the target language, and in the bilingual condition the system prompt is English while the user prompt is in the target language; the valence task additionally presents pinyin words to the Dutch persona and Dutch words to the Chinese persona. Consequently Tables 2–4 and Figures 3–4 cannot attribute the observed differences to the assigned linguistic identity as opposed to the language of the instruction text or the language of the stimulus word. The central abstract claim requires at least one condition that varies the persona while holding the task text and stimulus language fixed, or a full crossing of persona language and stimulus language.","section":"§4.2, Appendix A"},{"comment":"In Task 2 the persona language and the stimulus language are perfectly confounded. A Dutch persona always judges a romanized Chinese word (e.g., niao) and a Chinese persona always judges a Dutch word (e.g., vogel), so the large Dl values in Table 4 (e.g., +30.50 for Llama monolingual, −41.76 for Qwen bilingual) may reflect stimulus-language effects rather than identity effects. Without crossing persona language with stimulus language, the qualitative claims in Section 4.4 about language conditioning are underdetermined.","section":"§3.2, Appendix A"},{"comment":"All behavioral accuracy scores are computed against English-derived ground truth (ANEW norms and Alper and Averbuch-Elor pseudoword labels). The paper acknowledges in the Table 2 note and Section 4.3 that these scores 'should not be interpreted as absolute accuracy,' but Dm and Dl are still interpreted as alignment shifts. If native Dutch or Chinese human judgments differ from the English labels, these discrepancies measure divergence from an English reference, not psycholinguistic alignment. The paper should report raw response distributions (e.g., proportion 'positive' or 'round' per condition) or use language-specific human norms, and it should at least quantify how many of the observed differences survive when the evaluation is rerun with an alternative reference.","section":"§4.3, Tables 2–4"},{"comment":"The discrepancies Dm and Dl are reported without any measure of uncertainty. Although generation is deterministic at temperature 0, the estimates are computed over 648 and 1034 items, so binomial confidence intervals or bootstrap intervals are needed before claims such as 'Qwen's behavior reverses' (Section 4.3.2) or '+46.5%' (Table 3) can be evaluated. The probing results in Figures 3–6 likewise lack error bars or significance tests across random restarts of the probe.","section":"§4.3, Tables 3–4"}],"minor_comments":[{"comment":"The paper notes that synonymous responses such as 'joyful' instead of 'positive' are penalized, but it does not quantify how many responses are affected or provide an alternative evaluation (e.g., semantic equivalence matching). This could systematically lower accuracy in specific conditions and should be reported.","section":"§4.3.1"},{"comment":"The English translation in the valence example reads 'The word is bird,' which does not correspond to the Dutch/Chinese words shown; this appears to be a typo and should be corrected to reflect the actual stimulus (e.g., 'niao' or 'vogel').","section":"Figure 2"},{"comment":"The qualitative claim that nasal-initial words are perceived more positively in Chinese and more negatively in Dutch is not consistently supported by Table 5; for example, 'fear' (hai pa) is labeled negative in Chinese and 'problem' (ma fan) is labeled negative in Chinese despite the nasal onset. The text should either present a quantitative test of the nasal-onset hypothesis or temper the generalization.","section":"§4.4, Table 5"},{"comment":"The statement that translations were 'manually verified and cross-checked' is not accompanied by any procedure or inter-annotator detail; a brief description of the verification process would strengthen reproducibility.","section":"§1"},{"comment":"The probe hyperparameters are listed, but there is no mention of how many random seeds were used or whether the reported probing accuracies are averaged over seeds; this is needed because probing accuracy can vary with initialization.","section":"Appendix C"},{"comment":"The limitations listed (three languages, two models, narrow task set) do not include the English-based ground truth or the prompt-language confound; these should be explicitly acknowledged, or the claims in the abstract should be weakened accordingly.","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the journal's psycholinguistics/NLP audience, but the central claim is currently over-stated relative to the design. I would advise the editor that the revision must require either an identity-only manipulation or a substantive weakening of the abstract's 'language identity' claim; a purely textual revision may not be sufficient."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: the behavioral effects are real, but the paper's headline claim that prompted language identity conditions psycholinguistic judgments and internal representations is not supported by the design. The stress-test note is right, and it's the load-bearing issue.\n\nWhat's new and worth credit: the specific combination of monolingual vs bilingual persona prompting with layer-wise probing for sound symbolism and word valence appears not to have been done before. The authors use two models with different multilingual training (Llama-3.3-70B vs Qwen2.5-72B), temperature-0 generation, external human benchmarks, and manual verification of translations. The deterministic output differences across prompts are solid observations. No circularity: they don't fit to their own outputs.\n\nThe soft spot is experimental control. In the valence task, the Dutch persona always judges romanized Chinese words and the Chinese persona always judges Dutch words. Persona language and stimulus word language are perfectly confounded. In the sound symbolism task, the pseudowords are the same, but the instruction language changes with the persona (monolingual conditions use Dutch or Chinese instructions; bilingual conditions use English system prompts with Dutch or Chinese user messages). So in every contrast, \"language identity\" co-varies with the surface tokens the model sees. Even with correct Dutch and Chinese human norms, these comparisons couldn't separate an effect of the assigned persona from an effect of the prompt or stimulus language. That's a fixable flaw: a full crossing of persona by stimulus language, or holding the instruction text fixed while varying only the persona, would support the claim. As it stands, the abstract's \"language identity conditions\" statement overreaches.\n\nSecondary issues are real but lower stakes. The English-derived ground truth is acknowledged but not controlled for; exact-match scoring penalizes synonyms; and there are no significance tests or item-level statistics, so the magnitude claims (e.g., Qwen's \"sharper distinctions\") aren't quantified beyond point estimates.\n\nWho's this for? Researchers probing multilingual LLM behavior and anyone designing persona-prompting experiments. It's a useful reminder that persona effects are easy to confound with prompt-surface effects. I'd send it to review because the question is worth asking, the execution is mostly clean, and a revised design with proper controls could turn this into a solid contribution. But on the current evidence, I would not cite the central claim.","headline":"Real deterministic prompt effects, but the 'language identity' framing is undercut by a design that never isolates identity from prompt or stimulus language.","tokens_in":14502,"tokens_out":4497,"would_cite":false,"duration_ms":51622,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that prompting an LLM as a Dutch, Chinese, or English speaker changes both the psycholinguistic judgments it produces and how its internal layers encode sound-symbolic and valence information, so these models are not…","keywords":["psycholinguistics","sound symbolism","word valence","language conditioning","multilingual LLMs","Bouba-Kiki effect","probing analysis","cross-linguistic cognition"],"falsifier":"Collect native-speaker valence and sound-shape norms from Dutch and Mandarin participants for the same 648 pseudowords and 1,034 translated words, and compare the models' Dutch- and Chinese-prompted judgments against those native norms. If Dutch-prompted outputs align better with English norms than with Dutch human ratings, the claim that language identity changes psycholinguistic cognition in the model would be weakened.","tokens_in":13577,"feed_emoji":"🧠","tokens_out":8129,"duration_ms":88286,"temperature":0.7,"pith_summary":"This paper asks whether large language models behave like language-specific minds when told to adopt a linguistic identity. Using two psycholinguistic tasks, sound symbolism (round versus spiky pseudowords) and word valence (positive versus negative real words), the authors prompt Llama-3.3-70B and Qwen2.5-72B as monolingual or bilingual speakers of English, Dutch, or Chinese. The central claim is that language identity conditions both output behavior and internal representations: the same model changes its judgments and its layer-by-layer decodability depending on the prompted language. A reader should care because it bears on whether LLMs can serve as models of cross-linguistic human cognition, and because it shows that assumptions of language-neutral LLM behavior are false.","feed_headline":"Told to be Dutch or Chinese, LLMs judge words differently","feed_subtitle":"Sound-symbolism and valence judgments shift with prompted language, and so do the models' internal layers.","key_machinery":"The central mechanism is prompt-based language conditioning: a system prompt assigns the model a persona such as 'You are a native speaker of Dutch and do not speak any other language' or 'bilingual speaker of English and Mandarin Chinese', with the user message written in the matching language. Two psycholinguistic tasks carry the evaluation: 648 pseudowords with round/spiky shape labels (the Bouba–Kiki effect) and 1,034 real words translated from the ANEW valence norms into Dutch and Chinese, with Pinyin provided for Chinese. Two discrepancy metrics, $D_m$ and $D_l$, turn raw accuracy into measures of how much the prompted language moves behavior. For internal evidence, the paper trains a frozen MLP probe on hidden states at eight layers of the 80-layer Llama model and compares how decodable sound-symbolic and valence labels are across language conditions.","core_discovery":"On the paper's own terms, LLMs are not language-neutral. In the sound symbolism task, both models classify pseudowords differently when prompted as Dutch, Chinese, or English speakers, and in the word valence task they produce divergent positive/negative judgments for the same translated words. The behavioral differences are quantified by two discrepancy metrics: $D_m$ compares bilingual versus monolingual prompting, and $D_l$ compares Chinese versus Dutch prompts. Layer-wise probing of Llama-3.3-70B on hidden states from every tenth layer shows that psycholinguistic signals become more linearly decodable in deeper layers, that bilingual prompts can delay this emergence, and that Chinese prompts yield stronger, more stable valence representations than Dutch prompts. The authors read these results as evidence that prompt-based language conditioning modulates both output and internal encoding, not merely surface response style.","pith_inferences":["An implication the authors leave implicit is that any cross-lingual LLM evaluation fixing one language's human norms as ground truth confounds language conditioning with norm divergence.","A natural testable extension is to collect native Dutch and Mandarin norms and check whether models prompted in those languages track those norms better than they track English norms.","The layer-wise curves predict that languages whose phonology and script differ sharply from English should alter early-layer decodability trajectories more strongly, a pattern that could be checked in Arabic, Hindi, or Japanese.","Because the paper notes that exact-match scoring penalized synonymous answers such as 'joyful' for 'positive', part of Qwen's Chinese-bilingual instability may be an evaluation artifact; a soft-match rerun would isolate the representational effect."],"forward_implications":["Prompt language should be treated as a controlled variable in psycholinguistic LLM studies, because a single English-only accuracy score does not describe the model's behavior under other linguistic identities.","Monolingual prompts produce earlier and more stable sound-symbolism representations than bilingual prompts in Llama, so bilingual conditioning can delay or diffuse phonological signal in the network.","Chinese prompts yield stronger and more stable valence probing accuracy than Dutch prompts, so valence encoding is language-dependent even when the ground-truth labels are English-derived.","The English-centric Llama is more stable across languages on the behavioral tasks, while the multilingual Qwen shows larger and sometimes reversed language effects, so training-language coverage is a plausible driver of psycholinguistic sensitivity.","Bilingual prompting can either improve or degrade alignment with English labels depending on model, language, and task, as seen in Qwen's large Task-2 gain for Dutch and its loss for Chinese."],"supporting_citations":[{"why":"Provides the 648 pseudowords with round/spiky shape labels that form Task 1's data and English reference labels.","marker":"Alper and Averbuch-Elor, 2023"},{"why":"Supplies the English word set and mean valence ratings used as Task 2's English-derived ground truth after translation.","marker":"Bradley and Lang, 1999"},{"why":"Motivates the Dutch/Chinese comparison and supplies the cross-linguistic valence-from-phonology findings the tasks build on.","marker":"Louwerse and Qu, 2017"},{"why":"Documents Llama-3.3-70B-Instruct, the English-centric model whose behavior and hidden states are tested.","marker":"Grattafiori et al., 2024"},{"why":"Documents Qwen2.5-72B-Instruct, the multilingual model with Dutch and Chinese training data used as the contrast case.","marker":"Qwen et al., 2025"},{"why":"Provides the persona-based system-prompt method used to impose monolingual and bilingual linguistic identities.","marker":"Yuan et al., 2025"},{"why":"Supports the assumption that English-centric models can show emergent multilingual ability, justifying the Llama comparison.","marker":"Nie et al., 2024"}],"fun_headline_variants":["LLM word judgments shift with prompted language identity","Dutch vs Chinese prompts rewire LLM psycholinguistic ratings","Language prompts alter how LLMs rate word sounds and valence","LLMs act human-like when told to speak Dutch or Chinese","Prompted language changes LLM internal encoding of word meaning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the English-derived labels are a valid reference for the Dutch and Chinese conditions; if native Dutch or Chinese speakers would rate these pseudowords and translated words differently than English speakers do, the reported accuracy scores measure deviation from English norms rather than alignment with the prompted language.","fun_headline_variants_meta":{"raw":{"variants":["LLM word judgments shift with prompted language identity","Dutch vs Chinese prompts rewire LLM psycholinguistic ratings","Language prompts alter how LLMs rate word sounds and valence","LLMs act human-like when told to speak Dutch or Chinese","Prompted language changes LLM internal encoding of word meaning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1425,"prompt_tokens":881,"completion_tokens":544,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":463}},"tokens_in":497,"tokens_out":544,"duration_ms":7169,"temperature":1.0,"reasoning_tokens":463,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:56:54.070902+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect native-speaker valence and sound-shape norms from Dutch and Mandarin participants for the same 648 pseudowords and 1,034 translated words, and compare the models' Dutch- and Chinese-prompted judgments against those native norms. If Dutch-prompted outputs align better with English norms than with Dutch human ratings, the claim that language identity changes psycholinguistic cognition in the model would be weakened.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the 648 pseudowords with round/spiky shape labels that form Task 1's data and English reference labels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the English word set and mean valence ratings used as Task 2's English-derived ground truth after translation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the Dutch/Chinese comparison and supplies the cross-linguistic valence-from-phonology findings the tasks build on."}],"review_version":1}