{"id":"25f54f04-d24c-44dc-a334-3f900eefa72f","arxiv_id":"2506.20489","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Language models and human brains encode story meaning in a shared conceptual space that generalizes across English, Chinese, and French.","lead":"This paper asked whether speakers of very different languages share the same neural representation of meaning. Using fMRI and language models for English, Chinese, and French, the authors found that brain-activity models trained on one language can predict brain responses to the same story in other languages.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The uBERT cross-language encoding transfer (Fig. 3C) requires a coordinate alignment between separately trained BERT embedding spaces, but no such alignment is described in the encoding-model Methods; without it the reported transfer is not expected to occur.","rationale":"The reader's concern about translation equivalence and low-level confound regression is valid and worth testing, but the missing uBERT alignment is a more immediate threat to the main claim. The paper's own Results explicitly state that unilingual BERT spaces are oriented in arbitrary directions, so cross-language transfer with the same encoding weights is not expected unless the spaces are aligned. The Methods describe Procrustes alignment only for the embedding-similarity comparison, not for the encoding models, leaving a gap between the claim and the reported procedure. This is a correctable but decisive detail: if the code confirms alignment, the result stands and the reader's confound concern remains the main caveat; if the code does not confirm alignment, the central generalization result would need to be reevaluated. The paper has real strengths—open data, code availability, held-out run evaluation, and multiple control analyses—so the appropriate disposition is conditional rather than outright rejection.","tokens_in":33776,"tokens_out":5490,"duration_ms":66737,"concrete_test":"Inspect the public code at github.com/zaidzada/crosslingual-convergence to confirm whether the uBERT cross-language encoding analysis (Fig. 3C) applies a Procrustes/linear alignment to bring Chinese and French unilingual BERT embeddings into the English embedding space before training/testing. If alignment is absent, rerun the analysis with and without Procrustes alignment; if the result disappears without alignment, the claim is unsupported. If alignment is present, check that the rotation is fit only on training folds (e.g., using the held-out run's test sentences or a separate sentence split) and not on the evaluation sentences, to rule out leakage.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central cross-language result (Fig. 3C, 'Unilingual models and brains converge') uses uBERT embeddings from separately trained, per-language BERT models. Each uBERT occupies an arbitrarily oriented embedding space, so a voxelwise weight vector W learned to map English embeddings to English BOLD cannot transfer to Chinese/French embeddings without first aligning those spaces. The paper describes a Procrustes rotation only for the sentence-embedding similarity analysis (Methods, 'Computing similarity between word embeddings'), not for the encoding-model pipeline (Methods, 'Voxelwise encoding models' / 'Evaluating encoding models across languages'). If no alignment is applied before training/testing, the cross-language generalization should be near zero a priori; if alignment is applied, the Methods omit a step essential to the claim, and the alignment's train/test separation must be verified. This is more load-bearing than the translation-equivalence assumption: even with perfect translations, the uBERT transfer requires a shared coordinate frame.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses the Le Petit Prince multilingual fMRI corpus (English, Chinese, and French listeners) together with unilingual BERT, multilingual BERT, and Whisper embeddings to ask whether a conceptual space is shared across languages, both in language models and in the brain. The authors report three main results: (1) sentence embeddings from separately trained unilingual BERT models can be aligned by a learned rotation and show above-baseline cross-language similarity, peaking in middle layers; (2) voxelwise encoding models trained on one language's LM embeddings and brain responses generalize to brain responses of listeners of the other two languages, especially in high-level language and default-mode regions; and (3) in a multilingual model, embeddings from languages more similar to a listener's native language better predict that listener's brain activity, and Whisper speech embeddings contain cross-language phoneme information. The discussion interprets these findings as evidence for a partially shared conceptual representation across languages and across human brains and language models.","tokens_in":33942,"tokens_out":7125,"duration_ms":83939,"significance":"If the central claims hold, the paper would be a valuable contribution to the cross-language neuroscience of language: it uses a naturalistic, three-language fMRI dataset, proposes a concrete encoding-model framework for testing cross-language generalization, and includes a genuine held-out Procrustes alignment for embedding comparison. The phoneme-probing analysis is a true cross-language test, and the mBERT zero-shot generalization across 58 languages is a novel and falsifiable prediction. The paper also ships its code and uses an open dataset, which aids reproducibility. However, the reported unilingual embedding similarity is quite small (average r = 0.115), and the exact procedure for the central uBERT cross-language encoding transfer is ambiguous, which bears directly on the strength of the headline claim. These issues are addressable but currently leave the main conclusion less crisply supported than the abstract suggests.","major_comments":[{"comment":"The central uBERT cross-language result (Fig. 3C) is ambiguous about what is actually fed to the trained encoding weights at test time. The text says 'we evaluated predictions from the English model' against French and Chinese BOLD, which suggests the predictions are generated from English uBERT embeddings of the held-out English sentences, not from French or Chinese uBERT embeddings. If that is the case, no Procrustes alignment of separately trained uBERT spaces is needed, but then the claim should be stated as 'English LM features predict French/Chinese BOLD,' not as 'encoding models trained on one language generalize to another language.' If instead the weights were applied to French/Chinese uBERT embeddings, then the Methods omit the required rotation alignment for the encoding pipeline, and the train/test separation of that alignment would need to be verified. Because the abstract and Fig. 3C hinge on this point, please state explicitly which embeddings were used to generate the cross-language predictions and, if alignment was used, describe it.","section":"Methods, 'Evaluating encoding models across languages'"},{"comment":"The evidence for the claim that unilingual LMs 'converge on a similar embedding space' rests on a reported average test correlation of r = 0.115 across layers and language pairs after subtracting an untrained baseline. This is a small effect, and the paper does not report confidence intervals, permutation tests, or the distribution across the 825 held-out sentences and 768 dimensions. Since this is one of the two pillars of the title claim, please provide inferential statistics and effect sizes that establish that the held-out similarity is reliably above baseline, rather than only reporting the mean.","section":"Results, 'Unilingual models and brains converge'"},{"comment":"The difference map between within- and across-language encoding performance (Fig. S2) is used to argue that the across-language effect is only slightly weaker than the within-language effect. However, the whole-brain correlation between within- and across-language unthresholded maps (r = 0.974) is reported without any uncertainty or test of whether the small differences are reliable. More importantly, the evaluation against group-averaged BOLD from a different scanner (French participants) may introduce systematic noise; the authors do not discuss how scanner differences could affect the cross-language generalization magnitudes. Please address this explicitly, for example by reporting the correlation separately for each language pair and scanner.","section":"Results, Fig. 3C and Fig. S2"}],"minor_comments":[{"comment":"The caption states 'pFDR > .05' for thresholded encoding maps; this should presumably be 'pFDR < .05.'","section":"Fig. 3 caption"},{"comment":"The paper acknowledges the use of a single children's book and professionally translated audiobooks as limitations, which is appropriate; please also mention that the sentence-level averaging to 1,649 points limits temporal resolution and that the cross-language comparison is therefore at a coarse timescale.","section":"Discussion, limitations"},{"comment":"The GPT-4o translation into 55 languages is described, but no check of translation quality beyond manual inspection of samples is reported; a sentence-level back-translation consistency measure would strengthen the 58-language analysis in Fig. 4.","section":"Methods, 'Translating the story'"},{"comment":"The correlations between language-family closeness and encoding performance (r = 0.786, r = 0.869) are reported with p-values but without correction for the multiple language-family comparisons or for the fact that the same subjects and voxel mask are reused across the 58 languages.","section":"Results, Fig. 4C"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong candidate for a journals in cognitive neuroscience if the ambiguity about the uBERT transfer is resolved. My reading of the Methods is that the cross-language uBERT evaluation likely used source-language (e.g., English) embeddings to generate predictions, which avoids the alignment problem but weakens the claim that 'LMs trained on different languages' are doing the transfer. The mBERT 58-language analysis is the cleaner test of model-side convergence and could be foregrounded. I would ask for a clear statement of the test-time feature pipeline and for inferential statistics on the r = 0.115 embedding similarity before considering acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know that the paper's headline result—encoding models trained on one language's BERT embeddings generalizing to another language's brain responses—has a potentially serious methods gap. The Methods describe a Procrustes rotation only for the sentence-similarity analysis, not for the voxelwise encoding pipeline. The three unilingual BERT models live in arbitrarily oriented embedding spaces, so a linear regression trained on English embeddings (and their axes) cannot be applied directly to French or Chinese embeddings without first aligning those spaces. As written, the cross-language transfer should be near zero. Either the authors aligned the spaces and omitted that step, or the reported generalization is an artifact. This is more load-bearing than the translation-equivalence caveat, because even perfect translations leave the coordinate-frame problem unsolved.\n\nThe paper does several things well. The 58-language family-structure analysis is genuinely new and nicely executed, and the Whisper phoneme probing is a true cross-language test. The cross-validation uses held-out runs, and regressing out word rate, syllable rate, and acoustic confounds is careful. The writing is clear, and the authors are honest about the single-story limitation and Indo-European bias. Code and data are available, which is a real plus.\n\nThe mBERT and Whisper results are on firmer ground because those embeddings come from a single model, so no rotation is needed. The low embedding similarity (r = 0.115) and the assumption that professionally translated audiobooks are meaning-equivalent are real but minor by comparison. The pFDR typo in Fig. 3's caption is trivial.\n\nMy bottom line: this deserves a serious referee, but the first question must be about the missing alignment. If the authors can clarify or correct the encoding pipeline, the paper is a solid contribution. As it stands, I would not cite the uBERT transfer claim without checking the code. Send it to review, but flag the rotation issue as a major revision, not a copy-edit.","headline":"The cross-language uBERT encoding transfer as written is missing the embedding-space alignment it logically requires, which is a bigger problem than the translation-equivalence concern.","tokens_in":34463,"tokens_out":3615,"would_cite":false,"duration_ms":41708,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Encoding models for one language predict brain responses to the same story in two other languages, implying a shared conceptual space across English, Chinese, and French.","keywords":["cross-language neural representation","voxelwise encoding models","multilingual language models","naturalistic fMRI","shared conceptual space","narrative comprehension","intersubject correlation","default-mode network"],"falsifier":"A concrete test would use the same three languages with the story's sentences translated to carry different meanings while matching low-level acoustics; if encoding models trained on one language still predict the other-language brains as well as they do for the original story, cross-language generalization would not be evidence for a shared conceptual space.","tokens_in":33562,"feed_emoji":"🧠","tokens_out":8207,"duration_ms":79279,"temperature":0.7,"pith_summary":"This paper asks whether the neural representation of story meaning is shared across speakers of different languages, rather than merely occupying overlapping brain regions. Using fMRI recordings of native listeners who heard the same audiobook in English, Chinese, or French, the authors trained voxelwise encoding models that map language-model word embeddings onto each listener's brain activity. The central result is that a model trained on one language generalizes to listeners hearing the same story in another language, with significant predictive performance in high-level language and default-mode regions. The paper also shows that separately trained unilingual language models converge on a common embedding geometry, strongest in middle layers, and that languages closer to a listener's native language predict that listener's brain activity better. If correct, these findings indicate that speakers of different languages share a partly language-agnostic conceptual space, and that language models trained on different languages discover the same structure.","feed_headline":"Brain model for one language predicts listeners of two others","feed_subtitle":"English-trained brain models predict Chinese and French listeners of the same story, a sign of shared meaning.","key_machinery":"The central machinery is the voxelwise encoding model, a ridge-regression model that predicts each brain voxel's BOLD time series from contextual word embeddings extracted from a language model. To compare across languages, the authors downsample the actual and predicted BOLD responses to the sentence level and regress out low-level confounds including word rate, syllable rate, acoustic RMS energy, onset strength, framewise displacement, and TR count. For unilingual models, embedding spaces are aligned with a Procrustes rotation learned on half the sentences and evaluated on the other half; multilingual models are compared directly because they share a single embedding space. Whisper supplies a second stream of speech embeddings from its encoder and word embeddings from its decoder, allowing the paper to separate shared speech features from shared word-level meaning.","core_discovery":"The paper reports that language models trained on different languages converge onto a similar embedding space, especially in the middle layers, and that this shared geometry can be used to predict neural activity across language groups. In the brain analysis, a voxelwise encoding model trained on English embeddings and English listeners' BOLD responses predicts the BOLD responses of Chinese and French listeners hearing the same story, and the same holds for every language pair. The generalization is strongest in high-level language areas and default-mode regions, and is weaker in early auditory cortex and superior temporal gyrus, where language-specific speech processing likely dominates. The paper concludes that the neural representation of meaning is at least partly shared across speakers of different languages, and that language models trained on separate corpora converge on this shared meaning.","pith_inferences":["If the shared space is driven by plot and event structure rather than word-level semantics, cross-language prediction should survive within-sentence word-order scrambling but degrade when sentence order is shuffled; this is testable with the same corpus.","The same encoding framework could be turned into a translation-evaluation metric: candidate translations that better predict a target-language listener's brain activity would score higher on neural naturalness.","The 58-language analysis is biased toward Indo-European languages, so the family-tree gradient is most reliable for Germanic and Romance languages; extending to Sinitic, Dravidian, and Turkic families with native corpora is the direct next test.","Untrained or randomly initialized language models should not reproduce the cross-language generalization, making the trained-versus-untrained contrast a formal benchmark for any future claim of shared conceptual structure."],"forward_implications":["An encoding model trained on one language can be applied, without retraining, to predict where and when another language's listeners engage with the same narrative.","Unilingual language models trained on separate corpora contain a common geometric core, strongest in middle layers, and the degree of this geometric convergence predicts how well embeddings from one language transfer to another language's brain activity.","Brain responses to a story index linguistic relatedness: embeddings from languages closer to a listener's native language predict that listener's brain activity better, so neural data can serve as a behavioral measure of perceived language similarity.","Shared speech features across languages are present in Whisper's late encoder layers, and phoneme classifiers trained on one language's speech embeddings identify shared phonemes in other languages' speech.","The findings imply that high-level language and default-mode regions encode narrative meaning in a form that is partly independent of the particular sounds, scripts, and syntax of a given language."],"supporting_citations":[{"why":"Supplies the Le Petit Prince multilingual fMRI corpus: three groups of native listeners, the audiobook audio, and word-level transcripts in English, Chinese, and French; all neural analyses depend on this dataset.","marker":"Li et al., 2022"},{"why":"Provides the BERT architecture and the unilingual and multilingual pretrained models whose contextual embeddings drive the embedding convergence and encoding analyses.","marker":"Devlin et al., 2018"},{"why":"Establishes the prior demonstration of cross-language neural alignment that this paper extends from intersubject correlation to explicit computational encoding models.","marker":"Honey et al., 2012"},{"why":"Shows that language-model embeddings align with human brain activity in naturalistic listening, motivating the voxelwise encoding approach used here.","marker":"Goldstein et al., 2022"},{"why":"Provides the layer-wise interpretability result that multilingual conceptual features concentrate in middle layers, which the paper's layer analyses corroborate and rely on.","marker":"Lindsey et al., 2025"},{"why":"Defines the voxelwise encoding model framework used to map embeddings onto BOLD responses.","marker":"Naselaris et al., 2011"},{"why":"Introduces intersubject correlation as a way to isolate shared stimulus-driven neural responses, the evaluation target for the cross-language predictions.","marker":"Hasson et al., 2004"},{"why":"Shows that embedding spaces from different languages can be aligned by a rotation, providing the basis for the Procrustes alignment of unilingual models.","marker":"Mikolov et al., 2013"},{"why":"Supplies the Whisper speech-to-text model whose encoder and decoder embeddings drive the shared-speech and shared-word analyses.","marker":"Radford et al., 2022"}],"fun_headline_variants":["Cross-language brain activity shares a common meaning map","AI models reveal shared conceptual space across languages","One language's brain model predicts another's story listening","Shared meaning in brains emerges across English, Chinese, French","Language models and brains align on universal meaning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The three professionally translated audiobook versions are meaning-equivalent at the sentence level, and regressing out word rate, syllable rate, acoustic RMS, onset strength, framewise displacement, and TR count removes all language-specific low-level cues that could otherwise explain cross-language prediction.","fun_headline_variants_meta":{"raw":{"variants":["Cross-language brain activity shares a common meaning map","AI models reveal shared conceptual space across languages","One language's brain model predicts another's story listening","Shared meaning in brains emerges across English, Chinese, French","Language models and brains align on universal meaning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000258,"raw_usage":{"total_tokens":1575,"prompt_tokens":930,"completion_tokens":645,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":573}},"tokens_in":546,"tokens_out":645,"duration_ms":7923,"temperature":1.0,"reasoning_tokens":573,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:46:57.779583+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would use the same three languages with the story's sentences translated to carry different meanings while matching low-level acoustics; if encoding models trained on one language still predict the other-language brains as well as they do for the original story, cross-language generalization would not be evidence for a shared conceptual space.","supporting_citations":[],"review_version":1}