{"id":"833f454f-a5a2-4cd9-ab90-76c1ef22a85c","arxiv_id":"2505.06062","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Fine-tuning BERT-based models on syntactic versus semantic tasks changes the layer-wise attention they pay to idioms and microsyntactic units across six languages.","lead":"This paper measures how BERT-style language models change their attention to idioms and to microsyntactic units after fine-tuning on language tasks. It finds that task type shifts where attention peaks, a step toward understanding what fine-tuning does inside multilingual models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No statistical support is reported for the abstract's 'significantly influences': all central claims are read off averaged attention curves, so the main empirical conclusion is underdetermined.","rationale":"I considered the layer-to-function assumption highlighted by the reader. It is real, but it only affects the explanatory gloss 'corresponding with syntactic processing requirements'; the raw attention shifts would remain. The more load-bearing issue is whether those shifts are real at all. The paper reports no dispersion or significance for any attention comparison, and the plotted differences in Figure 3 are small relative to plausible sampling noise. The POS row in Table 3(b) provides internal evidence that the 'syntactic tasks' category is not coherent for the MSU claim, reinforcing the need for per-task quantitative analysis. A paired bootstrap/permutation analysis over the 227 MWE contexts is a minimal, feasible check that would settle whether the abstract's 'significantly influences' is warranted. This does not change the reader's CONDITIONAL verdict; it sharpens the condition that should be met before the central claim is accepted.","tokens_in":11116,"tokens_out":7567,"duration_ms":78778,"concrete_test":"For each language, MWE type, layer, and fine-tuned model, recompute the mean attention difference versus the pretrained model using paired bootstrap over the 227 MWE contexts (or available n). Report 95% bootstrap confidence intervals and run paired permutation tests (e.g., 10,000 permutations) for the two key claims: (i) syntactic fine-tuning increases mean attention to MSUs in lower layers (3–4); (ii) semantic fine-tuning makes idiom attention flatter across layers (e.g., variance of layer means decreases). Apply multiple-comparison correction across layers, tasks, and languages. If the lower-layer MSU increase is not significant after correction, or is driven only by DepRel/Slavic, the abstract's 'significantly influences' and the blanket syntactic-task statement must be weakened accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is that fine-tuning 'significantly influences' attention to MWEs, with syntactic tasks increasing lower-layer attention to MSUs and semantic tasks flattening idiom attention (abstract; Section 5). The evidence for this is entirely visual: conclusions in Section 5 are drawn from layer-wise means plotted in Figures 1–3, with no confidence intervals, per-instance variance, or significance tests. With 24 layers, two MWE types, four fine-tuned tasks, and six languages, the number of implicit comparisons is large; differences of a few percentage points (e.g., Figure 3, where most Russian changes are within roughly +/-2–5%) could easily be sampling noise over 227 MWE contexts. This is load-bearing because the paper's headline is a claim about how fine-tuning changes attention, not merely a description of these particular curves. The internal data also hint that the effect is not uniform: in Table 3(b), Russian MSUs fine-tuned on POS have top attention at layers 11 and 12, not lower layers, so the abstract's blanket 'syntactic tasks ... lower layers' claim is already contradicted by one of the two syntactic tasks in one language. Without a quantitative test of the average differences, the central conclusion is underdetermined.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript examines whether fine-tuning BERT-based models on syntactic (DepRel, POS) and semantic (NER, Topic) tasks changes their attention to two types of multiword expressions across six Indo-European languages. Attention weights are extracted from 24-layer monolingual BERT models, averaged over heads, and summarized as layer-wise percentages of attention from context to MWEs and within MWEs. The authors report that fine-tuning significantly changes attention: semantic fine-tuning distributes idiom attention more evenly across layers, while syntactic fine-tuning increases attention to microsyntactic units in lower layers, with Russian identified as an exception. The datasets and fine-tuned models are released.","tokens_in":11359,"tokens_out":6430,"duration_ms":63310,"significance":"If the main findings were supported by quantitative evidence, the paper would provide a useful multilingual contribution to interpretability of MWE processing in transformers, extending Jang et al. (2024) to idioms versus microsyntactic units. The release of datasets and fine-tuned models, the use of six languages and two MWE types, and the distinction between context-to-MWE and within-MWE attention are strengths. However, the current evidence is primarily visual, the abstract's 'significantly' is not backed by statistical tests, the Russian results in Figure 3 and Table 3 already contradict the blanket claims, and the interpretation relies on an unverified mapping from attention peaks to syntactic processing. The contribution is therefore potentially valuable but not yet established.","major_comments":[{"comment":"The central claim that fine-tuning 'significantly influences' attention to MWEs is supported only by visual inspection of layer-wise averages. No significance tests, confidence intervals, or variation across random seeds are reported, and with 24 layers, two MWE types, four fine-tuned tasks, and six languages the number of implicit comparisons is large. Differences of a few percentage points (e.g., Figure 3) could arise from sampling variation over 227 contexts. A quantitative analysis—for example, paired per-layer tests with multiple-comparison correction, bootstrap confidence intervals, and effect sizes—is needed before the headline claim can be accepted.","section":"Abstract; Section 5; Figures 1–3"},{"comment":"The abstract and Conclusion state that models fine-tuned on syntactic tasks show increased attention to MSUs in lower layers, but the paper's own Russian results contradict this. Section 5.2 reports that for Russian MSUs the POS task increases attention in middle and upper layers and that DepRel does not show the same pattern; Table 3(b) lists layers 11 and 12 among the top three for Russian MSUs under POS, and Figure 3 shows mostly decreased attention in Russian after fine-tuning. The claims should be revised to language- and task-specific statements, or the generality of the pattern should be demonstrated statistically.","section":"Section 5.2; Figure 3; Table 3(b); Conclusion"},{"comment":"The interpretation of lower-layer attention peaks as evidence of 'syntactic processing requirements' rests on the premise that lower BERT layers encode syntax and higher layers encode semantics, cited to Tenney et al. (2019). That work uses linear probes on contextual representations, not attention weights, and no analysis here links attention to representations or to task performance. Without such evidence, the layer-to-function mapping is an unsupported interpretive assumption; at minimum it should be labeled as a hypothesis, and ideally tested, e.g., by comparing layer-wise attention peaks with layer-wise probing accuracy for the same MWE instances.","section":"Section 5.2; Section 2 (Tenney et al.)"},{"comment":"The cross-linguistic comparisons in Section 5.3 are confounded with dataset source and model architecture. The MSU dataset covers only Slavic languages, and the idiom datasets come from different sources per language (ID10M for EN/DE/NL/PL, a Russian idiom corpus plus dictionary for RU, and ChatGPT-generated data for UK). Differences between Germanic and Slavic attention patterns could therefore reflect dataset composition, annotation guidelines, or model choice rather than language properties. The Limitations section acknowledges the sources differ, but Section 5.3 still attributes the observed differences to morphological complexity; this attribution should be removed or supported with a matched analysis.","section":"Section 3.1; Section 5.3"}],"minor_comments":[{"comment":"The sentence 'Fine-tuning on syntactic tasks (Topic and DepRel)...' misclassifies Topic as syntactic, although Section 3.3.2 defines Topic as a semantic task; this should be corrected to avoid confusing the reported pattern.","section":"Section 5.1"},{"comment":"The color coding of lower and middle layers is not reproducible in plain text; consider replacing colors with an unambiguous notation such as superscripts or labels.","section":"Table 3 caption"},{"comment":"The exact definition of 'attention percentage' is not fully specified: the denominator, the treatment of [CLS]/[SEP] tokens, and the aggregation of subword tokens should be stated explicitly so that the reported percentages are reproducible.","section":"Section 4.2–4.3"},{"comment":"The layer grouping thresholds 'lower 1–8, middle 9–16, upper 17–24' are introduced without justification and are used in the discussion; the authors should state whether the conclusions depend on this particular partition.","section":"Table 3 caption; Section 5.2"},{"comment":"The Figure 1 caption says results are shown for English and Ukrainian, while Section 5.1 says the figure shows PL, UK, and EN; the coverage should be reconciled.","section":"Figure 1 caption; Section 5.1"},{"comment":"The text says ID10M test sets were available for EN and DE, yet PL is used in the experiments; please clarify which PL idiom data were used and how they were validated.","section":"Section 3.1.2"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the public release of data and models is a positive feature. The main risk is statistical underdetermination; if the authors add quantitative tests, hedge the Russian exception, and temper the interpretive claims, the revision should be straightforward. The use of the authors' own MSU dataset is not a circularity concern because the attention measurements are independent of how that dataset was constructed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is a reasonable extension of Jang et al. (2024) to idioms and microsyntactic units across six languages, with new multilingual comparisons and released data/models. But the central claim—that fine-tuning significantly influences attention—is read off averaged attention curves without any significance testing, and the Russian results contradict the abstract's blanket statement. The paper deserves peer review, but only with a requirement to add quantitative support and qualify the exceptions.\n\nWhat's actually new: applying Jang's attention-extraction method to MWEs, contrasting idioms versus MSUs, and doing this consistently across EN/DE/NL/PL/RU/UK. The subword token aggregation is careful, and the datasets and fine-tuned models are publicly available. That is a useful contribution for anyone working on MWE processing or BERT interpretability.\n\nThe soft spots are real but mostly fixable. There are no confidence intervals, no per-instance variance, no tests across the many layer/task/language comparisons. The abstract says 'significantly influences' but the evidence is visual, and some differences are only a few percentage points. The Russian exception is acknowledged in the text, but it is exactly the kind of counterexample that the abstract's claim needs to be explicitly qualified for. Also, the idiom/MSU comparison is confounded: MSUs exist only for Slavic languages, while idioms come from heterogeneous sources (including ChatGPT for Ukrainian), so language family and dataset provenance are tangled. The reliance on 'lower layers = syntax' from Tenney et al. is an interpretive assumption; if that mapping does not hold in a given language or model, the 'syntactic processing requirements' framing is unsupported.\n\nOn balance, I would send this to peer review. The raw observations are probably real, the resources are valuable, and the flaws are addressable rather than fatal. I would ask for significance testing, per-language and per-task quantitative summaries, and a more careful treatment of the Russian deviations.\n\nReading group: only if someone in your group works on MWE interpretability; otherwise it is skimmable. I would not cite it for the claims, but I might cite the dataset if I worked in this niche.","headline":"A legitimate multilingual attention study whose headline claim outruns the evidence: the 'significantly' is not statistically backed, but the raw patterns and released resources are worth a serious look.","tokens_in":11899,"tokens_out":1720,"would_cite":false,"duration_ms":17575,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning steers BERT's attention to multiword expressions: syntax tasks focus lower layers on microsyntactic units, semantic tasks spread idiom attention evenly across layers.","keywords":["BERT attention","multiword expressions","idioms","microsyntactic units","fine-tuning","layer-wise attention","multilingual NLP","syntactic versus semantic tasks"],"falsifier":"Take one language and model, fine-tune several random seeds on a syntactic task and a semantic task, and measure attention to idioms and microsyntactic units layer by layer. The claim predicts that syntactic fine-tuning raises lower-layer attention to microsyntactic units and semantic fine-tuning flattens idiom attention; if either pattern fails to replicate across seeds, or if a second Russian-initialized model reproduces the Russian decrease rather than the cross-lingual pattern, the causal story about task type collapses.","tokens_in":10947,"feed_emoji":"🧠","tokens_out":6498,"duration_ms":59003,"temperature":0.7,"pith_summary":"The paper asks whether fine-tuning a BERT-style model on a supervised task changes where the model's attention lands on multiword expressions, and whether the direction of the change depends on whether the task is syntactic or semantic. It claims yes: fine-tuning on syntactic tasks (dependency-relation classification and part-of-speech tagging) raises attention to microsyntactic units in the lower layers, while fine-tuning on semantic tasks (named-entity recognition and topic classification) spreads attention to idioms more evenly across layers. The pattern is tested in six Indo-European languages with monolingual 24-layer BERT-based encoders. If the claim holds, attention is not a fixed residue of pretraining but a task-adaptable resource, and layer-wise attention profiles become a diagnostic for which linguistic demands a model has learned to meet. Russian is a reported exception, with attention to both idioms and MSUs mostly decreasing after fine-tuning.","feed_headline":"Fine-tuning moves BERT's attention into syntax or semantics","feed_subtitle":"Six languages show syntax tasks focusing lower layers on microsyntactic units while semantic tasks spread idiom attention.","key_machinery":"The load-bearing object is the layer-wise attention profile: for each of the model's 24 layers, attention matrices are averaged across heads and reduced to two numbers, the mean attention that context tokens pay to MWE tokens and the mean attention MWE tokens pay to each other. Subword tokens belonging to an MWE are aggregated so that each expression acts as one unit. The category contrast between idioms and microsyntactic units carries the argument, because the two are chosen to isolate semantic non-compositionality and syntactic unpredictability respectively. The interpretation step relies on a previously established mapping, cited by the paper, that lower BERT layers encode syntactic information and higher layers encode semantic information; attention peaks in lower layers are therefore read as a sign of syntactic processing requirements.","core_discovery":"On the paper's own terms, the discovery is that the division between semantic and syntactic fine-tuning leaves a measurable trace in BERT's attention. For idioms, semantic tasks make the model distribute attention across more layers instead of concentrating it, consistent with the idea that a non-compositional expression has to be assembled from distributed semantic cues. For microsyntactic units, syntactic tasks push attention toward the lower layers where the model is presumed to do syntactic work, consistent with the idea that these expressions demand non-standard grammatical processing. The authors report the same broad contrast in Germanic and Slavic languages, with Russian as a consistent exception where fine-tuning mostly lowers attention to both categories.","pith_inferences":["If lower-layer attention to microsyntactic units is genuinely tied to syntactic irregularity, then fine-tuning on dependency relations or part-of-speech tagging should measurably improve a downstream parser's accuracy on sentences containing these units; the paper does not test that, but it follows directly and is testable.","The Russian anomaly may have a model- or tokenizer-level cause rather than a language-level one; comparing a second Russian-initialized architecture on the same tasks would separate those possibilities.","An even layer distribution of idiom attention under semantic tasks could be an attention-entropy increase rather than targeted semantic integration; re-analyzing the same data with entropy or headwise measures would tell which.","The same method could be applied to other multiword-expression families, such as collocations or light-verb constructions, predicting an attention profile intermediate between idioms and microsyntactic units."],"forward_implications":["Fine-tuning on a syntactic task can be used deliberately to sharpen a model's lower-layer attention to syntactically irregular expressions, and fine-tuning on a semantic task to broaden attention to non-compositional ones.","Layer-wise attention to multiword expressions can serve as a diagnostic for what a fine-tuned model has learned, alongside task accuracy.","The syntactic-versus-semantic distinction is a genuine axis of attention behavior, not just a surface difference between datasets.","Because Russian breaks the pattern, task type alone does not determine attention; language-specific models and datasets must be part of any account."],"supporting_citations":[{"why":"Defines the BERT architecture and attention mechanism that all models in the study are based on.","marker":"Devlin et al., 2019"},{"why":"Supplies the lower-layer-syntax and upper-layer-semantics mapping used to interpret attention peaks.","marker":"Tenney et al., 2019"},{"why":"Provides the attention-extraction and fine-tuning comparison methodology the paper extends to multiword expressions.","marker":"Jang et al., 2024"},{"why":"Supplies the Slavic microsyntactic-unit dataset with 227 Russian MSUs and parallel translations.","marker":"Zaitova et al., 2023"},{"why":"Supplies the ID10M idiom corpus used for English, German, Dutch, and Polish idiom contexts.","marker":"Tedeschi et al., 2022"},{"why":"Supplies the 85 Russian idioms used in the Russian idiom subset.","marker":"Aharodnik et al., 2018"},{"why":"Defines microsyntactic units as syntactically unpredictable expressions, the object class studied.","marker":"Iomdin, 2016"},{"why":"Provides the Universal Dependencies datasets used to fine-tune the syntactic tasks.","marker":"Nivre et al., 2020"}],"fun_headline_variants":["Fine-tuning steers BERT attention to syntax or semantics","Idioms spread BERT attention; microsyntax sinks it low","Semantic tasks widen BERT's idiom focus; syntax drives microsyntax deep","Task type reroutes BERT attention: idioms diffuse, microsyntax concentrates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that lower BERT layers encode syntax and higher layers encode semantics; if that mapping does not hold for a given model or language, the interpretation of attention peaks as syntactic or semantic processing loses its footing even if the raw attention differences are real.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuning steers BERT attention to syntax or semantics","Idioms spread BERT attention; microsyntax sinks it low","Semantic tasks widen BERT's idiom focus; syntax drives microsyntax deep","Task type reroutes BERT attention: idioms diffuse, microsyntax concentrates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000356,"raw_usage":{"total_tokens":1896,"prompt_tokens":876,"completion_tokens":1020,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":941}},"tokens_in":492,"tokens_out":1020,"duration_ms":9754,"temperature":1.0,"reasoning_tokens":941,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:48:39.380138+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one language and model, fine-tune several random seeds on a syntactic task and a semantic task, and measure attention to idioms and microsyntactic units layer by layer. The claim predicts that syntactic fine-tuning raises lower-layer attention to microsyntactic units and semantic fine-tuning flattens idiom attention; if either pattern fails to replicate across seeds, or if a second Russian-initialized model reproduces the Russian decrease rather than the cross-lingual pattern, the causal story about task type collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the attention-extraction and fine-tuning comparison methodology the paper extends to multiword expressions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Slavic microsyntactic-unit dataset with 227 Russian MSUs and parallel translations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 85 Russian idioms used in the Russian idiom subset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines microsyntactic units as syntactically unpredictable expressions, the object class studied."}],"review_version":1}