{"id":"79f61482-503d-4801-9916-0e4855e22803","arxiv_id":"2508.12803","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"Projecting away the estimated Modern Standard Arabic subspace during fine-tuning improves generation across 25 Arabic dialects by up to +4.9 chrF++, evidence that subspace dominance by a high-resource variety restricts related-variety capacity.","lead":"This study argues that when a dominant standard language such as Modern Standard Arabic dominates an LLM's internal representations, generation quality on related Arabic dialects suffers, and that projecting those representations away from the MSA subspace during fine-tuning improves dialect translation by up to +4.9 chrF++ at a cost in standard-language performance. The key deliverable is an online variational probing framework that estimates and removes the dominant variety","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The supplied full text is an unrelated cs.LO survey, so the abstract's causal claim about MSA-subspace dominance and its reported improvements are unsupported; the central claim cannot be verified from the submitted materials.","rationale":"The reader's verdict is UNVERDICTED with LOW confidence, primarily because the supplied full text is unrelated to the abstract. My pass agrees with that core problem. The reader's stated weakest assumption, however, is about a possible confound in the variational probe (e.g., MSA-dialect overlap removing dialect-relevant features). That is a legitimate scientific concern, but it is secondary here: before evaluating whether the probe is confounded, there must be an actual described and testable probe. The load-bearing issue is that the submission contains no evidence for the existence of the claimed experiments, let alone their correctness. 'UNCHANGED' is appropriate because the correct verdict remains UNVERDICTED: the materials supplied do not permit assessment of the central claim.","tokens_in":28116,"tokens_out":1964,"duration_ms":22665,"concrete_test":"Obtain the complete arXiv:2508.12803 full text and confirm that it contains an experimental section describing the variational probing framework, the 25-dialect dataset, the fine-tuning procedure, and the chrF++ evaluation. Then, using the released code, reproduce the reported +4.9 chrF++ improvement on the stated dialect test set. If the full text is only the cs.LO survey or lacks the described experimental apparatus, the central causal claim is unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims a causal result: decoupling representations from the estimated MSA subspace during fine-tuning improves dialectal generation quality by up to +4.9 chrF++ and +2.0 on average across 25 dialects, with a measured tradeoff on standard Arabic. These claims depend on an 'online variational probing framework' that continuously estimates the standard-variety subspace and enables projection-based decoupling. The full text supplied with the submission is an unrelated survey chapter on Craig interpolation and separation of formal languages (arXiv:2508.12805v2, cs.LO), containing no NLP methods, datasets, experimental tables, code, or even a description of the variational probe. Treating the full text as in-scope evidence, the central argument is missing its evidential backbone: there is no way to check whether the probe correctly identifies MSA-dominance, whether orthogonal projection removes only harmful shared structure, whether the reported gains are significant or confounded, or whether the proxy claim about dialectal MT as a controlled generative proxy is even operationalized. This is not a dispute with consensus; it is a missing-support problem. The mismatch itself is the load-bearing concern: either the wrong manuscript was attached, or the claimed study is not present. In either case, the core causal claim is unverifiable from this submission.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submitted abstract (arXiv:2508.12803, cs.CL) claims a causal study of representational entanglement in multilingual LLMs: an online variational probing framework estimates the Modern Standard Arabic (MSA) subspace during fine-tuning, and projection-based decoupling from that subspace improves dialectal generation quality by up to +4.9 chrF++ and +2.0 on average across 25 dialects, at the cost of a measured tradeoff in standard-language performance. The full text attached to the submission, however, is an unrelated cs.LO survey chapter on Craig interpolation and separation of formal languages (arXiv:2508.12805v2). It contains no NLP methodology, no description of the variational probe, no datasets, no experimental tables, and no results. The manuscript's central claims are therefore entirely unsupported in the submitted materials.","tokens_in":28308,"tokens_out":2057,"duration_ms":27172,"significance":"If the claimed results were backed by a full experimental study, the paper would address a timely and practically important question: whether alignment with a high-resource standard variety can actively harm modeling of related low-resource varieties. The proposed online variational probing framework and the reported positive gains on 25 Arabic dialects would constitute a substantive contribution, and the explicit reporting of a tradeoff in standard-language performance is a sign of good-faith empirical reporting. However, none of this supporting evidence is present. There is no machine-checked proof, no reproducible code (despite the abstract's promise), no parameter-free derivation, and no experimental protocol. As submitted, the paper is only an abstract with a different manuscript attached. The central causal claim and the framework are unverifiable, so the scientific contribution cannot be assessed.","major_comments":[{"comment":"The full text supplied with the submission is not the paper described in the abstract. It is a survey on Craig interpolation and formal-language separation by different authors, with no connection to multilingual LLMs, Arabic dialects, probing, or fine-tuning. This is a load-bearing mismatch: every claimed result in the abstract—the online variational probe, the projection-based decoupling, the 25-dialect evaluation, the +4.9 chrF++ improvement, and the standard-language tradeoff—is unsupported by the attached manuscript. The submission cannot be evaluated as a research paper until the correct full text is provided.","section":"Full Text (all sections)"},{"comment":"The abstract reports precise quantitative results ('up to +4.9 chrF++ and +2.0 on average', 'a measured tradeoff') without any error bars, per-dialect breakdown, model sizes, training data, or evaluation protocol. Even with the correct full text, these numbers would need to be accompanied by variance estimates and significance testing to support the causal interpretation. As submitted, the numbers have no methodological context and cannot be checked.","section":"Abstract"},{"comment":"The abstract's central causal claim is that 'subspace dominance by high-resource varieties can restrict generative capacity.' This requires that the estimated MSA subspace captures dominance-related structure and that orthogonal projection away from it removes only harmful shared structure while preserving dialect-specific information. Because MSA and Arabic dialects overlap heavily in vocabulary and morphology, the probe subspace could be confounded with dialect-relevant features. The manuscript provides no analysis, control experiment, or ablation to rule out this confound; in fact, it provides no experimental content at all.","section":"Abstract (causal claim)"},{"comment":"The claim that 'dialectal MT serves as a controlled proxy for generative tasks where comparable multi-variety corpora are unavailable' is not operationalized. No argument or evidence is given that the proposed intervention transfers beyond the Arabic MT setting, and no comparison to other generative tasks or language families appears anywhere in the submission.","section":"Abstract (methodological proxy claim)"}],"minor_comments":[{"comment":"The abstract states 'Code will be released,' but no code, repository link, or supplement is included. If the full text is resubmitted, this promise should be fulfilled or clarified.","section":"Abstract"},{"comment":"The arXiv identifier cited for the full text (2508.12805v2, cs.LO) differs from the submission's identifier (2508.12803, cs.CL). This suggests a packaging error that should be corrected before resubmission.","section":"Metadata"},{"comment":"The term 'causal evidence' is used loosely. The described intervention is a geometric manipulation during fine-tuning; without a formal causal identification argument or adequate controls, 'evidence for a causal role' would be more appropriate.","section":"Abstract"}],"recommendation":"reject","confidential_remarks":"To the editor: the decisive issue is the mismatch between the abstract and the attached full text. The full text is an unrelated cs.LO survey, so the submission does not contain the claimed study. This could be a submission error, but as it stands the paper is not assessable. Should a corrected version be resubmitted, the reviewers would need to evaluate the actual experiments, the variational probe's identifiability, and the confound of MSA-subspace overlap with dialect-relevant features. I recommend rejection of the current submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The abstract points at a real and useful idea: estimate the MSA-dominated subspace online during fine-tuning and project it out, to free up representation capacity for Arabic dialects. If the reported numbers hold (up to +4.9 chrF++, +2.0 average across 25 dialects), this is a practical intervention that costs no extra data, and the framing as a control on representational allocation is genuinely interesting. The explicit tradeoff against standard Arabic is also a plus; it makes the causal claim falsifiable in principle.\n\nBut the submission as we received it is broken. The full text attached is an unrelated survey chapter on Craig interpolation and separation of formal languages, by different authors, with a different arXiv ID. There are no methods, no tables, no code, no evaluation protocol, nothing about the variational probe. The abstract's numbers are unverifiable from this bundle. That is the load-bearing problem, not a dispute about consensus or a theoretical objection.\n\nThe soft spots beyond the mismatch are the ones you would expect from the abstract alone. No error bars or per-dialect breakdown; no model sizes or training details. More substantively, the causal interpretation depends on the probe isolating the MSA-dominance subspace without also removing dialect-relevant shared structure. Given how much vocabulary and morphology Arabic dialects share with MSA, that is a real assumption to check. The abstract's proxy claim about dialectal MT as a controlled proxy for other generative tasks is plausible but not defended.\n\nThat said, none of these concerns constitute a verdict that the underlying work is wrong. They are all missing-support problems. The right response is to get the correct full text and then send it out. If the paper matches the abstract, it deserves serious referee time and probably would be a useful contribution to the multilingual and low-resource MT communities. If the authors cannot produce the matching manuscript, desk reject.\n\nBottom line: worth chasing down, but not citable from this submission.","headline":"The abstract describes a plausible and potentially useful intervention for dialectal MT, but the supplied full text is a different cs.LO paper, so the empirical claims are unverifiable from this submission.","tokens_in":28948,"tokens_out":2014,"would_cite":false,"duration_ms":23618,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that excessive entanglement with the high-resource standard variety actively harms generation for related low-resource varieties, and that projection-based decoupling from the estimated MSA subspace during fine-tuning impro","keywords":["multilingual models","representation geometry","subspace dominance","online variational probing","Arabic dialects","machine translation","fine-tuning","causal intervention"],"falsifier":"Run the same decoupling intervention with a random or mis-specified projection subspace of the same dimensionality; if the +2.0 chrF++ gain persists unchanged, the improvement is not attributable to removing MSA dominance. Alternatively, inspect the learned probe subspace: if it assigns high probability to dialect tokens that share vocabulary with MSA, the removed subspace is partly dialect-relevant, and the gains could come from removing shared information rather than a true dominance effect.","tokens_in":27883,"feed_emoji":"🌐","tokens_out":5459,"duration_ms":56706,"temperature":0.7,"pith_summary":"The paper sets out to prove that too much representational alignment with a dominant high-resource variety—Modern Standard Arabic (MSA) relative to Arabic dialects—actively harms generative modeling of the related varieties. It introduces an online variational probing framework that continuously estimates the MSA subspace during fine-tuning and enables projection-based decoupling from that subspace. Across 25 Arabic dialects, this intervention improves machine translation quality by up to +4.9 chrF++ and +2.0 on average compared with standard fine-tuning, at a measured cost to standard-language performance. The authors read these results as causal evidence that subspace dominance by high-resource varieties restricts generative capacity for related varieties, and they offer dialectal MT as a controlled proxy for other generative tasks with comparable multi-variety corpora.","feed_headline":"Decoupling from standard Arabic boosts dialect translation 4.9 points.","feed_subtitle":"Projecting away from the MSA subspace during fine-tuning gains +2.0 points across 25 dialects.","key_machinery":"The load-bearing object is the MSA-dominance subspace, estimated by an online variational probe during fine-tuning. The probe predicts whether a hidden representation comes from the standard variety; the direction that best separates MSA from dialects defines the subspace to be removed. The intervention projects hidden representations orthogonally away from that subspace before generation, stripping dominant-variety structure while leaving dialect-specific structure intact. The variational formulation makes the estimate continuous and trainable, so the probe tracks the representation space as it shifts during fine-tuning rather than relying on a static, pre-training estimate.","core_discovery":"The central claim is that entanglement of dialect representations with the MSA subspace is not benign: it actively suppresses dialect-specific generative capacity. The discovery combines a mechanism and a method. The mechanism is subspace dominance: because MSA is overwhelmingly represented in training, the learned representation space becomes dominated by a subspace encoding MSA-specific structure, and dialect tokens that pass through this subspace during generation inherit restrictions that hurt dialect output. The method is an online variational probing framework that estimates this subspace continuously during fine-tuning and then removes it by orthogonal projection. In controlled fine-t","pith_inferences":["We infer a natural testable extension: if subspace dominance is the mechanism, the gains from decoupling should be largest for dialects with high lexical and structural overlap with MSA, and smallest for distant varieties.","The results suggest a dynamic-probe requirement: a probe estimated only before fine-tuning might miss subspace drift during training, so the online variational formulation may be essential rather than merely convenient.","We infer that the same projection logic could apply to any high-resource 'hub' variety dominating its neighbors—for example, a standard language dominating its regional variants—with gains scaled by measured subspace dominance.","The tradeoff opens the possibility of interpolating between full alignment and full decoupling to set a desired balance between standard-variety performance and related-variety quality, though the paper does not test this interpolation."],"forward_implications":["If the claim is right, standard fine-tuning on mixed multi-variety data is not optimal for low-resource varieties; removing the dominant variety's subspace during fine-tuning yields consistent gains across 25 dialects.","The measured tradeoff in standard-language performance means there is a controllable knob between fidelity to the standard variety and quality on related varieties, not a free lunch.","Dialectal machine translation can serve as a controlled proxy for other generative tasks with comparable multi-variety corpora, so the same decoupling approach could transfer to other language families and domains.","Online variational probing during fine-tuning gives a practical way to identify and remove a dominant subspace, which could be extended to controlling representational allocation in multilingual and multi-domain LLMs."],"supporting_citations":[],"fun_headline_variants":["Decoupling from MSA gains up to +4.9 chrF++ on dialect generation","Subtracting MSA subspace boosts dialect translation by 4.9 chrF++","Alignment hurts: Separating dialects from MSA adds +4.9 chrF++","Decouple dialect space from Arabic for +4.9 chrF++ gain"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The causal story stands on the online variational probe correctly isolating the MSA-dominance subspace, so that projecting away from it removes mostly harmful shared structure and not dialect-specific information.","fun_headline_variants_meta":{"raw":{"variants":["Decoupling from MSA gains up to +4.9 chrF++ on dialect generation","Subtracting MSA subspace boosts dialect translation by 4.9 chrF++","Alignment hurts: Separating dialects from MSA adds +4.9 chrF++","Decouple dialect space from Arabic for +4.9 chrF++ gain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000988,"raw_usage":{"total_tokens":4031,"prompt_tokens":752,"completion_tokens":3279,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":3190}},"tokens_in":496,"tokens_out":3279,"duration_ms":28987,"temperature":1.0,"reasoning_tokens":3190,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:19:38.597331+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same decoupling intervention with a random or mis-specified projection subspace of the same dimensionality; if the +2.0 chrF++ gain persists unchanged, the improvement is not attributable to removing MSA dominance. Alternatively, inspect the learned probe subspace: if it assigns high probability to dialect tokens that share vocabulary with MSA, the removed subspace is partly dialect-relevant, and the gains could come from removing shared information rather than a true dominance effect.","supporting_citations":[],"review_version":1}