{"id":"4372373a-29bf-4f89-a68d-a294228e3cfa","arxiv_id":"2607.23545","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Instruction-hierarchy compliance in LLMs is asymmetric by language and position, and cross-language conflicts yield systematically higher compliance than same-language ones (Language Boundary Effect).","lead":"Multilingual LLMs do not treat instruction hierarchy the same way in every language: a language that helps when it is the boss can hurt when it should be ignored. The paper’s benchmark shows cross-language conflicts are easier to resolve than same-language ones, with real safety stakes for specialized models.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The per-language authority rankings (EN strongest, HI weakest) come from the three domains with language-sensitive metrics; in Safety — the only domain scored by language-invariant exact match — the HCR_H ranking reverses (HI top). Rankings may partly encode metric sensitivity.","rationale":"Read in good faith, this is a careful benchmark paper whose central empirical regularity — cross-language conflicts yield higher HCR than same-language ones — is unusually well supported for this domain: 15/15 language pairs positive in a narrow band, 12/13 models, and a covariate-adjusted item-level regression (Table 7) that survives language-position, model, setting, and domain fixed effects. The paper's own matrices also rebut the obvious alternative that LBE merely reflects weaker comprehension of lower-priority instructions: HI-HI sits *below* EN-EN on the diagonal, and cross pairs with an English (maximally potent) lower instruction still beat EN-EN. Promised public code/data and honest Limitations count in its favor. My concern lands on the second pillar — the per-language authority rankings and the \"implicit authority\" security claim — and it is the same locus the reader flagged (τ=0.15, English-only judge personas, translation pipeline). My addition is internal evidence that sharpens it from \"could confound\" to \"shows the expected signature\": the EN-authority pattern holds exactly in the three domains with language-sensitive scoring and reverses in Safety, the one domain scored by language-invariant exact match, with no mechanism offered for the exception. Because HCR normalizes by a Reference measured with the same language-sensitive instrument, and because binarization (τ, discrete judge labels) affects Ref and Conf nonlinearly, the ratio does not neutralize the confound. Even so, this would reshuffle language rankings and shrink the implicit-authority magnitudes rather than overturn the qualitative position-dependence or the LBE — which is why I keep the reader's CONDITIONAL rather than moving to REJECT. Acceptance should be gated on the metric-robustness sweep (and ideally the judge-language check), which is cheap given the released outputs and would either retire the confound or correctly re-scope the ranking claims.","tokens_in":48686,"tokens_out":13490,"duration_ms":1056142,"concrete_test":"Using the released Task-Execution outputs, recompute per-language HCR_H/HCR_L (Table 3) while sweeping τ over [0.05, 0.30], plus a threshold-free variant using mean chrF++ recall as the score. Decision criterion: if the ordering (EN highest, HI/ZH lowest HCR_H) is invariant across the sweep and matches the threshold-free variant, the threshold is innocuous and the asymmetry ranking survives in that domain; if HI/ZH HCR_H rises sharply as τ drops toward 0.05 — evidence their reference-condition translations cluster near 0.15 — the ranking is partly a threshold artifact and the Safety reversal is the cleaner signal. A parallel check would re-judge a stratified Persona sample with persona text in the response language.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The §4.3/§5.1/§5.2 claims — English is the strongest higher-priority authority, Hindi the most easily overridden lower-priority language, model-favored languages hold \"implicit authority\" — rest on per-language HCR differences of ~5–15pp. Three of four domains are scored by language-sensitive instruments: IFEval-derived rule checkers adapted from English (§B.2 patches punctuation, but sentence/paragraph segmentation for ZH/HI remains English-calibrated); a fixed chrF++-recall threshold τ=0.15 against Korean references (§B.3), under which genuine translations from lower-resource source languages plausibly sit closer to the boundary; and a GPT-5-mini judge that reads persona descriptions only in English while responses span six languages (§B.5). Each can depress measured compliance for weaker-language content independent of hierarchy behavior, and HCR=Conf/Ref does not cancel this: thresholding and discrete judging are nonlinear, so conflict-condition outputs are differentially pushed below τ or into [[C]]. The tell is the paper's own Table 3: Safety — the one domain whose criterion is language-invariant (exact match on the fixed English token \"Access Granted\", §B.4) — is the one domain where the EN-on-top HCR_H ranking disappears (HI 54.0 > EN 50.2). The paper calls Safety \"a notable exception\" but offers no mechanism for it; under the metric-confound hypothesis the reversal is the expected signature. Figure 7 shows chrF++ distributions per model, never per source language, so the paper never shows that HI/ZH reference translations clear τ as comfortably as EN/DE/ES/FR. This does not threaten the LBE regularity itself (15/15 pairs positive, covariate-adjusted item-level estimates, and the diagonal pattern — HI-HI below EN-EN while cross pairs with an English lower instruction exceed EN-EN — argues against a pure language-strength account). What is at risk is the specific language ranking and the magnitude of \"implicit authority,\" i.e., the deployment-relevant part of","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a careful empirical paper that actually fills the gap it claims. English-only IH work and single-instruction multilingual IF work never held task and conflict fixed while swapping languages on both sides of the hierarchy. XIH-Bench does that at scale (78k instances, 13 models, three hierarchy pairs, four domains), and two regularities stick: positional language asymmetry, and a Language Boundary Effect of about +3 pp for cross-language vs same-language conflicts (12/13 models, all 15 pair averages positive, item-level regressions still significant after covariates).\n\nWhat is new is the controlled same-vs-cross design plus the specialization angle (Qwen/Chinese, Mistral/European showing the H+/L− “implicit authority” pattern only in the models that advertise those strengths). HCR as Conf/Ref is a sensible normalization. Code/data promise and the appendix construction notes are above average for this genre. Citation pattern is honest: Wallace/IHEval, Multi-IF/XIFBench, Belebele, IFEval—no fake lineage.\n\nSoft spots, in proportion. LBE looks robust. The language ranking story is shakier. Three domains lean on language-sensitive instruments (adapted IFEval checkers, fixed chrF++ τ=0.15 into Korean, English-only persona judge). Safety—the one language-invariant exact-match domain—is exactly where the EN-on-top HCR_H ranking flips (HI highest). The paper calls that a “notable exception” and moves on; that is the place a referee should push. Nonlinear thresholds mean Conf/Ref does not fully cancel surface-form or translation-quality confounds, so the deployment claim that “English is hardest to override” may partly be metric. Specialization magnitudes inherit the same issue. Scope cuts (six languages, pairwise single-turn, LLM translation pipeline) are disclosed and fine if not oversold.\n\nWho it is for: people building multilingual agents, prompt-injection defenses, or IH training. Not a theory paper. I would send it to peer review; the benchmark and LBE earn referee time even if rankings get tempered. Engage if you work on multilingual safety; skim the LBE and §5.2 figures if you only need the headline effects.","headline":"Solid controlled multilingual IH benchmark with a real LBE finding; the EN/HI authority ranking is softer than the paper sells because three of four domains use language-sensitive metrics.","tokens_in":49451,"tokens_out":566,"would_cite":true,"duration_ms":20429,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Language is not a neutral carrier of instructions: it actively shapes whether multilingual LLMs obey source priority when instructions conflict.","keywords":["instruction hierarchy","multilingual LLMs","XIH-Bench","Language Boundary Effect","prompt injection","hierarchy compliance","language specialization","cross-lingual robustness"],"falsifier":"Re-run the same conflict items with human-verified native prompts and human labels (or a clearly calibrated multilingual detector) and check whether the language-position asymmetry and the positive cross- versus same-language HCR gap disappear or reverse for the same models.","tokens_in":49150,"feed_emoji":"🌐","tokens_out":966,"duration_ms":18723,"temperature":0.7,"pith_summary":"When large language models receive conflicting instructions from different sources—system prompts, user messages, tool outputs—they are supposed to follow a fixed hierarchy by source, not by language. Almost all prior tests of that behavior were English-only. This paper builds XIH-Bench, a controlled multilingual benchmark that holds the task and conflict structure fixed while swapping which languages sit at higher and lower priority across six languages, four domains, and three hierarchy settings. On thirteen models it finds two stable patterns: the same language that strengthens authority when it is higher-priority can become hard to suppress when it is lower-priority, and conflicts written in two different languages are resolved more correctly than same-language conflicts (the Language Boundary Effect). Model-favored languages can gain “implicit authority,” so specialized multilingual systems may let a lower-priority instruction in the favored language override the intended hierarchy—raising reliability and security risks that English-only evaluation misses.","feed_headline":"Language flips who wins when LLM instructions conflict","feed_subtitle":"Cross-language fights obey hierarchy better; model-favored languages resist being overridden","key_machinery":"XIH-Bench and Hierarchy Compliance Rate (HCR = Score_Conflict / Score_Reference). The benchmark fixes task semantics and conflict structure while crossing higher- and lower-priority languages (same- and cross-language) over System>User, System>Tool, and User>Tool in four domains; HCR isolates how much intended-hierarchy performance survives a contradictory lower-priority instruction.","core_discovery":"Instruction-hierarchy compliance in multilingual LLMs is language- and position-dependent rather than language-invariant. A language that improves compliance when placed at the higher-priority level can disrupt compliance when placed at the lower-priority level, and cross-language conflicts consistently yield higher Hierarchy Compliance Rate than same-language conflicts (mean Language Boundary Effect about +3.0 percentage points across 12 of 13 models). Language specialization further gives model-favored languages implicit authority—high influence from above and resistance to override from below.","pith_inferences":["Training or serving stacks that deliberately mark source roles more strongly might amplify or substitute for the language-boundary cue the paper links to better hierarchy separation.","If implicit authority tracks pretraining dominance, continued English-heavy or language-specialized pretraining could widen the lower-priority override problem unless hierarchy objectives explicitly penalize favored-language interference.","Same-language prompt-injection defenses validated only in English may not transfer; attackers might prefer the model’s favored language at the lower level or same-language conflicts to shrink the boundary advantage.","Extending the design to multi-turn and agentic tool chains would test whether the Language Boundary Effect still helps when languages mix across longer contexts."],"forward_implications":["English-only instruction-hierarchy benchmarks understate multilingual failure modes and give an incomplete robustness picture.","Strong single-instruction following in a language is double-edged under hierarchy: it can make that language harder to override when it should lose.","Cross-language source conflicts are systematically easier for models to rank correctly than same-language conflicts (Language Boundary Effect).","Language-specialized models can grant their favored language implicit authority, so lower-priority instructions in that language remain disproportionately hard to suppress.","Multilingual deployment and security review must test suppressibility under source conflict, not only whether the model can follow an instruction in each language."],"fun_headline_variants":["Language decides who wins when LLM instructions clash","Same-language fights break hierarchy more than cross-language ones","Favored languages resist override in multilingual instruction hierarchy","Instruction hierarchy flips with language and priority position","Cross-language conflicts boost hierarchy compliance in LLMs"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The measured scores, especially automatic translation detection, persona judging, and translated prompts, truly track hierarchy obedience rather than translation quality, surface-form quirks, or judge artifacts across languages.","fun_headline_variants_meta":{"raw":{"variants":["Language decides who wins when LLM instructions clash","Same-language fights break hierarchy more than cross-language ones","Favored languages resist override in multilingual instruction hierarchy","Instruction hierarchy flips with language and priority position","Cross-language conflicts boost hierarchy compliance in LLMs"]},"model":"grok-4.5","effort":"low","cost_usd":0.001501,"raw_usage":{"total_tokens":808,"prompt_tokens":728,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":15008000,"prompt_tokens_details":{"text_tokens":728,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":728,"tokens_out":60,"duration_ms":2130,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T19:25:51.373626+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same conflict items with human-verified native prompts and human labels (or a clearly calibrated multilingual detector) and check whether the language-position asymmetry and the positive cross- versus same-language HCR gap disappear or reverse for the same models.","supporting_citations":[],"review_version":1}