{"id":"f6d71092-d8b1-45da-97b7-97ab33b170fa","arxiv_id":"2505.11665","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that categorizes multilingual prompting techniques by NLP task and language family, and designates potential state-of-the-art prompting methods for each dataset.","lead":"This paper surveys 36 studies of multilingual prompt engineering, sorting 39 prompting techniques across 30 NLP tasks and roughly 250 languages. It maps which prompting methods are state-of-the-art per dataset and analyzes research coverage by language family and resource level.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SoTA column is unfalsifiable: metrics are omitted, dataset versions vary, and 'informed judgment' cannot support a single best method per dataset.","rationale":"The reader's weakest assumption identifies the same core issue: the SoTA designations depend on comparing studies that use incompatible metrics, dataset versions, and model settings, while the paper explicitly relies on informed judgment rather than a reproducible protocol. My stress-test sharpens this with concrete internal evidence: Section 3 admits metrics are omitted, and the tables show cross-paper SoTA calls such as CLSP for MGSM and X-InSTA for MARC. I also add a supporting inconsistency in the taxonomy claim: PAWS-X and XNLI each appear under two different task categories despite the stated single-task assignment rule. None of this changes the reader's conditional verdict; it reinforces it. The proposed check would settle whether any SoTA entry survives a common-protocol comparison, and if it does not, the appropriate fix is to relabel or remove the SoTA column rather than present it as a verified ranking.","tokens_in":43716,"tokens_out":4845,"duration_ms":50388,"concrete_test":"Choose three cross-paper SoTA claims (e.g., MGSM->CLSP in Table 2, MARC->X-InSTA in Table 8, XNLI->XLT/Translate-En in Table 22). From the cited source papers, extract each method's score, metric, dataset version, LLM, and few-shot example count. Restrict to entries that use the same LLM and the same metric (e.g., GPT-3.5-Turbo accuracy for MGSM, BLOOMZ-7.1B accuracy for MARC), then re-rank. If the surveyed SoTA method is not top-ranked under the restricted comparison for at least two of the three datasets, the SoTA column should be rewritten as 'reported in a single study' or removed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing component of the survey is the per-dataset SoTA designation (abstract; Section 3). The paper itself states in Section 3 that 'Evaluation metrics are omitted, as they differ across studies' and that SoTA is assigned by 'informed judgment' because 'varying versions of the same dataset further complicate direct performance comparisons.' That disclosure is accurate, but it means the SoTA column cannot be independently checked: entries such as CLSP for MGSM (Table 2) and X-InSTA for MARC (Table 8) aggregate results across different LLMs, few-shot counts, decoding settings, and unstated metrics from different source papers. Nothing in the survey demonstrates that the named method would win under a single common protocol. The same issue weakens the taxonomy: despite the Section 3 claim that each dataset is assigned to one task, PAWS-X appears in both Task Understanding Consistency (Table 25) and Paraphrasing (Table 28), and XNLI appears in both Natural Language Inference (Table 22) and Task Understanding Consistency (Table 25). The 'potential SoTA' label is hedged, but the survey is presented as a resource for choosing methods, and the current presentation gives readers no basis to verify or reproduce those choices.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys multilingual prompt engineering for large language models, reviewing 36 research papers, 39 prompting techniques, 30 NLP tasks, and roughly 250 languages. It organizes the literature by NLP task, presenting per-task tables that list prompting strategies, LLMs, language counts, references, and a designated 'potential SoTA' method for each dataset. It then derives descriptive insights about the distribution of tasks and prompting techniques across language families and high- vs. low-resource languages. The central claims are that multilingual prompt engineering can be systematically categorized by NLP task and that per-dataset SoTA prompting methods can be identified from the surveyed literature.","tokens_in":43955,"tokens_out":4032,"duration_ms":42706,"significance":"If the taxonomy and SoTA designations were adequately supported, this survey would be a useful reference for practitioners selecting prompting methods in multilingual settings, and the language-family/resource-level analyses would help identify coverage gaps. A notable strength is that the descriptive statistics in Section 4 are transparent tallies of the authors' own curated tables rather than fitted or predicted quantities, and the paper states its selection criteria for included papers. However, the usefulness of the survey as a method-selection resource depends on the reliability of the SoTA column and on the consistency of the task taxonomy, both of which currently need substantial revision.","major_comments":[{"comment":"The SoTA designations are not verifiable as presented. The paper states that 'Evaluation metrics are omitted, as they differ across studies' and that 'the use of varying versions of the same dataset further complicates direct performance comparisons,' with SoTA chosen by 'informed judgment.' Because each SoTA entry aggregates results across different LLMs, few-shot counts, decoding settings, and unstated metrics from different source papers, a reader cannot independently check entries such as CLSP for MGSM (Table 2), XLT for XNLI (Table 22), or X-InSTA for MARC (Table 8). To support the central claim, the authors should either report the metric and experimental protocol behind each SoTA designation, or explicitly restrict the claim to 'best among methods compared under a common protocol' and remove designations that aggregate incomparable results.","section":"Section 3 (intro) and Tables 2-31"},{"comment":"The claim that 'we ensure each dataset is associated with a single NLP task' is contradicted by the survey's own tables. XNLI appears in both Natural Language Inference (Table 22) and Task Understanding Consistency (Table 25), while PAWS-X appears in both Task Understanding Consistency (Table 25) and Paraphrasing (Table 28). This undermines the stated principle of the taxonomy and complicates the descriptive analyses that count tasks per dataset. The authors should either allow explicit multi-label assignments with justification, or remove one of the duplicate assignments and clarify the boundaries between the affected task definitions.","section":"Section 3 taxonomy; Tables 22, 25, 28"},{"comment":"The literature search process is not reproducible as described. The paper lists 11 Google Scholar queries, says manual filtering produced 189 articles, and then applies two selection criteria to reach 36 papers, but it does not report the search date, the exact query strings used, the number of papers retrieved per query, or the screening decisions that removed papers. Since the language-family and resource-level statistics in Section 4 depend entirely on this selected corpus, the authors should provide a more complete and reproducible selection protocol, including a flowchart or exclusion log.","section":"Section 1.2"}],"minor_comments":[{"comment":"The text says 'we found two datasets which are XGLUE-POS, UDPOS and XTREME,' but lists three datasets; the table also contains three rows. Please correct the count or the wording.","section":"Section 3.15, Table 16"},{"comment":"The abstract reports 39 prompting techniques, while Section 4.1 and Figure 7 state '38 distinct prompting techniques.' Please reconcile the counts.","section":"Abstract and Section 4"},{"comment":"There are typographical errors in language-pair names: 'Tamil-Finish,' 'Finish-French,' 'Finish-German,' 'Finish-Arabic,' and similar entries should read 'Finnish'; 'Gujrati-Hindi' should be 'Gujarati-Hindi.'","section":"Table 13 and Section 3.12"},{"comment":"Several LaTeX artifacts appear in the text, including 'Y oruba,' 'V osoughi,' and 'ty pologically.' A copyediting pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a survey with a useful organizational structure, but the SoTA column is the main load-bearing weakness: without evaluation metrics or a common protocol, those designations cannot be checked. The taxonomy's duplicate dataset assignments also need resolution. These issues are fixable within the manuscript's scope, so I recommend major revision rather than rejection. I would also encourage the editor to ask the authors to make the literature search protocol reproducible before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know first: this is a useful survey, and its main weakness is the one the reader flagged. The authors organize 36 papers and 39 prompting techniques by NLP task across roughly 250 languages, and they standardize names like En-Basic, Native-CoT, and Translate-En-CoT, which genuinely helps when reading this scattered literature. The language-family and resource-level analyses are a reasonable new framing, not just a pile of citations. I'd trust the taxonomy as a starting map.\n\nThe soft spot is exactly where the stress-test lands. The per-dataset SoTA column in Tables 2–31 is presented as a usable result, but the paper itself says evaluation metrics are omitted and dataset versions vary, and that SoTA rests on 'informed judgment.' That isn't a fixable detail; it means the column cannot be independently verified or reproduced. A reader wanting to pick a method for MGSM or MARC gets a recommendation with no stated basis for comparison. The authors are honest about the limitation, but honesty doesn't make the claim load-bearing. I'd either rename these 'reported best under original study conditions' or drop them entirely.\n\nThere's also a contradiction with their stated methodology: Section 3 says each dataset is assigned to a single NLP task for clarity, yet PAWS-X appears in both Task Understanding Consistency (Table 25) and Paraphrasing (Table 28), and XNLI appears in both NLI (Table 22) and Task Understanding Consistency (Table 25). That double-assignment weakens the taxonomy's clean structure.\n\nThe selection process (11 Google Scholar queries plus manual filtering) is documented but not reproducible in a strict sense; it's a reasonable effort, though, and the field doesn't demand systematic review standards here. The descriptive statistics are tallies of their own curated tables, so numbers like '38 techniques in high-resource vs. 20 in low-resource' are only as solid as the table curation. I haven't re-counted every row, so I won't pile on beyond the reader's note about internal inconsistencies.\n\nWho gets value from this? A practitioner who wants a quick, task-organized map of what multilingual prompting techniques exist and which papers use which datasets. A researcher looking for a citable survey of the subfield will also find it useful, provided they don't quote the SoTA column as gospel. It deserves a serious referee; the right fix is to strip or heavily qualify the SoTA designations, resolve the dataset overlaps, and add a note on how the manual filtering was applied. I'd send it for review with major revisions rather than desk-reject it.","headline":"A genuinely useful task-organized map of multilingual prompting, undercut by a SoTA column the authors admit cannot be checked.","tokens_in":44418,"tokens_out":1605,"would_cite":true,"duration_ms":18953,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey sorts 39 multilingual prompting techniques by NLP task and names the best method for each benchmark dataset.","keywords":["multilingual prompt engineering","large language models","cross-lingual transfer","chain-of-thought prompting","state-of-the-art prompting methods","low-resource languages","NLP task taxonomy","in-context learning"],"falsifier":"Run the designated SoTA method and its strongest competitor on the same dataset, with the same metric and decoding settings; if the competitor wins on MGSM, XNLI, or FLORES, that dataset's SoTA designation is wrong. A milder check: for any SoTA entry backed by a single study, reproduction by an independent group would be the first test.","tokens_in":43557,"feed_emoji":"🧭","tokens_out":5688,"duration_ms":52034,"temperature":0.7,"pith_summary":"This survey claims that multilingual prompt engineering is now systematic enough to be organized by NLP task, and that a task-based taxonomy is the right way to compare methods. It reviews 36 papers, 39 prompting techniques, and 30 multilingual tasks spanning roughly 250 languages, and assigns each dataset a potential best (SoTA) prompting method. The point of the exercise is practical: a researcher or practitioner who wants to prompt an LLM in multiple languages can look up which technique has worked best for their task and language, rather than re-deriving the field from scratch. The survey also documents a clear imbalance: most methods and studies concentrate on high-resource languages, while the wide low-resource coverage is driven almost entirely by machine translation.","feed_headline":"Survey names the best prompt for each multilingual task","feed_subtitle":"A task-by-task guide to prompting LLMs across 250 languages, with state-of-the-art picks for 30 benchmarks.","key_machinery":"The carrying apparatus is a task-first taxonomy plus a standardized naming scheme for prompting techniques. The survey groups 30 NLP tasks, from reasoning and question answering to translation, sequence labeling, and dialogue evaluation, and for each dataset it tabulates the prompting strategies tested, the LLMs used, the language count, and the designated SoTA method. The standardization of names, which collapses Basic, Standard, Vanilla, and Direct prompting into En-Basic or Native-Basic and merges variants into {Technique} + Variations, is what lets results from different papers sit side by side in one comparison.","core_discovery":"On the paper's own terms, the discovery is that multilingual prompt engineering can be productively standardized and catalogued: methods with different names across the literature are consolidated under common labels such as En-Basic, Native-Basic, X-Basic, XLT, CLP, CLSP, X-InSTA, SAP, DIPMT, DecoMT, MAPS, and MEEP, and each is mapped to the NLP task and datasets where it has been tested. The survey then identifies a potential state-of-the-art prompting method for each dataset, for instance CLSP for MGSM, XLT for XNLI and PAWS-X, X-InSTA for MARC and CLS, DecoMT for FLORES, and MAPS for WMT-22. It further argues that cross-lingual reasoning methods that route through English or align a native language with English generally beat plain English or native baselines on reasoning-heavy tasks, and that this whole field is recent and fast-moving.","pith_inferences":["Read the SoTA column as provisional: the survey itself notes that evaluation metrics are omitted and that designations rely on informed judgment, so the real ranking is an empirical question until methods meet on identical benchmarks.","A testable extension would sort the same prompts by language family rather than by task, asking whether English-routing methods hold their edge in typologically distant families such as Dravidian or Niger-Congo.","The high-resource versus low-resource gap suggests an immediate research target: port the twenty prompting techniques already used on high-resource languages to low-resource settings beyond translation, where only a subset has been tried.","The taxonomy could be extended into a living leaderboard, replacing informed judgment with standardized evaluation across the same model family."],"forward_implications":["A practitioner can select a starting-point prompt for a multilingual dataset from the SoTA column, such as CLSP for MGSM or X-InSTA for MARC, instead of testing every published variant.","Cross-lingual chain-of-thought and translation-anchored methods are the strongest family for reasoning and inference tasks, while dictionary-based and memory-based prompting dominates translation for low-resource languages.","The taxonomy gives future surveys and benchmarks a common vocabulary, since it consolidates multiple names for the same core prompting idea.","Low-resource languages are covered more broadly than commonly assumed, but almost entirely inside machine translation; other tasks remain high-resource dominated.","Because most included studies appeared within the last two years, the SoTA designations are expected to shift quickly as new methods are published."],"supporting_citations":[{"why":"Provides MGSM and the Native-CoT, En-CoT, and Translate-En-CoT baselines behind the reasoning-task comparisons.","marker":"Shi et al. (2022)"},{"why":"Introduces XLT and Translate-En, which the survey credits as leading methods on MKQA, XNLI, and PAWS-X.","marker":"Huang et al. (2023)"},{"why":"Introduces CLP and CLSP, the designated SoTA methods on MGSM and XCOPA.","marker":"Qin et al. (2023)"},{"why":"Introduces alignment-based prompting, with X-InSTA named SoTA on MARC, CLS, and HatEval.","marker":"Tanwar et al. (2023)"},{"why":"Contributes SAP, a few-shot prompting method for bidirectional LLMs used in the FLORES and XQuAD comparisons.","marker":"Patel et al. (2022)"},{"why":"Contributes dictionary-based DIPMT, the designated SoTA on FLORES-101.","marker":"Ghazvininejad et al. (2023)"},{"why":"Contributes DecoMT, the designated SoTA on FLORES for related-language pairs.","marker":"Puduppully et al. (2023a)"},{"why":"Contributes Rerank and MAPS, with MAPS designated SoTA on WMT-22.","marker":"He et al. (2024)"},{"why":"Introduces MEEP, the designated SoTA on FED, SEE, KdConv, and LCCC.","marker":"Ferron et al. (2023)"},{"why":"Supplies cross-lingual few-shot results behind many Native-Basic and X-Basic SoTA entries.","marker":"Asai et al. (2023)"}],"fun_headline_variants":["Survey names top prompt for each multilingual benchmark","Best prompts for 30 NLP tasks across 250 languages","Multilingual prompt guide: 39 techniques, winners per task","Cross-lingual reasoning prompts outperform English baselines","Prompt survey: state-of-the-art methods for multilingual LLMs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole per-dataset SoTA ranking rests on the assumption that results from different studies can be compared at all, even though evaluation metrics differ and dataset versions vary across the surveyed papers.","fun_headline_variants_meta":{"raw":{"variants":["Survey names top prompt for each multilingual benchmark","Best prompts for 30 NLP tasks across 250 languages","Multilingual prompt guide: 39 techniques, winners per task","Cross-lingual reasoning prompts outperform English baselines","Prompt survey: state-of-the-art methods for multilingual LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000261,"raw_usage":{"total_tokens":1620,"prompt_tokens":1002,"completion_tokens":618,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":539}},"tokens_in":618,"tokens_out":618,"duration_ms":6099,"temperature":1.0,"reasoning_tokens":539,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:49:18.508198+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the designated SoTA method and its strongest competitor on the same dataset, with the same metric and decoding settings; if the competitor wins on MGSM, XNLI, or FLORES, that dataset's SoTA designation is wrong. A milder check: for any SoTA entry backed by a single study, reproduction by an independent group would be the first test.","supporting_citations":[],"review_version":1}