{"id":"d954e58e-76ad-48ee-a5e5-c790136524f0","arxiv_id":"2607.06008","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A hand-curated 67-task multilingual long-horizon agent benchmark finds large domain, language, and harness-driven performance gaps for current LLM agents.","lead":"PolyWorkBench is a 67-task benchmark for testing LLM agents on multilingual, multi-step workplace workflows across five domains. It shows frontier agents drop sharply under mixed-language inputs and that agent scaffolding can swing scores by double-digit points.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The abstract's causal claim of multilingual degradation vs monolingual counterparts lacks a controlled same-task ablation in the reported experiments.","rationale":"The reader's weakest_assumption correctly isolates the load-bearing gap: the abstract asserts degradation versus monolingual counterparts, yet Section 4 only characterizes performance inside PolyWorkBench. Harness sensitivity, hand-curated anchors, and language imbalance are disclosed but still block causal attribution to multilingual trajectory coupling. The within-suite findings (Commerce arithmetic/schema failures, Judge–Grade disagreement, harness matrix) remain coherent and useful, so the verdict stays CONDITIONAL rather than REJECT; the required tightening is precisely a same-task monolingual control plus clearer scoping of the comparative claim. No stronger internal inconsistency was found.","tokens_in":16388,"tokens_out":534,"duration_ms":5888,"concrete_test":"Construct a monolingual English-only mirror of a stratified subset of ≥15 tasks (covering COM/KNW/LEG/LOC/MFG and both pools), keeping tools, file schemas, step budgets, and ground-truth anchors identical while translating all roles to English. Re-run the top 3 model×harness pairs under the same protocol; if mean Grade rises by <0.05 relative to the multilingual originals, the degradation claim is unsupported and should be rephrased as within-suite characterization.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's strongest claim (abstract + §1) is that SOTA agents suffer significant degradation in multilingual long-horizon workflows relative to monolingual counterparts, with multilinguality introducing compounding trajectory effects. Section 4, however, reports only within-PolyWorkBench variation: domain (Commerce dip), language (RU/ES/DE drops), harness (8–21 Pass@1 points on the same model), and sampling headroom. No controlled monolingual counterpart of the same 67 tasks (identical tools, schemas, step counts, and back-injected anchors, with all instruction/source/output roles collapsed to one language) appears in the main results. Without that ablation, observed low Pass@1 (best 0.921; most <0.77) and non-uniform drops cannot be cleanly attributed to cross-lingual trajectory coupling rather than task hardness, hand-authored numerical/schema strictness (Table 2), harness scaffolding, or language-coverage imbalance (AR n=1). The paper itself surfaces these confounds; the comparative claim therefore rests on an untested contrast.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces PolyWorkBench, a 67-task benchmark spanning five workplace domains (commerce, knowledge, legal, localization, manufacturing) for evaluating LLM agents on multilingual long-horizon workflows. Tasks require heterogeneous multilingual inputs, multi-step tool use, and structured outputs, with language variation embedded across instruction, source, and output roles (88% of tasks involve three or more languages). Evaluation combines structural Grade, executable Pytest suites, and LLM-as-Judge. Across 18 model×harness entries, the best Pass@1 is 0.921 (Claude Opus 4.8 + ClaudeCode); most models fall below 0.77. The authors report large harness effects (8–21 Pass@1 points on the same model), a systematic Commerce dip, non-uniform per-language performance, weak Grade–Judge correlation (r=0.18), and sampling headroom that grows for mid-tier models. They conclude that multilinguality introduces compounding effects across reasoning and execution and that agents degrade relative to monolingual counterparts.","tokens_in":16723,"tokens_out":1601,"duration_ms":19589,"significance":"If the comparative claim holds, the work would fill a genuine gap between monolingual agent benchmarks (WebArena, OSWorld, SWE-bench, OdysseyBench) and static multilingual suites (MGSM, M-MMLU, FLORES) by treating language variation as a trajectory-level factor rather than an input attribute. The hybrid Grade/Pytest/Judge design, multi-harness matrix, Pass@1 vs Pass@3 analysis, and honest disclosure of Judge bimodality and harness sensitivity are concrete methodological contributions that other agent benchmarks can reuse. Even without a clean monolingual ablation, a carefully curated, executable multilingual workplace suite of this kind is useful for the community. The significance of the causal story about “cross-lingual trajectory coupling,” however, depends on isolating multilinguality from task hardness, schema strictness, and scaffolding—something the current experiments only partially achieve.","major_comments":[{"comment":"Abstract and §1 claim that SOTA agents “suffer significant performance degradation in multilingual workflow settings compared to monolingual counterparts” and that multilinguality introduces compounding trajectory effects. Section 4 reports only within-PolyWorkBench variation (domain, language, harness, sampling). No controlled same-task monolingual ablation—identical tools, schemas, step counts, and back-injected anchors, with instruction/source/output collapsed to one language—appears in the main results. Without that contrast, low Pass@1 and domain/language drops cannot be cleanly attributed to cross-lingual trajectory coupling rather than numerical/schema strictness (Table 2), harness scaffolding (Fig. 6), or hand-authored difficulty. This is load-bearing for the paper’s central comparative claim and should be added or the claim should be narrowed to within-benchmark multilingual dif","section":"Abstract, §1, §4"},{"comment":"Language coverage is highly imbalanced (Fig. 2(b); Appendix A.2): English touches 66 tasks while Arabic has n=1 (uniform Grade 0.850 by construction). Per-language means for RU/ES/DE are therefore informative, but the “ten languages” framing and any claim of broad multilingual generalization overstate coverage. Either expand low-resource languages or report primary analyses only on languages with adequate task counts and treat AR as a pilot case.","section":"§3.3, Fig. 2(b), Appendix A.2"},{"comment":"Harness choice moves Pass@1 by 0.08–0.21 on the same model (Fig. 6; Table 1), and ClaudeCode is best or tied-best whenever available. The paper correctly treats harness as a first-class variable, but the leaderboard and abstract-level “agent” claims still risk being read as model rankings. The primary reported comparison should either fix one harness for all models or report model-level scores only after harness-normalized aggregation, with harness effects relegated to a sensitivity analysis rather than mixed into the main ranking narrative.","section":"§4.1–4.2, Table 1, Fig. 6"},{"comment":"The hybrid evaluation axiom—that Grade + Pytest + Judge jointly capture functional correctness and linguistic consistency—is only partially supported. Grade and Pytest align well (r=0.85), but Judge is weakly correlated (r=0.18 overall; r=−0.04 when Grade≥0.5) and heavily bimodal (§4.3, Fig. 5(a), Appendix A.3–A.4). Pass@1 is defined as mean Grade alone. If linguistic consistency is a core claim of the benchmark, either (i) define a composite metric that includes Judge under conditions where it is reliable, or (ii) demote Judge to a diagnostic secondary signal and revise claims about “linguistic consistency” accordingly. The current design honestly discloses the gap but still markets a three-axis framework whose third axis does not rank models.","section":"§3.4–3.5, §4.3, Fig. 5"}],"minor_comments":[{"comment":"Estimated step counts (mean 8.54) and difficulty scale 3–6 are used throughout §3.3 and Fig. 2(c) but are not operationally defined (tool calls? human annotation? agent traces?). A short definition or measurement protocol would help reproducibility.","section":"§3.3, Fig. 2(c)"},{"comment":"Several cited “2026” arXiv preprints (Claw-Eval, WildClawBench, CoffeeBench, MAPS, etc.) are contemporaneous or future-dated relative to the paper’s July 2026 date. Ensure citation status and availability are accurate at camera-ready time.","section":"§2, References"},{"comment":"Figure 1 and Figure 3 are dense overview diagrams; axis labels and small text in the language polar plot (Fig. 2(b)) may not reproduce well in print. Consider simplifying or enlarging key panels.","section":"Fig. 1–3"},{"comment":"Table 1 lists only one Codex entry; the harness comparison for Codex is underpowered relative to ClaudeCode/OpenClaw/Hermes. Note this limitation explicitly when discussing harness ordering.","section":"Table 1, Fig. 6"},{"comment":"Typographical consistency: “artefacts” vs “artifacts,” and occasional spacing anomalies around decimals (e.g., “0 .921”) appear in the compiled text and should be cleaned.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The benchmark construction and multi-harness reporting are above average for this subfield; I would not reject on novelty of the suite alone. The main risk is overclaiming a causal multilingual-vs-monolingual result that the experiments do not isolate. If the authors add a same-task monolingual control (even on a subset of domains) and tighten the abstract, this becomes a solid contribution. Fit for a serious AI/ML venue is good if the comparative claim is fixed; without it, the paper is closer to a useful resource release than a full research result on multilingual agent dynamics."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a useful evaluation artifact, not a theory paper. The real product is 67 hand-authored workplace workflows that force cross-lingual transfer across instruction, sources, and outputs, plus a hybrid Grade/Pytest/Judge stack and a full model×harness matrix. That combination is new enough to matter for people building enterprise agents.\n\nWhat they did well is concrete. Tasks are curated, not bulk-translated; anchors are back-injected so string-copying fails; 88% of tasks are trilingual by construction. The experiments are unusually honest about scaffolding: same model moves 8–21 Pass@1 points across harnesses, ClaudeCode dominates when available, and they publish the matrix instead of collapsing it. Domain and language heatmaps are readable; Commerce’s arithmetic/schema cliff and the RU/ES/DE drops are real patterns inside the suite. They also own the metric mess—Grade tracks Pytest (r≈0.85) while Judge is weak and bimodal (r≈0.18)—and keep Judge as a semantic check rather than the ranking signal. That is good experimental hygiene.\n\nSoft spots, in proportion. The abstract and intro claim significant degradation versus monolingual counterparts and “compounding” cross-lingual trajectory effects. Section 4 only characterizes variation inside PolyWorkBench (domain, language, harness, sampling). There is no controlled same-task monolingual ablation with identical tools, schemas, and anchors. So low Pass@1 (best 0.921; most <0.77) and the Commerce collapse could be hardness, strict numerical grading, or harness choice as much as multilingual coupling. Language coverage is uneven (Arabic n=1). Hand-authored difficulty is a feature for realism and a confound for causal claims. None of that kills the benchmark; it bounds how hard you can lean on the causal story.\n\nWho it’s for: agent-eval and multilingual-systems people who need a workplace suite harder than static XGLUE/MGSM and more language-mixed than WebArena/OSWorld. Math is light (means, Pass@k, correlations); citations look standard and fair. I’d bring it to reading group if the group cares about agent reliability. I’d cite the suite and the harness-sensitivity result. Send it to peer review—tighten the mono control or soften the abstract, expand low-resource coverage, ship the artifacts—and it will be a durable reference, not a desk reject.","headline":"Solid hand-curated multilingual agent bench with careful multi-harness reporting; the abstract’s mono-vs-multi degradation claim runs ahead of the controlled evidence in §4.","tokens_in":17393,"tokens_out":598,"would_cite":true,"duration_ms":11856,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"State-of-the-art LLM agents lose substantial accuracy when workplace workflows mix languages across reasoning, tools, and final outputs.","keywords":["LLM agents","multilingual benchmarks","long-horizon workflows","tool use","workplace tasks","hybrid evaluation","cross-lingual reasoning","structured outputs"],"falsifier":"Rerun the same 67 tasks under identical models, harnesses, timeouts, and scoring, but force instruction, sources, and required outputs into a single language; if Pass@1 and domain/language gaps do not close toward strong monolingual agent performance, the multilingual-compounding claim is weakened.","tokens_in":17246,"feed_emoji":"🌐","tokens_out":961,"duration_ms":18794,"temperature":0.7,"pith_summary":"Most agent benchmarks test long-horizon planning and tool use in a single language, while most multilingual benchmarks test static questions without multi-step execution. Real workplace work often mixes languages inside one workflow: instructions in one language, source documents in another, and deliverables in a third. This paper builds PolyWorkBench—67 hand-authored tasks across commerce, knowledge work, legal analysis, localization, and manufacturing—so agents must keep meaning aligned while retrieving, reasoning, calling tools, and producing structured artefacts. With hybrid scoring that checks structure, executable correctness, and semantic quality, the authors find that even strong models leave large room for failure, with uneven drops by domain and language. The practical claim is that language variation is not a side detail of inputs; it compounds along the agent’s trajectory, so agent evaluation has to treat multilinguality and procedural decision-making together.","feed_headline":"LLM agents drop hard on mixed-language workplace tasks","feed_subtitle":"A 67-task bench shows language mixing compounds errors across planning, tools, and outputs.","key_machinery":"PolyWorkBench: 67 end-to-end workplace tasks (five domains, ten languages) that embed language variation into the full execution trajectory, scored by a hybrid framework of structural Grade, executable Pytest checks, and LLM-as-Judge semantic assessment, with Pass@1 defined as mean Grade over all tasks.","core_discovery":"On multilingual long-horizon workplace workflows, current LLM agents suffer significant performance degradation relative to monolingual settings. Multilinguality introduces compounding effects across reasoning and execution steps—not only comprehension errors but also failures of planning stability, tool reliability, and cross-lingual coordination—so language variation and multi-step decision-making must be evaluated jointly rather than as separate axes.","pith_inferences":["Training and scaffolding that only optimize monolingual tool loops will systematically under-prepare agents for mixed-language enterprise pipelines even if static multilingual QA scores look strong.","Commerce-like workflows with unforgiving arithmetic and schema checks may be a sharper stress test for agent reliability than long-form legal or knowledge writing that awards partial structural credit.","Balancing rare languages and adding matched monolingual twins of each task would turn the benchmark into a cleaner causal test of trajectory coupling versus base capability.","Harness-invariant agent interfaces may matter as much as model scale for closing the multilingual gap the paper reports."],"forward_implications":["Agent leaderboards that ignore harness choice will mis-rank models, because the same model can shift by roughly 8–21 Pass@1 points across scaffolds.","Overall Pass@1 will overstate reliability for enterprise work that demands strict end-to-end numerical or schema correctness, especially commerce-style reconciliation tasks.","Mid-tier models can recover large Grade gains via multi-sample best-of-N, while top models are already near saturation on Pass@1.","Deterministic graders and LLM judges measure different axes; reporting Grade, Pytest, and Judge together is required to catch both functional failure and fluent-but-wrong or structure-correct-but-semantically-poor outputs.","Future agent evaluation should treat cross-lingual consistency as part of the trajectory, not only as input translation or final-output language choice."],"fun_headline_variants":["LLM agents degrade sharply on multilingual long-horizon workflows","Mixed-language tasks compound LLM agent planning and tool errors","PolyWorkBench shows agents drop vs monolingual workplace settings","Multilinguality hits LLM agents across reasoning and execution","Agents fail more on cross-lingual long-horizon workplace tasks"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That the measured shortfalls are mainly caused by multilingual trajectory coupling, rather than by large harness effects, hand-built tasks with planted anchors, uneven language coverage, or the lack of a controlled same-task monolingual ablation in the main results.","fun_headline_variants_meta":{"raw":{"variants":["LLM agents degrade sharply on multilingual long-horizon workflows","Mixed-language tasks compound LLM agent planning and tool errors","PolyWorkBench shows agents drop vs monolingual workplace settings","Multilinguality hits LLM agents across reasoning and execution","Agents fail more on cross-lingual long-horizon workplace tasks"]},"model":"grok-4.5","effort":"low","cost_usd":0.006126,"raw_usage":{"total_tokens":1579,"prompt_tokens":789,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":61260000,"prompt_tokens_details":{"text_tokens":789,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":707,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":789,"tokens_out":83,"duration_ms":8486,"temperature":1.0,"reasoning_tokens":707,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T01:36:03.480897+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rerun the same 67 tasks under identical models, harnesses, timeouts, and scoring, but force instruction, sources, and required outputs into a single language; if Pass@1 and domain/language gaps do not close toward strong monolingual agent performance, the multilingual-compounding claim is weakened.","supporting_citations":[],"review_version":2}