Pith. sign in

REVIEW 1 major objections 2 minor 185 cited by

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

T0 review · 1 major / 2 minor · reviewed 2026-05-11 · grok-4.3

Pith's one-line read GLM-4.5 reaches 70.1 percent on TAU-Bench and 91 percent on AIME 24 using an open-source 355B-parameter MoE model with only 32B parameters active at once.

desk verdict GLM-4.5 is a practical open MoE release with competitive ARC benchmark numbers, but the evaluation details need checking before the rankings can be taken as settled. read the letter →

arxiv 2508.06471 v1 submitted 2025-08-08 cs.CL

GLM-4.5 Team: Aohan Zeng , Xin Lv , Qinkai Zheng , Zhenyu Hou , Bin Chen , Chengxing Xie , Cunxiang Wang , Da Yin
show 161 more authors
Hao Zeng Jiajie Zhang Kedong Wang Lucen Zhong Mingdao Liu Rui Lu Shulin Cao Xiaohan Zhang Xuancheng Huang Yao Wei Yean Cheng Yifan An Yilin Niu Yuanhao Wen Yushi Bai Zhengxiao Du Zihan Wang Zilin Zhu Bohan Zhang Bosi Wen Bowen Wu Bowen Xu Can Huang Casey Zhao Changpeng Cai Chao Yu Chen Li Chendi Ge Chenghua Huang Chenhui Zhang Chenxi Xu Chenzheng Zhu Chuang Li Congfeng Yin Daoyan Lin Dayong Yang Dazhi Jiang Ding Ai Erle Zhu Fei Wang Gengzheng Pan Guo Wang Hailong Sun Haitao Li Haiyang Li Haiyi Hu Hanyu Zhang Hao Peng Hao Tai Haoke Zhang Haoran Wang Haoyu Yang He Liu He Zhao Hongwei Liu Hongxi Yan Huan Liu Huilong Chen Ji Li Jiajing Zhao Jiamin Ren Jian Jiao Jiani Zhao Jianyang Yan Jiaqi Wang Jiayi Gui Jiayue Zhao Jie Liu Jijie Li Jing Li Jing Lu Jingsen Wang Jingwei Yuan Jingxuan Li Jingzhao Du Jinhua Du Jinxin Liu Junkai Zhi Junli Gao Ke Wang Lekang Yang Liang Xu Lin Fan Lindong Wu Lintao Ding Lu Wang Man Zhang Minghao Li Minghuan Xu Mingming Zhao Mingshu Zhai Pengfan Du Qian Dong Shangde Lei Shangqing Tu Shangtong Yang Shaoyou Lu Shijie Li Shuang Li Shuang-Li Shuxun Yang Sibo Yi Tianshu Yu Wei Tian Weihan Wang Wenbo Yu Weng Lam Tam Wenjie Liang Wentao Liu Xiao Wang Xiaohan Jia Xiaotao Gu Xiaoying Ling Xin Wang Xing Fan Xingru Pan Xinyuan Zhang Xinze Zhang Xiuqing Fu Xunkai Zhang Yabo Xu Yandong Wu Yida Lu Yidong Wang Yilin Zhou Yiming Pan Ying Zhang Yingli Wang Yingru Li Yinpei Su Yipeng Geng Yitong Zhu Yongkun Yang Yuhang Li Yuhao Wu Yujiang Li Yunan Liu Yunqing Wang Yuntao Li Yuxuan Zhang Zezhen Liu Zhen Yang Zhengda Zhou Zhongpei Qiao Zhuoer Feng Zhuorui Liu Zichen Zhang Zijun Yao Zikang Wang Ziqiang Liu Ziwei Chai Zixuan Li Zuodong Zhao Wenguang Chen Jidong Zhai Bin Xu Minlie Huang Hongning Wang Juanzi Li Yuxiao Dong Jie Tang
This is my paper · ORCID
classification cs.CL
keywords mixtureofexpertslargelanguagemodelagentictasksreasoningcodingopensourcebenchmarkevaluationreinforcementlearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents GLM-4.5 as an open-source Mixture-of-Experts model built to handle agentic tasks, complex reasoning, and coding problems. It describes a multi-stage training process on 23 trillion tokens followed by expert iteration and reinforcement learning that produces a hybrid reasoning capability allowing both extended thinking traces and direct answers. The model posts the listed benchmark scores and ranks near the top of evaluated systems despite activating far fewer parameters than some denser competitors. A smaller 106B-parameter variant is also released to broaden access. The work aims to supply capable tools for building practical AI agents and technical problem solvers.

What carries the argument

The hybrid reasoning method that supports both thinking and direct response modes, built inside a Mixture-of-Experts architecture with 355 billion total parameters but only 32 billion activated per token.

What would settle it

Independent re-evaluation of the model on the same benchmark problems using fresh, publicly documented prompts and code, or testing on a new suite of problems created after the training cutoff, would confirm or refute the claimed scores.

Watch

Extended reading notes

Core claim

GLM-4.5 achieves strong performance across agentic, reasoning, and coding (ARC) tasks, scoring 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. With much fewer parameters than several competitors, GLM-4.5 ranks 3rd overall among all evaluated models and 2nd on agentic benchmarks.

Load-bearing premise

That the reported benchmark scores reflect genuine capabilities measured through fair, standardized, and uncontaminated evaluations that allow direct comparison to other models.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 2 minor

Summary. The manuscript introduces GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters. It features a hybrid reasoning method that supports both thinking and direct response modes. The model undergoes multi-stage training on 23T tokens and post-training with expert model iteration and reinforcement learning. GLM-4.5 reports strong results across agentic, reasoning, and coding (ARC) tasks, including 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. It ranks 3rd overall among evaluated models and 2nd on agentic benchmarks despite having fewer parameters than several competitors. A compact variant, GLM-4.5-Air (106B parameters), is also released, with code and models made available at a GitHub repository.

Significance. If the benchmark results hold under verifiable and standardized conditions, the work advances open-source models for agentic and reasoning tasks by demonstrating competitive performance with an efficient MoE architecture and hybrid reasoning. The public release of both the full and compact models, along with code, is a clear strength that enables reproducibility and community follow-up research on ARC capabilities.

major comments (1)
  1. [Abstract] Abstract: The central performance claims, including the specific scores of 70.1% on TAU-Bench and 64.2% on SWE-bench Verified together with the 3rd overall and 2nd agentic ranking, are presented without any description of the evaluation methodology. Details on agent scaffolding, tool-use protocols, attempt limits, prompting consistency, use of the hybrid thinking mode, and data-contamination controls are required to establish that the results are comparable to those of competing models; their absence undermines confidence in the headline rankings.
minor comments (2)
  1. [Abstract] The phrase 'expert model iteration' in the abstract is used without definition or reference to a methods section; a brief clarification would improve readability.
  2. The efficiency claim ('much fewer parameters than several competitors') would be strengthened by explicitly listing the parameter counts of the referenced competing models in a comparison table.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the constructive comments on our manuscript. The feedback highlights an important point about ensuring transparency in the abstract for benchmark results. We address this directly below.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The central performance claims, including the specific scores of 70.1% on TAU-Bench and 64.2% on SWE-bench Verified together with the 3rd overall and 2nd agentic ranking, are presented without any description of the evaluation methodology. Details on agent scaffolding, tool-use protocols, attempt limits, prompting consistency, use of the hybrid thinking mode, and data-contamination controls are required to establish that the results are comparable to those of competing models; their absence undermines confidence in the headline rankings.

    Authors: We agree that the abstract, constrained by length, omits explicit methodology details, which can affect immediate assessment of comparability. The full manuscript contains sections on evaluation protocols that cover agent scaffolding (standard setups for TAU-Bench and SWE-bench), tool-use protocols, attempt limits, prompting strategies, selective use of the hybrid thinking mode, and data-contamination controls via held-out test sets and decontamination procedures. In the revision, we will expand the abstract with a concise clause summarizing these elements and add cross-references to the detailed methodology sections. This change will improve clarity while preserving the abstract's brevity. We do not believe the core results or rankings require alteration, only better contextualization. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: purely empirical benchmark reporting

full rationale

The paper describes training GLM-4.5 (355B MoE) on 23T tokens with post-training and RL, then reports measured benchmark scores (70.1% TAU-Bench, 91.0% AIME 24, 64.2% SWE-bench Verified). No mathematical derivations, equations, fitted predictions, or first-principles results exist. Claims rest on independent empirical evaluations with no self-definitional loops, fitted-input predictions, or load-bearing self-citations that reduce the central results to inputs by construction. Standard model-release structure; derivation chain is absent.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

As an empirical report on a trained foundation model, the central claims rest on standard machine learning assumptions including the validity of benchmark evaluations and the effectiveness of the described training pipeline; no novel axioms, free parameters, or invented entities are introduced beyond typical hyperparameter choices in LLM training.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models." pith.science (2026). https://pith.science/paper/2508.06471

@misc{pith2026250806471,
  author       = {Pith},
  title        = {Pith review of: GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2508.06471}},
  note         = {Machine review of arXiv:2508.06471}
}
read the original abstract

We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes. Through multi-stage training on 23T tokens and comprehensive post-training with expert model iteration and reinforcement learning, GLM-4.5 achieves strong performance across agentic, reasoning, and coding (ARC) tasks, scoring 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. With much fewer parameters than several competitors, GLM-4.5 ranks 3rd overall among all evaluated models and 2nd on agentic benchmarks. We release both GLM-4.5 (355B parameters) and a compact version, GLM-4.5-Air (106B parameters), to advance research in reasoning and agentic AI systems. Code, models, and more information are available at https://github.com/zai-org/GLM-4.5.

Discussion (0). Continue with ORCID to comment.

Lean theorems connected to this paper

Citations machine-checked in the Pith Canon. Every link opens the source theorem in the public Lean library.

  • IndisputableMonolith.Foundation.DAlembert.Inevitability bilinear_family_forced unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    GLM-4.5 achieves strong performance across agentic, reasoning, and coding (ARC) tasks, scoring 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. With much fewer parameters than several competitors, GLM-4.5 ranks 3rd overall among all evaluated models and 2nd on agentic benchmarks.

  • IndisputableMonolith.Foundation.PhiForcing phi_equation unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes.

  • IndisputableMonolith.Foundation.LedgerForcing conservation_from_balance unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    Through multi-stage training on 23T tokens and comprehensive post-training with expert model iteration and reinforcement learning

What do these tags mean?
matches
The paper's claim is directly supported by a theorem in the formal canon.
supports
The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
extends
The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
uses
The paper appears to rely on the theorem as machinery.
contradicts
The paper's claim conflicts with a theorem or certificate in the canon.
unclear
Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.

Forward citations

Showing 60 of 185 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 185 Pith citations

  1. UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing

    cs.DC 2026-06 unverdicted novelty 8.0 of 10

    UltraEP is the first exact-load real-time expert balancer for large-EP MoE training and serving on rack-scale nodes, reaching 94.3% of ideal throughput and 1.49x over no-balancing.

  2. Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models

    cs.AR 2026-05 conditional novelty 8.0 of 10

    Sieve dynamically schedules MoE experts across GPU and PIM hardware to handle bimodal token distributions, achieving 1.3x to 1.6x gains in throughput and interactivity over static prior PIM systems on three large models.

  3. ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning

    cs.LG 2026-05 conditional novelty 8.0 of 10

    ReLibra uses pre-known token-to-expert routing from RL rollouts to perform inter-batch expert reordering and intra-batch replication, delivering up to 1.6x higher throughput than Megatron-LM and 1.2x over oracle-equip...

  4. WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

    cs.CV 2026-05 unverdicted novelty 8.0 of 10

    WildTableBench is the first QA benchmark for naturally occurring table images, where 21 multimodal models were evaluated and only one exceeded 50% accuracy.

  5. Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models

    cs.AI 2026-04 unverdicted novelty 8.0 of 10

    User-turn generation reveals that LLMs' interaction awareness is largely decoupled from task accuracy, remaining near zero in deterministic settings even as accuracy scales to 96.8% on GSM8K.

  6. SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

    cs.LG 2026-09 conditional novelty 7.0 of 10

    SEAL aligns shared experts in MoE models via DPO LoRA to provide a router-independent safety surface, reducing attack success by up to 60% with negligible capability cost.

  7. From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

    cs.CL 2026-09 accept novelty 7.0 of 10

    A modular post-training recipe with separate GRPO experts per weakness axis and two-stage SLERP merging produces a single Qwen3-32B model that outperforms a roughly 7x larger baseline on in-house production benchmarks...

  8. SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue

    cs.CL 2026-09 conditional novelty 7.0 of 10

    The paper introduces SDARE-Bench, a scenario-based benchmark showing that LLMs struggle to detect and appropriately respond to stigma in conversational contexts, particularly in multi-speaker dialogues with group pressure.

  9. Controllable Image Captioning with Prompt-Conditioned Scene Rewards

    cs.CV 2026-09 accept novelty 7.0 of 10

    FoCUS uses prompt-conditioned signed weights over scene-graph components to steer image captions toward user-specified semantic emphases (attributes, relations, foreground, background).

  10. EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants

    cs.AI 2026-08 accept novelty 7.0 of 10

    A benchmark of 150 five-turn UI maintenance tasks shows that even the strongest model completes only 37.3% of five-turn episodes, with tool-grounded tasks proving especially difficult (52.4% adjacent retention).

  11. SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

    cs.CL 2026-08 conditional novelty 7.0 of 10

    SWE Refactor Bench grades coding agents on 20 whole-repository stack migrations with a three-stage protocol, and finds 5.4% of 520 runs passed all stages, with 13 of 20 tasks unsolved.

  12. MidTool: Mid-training Data Synthesis for Agentic Tool Use

    cs.AI 2026-08 conditional novelty 7.0 of 10

    A 20.3B-token mid-training mixture for general tool use improves downstream function-calling and agentic tool-use performance on BFCL, tau2-Bench, and MCP-Universe.

  13. Synthetic Persona Pretraining: Alignment from Token Zero

    cs.LG 2026-08 conditional novelty 7.0 of 10

    Injecting first-person, value-laden reflections into pretraining text improves constitution following, jailbreak resistance, and out-of-distribution moral choices in small language models, with the largest gains when ...

  14. TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

    cs.CL 2026-08 conditional novelty 7.0 of 10

    TreeProbe, a 4,719-item Tibetan-medicine benchmark, shows LLMs score 40–60% and systematically drift to TCM or biomedical reasoning.

  15. MemTX: Transactional Belief Commit for Stateful Agent Memory

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Staging agent-memory writes through a validate-and-commit pipeline with maturity-gated irreversible actions and typed cascading repair yields zero realized downstream harm on five LLM backbones, where eight baselines ...

  16. Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

    cs.AI 2026-07 unverdicted novelty 7.0 of 10

    Static SFT and RL training for tool-use agents leads to performance drops under open-world distributional shifts across perception, interaction, reasoning and internalization; perturbation-augmented fine-tuning is pro...

  17. MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

    cs.CL 2026-07 unverdicted novelty 7.0 of 10

    Multilingual LLMs show a reproducible 'Illusion of Cultural Alignment': they can be fluent in a language while lacking the culture's factual knowledge, and confidence, sampling, and retrieval do not fix it.

  18. The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    Proposes Monotonic Inference Policy Improvement (MIPI) objective and MIPU two-step update framework to address objective misalignment between training and inference policies in LLM reinforcement learning.

  19. MacroLens: A Multi-Task Benchmark for Contextual Financial Reasoning under Macroeconomic Scenarios

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    MacroLens is a point-in-time multi-signal benchmark dataset and seven tasks for evaluating contextual financial reasoning models under macroeconomic scenarios.

  20. Is Agent Code Less Maintainable Than Human Code?

    cs.SE 2026-06 unverdicted novelty 7.0 of 10

    Agent-generated code produces up to 13.1% lower follow-on task resolution rates than human code in chained repository-level experiments, with differences linked to behavioral patterns rather than conventional metrics.

  21. When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    Empirical evaluation on LiveCodeBench shows certainty-based RLIF yields early gains followed by output shortening and reasoning collapse, providing no advantage for RLVR initialization on code tasks.

  22. LegalWorld: A Life-Cycle Interactive Environment for Legal Agents

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    LegalWorld is a life-cycle interactive environment modeling Chinese civil litigation as five causally connected stages grounded in 75,309 judgments, paired with LongJud-Bench for cross-stage agent evaluation.

  23. daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    daVinci-kernel trains one LLM to select, use, and summarize reusable GPU-kernel optimization skills in a single RL loop, reaching 37.2%, 70.6%, and 32.2% Fast-1 pass rates on KernelBench Levels 1-3 at 14B.

  24. ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm

    cs.SE 2026-06 unverdicted novelty 7.0 of 10

    Proposes COM-as-Action paradigm for deterministic software manipulation, introduces ComCADBench benchmark and ComActor agent that achieves SOTA performance over GUI baselines.

  25. LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    LoopMoE is a looped MoE language model that outperforms matched vanilla MoE on 8 of 9 downstream benchmarks at 3B scale and continues to outperform at 9B scale under strictly controlled budgets.

  26. Stateful Visual Encoders for Vision-Language Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Stateful visual encoders condition each visual representation on prior features, yielding consistent gains on multi-image tasks under supervised finetuning across model sizes and domains.

  27. OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

    cs.CV 2026-06 accept novelty 7.0 of 10

    OVO-S-Bench provides 1680 human-annotated questions on 348 videos to measure streaming spatial intelligence in MLLMs across instantaneous perception, spatiotemporal tracking, spatial simulation, and allocentric mapping.

  28. Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    Moral Trolley Arena shows frontier LLMs produce composite moral preferences that are compressed rather than additive functions of calibrated component act strengths across Moral Foundations Theory.

  29. VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    VitaBench 2.0 introduces a benchmark for long-term personalized and proactive agent behavior, with results indicating substantial gaps in current frontier LLMs.

  30. LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    LatentOmni proposes a latent-space cross-modal reasoning framework that uses feature-level supervision and Omni-Sync Position Embedding to align and synchronize audio-visual latents, supported by a new 35K interleaved...

  31. CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    CopT reverses CoT by eliciting a draft answer first then using continuous-embedding contrastive verification and on-policy thinking to reflect and correct, yielding up to 23% higher accuracy and 57% fewer tokens witho...

  32. PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    PRISM benchmark of over 10k pairs shows LLMs have a 41% average drop from code execution success to spatial correctness in programmatic video generation.

  33. Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers

    math.OC 2026-05 unverdicted novelty 7.0 of 10

    Proposes equivariant optimizer updates matched to layer symmetries for embeddings, SwiGLU MLPs, and MoE routers, with reported gains in validation loss and training stability on several language model architectures.

  34. A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    PARAMΔ upcycles dense models to MoE for per-language experts and grafts post-training deltas to enable data-efficient language expansion while preserving original capabilities.

  35. BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    Introduces BacktestBench benchmark with 18k QA pairs across four backtesting tasks and evaluates 23 LLMs via the AutoBacktest multi-agent system.

  36. GGBound: A Genome-Grounded Agent for Microbial Life-Boundary Prediction

    cs.CY 2026-05 unverdicted novelty 7.0 of 10

    A genome-conditioned 4B LLM agent predicts microbial life boundaries and matches larger frontier models via token fusion, tool use, and a counterfactual gene-grounding reward.

  37. CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    CUDABeaver shows LLM CUDA debuggers often degenerate code for test-passing at the cost of speed, with protocol-aware metrics shifting success rates by up to 40 percentage points.

  38. Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL

    cs.AI 2026-05 conditional novelty 7.0 of 10

    Masked diffusion language models, not larger autoregressive LLMs, are the better building block for text-based world models in agentic RL, improving rollout fidelity, diversity, and downstream task success.

  39. Towards Temporal Compositional Reasoning in Long-Form Sports Videos

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    SportsTime plus Chain-of-Time Reasoning (temporal-reward GRPO and anchor-observe-infer) modestly lifts open-ended sports VideoQA and step-wise temporal grounding over 4B–8B MLLM baselines.

  40. Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    Expert upcycling expands MoE models by duplicating experts and continuing pre-training, matching baseline performance while saving 32% GPU hours in 7B-13B experiments.

  41. OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset

    cs.CL 2026-03 unverdicted novelty 7.0 of 10

    OmniCompliance-100K supplies 12,985 distinct rules and 106,009 associated real-world cases from 74 multi-domain regulations to benchmark LLM safety and compliance.

  42. SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    SPIRAL is a closed-loop think-act-reflect framework using PlanAgent, VideoGenerator, and CriticAgent plus GRPO self-evolution to improve long-horizon action-conditioned video generation, with new dataset and benchmark...

  43. EvoESAP: Non-Uniform Expert Pruning for Sparse MoE

    cs.LG 2026-03 conditional novelty 7.0 of 10

    EvoESAP uses evolutionary search guided by a speculative-decoding-inspired ESAP metric to discover non-uniform layer-wise sparsity allocations for MoE expert pruning, improving generation accuracy up to 19.6% at 50% sparsity.

  44. SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise

    cs.IR 2026-02 unverdicted novelty 7.0 of 10

    SQuTR is a large bilingual benchmark of 37,317 synthesized spoken queries under clean/low/medium/high noise, showing that retrieval quality steadily degrades as noise increases.

  45. MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering

    cs.SE 2026-01 conditional novelty 7.0 of 10

    An automated multi-agent pipeline constructs and reuses executable Docker environments for verifiable software-engineering tasks across 10 languages, and fine-tuning on its 3,005-task dataset improves several code models.

  46. A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents

    cs.AI 2025-12 unverdicted novelty 7.0 of 10

    A new benchmark of 40 scenarios finds state-of-the-art LLMs exhibit outcome-driven constraint violations in 0-62.8% of cases under KPI pressure, with no consistent safety gains across model generations.

  47. SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios

    cs.SE 2025-12 unverdicted novelty 7.0 of 10

    SWE-EVO shows GPT-5.4 with OpenHands reaching only 25% success on complex multi-file evolution tasks versus 72.8% on SWE-Bench Verified, and introduces Fix Rate as a partial-progress metric.

  48. Dynamic Tool Dependency Retrieval for Lightweight Function Calling

    cs.LG 2025-12 unverdicted novelty 7.0 of 10

    DTDR dynamically retrieves relevant tools by modeling dependencies from demonstrations and conditioning on the evolving agent plan, improving function calling success rates by 23-104% over static retrievers across benchmarks.

  49. MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning

    cs.CL 2025-11 unverdicted novelty 7.0 of 10

    MemSearcher trains LLMs to manage compact memory in multi-turn searches via multi-context GRPO for end-to-end RL, outperforming ReAct-style baselines with stable token counts.

  50. SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents

    cs.CR 2025-10 unverdicted novelty 7.0 of 10

    SecureWebArena is a new benchmark suite for holistic security evaluation of LVLM-based web agents using diverse simulated environments, attack taxonomies, and multi-layered failure analysis across reasoning, behavior,...

  51. Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data

    cs.LG 2025-09 unverdicted novelty 7.0 of 10

    Reasoning LLMs with minimal tools for tree construction and analysis induce decision trees that outperform CART, compete with ensembles on low-resource tabular data, and provide human-readable reasoning traces.

  52. Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?

    cs.CL 2026-07 conditional novelty 6.5 of 10

    CreditCardQA shows LLMs err mainly on credit-card contractual conditions and comparisons, not arithmetic, with Program-of-Thought narrowing open–closed model gaps.

  53. K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos

    cs.CV 2026-07 accept novelty 6.5 of 10

    Frontier multimodal LLMs reach only ~40% accuracy on long-horizon canine video reasoning, well below human performance, despite bias-mitigated automated dataset construction.

  54. Spatio-Temporal Attention Graph Neural Network: Explaining Causalities With Attention

    cs.LG 2026-03 conditional novelty 6.5 of 10

    A workflow-aligned diagnostic agent that acquires evidence interactively and improves via retrieved diagnostic cognition primitives reaches ~90% accuracy and double-digit gains over baselines on MIMIC-CDM and an exter...

  55. GANDR: Claim Auditing for Verifiable Legal Answer Generation

    cs.CL 2026-09 conditional novelty 6.0 of 10

    GANDR, a drafter-plus-critic system with a deterministic citation gate and a per-claim audit trace, reached 70.8% strict citation-correct accuracy on a 185-item legal benchmark, 11.3 points above the strongest control...

  56. Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

    cs.LG 2026-09 conditional novelty 6.0 of 10

    CANOPY, a minimalist protocol using larger group sizes, strictly on-policy updates, KL anchoring, and token-level loss, shows that outcome-only RL can suffice for long-horizon interactive coding agents.

  57. Apodex 1.1: Scaling Agentic Intelligence for Complex Work

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Apodex 1.1 reports that training a general-purpose language model across executable file, search, and code environments plus coordination traces yields frontier-band agentic performance in a 397B model and a competiti...

  58. Benchmarking Composable Compression Techniques in Mixture-of-Experts LLMs

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A systematic benchmark of composable MoE compression shows that expert pruning dominates quality loss and that compression rate alone does not predict runtime or accuracy effects.

  59. Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A two-step framework uses muP width transfer plus a log-log token scaling law to predict the optimal learning rate for a 155B-parameter MoE over 10T tokens from small proxy runs.

  60. When Deep Research Agents Stagnate: Enhancing Reasoning with Retrieval-Aware Agent Control

    cs.IR 2026-08 conditional novelty 6.0 of 10

    A retrieval-aware controller using document novelty, criteria coverage, and query diversity reduces redundant search steps in seven deep research agents while improving or maintaining answer accuracy.

See all 185 Pith citations

Reference graph

Works this paper leans on

52 extracted references · 52 canonical work pages · cited by 185 Pith papers (see all)

  1. [1]

    SemDeDup: Data-efficient learning at web-scale through semantic deduplication

    A. Abbas, K. Tirumala, D. Simig, S. Ganguli, and A. S. Morcos. Semdedup: Data-efficient learning at web-scale through semantic deduplication. arXiv preprint arXiv:2303.09540, 2023

  2. [2]

    C. An, Z. Xie, X. Li, L. Li, J. Zhang, S. Gong, M. Zhong, J. Xu, X. Qiu, M. Wang, and L. Kong. Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models, 2025

  3. [3]

    Y . Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, et al. Longbench: A bilingual, multitask benchmark for long context understanding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3119–3137, 2024

  4. [4]

    Y . Bai, S. Tu, J. Zhang, H. Peng, X. Wang, X. Lv, S. Cao, J. Xu, L. Hou, Y . Dong, J. Tang, and J. Li. LongBench v2: Towards deeper understanding and reasoning on realistic long-context multitasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3639–3664, Vienna, Austria, July 202...

  5. [5]

    Bavarian, H

    M. Bavarian, H. Jun, N. Tezak, J. Schulman, C. McLeavey, J. Tworek, and M. Chen. Efficient training of language models to fill in the middle, 2022

  6. [6]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020

  7. [7]

    A. Chen, A. Li, B. Gong, B. Jiang, B. Fei, B. Yang, B. Shan, C. Yu, C. Wang, C. Zhu, et al. Minimax-m1: Scaling test-time compute efficiently with lightning attention. arXiv preprint arXiv:2506.13585, 2025

  8. [8]

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021

Show all 52 references
  1. [9]

    Cheng, Y

    S. Cheng, Y . Bao, Q. Cao, L. Huang, L. Kang, Z. Liu, Y . Lu, W. Zhu, Z. Huang, T. Li, et al. Seed-x: Building strong multilingual translation llm with 7b parameters. arXiv preprint arXiv:2507.13618, 2025

  2. [10]

    Deshpande, V

    K. Deshpande, V . Sirdeshmukh, J. B. Mols, L. Jin, E.-Y . Hernandez-Cardona, D. Lee, J. Kritz, W. E. Primack, S. Yue, and C. Xing. Multichallenge: A realistic multi-turn conversation evalua- tion benchmark challenging to frontier llms. In Findings of the Association for Comput...

  3. [11]

    H. Ding, Z. Wang, G. Paolini, V . Kumar, A. Deoras, D. Roth, and S. Soatto. Fewer truncations improve language modeling. In Proceedings of the 41st International Conference on Machine Learning, pages 11030–11048, 2024

  4. [12]

    Gloeckle, B

    F. Gloeckle, B. Y . Idrissi, B. Rozière, D. Lopez-Paz, and G. Synnaeve. Better & faster large language models via multi-token prediction. arXiv preprint arXiv:2404.19737, 2024

  5. [13]

    D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025

  6. [14]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)

  7. [15]

    Henry, P

    A. Henry, P. R. Dachapally, S. Pawar, and Y . Chen. Query-key normalization for transformers, 2020

  8. [16]

    Hsieh, S

    C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, and B. Ginsburg. Ruler: What’s the real context size of your long-context language models? In First Conference on Language Modeling. 23

  9. [17]

    S. Hu, Y . Tu, X. Han, G. Cui, C. He, W. Zhao, X. Long, Z. Zheng, Y . Fang, Y . Huang, et al. Minicpm: Unveiling the potential of small language models with scalable training strategies. In First Conference on Language Modeling

  10. [18]

    Jaech, A

    A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024

  11. [19]

    N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. In The Thirteenth International Conference on Learning Representations

  12. [20]

    C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770, 2023

  13. [21]

    Jordan, Y

    K. Jordan, Y . Jin, V . Boza, Y . Jiacheng, F. Cecista, L. Newhouse, and J. Bern- stein. Muon: An optimizer for hidden layers in neural networks, 2024. URL https://kellerjordan.github.io/posts/muon, 6

  14. [22]

    Joulin, E

    A. Joulin, E. Grave, P. Bojanowski, and T. Mikolov. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Compu- tational Linguistics: Volume 2, Short Papers, pages 427–431. Association for Computa...

  15. [23]

    A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024

  16. [24]

    J. Liu, J. Su, X. Yao, Z. Jiang, G. Lai, Y . Du, Y . Qin, W. Xu, E. Lu, J. Yan, et al. Muon is scalable for llm training. arXiv preprint arXiv:2502.16982, 2025

  17. [25]

    M. Luo, S. Tan, J. Wong, X. Shi, W. Y . Tang, M. Roongta, C. Cai, J. Luo, L. E. Li, R. A. Popa, and I. Stoica. Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl. https://pretty-radio-b75.notion.site/DeepScaleR-Surpassing-O1-Preview-with-a-1-5B-Model- by-Scaling...

  18. [26]

    S. G. Patil, H. Mao, C. Cheng-Jie Ji, F. Yan, V . Suresh, I. Stoica, and J. E. Gonzalez. The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models. In Forty-second International Conference on Machine Learning, 2025

  19. [27]

    Penedo, H

    G. Penedo, H. Kydlí ˇcek, V . Sabolˇcec, B. Messmer, N. Foroutan, A. H. Kargaran, C. Raffel, M. Jaggi, L. V on Werra, and T. Wolf. Fineweb2: One pipeline to scale them all–adapting pre-training data processing to every language. arXiv preprint arXiv:2506.20920, 2025

  20. [28]

    L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al. Humanity’s last exam. arXiv preprint arXiv:2501.14249, 2025

  21. [29]

    Y . Qin, T. Zhang, Y . Shen, W. Luo, Y . Zhang, Y . Qiao, Z. Zhou, W. Zhang, B. CUI, et al. Sysbench: Can llms follow system message? In The Thirteenth International Conference on Learning Representations, 2024

  22. [30]

    D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y . Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling, 2024

  23. [31]

    Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y . Li, Y . Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024

  24. [32]

    D. Su, K. Kong, Y . Lin, J. Jennings, B. Norick, M. Kliegl, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset. arXiv preprint arXiv:2412.02595, 2024

  25. [33]

    G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. 24

  26. [34]

    K. Team, Y . Bai, Y . Bao, G. Chen, J. Chen, N. Chen, R. Chen, Y . Chen, Y . Chen, Y . Chen, et al. Kimi k2: Open agentic intelligence. arXiv preprint arXiv:2507.20534, 2025

  27. [35]

    T. T.-B. Team. Terminal-bench: A benchmark for ai agents in terminal environments, Apr 2025

  28. [36]

    M. Tian, L. Gao, S. Zhang, X. Chen, C. Fan, X. Guo, R. Haas, P. Ji, K. Krongchon, Y . Li, et al. Scicode: A research coding benchmark curated by scientists. Advances in Neural Information Processing Systems, 37:30624–30650, 2024

  29. [37]

    Touvron, T

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023

  30. [38]

    V odrahalli, S

    K. V odrahalli, S. Ontanon, N. Tripuraneni, K. Xu, S. Jain, R. Shivanna, J. Hui, N. Dikkala, M. Kazemi, B. Fatemi, R. Anil, E. Dyer, S. Shakeri, R. Vij, H. Mehta, V . Ramasesh, Q. Le, E. Chi, Y . Lu, O. Firat, A. Lazaridou, J.-B. Lespiau, N. Attaluri, and K. Olszewska. Michela...

  31. [39]

    F. Wan, W. Shen, S. Liao, Y . Shi, C. Li, Z. Yang, J. Zhang, F. Huang, J. Zhou, and M. Yan. Qwenlong-l1: Towards long-context large reasoning models with reinforcement learning. arXiv preprint arXiv:2505.17667, 2025

  32. [40]

    L. Wang, H. Gao, C. Zhao, X. Sun, and D. Dai. Auxiliary-loss-free load balancing strategy for mixture-of-experts. arXiv preprint arXiv:2408.15664, 2024

  33. [41]

    S. Wang, L. Yu, C. Gao, C. Zheng, S. Liu, R. Lu, K. Dang, X. Chen, J. Yang, Z. Zhang, et al. Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for llm reasoning. arXiv preprint arXiv:2506.01939, 2025

  34. [42]

    X. Wang, B. Li, Y . Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y . Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y . Shao, N. Muennighoff, Y . Zhang, B. Hui, J. Lin, R. Brennan, H. Peng, H. Ji, and G. Neubig. Openhands: An open platform for AI software de...

  35. [43]

    Y . Wang, X. Ma, G. Zhang, Y . Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems, 37:95266–95290, 2024

  36. [44]

    J. Wei, N. Karina, H. W. Chung, Y . J. Jiao, S. Papay, A. Glaese, J. Schulman, and W. Fedus. Measuring short-form factuality in large language models, 2024

  37. [45]

    J. Wei, Z. Sun, S. Papay, S. McKinney, J. Han, I. Fulford, H. W. Chung, A. T. Passos, W. Fedus, and A. Glaese. Browsecomp: A simple yet challenging benchmark for browsing agents. arXiv preprint arXiv:2504.12516, 2025

  38. [46]

    Z. Xi, Y . Ding, W. Chen, B. Hong, H. Guo, J. Wang, D. Yang, C. Liao, X. Guo, W. He, et al. Agentgym: Evolving large language model-based agents across diverse environments. arXiv preprint arXiv:2406.04151, 2024

  39. [47]

    A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025

  40. [48]

    S. Yao, N. Shinn, P. Razavi, and K. Narasimhan. tau-bench: A benchmark for tool-agent-user interaction in real-world domains. arXiv preprint arXiv:2406.12045, 2024

  41. [49]

    Q. Yu, Z. Zhang, R. Zhu, Y . Yuan, X. Zuo, Y . Yue, W. Dai, T. Fan, G. Liu, L. Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025

  42. [50]

    A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y . Xu, W. Zheng, X. Xia, et al. Glm-130b: An open bilingual pre-trained model. In The Eleventh International Conference on Learning Representations. 25

  43. [51]

    Zhang, L

    Z. Zhang, L. Lei, L. Wu, R. Sun, Y . Huang, C. Long, X. Liu, X. Lei, J. Tang, and M. Huang. Safetybench: Evaluating the safety of large language models with multiple choice questions. arXiv preprint arXiv:2309.07045, 2023

  44. [52]

    J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y . Luan, D. Zhou, and L. Hou. Instruction- following evaluation for large language models. arXiv preprint arXiv:2311.07911, 2023. 26

Pith tools

Reviewed May 11, 2026 · model on record in the stance chip above.