REVIEW 1 major objections 2 minor 185 cited by
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
T0 review · 1 major / 2 minor · reviewed 2026-05-11 · grok-4.3
Pith's one-line read GLM-4.5 reaches 70.1 percent on TAU-Bench and 91 percent on AIME 24 using an open-source 355B-parameter MoE model with only 32B parameters active at once.
desk verdict GLM-4.5 is a practical open MoE release with competitive ARC benchmark numbers, but the evaluation details need checking before the rankings can be taken as settled. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The hybrid reasoning method that supports both thinking and direct response modes, built inside a Mixture-of-Experts architecture with 355 billion total parameters but only 32 billion activated per token.
What would settle it
Independent re-evaluation of the model on the same benchmark problems using fresh, publicly documented prompts and code, or testing on a new suite of problems created after the training cutoff, would confirm or refute the claimed scores.
Extended reading notes
Core claim
GLM-4.5 achieves strong performance across agentic, reasoning, and coding (ARC) tasks, scoring 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. With much fewer parameters than several competitors, GLM-4.5 ranks 3rd overall among all evaluated models and 2nd on agentic benchmarks.
Load-bearing premise
That the reported benchmark scores reflect genuine capabilities measured through fair, standardized, and uncontaminated evaluations that allow direct comparison to other models.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters. It features a hybrid reasoning method that supports both thinking and direct response modes. The model undergoes multi-stage training on 23T tokens and post-training with expert model iteration and reinforcement learning. GLM-4.5 reports strong results across agentic, reasoning, and coding (ARC) tasks, including 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. It ranks 3rd overall among evaluated models and 2nd on agentic benchmarks despite having fewer parameters than several competitors. A compact variant, GLM-4.5-Air (106B parameters), is also released, with code and models made available at a GitHub repository.
Significance. If the benchmark results hold under verifiable and standardized conditions, the work advances open-source models for agentic and reasoning tasks by demonstrating competitive performance with an efficient MoE architecture and hybrid reasoning. The public release of both the full and compact models, along with code, is a clear strength that enables reproducibility and community follow-up research on ARC capabilities.
major comments (1)
- [Abstract] Abstract: The central performance claims, including the specific scores of 70.1% on TAU-Bench and 64.2% on SWE-bench Verified together with the 3rd overall and 2nd agentic ranking, are presented without any description of the evaluation methodology. Details on agent scaffolding, tool-use protocols, attempt limits, prompting consistency, use of the hybrid thinking mode, and data-contamination controls are required to establish that the results are comparable to those of competing models; their absence undermines confidence in the headline rankings.
minor comments (2)
- [Abstract] The phrase 'expert model iteration' in the abstract is used without definition or reference to a methods section; a brief clarification would improve readability.
- The efficiency claim ('much fewer parameters than several competitors') would be strengthened by explicitly listing the parameter counts of the referenced competing models in a comparison table.
Simulated Author's Rebuttal
We thank the referee for the constructive comments on our manuscript. The feedback highlights an important point about ensuring transparency in the abstract for benchmark results. We address this directly below.
read point-by-point responses
-
Referee: [Abstract] Abstract: The central performance claims, including the specific scores of 70.1% on TAU-Bench and 64.2% on SWE-bench Verified together with the 3rd overall and 2nd agentic ranking, are presented without any description of the evaluation methodology. Details on agent scaffolding, tool-use protocols, attempt limits, prompting consistency, use of the hybrid thinking mode, and data-contamination controls are required to establish that the results are comparable to those of competing models; their absence undermines confidence in the headline rankings.
Authors: We agree that the abstract, constrained by length, omits explicit methodology details, which can affect immediate assessment of comparability. The full manuscript contains sections on evaluation protocols that cover agent scaffolding (standard setups for TAU-Bench and SWE-bench), tool-use protocols, attempt limits, prompting strategies, selective use of the hybrid thinking mode, and data-contamination controls via held-out test sets and decontamination procedures. In the revision, we will expand the abstract with a concise clause summarizing these elements and add cross-references to the detailed methodology sections. This change will improve clarity while preserving the abstract's brevity. We do not believe the core results or rankings require alteration, only better contextualization. revision: yes
Circularity Check
No circularity: purely empirical benchmark reporting
full rationale
The paper describes training GLM-4.5 (355B MoE) on 23T tokens with post-training and RL, then reports measured benchmark scores (70.1% TAU-Bench, 91.0% AIME 24, 64.2% SWE-bench Verified). No mathematical derivations, equations, fitted predictions, or first-principles results exist. Claims rest on independent empirical evaluations with no self-definitional loops, fitted-input predictions, or load-bearing self-citations that reduce the central results to inputs by construction. Standard model-release structure; derivation chain is absent.
Assumptions & free parameters
Cite this review
Pith. "Pith review of GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models." pith.science (2026). https://pith.science/paper/2508.06471
@misc{pith2026250806471,
author = {Pith},
title = {Pith review of: GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/2508.06471}},
note = {Machine review of arXiv:2508.06471}
}
read the original abstract
We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes. Through multi-stage training on 23T tokens and comprehensive post-training with expert model iteration and reinforcement learning, GLM-4.5 achieves strong performance across agentic, reasoning, and coding (ARC) tasks, scoring 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. With much fewer parameters than several competitors, GLM-4.5 ranks 3rd overall among all evaluated models and 2nd on agentic benchmarks. We release both GLM-4.5 (355B parameters) and a compact version, GLM-4.5-Air (106B parameters), to advance research in reasoning and agentic AI systems. Code, models, and more information are available at https://github.com/zai-org/GLM-4.5.
Lean theorems connected to this paper
-
IndisputableMonolith.Foundation.DAlembert.Inevitabilitybilinear_family_forced unclear?
unclearRelation between the paper passage and the cited Recognition theorem.
GLM-4.5 achieves strong performance across agentic, reasoning, and coding (ARC) tasks, scoring 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. With much fewer parameters than several competitors, GLM-4.5 ranks 3rd overall among all evaluated models and 2nd on agentic benchmarks.
-
IndisputableMonolith.Foundation.PhiForcingphi_equation unclear?
unclearRelation between the paper passage and the cited Recognition theorem.
We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes.
-
IndisputableMonolith.Foundation.LedgerForcingconservation_from_balance unclear?
unclearRelation between the paper passage and the cited Recognition theorem.
Through multi-stage training on 23T tokens and comprehensive post-training with expert model iteration and reinforcement learning
What do these tags mean?
- matches
- The paper's claim is directly supported by a theorem in the formal canon.
- supports
- The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
- extends
- The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
- uses
- The paper appears to rely on the theorem as machinery.
- contradicts
- The paper's claim conflicts with a theorem or certificate in the canon.
- unclear
- Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.
Forward citations
Showing 60 of 185 Pith papers that cite this
-
UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing
UltraEP is the first exact-load real-time expert balancer for large-EP MoE training and serving on rack-scale nodes, reaching 94.3% of ideal throughput and 1.49x over no-balancing.
-
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
Sieve dynamically schedules MoE experts across GPU and PIM hardware to handle bimodal token distributions, achieving 1.3x to 1.6x gains in throughput and interactivity over static prior PIM systems on three large models.
-
ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning
ReLibra uses pre-known token-to-expert routing from RL rollouts to perform inter-batch expert reordering and intra-batch replication, delivering up to 1.6x higher throughput than Megatron-LM and 1.2x over oracle-equip...
-
WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild
WildTableBench is the first QA benchmark for naturally occurring table images, where 21 multimodal models were evaluated and only one exceeded 50% accuracy.
-
Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models
User-turn generation reveals that LLMs' interaction awareness is largely decoupled from task accuracy, remaining near zero in deterministic settings even as accuracy scales to 96.8% on GSM8K.
-
SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment
SEAL aligns shared experts in MoE models via DPO LoRA to provide a router-independent safety surface, reducing attack success by up to 60% with negligible capability cost.
-
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
A modular post-training recipe with separate GRPO experts per weakness axis and two-stage SLERP merging produces a single Qwen3-32B model that outperforms a roughly 7x larger baseline on in-house production benchmarks...
-
SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue
The paper introduces SDARE-Bench, a scenario-based benchmark showing that LLMs struggle to detect and appropriately respond to stigma in conversational contexts, particularly in multi-speaker dialogues with group pressure.
-
Controllable Image Captioning with Prompt-Conditioned Scene Rewards
FoCUS uses prompt-conditioned signed weights over scene-graph components to steer image captions toward user-specified semantic emphases (attributes, relations, foreground, background).
-
EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
A benchmark of 150 five-turn UI maintenance tasks shows that even the strongest model completes only 37.3% of five-turn episodes, with tool-grounded tasks proving especially difficult (52.4% adjacent retention).
-
SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
SWE Refactor Bench grades coding agents on 20 whole-repository stack migrations with a three-stage protocol, and finds 5.4% of 520 runs passed all stages, with 13 of 20 tasks unsolved.
-
MidTool: Mid-training Data Synthesis for Agentic Tool Use
A 20.3B-token mid-training mixture for general tool use improves downstream function-calling and agentic tool-use performance on BFCL, tau2-Bench, and MCP-Universe.
-
Synthetic Persona Pretraining: Alignment from Token Zero
Injecting first-person, value-laden reflections into pretraining text improves constitution following, jailbreak resistance, and out-of-distribution moral choices in small language models, with the largest gains when ...
-
TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs
TreeProbe, a 4,719-item Tibetan-medicine benchmark, shows LLMs score 40–60% and systematically drift to TCM or biomedical reasoning.
-
MemTX: Transactional Belief Commit for Stateful Agent Memory
Staging agent-memory writes through a validate-and-commit pipeline with maturity-gated irreversible actions and typed cascading repair yields zero realized downstream harm on five LLM backbones, where eight baselines ...
-
Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
Static SFT and RL training for tool-use agents leads to performance drops under open-world distributional shifts across perception, interaction, reasoning and internalization; perturbation-augmented fine-tuning is pro...
-
MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark
Multilingual LLMs show a reproducible 'Illusion of Cultural Alignment': they can be fluent in a language while lacking the culture's factual knowledge, and confidence, sampling, and retrieval do not fix it.
-
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
Proposes Monotonic Inference Policy Improvement (MIPI) objective and MIPU two-step update framework to address objective misalignment between training and inference policies in LLM reinforcement learning.
-
MacroLens: A Multi-Task Benchmark for Contextual Financial Reasoning under Macroeconomic Scenarios
MacroLens is a point-in-time multi-signal benchmark dataset and seven tasks for evaluating contextual financial reasoning models under macroeconomic scenarios.
-
Is Agent Code Less Maintainable Than Human Code?
Agent-generated code produces up to 13.1% lower follow-on task resolution rates than human code in chained repository-level experiments, with differences linked to behavioral patterns rather than conventional metrics.
-
When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study
Empirical evaluation on LiveCodeBench shows certainty-based RLIF yields early gains followed by output shortening and reasoning collapse, providing no advantage for RLVR initialization on code tasks.
-
LegalWorld: A Life-Cycle Interactive Environment for Legal Agents
LegalWorld is a life-cycle interactive environment modeling Chinese civil litigation as five causally connected stages grounded in 75,309 judgments, paired with LongJud-Bench for cross-stage agent evaluation.
-
daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization
daVinci-kernel trains one LLM to select, use, and summarize reusable GPU-kernel optimization skills in a single RL loop, reaching 37.2%, 70.6%, and 32.2% Fast-1 pass rates on KernelBench Levels 1-3 at 14B.
-
ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm
Proposes COM-as-Action paradigm for deterministic software manipulation, introduces ComCADBench benchmark and ComActor agent that achieves SOTA performance over GUI baselines.
-
LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling
LoopMoE is a looped MoE language model that outperforms matched vanilla MoE on 8 of 9 downstream benchmarks at 3B scale and continues to outperform at 9B scale under strictly controlled budgets.
-
Stateful Visual Encoders for Vision-Language Models
Stateful visual encoders condition each visual representation on prior features, yielding consistent gains on multi-image tasks under supervised finetuning across model sizes and domains.
-
OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs
OVO-S-Bench provides 1680 human-annotated questions on 348 videos to measure streaming spatial intelligence in MLLMs across instantaneous perception, spatiotemporal tracking, spatial simulation, and allocentric mapping.
-
Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs
Moral Trolley Arena shows frontier LLMs produce composite moral preferences that are compressed rather than additive functions of calibrated component act strengths across Moral Foundations Theory.
-
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions
VitaBench 2.0 introduces a benchmark for long-term personalized and proactive agent behavior, with results indicating substantial gaps in current frontier LLMs.
-
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
LatentOmni proposes a latent-space cross-modal reasoning framework that uses feature-level supervision and Omni-Sync Position Embedding to align and synchronize audio-visual latents, supported by a new 35K interleaved...
-
CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning
CopT reverses CoT by eliciting a draft answer first then using continuous-embedding contrastive verification and on-policy thinking to reflect and correct, yielding up to 23% higher accuracy and 57% fewer tokens witho...
-
PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
PRISM benchmark of over 10k pairs shows LLMs have a 41% average drop from code execution success to spatial correctness in programmatic video generation.
-
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
Proposes equivariant optimizer updates matched to layer symmetries for embeddings, SwiGLU MLPs, and MoE routers, with reported gains in validation loss and training stability on several language model architectures.
-
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE
PARAMΔ upcycles dense models to MoE for per-language experts and grafts post-training deltas to enable data-efficient language expansion while preserving original capabilities.
-
BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting
Introduces BacktestBench benchmark with 18k QA pairs across four backtesting tasks and evaluates 23 LLMs via the AutoBacktest multi-agent system.
-
GGBound: A Genome-Grounded Agent for Microbial Life-Boundary Prediction
A genome-conditioned 4B LLM agent predicts microbial life boundaries and matches larger frontier models via token fusion, tool use, and a counterfactual gene-grounding reward.
-
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging
CUDABeaver shows LLM CUDA debuggers often degenerate code for test-passing at the cost of speed, with protocol-aware metrics shifting success rates by up to 40 percentage points.
-
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Masked diffusion language models, not larger autoregressive LLMs, are the better building block for text-based world models in agentic RL, improving rollout fidelity, diversity, and downstream task success.
-
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
SportsTime plus Chain-of-Time Reasoning (temporal-reward GRPO and anchor-observe-infer) modestly lifts open-ended sports VideoQA and step-wise temporal grounding over 4B–8B MLLM baselines.
-
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
Expert upcycling expands MoE models by duplicating experts and continuing pre-training, matching baseline performance while saving 32% GPU hours in 7B-13B experiments.
-
OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset
OmniCompliance-100K supplies 12,985 distinct rules and 106,009 associated real-world cases from 74 multi-domain regulations to benchmark LLM safety and compliance.
-
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents
SPIRAL is a closed-loop think-act-reflect framework using PlanAgent, VideoGenerator, and CriticAgent plus GRPO self-evolution to improve long-horizon action-conditioned video generation, with new dataset and benchmark...
-
EvoESAP: Non-Uniform Expert Pruning for Sparse MoE
EvoESAP uses evolutionary search guided by a speculative-decoding-inspired ESAP metric to discover non-uniform layer-wise sparsity allocations for MoE expert pruning, improving generation accuracy up to 19.6% at 50% sparsity.
-
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise
SQuTR is a large bilingual benchmark of 37,317 synthesized spoken queries under clean/low/medium/high noise, showing that retrieval quality steadily degrades as noise increases.
-
MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering
An automated multi-agent pipeline constructs and reuses executable Docker environments for verifiable software-engineering tasks across 10 languages, and fine-tuning on its 3,005-task dataset improves several code models.
-
A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents
A new benchmark of 40 scenarios finds state-of-the-art LLMs exhibit outcome-driven constraint violations in 0-62.8% of cases under KPI pressure, with no consistent safety gains across model generations.
-
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
SWE-EVO shows GPT-5.4 with OpenHands reaching only 25% success on complex multi-file evolution tasks versus 72.8% on SWE-Bench Verified, and introduces Fix Rate as a partial-progress metric.
-
Dynamic Tool Dependency Retrieval for Lightweight Function Calling
DTDR dynamically retrieves relevant tools by modeling dependencies from demonstrations and conditioning on the evolving agent plan, improving function calling success rates by 23-104% over static retrievers across benchmarks.
-
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
MemSearcher trains LLMs to manage compact memory in multi-turn searches via multi-context GRPO for end-to-end RL, outperforming ReAct-style baselines with stable token counts.
-
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
SecureWebArena is a new benchmark suite for holistic security evaluation of LVLM-based web agents using diverse simulated environments, attack taxonomies, and multi-layered failure analysis across reasoning, behavior,...
-
Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data
Reasoning LLMs with minimal tools for tree construction and analysis induce decision trees that outperform CART, compete with ensembles on low-resource tabular data, and provide human-readable reasoning traces.
-
Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
CreditCardQA shows LLMs err mainly on credit-card contractual conditions and comparisons, not arithmetic, with Program-of-Thought narrowing open–closed model gaps.
-
K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
Frontier multimodal LLMs reach only ~40% accuracy on long-horizon canine video reasoning, well below human performance, despite bias-mitigated automated dataset construction.
-
Spatio-Temporal Attention Graph Neural Network: Explaining Causalities With Attention
A workflow-aligned diagnostic agent that acquires evidence interactively and improves via retrieved diagnostic cognition primitives reaches ~90% accuracy and double-digit gains over baselines on MIMIC-CDM and an exter...
-
GANDR: Claim Auditing for Verifiable Legal Answer Generation
GANDR, a drafter-plus-critic system with a deterministic citation gate and a per-claim audit trace, reached 70.8% strict citation-correct accuracy on a 185-item legal benchmark, 11.3 points above the strongest control...
-
Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents
CANOPY, a minimalist protocol using larger group sizes, strictly on-policy updates, KL anchoring, and token-level loss, shows that outcome-only RL can suffice for long-horizon interactive coding agents.
-
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
Apodex 1.1 reports that training a general-purpose language model across executable file, search, and code environments plus coordination traces yields frontier-band agentic performance in a 397B model and a competiti...
-
Benchmarking Composable Compression Techniques in Mixture-of-Experts LLMs
A systematic benchmark of composable MoE compression shows that expert pruning dominates quality loss and that compression rate alone does not predict runtime or accuracy effects.
-
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
A two-step framework uses muP width transfer plus a log-log token scaling law to predict the optimal learning rate for a 155B-parameter MoE over 10T tokens from small proxy runs.
-
When Deep Research Agents Stagnate: Enhancing Reasoning with Retrieval-Aware Agent Control
A retrieval-aware controller using document novelty, criteria coverage, and query diversity reduces redundant search steps in seven deep research agents while improving or maintaining answer accuracy.
Reference graph
Works this paper leans on
-
[1]
SemDeDup: Data-efficient learning at web-scale through semantic deduplication
A. Abbas, K. Tirumala, D. Simig, S. Ganguli, and A. S. Morcos. Semdedup: Data-efficient learning at web-scale through semantic deduplication. arXiv preprint arXiv:2303.09540, 2023
work page Pith review arXiv 2023
-
[2]
C. An, Z. Xie, X. Li, L. Li, J. Zhang, S. Gong, M. Zhong, J. Xu, X. Qiu, M. Wang, and L. Kong. Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models, 2025
work page 2025
-
[3]
Y . Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, et al. Longbench: A bilingual, multitask benchmark for long context understanding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3119–3137, 2024
work page 2024
-
[4]
Y . Bai, S. Tu, J. Zhang, H. Peng, X. Wang, X. Lv, S. Cao, J. Xu, L. Hou, Y . Dong, J. Tang, and J. Li. LongBench v2: Towards deeper understanding and reasoning on realistic long-context multitasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3639–3664, Vienna, Austria, July 202...
work page 2025
-
[5]
M. Bavarian, H. Jun, N. Tezak, J. Schulman, C. McLeavey, J. Tworek, and M. Chen. Efficient training of language models to fill in the middle, 2022
work page 2022
- [6]
-
[7]
A. Chen, A. Li, B. Gong, B. Jiang, B. Fei, B. Yang, B. Shan, C. Yu, C. Wang, C. Zhu, et al. Minimax-m1: Scaling test-time compute efficiently with lightning attention. arXiv preprint arXiv:2506.13585, 2025
work page Pith review arXiv 2025
-
[8]
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021
work page Pith review arXiv 2021
Show all 52 references
-
[9]
Cheng, Y
S. Cheng, Y . Bao, Q. Cao, L. Huang, L. Kang, Z. Liu, Y . Lu, W. Zhu, Z. Huang, T. Li, et al. Seed-x: Building strong multilingual translation llm with 7b parameters. arXiv preprint arXiv:2507.13618, 2025
2025
-
[10]
Deshpande, V
K. Deshpande, V . Sirdeshmukh, J. B. Mols, L. Jin, E.-Y . Hernandez-Cardona, D. Lee, J. Kritz, W. E. Primack, S. Yue, and C. Xing. Multichallenge: A realistic multi-turn conversation evalua- tion benchmark challenging to frontier llms. In Findings of the Association for Comput...
2025
-
[11]
H. Ding, Z. Wang, G. Paolini, V . Kumar, A. Deoras, D. Roth, and S. Soatto. Fewer truncations improve language modeling. In Proceedings of the 41st International Conference on Machine Learning, pages 11030–11048, 2024
2024
-
[12]
Gloeckle, B
F. Gloeckle, B. Y . Idrissi, B. Rozière, D. Lopez-Paz, and G. Synnaeve. Better & faster large language models via multi-token prediction. arXiv preprint arXiv:2404.19737, 2024
2024
-
[13]
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025
2025 arXiv
-
[14]
Hendrycks, C
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
-
[15]
Henry, P
A. Henry, P. R. Dachapally, S. Pawar, and Y . Chen. Query-key normalization for transformers, 2020
2020
-
[16]
Hsieh, S
C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, and B. Ginsburg. Ruler: What’s the real context size of your long-context language models? In First Conference on Language Modeling. 23
-
[17]
S. Hu, Y . Tu, X. Han, G. Cui, C. He, W. Zhao, X. Long, Z. Zheng, Y . Fang, Y . Huang, et al. Minicpm: Unveiling the potential of small language models with scalable training strategies. In First Conference on Language Modeling
-
[18]
Jaech, A
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024
2024 arXiv
-
[19]
N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. In The Thirteenth International Conference on Learning Representations
-
[20]
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770, 2023
2023 arXiv
-
[21]
Jordan, Y
K. Jordan, Y . Jin, V . Boza, Y . Jiacheng, F. Cecista, L. Newhouse, and J. Bern- stein. Muon: An optimizer for hidden layers in neural networks, 2024. URL https://kellerjordan.github.io/posts/muon, 6
2024
-
[22]
Joulin, E
A. Joulin, E. Grave, P. Bojanowski, and T. Mikolov. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Compu- tational Linguistics: Volume 2, Short Papers, pages 427–431. Association for Computa...
2017
-
[23]
A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024
2024 arXiv
-
[24]
J. Liu, J. Su, X. Yao, Z. Jiang, G. Lai, Y . Du, Y . Qin, W. Xu, E. Lu, J. Yan, et al. Muon is scalable for llm training. arXiv preprint arXiv:2502.16982, 2025
2025
-
[25]
M. Luo, S. Tan, J. Wong, X. Shi, W. Y . Tang, M. Roongta, C. Cai, J. Luo, L. E. Li, R. A. Popa, and I. Stoica. Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl. https://pretty-radio-b75.notion.site/DeepScaleR-Surpassing-O1-Preview-with-a-1-5B-Model- by-Scaling...
2025
-
[26]
S. G. Patil, H. Mao, C. Cheng-Jie Ji, F. Yan, V . Suresh, I. Stoica, and J. E. Gonzalez. The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models. In Forty-second International Conference on Machine Learning, 2025
2025
-
[27]
Penedo, H
G. Penedo, H. Kydlí ˇcek, V . Sabolˇcec, B. Messmer, N. Foroutan, A. H. Kargaran, C. Raffel, M. Jaggi, L. V on Werra, and T. Wolf. Fineweb2: One pipeline to scale them all–adapting pre-training data processing to every language. arXiv preprint arXiv:2506.20920, 2025
2025
-
[28]
L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al. Humanity’s last exam. arXiv preprint arXiv:2501.14249, 2025
2025 arXiv
-
[29]
Y . Qin, T. Zhang, Y . Shen, W. Luo, Y . Zhang, Y . Qiao, Z. Zhou, W. Zhang, B. CUI, et al. Sysbench: Can llms follow system message? In The Thirteenth International Conference on Learning Representations, 2024
2024
-
[30]
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y . Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling, 2024
2024
-
[31]
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y . Li, Y . Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024
2024 arXiv
-
[32]
D. Su, K. Kong, Y . Lin, J. Jennings, B. Norick, M. Kliegl, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset. arXiv preprint arXiv:2412.02595, 2024
2024
-
[33]
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. 24
2023 arXiv
-
[34]
K. Team, Y . Bai, Y . Bao, G. Chen, J. Chen, N. Chen, R. Chen, Y . Chen, Y . Chen, Y . Chen, et al. Kimi k2: Open agentic intelligence. arXiv preprint arXiv:2507.20534, 2025
2025 arXiv
-
[35]
T. T.-B. Team. Terminal-bench: A benchmark for ai agents in terminal environments, Apr 2025
2025
-
[36]
M. Tian, L. Gao, S. Zhang, X. Chen, C. Fan, X. Guo, R. Haas, P. Ji, K. Krongchon, Y . Li, et al. Scicode: A research coding benchmark curated by scientists. Advances in Neural Information Processing Systems, 37:30624–30650, 2024
2024
-
[37]
Touvron, T
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[38]
V odrahalli, S
K. V odrahalli, S. Ontanon, N. Tripuraneni, K. Xu, S. Jain, R. Shivanna, J. Hui, N. Dikkala, M. Kazemi, B. Fatemi, R. Anil, E. Dyer, S. Shakeri, R. Vij, H. Mehta, V . Ramasesh, Q. Le, E. Chi, Y . Lu, O. Firat, A. Lazaridou, J.-B. Lespiau, N. Attaluri, and K. Olszewska. Michela...
2024
-
[39]
F. Wan, W. Shen, S. Liao, Y . Shi, C. Li, Z. Yang, J. Zhang, F. Huang, J. Zhou, and M. Yan. Qwenlong-l1: Towards long-context large reasoning models with reinforcement learning. arXiv preprint arXiv:2505.17667, 2025
2025
-
[40]
L. Wang, H. Gao, C. Zhao, X. Sun, and D. Dai. Auxiliary-loss-free load balancing strategy for mixture-of-experts. arXiv preprint arXiv:2408.15664, 2024
2024
-
[41]
S. Wang, L. Yu, C. Gao, C. Zheng, S. Liu, R. Lu, K. Dang, X. Chen, J. Yang, Z. Zhang, et al. Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for llm reasoning. arXiv preprint arXiv:2506.01939, 2025
2025
-
[42]
X. Wang, B. Li, Y . Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y . Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y . Shao, N. Muennighoff, Y . Zhang, B. Hui, J. Lin, R. Brennan, H. Peng, H. Ji, and G. Neubig. Openhands: An open platform for AI software de...
2025
-
[43]
Y . Wang, X. Ma, G. Zhang, Y . Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems, 37:95266–95290, 2024
2024
-
[44]
J. Wei, N. Karina, H. W. Chung, Y . J. Jiao, S. Papay, A. Glaese, J. Schulman, and W. Fedus. Measuring short-form factuality in large language models, 2024
2024
-
[45]
J. Wei, Z. Sun, S. Papay, S. McKinney, J. Han, I. Fulford, H. W. Chung, A. T. Passos, W. Fedus, and A. Glaese. Browsecomp: A simple yet challenging benchmark for browsing agents. arXiv preprint arXiv:2504.12516, 2025
2025
-
[46]
Z. Xi, Y . Ding, W. Chen, B. Hong, H. Guo, J. Wang, D. Yang, C. Liao, X. Guo, W. He, et al. Agentgym: Evolving large language model-based agents across diverse environments. arXiv preprint arXiv:2406.04151, 2024
2024
-
[47]
A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025
2025 arXiv
-
[48]
S. Yao, N. Shinn, P. Razavi, and K. Narasimhan. tau-bench: A benchmark for tool-agent-user interaction in real-world domains. arXiv preprint arXiv:2406.12045, 2024
2024 arXiv
-
[49]
Q. Yu, Z. Zhang, R. Zhu, Y . Yuan, X. Zuo, Y . Yue, W. Dai, T. Fan, G. Liu, L. Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025
2025 arXiv
-
[50]
A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y . Xu, W. Zheng, X. Xia, et al. Glm-130b: An open bilingual pre-trained model. In The Eleventh International Conference on Learning Representations. 25
-
[51]
Zhang, L
Z. Zhang, L. Lei, L. Wu, R. Sun, Y . Huang, C. Long, X. Liu, X. Lei, J. Tang, and M. Huang. Safetybench: Evaluating the safety of large language models with multiple choice questions. arXiv preprint arXiv:2309.07045, 2023
2023
-
[52]
J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y . Luan, D. Zhou, and L. Hou. Instruction- following evaluation for large language models. arXiv preprint arXiv:2311.07911, 2023. 26
2023 arXiv
Reviewed May 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.