Pith. sign in

REVIEW 2 major objections 5 minor 2 cited by

Gemma 4 open multimodal models leap on STEM, vision, audio and long-context tasks while the 31B dense version ranks as the leading dense open model on human Arena evaluations.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 07:08 UTC pith:R2OUDRBF

load-bearing objection Solid open multimodal release with real efficiency engineering and competitive Arena numbers; attribution of the leap to specific design choices is under-supported by missing ablations. the 2 major comments →

arxiv 2607.02770 v1 pith:R2OUDRBF submitted 2026-07-02 cs.CL cs.AI

Gemma 4 Technical Report

Gemma Team: Sherif El Abd , Vaibhav Aggarwal , Robin Algayres , Alek Andreev , Olivier Bachem , Ian Ballantyne , Cormac Brick , Victor C\u{a}rbune
show 292 more authors
Michelle Casbon Mayank Chaturvedi Victor Cotruta Alice Coucke Phil Culliton Robert Dadashi Lucas Dixon Mohamed Elhawaty Utku Evci Cl\'ement Farabet Johan Ferret Filippo Galgani Sertan Girgin Jean-Bastien Grill Maarten Grootendorst Jiaxian Guo Cassidy Hardin Yanzhang He Steven M. Hernandez Omri Homburger L\'eonard Hussenot Juyeong Ji Armand Joulin Aishwarya Kamath Parnian Kassraie Olivier Lacombe Preethi Lahoti Ga\"el Liu Gus Martins Luciano Martins Tatiana Matejovicova Ramona Merhej Nikola Momchev Sneha Mondal Ryan Mullins Sindhu Raghuram Panyam Shreya Pathak Sarah Perrin Andr\'e Susano Pinto Etienne Pot Ang\'eline Pouget Alexandre Ram\'e Sabela Ramos Douglas Reid David Rim Morgane Rivi\`ere Karsten Roth Louis Rouillard Omar Sanseviero Pier Giuseppe Sessa Shane Settle Danila Sinopalnikov Sara Smoot Piotr Stanczyk Andreas Steiner Lawrence Stewart Ilya Tolstikhin Michael Tschannen Anton Tsitsulin Nino Vieillard Renjie Wu Pingmei Xu Haichuan Yang Edouard Yvinec Li Zhang Joe Zou Nicolas Aagnes Abdelrahman Abdelhamed Shivani Agrawal Shubham Agrawal Ibrahim Alabdulmohsin Jean Baptiste Alayrac Uri Alon Chandramouli Amarnath Ankesh Anand Chrysovalantis Anastasiou Setareh Ariafar Fran\c{c}ois-Xavier Aubet Kyriakos Axiotis Federico Barbero Joelle Barral Alexei Bendebury Urs Bergmann Stanley Bileschi Kat Black Mathieu Blondel Sebastian Borgeaud Arthur Bra\v{z}inskas Ryan Burnell Robert Busa-Fekete Mu Cai Glenn Cameron Charlotte Caucheteux Garima Chadha Jetha Chan Aditya Chawla Blake Jianhang Chen Jesse Chen Lin Chen Xu Chen Derek Cheng Tzu-hsiang Chien Nikolai Chinaev Yi Chou Zhaohui Chu Benjamin Coleman Pooja Consul Sam Conway-Rahman Scott Crowell Dylan Cutler Vivek Dani Samira Daruki Anil Das Daniel Deutsch Nishanth Dikkala Li Ding Qiuhan Ding Shenil Dodhia Konstantin Donhauser Tulsee Doshi Anca Dragan Alex Druinsky Sahil Dua Zoltan Egyed Danielle Eisenbud Daniel Eppens Cindy Fan Bahare Fatemi Yassir Fathullah Vlad Feinberg Milen Ferev Takumi Fujimoto Isaac Galatzer-Levy Jo\~ao Gante Simon Geisler Soham Ghosal Antonious M. Girgis Alec Go Alhaad Gokhale Alex Grills Yiming Gu Pramod Gupta Guru Guruganesh Raia Hadsell Hamza Harkous Jitendra Harlalka Demis Hassabis Anja Hauth Joe Heyward Arian Hosseini Chih-Yang Hsia I-Hung Hsu Xiaopeng Huang Yangsibo Huang Kevin Hui Adrian Hutter Te I Fotis Iliopoulos Advait Jain Ganesh Jawahar Ziwei Ji Qilin Jin Melvin Johnson Kandarp Joshi Arun Kandoor Wang-Cheng Kang Koray Kavukcuoglu Mehran Kazemi Kathleen Kenealy Amr Khalifa Phoebe Kirk Suraj Kothawade Vitaly Kovalev Neel Kovelamudi Adam Kraft Ravin Kumar Harish Kuppam Justin Lannin Chen-Yu Lee Seungji Lee Dmitry Lepikhin Dongdong Li Qiujia Li Valentin Li\'evin Ethan Lin Ziqian Lin Casper Liu Tianlin Liu Tianqi Liu Xin Liu Mayank Lunayach Min Ma Gagan Madan Andrii Maksai Eric Malmi Michal Matuszak Daniel McDuff Gaurav Menghani Daniil Mirylenka Karolis Misiunas Vedant Misra Andreea Mitran Kareem Mohamed Maksim Mukha Eric Noland James O'Donnell Kate Olszewska Bernett Orlando Wanqiong Pan Rina Panigrahy Unnati Parekh Chunjong Park Eric Paskie Liqian Peng Bryce Petrini Slav Petrov Jonas Pfeiffer Bilal Piot Martyna Plomecka Siim Poder Octavio Ponce Arijit Pramanik David Racz Anish Rajan Michelle Ramanovich Anand Rao Marvin Ritter Vitor Rodrigues Evan Rosen Miko{\l}aj Rybi\'nski Noveen Sachdeva Micha\"el E. Sander Rohit Sathyanarayana Sagar Savla Samuel Schmidgall Tal Schuster Benoit Seguin Andrew Sellergren Aliaksei Severyn Izhak Shafran Dhruv Shah Yuan Shangguan Ashish Shenoy Pradeep Shenoy Rakesh Shivanna Pauline Sho Lucas Spangher Wojciech Stokowiec Tim Strother Yao Su Yinghao Sun Mukund Sundararajan Andrea Tacchetti Mor Hazan Taege Pouya Tafti Chetan Tekur Rahul Thapa Madeleine Traverse Lenart Treven Tao Tu Chien Te Tung Petar Veli\v{c}kovi\'c Malini Pooni Venkat Sagar Gubbi Venkatesh Vidya Venkiteswaran Francesco Visin Alex Vitvitskyi Kiran Vodrahalli Weiyi Wang Xin Wang Tris Warkentin Jan Wassenberg John Wieting Lechao Xiao Hao Xu Yuhui Xu Fuzhao Xue Arun Yadav Jun Yan Antoine Yang Lin Yang Ming-Hsuan Yang Ziyu Ying Jae Hyeon Yoo Sajjad Zafar Fred Zhang Jiageng Zhang Jianyi Zhang Xiaofan Zhang Chao Zhao David Zhou Chen Zou
This is my paper
classification cs.CL cs.AI
keywords multimodal language modelsopen-weight modelsmixture-of-expertsthinking modelong-context efficiencyquantization-aware trainingencoder-free architecturespeculative decoding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Gemma 4 is a family of open-weight, natively multimodal language models ranging from roughly 2B to 31B parameters, offered in both dense and mixture-of-experts forms. The paper establishes that these models, released under Apache 2.0, deliver a clear performance jump over the prior Gemma generation across STEM reasoning, vision, audio, and long-context benchmarks, and that the largest members rival much larger frontier open models in blind human preference tests. The suite adds a thinking mode that first emits reasoning traces, an encoder-free 12B architecture that ingests raw image patches and audio chunks, and a set of attention, caching, drafting and quantization choices that shrink memory and raise decoding speed. A sympathetic reader cares because the result puts competitive multimodal reasoning and long-context ability into freely available, on-device-friendly packages.

Core claim

The central claim is that the combination of thinking-mode generation, 5-to-1 local-to-global attention with p-RoPE and key-value reuse, multi-token-prediction drafting, quantization-aware training, and a unified encoder-free backbone for the 12B model produces open multimodal models that substantially outperform Gemma 3 counterparts of similar or larger size and place the 31B dense model at the top of the dense open category on Arena while smaller variants match earlier 27B-class results with far fewer parameters.

What carries the argument

Thinking mode (the model emits an explicit reasoning trace before the final answer) together with the local/global attention ratio, p-RoPE positional encoding, and KV-cache sharing that together cut the global KV footprint by up to 37.5 percent; these are the mechanisms the paper credits for the reasoning and efficiency gains.

Load-bearing premise

The reported gains are produced by the listed architectural and training choices rather than by undisclosed differences in data mixture, filtering, or evaluation protocol.

What would settle it

An independent re-run of the same Arena blind comparisons and of the public STEM/multimodal/long-context suites under identical protocols and without the new design choices (or with thinking mode ablated) that fails to reproduce the claimed ranking and numerical leaps.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Open 2–4 B class models become practical substitutes for earlier 27 B models on many STEM and multimodal tasks.
  • Dense open models at 31 B can occupy the same human-preference tier as much larger mixture-of-experts systems.
  • Encoder-free ingestion of raw patches and audio chunks reduces memory fragmentation and simplifies on-device multimodal stacks.
  • Thinking traces become a standard, controllable feature of open instruction-tuned models rather than a closed-model exclusive.
  • Long-context workloads at 128 k tokens become feasible under tighter KV-cache budgets.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The encoder-free 12 B design may lower the barrier for community fine-tuning of multimodal adapters because there is no frozen external encoder to keep in sync.
  • If thinking mode transfers cleanly to the smallest models, on-device agents could gain chain-of-thought reliability without cloud round-trips.
  • The same local/global and p-RoPE recipe could be ported to other open dense families to obtain similar cache reductions with modest re-training.
  • Public ablations that isolate thinking mode from data scale would clarify how much of the STEM leap is algorithmic versus corpus-driven.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper presents Gemma 4, a family of open-weight natively multimodal decoder-only Transformers (dense E2B/E4B/12B/31B and MoE 26B-A4B) that process text, images and audio. Key design choices include a thinking mode that emits reasoning traces, 4:1/5:1 local-to-global attention with p-RoPE (p=0.25), key-as-value reuse and KV-cache sharing for long-context memory reduction (up to 37.5 %), an encoder-free 12B architecture that linearly projects raw 48 imes48 image patches and 40 ms audio chunks, an autoregressive multi-token-prediction drafter for speculative decoding, and quantization-aware training. High-level pre-training (Jan 2025 cutoff corpus, decontamination) and post-training recipes are given. Extensive automatic and human evaluations (Arena Elo with 95 % CIs, MMLU-Pro, AIME, LiveCodeBench, GPQA, vision suites, CoVoST/FLEURS audio, RULER/LOFT/MTOB long-context) claim large gains over Gemma 3 and parity with far larger open models, with the 31B dense model leading the dense open category on Arena.

Significance. If the reported numbers hold under independent scrutiny, the work supplies a valuable open multimodal baseline family that advances practical efficiency and reasoning at edge-relevant sizes. Concrete engineering contributions include the quantified KV-cache and QAT memory reductions (Table 3), the encoder-free 12B projection design, the MTP drafter, and the thinking-mode integration. Arena leadership of the 31B dense model (Elo 1451) and the observation that E2B roughly matches Gemma-3 27B at ~10 imes fewer parameters are high-utility empirical facts for the community. Release under Apache 2.0 together with quantized checkpoints further increases impact. The absence of full causal ablations is a limitation common to industrial technical reports and does not erase the utility of the released artifacts.

major comments (2)
  1. [Table 5, §4.2] The central claim of a performance leap over Gemma 3 (abstract, §4.2) rests on Tables 5–9, yet evaluation protocols are not matched: Gemma-4 models run in thinking mode while Gemma-3 27B is non-thinking (explicit header note in Table 5); vision results use different maximum token budgets and resizing (Table 6 vs. Pan & Scan). Without non-thinking Gemma-4 numbers, matched-resolution ablations, or data-mixture controls, the contribution of the architectural innovations listed in §2 cannot be isolated from the thinking protocol or later data. This is load-bearing for the attribution language used throughout the abstract and introduction.
  2. [§2.4] §2.4 describes the pre-training corpus only as a “large-scale, diverse collection o cutoff January 2025” with high-level decontamination and safety filtering. Because the leap claim is presented as arising from the design choices of §2, the lack of mixture proportions, decontamination procedure details, or any ablation that holds data fixed while toggling architecture leaves the causal story under-supported. A short caveat or additional controlled experiment would strengthen the manuscript.
minor comments (5)
  1. [Table 1] Table 1 and the surrounding text should more prominently flag that E2B/E4B “effective” parameter counts exclude the large per-layer embeddings; the distinction is easy to miss when comparing against other open models.
  2. [Figure 1] Figure 1 (MTP drafter) would benefit from explicit dimension annotations on the cross-attention and embedder blocks so that the claimed decoding-speed advantage can be verified at a glance.
  3. [Tables 5–9] Most automatic metrics lack error bars or multiple-run statistics (Arena is the exception). Adding even simple standard deviations or bootstrap intervals for the key STEM and long-context numbers would improve interpretability.
  4. [§2.1, Algorithm 1] Algorithm 1 and Figure 2 are clear, yet the precise mapping from N_max values (70 o 1120) to the final soft-token counts after 3 imes3 pooling is left implicit; a short formula or extra column would help re-implementers.
  5. [References] A few concurrent 2026 citations (e.g., Kayyam et al.) appear before their public availability; a note on arXiv versioning would avoid confusion for readers.

Circularity Check

0 steps flagged

No significant circularity: empirical model-release report whose performance claims rest on external community benchmarks, not on self-defined or fitted quantities.

full rationale

Gemma 4 is an architectural and training technical report. Its central claims (leap on STEM/multimodal/long-context suites; Arena Elo leadership for the 31B dense model; E2B roughly matching Gemma-3 27B at ~10 imes fewer parameters) are supported by tabulated scores on external, community-defined suites (Arena, MMLU-Pro, AIME, LiveCodeBench, GPQA, RULER, CoVoST, FLEURS, MMMU-Pro, etc.). These metrics are not redefined in terms of any parameter fitted inside the paper, nor are they obtained by construction from the listed design choices (thinking mode, local-global attention ratios, p-RoPE, key-as-value reuse, encoder-free projections, MTP drafter, QAT). Self-citations appear only for architectural lineage (prior Gemma reports) and are not load-bearing for the new numbers. There is no uniqueness theorem, no ansatz smuggled via self-citation, no renaming of a known empirical pattern as a derived result, and no equation that reduces a claimed prediction to its own input. Attribution of gains to specific innovations is under-supported by missing ablations (a correctness/attribution concern, not circularity). The derivation chain is therefore self-contained against external benchmarks; circularity score is zero.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

As an empirical systems paper the load-bearing premises are standard Transformer design choices, external benchmark definitions, and the unstated but conventional assumption that the (undisclosed) training mixture plus the listed architectural modifications produce the observed gains. No new physical entities or free parameters fitted to a scientific constant are introduced; free parameters are the usual architectural hyper-parameters and the decision to freeze encoders.

free parameters (4)
  • local-to-global attention ratio = 4:1 / 5:1
    Fixed at 4:1 (E2B) or 5:1 (others); chosen by design rather than derived, directly affects KV-cache size claims.
  • p-RoPE fraction p = 0.25
    Set to 0.25 on global layers; an engineering choice that reduces global KV footprint by the claimed 37.5 %.
  • max vision tokens N_max = up to 1120
    Discrete set {70,140,280,560,1120}; controls resolution and compute, selected rather than derived.
  • MTP drafter depth and width = 4 layers, dim 256/1024
    4-layer Transformer with model dim 256/1024; sized for acceptance-rate vs. overhead trade-off.
axioms (4)
  • domain assumption Decoder-only Transformer with RMSNorm, QKNorm, and the listed attention patterns is a sufficient backbone for multimodal next-token prediction.
    Standard in modern LLMs; invoked throughout Section 2 without re-derivation.
  • domain assumption Frozen vision/audio encoders (or their lightweight projections) plus continuous embeddings preserve enough information for the LLM to solve the reported multimodal tasks.
    Stated in Sections 2.1–2.3; no proof that freezing is optimal, only empirical results.
  • domain assumption Public benchmarks (Arena Elo, MMLU-Pro, AIME, CoVoST, RULER, etc.) are valid proxies for the capabilities claimed.
    Used as the sole quantitative evidence in Sections 4.1–4.2.
  • ad hoc to paper Decontamination and safety filtering of the January 2025 cutoff corpus remove benchmark leakage and harmful content sufficiently for the reported scores to be meaningful.
    Asserted in Sections 2.4 and 5; no public verification set or leakage statistics provided.
invented entities (2)
  • Unified encoder-free 12B architecture independent evidence
    purpose: Replace separate 550 M vision + 305 M audio encoders with direct patch/chunk projections into the LLM embedding space to reduce memory fragmentation.
    Introduced in Section 2.3; performance tables show it remains competitive, but the entity itself is a new design choice without independent theoretical derivation.
  • Autoregressive multi-token prediction (MTP) drafter head with cross-attention to main-model KVs independent evidence
    purpose: Enable speculative decoding of arbitrary draft length without separate prefill.
    Described in Section 2.6 and Figure 1; acceptance rates are claimed but not tabulated in detail.

pith-pipeline@v1.1.0-grok45 · 21721 in / 3287 out tokens · 31338 ms · 2026-07-12T07:08:37.948045+00:00 · methodology

0 comments
read the original abstract

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. COBS: Cumulant Order Block Sparse Attention

    cs.LG 2026-07 conditional novelty 7.0

    A compressed within-block key covariance lets block-sparse attention recover most of dense long-context retrieval quality at roughly first-order selector traffic.

  2. Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems

    cs.LG 2026-07 conditional novelty 6.0

    AdaPrefix-GRPO treats solution-prefix length as a feedback controller targeting 50% rollout success rate during GRPO training, then anneals to zero prefix, yielding 1.6–2.1× accuracy gains over vanilla GRPO at matched...

Reference graph

Works this paper leans on

23 extracted references · 20 linked inside Pith · cited by 2 Pith papers

  1. [1]

    J. L. Ba, J. R. Kiros, and G. E. Hinton. Layer nor- malization.arXiv preprint arXiv:1607.06450,

  2. [2]

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabili- ties.arXiv preprint arXiv:2507.06261,

    Gemini Team. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabili- ties.arXiv preprint arXiv:2507.06261,

  3. [3]

    Gemma: Open models based on gemini research and technology, 2024a

    Gemma Team. Gemma: Open models based on gemini research and technology, 2024a. Gemma Team. Gemma 2: Improving open lan- guage models at a practical size.arXiv preprint arXiv:2408.00118, 2024b. Gemma Team. Gemma 3: Technical report.arXiv preprint arXiv:2503.19786, 2025a. Gemma Team. Gemma 3n. https://deepmi nd.google/models/gemma/gemma-3n/ , 2025b. Google ...

  4. [4]

    Gulati, J

    A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, et al. Conformer: Convolution-augmented transformer for speech recognition.arXiv preprint arXiv:2005.08100,

  5. [5]

    Henry, P

    A. Henry, P. R. Dachapally, S. S. Pawar, and Y. Chen. Query-key normalization for trans- formers. InFindings of the Association for Computational Linguistics: EMNLP 2020, pages 4246–4253,

  6. [6]

    Hsieh, S

    C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, Y. Zhang, and B. Ginsburg. Ruler: What’s the real context size of your 9 Gemma 4 Technical Report long-context language models?arXiv preprint arXiv:2404.06654,

  7. [7]

    N. Jain, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica. Live- codebench: Holistic and contamination free evaluation of large language models for code. InInternational Conference on Learning Repre- sentations, volume 2025, pages 58791–58831,

  8. [8]

    Kayyam, A

    A. Kayyam, A. M. Gopal, and M. A. Lewis. Do transformers need three projections? system- atic study of qkv variants.arXiv preprint arXiv:2606.04032,

  9. [9]

    Kazemi, B

    M. Kazemi, B. Fatemi, H. Bansal, J. Palowitch, C.Anastasiou, S.V.Mehta, L.K.Jain, V.Aglietti, D. Jindal, P. Chen, et al. Big-bench extra hard. arXiv preprint arXiv:2502.19187,

  10. [10]

    M. A. Merrill, A. G. Shaw, N. Carlini, B. Li, H. Raj, I. Bercovich, L. Shi, J. Y. Shin, T. Wal- she, E. K. Buchanan, et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces.arXiv preprint arXiv:2601.11868,

  11. [11]

    Openai o1 system card.arXiv preprint arXiv:2412.16720,

    OpenAI. Openai o1 system card.arXiv preprint arXiv:2412.16720,

  12. [12]

    L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al. Humanity’s last exam.arXiv preprint arXiv:2501.14249,

  13. [13]

    D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bow- man. Gpqa: A graduate-level google-proof q&a benchmark.ArXiv, abs/2311.12022,

  14. [14]

    N. Shazeer. Fast transformer decoding: One write- head is all you need.CoRR, abs/1911.02150,

  15. [15]

    K. Team, T. Bai, Y. Bai, Y. Bao, S. Cai, Y. Cao, Y. Charles, H. Che, C. Chen, G. Chen, et al. Kimi k2. 5: Visual agentic intelligence.arXiv preprint arXiv:2602.02276,

  16. [16]

    Q. Team. Qwen3. 5-omni technical report.arXiv preprint arXiv:2604.15804,

  17. [17]

    Vodrahalli, S

    K. Vodrahalli, S. Ontanon, N. Tripuraneni, K. Xu, S. Jain, R. Shivanna, J. Hui, N. Dikkala, M. Kazemi, B. Fatemi, et al. Michelangelo: Long context evaluations beyond haystacks via latent structure queries.arXiv preprint arXiv:2409.12640,

  18. [18]

    B. Xiao, B. Xia, B. Yang, B. Gao, B. Shen, C. Zhang, C. He, C. Lou, F. Luo, G. Wang, et al. Mimo-v2-flash technical report.arXiv preprint arXiv:2601.02780,

  19. [19]

    A. Xu, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong, C. Ling, et al. Deepseek-v4: Towards highly efficient million- token context intelligence.arXiv preprint arXiv:2606.19348,

  20. [20]

    A. Zeng, X. Lv, Z. Hou, Z. Du, Q. Zheng, B. Chen, D. Yin, C. Ge, C. Huang, C. Xie, et al. Glm-5: from vibe coding to agentic engineering.arXiv preprint arXiv:2602.15763,

  21. [21]

    Zhang, W

    Y. Zhang, W. Han, J. Qin, Y. Wang, A. Bapna, Z. Chen, N. Chen, B. Li, V. Axelrod, G. Wang, et al. Google usm: Scaling automatic speech recognition beyond 100 languages.arXiv preprint arXiv:2303.01037,

  22. [22]

    J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou. Instruction- following evaluation for large language models. arXiv preprint arXiv:2311.07911,

  23. [23]

    11 Gemma 4 Technical Report Y. Zuo, S. Qu, Y. Li, Z. Chen, X. Zhu, E. Hua, K. Zhang, N. Ding, and B. Zhou. Medx- pertqa: Benchmarking expert-level medical reasoning and understanding.arXiv preprint arXiv:2501.18362,