Pith. sign in

REVIEW 3 major objections 2 minor 2 cited by

ReQuestNet: A Foundational Learning model for Channel Estimation

T0 review · 3 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read ReQuestNet claims a single neural estimator beats genie MMSE channel estimation by up to 10 dB in 5G systems.

desk verdict The abstract describes a potentially important neural channel estimator, but the visible full text is an unrelated LLM-tool-use paper, so the claims are unverifiable as submitted. read the letter →

arxiv 2508.08790 v1 pith:XTM3AZ2I submitted 2025-08-12 eess.SP cs.AIstat.ML

classification eess.SPcs.AIstat.ML
keywords channelestimation5GMMSEneuralnetworkMIMODMRSresourceblockgroupsprecodingblindness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes ReQuestNet, a neural architecture that estimates wireless channels directly from reference signals, and claims it outperforms the genie MMSE estimator—the classic linear benchmark that knows the true channel statistics—by up to 10 dB at high SNR. A single trained model is said to handle variable numbers of resource blocks, dynamic transmit layers, different PRG bundling sizes, and different DMRS patterns, while jointly estimating MIMO layers even when the transmitter's precoding is unknown. The practical bet is that one learned estimator can replace the pipeline of per-configuration MMSE filters, which need accurate second-order statistics and knowledge of overheads. If true, this would simplify 5G receivers and generalize across configurations without retraining.

What carries the argument

The two-stage architecture of CoarseNet followed by RefinementNet. CoarseNet estimates the channel per PRG per stream using the available DMRS; RefinementNet applies learned correlation structure across precoded PRGs and across MIMO spatial dimensions, which is what lets the model handle varying PRG bundling sizes, transmit layers, and unknown precoding with one set of weights.

What would settle it

Evaluate ReQuestNet on a standardized channel model or measured dataset whose delay-Doppler profile, mobility, and SNR range differ from the training simulation, and compare end-to-end against genie MMSE computed with the exact covariance and noise variance. If the reported 10 dB gain at high SNR shrinks to statistical noise, or if the baseline is found to use estimated rather than true statistics, the central claim is refuted.

Watch

Extended reading notes

Core claim

ReQuestNet is a two-stage neural estimator. CoarseNet produces a first channel estimate per PRG and per transmit-receive stream; RefinementNet then improves that estimate by learning correlations across differently precoded PRG bundles and across MIMO spatial dimensions. The paper's central claim is that this architecture, trained once, outperforms genie MMSE across a wide range of channel conditions and delay-Doppler profiles, with gains up to 10 dB at high SNR, and that it generalizes to unseen channel profiles while remaining blind to the precoding and to the specific DMRS/resource-block configuration.

Load-bearing premise

The entire performance claim rests on the simulated channel data used for training and evaluation—the paper does not describe that generative model in the abstract—so the gains are only as real as the simulation's fidelity to deployed channels, and the genie MMSE baseline must be implemented with true statistics for the comparison to be meaningful.

Editorial extensions

If this is right

  • A single trained ReQuestNet can serve multiple 5G configurations—variable RB count, transmit layers, PRG bundling size, and DMRS patterns—removing per-configuration retraining or filter design.
  • Receivers can estimate channels from DMRS alone, without relying on other reference signals or on knowledge of the precoder.
  • Up to 10 dB SNR gains over genie MMSE at high SNR are claimed across channel conditions and delay-Doppler profiles, with generalization to unseen profiles.
  • Joint handling of MIMO layers and differently precoded PRGs is claimed to convert cross-layer and inter-PRG correlation into estimation gain rather than interference.
  • The approach positions learned estimators as a replacement for optimal-linear (MMSE) estimation in physical-layer pipelines.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the simulated training prior matches deployment, ReQuestNet could make channel estimation robust to inaccurate or stale covariance information, a known weakness of MMSE in fast-varying channels.
  • The per-PRG-then-refine design suggests a general recipe for other variable-grid estimation tasks in wireless systems, such as CSI feedback or beam management, where local estimates are refined with global correlation structure.
  • A rigorous test of the 10 dB claim would compare against genie MMSE on standardized channel models and measured data; the gain is meaningful only if the baseline is implemented with the true channel covariance and noise variance.
  • Because the abstract does not describe the simulation prior, transferability of the gain to real deployments is an open question that only field or standard-model evaluation can settle.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The abstract of the submitted manuscript describes ReQuestNet, a neural architecture for channel estimation in 5G and beyond, claiming a single unified model that handles variable resource-block counts, dynamic transmit layers, PRG bundling sizes, DMRS patterns, and unknown precoding, and that it significantly outperforms genie MMSE by up to 10 dB at high SNR while generalizing to unseen channel profiles. The full text provided, however, is a completely different paper: 'Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments' (arXiv:2508.08791v3), which concerns reinforcement learning for LLM tool invocation. The body contains no mention of ReQuestNet, channel estimation, CoarseNet, RefinementNet, genie MMSE, delay-Doppler profiles, PRG bundling, DMRS, or any related equations, simulations, or baseline implementations. Thus the manuscript as submitted does not contain the technical content that would support the abstract's central claims.

Significance. If the abstract's claims were substantiated, the contribution would be significant: a learned estimator that beats the statistics-aware linear MMSE baseline under precoding blindness and portability across 5G configurations would simplify the channel-estimation pipeline and offer practical gains. The claimed up-to-10 dB improvement over genie MMSE is a strong, falsifiable prediction. However, the submitted full text provides none of the evidence or derivations needed to assess validity. There are no machine-checked proofs, no reproducible code, no parameter-free derivations, and no empirical results in the body of the paper as provided. Consequently, the significance of the work cannot be evaluated at this time; the manuscript is unverdictable as submitted.

major comments (3)
  1. [Full text (entire body)] The submitted full text is arXiv:2508.08791v3, an unrelated paper on LLM tool use. It contains no description of ReQuestNet, CoarseNet, RefinementNet, channel estimation, MMSE, delay-Doppler profiles, DMRS, PRG bundling, or precoding. The abstract's central claim—'ReQuestNet significantly outperforms genie MMSE CE ... achieving up to 10dB gain at high SNRs'—has no supporting architecture, derivation, or simulation results in the manuscript. This is a load-bearing omission: without the technical body, the paper cannot be evaluated for correctness.
  2. [Abstract (simulation claims)] The abstract reports up to 10 dB gain over genie MMSE and generalization to unseen channel profiles, but no simulation methodology is present. There is no description of the channel model, training/test data generation, SNR ranges, number of channel realizations, error bars, or the genie MMSE implementation. Since genie MMSE is the optimal linear estimator under true statistics, the claim of substantial improvement requires exact specification of the simulation prior and baseline; absent these, the headline result is unverifiable and irreproducible.
  3. [Limitations (full text)] The limitations statement in the full text ('our approach primarily focuses on improving tool invocation rather than optimizing the model's underlying reasoning process') pertains to the LLM tool-use paper and has no relevance to ReQuestNet. Thus even the self-reported limitations of the submitted work do not apply to the claimed contribution, further confirming that the body of the paper does not match its abstract.
minor comments (2)
  1. [Title/abstract vs. full text] The abstract (arXiv:2508.08790) and the full text (arXiv:2508.08791v3) have different titles, author lists, and arXiv identifiers, indicating a submission mismatch. The authors should verify the correct manuscript is uploaded.
  2. [Text rendering] The full text contains numerous garbled characters (e.g., the reward formula in Section 3.2) and missing figure labels. These presentation issues would need correction even in the intended manuscript.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity analyzable: the provided full text is an unrelated arXiv paper, so the ReQuestNet derivation chain cannot be assessed.

full rationale

The abstract describes ReQuestNet, a learned channel estimator claimed to outperform genie MMSE by up to 10 dB and to generalize across channel profiles, PRG bundling sizes, and transmit-layer allocations. However, the supplied full text is arXiv:2508.08791v3, 'Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments,' which contains no ReQuestNet architecture, no CoarseNet/RefinementNet equations, no channel model or delay-Doppler simulation setup, no precoding model, and no genie MMSE baseline implementation. There is therefore no derivation chain to walk and no equation that can be exhibited as reducing to an input. The visible limitations statement ('our approach primarily focuses on improving tool invocation rather than optimizing the model's underlying reasoning process') belongs to the unrelated tool-use paper and does not comment on ReQuestNet. Absence of evidence for the headline claim is a verifiability and manuscript-integrity concern, not circularity: the provided text does not permit the claim to be checked, but also does not show that any prediction is equivalent to its inputs by construction or that any load-bearing step rests on a self-citation. Under the hard rule that circularity must be demonstrated by quoted reduction, no circular step is identifiable. The honest finding is therefore no significant circularity (score 0), with the caveat that the ReQuestNet claims remain unverifiable from this submission.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

No new physical entities (particles, forces, dimensions) are introduced; ReQuestNet is a software architecture, not an invented physical entity, and no falsifiable physical handle is claimed beyond wireless simulation results. The free parameters are the network weights and undisclosed training hyperparameters; the axioms are the simulation-to-deployment transfer assumption, the DMRS availability assumption, and the correct-implementation assumption for the genie MMSE baseline.

free parameters (2)
  • Trained network weights and architecture hyperparameters = not reported
    ReQuestNet's performance is obtained by fitting network weights to simulated channel data; architecture choices (layer counts, equivariance group, loss weighting) are not given in the abstract, so the fitted content behind the 10 dB claim is not auditable.
  • Loss weighting between CoarseNet and RefinementNet objectives = not reported
    The abstract does not state the training loss or its weighting; any such weights are fitted hyperparameters that affect the reported gain.
assumptions (3)
  • domain assumption Simulated delay-Doppler and channel profiles are representative of the deployment channels where the 10 dB gain is claimed
    The abstract's evaluation is simulation-based ('wide range of channel conditions, delay-Doppler profiles'); the transfer of learned estimators from simulation to field is assumed.
  • domain assumption Receiver has access to the assumed DMRS/reference-signal structure and no other reference signals
    The abstract claims independence from other reference signals and dependence on DMRS patterns; the availability and correct patterning of DMRS is an input assumption.
  • domain assumption Genie MMSE baseline was given correct channel statistics and noise statistics
    The headline comparison is against genie MMSE; the comparison is only informative if the genie baseline is correctly implemented with true second-order statistics, which the abstract does not document.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ReQuestNet: A Foundational Learning model for Channel Estimation." pith.science (2026). https://pith.science/paper/XTM3AZ2I

@misc{pith2026250808790,
  author       = {Pith},
  title        = {Pith review of: ReQuestNet: A Foundational Learning model for Channel Estimation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XTM3AZ2I}},
  note         = {Machine review of arXiv:2508.08790}
}
read the original abstract

In this paper, we present a novel neural architecture for channel estimation (CE) in 5G and beyond, the Recurrent Equivariant UERS Estimation Network (ReQuestNet). It incorporates several practical considerations in wireless communication systems, such as ability to handle variable number of resource block (RB), dynamic number of transmit layers, physical resource block groups (PRGs) bundling size (BS), demodulation reference signal (DMRS) patterns with a single unified model, thereby, drastically simplifying the CE pipeline. Besides it addresses several limitations of the legacy linear MMSE solutions, for example, by being independent of other reference signals and particularly by jointly processing MIMO layers and differently precoded channels with unknown precoding at the receiver. ReQuestNet comprises of two sub-units, CoarseNet followed by RefinementNet. CoarseNet performs per PRG, per transmit-receive (Tx-Rx) stream channel estimation, while RefinementNet refines the CoarseNet channel estimate by incorporating correlations across differently precoded PRGs, and correlation across multiple input multiple output (MIMO) channel spatial dimensions (cross-MIMO). Simulation results demonstrate that ReQuestNet significantly outperforms genie minimum mean squared error (MMSE) CE across a wide range of channel conditions, delay-Doppler profiles, achieving up to 10dB gain at high SNRs. Notably, ReQuestNet generalizes effectively to unseen channel profiles, efficiently exploiting inter-PRG and cross-MIMO correlations under dynamic PRG BS and varying transmit layer allocations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diffusion-Based Noise-Adaptive Null-Space Channel Estimation for OFDM Systems

    cs.IT 2026-07 conditional novelty 5.0 of 10

    A diffusion estimator with noise-adaptive null-space correction recovers sparse-pilot OFDM channels at lower NMSE than MMSE, toolbox, DPS, and DMPS baselines on 5G TDL/CDL simulations.

  2. Never Compromise to Vulnerabilities: A Comprehensive Survey on AI Governance

    cs.CR 2025-08 unverdicted novelty 4.0 of 10

    A survey of AI governance proposing three pillars, Intrinsic Security, Derivative Security, and Social Ethics, and identifying generalization, evaluation, and regulatory gaps.

Reference graph

Works this paper leans on

58 extracted references · 9 canonical work pages · cited by 2 Pith papers

  1. [1]

    Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J

    Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. Program synthesis with large language models.CoRR, abs/2108.07732, 2021. URLhttps://arxiv.org/abs/2108.07732

  2. [2]

    Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, K...

  3. [3]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert- Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwi...

  4. [4]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian...

  5. [5]

    MUC-4 evaluation metrics

    Nancy Chinchor. MUC-4 evaluation metrics. In Proceedings of the 4th Conference on Message Understanding, MUC1992, McLean, Virginia, USA, June 16-18, 1992, pages 22–29. ACL, 1992. doi: 10.3115/1072064.1072067. URL https://doi.org/ 10.3115/1072064.1072067

  6. [6]

    Training verifiers to solve math word problems

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021. URL https://ar xiv.org/abs/2110.14168

  7. [7]

    Gemini 2.5: Pushing the fron- tier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,

    Google Gemini Team. Gemini 2.5: Pushing the fron- tier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,

  8. [8]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Deli Chen, Dongjie Ji, ...

Show all 58 references
  1. [9]

    Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings

    Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advancesin Neural Information Processi...

  2. [10]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net,

  3. [11]

    Measuring mathematical problem solving with the MATH dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In Joaquin Vanschoren and Sai-Kit Yeung, editors,Proceedings of the Neural Information Processing ...

  4. [12]

    Tool documentation enables zero-shot tool-usage with large language models

    Cheng-Yu Hsieh, Si-An Chen, Chun-Liang Li, Yasuhisa Fujii, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, and Tomas Pfister. Tool documentation enables zero-shot tool-usage with large language models. CoRR, abs/2308.00675,

  5. [13]

    Reinforce++: An efficient rlhf algorithm with robustness to both prompt and reward models

    Jian Hu, Xibin Wu, Wei Shen, Jason Klein Liu, Zilin Zhu, Weixun Wang, Songlin Jiang, Haoran Wang, Hao Chen, Bin Chen, Weikai Fang, Xianyu, Yu Cao, and Haotian Xu. Reinforce++: An efficient rlhf algorithm with robustness to both prompt and reward models. CoRR, abs/2501.03262, 2...

  6. [14]

    URL https://datasets-benchmarks-pro ceedings.neurips.cc/paper/2021/hash/be83ab 3ecd0db773eb2dc1b0a17836a1-Abstract-round2. html. 10

  7. [15]

    Controlllm: Augment language models with tools by searching on graphs

    Zhaoyang Liu, Zeqiang Lai, Zhangwei Gao, Erfei Cui, Ziheng Li, Xizhou Zhu, Lewei Lu, Qifeng Chen, Yu Qiao, Jifeng Dai, and Wenhai Wang. Controlllm: Augment language models with tools by searching on graphs. In Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten...

  8. [16]

    Inference-time scaling for generalist reward modeling

    Zijun Liu, Peiyi Wang, Runxin Xu, Shirong Ma, Chong Ruan, Peng Li, Yang Liu, and Yu Wu. Inference-time scaling for generalist reward modeling. CoRR, abs/2504.02495, 2025. doi: 10.48550/ARX IV.2504.02495. URL https://doi.org/10.48550/a rXiv.2504.02495

  9. [17]

    GPT-4 technical report

    OpenAI. GPT-4 technical report. CoRR, abs/2303.08774, 2023. doi: 10.48550/ARXIV.2 303.08774. URL https://doi.org/10.48550/arX iv.2303.08774

  10. [18]

    Metatool benchmark for large language models: Deciding whether to use tools and which to use

    Yue Huang, Jiawen Shi, Yuan Li, Chenrui Fan, Siyuan Wu, Qihui Zhang, Yixin Liu, Pan Zhou, Yao Wan, Neil Zhenqiang Gong, and Lichao Sun. Metatool benchmark for large language models: Deciding whether to use tools and which to use. In The Twelfth International Conference on Lear...

  11. [19]

    Toolrl: Reward is all tool learning needs

    Cheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang, Xiusi Chen, Dilek Hakkani-Tür, Gokhan Tur, and Heng Ji. Toolrl: Reward is all tool learning needs. CoRR, abs/2504.13958, 2025. doi: 10.48550 /ARXIV.2504.13958. URL https://doi.org/10.4 8550/arXiv.2504.13958

  12. [20]

    Toolllm: Facilitating large language models to master 16000+ real-world apis

    Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. Toolllm: Facilitating large language models to m...

  13. [21]

    Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Xuanhe Zhou, Yufei Huang, Chaojun Xiao, Chi Han, Yi R. Fung, Yusheng Su, Huadong Wang, Cheng Qian, Runchu Tian, Kunlun Zhu, Shihao Liang, Xingyu Shen, Bokai Xu, Zhen Zhang, Yining Ye, Bowen Li, ...

  14. [22]

    Lundberg, Sameer Singh, Hannaneh Hajishirzi, Luke Zettlemoyer, and Marco Túlio Ribeiro

    Bhargavi Paranjape, Scott M. Lundberg, Sameer Singh, Hannaneh Hajishirzi, Luke Zettlemoyer, and Marco Túlio Ribeiro. ART: automatic multi-step reasoning and tool-use for large language models. CoRR, abs/2303.09014, 2023. doi: 10.48550/ARX IV.2303.09014. URL https://doi.org/10....

  15. [23]

    Lillicrap, Jean- Baptiste Alayrac, Radu Soricut, Angeliki Lazari- dou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew M

    Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy P. Lillicrap, Jean- Baptiste Alayrac, Radu Soricut, Angeliki Lazari- dou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew M. Dai, Katie Millican, Ethan Dyer, ...

  16. [24]

    Toolformer: Language models can teach themselves to use tools

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Har...

  17. [26]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. CoRR, abs/2402.03300, 2024. doi: 10.48550/ARX IV.2402.03300. URL https://doi.org/10...

  18. [27]

    Tool learning with large language models: a survey

    Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong Wen. Tool learning with large language models: a survey. FrontiersComput. Sci., 19(8):198343, 2025. doi: 10.1007/S11704-024-40678-2. URL https: //doi.org/10.1007/s11704-024-40678-2

  19. [28]

    Restgpt: Connecting large language models with real-world applications via restful apis.CoRR, abs/2306.06624,

    Yifan Song, Weimin Xiong, Dawei Zhu, Cheng Li, Ke Wang, Ye Tian, and Sujian Li. Restgpt: Connecting large language models with real-world applications via restful apis.CoRR, abs/2306.06624,

  20. [29]

    Le, Ed H

    Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. Challenging big- bench tasks and whether chain-of-thought can solve them. In Anna Rogers, Jordan L. Boyd- Graber,...

  21. [30]

    Toolalpaca: Generalized tool learning for language models with 3000 simulated cases.CoRR, abs/2306.05301, 2023

    Qiaoyu Tang, Ziliang Deng, Hongyu Lin, Xianpei Han, Qiao Liang, and Le Sun. Toolalpaca: Generalized tool learning for language models with 3000 simulated cases.CoRR, abs/2306.05301, 2023. doi: 10.48550/ARXIV.2306.05301. URL https: //doi.org/10.48550/arXiv.2306.05301

  22. [31]

    Introducing claude 4, 2025

    Anthropic Team. Introducing claude 4, 2025. URL https://www.anthropic.com/news/claude-4

  23. [32]

    Appworld: A controllable world of apps and people for benchmarking interactive coding agents

    Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, and Niranjan Balasubra- manian. Appworld: A controllable world of apps and people for benchmarking interactive coding agents. In Lun-Wei Ku, Andre Martins, and ...

  24. [33]

    Hybridflow: A flexible and efficient RLHF framework

    Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient RLHF framework. In Proceedings of the TwentiethEuropean Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherla...

  25. [34]

    The rise and potential of large language model based agents: a survey

    Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, Rui Zheng, Xiaoran Fan, Xiao Wang, Limao Xiong, Yuhao Zhou, Weiran Wang, Changhao Jiang, Yicheng Zou, Xiangyang Liu, Zhangyue Yin, Shihan Dou, Rongxiang Weng, W...

  26. [35]

    URL https://doi.org/10.48550/arXiv.2306.06624

    doi: 10.48550/ARXIV.2306.06624. URL https://doi.org/10.48550/arXiv.2306.06624

  27. [37]

    Gpt4tools: Teaching large language model to use tools via self-instruction

    Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, and Ying Shan. Gpt4tools: Teaching large language model to use tools via self-instruction. CoRR, abs/2305.18752, 2023. doi: 10.48550/ARX IV.2305.18752. URL https://doi.org/10.48550/a rXiv.2305.18752

  28. [38]

    Narasimhan, and Yuan Cao

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview...

  29. [39]

    Narasimhan

    Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik R. Narasimhan. �-bench: A benchmark for tool-agent-user interaction in real-world domains. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. ...

  30. [40]

    Toolgen: Unified tool retrieval and calling via generation

    Renxi Wang, Xudong Han, Lei Ji, Shu Wang, Timothy Baldwin, and Haonan Li. Toolgen: Unified tool retrieval and calling via generation. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL http...

  31. [41]

    Rotbench: A multi- level benchmark for evaluating the robustness of large language models in tool learning

    Junjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang, SixianLi, GuanyuLi, XiaoranFan, QiZhang, Tao Gui, and Xuanjing Huang. Rotbench: A multi- level benchmark for evaluating the robustness of large language models in tool learning. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nun...

  32. [42]

    Qwen2.5 technical report.CoRR, abs/2412.15115, 2024

    An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Me...

  33. [43]

    A multi-dimensional constraint framework for evaluating and improving instruction following in large language models

    Junjie Ye, Caishuang Huang, Zhuohan Chen, Wenjie Fu, Chenyuan Yang, Leyi Yang, Yilong Wu, Peng Wang, Meng Zhou, Xiaolong Yang, Tao Gui, Qi Zhang, Zhongchao Shi, Jianping Fan, and Xuanjing Huang. A multi-dimensional constraint framework for evaluating and improving instruction ...

  34. [44]

    URLhttps://arxiv.org/abs/2505.09388

  35. [45]

    Tl-training: A task-feature-based framework for training large language models in tool use

    Junjie Ye, Yilong Wu, Sixian Li, Yuming Yang, Zhi- heng Xi, Tao Gui, Qi Zhang, Xuanjing Huang, Peng Wang, Zhongchao Shi, Jianping Fan, and Zhengyin Du. Tl-training: A task-feature-based framework for training large language models in tool use. In Christos Christodoulopoulos, T...

  36. [46]

    Steptool: A step- grained reinforcement learning framework for tool learning in llms

    Yuanqing Yu, Zhefan Wang, Weizhi Ma, Zhicheng Guo, Jingtao Zhan, Shuai Wang, Chuhan Wu, Zhiqiang Guo, and Min Zhang. Steptool: A step- grained reinforcement learning framework for tool learning in llms. CoRR, abs/2410.07745, 2024. doi: 10.48550/ARXIV.2410.07745. URL https: //d...

  37. [47]

    Agent-r: Training language model agents to reflect via iterative self- training

    Siyu Yuan, Zehui Chen, Zhiheng Xi, Junjie Ye, Zhengyin Du, and Jiecao Chen. Agent-r: Training language model agents to reflect via iterative self- training. CoRR, abs/2501.11425, 2025. doi: 10.485 50/ARXIV.2501.11425. URL https://doi.org/10 .48550/arXiv.2501.11425

  38. [48]

    Toolsword: Unveiling safety issues of large language models in tool learning across three stages

    Junjie Ye, Sixian Li, Guanyu Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang, Tao Gui, and Xuanjing Huang. Toolsword: Unveiling safety issues of large language models in tool learning across three stages. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Proceed...

  39. [49]

    Toolqa: A dataset for LLM question answering with external tools

    Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang. Toolqa: A dataset for LLM question answering with external tools. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advancesin Neural Information Processing System...

  40. [50]

    political_figure

    Yuchen Zhuang, Xiang Chen, Tong Yu, Saayan Mitra, Victor S. Bursztyn, Ryan A. Rossi, Somdeb Sarkhel, and Chao Zhang. Toolchain*: Efficient ac- tion space navigation in large language models with a* search. InThe TwelfthInternational Conference on Learning Representations, ICLR...

  41. [52]

    Toolhop: A query-driven benchmark for evaluating large language models in multi-hop tool use

    Junjie Ye, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zehui Chen, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Siyu Yuan, Tao Gui, Qi Zhang, Xuanjing Huang, and Jiecao Chen. Toolhop: A query-driven benchmark for evaluating large language models in multi-hop tool use. In Wanxian...

  42. [54]

    Tooleyes: Fine-grained evaluation for tool learning capabilities of large language models in real-world scenarios

    Junjie Ye, Guanyu Li, Songyang Gao, Caishuang Huang, Yilong Wu, Sixian Li, Xiaoran Fan, Shihan Dou, Tao Ji, Qi Zhang, Tao Gui, and Xuanjing Huang. Tooleyes: Fine-grained evaluation for tool learning capabilities of large language models in real-world scenarios. In Owen Rambow,...

  43. [58]

    Opennovelty: An llm-powered agentic system for verifiable scholarly novelty assessment

    Ming Zhang, Kexin Tan, Yueyuan Huang, Yujiong Shen, Chunchun Ma, Li Ju, Xinran Zhang, Yuhui Wang, Wenqing Jing, Jingyi Deng, Huayu Sha, Binze Hu, Jingqi Tong, Changhao Jiang, Yage Geng, Yuankai Ying, Yue Zhang, Zhangyue Yin, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang, and Xuanj...

  44. [435]

    URLhttps://doi.org/10.1145/3704435

  45. [2017]

    URLhttp://arxiv.org/abs/1707.06347

  46. [2021]

    URL https://openreview.net/forum?i d=d7KBjmI3GmQ

  47. [2023]

    URL https://doi.org/10.48550/arXiv.2308.00675

    doi: 10.48550/ARXIV.2308.00675. URL https://doi.org/10.48550/arXiv.2308.00675

  48. [2024]

    URL https://doi.org/10.18653/v1/2024.acl-l ong.119

    doi: 10.18653/V1/2024.ACL-LONG.119. URL https://doi.org/10.18653/v1/2024.acl-l ong.119

  49. [2025]

    URL https://storage.googleapis.com/dee pmind-media/gemini/gemini_v2_5_report.pdf

  50. [2211]

    Association for Computational Linguistics,

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.