Pith. sign in

REVIEW 3 major objections 6 minor 53 references

Scaling Scientific Discovery Environments for Turn-Level Agentic RL

T0 review · 3 major / 6 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read The paper claims that a 14B open-weight model can reach leading results on hypothesis-driven scientific data analysis if it is trained inside process-verifiable environments that reward each turn producing verifiable analytical evidence, ra

desk verdict A genuinely new pipeline for turn-level RL in scientific discovery, with a clean ablation and honest limitations, but the SOTA headline rests on a 0.5-point margin and a leaky overlap audit. read the letter →

arxiv 2607.28990 v1 pith:EANVAIKU submitted 2026-07-31 cs.AI

classification cs.AI
keywords scientificdiscoveryagentsprocess-verifiableenvironmentshiddenevidenceDAGsturn-levelreinforcementlearningverifier-groundedcreditassignmenthypothesis-drivendataanalysisagenticRLforsciencesynthetictrajectoryfiltering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that the missing ingredient for reliable data-driven scientific-discovery agents is process supervision over real scientific data, not larger models or better prompts. It introduces SciDisco, a three-stage pipeline: SciThèque compiles 1,686 open scientific datasets into sandboxed tasks where a hidden evidence DAG lets a verifier check intermediate progress; DAG-grounded trajectory synthesis produces 5,620 verifier-filtered multi-turn demonstrations for cold-start SFT; and DiscoPO assigns turn-level rewards equal to the increase in verified DAG progress, then applies group-relative policy optimization at the turn level. The reported result is that SciDisco-14B reaches 35.2% average HMS on DiscoveryBench, ahead of stronger proprietary and open-source baselines, and outperforms trajectory-level GRPO on DataSciBench. The authors present this as improved performance on executable hypothesis-driven analysis, and explicitly note that environment-accepted progress is a proxy for scientific quality rather than proof of novel discovery.

What carries the argument

The hidden evidence DAG gj=(Vj,Aj) is the central object: each node is a verifiable scientific state transition typed by a process primitive (data inspection, model fitting, diagnostic checking, robustness analysis), and directed edges encode prerequisites so only frontier nodes are eligible at each turn. SciThèque uses the DAG to define tasks and verifiers; trajectory synthesis uses it to schedule and filter demonstrations; DiscoPO uses it to compute progress potential Φj(C)=|C∩Pj|/|Pj| and turn rewards as potential differences. The same graph thus carries task construction, imitation cold-start, and RL credit assignment.

What would settle it

Take SciDisco-14B and run it on a held-out evaluation set built by independent domain experts who define hypotheses, evidence prerequisites, and verifiers without using the SciThèque template library or reference analyses. If the model's accuracy on that set is no better than the SFT-only checkpoint, then the DiscoPO turn-level signal—rather than verifier content or corpus overlap—is not what produces the reported gains.

Watch

Extended reading notes

Core claim

The central claim is that scientific discovery can be trained as an interactive process if the environment can verify each analytical step, not just the final answer. To that end, SciThèque materializes 1,686 tasks from open datasets, each with a hidden evidence DAG whose nodes are typed process primitives—data setup, EDA, feature construction, model fitting, diagnostics, uncertainty, robustness, and submission—and whose edges enforce a prerequisite order. A verifier accepts a turn only when it produces new evidence for exactly one eligible frontier node, refusing bundled, repeated, premature, or unverifiable actions. DAG-grounded trajectory synthesis uses the DAG to schedule multi-turn demo

Load-bearing premise

The load-bearing premise is that the hidden evidence DAGs and verifier acceptance criteria—built from the authors' template libraries and reference analyses—faithfully measure scientific progress; if they only encode the authors' task-specific heuristics, the training signal and benchmark gains may reflect conformance to those heuristics (or topical overlap with the benchmark data) rather than improved scientific reasoning.

Editorial extensions

If this is right

  • If the reported numbers hold, a 14B open-weight model can outperform much larger closed models on hypothesis formation from scientific datasets, implying process-verifiable environments can substitute for model scale in this regime.
  • Turn-level verifier credit scales directly: the same environments that emit SFT demonstrations also supply RL rewards, so expanding the environment corpus expands the training signal without human trajectory annotation.
  • The ablation ordering SFT < SFT+GRPO < SFT+DiscoPO on DataSciBench indicates that fine-grained process reward, not just outcome verification, drives the largest gains, especially in data-exploration and data-modeling steps.
  • Because the training distribution is hypothesis-driven scientific analysis, the weaker DABStep result shows the gains will not automatically transfer to business-analytics workflows; transfer is selective.
  • The hidden-DAG design makes training auditable: leakage scans and verifier traces can be checked, and failure modes such as premature submission or bundled analyses are explicitly excluded from both SFT and RL signal.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the same hidden-DAG credit mechanism could generalize to other long-horizon agentic domains—code repair, lab automation, literature search—wherever a verifier can define typed, prerequisite-ordered evidence nodes; the authors do not claim this extension.
  • The paper's own limitation section concedes that environment-accepted progress is a proxy, and its overlap audit checks exact file reuse only. A content-level audit of topical similarity between training sources and DiscoveryBench/DataSciBench would be the natural next test of whether the benchmark gains reflect general reasoning or corpus overlap.
  • Because hypotheses are generated from template libraries, the discovery space is bounded by the templates; replacing templates with open-ended hypothesis generation would test whether the pipeline can move beyond template-shaped discoveries toward genuinely novel claims.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents SciDisco, a three-stage post-training pipeline for data-driven scientific discovery agents. It introduces SciThèque, which compiles open scientific datasets and hypothesis templates into 1,686 sandboxed task environments, each with a hidden evidence DAG that defines prerequisite-ordered scientific analysis steps and a verifier that accepts a turn only when it produces evidence for exactly one frontier node. DAG-grounded trajectory synthesis generates 5,620 verifier-filtered multi-turn demonstrations used for SFT. DiscoPO extends GRPO with turn-level rewards computed from increases in the DAG-progress potential, assigning token-level advantages at each interaction turn. The paper reports experiments on DiscoveryBench, DABStep, and DataSciBench, claiming state-of-the-art on DiscoveryBench for SciDisco-14B, and an ablation showing that SFT+DiscoPO outperforms SFT+GRPO on DataSciBench.

Significance. If the empirical claims withstand scrutiny, the framework is a solid contribution: it provides a reusable, executable, and verifiable training substrate for scientific discovery agents, with a clear mechanism for process supervision and turn-level credit assignment. The ablation in Table 2 is clean and supports the central algorithmic claim. Strengths include a detailed environment contract, leakage controls, an explicit benchmark-overlap audit, and a verifier-grounded trajectory synthesis pipeline. However, the headline SOTA claim rests on a 0.5-point margin on DiscoveryBench, a single-run evaluation with unreported judge variance, and an overlap audit that explicitly excludes semantic/template overlap. These methodological weaknesses make the empirical conclusions provisional.

major comments (3)
  1. [§5.1, Table 1, Appendix A] The abstract's 'state-of-the-art' claim rests on a 0.5-point margin over Intern-S1-Pro on DiscoveryBench (35.2 vs 34.7, Table 1). No error bars, multiple seeds, or judge-reliability statistics are reported, and the HMS judge is GPT-5 Mini with no human-agreement measure. Appendix A's overlap audit compares only SHA-256 hashes and normalized source identifiers, and explicitly states it is scoped to exact file reuse while 'broad topical similarity is expected.' Since SciThèque is built from the same public repositories (UCI, OpenML, FRED, CDC, etc.) and the same hypothesis-template space that DiscoveryBench and DataSciBench sample, the reported gains may reflect distributional familiarity rather than improved scientific reasoning. Please either add a semantic/template-level overlap analysis or a held-out source/template evaluation, and temper the SOTA claim accordingly.
  2. [§4.1, Appendix B, Limitations] The training reward is defined by the authors' hidden DAGs and verifier acceptance criteria, and the Limitations section concedes that 'environment-accepted progress is still a proxy for scientific quality.' The 10-expert review in Appendix B inspects verifier specifications, not rollout-level verifier decisions. No empirical evidence is given that accepted steps correlate with human-judged scientific quality or final-answer correctness. This is a correctness risk for the RL signal: the policy could optimize verifier compliance rather than scientific progress. Please report verifier precision/recall on a human-labeled rollout sample and analyze whether accepted intermediate turns predict final score on held-out benchmarks.
  3. [§5.4, Table 1] The DABStep average of 17.8% is far below DeepAnalyze-8B (38.9%) and several proprietary models. The paper attributes this to a training-distribution boundary, which is plausible but not tested. Given the abstract's broad 'scientific data analysis benchmarks' phrase, the authors should either restrict the headline claim to DiscoveryBench/DataSciBench or provide an analysis (e.g., per-task-type breakdown) showing the cause is domain shift rather than a general inability to do multi-step analysis.
minor comments (6)
  1. [Figure 1] The pipeline diagram is dense, and the labels for DiscoPO's advantage calculation are difficult to read. Consider enlarging or splitting into separate figures.
  2. [§4.1] The term 'materialization pool' is used before it is defined; define it at first use.
  3. [Appendix A, Stress variants] The 'Stress variants' are counted separately from aggregate totals, but no further detail is given. A sentence describing what perturbations are applied would help reproducibility.
  4. [Table 2] The VLM column is left empty with a note. Consider removing the column or adding 'n/a' to avoid confusion.
  5. [References] Several 2026 preprints central to the baseline comparisons (e.g., Intern-S1-Pro Team, DeepSeek-V4-Flash) should have stable identifiers or URLs to allow verification. The 'slime' software is cited as a repository without a version.
  6. [Typographical] The conclusion contains 'SciTh‘eque' with a stray quote; spelling of 'materialised/materialized' is inconsistent between appendix and main text.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the training loop uses self-built verifiers, but the headline claims are tested on external benchmarks.

full rationale

Walked the claimed derivation chain: SciThèque builds environments from open datasets, template libraries, hidden evidence DAGs, and verifiers; DAG-grounded synthesis filters SFT trajectories by the same hidden graph; DiscoPO assigns turn-level reward from DAG-progress potential differences (Eqs. 7-8). The reward definition is a training-signal design choice, not a "prediction" that reduces to its own inputs. The central empirical claim—state-of-the-art on hypothesis-driven scientific data analysis—is evaluated on external benchmarks (DiscoveryBench, DABStep, DataSciBench), giving an independent, falsifiable target outside the authors' own environment corpus. No equation in the paper is shown to equal another by construction, and no load-bearing premise rests on a self-citation: the citations to GRPO, PPO, Slime, SGLang, and baseline systems are standard prior work, not an unverified uniqueness claim. The only substantive concern is train/eval distribution overlap: Appendix A's overlap audit compares only SHA-256 hashes and normalized source identifiers and explicitly states that 'broad topical similarity is expected across independently sourced scientific datasets.' That is a generalization/validity risk, not a circular derivation. The paper's own Limitations concede that 'environment-accepted progress is still a proxy for scientific quality,' which further confirms the authors treat the verifier as a proxy rather than as ground truth. Under the required standard—exhibiting a specific reduction of the claimed result to its inputs—no circular step can be identified, so the score is 0.

Assumptions & free parameters 4 free parameters · 4 assumptions · 2 invented entities

The central claim rests on hand-authored environmental components (DAGs, verifiers, templates) rather than fitted constants; these are the paper's main free design choices. No code or corpus is shipped, so the reader cannot audit the verifier's behavior or the distributional overlap.

free parameters (4)
  • Hidden evidence DAG structure (node primitives, prerequisite edges, required nodes N_j)
    Authors design a DAG per task; Eq. 8-9's reward and therefore the training signal depend entirely on this hand-authored structure. No external standard verifies that the DAG matches real scientific practice.
  • Verifier acceptance criteria and hidden tolerances
    The progress rule (exactly one frontier node, evidence fields, tolerance policy) is set by the authors; the paper gives no quantitative validation of verifier precision/recall, only a qualitative expert review.
  • Template library T_m and materialization pool
    The set of hypotheses/types is chosen by the authors; this controls the training distribution and limits generalization, as visible in the DABStep gap.
  • Training hyperparameters (rollout group size 8, clip bounds, etc.)
    Chosen by hand in Table 12; standard for RL but they influence the reported results.
assumptions (4)
  • ad hoc to paper The verifier v_j correctly detects whether a turn produces evidence for exactly one frontier node and whether the final submission is valid.
    No quantitative precision/recall evaluation of the verifier is reported; the expert review (App. B) is qualitative. If the verifier misfires, the turn-level rewards are misaligned.
  • domain assumption The hidden evidence DAG is a faithful model of scientific progress for each task domain.
    The paper's own Limitations state 'environment-accepted progress is still a proxy for scientific quality'; the design assumes that the selected primitives and prerequisites capture what a scientist should do.
  • ad hoc to paper The hash-based benchmark-overlap audit is sufficient to rule out distributional contamination.
    The audit checks exact file hashes and source IDs only; the paper explicitly allows 'broad topical similarity', so style/task overlap with DiscoveryBench could inflate the SOTA claim.
  • standard math Potential-difference rewards (Eq. 8) provide a valid training signal for the POMDP.
    This is standard reward shaping; the potential difference is well-defined and does not change the optimal policy under standard conditions.
invented entities (2)
  • Hidden evidence DAG (g_j)
    purpose: Defines verifiable scientific state transitions and the progress state used for turn-level reward (Eq. 7-9).
    The DAG is constructed by the authors from their template libraries; no external falsifiable handle or benchmark validates it.
  • Process primitives (setup, EDA, feature extraction, model fit, diagnostics, uncertainty, robustness)
    purpose: Typed node constraints in the DAG; define which operations count as progress.
    A hand-chosen taxonomy; no independent evidence that this taxonomy spans meaningful scientific analyses.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Scaling Scientific Discovery Environments for Turn-Level Agentic RL." pith.science (2026). https://pith.science/paper/EANVAIKU

@misc{pith2026260728990,
  author       = {Pith},
  title        = {Pith review of: Scaling Scientific Discovery Environments for Turn-Level Agentic RL},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EANVAIKU}},
  note         = {Machine review of arXiv:2607.28990}
}
read the original abstract

Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains constrained by the lack of process supervised environments over real-world scientific data. This paper introduces SciDisco, a scalable framework for training Scientific Discovery agents in process-verifiable environments. SciTh\`eque compiles hypotheses, datasets, hidden evidence graphs, and verifiers into task environments where analytical progress can be checked during interaction. DAG-grounded trajectory synthesis uses these environments to construct verifier-filtered multi-turn demonstrations. DiscoPO then uses the environment as the source of training signal, assigning turn-level credit to actions that produce verifiable analytical evidence. Experiments show that SciDisco-14B reaches state-of-the-art on hypothesis-driven scientific data analysis benchmarks.

Figures

Figures reproduced from arXiv: 2607.28990 by the authors.

Figure 1
Figure 1. The pipeline of SciDisco: 1) SciThèque builds a scalable collection of data–hypothesis–verifier environ￾ments from scientific datasets; 2) Trajectory Synthesis traverses the evidence DAG through executable analysis primitives and keeps accepted state transitions as multi-turn SFT demonstrations; 3) DiscoPO defines turn-level rewards by increases in verified DAG progress, assigning credit to evidence-producing turns.… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 3 linked inside Pith

  1. [1]

    2025 , url =

    Majumder, Bodhisattwa Prasad and Surana, Harshit and Agarwal, Dhruv and Dalvi Mishra, Bhavana and Meena, Abhijeetsingh and Prakhar, Aryan and Vora, Tirth and Khot, Tushar and Sabharwal, Ashish and Clark, Peter , booktitle =. 2025 , url =

  2. [2]

    2025 , url =

    Egg, Alex and Iglesias Goyanes, Martin and Kingma, Friso and Mora, Andreu and von Werra, Leandro and Wolf, Thomas , journal =. 2025 , url =

  3. [3]

    2025 , url =

    Zhang, Dan and Zhoubian, Sining and Cai, Min and Li, Fengzu and Yang, Lekang and Wang, Wei and Dong, Tianjiao and Hu, Ziniu and Tang, Jie and Yue, Yisong , journal =. 2025 , url =

  4. [4]

    2025 , howpublished =

    Introducing. 2025 , howpublished =

  5. [5]

    arXiv preprint arXiv:2410.21276 , year =

  6. [6]

    2026 , howpublished =

  7. [7]

    arXiv preprint arXiv:2603.25040 , year =

  8. [8]

    2025 , url =

    Yang, An and Li, Anfeng and Yang, Baosong and others , journal =. 2025 , url =

Show all 53 references
  1. [9]

    and Barrett, Clark W

    Zheng, Lianmin and Yin, Liangsheng and Xie, Zhiqiang and Sun, Chuyue and Huang, Jeff and Yu, Cody Hao and Cao, Shiyi and Kozyrakis, Christos and Stoica, Ion and Gonzalez, Joseph E. and Barrett, Clark W. and Sheng, Ying , booktitle =. 2024 , url =

  2. [10]

    2025 , howpublished =

  3. [11]

    2025 , url =

    Zhang, Shaolei and Fan, Ju and Fan, Meihao and Li, Guoliang and Du, Xiaoyong , journal =. 2025 , url =

  4. [12]

    and Payani, Ali and Sun, Huan , booktitle =

    Li, Yifei and Moussa, Hanane Nour and Chen, Ziru and Chen, Shijie and Yu, Botao and Xue, Mingyi and Burns, Benjamin and Chiu, Tzu-Yao and Dey, Vishal and Lu, Zitong and Wei, Chen and Zhang, Qianheng and Zhang, Tianyu and Gao, Song and Huang, Xuhui and Ning, Xia and Ahmed, Nesr...

  5. [13]

    arXiv preprint arXiv:2509.25084 , year =

    Scaling Generalist Data-Analytic Agents , author =. arXiv preprint arXiv:2509.25084 , year =

  6. [14]

    Advances in Neural Information Processing Systems , year =

    Are Large Language Models Good Statisticians? , author =. Advances in Neural Information Processing Systems , year =

  7. [15]

    Liu, Xiao and Wu, Zirui and Wu, Xueqing and Lu, Pan and Chang, Kai-Wei and Feng, Yansong , booktitle =. Are. 2024 , pages =

  8. [16]

    Zhang and Zhu, Lanyi and Merrill, Mike A

    Gu, Ken and Shang, Ruoxi and Jiang, Ruien and Kuang, Keying and Lin, Richard-John and Lyu, Donghe and Mao, Yue and Pan, Youran and Wu, Teng and Yu, Jiaqian and Zhang, Yikun and Tianmai M. Zhang and Zhu, Lanyi and Merrill, Mike A. and Heer, Jeffrey and Althoff, Tim , booktitle ...

  9. [17]

    and Burns, Benjamin and Adu-Ampratwum, Daniel and Huang, Xuhui and Ning, Xia and Gao, Song and Su, Yu and Sun, Huan , booktitle =

    Chen, Ziru and Chen, Shijie and Ning, Yuting and Zhang, Qianheng and Wang, Boshi and Yu, Botao and Li, Yifei and Liao, Zeyi and Wei, Chen and Lu, Zitong and Dey, Vishal and Xue, Mingyi and Baker, Frazier N. and Burns, Benjamin and Adu-Ampratwum, Daniel and Huang, Xuhui and Nin...

  10. [18]

    2023 , url =

    Abdulhai, Marwa and White, Isadora and Snell, Charlie and Sun, Charles and Hong, Joey and Zhai, Yuexiang and Xu, Kelvin and Levine, Sergey , journal =. 2023 , url =

  11. [19]

    2025 , url =

    Qian, Cheng and Acikgoz, Emre Can and He, Qi and Wang, Hongru and Chen, Xiusi and Hakkani-Tur, Dilek and Tur, Gokhan and Ji, Heng , journal =. 2025 , url =

  12. [20]

    2024 , url =

    Liu, Xiao and Yu, Hao and Zhang, Hanchen and Xu, Yifan and Lei, Xuanyu and Lai, Hanyu and Gu, Yu and Ding, Hangliang and Men, Kaiwen and Yang, Kejuan and others , booktitle =. 2024 , url =

  13. [21]

    2024 , url =

    Chan, Jun Shern and Chowdhury, Neil and Jaffe, Oliver and Aung, James and Sherburn, Dane and Mays, Evan and Starace, Giulio and Liu, Kevin and Maksin, Leon and Patwardhan, Tejal and Weng, Lilian and Madry, Aleksander , journal =. 2024 , url =

  14. [22]

    2025 , url =

    Wang, Zihan and Wang, Kangrui and Wang, Qineng and Zhang, Pingyue and Li, Linjie and Yang, Zhengyuan and Jin, Xing and Yu, Kefan and Nguyen, Minh Nhat and Liu, Licheng and Gottlieb, Eli and Lu, Yiping and Cho, Kyunghyun and Wu, Jiajun and Fei-Fei, Li and Wang, Lijuan and Choi,...

  15. [23]

    Reinforcing Multi-Turn Reasoning in

    Wei, Quan and Zeng, Siliang and Li, Chenliang and Brown, William and Frunza, Oana and Deng, Wei and Schneider, Anderson and Nevmyvaka, Yuriy and Zhao, Yang Katie and Garcia, Alfredo and Hong, Mingyi , journal =. Reinforcing Multi-Turn Reasoning in. 2025 , url =

  16. [24]

    2026 , url =

    Li, Peiji and Li, Linyang and Sun, Handa and Mai, Wenjin and Chen, Yongkang and Li, Xiaozhe and Shen, Yue and Ma, Yichuan and Sun, Yiliu and Cao, Jiaxi and He, Zhishu and Wang, Bo and Zheng, Xiaoqing and Bi, Zhaori and Qiu, Xipeng and Guo, Qipeng and Chen, Kai and Lin, Dahua ,...

  17. [25]

    2026 , url =

    Xie, Yutao and Thomas, Nathaniel and Hansen, Nicklas and Fu, Yang and Li, Li Erran and Wang, Xiaolong , journal =. 2026 , url =

  18. [26]

    2024 , url =

    Guo, Siyuan and Deng, Cheng and Wen, Ying and Chen, Hechang and Chang, Yi and Wang, Jun , booktitle =. 2024 , url =

  19. [27]

    2023 , url =

    Zhang, Wenqi and Shen, Yongliang and Lu, Weiming and Zhuang, Yueting , journal =. 2023 , url =

  20. [28]

    Data Interpreter: An

    Hong, Sirui and Lin, Yizhang and Liu, Bang and Liu, Bangbang and Wu, Binhao and Zhang, Ceyao and Wei, Chenxing and Li, Danyang and Chen, Jiaqi and Zhang, Jiayi and Wang, Jinlin and Zhang, Li and Zhang, Lingyao and Yang, Min and Zhuge, Mingchen and Guo, Taicheng and Zhou, Tuo a...

  21. [29]

    2024 , url =

    Li, Ziming and Zang, Qianbo and Ma, David and Guo, Jiawei and Zheng, Tuney and Liu, Minghao and Niu, Xinyao and Wang, Yue and Yang, Jian and Liu, Jiaheng and Zhong, Wanjun and Zhou, Wangchunshu and Huang, Wenhao and Zhang, Ge , journal =. 2024 , url =

  22. [30]

    2022 , url =

    Wang, Ruoyao and Jansen, Peter and Cote, Marc-Alexandre and Ammanabrolu, Prithviraj , journal =. 2022 , url =

  23. [31]

    , journal =

    O'Sullivan, John and Lindsay, Alan and Magerko, Brian and Lieto, Antonio and Goel, Ashok K. , journal =. 2024 , url =

  24. [32]

    and Zhu, Hao and Zhou, Xuhui and Lo, Robert and Sridhar, Abishek and Cheng, Xianyi and Bisk, Yonatan and Fried, Daniel and Alon, Uri and Neubig, Graham , booktitle =

    Zhou, Shuyan and Xu, Frank F. and Zhu, Hao and Zhou, Xuhui and Lo, Robert and Sridhar, Abishek and Cheng, Xianyi and Bisk, Yonatan and Fried, Daniel and Alon, Uri and Neubig, Graham , booktitle =. 2024 , url =

  25. [33]

    2025 , url =

    Qiang, Rushi and Zhuang, Yuchen and Li, Yinghao and Dingu Sagar, V K and Zhang, Rongzhi and Li, Changhao and Wong, Ian Shu-Hei and Yang, Sherry and Liang, Percy and Zhang, Chao and Dai, Bo , booktitle =. 2025 , url =

  26. [34]

    arXiv preprint arXiv:2604.04872 , year =

    Synthetic Sandbox for Training Machine Learning Engineering Agents , author =. arXiv preprint arXiv:2604.04872 , year =

  27. [35]

    arXiv preprint arXiv:1707.06347 , year =

    Proximal Policy Optimization Algorithms , author =. arXiv preprint arXiv:1707.06347 , year =

  28. [36]

    Shao, Zhihong and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Song, Junxiao and Bi, Xiao and Zhang, Haowei and Zhang, Mingchuan and Li, Y. K. and Wu, Y. and Guo, Daya , journal =. 2024 , url =

  29. [37]

    2025 , url =

    Guo, Daya and Yang, Dejian and Zhang, Haowei and Song, Junxiao and Zhang, Ruoyu and Xu, Runxin and Zhu, Qihao and Ma, Shirong and Wang, Peiyi and Bi, Xiao and others , journal =. 2025 , url =

  30. [38]

    2025 , url =

    Yu, Qiying and Zhang, Zheng and Zhu, Ruofei and Yuan, Yufeng and Zuo, Xiaochen and Yue, Yu and Fan, Tiantian and Liu, Gaohong and Liu, Lingjun and Liu, Xin and others , journal =. 2025 , url =

  31. [39]

    Understanding

    Liu, Zichen and Liu, Chang and Liu, Xueqing and Guo, Daya and Chen, Jun and Zhou, Chunting , journal =. Understanding. 2025 , url =

  32. [40]

    arXiv preprint arXiv:2604.02288 , year =

    Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing , author =. arXiv preprint arXiv:2604.02288 , year =

  33. [41]

    Group-in-Group Policy Optimization for

    Feng, Lang and Xue, Zhenghai and Liu, Tingcong and An, Bo , journal =. Group-in-Group Policy Optimization for. 2025 , url =

  34. [42]

    Segment Policy Optimization: Effective Segment-Level Credit Assignment in

    Guo, Yiran and Xu, Lijie and Liu, Jie and Ye, Dan and Qiu, Shuang , journal =. Segment Policy Optimization: Effective Segment-Level Credit Assignment in. 2025 , url =

  35. [43]

    2026 , url =

    Lu, Zhicong and Lin, Zichuan and Jia, Wei and Tian, Changyuan and Ye, Deheng and Li, Peiguang and Jin, Li and Liu, Nayu and Xu, Guangluan and Feng, Wei , journal =. 2026 , url =

  36. [44]

    2025 , url =

    Wang, Hanlin and Wang, Jian and Leong, Chak Tou and Li, Wenjie , booktitle =. 2025 , url =

  37. [45]

    2025 , url =

    Wang, Ziliang and Zheng, Xuhui and An, Kang and Ouyang, Cijun and Cai, Jialu and Wang, Yuhang and Wu, Yichao , journal =. 2025 , url =

  38. [46]

    2026 , url =

    Zong, Zefang and Chen, Dingwei and Li, Yang and Yi, Qi and Zhou, Bo and Li, Chengming and Qian, Bo and Chen, Peng and Jiang, Jie , journal =. 2026 , url =

  39. [47]

    Kelly, Markelle and Longjohn, Rachel and Nottingham, Kolby , year =. The

  40. [48]

    and Bischl, Bernd and Torgo, Luis , journal =

    Vanschoren, Joaquin and van Rijn, Jan N. and Bischl, Bernd and Torgo, Luis , journal =. 2013 , doi =

  41. [49]

    2025 , doi =

    Nucleic Acids Research , volume =. 2025 , doi =

  42. [50]

    Advances in Neural Information Processing Systems Datasets and Benchmarks Track , year =

    Monash Time Series Forecasting Archive , author =. Advances in Neural Information Processing Systems Datasets and Benchmarks Track , year =

  43. [51]

    and Fang, Tao and Doncheva, Nadezhda T

    Szklarczyk, Damian and Kirsch, Rebecca and Koutrouli, Mikaela and Nastou, Katerina and Mehryary, Farrokh and Hachilif, Rayan and Gable, Annika L. and Fang, Tao and Doncheva, Nadezhda T. and Pyysalo, Sampo and Bork, Peer and Jensen, Lars J. and von Mering, Christian , journal =...

  44. [52]

    Advances in Neural Information Processing Systems , year =

    Open Graph Benchmark: Datasets for Machine Learning on Graphs , author =. Advances in Neural Information Processing Systems , year =

  45. [53]

    2026 , howpublished =

    Yelp Open Dataset , author =. 2026 , howpublished =

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.