Pith. sign in

REVIEW 4 major objections 5 minor 26 references

The paper argues that AI-native biotechs should be organized around a Company World Model—a shared asset-to-value state with transition operators, an explicit value function, and a planner—rather than around department-shaped agent roles.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 14:36 UTC pith:GJ7ZNNFV

load-bearing objection Useful benchmark, honest stress-testing, but the headline claim is closer to rubric-matching than to a demonstrated advantage of Company World Models. the 4 major comments →

arxiv 2607.18696 v1 pith:GJ7ZNNFV submitted 2026-07-21 cs.AI

Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development

classification cs.AI
keywords AI-native biotechCompany World Modeldrug development decisionsLLM agentsorganizational designretrospective benchmarkasset-to-value stateobjective sensitivity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that AI-native biotechs should not be built by copying human biotech org charts into agent roles; the core design primitive should instead be a Company World Model—a shared, auditable representation of an asset's state and its path to enterprise value, with transition operators, an explicit value function, and a planner. To test this, it builds a dry-lab benchmark of 45 retrospective drug-development decisions with time cutoffs, hidden outcomes, and blinded judging, then compares department-mimicking and asset-centric agent architectures with a value-conversion architecture that approximates the Company World Model. Under an objective defined by external business development, regulatory approval and launch, and revenue discipline, the value-conversion architecture scores highest and is strongly preferred by value-specific blinded judges. Stress tests narrow the claim: a stronger human baseline stays competitive, and a neutral judge does not confirm the advantage. The paper's central conclusion is objective-sensitive: an AI-native biopharma should expose its value function, and departments remain useful governance views but not the deepest computational layer.

Core claim

On the paper's own terms, the discovery is that organizational abstraction matters more than role labels for an AI-native biopharma. The authors define a Company World Model as a live asset-to-value state, a transition model for how actions change that state, a value function tied to external BD, approval, and revenue, and a planner that selects next actions. Their prompt-level approximation—a Live Asset Value Record updated by Deal, Approval, Revenue, and Investment Arbiter loops—outperformed the original human-org-mimic and asset-centric architectures on automatic scoring (4.68 vs. 4.33 and 4.26) and was preferred 42–3 and 44–1 by value-specific blinded judges. The authors are explicit tha

What carries the argument

The central object is the Company World Model: a persistent, auditable asset-to-value state, together with a transition model, an explicit value function, and a planner. It is instantiated here at prompt level as a Live Asset Value Record that Deal, Approval, Revenue, and Investment Arbiter loops update—each loop acts as a domain transition and value operator over the same shared state rather than as a department. The machinery forces every agent output to reason from one state object and makes the optimization objective explicit so the planner can recommend go, no-go, watch, or partner. The ablation design, removing one loop at a time while keeping the state record, is what lets the authors

Load-bearing premise

The load-bearing premise is that the automatic value-conversion proxy and the value-specific judges measure genuine decision quality for the stated objective rather than rewarding outputs that merely match the architecture's template and judging prompt—and the authors recognize the value-conversion architecture is more aligned with its judging prompt than the baselines, while the neutral judge did not confirm the advantage.

What would settle it

Take the same 45 hidden-outcome cases and the same generated outputs, then run a pre-registered neutral judge from a different model family with all organizational terminology stripped and output lengths equalized; if value-conversion outputs are not preferred over human-org-mimic-plus at or above chance, or if the automatic proxy scores re-order, the 42–3/44–1 claims are attributable to prompt-template alignment rather than organizational design.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the central claim holds, an AI-native drug-development company should be designed around a shared Live Asset Value Record and value-conversion operators, not a chain of department-like agent roles.
  • Departments can remain as human governance, audit, and expert-review interfaces, but they are no longer the core computational primitive.
  • The value function must be explicit and reviewable; an organization optimized for BD, approval, and revenue conversion will not automatically maximize neutral diligence, evidence quality, or auditability.
  • Deal, approval, and revenue reasoning cannot be appended after scientific diligence; they need to be part of the state model from the start.
  • The 45-case benchmark can be reused as a dry-lab evaluation suite for future agent-organization designs, since outcomes and scoring notes are hidden from the public case files.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The authors leave implicit that the surviving evidence may be an objective-framing effect: the strongest wins came from a judge prompt aligned with the architecture, and the neutral judge did not reproduce them, so the portable lesson may be the need for an explicit value function rather than the specific loop structure.
  • A direct extension would be longitudinal: run the same 45 cases over several simulated decision points, updating the Live Asset Value Record after each new public observation, and test whether multi-step updates track retrospectively known outcomes better than one-shot outputs—that would separate world-model updating from static template effects.
  • The abstraction should port to other siloed sequential-decision organizations—startups, research consortia, investment firms—where the same retrospective-case method could benchmark departments against a shared asset-to-value state.
  • A stronger interpretation, not tested here, is that the value-conversion advantage will grow when transition operators become learned numerical simulators of evidence, regulatory, and market dynamics rather than prompt-level role descriptions; the paper only claims the prompt-level version is a first approximation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper argues that AI-native biopharmaceutical organizations should not simply copy human department charts into agent roles. Instead, it introduces a 'Company World Model' abstraction: a persistent asset-to-value state, transition operators for science/regulatory/BD/commercial/execution dynamics, an explicit value function, and a planner. To test this, the authors built a 45-case retrospective dry-lab benchmark with time cutoffs, hidden outcomes, and blinded LLM judging, comparing human-org-mimic, human-org-mimic-plus, AI-native asset-centric, and AI-native value-conversion architectures. The value-conversion architecture, a prompt-level approximation of the Company World Model, achieves the highest automatic value-conversion score and is strongly preferred by value-specific judges over the original baselines, but loses to or ties the stronger human baseline and neutral judge. Ablations remove individual rooms (Deal, Approval, Revenue, Investment Arbiter) and show the full architecture is preferred. The paper is unusually candid about limitations, including the circularity risk that the value-conversion architecture is aligned with the value-conversion judging prompt.

Significance. If the central design recommendation were empirically validated, this would be a useful contribution to the emerging field of AI-agent organization design in regulated, high-stakes industries. The paper's dry-lab benchmark, public no-label dataset, matched architectures, and explicit stress-testing methodology are valuable engineering contributions, and the authors deserve credit for reporting null/weak results under neutral judging rather than only the favorable pairwise wins. However, the key empirical evidence for the Company World Model is currently confounded by rubric-architecture alignment: the automatic scoring is built from the same success function the architecture optimizes, and the strongest pairwise wins come from a judge prompt that mirrors the architecture's own vocabulary. The neutral judge and strong-baseline results undermine the central positive claim. The paper is therefore a promising systems-design proposal with a useful benchmark, but it does not yet establish the advertised design conclusion.

major comments (4)
  1. [§4.2–4.3, Tables 1 and 3] The headline value-conversion wins (42–3, 44–1) come from a value-specific LLM judge that asks which output 'better converts the asset into BD/approval/revenue value'; the winning architecture contains Deal, Approval, and Revenue Rooms designed for exactly that objective, and §6 admits it is 'more aligned' with the prompt. Under the neutral judge the value-conversion architecture loses to human-org-mimic 23–22 and to human-org-mimic-plus 25–20, and against the stronger baseline the value-specific Codex judge gives only 26–19 (p=0.371). The pairwise advantage is therefore confounded with template alignment. This is load-bearing for the conclusion that value-conversion organization is useful; without an outcome-based or expert-adjudicated measure, the central positive result reduces to rubric matching.
  2. [§3.3, §3.6, Table 1] The automatic value-conversion proxy score is constructed from the same success function (BD, approval/launch, revenue discipline) that the value-conversion architecture is explicitly designed to optimize. Thus 'highest automatic score' is not an independent confirmation. The decision-credit differences (0.91 vs 0.88 vs 0.86) are reported as point estimates on 45 cases with no significance test or interval; these differences are likely within noise. The paper should either provide a statistically justified outcome metric (e.g., acceptable-decision match with McNemar tests) or remove the automatic score from the evidence for the architectural recommendation.
  3. [§4.4, Table 3] The mechanistic ablations are interpreted as showing that Revenue, Deal, and Approval Rooms 'carry useful work,' but the same value-specific Codex judge is used and the full architecture is even more aligned with that judge than the ablations. Removing a room removes both a computational component and the related vocabulary/structure; the 44–1 and 41–4 preferences may reflect judge recognition of missing template elements rather than degraded decision quality. The manuscript already notes the Codex-only limitation, but the sentence in §4.4 that 'these ablations support a mechanistic reading' goes beyond what the data can show. Cross-model or human expert judging, or pre-registered outcome-based evaluation, is needed before these ablations can support the world-model interpretation.
  4. [§7 Conclusion] The concluding recommendation that an AI-native biopharma 'should be designed around a Company World Model' is not entailed by the empirical results. The benchmark supports at most an objective-relative statement: under an explicit BD/approval/revenue value-conversion objective, the value-conversion prompt architecture is preferred by a judge aligned with that objective, but this preference does not transfer to neutral decision quality. Given the neutral-judge and strong-baseline failures, the manuscript should reframe the CWM design recommendation as a testable hypothesis for future work rather than a 'best current conclusion,' unless additional evidence is supplied.
minor comments (5)
  1. [Abstract / §4.2] The abstract's '42–3 / 44–1' refers to the Codex judge only; Claude had 40–5 and 42–3. Please specify the judge in the abstract to avoid ambiguity.
  2. [§3.5] Please clarify the exact model version and inference settings for gpt-5.5 used with Codex CLI, including temperature and any other sampling parameters, to improve reproducibility.
  3. [Table 3] The pairwise judging table reports only wins and losses with no ties. If ties were possible, please state how they were handled and whether the binomial tests exclude them; if ties were impossible by construction, say so.
  4. [§4.5] The output-length check reports means only. Reporting standard deviations or a formal comparison would strengthen the claim that the pairwise results are not driven by length differences.
  5. [Artifact Availability] The manuscript states that the public no-label benchmark case file is included in the submission package. Please verify that the ancillary file is actually present and listed in the submission manifest, since it is central to reproducibility.

Circularity Check

2 steps flagged

Value-conversion benchmark wins are partly rubric-architecture circularity; the paper's own neutral-judge stress test fails to confirm, leaving only an objective-sensitive, partial empirical case for the Company World Model.

specific steps
  1. self definitional [Abstract; §3.3; §3.4; §3.6; §4.2; §6]
    "Under a success function defined by external BD, regulatory approval and launch, and revenue discipline, this architecture achieved the highest automatic value-conversion score and was strongly preferred over the original baselines by value-specific blinded judges. ... Second, value-specific blinded pairwise judges compare outputs A and B without architecture labels and are asked which output better converts the asset into BD, approval, and revenue value under cutoff-appropriate uncertainty. ... This is the strongest positive evidence for value-conversion organization, but it is also the resul"

    The benchmark is a closed loop by construction: the success function (external BD, approval/launch, revenue), the architecture (Deal, Approval, Revenue Rooms), the automatic 'value-conversion proxy', and the value-specific judge prompt are all defined in terms of the same three outcome categories. An architecture organized around those exact loops is therefore scored by a rubric that restates its own internal objective; the 42–3 / 44–1 wins measure alignment with the judging prompt rather than an independent external outcome. §6 concedes this: 'the value-conversion architecture is more aligned with the value-conversion judging prompt than the baselines,' and the neutral judge removes the advantage (25–20 vs human-org-mimic-plus). Thus the strongest empirical support for the Company World M

  2. fitted input called prediction [§3.6; §4.4; §5.3]
    "Mechanistic ablations comparing the full value-conversion architecture against four versions that remove exactly one loop: Deal Room, Approval Room, Revenue Room, or Investment Arbiter. ... Removing Revenue Room mostly damaged revenue discipline, removing Deal Room damaged BD conversion, and removing Approval Room damaged approval conversion."

    The ablation outcomes are read off the same value-conversion proxy and value-specific judge that are constructed from the very categories being removed. Removing the Revenue Room and then observing a drop in 'revenue discipline' is expected when the scorer explicitly prompts for revenue value conversion; it does not independently validate the Company World Model's transition operators. The paper's own hedge—'consistent with, but do not prove, the world-model interpretation'—confirms that this evidence is not independent of the rubric.

full rationale

The paper's derivation chain does not rest on self-citation: the world-model references are external prior work, and no uniqueness theorem is imported from the authors' own publications. The 'Company World Model' is explicitly a prompt-level approximation, and the paper disclaims a learned numerical simulator. The remaining circularity is in the empirical support: the value-conversion architecture is evaluated with a value-conversion judge and a value-conversion proxy, both built from the same three-outcome objective that the architecture instantiates. The paper transparently flags this and runs a neutral-judge stress test, which fails to confirm the advantage (human-org-mimic-plus beats it 25–20). This mitigates overclaiming, but it also demonstrates that the headline wins are not independent evidence. The ablations suffer from the same judge confound. No external outcome data (real BD terms, realized approvals, actual revenue) are used. Because the paper narrows its conclusion to 'objective-sensitive' and the neutral judge remains a real counter-test, the circularity is partial rather than total: score 6.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 3 invented entities

The central claim depends on several unvalidated modeling choices: the success function, the unpublished scoring rubric, the hand-picked case mix, LLM-judge validity, and the representativeness of the Codex harness. There are no fitted numerical parameters, but the benchmark's evaluation stack is itself a free design choice.

free parameters (3)
  • Value-conversion success function components
    The objective (external BD, regulatory approval/launch, revenue discipline) is chosen by the authors and hard-coded into both the value-conversion architecture and the evaluation rubric (§3.3). It is not derived from data and is the entire basis for the headline score.
  • Automatic value-conversion score rubric
    The composite 0–5 score and its BD/Approval/Revenue/Decision sub-scores come from an unpublished scoring pipeline; weights and scoring rules are not reported, so the 4.68 vs. 4.33 gap cannot be independently recomputed (§3.6, Table 1).
  • 45-case mix = 13 BD/competitive, 10 failure, 12 mixed, 10 success
    Case selection is hand-picked, not randomized or pre-registered; results may depend on this composition (§3.2).
axioms (5)
  • domain assumption LLM agents can share state and dynamically load skills, so human org-chart design is not the only or best default.
    This premise motivates the entire comparison; it is asserted in §1 but not proven in this paper.
  • domain assumption Retrospective public-information cases with time cutoffs are a meaningful proxy for biopharma decision quality.
    The benchmark evaluates written decision outputs, not real actions or outcomes (§3.2, §6).
  • domain assumption LLM judges provide valid pairwise preference judgments for asset-to-value conversion.
    All headline preference results use LLM judges; the paper acknowledges value-specific judging is aligned with the architecture (§3.6, §6).
  • domain assumption Codex CLI with gpt-5.5 is a representative agent harness for AI-native organizations.
    All outputs and most stress-test judges are Codex-based; Claude replication was blocked (§3.5, §6).
  • domain assumption Revenue discipline reasoning (payer, access, adoption, forecast assumptions) approximates the revenue outcome.
    The paper explicitly says revenue is measured as reasoning discipline, not forecast accuracy (§3.3, §6).
invented entities (3)
  • Company World Model no independent evidence
    purpose: Proposed core computational abstraction for AI-native biotech: persistent asset-to-value state, transition model, value function, planner, update process.
    Defined in §3.3 and implemented only as a prompt-level architecture; no learned simulator or real-world validation, so independent evidence is absent.
  • Live Asset Value Record no independent evidence
    purpose: Shared state object that Deal/Approval/Revenue/Investment Arbiter loops update.
    Introduced in §3.4/§4.4; evidence for its usefulness comes only from the same benchmark and Codex-only ablations.
  • Deal, Approval, Revenue Rooms and Investment Arbiter no independent evidence
    purpose: Domain transition/value operators over the shared state.
    Ablations show degradation when removed, but these are internal, same-pipeline comparisons; no external benchmark or human expert validation.

pith-pipeline@v1.3.0-alltime-deepseek · 10272 in / 14784 out tokens · 130700 ms · 2026-08-01T14:36:10.830099+00:00 · methodology

0 comments
read the original abstract

AI-native biotechnology companies are often designed by copying human biotech org charts into agent roles. We argue for a different abstraction: a Company World Model, defined as a persistent asset-to-value state representation with transition models, explicit value functions, planning, and updating across scientific, regulatory, BD, commercial, financial, and execution constraints. We introduce a dry-lab benchmark for testing whether AI-agent organizations should mimic departments or operate around such a world model. The benchmark contains 45 retrospective public-information decision cases with strict time cutoffs, hidden outcomes, common schemas, automatic scoring, and blinded pairwise judging. We compare human-org-mimic, stronger human-org-mimic-plus, AI-native asset-centric, and AI-native value-conversion architectures. The value-conversion architecture is a prompt-level approximation of a Company World Model: a Live Asset Value Record updated by Deal, Approval, Revenue, and Investment Arbiter loops. Under a success function defined by external BD, regulatory approval and launch, and revenue discipline, it achieved the highest automatic value-conversion score and was strongly preferred over the original baselines by value-specific blinded judges. Stress tests narrowed the claim: a stronger human baseline remained competitive, and a neutral judge did not show robust value-conversion dominance. Codex-only mechanistic ablations suggest that Revenue Room, Deal Room, and Approval Room carry useful work under the target objective. The central finding is objective-sensitive: departments may remain useful governance views, but the core AI-native operating primitive should be a shared, predictive asset-to-value state rather than a static human org chart. The study is dry-lab only and does not establish real-world drug success, clinical benefit, or revenue prediction accuracy.

Figures

Figures reproduced from arXiv: 2607.18696 by Yinan Wang.

Figure 1
Figure 1. Figure 1: The benchmark compares a human-org-mimic architecture, in which functional agents [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Automatic value-conversion score after adding the stronger human-org-mimic-plus base [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Codex-only blinded value-conversion judge preferences for the full architecture versus [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

26 extracted references · 14 linked inside Pith

  1. [1]

    David Ha and J¨ urgen Schmidhuber. 2018. World Models.arXiv:1803.10122

  2. [2]

    Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal Behbahani, Stephanie Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott Reed, Jingwei Zhang, Konrad Zolna, Jeff Clune, Nando de Freitas, Satinder Si...

  3. [3]

    Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Ar- naud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xi...

  4. [4]

    Meng Chu, Xuan Billy Zhang, Kevin Qinghong Lin, Lingdong Kong, Jize Zhang, Teng Tu, Weijian Ma, Ziqi Huang, Senqiao Yang, Wei Huang, Yeying Jin, Zhefan Rao, Jinhui Ye, Xinyu Lin, Xichen Zhang, Qisheng Hu, Shuai Yang, Leyang Shen, Wei Chow, Yifei Dong, Fengyi Wu, Quanyu Long, Bin Xia, Shaozuo Yu, Mingkang Zhu, Wenhu Zhang, Jiehui Huang, Haokun Gui, Runyi L...

  5. [5]

    O’Brien, Carrie J

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. arXiv:2304.03442

  6. [6]

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023. Voyager: An Open-Ended Embodied Agent with Large Language Models.arXiv:2305.16291

  7. [7]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models.International Con- ference on Learning Representations

  8. [8]

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao

  9. [9]

    Griffiths, Yuan Cao, and Karthik Narasimhan

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Lan- guage Models.arXiv:2305.10601

  10. [10]

    Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang

  11. [11]

    Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Yuxian Gu, Hangliang Ding, Kai Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. 2023. AgentBench: Evaluating LLMs as Agents.arXiv:2308.03688

  12. [12]

    arXiv:2303.17580

    HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. arXiv:2303.17580. 13

  13. [13]

    Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?International Conference on Learning Representations

  14. [14]

    Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, and Jie Tang. 2024. Agent- Board: An Analytical Evaluation Board of Multi-turn LLM Agents.arXiv:2401.13178

  15. [15]

    John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ron- neberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Zidek, Anna Potapenko, and oth- ers. 2021. Highly accurate protein structure prediction with AlphaFold.Nature596, 583–589

  16. [16]

    Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes

    Daniil A. Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. 2023. Autonomous chemical research with large language models.Nature624, 570–578

  17. [17]

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, and others. 2022. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.arXiv:2206.04615

  18. [18]

    Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. 2024. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.arXiv:2408.06292

  19. [19]

    Gonzalez, and Ion Stoica

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica

  20. [20]

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, and others. 2022. Holistic Eval- uation of Language Models.arXiv:2211.09110

  21. [21]

    Cohen, and Xinghua Lu

    Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. 2019. Pub- MedQA: A Dataset for Biomedical Research Question Answering.Empirical Methods in Nat- ural Language Processing

  22. [22]

    Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.Advances in Neural In- formation Processing Systems

  23. [23]

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.Empirical Methods in Natural Language Processing

  24. [25]

    Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams.Applied Sciences11(14), 6421

  25. [26]

    Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, and others

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, and others. 2023. Large language models encode clinical knowledge.Nature620, 172–180. 14

  26. [2023]

    Reflexion: Language Agents with Verbal Reinforcement Learning.arXiv:2303.11366