Pith. sign in

Paper Citation Record · LEDGER

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design

As of 23 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2605.15341.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.15341 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T15:50:01.064709Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T01:16:50.927415Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact9
  • verified fuzzy24
  • unresolved4
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5cf1daf-ba07-4fe6-88c1-0bd3449b9607 · outbound

This paper cites Gonzalez.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Gonzalez

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:39.013931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:3eb63101cd082f7343b2665d43071a880caebab051bac063463c190a1f1ee019

Observation e631a623-c8ed-47e4-8268-51ee34bb6aa5 · outbound

This paper cites Autonomous chemical research with large language models.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Autonomous chemical research with large language models

Reference 2

Resolution
verified exact
doi, observed 2026-05-19T15:52:37.969723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:95413a6724bc935a57998b08515b8c2d0b24b9ae6a7bfdb458c7457dae002733

Observation 269297ba-250e-419b-8736-9973f737c78f · outbound

This paper cites On the Measure of Intelligence.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design On the Measure of Intelligence

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:52:38.080243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:819b463a5a95234e54944fdbf75fb35b29ec065999b16b6998a3ed048a033061

Observation 651285ec-1fe0-4e4b-9cc6-6f64dde5da37 · outbound

This paper cites Towards an AI co-scientist.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Towards an AI co-scientist

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:52:38.089188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:2ec64a1537b2c908f600f9741fdec5b0ee11a45e0a4ae9cf3b720f3762a88b5f

Observation 60020a3d-142f-4cbe-82ff-c3a849976384 · outbound

This paper cites Ideabench: Benchmarking large language models for research idea generation.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Ideabench: Benchmarking large language models for research idea generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:39.004715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:9e428618409d80ac396ceaf44009d7d2fcf4a0c05174920fc1365fe5df24e223

Observation 77e0de7a-30d0-4f0e-b4b2-233d509d1c14 · outbound

This paper cites BurstGPT: A real-world workload dataset to optimize LLM serving systems.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design BurstGPT: A real-world workload dataset to optimize LLM serving systems

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T15:52:37.960886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:41ad221db666bad1e866e987eca84107f6684facf532e723cb38a5527ae1ba30

Observation f0908c49-759b-44bc-9412-464948262429 · outbound

This paper cites LAB-Bench: Measuring Capabilities of Language Models for Biology Research.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design LAB-Bench: Measuring Capabilities of Language Models for Biology Research

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:52:38.083379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:ae6f35c98c6cd14170d92cc0ee9bd408f42271d75a26a2ed19fcc67990d4d58b

Observation 221c4c7c-b92b-4040-b84f-f23b649768d4 · outbound

This paper cites ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:52:38.068813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:bfb3b9642d174b8a45e1b01aa5db92a95a697508de56daa15540e4ca47358400

Observation 0a6921a4-010e-4a87-9dec-fba740d5baaa · outbound

This paper cites A framework for evaluating the chemical knowledge and reasoning abilities of large language models against the expertise of chemists.Nature Chemistry.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design A framework for evaluating the chemical knowledge and reasoning abilities of large language models against the expertise of chemists.Nature Chemistry

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.961604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:e35b9b774742199688c85863b8af3bee4563c71436cc6ff20da3ba90d5ed2604

Observation 4455a4df-0ea0-4926-bf7f-f6f2bb1f71fd · outbound

This paper cites Laurent, Alex Andonian, Benjamin Tenmann, Siddharth Narayanan, Geemi P.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Laurent, Alex Andonian, Benjamin Tenmann, Siddharth Narayanan, Geemi P

Reference 10

Resolution
malformed identifier
arxiv_id, observed 2026-05-19T15:52:38.075556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:88ba1a26c0421daa5ef72434d90f9beeb5df9153dfebc6241492110fa8f054ec

Observation c3b4dfc1-d676-44b8-9b68-6a027b65d602 · outbound

This paper cites Journal of Open Psychology Data , author =.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Journal of Open Psychology Data , author =

Reference 11

Resolution
metadata mismatch
doi, observed 2026-05-19T15:52:37.962911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:5461d7e553af9ced4f92307ad30ed9b8e9ba1517418365d86395125fcbc838b1

Observation ddd4b8dd-6b70-458c-a819-e05b7575915e · outbound

This paper cites Quantifying language models’ sensitivity to spurious features in prompt design, or: How i learned to start worrying about prompt formatting.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Quantifying language models’ sensitivity to spurious features in prompt design, or: How i learned to start worrying about prompt formatting

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.967808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:0c16ab4e0504f1917b285a35834f42e7f0cc4f92c4f3e9a9617febdaedf53d47

Observation 7a375841-1d3c-4874-9943-8543ce926562 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:52:38.086290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:3085ad39d7bc5d61c57ca821fbd1bc0606d281a4a596aebd7f315bc27d6c5279

Observation 96707ca3-2673-4e7d-be44-270da38de506 · outbound

This paper cites McConnell Rooney, M.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design McConnell Rooney, M

Reference 14

Resolution
malformed identifier
doi_truncated, observed 2026-05-19T15:52:37.967345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:c013b2d859a8fd0c55690be696925863e9b0478b0042a25c1dea6e431c42d7b2

Observation fb733375-8091-49ae-b246-0ba0c42c5cf7 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Solving math word problems with process- and outcome-based feedback

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T15:52:38.071639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:0ec50575a171ff4636d5b8ce87565fdd3ace3725327a41be4ae26b88b8abb48f

Observation 3c590227-1b47-4efa-bf25-fe5d64cd14d9 · outbound

This paper cites Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:39.017783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:f0fe1b8e630c921bad767768b4a490c123b656467fdd3e87fe6f973a95a75634

Observation fd5f72ee-db10-4a12-b720-b7da73af6b89 · outbound

This paper cites Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:39.008664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:68c8d7e91a09bd102113cb2cbf0d723be9b8532adad843f22aee42e6878237eb

Observation 8188bd6c-ad1b-44fd-9b47-fd4cf14f668e · outbound

This paper cites Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin

Reference 18

Resolution
verified exact
doi, observed 2026-05-19T15:52:37.971766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:f518bb75ea100127c0a263b2c2a2b8517fb8d925087c0dc739a90bb05e05f189

Observation 244412da-2759-48b4-8b3b-deecc4444c45 · outbound

This paper cites rank-2, 50%).

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design rank-2, 50%)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:39.001109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:9c0aefae13106e5222b79103a1a72429c7fc00d369c17bff0d7fa75dfa6c2978

Observation 568380db-c8d6-437e-a0ea-009232258018 · outbound

This paper cites an unresolved cited work.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-19T15:52:39.002833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:52643ff92e0cc09672fd1df3d4909baa0e6a933e8c1d36033f677fd6dd84ddf2

Observation bd4eff59-07d9-4195-80b5-acc2fd6b458a · outbound

This paper cites You are optimizing CRISPR HDR efficiency.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design You are optimizing CRISPR HDR efficiency

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:52:37.957631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:3f0da9ac77a6905cfec81852440039b66f28a99f4950c6ee31efac5864c7fbff

Observation 10c2a073-92eb-4340-8691-877eeaf00b1e · outbound

This paper cites no model clears 50%.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design no model clears 50%

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:39.006745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:25761f6232449ab573c2e0dabf496f98eae6f8d5ba6525df9b5e8bf1393ffdaf

Observation bf6af465-1330-42cc-9510-e96c6c9f4b76 · outbound

This paper cites Task-clustered 95% CIs shown.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Task-clustered 95% CIs shown

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:39.012193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:0bc254b0db3b839d93f5dd61ee8e9ba275d7bd982676f0b3895020edb72e1748

Observation e8f92d74-ef9c-4fa0-a83c-6c268a324c23 · outbound

This paper cites Error bars are task-clustered bootstrap 95% CIs (𝐵=2000).

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Error bars are task-clustered bootstrap 95% CIs (𝐵=2000)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:39.010438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:64a6f2a9adb546bff416dab018c2c0d3ad232027088ed38227b108a97ad34900

Observation d62e5ed4-d65d-4a05-a940-6132584f5404 · outbound

This paper cites GP-UCB, across horizons.Each model is ranked by the fraction of 45 biology tasks where its median bsf-AUC@𝑘 outperforms GP-UCB.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design GP-UCB, across horizons.Each model is ranked by the fraction of 45 biology tasks where its median bsf-AUC@𝑘 outperforms GP-UCB

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:39.015969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:1eae21444fb79f4c21af6451a953fdb606a2ade8a8a0a386ad433a3381e0b653

Observation a02cde04-87ef-4fa5-9f29-fab0766f1a0e · outbound

This paper cites Error bars are task-clustered bootstrap 95% CIs (𝐵= 2000).

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Error bars are task-clustered bootstrap 95% CIs (𝐵= 2000)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.997609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:6c7147a764d0c9e069f82060fd2dbc2b6e0d7a4b92afa4e8d6af770f599a3e91

Observation 68bcc7b3-c85f-4de1-8d33-f843a07bed8d · outbound

This paper cites known-good.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design known-good

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.999342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:5c0042f51bd611a3bc5076813ff01e633ea6598e15f7f4ba04a9a313c68c6a7e

Observation e161c94b-b99f-46fb-90df-38db291771b3 · outbound

This paper cites A.15 Robustness of Figure 4: alignment, match definition, and inferential tests This appendix gives the full robustness battery for the oracle-aligned match-rate finding in §4.3.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design A.15 Robustness of Figure 4: alignment, match definition, and inferential tests This appendix gives the full robustness battery for the oracle-aligned match-rate finding in §4.3

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.993657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:45e9a0a78dd52d58619e6d373bce8c1f6c03dc4061bc3d627d8052b3efe3a67f

Observation 493bdb43-fdf3-4258-bb49-ab121992058e · outbound

This paper cites an unresolved cited work.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-19T15:52:38.991676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:7e30a87f4d72cd4bc830b6356fa28d20893bce67ed24ee37dfe59b70a78031a1

Observation 21760784-1ab0-492c-9988-38325e2d6c39 · outbound

This paper cites domain-aware’s prior matches the RCT-confirmed mechanism.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design domain-aware’s prior matches the RCT-confirmed mechanism

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.995696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:305d29823d796f88491c6f5427a8ff58dde3d19dc628d539c9903867c1d68871

Observation f8bedd77-a901-4d1c-ba1b-7db44488141b · outbound

This paper cites A.17 Threshold sensitivity of the audit’s main result The audit selects 6 literature-divergent biology tasks using two gap criteria connected by an OR: (R)top-1 vs.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design A.17 Threshold sensitivity of the audit’s main result The audit selects 6 literature-divergent biology tasks using two gap criteria connected by an OR: (R)top-1 vs

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.986533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:0a1ce584e047a3b80d099e1320be4e94178b7b2c856cdca5d6b88712dfe02f59

Observation 27dfe25f-6af0-4c3b-a7f8-63fb2735d900 · outbound

This paper cites literature-typical and best-result diverge by a meaningful margin.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design literature-typical and best-result diverge by a meaningful margin

Reference 32

Resolution
malformed identifier
raw_fallback, observed 2026-05-19T15:52:38.984285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:a49fa55bc40ef5ffa311229599585e18c7210e6292073963b08267e977daed39

Observation a7040760-9283-4b7d-b45c-54bf6e4ae5c0 · outbound

This paper cites domain-agnostic minus domain-aware.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design domain-agnostic minus domain-aware

Reference 33

Resolution
malformed identifier
raw_fallback, observed 2026-05-19T15:52:38.980650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:959c7b8531f964c776b3a5aa529840f967740b128d0a3283a563ccf8e214fc59

Observation d41a908a-0de6-4e8c-b05f-e21a9c50be38 · outbound

This paper cites an unresolved cited work.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-19T15:52:38.982469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:7c74d9b5634cb260b945a8e3d9b9d5a68ced5795a960baef9ede2515bce138d6

Observation f5562413-fe58-412f-af88-0900c28d764c · outbound

This paper cites an unresolved cited work.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Unresolved cited work

Reference 35

Resolution
malformed identifier
raw_fallback, observed 2026-05-19T15:52:38.988269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:8c08edebd2e58c44444254fc9a82175af1f85432c754968e09d5c21f8ae9051f

Observation c52728a9-8f27-4904-9fc3-2dd0e883dc42 · outbound

This paper cites an unresolved cited work.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-19T15:52:38.975338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:aa693aa14a147399bf8c4627fe64932e3dcc26d00f1cc83597d009470e968cf6

Observation dc820085-b111-4363-8a25-63d694055bf3 · outbound

This paper cites Bottom row: domain-agnostic.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Bottom row: domain-agnostic

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.978869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:e593c6efe8abbba71f04ac001f0a45c2fecdff6958b640eaaadc2bfaa08c1a65

Observation 8ee2c146-d29d-45df-a3a2-c8300bd4833a · outbound

This paper cites GP-UCB on biology under thedomain-agnosticcondition.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design GP-UCB on biology under thedomain-agnosticcondition

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.977127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:1a9ab3f3c3ace245b09ee570332f6cf292c315997b5c86310464ebc5d2f51c3c

Observation b131d0ea-b6cc-451a-b092-266acad0f02c · outbound

This paper cites fall back to common options.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design fall back to common options

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.989983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:9b9e233483d44bfd032b711e1b06e5ae5819be418ba62ef66246d7bd9c473c76

Observation 76a29c78-051a-4d53-838a-05832fbd1800 · outbound

This paper cites Oracle models are derived predictors trained on these data, released alongside the benchmark for reproducibility.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Oracle models are derived predictors trained on these data, released alongside the benchmark for reproducibility

Reference 40

Resolution
verified exact
doi, observed 2026-05-19T15:52:37.965048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:743f3861f60e230cfe6d3a648d05e91c2100874be061f2279ec2eed573f3d03c

Observation f8a61c0f-4a65-4baa-96ab-c1ffd58b532d · outbound

This paper cites Offline GRPO with KL penalty 𝛽=0.1, group size 8, learning rate 5 × 10−6, 2 epochs over the fixed trajectory pool.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Offline GRPO with KL penalty 𝛽=0.1, group size 8, learning rate 5 × 10−6, 2 epochs over the fixed trajectory pool

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.965819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:ec672121efc4001f1a747903044b46f92fafba57b268b17e67b4b13d82537e01

Observation 3a152dbc-7de6-42ae-b2d8-4957d8f33a28 · outbound

This paper cites trainable property.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design trainable property

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.973338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:e194a60c7979a70e7c6abe1f1971180648020dd435404155082169fff6a79823

Observation aff0ffc9-b847-40fa-bccd-7977240d7171 · outbound

This paper cites persisting with 47 Pareto.ai LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design May 2026 similar or domain-typical designs despite weak or stagnant outcomes.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design persisting with 47 Pareto.ai LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design May 2026 similar or domain-typical designs despite weak or stagnant outcomes

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.963578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:5a13ba522ad80e8b2a64352bc6e38c073a1690f90648f97d67ee0008e248e290

Observation 62dd09dd-052e-48bf-93d9-079de9b7490e · outbound

This paper cites A couple did show quite firm anchoring, persisting with changing one tiny thing that clearly wasn’t working.

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design A couple did show quite firm anchoring, persisting with changing one tiny thing that clearly wasn’t working

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.971581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:8437f357a4b3f042cb5dabc8be1d6f29ce3906eea209cec23d1a07b186d5cac8

Observation dd25b1d5-1d36-4a05-8c0c-2642ef900988 · outbound

This paper cites Fleiss’𝜅= 0.33across the three raters (fair agreement, supporting evidence, not confirmatory).

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design Fleiss’𝜅= 0.33across the three raters (fair agreement, supporting evidence, not confirmatory)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T15:52:38.969689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T15:50:01.064709Z digest=sha256:33567bec61ab3092c57ea3fd5c302e549de66ed3a9450a9be6f8d31b234e9ae7

Pith citing papers

Observation 220371d5-9ce2-4ba8-868b-da4c7c863555 · inbound

ArchEval: Measuring AI Agents as Computer Architects cites this paper.

ArchEval: Measuring AI Agents as Computer Architects LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T01:16:50.927415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:16:50.927415Z digest=sha256:e8add7928e1ab260d722dc6e71f704c491b8555d0150c49a1a3cc1b647eca502