Pith. sign in

Paper Citation Record · LEDGER

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 61 inbound Pith citation observations for arXiv:2411.15114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15114 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:33:13.231482Z

measured 125 of 125 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:27:30.618243Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved44
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 93e50550-a810-40fc-a4a7-8283917db63c · outbound

This paper cites Evaluating Large Language Models Trained on Code.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.069528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.069528Z digest=sha256:b937ae987d237c13d46a9fba119845e82d6f2c964e06dec5f28a381f38c84c73

Observation 9214f6e2-9035-4c84-9c81-1fd6d4366a98 · outbound

This paper cites Competition-level code generation with AlphaCode.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Competition-level code generation with AlphaCode

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.073102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.073102Z digest=sha256:8ecbf3d8515d25121a6c8642fbc19d52b05e2bc107a9b2f9ae099129fcf2bb3e

Observation d68dc253-f549-4593-aebf-c77561f0f0ee · outbound

This paper cites Textbooks Are All You Need.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Textbooks Are All You Need

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.076189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.076189Z digest=sha256:49a4c55c8860c1ff4301a301a51b8c89b6b22be103b91c987f90a3373a2b9a3d

Observation bea0344b-a812-44b7-aef2-106eb00a4911 · outbound

This paper cites The Llama 3 Herd of Models.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.078973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.078973Z digest=sha256:1162a5ad9d29fa5eeb9110ea128438ebfaac15bae098f585f82c6d8bb3188e5a

Observation 419e9304-59a9-477e-88c0-8b9dfed02aa1 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.703896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.081910Z digest=sha256:7c452881ff769c54529ce262ea13d0a2db828a7687164edf5d370dfb85d852c4

Observation 2f4af9d2-ff40-4e0e-9bec-fe388e460293 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.084827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.084827Z digest=sha256:6b2887b127145c07c5d5fff950b4408949de3e3a6d6ca6975477d7639e3d9add

Observation fb9d432a-3125-41da-b3f2-89fed418c256 · outbound

This paper cites Interviewing AI researchers on automation of AI r&d (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Interviewing AI researchers on automation of AI r&d (2024)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.697544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.088517Z digest=sha256:18c1d2106a4eb059ea9a6e8da883d28f819b068dcb8186ccffd1b3471d11c9a8

Observation d94f04f1-6f4d-4ba2-b655-c0a96dd5ba44 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 8

Resolution
verified exact
doi, observed 2026-08-12T14:33:13.252997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.090814Z digest=sha256:a0ff649848e5eba3ff5401bca1bb2f4a189d034593f8dc80ed4a2a209164ee0e

Observation 13b94930-6795-4441-aff6-0483157c500f · outbound

This paper cites Explosive growth from AI automation: A review of the arguments.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Explosive growth from AI automation: A review of the arguments

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.093306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.093306Z digest=sha256:3e4c951a38250af4337a286f78b54693256f952e2b565802cd2f6fd968514ff7

Observation 9108dd92-eba2-470c-b99b-80ec96d37abd · outbound

This paper cites OpenAI preparedness framework (beta).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OpenAI preparedness framework (beta)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.691722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.095724Z digest=sha256:6f82b23d72ca67927841386121c65d12e44c07d8efd170ee699628a9935c9f9f

Observation daad44ba-e91c-40f9-8909-3730c0508d4a · outbound

This paper cites Frontier safety framework, version 1.0.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Frontier safety framework, version 1.0

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.685477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.097831Z digest=sha256:34b7210d7cfbf0ad6f33cb77233edb1815b1e56c94bc61b9d3cf4c5ab1d92e02

Observation caa42555-43ef-46bc-9dda-6754402a07df · outbound

This paper cites Anthropic responsible scaling policy.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Anthropic responsible scaling policy

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.678875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.099932Z digest=sha256:ff6ac55391fe124249404bc39a5142c11c645bcab465408f75d9bf7b0c46f058

Observation cfac877e-9688-40de-b039-a5b053e7a108 · outbound

This paper cites Recital 110 of the eu artificial intelligence act (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Recital 110 of the eu artificial intelligence act (2024)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.671828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.101999Z digest=sha256:de5b559afd8229370a12b7d7e2495148f152fc626d3484e9e94d1c7c5e480d37

Observation 8825082f-4866-415c-8950-0e5c94ab0ab6 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.663824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.103890Z digest=sha256:ae2e02151083f0565bd24224205ac889026fc8b668065b07ca409896b40e9273

Observation a2741764-7fae-47ce-9f96-def74a3d4ac0 · outbound

This paper cites The bletchley declaration on AI safety (2023).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts The bletchley declaration on AI safety (2023)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.656928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.105870Z digest=sha256:c6565b175758af726e8ae49aa93f91ee99dbcbb0c5e653724de5586eb45b842a

Observation 2aebe51c-8d48-42b4-a261-88fb47af02e5 · outbound

This paper cites Frontier AI safety commitments, AI seoul summit 2024.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Frontier AI safety commitments, AI seoul summit 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.650122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.107747Z digest=sha256:f7b14b9f459f9e7a10ada9388bbc90294e60432f66ab5147df28cdfcc3dafd76

Observation 2d114eb8-0e65-4d24-81f6-aee212b8eaf4 · outbound

This paper cites Building an early warning system for llm-aided biological threat creation (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Building an early warning system for llm-aided biological threat creation (2024)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.643125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.109734Z digest=sha256:2b296e1bb7439d2efb71eb06404fa51544ccfbcd82a7e266358d6e94e382100d

Observation bfcb303c-7e77-41d0-9a92-19acb60f329a · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.112305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.112305Z digest=sha256:317872a5d433faa9bcd8e7ae61772375007112e0ce8e61b7441b3685cb506deb

Observation 21553f6f-d87d-4220-8f52-d8c10669d594 · outbound

This paper cites What a compute-centric framework says about takeoff speeds (2023).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts What a compute-centric framework says about takeoff speeds (2023)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.636145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.115443Z digest=sha256:10c3b21f5ed6c41a6a91aa688c431d2d52efac791675316f7ffadf72fbda8e37

Observation 049cc47b-1736-4ae4-ba5f-dd9eddbf9ed7 · outbound

This paper cites Sabotage Evaluations for Frontier Models.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Sabotage Evaluations for Frontier Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.117927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.117927Z digest=sha256:74b0205afc7fa7f9ead4af46ed44128133545d1960b23d5136cb3a765b7e1e7b

Observation 43984ef0-bde5-411f-8d9a-bb6bf903d172 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.628799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.120597Z digest=sha256:8dc2aa381b9a8ab43585ba987c914a9bb90304085385563b1ef06eda4b7c61d3

Observation b52f8e02-5694-430a-881c-ebe71f379032 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.622601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.122748Z digest=sha256:30af2210b07fe3515328690a958d83e567841df2606315aa913639f1b2d881bf

Observation dda5e2a4-9525-41fa-bd72-aba38ab4de70 · outbound

This paper cites Raising the bar on swe-bench verified with claude 3.5 sonnet (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Raising the bar on swe-bench verified with claude 3.5 sonnet (2024)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.616633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.125629Z digest=sha256:b4f8f84ff4c67f79e1337ec769236f53f4c7112333e21e104b009ed80e038ec5

Observation 055c5b1d-80d5-4728-9bcd-93629c551826 · outbound

This paper cites Artificial Intelligence Index Report 2024.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Artificial Intelligence Index Report 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.127921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.127921Z digest=sha256:57a3feefc7a28da5904e03bab8503addf0876a2891309810b0f0b861a0d59602

Observation 26c207d8-7c32-4898-8e65-e5ff9de0db1e · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.130371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.130371Z digest=sha256:d727999ecd42d66991258fe7fa5a2400d23963bbf1b864d0ad0053cea4c44018

Observation 357ddd75-67dc-4f11-8700-e9dae2fa006f · outbound

This paper cites SciCode: A Research Coding Benchmark Curated by Scientists.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts SciCode: A Research Coding Benchmark Curated by Scientists

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.132899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.132899Z digest=sha256:e4ec4f22370e52d9bac1d2d0fe548a39201aa32cf8aaad62b881aba2e115d2ee

Observation 1f305af8-669c-4c5f-bd9c-842163806679 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.135584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.135584Z digest=sha256:7066c0660fd125e32371a0efded45f6971267f7e3efbec269d4e91eec2e91b86

Observation 1d91d83a-f5ba-4e82-9d6a-ddcc1cad57c0 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts GAIA: a benchmark for General AI Assistants

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.138150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.138150Z digest=sha256:fa0feceaf6c11e1c76b9ee2bbbc983316c5c04dc729fd32095e76f4ac30a6994

Observation 5db7ccbf-5e43-407d-99a4-bbdb04f4c4d7 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.610070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.140723Z digest=sha256:ffced5d128196e2ecbdf6c6d3bffa3b4e374624fc3a7d20c3a72fe38318671a0

Observation 33f1df9b-aaf8-44c7-b2d6-80125436611e · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.143333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.143333Z digest=sha256:66604136a37280a46aabef0b2705bbf2bbe62b38701f21352d466295054c2071

Observation 482e5e49-c928-48db-80dd-957436abcce4 · outbound

This paper cites MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.146051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.146051Z digest=sha256:09134d88194fa74811bf5f4ea0e962c6fd4f409084847155082ad055fe0ec9d1

Observation 70053717-4393-4d76-a4f4-d829c6c4311b · outbound

This paper cites DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.148692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.148692Z digest=sha256:9e65303d56d8613ad079300e9409ed23a92bdbf204fe58d3bd84004484022b03

Observation ab39c1ef-c811-4eb3-9f1c-5e419341688d · outbound

This paper cites H-ARC: A Robust Estimate of Human Performance on the Abstraction and Reasoning Corpus Benchmark.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts H-ARC: A Robust Estimate of Human Performance on the Abstraction and Reasoning Corpus Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.151881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.151881Z digest=sha256:fa80d02c640622d9e9efc1d7ff58cfe1a0b78f0892dbae6035c47cdd7b922c38

Observation 7add920a-c7e6-4f45-bd36-39f67fc91252 · outbound

This paper cites Training language models to follow instructions with human feedback.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.154509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.154509Z digest=sha256:652905cfc9fd7bb3c20328b576a5dd909f6a6b59abd31c43c0f03325e74efed1

Observation 9a730257-0cea-427b-94c2-65df035f2290 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Constitutional AI: Harmlessness from AI Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.157074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.157074Z digest=sha256:7ae9a6c6ad3190f5da7f4fc2422cf6853df59d9e333f7f51fbd898f9a8c23e17

Observation 36689a9d-015e-49bd-955b-1b15cf741994 · outbound

This paper cites Nemotron-4 340B Technical Report.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Nemotron-4 340B Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.159698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.159698Z digest=sha256:72876e061120c5f97ec863460619f21cb64b609db2995853354e22a70426339d

Observation be93dd10-ed34-4eaf-a0ed-340d20b151c3 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.162197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.162197Z digest=sha256:381fe57b4c97f5582340eebba5c0f321859c491c20cd5db909bbe08c2666d1a4

Observation 9306505f-d8a0-4dff-aa71-b43aee00c239 · outbound

This paper cites EvoPrompting: Language Models for Code-Level Neural Architecture Search.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts EvoPrompting: Language Models for Code-Level Neural Architecture Search

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.165447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.165447Z digest=sha256:95f9a8308d2585396854c738debe88ce597ecb01a1ca8c3f430d77370fdb7aaa

Observation 4cf5c454-0308-43ae-8b62-b4ad963cb9ed · outbound

This paper cites DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.167944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.167944Z digest=sha256:9f7d268f12ecf70647ea02bd2fa98ff3cf4bb1524d089bd27244d2846e61ce29

Observation 0306d3bd-d6ea-4cfb-9ff7-ef2b4388d32f · outbound

This paper cites SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.170581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.170581Z digest=sha256:3ff571f7c4b5a7d42b194491900e679ffdf16c509fe90d92ef6132e67f8a8b3e

Observation 97edb651-32a6-4773-b83e-73d493d1b791 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.173140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.173140Z digest=sha256:5b544c464064e44b69d398700189e8a402e00eae395c4a9a645a19c8617ce47a

Observation d4886c2d-9919-4e81-a473-72b7b55735e7 · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.176588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.176588Z digest=sha256:b0805398acfab605a4c34f5669a823acb24d87ad5149f8147b7d480eecfc8131

Observation 619dfc4e-82ba-4856-ba1e-e93636af51f6 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.179246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.179246Z digest=sha256:745f68cbfe376aeb898f0762cbb81e6c97a10d908e39efa2bbe5849bc05b2d5d

Observation 943925d1-4f81-4b0c-8197-e482e3e8d8bd · outbound

This paper cites Discovering Preference Optimization Algorithms with and for Large Language Models.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Discovering Preference Optimization Algorithms with and for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.181972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.181972Z digest=sha256:a03d6ccc262f1cb39456512aae70b591dd786ebc7c2e2f941c6ab18da6e621fc

Observation 374b2181-0d01-4879-a464-a963c73c8b91 · outbound

This paper cites SciAgent: Tool-augmented Language Models for Scientific Reasoning.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts SciAgent: Tool-augmented Language Models for Scientific Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.184324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.184324Z digest=sha256:0dc55ebe15c5d4171faffc0442894d686e4e05707acd6bb45fc741999730ae0d

Observation ce8cc2a6-f7fa-485e-af96-2020ce6cf0ce · outbound

This paper cites ChemCrow: Augmenting large-language models with chemistry tools.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts ChemCrow: Augmenting large-language models with chemistry tools

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.186866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.186866Z digest=sha256:0246904847c16df76e5d094e8fac603217750a9933ba9df86377bbc0f2fe84dd

Observation b46ed719-5073-413b-b47a-0747638c7318 · outbound

This paper cites Scientific Large Language Models: A Survey on Biological & Chemical Domains.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Scientific Large Language Models: A Survey on Biological & Chemical Domains

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.189430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.189430Z digest=sha256:a34f2165179ec373169c6d69d55f27176f5b05c2323de2e451eb104b3e3ac7b3

Observation 9a49e1d6-cb3f-4d06-a864-f12545dc0181 · outbound

This paper cites AutoML in the Age of Large Language Models: Current Challenges, Future Opportunities and Risks.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts AutoML in the Age of Large Language Models: Current Challenges, Future Opportunities and Risks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.191545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.191545Z digest=sha256:57542637cd0481872dfa5457d6efeb7a460433ddb64cd6288bc92c11026cc6af

Observation d7a33066-7114-4595-95a8-d5e6ea53ad7a · outbound

This paper cites Chip Placement with Deep Reinforcement Learning.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Chip Placement with Deep Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.194082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.194082Z digest=sha256:3fc6e40eaff2db93ab99b0730843644866233160897e67816622403c518e80e0

Observation bd8aacd5-e317-4e49-aa2b-d1e8e4d17979 · outbound

This paper cites & Aydos, G.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts & Aydos, G

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.603105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.196795Z digest=sha256:28156141ef38edcd2a01ccce3980764a0c55a34a1680d8b5045593b959197f30

Observation 27a19e35-bd54-415a-8ae2-b056bd5b6fbc · outbound

This paper cites Vivaria: Open-source platform for agent evaluations (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Vivaria: Open-source platform for agent evaluations (2024)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.596263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.199205Z digest=sha256:d736e1b2a47447a599a52589a6d971cf1fdea1b186086b57a00ecf5b777abedb

Observation 2a51f86b-eb88-4539-b1a0-ff870c907679 · outbound

This paper cites Claude 3.5 Sonnet model card addendum (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Claude 3.5 Sonnet model card addendum (2024)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.589134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.201630Z digest=sha256:0d90e5c4329196f5a693887ab59196db79ed08e24031a7b1fd2a1c90cc7dd454

Observation b537316d-67d4-46f4-a304-82e7eff5addb · outbound

This paper cites OpenAI o1 system card (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OpenAI o1 system card (2024)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.580503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.204017Z digest=sha256:f89113eaeb23d9519db3930608f00395f4585d25715260602ee88d39985c7440

Observation b7390dc8-00a2-4757-98b3-64235c222ea9 · outbound

This paper cites AIDE: Data science automation technical report (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts AIDE: Data science automation technical report (2024)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.573410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.206318Z digest=sha256:1ceeaa1eb3db9e89b93c302e0b5ca1f63d35a8fb959a2c191c08d233fe7d4c5c

Observation 7693bafe-b60a-40ed-85b6-4579cff93977 · outbound

This paper cites Details about METR’s preliminary evaluation of OpenAI o1-preview (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Details about METR’s preliminary evaluation of OpenAI o1-preview (2024)

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.565835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.208642Z digest=sha256:8ee0f32ad5a0bf3a2f3aea036a27eb26a8205cd6da5058d0f6c19aeb78eb3324

Observation a4dbaa28-eb60-414f-96d8-590c9bce99bc · outbound

This paper cites Claude 3.5 Sonnet: Quality, performance & price analysis (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Claude 3.5 Sonnet: Quality, performance & price analysis (2024)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.558707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.210938Z digest=sha256:dec31a58e635b22ab5ed207fdebe682eb8556d8d5eb04ef6b34b36252037d5ae

Observation 4008e77c-c644-4c3c-a78f-76a5be9c7e79 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.552270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.213202Z digest=sha256:9e83ddae490fd8ee7130bf02ac357e336a2a5c383d5da2c645647d19a18ae034

Observation 02f7ff6e-504e-4a93-bc58-097c5a6ada3c · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.215554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.215554Z digest=sha256:0e15506b858a2792f0b65ec75222c00077f049bdc41fa1f2835095a3dfa60383

Observation b023ccb1-9588-4ffa-bd1e-97c18c884f76 · outbound

This paper cites Is this score similar to what you predicted or measured yourself, or does it come as a surprise?.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Is this score similar to what you predicted or measured yourself, or does it come as a surprise?

Reference 59

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T14:33:13.545394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.218367Z digest=sha256:be0b4dbc039165b9cf44ff87f8cac3061bab614d6d3a2a78c6477e1c4a87a6be

Observation 46c3d2b7-a08a-4178-868e-fe1ab97cb6d9 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.536710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.222060Z digest=sha256:bab840f586b3772a75e552d9ce852d72632aff9c0c6f6494ae262bfc341701f6

Observation 0f6e2c1e-2e63-4d88-9548-82a0c955aa1f · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.529316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.224373Z digest=sha256:c8c61ce8b30a8dd6cfc5b89d77a3af70c2146d4f640cc3346bc58c3decac57fb

Observation 6231b3b9-ee6a-4a25-baf1-34cc040b76dd · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.521984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.226781Z digest=sha256:46c5fae91dcd05e7fe67f34f661b0f61165ffbc189b8fae5ddedbbd45a3dfeb1

Observation 9a6f4ce8-f8af-4924-a545-1e699ff2598a · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.514722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.229141Z digest=sha256:572631e1b7cc72925b2e09181ff6245c16b52119aa96fd095c39a7042cab45f3

Observation f6629272-2716-425d-b21e-14db69ae21ed · outbound

This paper cites cuda" "cuda.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts cuda" "cuda

Reference 64

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T14:33:13.507078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:33:13.231482Z digest=sha256:2a74e5f78d51f482f4203ade7d70f166e04ce16bec0e553101e4af7321ae5fb6

Pith citing papers

Observation 4b89afad-6c15-4cd6-a734-fdf2480c86bc · inbound

Frontier Models are Capable of In-context Scheming cites this paper.

Frontier Models are Capable of In-context Scheming RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:22:01.681340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-16T14:22:01.616448Z digest=sha256:3738b5416d27a5529726d699798f1d58cc4972898831e1f6ad91fbc94b66c114

Observation d9fa4ac5-63e7-45de-986b-c9ba862334fb · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.326504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:d6fef5d853a3bda0ee0881701784dffff99232ca3f81b42fb5f4735195d07335

Observation 7dddfdc1-9de0-4ea1-b0b6-ecb02f46e6d0 · inbound

The AI Agent Index cites this paper.

The AI Agent Index RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T14:48:35.872296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:48:35.872296Z digest=sha256:533822e44d4e1db6cb0c3280de09d9c07d791639f41ec8be9c96f0d57e6ab336

Observation 0db8ffb1-e533-489d-b388-a0a9ea0c6fec · inbound

KernelBench: Can LLMs Write Efficient GPU Kernels? cites this paper.

KernelBench: Can LLMs Write Efficient GPU Kernels? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:55:02.090533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T16:55:01.976356Z digest=sha256:179cf4ce037cd0dc855daa5bb2d8cddd57beeabc26b13a36dfb0556ed82e07b6

Observation 19dcc8fb-763b-47e7-8ec9-771e5de241fe · inbound

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions cites this paper.

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 195

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:30.618243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:27:30.618243Z digest=sha256:faab58e6d2a1918208e684066e0939361dff58e46f4f1b0727d88a7c7ff41727

Observation 35d19ee8-5d53-4ede-b2c2-bd9c4fc69316 · inbound

LLMs Outperform Experts on Challenging Biology Benchmarks cites this paper.

LLMs Outperform Experts on Challenging Biology Benchmarks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:50:49.704789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:50:49.704789Z digest=sha256:3fc2f93c196f5919eb36bc8ab84b960d0c072d6054abe80009989fcdacdad9dd

Observation 79ec15b4-8781-468b-8d3e-beeea998a9e1 · inbound

TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents cites this paper.

TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:59.964481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:18:59.964481Z digest=sha256:f52a3e1164ba3651f71c2f74d859db4f98e09ca76712cdba9215c01746b51bf1

Observation e5b67da1-0357-4675-8eb8-5547e5883c99 · inbound

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks cites this paper.

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.881071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.881071Z digest=sha256:4cc5d8bc5a3a3f1262d7fbd317103c3ffe5de0129c1ade473e1b4c9e482f8786

Observation 603c9afb-9026-40fe-9f5a-dc6febb693fb · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:58.052862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:58.052862Z digest=sha256:ea0f53e8097aff79c738d2abf5466eaf253cb4a5de07bda3d7cb9acb7385a98c

Observation f650702d-9842-4d40-99ca-3e58f8cf9c10 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.952422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.952422Z digest=sha256:f9dd06ecc215d37107a12abb1caec4b9da552f217b8e038cfef346afc44f71a7

Observation 36100b82-0a6a-453d-8fd8-7252d4bb2915 · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.421486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.421486Z digest=sha256:c55f1d4ffed103fe51932d2596c5d95bbbddec9c9476ff4cacf345aba7dce311

Observation b3799e82-49f1-4e93-b917-0b42741574ce · inbound

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework cites this paper.

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:59.279843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:59.279843Z digest=sha256:be7d02542125b24cba13ae833c8178249737998d7875623d1c8d38a19047086e

Observation a58694cb-4e26-4f00-876c-ee10cd3eb27b · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:52:07.806483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:70a5940926972b694076bc62ad7dfaba33d19c473d0377cf6343fde773e627da

Observation 9e58febf-03fe-4d5f-8648-635514eec14a · inbound

AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research cites this paper.

AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:00:08.585286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:00:08.585286Z digest=sha256:a01f14d84bc565edbfbc6fe054b7c5e2cf7be00afee00f657f9e1c5b3924b546

Observation 09cd1878-e82b-4063-bde5-4ff58152aeb2 · inbound

Exploring Design of Multi-Agent LLM Dialogues for Research Ideation cites this paper.

Exploring Design of Multi-Agent LLM Dialogues for Research Ideation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.153839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.153839Z digest=sha256:e674d241e47498c2cdb2b56ec9f01f1ab62a2b97f1d8430ff5312bd984072d71

Observation 203e8758-23fb-496e-bb2c-9753f4b9bdb4 · inbound

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity cites this paper.

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:29.939412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:29.939412Z digest=sha256:1204087be24a07a17285fe13174a21a21419c9548c76adc9935a5d93beaebeed

Observation 7eca4a44-634e-4288-828e-ad2fcd68d077 · inbound

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety cites this paper.

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:19:44.776446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T14:19:44.695462Z digest=sha256:20b02c9c35ab60cbfd4bec4308c459039c348b94ec208008735c8aea2ae5aae3

Observation f89e799f-9104-420c-9c93-7bfaa6b22443 · inbound

Scheming Ability in LLM-to-LLM Strategic Interactions cites this paper.

Scheming Ability in LLM-to-LLM Strategic Interactions RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:51:03.835431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T07:50:30.597108Z digest=sha256:ec10cc462ce37988575f0aecae656a666c36de3340bcaa0193187d03838d1b56

Observation 76b13409-8d2b-4366-afea-6a06c93a1a20 · inbound

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies cites this paper.

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:57:24.639150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T05:53:29.860037Z digest=sha256:5a17c9db12b95f67f6a1a01b75c15280331bca1014b2010d0f637608b7786238

Observation 9126949c-9665-45a0-b476-7436e91deb07 · inbound

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy cites this paper.

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T20:36:06.369914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:36:06.369914Z digest=sha256:3f70c4cc2a96fc10b42f5f195d0b64ba71d7893abe453de433f9a9b7b03d46a6

Observation b030444d-0020-438e-a874-0edb4351c374 · inbound

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization cites this paper.

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.739286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:17:32.290531Z digest=sha256:3fed2cece00446916607a3b230ab08a356c0ffc19e2548a67864375b45397b2b

Observation 8c7a22b9-3c70-4efa-97fd-886a14846ce8 · inbound

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration cites this paper.

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:46:34.653244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T12:34:29.808503Z digest=sha256:252e6a1b6c088f5670460e5003d594816edfd7642fac8e8fe96ef0e6ff8b4895

Observation c7d436cc-01f1-4907-92f6-2883a4edf091 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:16.444053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:b02f5a2a36cd6cf71310b477116367f5ab735f8e8eedfac20ae283ee24b21360

Observation 43f82950-7e26-4aa2-994c-4d56fa16925b · inbound

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation cites this paper.

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:05:36.486105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T18:55:02.645026Z digest=sha256:affd2240743be3dd4ca4c214769d8d3df0c43cd6fc64565ba51efe6697ada9d0

Observation b4e7854c-6681-4446-ae70-14c4cba12e9a · inbound

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction cites this paper.

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:05:06.762989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-15T06:04:03.605898Z digest=sha256:022fde0e5ee8ef3f50374a17436b472f450210b5abd228b533246c6e0dba5318

Observation 002c6092-2b3e-4304-a6b9-38cc4ca0c284 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:03:29.062735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:9b16db98067f435933df1fb0898e0c7f46b8c7317cba50cc71838fd7f59dd870

Observation f16d3fd5-9469-4b23-8d75-671c0be93582 · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.448976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T14:25:15.565386Z digest=sha256:5abaaf574e15262a7fcd4e8b5f9d0d121f7dfaf166347f8d19223792b20492b6

Observation 98f353b7-e769-4be9-910b-36a33c63fd71 · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.924661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T19:00:30.961402Z digest=sha256:bb06bfbf23d4f33821b1e374312427ec14f42c86a7c8ac1c541001a9892c3ffc

Observation f9630578-942a-41d2-9260-a9449fdd468e · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 222

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.557130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T10:30:50.256635Z digest=sha256:4f6e8b579369bccbd86ba90bde1d09774a103bd6f388e213f6efa72bf90c373c

Observation 5328db6c-78ad-45aa-86de-e3618d706e3f · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:49.951321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:49.951321Z digest=sha256:c1a0816d13b2963910603eafb33ba3e2166c69bf5fc9293fbcae77f4530df5cd

Observation 321ddc94-9675-4172-88d5-3f37ac4c7741 · inbound

How Far Are We From True Auto-Research? cites this paper.

How Far Are We From True Auto-Research? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T09:58:11.242622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T09:56:16.160551Z digest=sha256:c292869ec620f3fb728b5d0549938dc8d0f0a4ed9d65737d96db7d39f50e58d4

Observation 95e6e6dd-7a45-4c86-a98d-45990b1871f9 · inbound

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale cites this paper.

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:59:45.559572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T06:56:27.532299Z digest=sha256:3fd637a3074676eb1c6e2778f55b1dcd69c9002280dbd2c01a93b88765859f7d

Observation 23a23acf-0525-40c6-9917-a28501eda992 · inbound

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems cites this paper.

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:33:32.798892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T15:32:21.737028Z digest=sha256:69ecf41710f212c4b2566781d2c327be331b0d57c9f36fb1b6ba700bda86b44c

Observation abf9f5f6-6c96-483c-8388-2bd6d7e356b7 · inbound

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development? cites this paper.

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.690735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T06:29:42.665765Z digest=sha256:b6c91716920eb5f31b9611a839613c1e0dc7d6410423031ad31b06c785500746

Observation 52945ae3-1b6f-4313-b036-adb3f275750d · inbound

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks? cites this paper.

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:46:46.619488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T06:37:03.670571Z digest=sha256:ecddfdc045d4d98fce2d9644e48aaa9d8962a0eaa314cdf5d4cec1f8eaf702cf

Observation 9668af3d-f1e5-45f1-98eb-3e0245013815 · inbound

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees cites this paper.

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.255049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T14:06:59.471772Z digest=sha256:fd2804742000214a1bccdb55d7587b5a5b34106dffa0f42200dc8ff866ea8fcc

Observation efde55ab-60f9-4aae-8e1d-c230c3c7495a · inbound

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement cites this paper.

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.033546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T09:34:41.800309Z digest=sha256:8b546ed539f4bf930dd2006d3150c623044a060e7ddd2a1f927c6a724288e83e

Observation 4cc5987e-a720-497d-9c48-74c0d0c44ba6 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:08:58.899735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T23:48:58.497927Z digest=sha256:027a1ea90308be9b7462038c42cb5b12f0afb8c2d427e0592606c42118b08e9c

Observation 5d39a32f-a09d-4ee2-a9dc-45b587457b01 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T11:06:03.847203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:06:03.847203Z digest=sha256:6443cfded5d10bd389dd119423b42145a693f9226296a276aad2b57288a71cc0

Observation 4806bec1-2f3f-447a-b62b-400a84e06fa3 · inbound

Learning the ARTS of Search for Automated Discovery cites this paper.

Learning the ARTS of Search for Automated Discovery RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:41.833473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T12:04:13.117307Z digest=sha256:23765d668ec084525433b209fc3ac09802b3bd4d80aa1b005b0fcf85c094b069

Observation 1ac7c5ff-0cf2-4e78-b901-f62eaf2d6854 · inbound

Discovering Crystal Structure Prediction Algorithms with an AI Co-Scientist cites this paper.

Discovering Crystal Structure Prediction Algorithms with an AI Co-Scientist RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:46.987416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T09:03:00.675456Z digest=sha256:98027aa894299330a085bcbb459798da695c8bda4da03569ef6f9905a399c83a

Observation a0aaa883-453e-4868-ac34-d39a298bd0e0 · inbound

MirrorCode: AI can rebuild entire programs from behavior alone cites this paper.

MirrorCode: AI can rebuild entire programs from behavior alone RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.703971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T06:35:32.667865Z digest=sha256:08c084cf78f6bb508f4f8a3e122f74c1b19abdc0809e93ac5da118cce29dd9a3

Observation c974db13-d3e9-4d21-8d11-6794418d9f38 · inbound

MirrorCode: AI can rebuild entire programs from behavior alone cites this paper.

MirrorCode: AI can rebuild entire programs from behavior alone RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T09:35:02.886229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:35:02.886229Z digest=sha256:347dfdded85dc3447406893c417426bcd999ad0e1e2c16ebaa3fdd698a46ca7f

Observation 9c361421-912f-4c8e-8c9b-46ec8d5cb72a · inbound

Two AI Metrics Diverged: Will it Make All the Difference? cites this paper.

Two AI Metrics Diverged: Will it Make All the Difference? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:36:56.156858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-02T12:29:24.439779Z digest=sha256:bde5d7697fc0c6ea7d6b219777d7c6e96894f4ff5d065dd1e76f0be2c82a9f8f

Observation 4031574a-1bc3-4e7d-84a4-046c5a8f602e · inbound

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications cites this paper.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:04:44.520871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T06:55:40.830922Z digest=sha256:e589c9bbb1043d66bd9ae9238a167bad5ace77d4955be321afa40689482b71fd

Observation 58fcc192-7e31-4da4-944f-af8b18859fb3 · inbound

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications cites this paper.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.147998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.147998Z digest=sha256:562d408458868d27510957265461e65f5643ce2d41d8345d1a4fa17bfd5e8119

Observation 1296820b-a97b-4a1e-ab5e-dfdc4571568b · inbound

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading cites this paper.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T05:28:45.311405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:28:45.311405Z digest=sha256:cceb873daa8cdc955fc0d26aed7976928d3b69e65f01d19bef9108634a2932ce

Observation c4e9e2a6-d05e-4c08-b8cf-89f4e4b97842 · inbound

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading cites this paper.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:c70563468fd18bdcb420138639509f18e1250c2880c9e4d22cd79e90f2e193f1

Observation 4300af11-4056-4bf9-af7a-3b20f2f358d5 · inbound

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation cites this paper.

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T10:46:01.272433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:46:01.272433Z digest=sha256:0f4be3b2c32cec459ba7e8cfa0f4d0ed4cb3c9d3d9a7fb9caf70f6d1b5b08296

Observation 45bc81b3-20ed-4ca4-a3a7-bf39ecc86496 · inbound

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D cites this paper.

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T12:50:00.490138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:50:00.490138Z digest=sha256:8cc8e5151bf05d09b15ecd01c77b7d76427c992e83630d79418dcadda89e7e3c

Observation 4082e15f-b500-40b7-96f1-5feb288df4fa · inbound

Efficiency Matters in Autonomous Research cites this paper.

Efficiency Matters in Autonomous Research RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T09:14:06.207902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T09:14:06.207902Z digest=sha256:18cf6807109354c575e8e892fd664f87da03e2d9d305b1b1055fe4915d44ab9d

Observation d6f79981-5a6c-4191-b434-dd3ce80357e7 · inbound

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation cites this paper.

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T01:16:43.437896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:16:43.437896Z digest=sha256:3c759e2f5c25217297525ff98fe50b2369ccbecb48b2a370b7bc9afdef6829b1

Observation d0bf2f4c-2256-4be0-95f0-6a8f25daeea7 · inbound

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs cites this paper.

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:22.349139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:25:22.349139Z digest=sha256:470a87e5b3a508966db6266ddbf2946b3239a0183fee14e73c082fa879bec671

Observation c77e66a2-c16c-433a-b39f-7b6101b3ccfc · inbound

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing cites this paper.

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T01:30:17.866119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:30:17.866119Z digest=sha256:3a190f5903f150ff33446eed7d6f5cf1411363c26297ded3bb4e208d2079a693

Observation a0cc8a4e-0e7d-46b8-a1fd-d9b77f400289 · inbound

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning cites this paper.

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:16:45.749872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:16:45.749872Z digest=sha256:5d3302530c8dbacf002bbc6814e4d17aa64d955e16150bc4e0d658f9080ae424

Observation fd24008f-7d94-46d1-b45c-723447228ba8 · inbound

Predicting Task Difficulty Without Rollouts cites this paper.

Predicting Task Difficulty Without Rollouts RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.492888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.492888Z digest=sha256:699d4bd86bce87390c871b7e5b662208fd1b22258c7938054cd19acb69004c63

Observation c3ee986a-611a-463a-941e-2e856cb77cad · inbound

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization cites this paper.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.759327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.759327Z digest=sha256:1f0ccabce74ff1242802154dd571df80c35a997dccd14e9d1d37328075d7a37d

Observation b1c4f2d5-3f9b-4e58-a2ea-052cc44da010 · inbound

QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing cites this paper.

QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:24:44.676649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:24:44.676649Z digest=sha256:6e9fdfcff532a5bb8fa2d1aa26965c2577eb92554837f265e8a200e5066c3f6b

Observation 06f8b810-3b52-4410-a6bc-6f129843f1ff · inbound

Evo-Bench: Can Language Models Improve Agent Harness? cites this paper.

Evo-Bench: Can Language Models Improve Agent Harness? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:44:24.561432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:44:24.561432Z digest=sha256:5dbe6efe76a104ae59c9fa7b626f5904d5908ecd87f82ea23db77f56825c3642

Observation 40910431-7f02-4f95-ab99-34b730e701de · inbound

Evo-Bench: Can Language Models Improve Agent Harness? cites this paper.

Evo-Bench: Can Language Models Improve Agent Harness? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.843491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.843491Z digest=sha256:7371aeff7a20983d8283bacca281b2591ec7f359aa32fb7a7b886ae9de8a7ba1

Observation b631cdce-8a01-4068-a59f-eb8bc494689e · inbound

VALG: An Agentic System for ML Theory Research cites this paper.

VALG: An Agentic System for ML Theory Research RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:44:14.918108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:44:14.918108Z digest=sha256:8abaa9d0da5168238d4ce095d44601bbd7fdc46a00fcc011306b4e815b694148