Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:33:13.231482Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 61 inbound Pith citation observations for arXiv:2411.15114.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:33:13.231482Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:27:30.618243Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
64 of 64 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
Observation 93e50550-a810-40fc-a4a7-8283917db63c · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Evaluating Large Language Models Trained on Code
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9214f6e2-9035-4c84-9c81-1fd6d4366a98 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Competition-level code generation with AlphaCode
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68dc253-f549-4593-aebf-c77561f0f0ee · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Textbooks Are All You Need
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bea0344b-a812-44b7-aef2-106eb00a4911 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts The Llama 3 Herd of Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419e9304-59a9-477e-88c0-8b9dfed02aa1 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2f4af9d2-ff40-4e0e-9bec-fe388e460293 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9d432a-3125-41da-b3f2-89fed418c256 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Interviewing AI researchers on automation of AI r&d (2024)
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d94f04f1-6f4d-4ba2-b655-c0a96dd5ba44 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 13b94930-6795-4441-aff6-0483157c500f · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Explosive growth from AI automation: A review of the arguments
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9108dd92-eba2-470c-b99b-80ec96d37abd · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OpenAI preparedness framework (beta)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation daad44ba-e91c-40f9-8909-3730c0508d4a · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Frontier safety framework, version 1.0
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation caa42555-43ef-46bc-9dda-6754402a07df · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Anthropic responsible scaling policy
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cfac877e-9688-40de-b039-a5b053e7a108 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Recital 110 of the eu artificial intelligence act (2024)
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8825082f-4866-415c-8950-0e5c94ab0ab6 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a2741764-7fae-47ce-9f96-def74a3d4ac0 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts The bletchley declaration on AI safety (2023)
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2aebe51c-8d48-42b4-a261-88fb47af02e5 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Frontier AI safety commitments, AI seoul summit 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2d114eb8-0e65-4d24-81f6-aee212b8eaf4 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Building an early warning system for llm-aided biological threat creation (2024)
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bfcb303c-7e77-41d0-9a92-19acb60f329a · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21553f6f-d87d-4220-8f52-d8c10669d594 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts What a compute-centric framework says about takeoff speeds (2023)
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 049cc47b-1736-4ae4-ba5f-dd9eddbf9ed7 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Sabotage Evaluations for Frontier Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43984ef0-bde5-411f-8d9a-bb6bf903d172 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b52f8e02-5694-430a-881c-ebe71f379032 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dda5e2a4-9525-41fa-bd72-aba38ab4de70 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Raising the bar on swe-bench verified with claude 3.5 sonnet (2024)
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 055c5b1d-80d5-4728-9bcd-93629c551826 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Artificial Intelligence Index Report 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c207d8-7c32-4898-8e65-e5ff9de0db1e · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 357ddd75-67dc-4f11-8700-e9dae2fa006f · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts SciCode: A Research Coding Benchmark Curated by Scientists
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f305af8-669c-4c5f-bd9c-842163806679 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d91d83a-f5ba-4e82-9d6a-ddcc1cad57c0 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts GAIA: a benchmark for General AI Assistants
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db7ccbf-5e43-407d-99a4-bbdb04f4c4d7 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 33f1df9b-aaf8-44c7-b2d6-80125436611e · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 482e5e49-c928-48db-80dd-957436abcce4 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70053717-4393-4d76-a4f4-d829c6c4311b · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab39c1ef-c811-4eb3-9f1c-5e419341688d · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts H-ARC: A Robust Estimate of Human Performance on the Abstraction and Reasoning Corpus Benchmark
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7add920a-c7e6-4f45-bd36-39f67fc91252 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Training language models to follow instructions with human feedback
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a730257-0cea-427b-94c2-65df035f2290 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Constitutional AI: Harmlessness from AI Feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36689a9d-015e-49bd-955b-1b15cf741994 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Nemotron-4 340B Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be93dd10-ed34-4eaf-a0ed-340d20b151c3 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9306505f-d8a0-4dff-aa71-b43aee00c239 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts EvoPrompting: Language Models for Code-Level Neural Architecture Search
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf5c454-0308-43ae-8b62-b4ad963cb9ed · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0306d3bd-d6ea-4cfb-9ff7-ef2b4388d32f · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97edb651-32a6-4773-b83e-73d493d1b791 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4886c2d-9919-4e81-a473-72b7b55735e7 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Eureka: Human-Level Reward Design via Coding Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 619dfc4e-82ba-4856-ba1e-e93636af51f6 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 943925d1-4f81-4b0c-8197-e482e3e8d8bd · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Discovering Preference Optimization Algorithms with and for Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 374b2181-0d01-4879-a464-a963c73c8b91 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts SciAgent: Tool-augmented Language Models for Scientific Reasoning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce8cc2a6-f7fa-485e-af96-2020ce6cf0ce · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts ChemCrow: Augmenting large-language models with chemistry tools
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46ed719-5073-413b-b47a-0747638c7318 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Scientific Large Language Models: A Survey on Biological & Chemical Domains
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a49e1d6-cb3f-4d06-a864-f12545dc0181 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts AutoML in the Age of Large Language Models: Current Challenges, Future Opportunities and Risks
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a33066-7114-4595-95a8-d5e6ea53ad7a · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Chip Placement with Deep Reinforcement Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd8aacd5-e317-4e49-aa2b-d1e8e4d17979 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts & Aydos, G
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 27a19e35-bd54-415a-8ae2-b056bd5b6fbc · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Vivaria: Open-source platform for agent evaluations (2024)
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2a51f86b-eb88-4539-b1a0-ff870c907679 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Claude 3.5 Sonnet model card addendum (2024)
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b537316d-67d4-46f4-a304-82e7eff5addb · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OpenAI o1 system card (2024)
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b7390dc8-00a2-4757-98b3-64235c222ea9 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts AIDE: Data science automation technical report (2024)
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7693bafe-b60a-40ed-85b6-4579cff93977 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Details about METR’s preliminary evaluation of OpenAI o1-preview (2024)
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a4dbaa28-eb60-414f-96d8-590c9bce99bc · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Claude 3.5 Sonnet: Quality, performance & price analysis (2024)
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4008e77c-c644-4c3c-a78f-76a5be9c7e79 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 02f7ff6e-504e-4a93-bc58-097c5a6ada3c · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b023ccb1-9588-4ffa-bd1e-97c18c884f76 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Is this score similar to what you predicted or measured yourself, or does it come as a surprise?
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 46c3d2b7-a08a-4178-868e-fe1ab97cb6d9 · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0f6e2c1e-2e63-4d88-9548-82a0c955aa1f · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6231b3b9-ee6a-4a25-baf1-34cc040b76dd · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9a6f4ce8-f8af-4924-a545-1e699ff2598a · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f6629272-2716-425d-b21e-14db69ae21ed · outbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts cuda" "cuda
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4b89afad-6c15-4cd6-a734-fdf2480c86bc · inbound
Frontier Models are Capable of In-context Scheming RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d9fa4ac5-63e7-45de-986b-c9ba862334fb · inbound
Humanity's Last Exam RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7dddfdc1-9de0-4ea1-b0b6-ecb02f46e6d0 · inbound
The AI Agent Index RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db8ffb1-e533-489d-b388-a0a9ea0c6fec · inbound
KernelBench: Can LLMs Write Efficient GPU Kernels? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 19dcc8fb-763b-47e7-8ec9-771e5de241fe · inbound
AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 195
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d19ee8-5d53-4ede-b2c2-bd9c4fc69316 · inbound
LLMs Outperform Experts on Challenging Biology Benchmarks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79ec15b4-8781-468b-8d3e-beeea998a9e1 · inbound
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5b67da1-0357-4675-8eb8-5547e5883c99 · inbound
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 603c9afb-9026-40fe-9f5a-dc6febb693fb · inbound
TextAtari: 100K Frames Game Playing with Language Agents RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f650702d-9842-4d40-99ca-3e58f8cf9c10 · inbound
Deep Research Agents: A Systematic Examination And Roadmap RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 118
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36100b82-0a6a-453d-8fd8-7252d4bb2915 · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3799e82-49f1-4e93-b917-0b42741574ce · inbound
Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58694cb-4e26-4f00-876c-ee10cd3eb27b · inbound
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9e58febf-03fe-4d5f-8648-635514eec14a · inbound
AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09cd1878-e82b-4063-bde5-4ff58152aeb2 · inbound
Exploring Design of Multi-Agent LLM Dialogues for Research Ideation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 203e8758-23fb-496e-bb2c-9753f4b9bdb4 · inbound
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eca4a44-634e-4288-828e-ad2fcd68d077 · inbound
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f89e799f-9104-420c-9c93-7bfaa6b22443 · inbound
Scheming Ability in LLM-to-LLM Strategic Interactions RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 76b13409-8d2b-4366-afea-6a06c93a1a20 · inbound
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9126949c-9665-45a0-b476-7436e91deb07 · inbound
SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b030444d-0020-438e-a874-0edb4351c374 · inbound
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8c7a22b9-3c70-4efa-97fd-886a14846ce8 · inbound
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c7d436cc-01f1-4907-92f6-2883a4edf091 · inbound
Risk Reporting for Developers' Internal AI Model Use RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 43f82950-7e26-4aa2-994c-4d56fa16925b · inbound
Principles and Guidelines for Randomized Controlled Trials in AI Evaluation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b4e7854c-6681-4446-ae70-14c4cba12e9a · inbound
Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 002c6092-2b3e-4304-a6b9-38cc4ca0c284 · inbound
FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f16d3fd5-9469-4b23-8d75-671c0be93582 · inbound
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 98f353b7-e769-4be9-910b-36a33c63fd71 · inbound
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f9630578-942a-41d2-9260-a9449fdd468e · inbound
AI for Auto-Research: Roadmap & User Guide RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 222
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5328db6c-78ad-45aa-86de-e3618d706e3f · inbound
AI for Auto-Research: Roadmap & User Guide RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 221
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 321ddc94-9675-4172-88d5-3f37ac4c7741 · inbound
How Far Are We From True Auto-Research? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 95e6e6dd-7a45-4c86-a98d-45990b1871f9 · inbound
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 23a23acf-0525-40c6-9917-a28501eda992 · inbound
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation abf9f5f6-6c96-483c-8388-2bd6d7e356b7 · inbound
The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 52945ae3-1b6f-4313-b036-adb3f275750d · inbound
AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9668af3d-f1e5-45f1-98eb-3e0245013815 · inbound
InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation efde55ab-60f9-4aae-8e1d-c230c3c7495a · inbound
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4cc5987e-a720-497d-9c48-74c0d0c44ba6 · inbound
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5d39a32f-a09d-4ee2-a9dc-45b587457b01 · inbound
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4806bec1-2f3f-447a-b62b-400a84e06fa3 · inbound
Learning the ARTS of Search for Automated Discovery RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1ac7c5ff-0cf2-4e78-b901-f62eaf2d6854 · inbound
Discovering Crystal Structure Prediction Algorithms with an AI Co-Scientist RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a0aaa883-453e-4868-ac34-d39a298bd0e0 · inbound
MirrorCode: AI can rebuild entire programs from behavior alone RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c974db13-d3e9-4d21-8d11-6794418d9f38 · inbound
MirrorCode: AI can rebuild entire programs from behavior alone RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c361421-912f-4c8e-8c9b-46ec8d5cb72a · inbound
Two AI Metrics Diverged: Will it Make All the Difference? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4031574a-1bc3-4e7d-84a4-046c5a8f602e · inbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 58fcc192-7e31-4da4-944f-af8b18859fb3 · inbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1296820b-a97b-4a1e-ab5e-dfdc4571568b · inbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4e9e2a6-d05e-4c08-b8cf-89f4e4b97842 · inbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4300af11-4056-4bf9-af7a-3b20f2f358d5 · inbound
When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45bc81b3-20ed-4ca4-a3a7-bf39ecc86496 · inbound
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4082e15f-b500-40b7-96f1-5feb288df4fa · inbound
Efficiency Matters in Autonomous Research RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f79981-5a6c-4191-b434-dd3ce80357e7 · inbound
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0bf2f4c-2256-4be0-95f0-6a8f25daeea7 · inbound
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c77e66a2-c16c-433a-b39f-7b6101b3ccfc · inbound
To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0cc8a4e-0e7d-46b8-a1fd-d9b77f400289 · inbound
Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd24008f-7d94-46d1-b45c-723447228ba8 · inbound
Predicting Task Difficulty Without Rollouts RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 1978
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ee986a-611a-463a-941e-2e856cb77cad · inbound
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c4f2d5-3f9b-4e58-a2ea-052cc44da010 · inbound
QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f8b810-3b52-4410-a6bc-6f129843f1ff · inbound
Evo-Bench: Can Language Models Improve Agent Harness? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40910431-7f02-4f95-ab99-34b730e701de · inbound
Evo-Bench: Can Language Models Improve Agent Harness? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b631cdce-8a01-4068-a59f-eb8bc494689e · inbound
VALG: An Agentic System for ML Theory Research RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.