Pith. sign in

Paper Citation Record · LEDGER

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research

As of 4 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 3 inbound Pith citation observations for arXiv:2604.15411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.15411 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T11:10:21.639856Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:53:09.670362Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T17:51:14.781135Z

Reference resolution

11 of 11 outbound references displayed

  • verified exact10
  • verified fuzzy0
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee5cb94f-01d4-4bca-9bc1-b63cac032252 · outbound

This paper cites an unresolved cited work.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-19T15:13:08.570494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:6eb0748b1ccc85726d54a6ee13be8891478c5a897aa8d2d5f41b20448ee8b7e8

Observation b34068fe-58b4-4353-9f64-a9e7deff6fdf · outbound

This paper cites Robin: A multi-agent system for automating scientific discovery.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research Robin: A multi-agent system for automating scientific discovery

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:25:20.174033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:e700b57bec1342deb6d4706409985331c7d78f1263a9b970bf95361386d3f419

Observation a91e9023-2dc0-4170-abdc-89ab88c46288 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:38:21.105975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:f0f94dee301afedb0f48ab869613c74054ae7101340bdec5c34304263862dcaf

Observation 58259e96-f18d-4034-9f41-f31272cb8892 · outbound

This paper cites OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:11.398016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:0ee04ae3c40c7d557d450a49f53c91e593ef00d645c60b945d643d9bcccf2757

Observation 94790942-d746-47f7-945d-ea73bad8414b · outbound

This paper cites Kosmos: An AI Scientist for Autonomous Discovery.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research Kosmos: An AI Scientist for Autonomous Discovery

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:40:55.079629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:cfc7a1b28569107b1438a4845f1fb64031f5b4aa06928dfb847e856751ee1679

Observation 098caa3a-c53a-49b6-b98a-8ef6a88c2135 · outbound

This paper cites Towards an AI co-scientist.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research Towards an AI co-scientist

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:02:46.484476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:633935e6b2c7c4c3c6bee06810d1838a05275d5f04b8e68fe48e1cf4a8c65ce2

Observation 54cd7de7-1aa8-425c-b352-d2576954b524 · outbound

This paper cites Humanity's Last Exam.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research Humanity's Last Exam

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.715344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:e700e29f13315b805d81732265fd04ac34172708aed7cc004730e63b9fd8dbce

Observation aaaff15b-3695-4c8d-a275-dc55b59c8288 · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:11.381486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:6a4989c21a0e08bd741c27ff029ee2bf03c1385c7e6f0b118a446464c8c6bac2

Observation e0919d7e-2227-4579-9c43-2a422ebba608 · outbound

This paper cites An End-to-end Architecture for Collider Physics and Beyond.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research An End-to-end Architecture for Collider Physics and Beyond

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:11.387179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:d453a940778364e210ef9407038e14e38b4d5a9e04f6679bb37eac925e7bb508

Observation 36d311b2-cc18-4336-a36b-13591c8fe84f · outbound

This paper cites Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-10T11:20:11.407892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:264505aed18bd57729e90e9bf5bb739bda80c5c3e8a14dfe47d41c266326348e

Observation 9ddd4d9a-d334-4d14-a3bc-db2d4b627f4b · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:59:45.390593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:10:21.639856Z digest=sha256:58596870e70b43cce1cc4e77233831c05454be70e8b3552a2511fd180a52782b

Pith citing papers

Observation 07c2da68-61b4-4b50-8a06-b29b79dae545 · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-05T17:51:14.782728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-05T17:45:55.631459Z digest=sha256:bacb2554f142ff927b5eab56932ec2dda4786d9f406a306637c44ec0738e09dc

Observation 10f95797-1de5-40d8-bdfb-e3e95e05d4a3 · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:26:06.895857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:27a288c2ee01d038aa273d55533bbcb5ffc4c8fa821a4ca636fd4c201684c164

Observation 836dfa31-5bc2-417f-a0b0-3e2545b2ea7e · inbound

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review cites this paper.

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T01:53:09.670362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:53:09.670362Z digest=sha256:2f7b7e473d497daae95568874877925668f868c18ba26403515addd9239abb5a