Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models in Scientific Discovery

As of 5 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 19 inbound Pith citation observations for arXiv:2512.15567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.15567 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T21:47:09.588941Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:21:28.776336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T23:49:02.992717Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact53
  • verified fuzzy14
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f0640b1-528c-4d3d-83d2-ace7d6334936 · outbound

This paper cites Attention Is All You Need.

Evaluating Large Language Models in Scientific Discovery Attention Is All You Need

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.419764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:c9a78acd7fdd1c0e4b39b032ff2c30c623274897e9238758dfb67b62d1774d7e

Observation dbc7f627-0a59-479e-9d19-ea2b35641718 · outbound

This paper cites Language Models are Few-Shot Learners.

Evaluating Large Language Models in Scientific Discovery Language Models are Few-Shot Learners

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.423251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:3c90de70b7442d0c3b3e74c9d72aeb89ea5bde91de59c40343b1f3e41b414f30

Observation afff3a85-eff5-47df-80c2-c0a6c721b2b3 · outbound

This paper cites Scaling Laws for Neural Language Models.

Evaluating Large Language Models in Scientific Discovery Scaling Laws for Neural Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.430755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:2be4714cd2d2743b3a404207255139f6e373ea76828b78f2995f0ad8263e590c

Observation 30cff53c-4f69-4892-8b5e-842d005e7555 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Evaluating Large Language Models in Scientific Discovery ReAct: Synergizing Reasoning and Acting in Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.433996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:447a0257e5f808e87ad38c30d56a897caddaf7aea3915c418635d532c8bc63bc

Observation 0a7d8187-c1a2-4a04-81ef-2e55e789b84e · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 5

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.179219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:05d232a993eb963bcaa1330b3498d931f48877bacf89ee5ebf710637cd75c50b

Observation f70da868-bce8-4aea-b8da-340e2689ea60 · outbound

This paper cites Szczypiński, Jean-François Ayme, Ehsan Simaei, Thomas Fellowes, Rob Clowes, Lyubomir Kotopanov, Caitlin E.

Evaluating Large Language Models in Scientific Discovery Szczypiński, Jean-François Ayme, Ehsan Simaei, Thomas Fellowes, Rob Clowes, Lyubomir Kotopanov, Caitlin E

Reference 6

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.176942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:f65223a5fec518786fd2795ad68ce25f89b7560c91978d7aedba16e54d8da4d5

Observation e603c3ba-77f9-4cc4-a0e7-f64ac86c14d9 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.193408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:f8f0f65bca8d2a23dba624c8300a3318387e573306f5db3644d466baa3b13dce

Observation c444f2e5-1d85-4d02-9f60-f6bdf5a7ab5a · outbound

This paper cites M., Schwaller, P., Ortega-Guerrero, A.

Evaluating Large Language Models in Scientific Discovery M., Schwaller, P., Ortega-Guerrero, A

Reference 8

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.181289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:95f5fe1f92c93a8daafa36b31229c18baf7174ac48d882617f536fa66ac5e1ca

Observation ee537345-f752-4b81-b1d6-3cbb2cd771b6 · outbound

This paper cites Large language models for scientific discovery in molecular property prediction.

Evaluating Large Language Models in Scientific Discovery Large language models for scientific discovery in molecular property prediction

Reference 9

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.185238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:ac4f50953a1daf6fc0badaea15e0b45dc5e82e4dc41f4741acfefac2333aae44

Observation 8157938f-d834-4280-8055-3f6d762e87b7 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 10

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.183304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:2d07e74b5240898b498f7a5ac57db54d9dc5083e5f5d8aa51bc71ebdfb856700

Observation 6f83cdde-1ec6-4981-815c-6916632c7119 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Evaluating Large Language Models in Scientific Discovery Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.415900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:9fc639715862509f9c7497b26752481c76561a64aa6148ca6083dd067277c4d0

Observation efcefe1b-1ada-4635-ae5a-6f9bd7fd68b0 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Evaluating Large Language Models in Scientific Discovery Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.427363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:3b22f08c2f05406a4bc34bcc5bd3dae764f7f4495f5ace787a95d11c65e23091

Observation e596813c-3562-4a4c-a084-ee7faff8cbe0 · outbound

This paper cites OpenAI o1 System Card.

Evaluating Large Language Models in Scientific Discovery OpenAI o1 System Card

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.101539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:1cd70c29c7227c98501ab0b6752e04ba489262e27630703cf5eefd45d2edefed

Observation c0243ecb-d07d-4dc2-ba2e-84de42be422e · outbound

This paper cites doi: 10.1038/s41586-025-09422-z.

Evaluating Large Language Models in Scientific Discovery doi: 10.1038/s41586-025-09422-z

Reference 15

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.054580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:5a5b7fb5847909411f87d3b8253b28348648bc4351921f1539804963c8ae0c9a

Observation d7ad4896-b02d-4838-9b44-96f52eee144b · outbound

This paper cites Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q.

Evaluating Large Language Models in Scientific Discovery Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q

Reference 16

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.165587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:c90e6bb3c8ff709479d8467278ed2f6bbb2d836056a38458fb39c7212b6c24a1

Observation 21f1490d-d25f-4580-9681-a2d834e8ef9e · outbound

This paper cites Optimizing generative ai by backpropagating language model feedback.

Evaluating Large Language Models in Scientific Discovery Optimizing generative ai by backpropagating language model feedback

Reference 17

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.157366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:09dce6d211a1c39242da82b8c50476b45b0335be10241bd7a7bec6f1d27fb027

Observation 1aea955e-5fcc-421d-91df-1f631e346c4e · outbound

This paper cites Augmenting large language models with chemistry tools.

Evaluating Large Language Models in Scientific Discovery Augmenting large language models with chemistry tools

Reference 18

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.064570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:a5b561766749c95b45f01f6070e8da130c2c85ccc34a4135c865ead9450d0182

Observation 5755c56e-43b1-4364-a810-bcbbb10d99e9 · outbound

This paper cites Autonomous chemical research with large language models.

Evaluating Large Language Models in Scientific Discovery Autonomous chemical research with large language models

Reference 19

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.087275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:e6fe827a58213f9626981a57a12f98923f58b8a4360c0bb35a7f808f2212424f

Observation b9b766ca-b087-4907-b0eb-21f0987dcd2f · outbound

This paper cites Towards an AI co-scientist.

Evaluating Large Language Models in Scientific Discovery Towards an AI co-scientist

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.393084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:f60f59408256ed31a9a26ed698127ec5a4349090731c8cbd9e14ebdd95c56c40

Observation a6853eaa-e856-4e54-9fae-36d304c724da · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Evaluating Large Language Models in Scientific Discovery The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.389858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:831d4353657e3a0460bc7c689125078db8207324a1d0420652115d2ceba136c4

Observation 1e2fe954-0c13-4eb5-9516-6d7def2bb618 · outbound

This paper cites The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies.Nature.

Evaluating Large Language Models in Scientific Discovery The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies.Nature

Reference 22

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.139295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:eec268f2ba17cb1f6847c2c0dbec54867a01f7a6e3963a0a95355647bddd4047

Observation a5d05575-9a18-4e68-8b60-70b0c1812ca4 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.149082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:fc54b73632fc5e58e1e2d368f4c53fdb8ab4f13f95cf5e4932200f8a37cc785b

Observation 83c69c0a-5b55-4df2-8c2b-9d07af1bab00 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 24

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.077398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:be3cb22727d27269becb06b696f16e97ca03e8affd0e98d2976c29481eaa2fac

Observation d043dccf-8417-4b83-8917-3e4aa7469eaa · outbound

This paper cites Self-Driving Laboratories for Chemistry and Materials Science.Chemical Reviews.

Evaluating Large Language Models in Scientific Discovery Self-Driving Laboratories for Chemistry and Materials Science.Chemical Reviews

Reference 25

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.069283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:6c9ff95cb7715e1a5fa614a028b0a22e30d4912569429031af839346b7cac874

Observation 4dbfdc77-6178-4673-a072-62ec0212d1b2 · outbound

This paper cites author Kitchin, J.

Evaluating Large Language Models in Scientific Discovery author Kitchin, J

Reference 26

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.167862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:8eb8591a160b656b906760665bcd8ea6d0a1271ba73a2d5d73c205f233697d41

Observation 8183efdb-d861-421b-af41-32ab346a75d4 · outbound

This paper cites A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence.

Evaluating Large Language Models in Scientific Discovery A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.094503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:34892bf4998de6be29cf0396cf6525a3b629877da71b12258e5a707359450761

Observation ae2ce515-2ae1-4f8e-bccd-db46876d6b13 · outbound

This paper cites and Johnson, William A.

Evaluating Large Language Models in Scientific Discovery and Johnson, William A

Reference 28

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.081059Z

Source-reported events for the cited work

correction dated 2025-12-08. Source: crossref record 10.1038/s41551-025-01589-0->10.1038/s41551-025-01463-z:correction, observed 2026-07-11T03:07:47.241959+00:00. This notice travels one citation hop only.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:f0caafd421a43c11ce52e4b046bae8540c1a86a531d9009f4854346f390f2075

Observation 4f033503-7f58-4dd1-b20d-b547dfdd5da7 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 29

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.084003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:909f9dec876f0766584ac3f382f7b1dee9677da11be3aae89243fe4aa6a7ad86

Observation 07f03eb2-7fce-4928-8c9f-85b38e985c9c · outbound

This paper cites Democratizing ai scientists using tooluniverse.

Evaluating Large Language Models in Scientific Discovery Democratizing ai scientists using tooluniverse

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.133366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:3b4720bcbd0b0e6f81e754d3ecc6c65edfef601e14762c863dfe8f84ca884ac9

Observation 6622025f-b967-4f82-9024-63fa96084adf · outbound

This paper cites (10) Jiang, G.; Luo, Q.

Evaluating Large Language Models in Scientific Discovery (10) Jiang, G.; Luo, Q

Reference 31

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.136456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:e226badda46c660633cea88bd2fc7800ea850b8c2766e5cfdfcb734e77d3ead4

Observation ac5939bc-ad50-4f65-a51f-963cfec2128f · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.134115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:76deb71a62d4e5eaea1e185201971cdc9c8f56b540254594578da77ab80bd3fe

Observation 71b16236-543e-4da2-ac7e-8952cb2bdbe9 · outbound

This paper cites Kosmos: An AI Scientist for Autonomous Discovery.

Evaluating Large Language Models in Scientific Discovery Kosmos: An AI Scientist for Autonomous Discovery

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:40:55.079629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:e1ffd467117fd6a74899e567ad3474d4fe6b64cef135c60b2c037861afeac016

Observation 2db10bce-473a-4909-bfd8-7e9e01c06712 · outbound

This paper cites Carter, Xin Zhou, Matthew Wheeler, Jonathan A.

Evaluating Large Language Models in Scientific Discovery Carter, Xin Zhou, Matthew Wheeler, Jonathan A

Reference 34

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.160536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:787fd0407b523faf1b4a280c1885c64ef78a997b71b2787ba3143e4d77a9bc26

Observation 4665ccbb-6984-43f9-a616-04b528763551 · outbound

This paper cites Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025.

Evaluating Large Language Models in Scientific Discovery Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.380228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:d86e63bc6482a3f535968ec1b6fbe4859309daf0d29ca893c883ed48165b9bab

Observation 95a7e49d-c68f-48b6-b3ff-b27d0ccb3e79 · outbound

This paper cites Sciarena: An open evaluation platform for foundation models in scientific literature tasks.

Evaluating Large Language Models in Scientific Discovery Sciarena: An open evaluation platform for foundation models in scientific literature tasks

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.374071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:58ce84a46e19160fb43b68f0bdc41a6c0644506ee4d36b720023b0094461689c

Observation 8a894e2a-8a35-4514-b7e3-10727f84c871 · outbound

This paper cites title Swe-bench verified.

Evaluating Large Language Models in Scientific Discovery title Swe-bench verified

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.129593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:c8094b953b324a7090650f4bc6224e596865d61094ddabd2a2ba440390cd9302

Observation 5a14fbfa-cee3-4293-8761-74db70cfd251 · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Evaluating Large Language Models in Scientific Discovery MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:48:34.377051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:06aa902449b180ad19ba320cb3e48eec88be0c3cf7b1a78d9dee0e4208aca902

Observation dfe06954-e8e9-47ef-bd0a-e0287188fce7 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Evaluating Large Language Models in Scientific Discovery From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.399906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:9f6b990d199634f5026340e25796ade9c71b587bad27ee87f8418f4a548349af

Observation de5bfb54-62c9-4c9f-aa94-2c5bb8c53223 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Evaluating Large Language Models in Scientific Discovery $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.128012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:4e1773fb08e7c2bcdb7849997ee103ad34ec3883c94a5126cf5a454ac726ff37

Observation d6f5e6cc-05b8-4fe4-958a-dd25b98dc564 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Evaluating Large Language Models in Scientific Discovery GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.112821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:981fc3dd0f49b79cc4f658cddde5a5b957b69477e115b7fb2f70540f0bd98991

Observation cea9f1d9-0d6e-4e02-a93c-bde834d0e606 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

Evaluating Large Language Models in Scientific Discovery Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.106503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:d73051e9b82664c037370e15756e4912b0011c75916250174316ab28ca2416bf

Observation 57ed24f4-c6c1-4ae4-899f-03564cc70e2f · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Evaluating Large Language Models in Scientific Discovery MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.143394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:04fa7e0bcbea2a1f961f8d337cdae23792fdcd581fec77bdb5d1acf5e5fab134

Observation caa43568-44e2-46c8-a1c4-f85b36185768 · outbound

This paper cites Humanity's Last Exam.

Evaluating Large Language Models in Scientific Discovery Humanity's Last Exam

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:48:34.124182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:e79c96bea2a7037e3adac107a19efa077f042aa8200009eb64aa0a26e4f29132

Observation 6d2225c7-b4fc-4369-a909-f7fe10b3249c · outbound

This paper cites Intell.1, 14, DOI: 10.1038/s44387-025-00019-5 (2025).

Evaluating Large Language Models in Scientific Discovery Intell.1, 14, DOI: 10.1038/s44387-025-00019-5 (2025)

Reference 45

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.146674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:b0f7de4014c8bd2844a9dea7f7831daf70bffbef2ba2d5e1bab14670c830d8be

Observation eb9bae1d-f52c-4578-90fd-7daf419b1f52 · outbound

This paper cites Elbeheiry, María Victoria Gil, Christina Glaubitz, Maximilian Greiner, Caroline T.

Evaluating Large Language Models in Scientific Discovery Elbeheiry, María Victoria Gil, Christina Glaubitz, Maximilian Greiner, Caroline T

Reference 46

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.172602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:5f24f01af20f72e41352d870965b7aa9e719b3054fedda43198c90f600aac23e

Observation 836537d0-3232-4b45-9326-727cb0c321c3 · outbound

This paper cites Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning.

Evaluating Large Language Models in Scientific Discovery Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.403311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:aca9995b113695121c94521c402cb1b76dea9979784f84f4a84614cca378f06e

Observation 1c4242b2-3efd-4076-8f64-254f6546596c · outbound

This paper cites Alampara, et al.

Evaluating Large Language Models in Scientific Discovery Alampara, et al

Reference 48

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.151610Z

Source-reported events for the cited work

correction dated 2025-08-21. Source: crossref record 10.1038/s43588-025-00869-8->10.1038/s43588-025-00836-3:correction, observed 2026-07-11T03:08:42.754763+00:00. This notice travels one citation hop only.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:758f4ba21822f5c3469d94447487b3b18c396efda505d915860443e70da97381

Observation 41a47057-3d46-43e0-b4b9-53bb251861a3 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Evaluating Large Language Models in Scientific Discovery gpt-oss-120b & gpt-oss-20b Model Card

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.409689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:d6ca34e2788ec93d6152109f3dd96df7676785c0b90ef3014ee5dd4a264bc86b

Observation 737a01ef-60d1-4c6a-a5a2-ed90de0e7c8c · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Evaluating Large Language Models in Scientific Discovery Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.406506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:4519eb23f047ed32a7d91e1c7072c1868a36ff02510bdbfe5707c7967441c295

Observation 4be5f1ed-30b2-4ee9-a173-ced6d49e7de1 · outbound

This paper cites Reasoning with Sampling: Your Base Model is Smarter Than You Think.

Evaluating Large Language Models in Scientific Discovery Reasoning with Sampling: Your Base Model is Smarter Than You Think

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:16:05.617665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:e3ee48567933d191eef282a4b36d019c834587f0cb4c53498cd851bb62cc4b77

Observation 2e6f973f-fc99-4d59-a193-93b0630dcc3c · outbound

This paper cites Stress-testing model specs reveals character differences among language models, 2025a.

Evaluating Large Language Models in Scientific Discovery Stress-testing model specs reveals character differences among language models, 2025a

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.383323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:8fb2d9efee44f5cf3cf0687e934a12967a36d2cce1a23cb2b90004131a0d4171

Observation 82f3c756-502a-4990-92cb-1eba1dba1955 · outbound

This paper cites Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl.

Evaluating Large Language Models in Scientific Discovery Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.396240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:4394475d74c032fa3a636749d628dfd653b7f10166296fa6ae7e8c26cdfc0854

Observation 15e21fa9-ff4a-453e-b1e6-12c723136e4d · outbound

This paper cites Evaluating Frontier Models for Dangerous Capabilities,.

Evaluating Large Language Models in Scientific Discovery Evaluating Frontier Models for Dangerous Capabilities,

Reference 55

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.154150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:b5fa1d28ef7e58f2b9ad6faaea528cd23286cb2806e89c5cf489a3b658972c46

Observation 9441a68a-3c93-4aca-9ed3-77866b9dac84 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 56

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.116098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:cdc3448b3f5470b7e21c975764ae6a124b3189bc439ed3142f58138a4d04e413

Observation d518bbec-2b38-4316-b7d8-2d7f483bef08 · outbound

This paper cites author Li, C.

Evaluating Large Language Models in Scientific Discovery author Li, C

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.180952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:c1201c7032fc1812fbc86a31be34059be4b9ca8ce2388d5cb4346b84731f44df

Observation b540fc08-a9e4-4917-93b1-cd02b788d08c · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.189231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:a3e4b9150371301f4f8429ff7a9baa884902acaff1bea18b5e1152c7ac1b311b

Observation 4cdd9e5e-d59a-403d-bff6-a3779c8ff0f1 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.174873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:4c9ffc84b7aca68c218b0ca91334e01aedf9f1d2d0e43d22dbe26a1be30cddc4

Observation 8acfd890-e4d2-43bd-9724-059453008836 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 60

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.174737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:4c49404aee7b59d391f21ea294c42f218bfacbeb539e91255520602001d4291f

Observation aa83e4ec-e2c8-4be1-b880-5111b399c0e2 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.187216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:561b80b95bdbd903cd1c271588b86e7267b51a6f0545269a76d3db021f69035c

Observation 1427d0a9-b75c-4135-a5f7-fc1e77913a93 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.172831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:a0c23b416d754d32634e9c2d333289fbfeb2ada18af017515ee7a3cf5d8fd7c1

Observation 90cf00be-873a-41bf-97a7-dc7b4a812422 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.176980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:a5986630c3608b3b9902d96b57bdb30f7f0187e4b3db8d948d67ac47ddc53f4e

Observation 4e682a79-5aa7-44c9-99a1-b1aa09978c3d · outbound

This paper cites author Meidani, K.

Evaluating Large Language Models in Scientific Discovery author Meidani, K

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.178982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:599abb0a890fdc5ac0dd7c77bb448eba7942992e262c6b35ec67fdbcf4afd23e

Observation 70ed37a8-336f-4ec0-8141-743fb61a8a95 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.182967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:a46e5e2168da19429c6408d6a6bf78c8046fdb93d5d041658a92515030d1bebb

Observation 9d6fb357-95b7-41e6-b473-d86fa7b84956 · outbound

This paper cites doi:10.5281/zenodo.12608602 , url =.

Evaluating Large Language Models in Scientific Discovery doi:10.5281/zenodo.12608602 , url =

Reference 66

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.073825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:40e200c14bfbf1057a3f75ec9fea325ec722be76e2ad1b7ff0681872640a8b41

Observation ebb268b7-7bee-44f5-b9ff-6f6355780c66 · outbound

This paper cites title Sde-harness: Scientific discovery evaluation framework.

Evaluating Large Language Models in Scientific Discovery title Sde-harness: Scientific discovery evaluation framework

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.168351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:633b41c3c9e203c789af2d14c4a1fd3ad741c035e565a119975c757e7ef9a5d7

Observation 09fbd1e7-6c27-4c5c-aabd-4184207b210c · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.170499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:5aacd0f4bb66e6afefc37af33ebba62725098984b8a64fc6642993b8533c3b4a

Observation b62a95cb-d2e1-4003-9f64-437ec2e2c48c · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 69

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.163064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:460d3d272e998f861b17c0ebbcf9bd55a5a7a00de621d38dcba08bfdb569d607

Observation 59bd17c0-1467-4167-9fdb-013cf326f4bf · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.191371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:a3676ccde8cfa35236fd510ed67d7cc9a35a5a008270ea5e7e828c16baf39532

Observation d78db3f2-0ada-43d0-8bb9-2867cc6d431e · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.163570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:625c99199f4aef65c63691bb1ec5b480197f10cb4497558ee41b601e1fe16e19

Observation 517cb734-461b-4064-a829-58fcfc2c2652 · outbound

This paper cites & author Schuffenhauer, A.

Evaluating Large Language Models in Scientific Discovery & author Schuffenhauer, A

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.165952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:a99771bae4d2644d24069a93df7d89b4635715efb1a20eccc59a1a60eae28c33

Observation 12144205-038a-4174-a5c7-5a3d3d3b0598 · outbound

This paper cites title Pistachio (january 2024).

Evaluating Large Language Models in Scientific Discovery title Pistachio (january 2024)

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.185086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:9b1cdca209b24554a74904bd16ce3bf7492c4fb88bb9e3890d68a79050f02add

Observation fa129da7-8894-46e4-81fa-5510291f49ed · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.150585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:bd9bfa0c99bff755abc282fad6519481a0b8fdcf473b7e1166c20cb7e54ccccb

Observation 67e42e0a-bb22-4cc8-9db2-a26712416d61 · outbound

This paper cites author Yang, Z.

Evaluating Large Language Models in Scientific Discovery author Yang, Z

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.161678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:4fce9ef5948c9a711332910176bb4d7b8dcaa3233fc4fad86cd5b53bd662372c

Observation 96d09d07-bd51-4aeb-a748-4e732c520cc6 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.146636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:0e25f70bb29cd93ac541f8423f2ebcf5f3e631f09a5b1da434b66d1765e3275b

Observation ae19da77-8fe3-4826-b1b7-8da3819b1a3b · outbound

This paper cites & author Jung, Y.

Evaluating Large Language Models in Scientific Discovery & author Jung, Y

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.154718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:41a03f100c07cddf15afa4bd131a69ff0f931a595d4cf09853b06c6e8f122709

Observation a4a8c75a-e4ee-4849-822d-e43245495dc9 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 78

Resolution
verified exact
doi, observed 2026-05-16T21:48:34.170169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:ff7b0e5956afa594c961bd97cab8a909f4bd8f2e08f8b2f99f2452a25d297ea8

Observation 779c47a7-af95-4505-a389-534278128f4d · outbound

This paper cites author Wang, Q.

Evaluating Large Language Models in Scientific Discovery author Wang, Q

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.159437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:448a545921c79938284bdbf2c1cc989f87b14ac7160d34cfb32a7c7b10c7f0da

Observation a7a49271-78e7-4b55-9b20-332cdd763497 · outbound

This paper cites author Zhong, P.

Evaluating Large Language Models in Scientific Discovery author Zhong, P

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.136065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:f16b72859c18b64e7e70eff9cfff8c1af99f3cef56a05b985ac6766ce61041e9

Observation 6a46dd7d-eb19-4d8e-9b16-46c0b3af1ddb · outbound

This paper cites author Fu, X.

Evaluating Large Language Models in Scientific Discovery author Fu, X

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.148666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:9d6097432e3b4b1ebe31de580532b8e93e267729def18cd5cffbc66c6b7eaaeb

Observation 2ccc3f68-6995-4f33-85b2-6bf6d9788841 · outbound

This paper cites author Huang, W.

Evaluating Large Language Models in Scientific Discovery author Huang, W

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.140853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:2b818bd904794d1f535bc12cc7c469b4472c25fbdf54f68bf9c29928854a0901

Observation f9bb2b78-9725-483b-a704-adcfcf4934b2 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.142780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:f91a2c67717538a7283721f4dfb904231f60ef2e9b5bb0431b43434fcf8f3f38

Observation 91ae0adf-337c-449f-b83a-489870275ac4 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.138903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:52f1ed1318777d7aa9c36f00af17d7568d92f8594daede6e823a5f971d97b59d

Observation 8cbd6a15-06dc-479b-b0de-cfb0cbf610a4 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.144671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:6bb776c9fd233eb33a3270f70c4be9c0bb53ab7d354f95dafbd7e541a545a59f

Observation 8ce00b35-6e14-4894-9eb6-ff94de626cea · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.152686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:82c0a40d8c0d764e2c42ac1ccc0110024b4c298750f5bb6bdbd68ffd98eafb2a

Observation b74d55ca-6340-45fd-b9f4-725d0816f60a · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models in Scientific Discovery Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:48:35.157325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:2d439d17fd08d9c6f9ede3011a8afbfb7dd243ab50d703aad7f97bf75f6eb2d5

Observation f74df5ac-1fa3-44b4-8a86-c2296121324f · outbound

This paper cites " * write output.state after.block = add.period write newline.

Evaluating Large Language Models in Scientific Discovery " * write output.state after.block = add.period write newline

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.132252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:791dfcc2b153c5d0a08f06d82d2b252162aae49fc9ee5bfc78baf79083dfc2ff

Observation 502fdcde-8078-414f-a302-cbbb1b7595c1 · outbound

This paper cites write newline.

Evaluating Large Language Models in Scientific Discovery write newline

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:48:35.127502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:708afdafe4686b8cc1560faa6ec2f80c88b9f993f6f8a785b556f5d45b54f623

Pith citing papers

Observation f7630111-0b95-430b-b6d9-f8ecef3e64bb · inbound

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs cites this paper.

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs Evaluating Large Language Models in Scientific Discovery

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T14:21:28.776336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:21:28.776336Z digest=sha256:8514314168beedb0447510545edea90546c2de5d6df3a73cd0a1148bfe1037fb

Observation 7605bb7a-ab87-472b-b7ba-26d3659aaff0 · inbound

Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact cites this paper.

Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact Evaluating Large Language Models in Scientific Discovery

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T18:46:29.476304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:44:17.986764Z digest=sha256:b9d317aa88db00f7006f940971f5c87e2ca1a44ffcb83dbacd0f0a8b78fec4a2

Observation 3855307a-017b-4c80-925f-9a1f67bdd747 · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models Evaluating Large Language Models in Scientific Discovery

Reference 161

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:40:43.911729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:76ab7f4b3f5a4a419d7ba276c66226573e9b8702b518c43fa7b25fe5699dee2a

Observation 00abebb5-38ef-4d6c-9288-a2a26771887b · inbound

Heterogeneous Scientific Foundation Model Collaboration cites this paper.

Heterogeneous Scientific Foundation Model Collaboration Evaluating Large Language Models in Scientific Discovery

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:56:27.155571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T08:50:05.980191Z digest=sha256:715ce3bf7b15403bbd80ac0e9375904628258dce13a5373fc245c00461b15458

Observation b8ccea12-af6d-43b7-8907-4fa7533a808a · inbound

Agentic-imodels: Evolving agentic interpretability tools via autoresearch cites this paper.

Agentic-imodels: Evolving agentic interpretability tools via autoresearch Evaluating Large Language Models in Scientific Discovery

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:36:36.206604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:37:43.371592Z digest=sha256:3296fc592314ead82ab86de310d1215d359523bae372d7e565606796180db547

Observation d7efebd3-82f0-4765-80ab-fce504f930d4 · inbound

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning cites this paper.

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning Evaluating Large Language Models in Scientific Discovery

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:45:51.815524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:29:47.384341Z digest=sha256:367fe02fe75ac98d8da82f8ce3fcb9a6de745fdc86fbdef35a714cbb09a7fc15

Observation c0499d08-6750-449b-be24-7541edb5792b · inbound

TO-Agents: A Multi-Agent AI Pipeline for Preference-Guided Topology Optimization cites this paper.

TO-Agents: A Multi-Agent AI Pipeline for Preference-Guided Topology Optimization Evaluating Large Language Models in Scientific Discovery

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:34:46.896481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T09:33:43.733883Z digest=sha256:2a7ecea8d98386f9384449c765b1e7796e223c15c5347058778eae9859b1a7af

Observation 7ee86ef2-5369-4a25-b34a-c23c4a4a04a8 · inbound

Toward General Quantum Control with Physics-Informed Large Language Models cites this paper.

Toward General Quantum Control with Physics-Informed Large Language Models Evaluating Large Language Models in Scientific Discovery

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T22:04:00.500437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T21:55:50.253489Z digest=sha256:cc83d4c9f53d076893bc15f7a1eafa05552f40cf9d26824f12d114418b9e3ded

Observation d8d504e5-b600-4f94-83da-836be0cd4aa3 · inbound

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? cites this paper.

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? Evaluating Large Language Models in Scientific Discovery

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:23:24.413853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T12:18:36.854164Z digest=sha256:c0d68e6e61ba6512c494a31b463b697acf2331fed9c1b151e24f145d5a7d4dbe

Observation 8878780a-7799-4ae2-a725-7217eeedd4f6 · inbound

Quantifying Faithful Confidence Expression in Large Reasoning Models cites this paper.

Quantifying Faithful Confidence Expression in Large Reasoning Models Evaluating Large Language Models in Scientific Discovery

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:06:29.146508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:24:38.335417Z digest=sha256:73a6e4000e3bc405e90b12d7e56deb298b0dfffb942433398c2327180ce35ee4

Observation 08065057-5056-4ed9-9148-32a569d2b1fa · inbound

Ontology-constrained multi-LLM scoring of hypothesis support in the predictive processing literature cites this paper.

Ontology-constrained multi-LLM scoring of hypothesis support in the predictive processing literature Evaluating Large Language Models in Scientific Discovery

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:54:38.385810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T11:49:43.332490Z digest=sha256:15f77374e707db014d72e8d0cd210de9c7a453477c73de805de79f29a6bd27b7

Observation d946edb6-0b61-474c-8c58-9f034d5fa91f · inbound

Towards Diverse Scientific Hypothesis Search with Large Language Models cites this paper.

Towards Diverse Scientific Hypothesis Search with Large Language Models Evaluating Large Language Models in Scientific Discovery

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:07:36.986726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T14:09:31.837231Z digest=sha256:a757355b95ac5392f217161169e249335a940bcfb25b800499bcd037e46eec0a

Observation f04ef476-8bc0-4a0f-a835-959f2171633b · inbound

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales cites this paper.

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales Evaluating Large Language Models in Scientific Discovery

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:28:04.340950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:34:09.347912Z digest=sha256:240bc3509265322e698d28d2c1c8aebd68f222a305ab5c338bddbf01526c9548

Observation 16b80c3a-0704-43ac-ae4c-2d3b27c49580 · inbound

Automated reproducibility assessments in the social and behavioral sciences using large language models cites this paper.

Automated reproducibility assessments in the social and behavioral sciences using large language models Evaluating Large Language Models in Scientific Discovery

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:28:34.064022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:30:08.805659Z digest=sha256:8813c8819e8184d53df1f15c31f8be9ee81db3a76c4ebaefef09ba954079815d

Observation b0e6370f-d74a-4949-9890-b3c19e1a0373 · inbound

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization cites this paper.

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization Evaluating Large Language Models in Scientific Discovery

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:49:02.994186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:40:55.946788Z digest=sha256:802f08e871800fb1efd689d6ec726f8c35fcc80acf4b57c815978b8a6e693064

Observation cfc64460-1382-4d82-85fc-a380c1ef1072 · inbound

SFBench: The SciFy Scientific Feasibility Benchmark cites this paper.

SFBench: The SciFy Scientific Feasibility Benchmark Evaluating Large Language Models in Scientific Discovery

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:04:21.313418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T06:59:19.857066Z digest=sha256:ff64dc015bd8f79ed00ee18859eece6ce076f7ea4e302ae4b653e842f832baeb

Observation 141afa4a-8756-46a1-bb99-7ba67e15d60f · inbound

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs cites this paper.

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs Evaluating Large Language Models in Scientific Discovery

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:35:42.129582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T05:22:38.232552Z digest=sha256:3fe47be810b445abe5889c05323b2ac63e989f326c211c3e795fd49b799a9a3c

Observation 681e9fe6-0b58-4eca-869d-0cdd4b204582 · inbound

NMR Elucidation as an Agentic Search Problem, Not a Modeling Problem cites this paper.

NMR Elucidation as an Agentic Search Problem, Not a Modeling Problem Evaluating Large Language Models in Scientific Discovery

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:14:36.215489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:14:36.215489Z digest=sha256:d47504291e9de9de91427c60510b77387d5dd9ef43399f42fc132a5d580e345e

Observation cd74bc2b-cdb4-4fa4-aea9-1793100af5de · inbound

Generative Artificial Intelligence in Scientific Research: Individual Benefits, Collective Risks, and a Framework for Responsible Research with AI cites this paper.

Generative Artificial Intelligence in Scientific Research: Individual Benefits, Collective Risks, and a Framework for Responsible Research with AI Evaluating Large Language Models in Scientific Discovery

Reference 987

Resolution
unresolved
no resolver link, observed 2026-07-31T23:00:26.004815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:00:26.004815Z digest=sha256:b014966c6e4cd18e6a880308f1d783f5b2ed9bfa83703464e1066a5eb3ca08fb