Pith. sign in

Paper Citation Record · LEDGER

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

As of 18 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 8 inbound Pith citation observations for arXiv:2505.11855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11855 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:51:34.451656Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:44:25.941162Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact5
  • verified fuzzy21
  • unresolved63
  • parse uncertain1
  • malformed identifier2
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 2d0c873f-7ae6-4190-a326-0335623f73a2 · outbound

This paper cites Improving language understanding by generative pre-training.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Improving language understanding by generative pre-training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:33.971756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:33.971756Z digest=sha256:20d920fba9f4fa75f195861618992b40734c23e805e988ee4fd6bceaf81af60a

Observation f415ccc6-1207-45a5-904d-3aeb01b29bfc · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:33.977488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:33.977488Z digest=sha256:c634ea489decceb44bff25e5598ab82727cdbec08088fd97462dc2ceb47a84a6

Observation 7bb2e622-0bb7-408f-a4e1-f37537f1d806 · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:33.982308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:33.982308Z digest=sha256:2d5fd952dd1c019a6c12a2a1b3aae43ff5daedb2356cfe6814d44402d8b2e9d3

Observation ac595b49-9203-461e-833e-4da9b9cbcc4a · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Gpqa: A graduate-level google-proof q&a benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:33.987609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:33.987609Z digest=sha256:0f34da47afdb89ec4e595872a2860aa2abb696103dd423e8777543995f7de088

Observation 2d27e811-b29f-49ed-90d8-a9872ccb6b1f · outbound

This paper cites PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:33.992511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:33.992511Z digest=sha256:471ed83211bf698d2990c086e41fe310aee34700c68a7a49a607859992f6a910

Observation da9cf596-f628-4906-a4f4-a6e4f9b50e59 · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:33.997813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:33.997813Z digest=sha256:b769e76ed3eb561fa79e0ab690a75deef95428f2d91d48c7a9906d548d78acb0

Observation ceda3ba0-d88a-4be6-b092-16dbaea4c496 · outbound

This paper cites Can chatgpt be used to generate scientific hypotheses?Journal of Materiomics, 10(3):578–584, 2024.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Can chatgpt be used to generate scientific hypotheses?Journal of Materiomics, 10(3):578–584, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.003461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.003461Z digest=sha256:2319a7be3129b69f508cb5502d3ce38aa6829e7161fabb84e6e407434aa71c91

Observation 99711ad0-3f2b-43a4-be06-0088f70f0552 · outbound

This paper cites PaSa: An LLM Agent for Comprehensive Academic Paper Search.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research PaSa: An LLM Agent for Comprehensive Academic Paper Search

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.008442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.008442Z digest=sha256:c896ee936ff2091405ed427df2d0896a4671dd0edc11a82e2c62593caa17a6af

Observation 541aea6a-49a3-4b13-8ae6-08879e11df6c · outbound

This paper cites Generative ai in writing research papers: a new type of algorithmic bias and uncertainty in scholarly work.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Generative ai in writing research papers: a new type of algorithmic bias and uncertainty in scholarly work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.014130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.014130Z digest=sha256:6a77305d997f503efaf5230bc7fe79eb1e0ec5773a38105e9f6fc9196849af70

Observation a0197d09-e72a-4d56-8331-23c566057272 · outbound

This paper cites Towards an AI co-scientist.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Towards an AI co-scientist

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.019504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.019504Z digest=sha256:7a742b4ffbc2e7abb1eea096ae243da6a7a77112f87b08a67d97af6808948038

Observation a273955e-0333-4ad0-93ea-c6a6d6ae9f9b · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.025322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.025322Z digest=sha256:85fc9839e260e62325f83cf35b53b4677818cb1f2ab93cf7eaa113bfbdbc98f0

Observation dc2175da-2857-4477-a8b0-86e5b5516e6b · outbound

This paper cites Ai mirrors experimental science to uncover a novel mechanism of gene transfer crucial to bacterial evolu- tion.bioRxiv, pages 2025–02, 2025.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Ai mirrors experimental science to uncover a novel mechanism of gene transfer crucial to bacterial evolu- tion.bioRxiv, pages 2025–02, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.030797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.030797Z digest=sha256:8f508a28dadc493dfd8e65e2c0194caa1385f9b653800a57bdf2bfe5a91d82fa

Observation e86aaef8-191e-4d85-82f3-6ae296b78e44 · outbound

This paper cites Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.035772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.035772Z digest=sha256:2482f2998d4cf50e9e53235753a70027f67f220af989b498bd6aae80844d929e

Observation ade9501b-4d23-4c27-b912-466c7bcc6197 · outbound

This paper cites Quantum many-body physics calculations with large language models.Communications Physics, 8(1):49, 2025.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Quantum many-body physics calculations with large language models.Communications Physics, 8(1):49, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.041049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.041049Z digest=sha256:d09a2ef14b1f9888a1792580a72aac0258ff728ddb61467ba42d368e2f553f3d

Observation 1774546d-4a06-42b0-9f91-72d021a102e8 · outbound

This paper cites Alphaevolve: a gemini-powered coding agent for design- ing advanced algorithms.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Alphaevolve: a gemini-powered coding agent for design- ing advanced algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.047576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.047576Z digest=sha256:4e741f7878d6b26628b887bb5fd87fc762282d5f480b488b8cf8456931e09478

Observation 133dfc56-28c4-423d-9371-dfd4d3188fd4 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.056955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.056955Z digest=sha256:72228e82a8164c364a20ec4e3cf5932f076e0d6791ac1deb4de70db6b248f1d3

Observation 109c1cbc-3da2-454a-ba2b-ab0daef44d97 · outbound

This paper cites TabFact: A Large-scale Dataset for Table-based Fact Verification.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research TabFact: A Large-scale Dataset for Table-based Fact Verification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.061714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.061714Z digest=sha256:a83c03bea5d85d14e7553cc45fbeee3c2c2611b6ef34a66c9dfa5c2e2487ae94

Observation 8c58a3b7-a83f-492c-a3b7-642111f06c21 · outbound

This paper cites A review on fact extraction and verification.ACM Computing Surveys (CSUR), 55(1):1–35, 2021.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research A review on fact extraction and verification.ACM Computing Surveys (CSUR), 55(1):1–35, 2021

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.066723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.066723Z digest=sha256:bf428d3c747039b69653336cd52ccba2f62bd56d0f6ad375bb851aa98513e299

Observation 712017a8-56f3-49f3-8dd0-237e434a9bc4 · outbound

This paper cites Poly-FEVER: A Multilingual Fact Verification Benchmark for Hallucination Detection in Large Language Models.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Poly-FEVER: A Multilingual Fact Verification Benchmark for Hallucination Detection in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.072118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.072118Z digest=sha256:1c90914515d66b0a02b9577e349b39ba373fa48aca85f4ad883bbd211149a8eb

Observation 723fec5a-acf5-48d5-919e-c29c0c49e44b · outbound

This paper cites Sciclaims: An end-to-end generative system for biomedical claim analysis.arXiv preprint arXiv:2503.18526, 2025.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Sciclaims: An end-to-end generative system for biomedical claim analysis.arXiv preprint arXiv:2503.18526, 2025

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:51:35.374442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.077195Z digest=sha256:5bb8c7351e6a588bee0998d0d6403bb0834d8a37fa7c4974b63b86de5449873e

Observation 7388828d-055b-4328-8fa8-cd8913315ca5 · outbound

This paper cites SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:51:35.227421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.081901Z digest=sha256:b0143d7e191185985223e000f4ca141592be3f9b4ea70bdbbf6e5de6a2bf7258

Observation 242412c1-3f39-4abc-b6a7-120967c4bc60 · outbound

This paper cites CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.086961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.086961Z digest=sha256:57625b75c644aaa066a3c4c388b314c2840b96aade3c28f2fc275db74ee85083

Observation 2abce3ba-e442-40e7-9147-7a946b141602 · outbound

This paper cites NLPeer: A Unified Resource for the Computational Study of Peer Review.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research NLPeer: A Unified Resource for the Computational Study of Peer Review

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.092016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.092016Z digest=sha256:118fc828723876ddb45a5b090efe9315e22ea3fcef9d7ae1ae1ecce39d33a01b

Observation 89365f2c-70c9-488d-b475-3071e6e4861d · outbound

This paper cites PeerQA: A Scientific Question Answering Dataset from Peer Reviews.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research PeerQA: A Scientific Question Answering Dataset from Peer Reviews

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.097216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.097216Z digest=sha256:f02d829951dbec4aff6a8caca2db79d5f47800bfb901ff3a6cab4d74efde8982

Observation 995e37ec-49b5-4513-9f45-8e4c9112404b · outbound

This paper cites Make every example count: On the stability and utility of self-influence for learning from noisy NLP datasets.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Make every example count: On the stability and utility of self-influence for learning from noisy NLP datasets

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.103711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.103711Z digest=sha256:f78902de65c14c3da44356a36b0c16fe291a59d3bad9ceaf76e5c7201a1a0a79

Observation bcd1c5db-eafa-4072-8e72-fb88700f74ea · outbound

This paper cites FEVER: a large-scale dataset for fact extraction and VERification.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research FEVER: a large-scale dataset for fact extraction and VERification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.109973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.109973Z digest=sha256:b051671f3b96b5d2cd473881fc9e46f1140ae554be2caa2e54d1f25df79381ea

Observation 8bf741d7-66c3-47cd-8cf1-e978ac8fddc1 · outbound

This paper cites Fact or fiction: Verifying scientific claims.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Fact or fiction: Verifying scientific claims

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:36.070073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.115683Z digest=sha256:1b9b6a5ad0eaffa94b92def7603a97cdbed303f0ac8466b8baaf8031d42085ee

Observation 366cf880-d97a-4b64-bcc1-3e06231c9c8c · outbound

This paper cites Moprd: A multidisciplinary open peer review dataset.Neural Computing and Applications, 35(34): 24191–24206, 2023.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Moprd: A multidisciplinary open peer review dataset.Neural Computing and Applications, 35(34): 24191–24206, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:36.053820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.126149Z digest=sha256:831e367523f58e6d33b3169dae1adc40fdf8b5963f101dcae97df311dacd0c4a

Observation f32e4cb3-4ecd-4559-b0a5-c61869155b7c · outbound

This paper cites Automatically evaluating the paper reviewing capability of large language models.arXiv preprint arXiv:2502.17086, 2025.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Automatically evaluating the paper reviewing capability of large language models.arXiv preprint arXiv:2502.17086, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.131482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.131482Z digest=sha256:d2b6886df2c05b65113447bf311b71b1c2bb27d6c88db5b61aef22f4cfa63d26

Observation 8ddae998-2427-48ce-a8cb-3b0b85e0a906 · outbound

This paper cites Openai o3 and o4-mini system card.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Openai o3 and o4-mini system card

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:36.035418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.136486Z digest=sha256:1b97d6f4f24297326a0abf80a1ec1da184e873b76e753c8a340e7bf05f5ed52f

Observation 506c3022-380a-427c-adaa-db0cbd874010 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innova- tion.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research The llama 4 herd: The beginning of a new era of natively multimodal ai innova- tion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:36.018819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.141479Z digest=sha256:c190602fcdde82cd45c9332beb339873fcf57efa01a1aa97c1d3896473439e58

Observation 9b5f2646-e019-4f8d-bcdd-87cb61850f39 · outbound

This paper cites WithdrarXiv: A Large-Scale Dataset for Retraction Study.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research WithdrarXiv: A Large-Scale Dataset for Retraction Study

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.148285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.148285Z digest=sha256:8eef9702dd0f8a561244a6b9d1c9537ba1e590f0889602621ef7e3c888e7b632

Observation 47bf3033-2a12-4303-a279-2859e435febc · outbound

This paper cites Classification and analysis of pubpeer comments: How a web journal club is used.Journal of the Association for Information Science and Technology, 73(5):655–670, 2022.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Classification and analysis of pubpeer comments: How a web journal club is used.Journal of the Association for Information Science and Technology, 73(5):655–670, 2022

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:36.002842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.153530Z digest=sha256:85baff0fb54010b60b67469ff9599d0a81305cf399e224c43c19aa1227a3bab0

Observation 8861ffd9-a113-45a4-9202-8a98cf225664 · outbound

This paper cites American Invitational Mathematics Examination – AIME.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research American Invitational Mathematics Examination – AIME

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.986386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.163373Z digest=sha256:539f728f3dc606a23c37b9a89cfdafc34c3c492e1ac0963299cb97f87434556d

Observation 4fec5ec3-d608-4198-ba5c-cb03ba127e45 · outbound

This paper cites Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.167906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.167906Z digest=sha256:4e233e2279f5a87a4dd9569c8d7b9f5d55e5091fc46365de4f7b7872d2b408bb

Observation 360c5245-dd0b-46cf-acee-a4cb9596f492 · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.172803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.172803Z digest=sha256:da9a6b0bac3d2c190385566cfb7f3d2b1eaee96a8d0f17b25a5777b08fbe3868

Observation 20dcb9f9-441d-4f74-9372-316a4b4bfc04 · outbound

This paper cites Paper2code: Automating code generation from scientific papers in machine learning.arXiv preprint arXiv:2504.17192, 2025.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Paper2code: Automating code generation from scientific papers in machine learning.arXiv preprint arXiv:2504.17192, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.178088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.178088Z digest=sha256:0a7ddd6c2eab4007b17fbdeb363275bbda9164945215880be36fe13bc121e4a0

Observation b09c2bdc-cc80-4684-a8d9-a3148af549c7 · outbound

This paper cites GPT-4o System Card.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research GPT-4o System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.182811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.182811Z digest=sha256:04550f0eb6c36678155c77783eab335ba31e943167c480f801f610b45bd62175

Observation 030f2c9e-7b5b-4f77-98e1-baea604c75e7 · outbound

This paper cites tiktoken: A fast bpe tokeniser for use with openai’s models.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research tiktoken: A fast bpe tokeniser for use with openai’s models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.970185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.187951Z digest=sha256:94ec91bf66735416c979ece63e5a25ffcc173b50381321054eaa78c8b5b24712

Observation 31acb379-afae-4f1f-953f-aa1742758e0c · outbound

This paper cites MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.192526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.192526Z digest=sha256:13e936aa37d2c06726b235f90d670f9608ef656d3d9159a025034962f8d81f0a

Observation 972e8633-bd04-4fa4-a737-f3968ba35c81 · outbound

This paper cites Spoc: Search-based pseudocode to code.Advances in Neural Information Processing Systems, 32, 2019.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Spoc: Search-based pseudocode to code.Advances in Neural Information Processing Systems, 32, 2019

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.198426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.198426Z digest=sha256:12a7c24175d3065649099b51614dc7499cfb465dbb0e28de3fc3145e7872948f

Observation 8ca194c5-6b59-45a3-b1bd-b1d134dedf54 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Evaluating Large Language Models Trained on Code

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.203367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.203367Z digest=sha256:0158fe55cfcd775630e22d000dcffdd2251b0bbb25af0655035c7a7b4d0d0ac9

Observation 103b4b0a-850c-46c2-981d-ad473026c98b · outbound

This paper cites an unresolved cited work.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:35.943261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.209225Z digest=sha256:fa08af317cd9f93f437f23f2044a0409593bff45f3c1601f7791693df3777035

Observation c574a8dc-7f3b-4197-9b32-19d7efdcc432 · outbound

This paper cites Gemini 2.5 pro.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Gemini 2.5 pro

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.925752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.214041Z digest=sha256:7d32b6e5d67130569efcf7bfb4b301eac7edd90ddc288eec1d68542910cf2c36

Observation ce12ad1f-b99c-4e61-b13b-5634fd66017f · outbound

This paper cites Gemini 2.0 Flash Lite.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Gemini 2.0 Flash Lite

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.909191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.218794Z digest=sha256:6bb018e8d9f80760fbe49a3a907b191b7ca6364d5883e32cb21938652eda152a

Observation c7ff58c0-381e-407a-a38c-75c94493c9a1 · outbound

This paper cites Claude 3.7 Sonnet System Card.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Claude 3.7 Sonnet System Card

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.893522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.223686Z digest=sha256:814d676ced5bbf874eb37a76af4974af10c1260e4e96edbb2c7c72fd13745e1c

Observation 73454a53-a82b-4e39-ae00-4412a4cda119 · outbound

This paper cites Qwen2.5-VL Technical Report.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Qwen2.5-VL Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.228794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.228794Z digest=sha256:16315f3e52db60efca2d6353141cce60b0e8cfea4a36662b5a192ff90cfdbcd1

Observation 49f07a06-c9e9-45c9-8a18-05b0538ed520 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.233596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.233596Z digest=sha256:9d625642844b732bfda2a2c0efc834fdd00d83705694e182f27876b30669ce05

Observation 24328483-875b-4710-b6b5-f970455c3e0d · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.238568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.238568Z digest=sha256:14d545b1789be002021d5c9a5aeea959be19aa28e9e9dd47b40cf6d1f08c44eb

Observation 4257c18e-5322-4c01-b4d1-aa2024677069 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.243677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.243677Z digest=sha256:deb5414dddaa6dd25ef5ee5e60ea141e3f3d46d6ff4381625930ba73c68e31ac

Observation 4b3432f5-3308-43d9-b4bf-6bc420ca87d2 · outbound

This paper cites Humanity's Last Exam.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Humanity's Last Exam

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.248451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.248451Z digest=sha256:bc3e50917ad2647bf7e52a88a552243a86e03d6a65650616672b03c15b0769a9

Observation 8903a918-5c24-4ed9-a450-2ed0c00f237c · outbound

This paper cites On calibration of modern neural networks.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research On calibration of modern neural networks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.253284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.253284Z digest=sha256:e14f6c9509b3609fbb6885f645f3e0107d5afb16a9b26a688576bb4489ae92ef

Observation ac5c5aab-fb44-4f46-9f9a-5cb0aabc752b · outbound

This paper cites Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift.Advances in neural information processing systems, 32, 2019.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift.Advances in neural information processing systems, 32, 2019

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.257924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.257924Z digest=sha256:b556fd7765832b7ff694198c680d5dd6745b3cfad0c73fd5e1a555a939bf30c2

Observation 5ca77c99-a692-42ea-b748-f8df352518b5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.263198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.263198Z digest=sha256:4abf829239642467bd3831e8bcac7f6eba83b18ade84b515789f4778ca310aae

Observation ff70fb3a-935a-482e-bcb5-0503799b4a31 · outbound

This paper cites DeepSeek-V3 Technical Report.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research DeepSeek-V3 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.268390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.268390Z digest=sha256:48528a8bf5bc67dae664937d3faaee3fee8ddcf805164a77968069f1e14ad036

Observation 3adc1fc3-b475-46a6-af57-d820e9de4aa3 · outbound

This paper cites Qwen3 Technical Report.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Qwen3 Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.273333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.273333Z digest=sha256:573e93b6de83cb971ed1de306cb18f6e14af895ed577c324501741adf9cfbd8d

Observation 2ea95058-77e2-4f09-8066-12f48bc7a8c8 · outbound

This paper cites Multiplicative Chow-K\"unneth decomposition and homology splitting of configuration spaces.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Multiplicative Chow-K\"unneth decomposition and homology splitting of configuration spaces

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:51:34.816915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.278297Z digest=sha256:80a90bdac246b729e2d93f32e6e76fe5c422618be8be0dc50c9facd3c69e1373

Observation 90e9562a-eddf-451a-bc50-49e4330d0e09 · outbound

This paper cites Superacid in situ protected synthesis of covalent organic frameworks.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Superacid in situ protected synthesis of covalent organic frameworks

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.834835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.284626Z digest=sha256:7ffdbc0b2c9b14e12c955fa3c801fe2a6833730672a61a087f27dcc7d06ce9f0

Observation e891a275-7c4b-4f7a-83d0-57def7bc15a2 · outbound

This paper cites LLM-as-a-Judge & Reward Model: What They Can and Cannot Do.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.289576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.289576Z digest=sha256:c683cd6381c8f461141757029c21fe8240460bb713973a9061aeba9a11f00d7f

Observation 02dd2aac-e86c-4eb7-ab5b-9ef2fcce6734 · outbound

This paper cites Discourse-Based Objectives for Fast Unsupervised Sentence Representation Learning.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Discourse-Based Objectives for Fast Unsupervised Sentence Representation Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.294720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.294720Z digest=sha256:3ff3c760fc173d189c6b9991fd846249973f8000006e7e9ccb3350b6402efa5d

Observation dd724038-6b4c-47e3-8256-1cfea1e9f252 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.300238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.300238Z digest=sha256:498de269c899a7c370b552275d9af11f258fd8258e89cc1a7dadc9b0e6fec699

Observation 2429bfb3-af12-4b89-95a9-5758101ce73e · outbound

This paper cites Reviewer2: Optimizing Review Generation Through Prompt Generation.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Reviewer2: Optimizing Review Generation Through Prompt Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.304874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.304874Z digest=sha256:ef6a2135777d74e772807afcfdbe6b7982203d568fd0fed3e55fc0391081dcdc

Observation 6cd76ba9-7ae9-43ed-bd6b-2042e5c2b938 · outbound

This paper cites Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.818381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.309751Z digest=sha256:b4fec15617adb1f917c26831ddd7a2b58a01d69e22bdf111241d8c1ff214bd5b

Observation c4569237-b995-435a-98b8-c9a1d8fa30fa · outbound

This paper cites Inconsistency in Conference Peer Review: Revisiting the 2014 NeurIPS Experiment.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Inconsistency in Conference Peer Review: Revisiting the 2014 NeurIPS Experiment

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.314111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.314111Z digest=sha256:5a94e04635b3c65b7569f7e4adfd62b140f3646076749645369872e02f12e774

Observation 7f427bc4-e695-4e1a-ae83-40ad7caf3538 · outbound

This paper cites A noise audit of the peer review of a scientific article: a wpom journal case study.WPOM-Working Papers on Operations Management, 14(2): 137–166, 2023.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research A noise audit of the peer review of a scientific article: a wpom journal case study.WPOM-Working Papers on Operations Management, 14(2): 137–166, 2023

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.801028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.320209Z digest=sha256:8a0f7b5d1b84786aad1bcdced4073931a9c40b9fc05a86273f04afb65939f494

Observation c255511a-bdf6-4314-abc3-6df78478d8b4 · outbound

This paper cites Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.324998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.324998Z digest=sha256:b09dd165a68ee47c6f3836e72627b43402b74c24c385028c9bf6904cf81ef37b

Observation f87f38e3-97d1-4a75-9eb4-b6ee4a27b08c · outbound

This paper cites Scaling Scaling Laws with Board Games.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Scaling Scaling Laws with Board Games

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.330216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.330216Z digest=sha256:ffe3c6e2e6e5bb7335e1c3c4da88b818757c5ba6b1315153673a7541805cd0a0

Observation a0e4789b-faab-4219-b6d3-3ce4d310aef1 · outbound

This paper cites Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.335360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.335360Z digest=sha256:d98b45ddb643eb52e1ce4dea73e0343176e5ed0ab8af8d4cf188a57b54af660d

Observation f8ddd532-0db0-4ebf-a81e-546525a6a653 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.340516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.340516Z digest=sha256:698b312b557aeab9035edd45d1f1c64e2c5a2fea197d7eb7311dd25f3c9f0b34

Observation d1425ae0-1b37-46b6-8cdd-f23d5e7119fb · outbound

This paper cites s1: Simple test-time scaling.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research s1: Simple test-time scaling

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.346039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.346039Z digest=sha256:20ab6aa21c09665b6243fabdddf2004536c39921e3dea142fd990723cb6aaf56

Observation 3d261402-f240-46ea-bbf0-b62b886a6372 · outbound

This paper cites Finding flawed fictions: Evaluating complex reasoning in language models via plot hole detection.arXiv preprint arXiv:2504.11900, 2025.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Finding flawed fictions: Evaluating complex reasoning in language models via plot hole detection.arXiv preprint arXiv:2504.11900, 2025

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.351393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.351393Z digest=sha256:b27fe0a5fdd280eb85cf7a5b3a0cc08feec1dc12801bd94638460e4d46da6c6f

Observation 869e8f96-440a-4dcf-99f0-a66852e2da12 · outbound

This paper cites Algebraic description of complex conjugation on cohomology of a smooth projective hypersurface.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Algebraic description of complex conjugation on cohomology of a smooth projective hypersurface

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:51:34.553041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.356275Z digest=sha256:ba982a08c9a0b14f07fee17fc9fae8cfae00874b7e7c2e5193f6bfb0a424bc64

Observation 80dbaa64-993c-4bb8-a430-54c807394d48 · outbound

This paper cites an unresolved cited work.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:35.784986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.361179Z digest=sha256:badd67187821188fc10f6d3abccc1ea1ef20999d27fe635c07ee253b07869a8a

Observation b3ef3239-d156-4650-972f-a5b1c752b85c · outbound

This paper cites Limitations.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Limitations

Reference 75

Resolution
verified exact
doi, observed 2026-08-15T20:51:34.494849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.366339Z digest=sha256:98003ce5293c000c2bbc55df8497918033d2ac0693e252c7e51b50c58982bd88

Observation 49dfd2de-6d24-4031-a9f3-2a5335bd87ce · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.768343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.372242Z digest=sha256:7289fc233dbcaea9426dad5c9d9e79c2d9fbfeee5530b4636aeef9a3afa97deb

Observation d86d75e8-32a0-4d47-acde-21aa4cefa30b · outbound

This paper cites an unresolved cited work.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:35.753205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.378243Z digest=sha256:7c7e245972d8aa8c73b796593cb7cb0d78625cca87e1cfc8b584c08c1ff3ec9a

Observation cf924b38-2096-4192-8dcb-fdd0ce54593c · outbound

This paper cites Conversely, false positives may occur when:.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Conversely, false positives may occur when:

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.737536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.383718Z digest=sha256:5a57497e8028b4fe558c3376abfa36216aebe1b48b7150814ae6277581a4b658

Observation 853f51a8-fa0b-4256-b657-a14c7b494968 · outbound

This paper cites reasoning effort.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research reasoning effort

Reference 81

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:51:35.721107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.388879Z digest=sha256:116c64c180bbe0f7c19e464555dc73a0597177188bae29c9026d875e42009e58

Observation a727aaa3-0704-4346-b6d8-c7ac26147cf3 · outbound

This paper cites decreased from 0.116 eV. . . to 1.03 eV.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research decreased from 0.116 eV. . . to 1.03 eV

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.705316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.394801Z digest=sha256:560680f487c6a05fdad1927bc34844816088d4fa299902d5adad2c1132526f8f

Observation d7fe2f89-fdda-4110-8743-99f9400eafef · outbound

This paper cites an unresolved cited work.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:35.689383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.400945Z digest=sha256:4c12d1e6e2372967ab54c19057ee2d8ea6a11c471dac67f8cc0195cd21937bf3

Observation 26c88657-f69c-44f2-b740-e6053cc4655b · outbound

This paper cites an unresolved cited work.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:35.672481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.406201Z digest=sha256:da5421e264e2e2934dc1af839c0f289d9c872b5fa7e24dc5c4deeb93c9a80c81

Observation d57fc1df-fb11-4c79-8579-a1ad812c9dda · outbound

This paper cites an unresolved cited work.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:35.656237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.410893Z digest=sha256:8a916d4ea0cd56fceca64a82a5d5302370b46b23fe730e59b04ba4b3b503e045

Observation 5b4e1349-aff1-4836-aa97-054c13dee291 · outbound

This paper cites Generation Prompt.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Generation Prompt

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.639532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.415574Z digest=sha256:b79e07c91ada79f4e09241cbec83f64c933e71b75025804e626bf4d5f82cdea4

Observation a2214f1e-fd27-435e-be96-2c31d928e4fc · outbound

This paper cites annotations.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research annotations

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.620338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.420652Z digest=sha256:e86b11c3e781a5930bcc9a45f5de1d1f7d9ea10b460740a5382ddbaf6bd8d6d1

Observation 7c8e0159-fcb3-4911-8710-07d2dbf85ecc · outbound

This paper cites predictions.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research predictions

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.603384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.425891Z digest=sha256:3df87d7c29e11a8c762d00a391b67409b1b3a68410ca82815ef4f6b05c6ebe9e

Observation 4f5c2bb1-0003-4821-9c10-0f1600bac133 · outbound

This paper cites an unresolved cited work.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:35.585470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.431578Z digest=sha256:f724805915a6578f6b1b6c613169abe336f4dc3e7423da7f08eb9bb0a8a3de58

Observation e1db4614-0259-4a17-8ee2-b8433aed7f98 · outbound

This paper cites location.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research location

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.568770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.436601Z digest=sha256:d42d485b41f71475071a6239fdcbe55264293df77b67a5437d27e9bf435e3350

Observation d92ebf6b-0b6b-46cb-8fcd-3df1b91b7f45 · outbound

This paper cites matches": [ {.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research matches": [ {

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:35.553135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.441326Z digest=sha256:fcdf723583ef60cc26e3e178261822ce6dde7002b92a760ab638739d82d7ef9b

Observation 8ebfa0d1-ebaa-4ddd-8f6f-de226baa6c33 · outbound

This paper cites an unresolved cited work.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:35.536610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.446492Z digest=sha256:a99be1066fed483c28c7ad199e0f02dcf99f3545405c88e8df531da72746e618

Observation 3eb44cf9-32c6-442e-92f4-9daa85ffeafe · outbound

This paper cites Table 4:Mean and standard deviation of pass@K for o3 (K∈{1,2,4} ) by error category (left) and paper category (right).

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Table 4:Mean and standard deviation of pass@K for o3 (K∈{1,2,4} ) by error category (left) and paper category (right)

Reference 93

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:51:35.520694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:51:34.451656Z digest=sha256:9565ce9a45f141c6bdbcbb56d30e58257207b53647c855a544c9fa88d437269c

Observation df9bbf82-cc38-495a-8ad6-683aa4e36791 · outbound

This paper cites doi: 10.18653/v1/2020.emnlp-main.609.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research doi: 10.18653/v1/2020.emnlp-main.609

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:34.121025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.121025Z digest=sha256:23c05b6b801fbcfa8b6cba67dc9b530121509e70761eff131bfab21ba64c238f

Observation 8ffb36bc-e6e6-42e7-b077-2450dceb4fc6 · outbound

This paper cites an unresolved cited work.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work

Reference 2025

Resolution
parse uncertain
no resolver link, observed 2026-08-15T20:51:34.052419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:34.052419Z digest=sha256:a46213b6924c61b477e5ad63a1ade75e4777711b4af5fd6cddf3ed6010e0a463

Pith citing papers

Observation 72a1682b-7151-4ec9-9350-342cd7bd97bb · inbound

SciCoQA: Quality Assurance for Scientific Paper--Code Alignment cites this paper.

SciCoQA: Quality Assurance for Scientific Paper--Code Alignment When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:20:57.864061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T13:20:07.741347Z digest=sha256:2860c09f2766ae4cdc3452302c5e199df67172e16d342ee2eb1c193a4e782f0b

Observation ca655b53-07fd-4b61-973b-66e3a49045ae · inbound

Grounded autonomous scrutiny at scale: emergent critique from reproduction of published computational physics papers cites this paper.

Grounded autonomous scrutiny at scale: emergent critique from reproduction of published computational physics papers When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:30:31.294703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T14:27:49.721496Z digest=sha256:0160bee42e2bc3af84a8274f603f3c6aad35e2d668be7457b23dd33e4e701014

Observation 0e77ac46-83ba-4ad8-b712-fb2df2ad8082 · inbound

Toward an Engineering of Science: Rebalancing Generation and Verification in the Age of AI cites this paper.

Toward an Engineering of Science: Rebalancing Generation and Verification in the Age of AI When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:25.773935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T05:08:12.470413Z digest=sha256:3b5892d298deb0194a83c276aa5d3da1a5f183cf0767b460417c00a48b90287c

Observation 3d31b126-e3a3-473a-8a3b-f5fb6b4cfc6e · inbound

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery cites this paper.

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:21.657844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T04:46:43.679185Z digest=sha256:9973b22e9835026020311dbb6f0e7c91f620bd64c744ccd408a7425cab6f9b4d

Observation cdeecb52-f9fc-4dae-a1b3-c5278c3f73dd · inbound

ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues cites this paper.

ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:10:20.188905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T01:01:24.456670Z digest=sha256:79afcc6d53d7ec285a2d047d7d0d979142ef8bcd57acedaaae429bf312033af6

Observation f42c506e-c08d-415d-9e9e-82acc792eee6 · inbound

Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy cites this paper.

Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:50:00.401307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-25T23:24:55.361203Z digest=sha256:7858089a1b54e9a85c82a7e6660af081ee95963a8ec84e0e2aaa6896f117b6f1

Observation e71ca82e-6069-4fb4-9765-b6404ae24430 · inbound

Towards Automating Scientific Review with Google's Paper Assistant Tool cites this paper.

Towards Automating Scientific Review with Google's Paper Assistant Tool When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:05:50.908192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T04:15:47.487244Z digest=sha256:72dd52d6570ecb5972cba6a83b4f1a7743280fcd658eb8eb21b33db7b51a927b

Observation 5108484b-d59b-4269-bb6d-1d47841309d2 · inbound

Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System cites this paper.

Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T04:44:25.941162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:44:25.941162Z digest=sha256:1529ba28c9c0a49e1a9f66cf5efb1b2f248ae05bdce992d687d5c46ab19c1235