Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:51:34.451656Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 8 inbound Pith citation observations for arXiv:2505.11855.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:51:34.451656Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T04:44:25.941162Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
92 of 92 outbound references displayed
External citation measurements
4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 2d0c873f-7ae6-4190-a326-0335623f73a2 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Improving language understanding by generative pre-training
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f415ccc6-1207-45a5-904d-3aeb01b29bfc · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bb2e622-0bb7-408f-a4e1-f37537f1d806 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac595b49-9203-461e-833e-4da9b9cbcc4a · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Gpqa: A graduate-level google-proof q&a benchmark
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d27e811-b29f-49ed-90d8-a9872ccb6b1f · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da9cf596-f628-4906-a4f4-a6e4f9b50e59 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceda3ba0-d88a-4be6-b092-16dbaea4c496 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Can chatgpt be used to generate scientific hypotheses?Journal of Materiomics, 10(3):578–584, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99711ad0-3f2b-43a4-be06-0088f70f0552 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research PaSa: An LLM Agent for Comprehensive Academic Paper Search
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 541aea6a-49a3-4b13-8ae6-08879e11df6c · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Generative ai in writing research papers: a new type of algorithmic bias and uncertainty in scholarly work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0197d09-e72a-4d56-8331-23c566057272 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Towards an AI co-scientist
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a273955e-0333-4ad0-93ea-c6a6d6ae9f9b · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc2175da-2857-4477-a8b0-86e5b5516e6b · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Ai mirrors experimental science to uncover a novel mechanism of gene transfer crucial to bacterial evolu- tion.bioRxiv, pages 2025–02, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e86aaef8-191e-4d85-82f3-6ae296b78e44 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ade9501b-4d23-4c27-b912-466c7bcc6197 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Quantum many-body physics calculations with large language models.Communications Physics, 8(1):49, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1774546d-4a06-42b0-9f91-72d021a102e8 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Alphaevolve: a gemini-powered coding agent for design- ing advanced algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 133dfc56-28c4-423d-9371-dfd4d3188fd4 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109c1cbc-3da2-454a-ba2b-ab0daef44d97 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research TabFact: A Large-scale Dataset for Table-based Fact Verification
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c58a3b7-a83f-492c-a3b7-642111f06c21 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research A review on fact extraction and verification.ACM Computing Surveys (CSUR), 55(1):1–35, 2021
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712017a8-56f3-49f3-8dd0-237e434a9bc4 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Poly-FEVER: A Multilingual Fact Verification Benchmark for Hallucination Detection in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 723fec5a-acf5-48d5-919e-c29c0c49e44b · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Sciclaims: An end-to-end generative system for biomedical claim analysis.arXiv preprint arXiv:2503.18526, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7388828d-055b-4328-8fa8-cd8913315ca5 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 242412c1-3f39-4abc-b6a7-120967c4bc60 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2abce3ba-e442-40e7-9147-7a946b141602 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research NLPeer: A Unified Resource for the Computational Study of Peer Review
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89365f2c-70c9-488d-b475-3071e6e4861d · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research PeerQA: A Scientific Question Answering Dataset from Peer Reviews
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 995e37ec-49b5-4513-9f45-8e4c9112404b · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Make every example count: On the stability and utility of self-influence for learning from noisy NLP datasets
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd1c5db-eafa-4072-8e72-fb88700f74ea · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research FEVER: a large-scale dataset for fact extraction and VERification
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bf741d7-66c3-47cd-8cf1-e978ac8fddc1 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Fact or fiction: Verifying scientific claims
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 366cf880-d97a-4b64-bcc1-3e06231c9c8c · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Moprd: A multidisciplinary open peer review dataset.Neural Computing and Applications, 35(34): 24191–24206, 2023
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f32e4cb3-4ecd-4559-b0a5-c61869155b7c · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Automatically evaluating the paper reviewing capability of large language models.arXiv preprint arXiv:2502.17086, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ddae998-2427-48ce-a8cb-3b0b85e0a906 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Openai o3 and o4-mini system card
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 506c3022-380a-427c-adaa-db0cbd874010 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research The llama 4 herd: The beginning of a new era of natively multimodal ai innova- tion
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9b5f2646-e019-4f8d-bcdd-87cb61850f39 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research WithdrarXiv: A Large-Scale Dataset for Retraction Study
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47bf3033-2a12-4303-a279-2859e435febc · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Classification and analysis of pubpeer comments: How a web journal club is used.Journal of the Association for Information Science and Technology, 73(5):655–670, 2022
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8861ffd9-a113-45a4-9202-8a98cf225664 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research American Invitational Mathematics Examination – AIME
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4fec5ec3-d608-4198-ba5c-cb03ba127e45 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360c5245-dd0b-46cf-acee-a4cb9596f492 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20dcb9f9-441d-4f74-9372-316a4b4bfc04 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Paper2code: Automating code generation from scientific papers in machine learning.arXiv preprint arXiv:2504.17192, 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b09c2bdc-cc80-4684-a8d9-a3148af549c7 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research GPT-4o System Card
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 030f2c9e-7b5b-4f77-98e1-baea604c75e7 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research tiktoken: A fast bpe tokeniser for use with openai’s models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 31acb379-afae-4f1f-953f-aa1742758e0c · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 972e8633-bd04-4fa4-a737-f3968ba35c81 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Spoc: Search-based pseudocode to code.Advances in Neural Information Processing Systems, 32, 2019
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ca194c5-6b59-45a3-b1bd-b1d134dedf54 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Evaluating Large Language Models Trained on Code
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 103b4b0a-850c-46c2-981d-ad473026c98b · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c574a8dc-7f3b-4197-9b32-19d7efdcc432 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Gemini 2.5 pro
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ce12ad1f-b99c-4e61-b13b-5634fd66017f · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Gemini 2.0 Flash Lite
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c7ff58c0-381e-407a-a38c-75c94493c9a1 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Claude 3.7 Sonnet System Card
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 73454a53-a82b-4e39-ae00-4412a4cda119 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Qwen2.5-VL Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49f07a06-c9e9-45c9-8a18-05b0538ed520 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24328483-875b-4710-b6b5-f970455c3e0d · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4257c18e-5322-4c01-b4d1-aa2024677069 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b3432f5-3308-43d9-b4bf-6bc420ca87d2 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Humanity's Last Exam
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8903a918-5c24-4ed9-a450-2ed0c00f237c · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research On calibration of modern neural networks
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac5c5aab-fb44-4f46-9f9a-5cb0aabc752b · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift.Advances in neural information processing systems, 32, 2019
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ca77c99-a692-42ea-b748-f8df352518b5 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff70fb3a-935a-482e-bcb5-0503799b4a31 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research DeepSeek-V3 Technical Report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3adc1fc3-b475-46a6-af57-d820e9de4aa3 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Qwen3 Technical Report
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea95058-77e2-4f09-8066-12f48bc7a8c8 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Multiplicative Chow-K\"unneth decomposition and homology splitting of configuration spaces
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 90e9562a-eddf-451a-bc50-49e4330d0e09 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Superacid in situ protected synthesis of covalent organic frameworks
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e891a275-7c4b-4f7a-83d0-57def7bc15a2 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02dd2aac-e86c-4eb7-ab5b-9ef2fcce6734 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Discourse-Based Objectives for Fast Unsupervised Sentence Representation Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd724038-6b4c-47e3-8256-1cfea1e9f252 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2429bfb3-af12-4b89-95a9-5758101ce73e · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Reviewer2: Optimizing Review Generation Through Prompt Generation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cd76ba9-7ae9-43ed-bd6b-2042e5c2b938 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4569237-b995-435a-98b8-c9a1d8fa30fa · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Inconsistency in Conference Peer Review: Revisiting the 2014 NeurIPS Experiment
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f427bc4-e695-4e1a-ae83-40ad7caf3538 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research A noise audit of the peer review of a scientific article: a wpom journal case study.WPOM-Working Papers on Operations Management, 14(2): 137–166, 2023
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c255511a-bdf6-4314-abc3-6df78478d8b4 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f87f38e3-97d1-4a75-9eb4-b6ee4a27b08c · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Scaling Scaling Laws with Board Games
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e4789b-faab-4219-b6d3-3ce4d310aef1 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ddd532-0db0-4ebf-a81e-546525a6a653 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1425ae0-1b37-46b6-8cdd-f23d5e7119fb · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research s1: Simple test-time scaling
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d261402-f240-46ea-bbf0-b62b886a6372 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Finding flawed fictions: Evaluating complex reasoning in language models via plot hole detection.arXiv preprint arXiv:2504.11900, 2025
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 869e8f96-440a-4dcf-99f0-a66852e2da12 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Algebraic description of complex conjugation on cohomology of a smooth projective hypersurface
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 80dbaa64-993c-4bb8-a430-54c807394d48 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b3ef3239-d156-4650-972f-a5b1c752b85c · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Limitations
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 49dfd2de-6d24-4031-a9f3-2a5335bd87ce · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d86d75e8-32a0-4d47-acde-21aa4cefa30b · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cf924b38-2096-4192-8dcb-fdd0ce54593c · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Conversely, false positives may occur when:
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 853f51a8-fa0b-4256-b657-a14c7b494968 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research reasoning effort
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a727aaa3-0704-4346-b6d8-c7ac26147cf3 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research decreased from 0.116 eV. . . to 1.03 eV
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d7fe2f89-fdda-4110-8743-99f9400eafef · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 26c88657-f69c-44f2-b740-e6053cc4655b · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d57fc1df-fb11-4c79-8579-a1ad812c9dda · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5b4e1349-aff1-4836-aa97-054c13dee291 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Generation Prompt
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a2214f1e-fd27-435e-be96-2c31d928e4fc · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research annotations
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7c8e0159-fcb3-4911-8710-07d2dbf85ecc · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research predictions
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4f5c2bb1-0003-4821-9c10-0f1600bac133 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e1db4614-0259-4a17-8ee2-b8433aed7f98 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research location
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d92ebf6b-0b6b-46cb-8fcd-3df1b91b7f45 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research matches": [ {
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8ebfa0d1-ebaa-4ddd-8f6f-de226baa6c33 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3eb44cf9-32c6-442e-92f4-9daa85ffeafe · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Table 4:Mean and standard deviation of pass@K for o3 (K∈{1,2,4} ) by error category (left) and paper category (right)
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation df9bbf82-cc38-495a-8ad6-683aa4e36791 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research doi: 10.18653/v1/2020.emnlp-main.609
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffb36bc-e6e6-42e7-b077-2450dceb4fc6 · outbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72a1682b-7151-4ec9-9350-342cd7bd97bb · inbound
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ca655b53-07fd-4b61-973b-66e3a49045ae · inbound
Grounded autonomous scrutiny at scale: emergent critique from reproduction of published computational physics papers When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0e77ac46-83ba-4ad8-b712-fb2df2ad8082 · inbound
Toward an Engineering of Science: Rebalancing Generation and Verification in the Age of AI When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3d31b126-e3a3-473a-8a3b-f5fb6b4cfc6e · inbound
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cdeecb52-f9fc-4dae-a1b3-c5278c3f73dd · inbound
ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f42c506e-c08d-415d-9e9e-82acc792eee6 · inbound
Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e71ca82e-6069-4fb4-9765-b6404ae24430 · inbound
Towards Automating Scientific Review with Google's Paper Assistant Tool When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5108484b-d59b-4269-bb6d-1d47841309d2 · inbound
Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.