Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:07:19.897165Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2508.17580.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:07:19.897165Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T04:09:46.616019Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T12:06:04.092759Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1f01bbf6-b12f-46b1-b9ae-2e3b1a5845af · outbound
UQ: Assessing Language Models on Unsolved Questions Piqa: Reasoning about physical commonsense in natural language
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5a54b432-f1ee-46b1-bbef-b1f27ff77a20 · outbound
UQ: Assessing Language Models on Unsolved Questions Evaluating Large Language Models Trained on Code
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c1e60d2-ce2f-4c88-b390-652795b68d92 · outbound
UQ: Assessing Language Models on Unsolved Questions Chatbot arena: An open platform for evaluating llms by human preference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e4e6170e-de5a-4d34-b437-6df8349d1a78 · outbound
UQ: Assessing Language Models on Unsolved Questions ARC Prize 2024: Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2176fa22-ebda-4452-93cd-285dcdc6d287 · outbound
UQ: Assessing Language Models on Unsolved Questions Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b463a90-bfec-45d1-b90e-d542c6e823f5 · outbound
UQ: Assessing Language Models on Unsolved Questions Chain-of-Verification Reduces Hallucination in Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0069902-9909-4c98-8a27-57bfafb52b46 · outbound
UQ: Assessing Language Models on Unsolved Questions AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f96f91e-c4d8-41ae-a264-112d60ea3305 · outbound
UQ: Assessing Language Models on Unsolved Questions Are We Done with MMLU?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ab611a-d399-4435-a64c-f27784877ef4 · outbound
UQ: Assessing Language Models on Unsolved Questions Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 593bde5f-a009-44e0-89ef-0bb417085f93 · outbound
UQ: Assessing Language Models on Unsolved Questions Great Models Think Alike and this Undermines AI Oversight
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd69e84c-1409-4aa7-8517-78fe5e4d119c · outbound
UQ: Assessing Language Models on Unsolved Questions Measuring Coding Challenge Competence With APPS
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a09cdff2-45c3-4aa9-bf4b-68c6d2d6426a · outbound
UQ: Assessing Language Models on Unsolved Questions Measuring Massive Multitask Language Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcbcf013-b427-4ecb-95de-df825ee3cb95 · outbound
UQ: Assessing Language Models on Unsolved Questions Measuring Mathematical Problem Solving With the MATH Dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79c6791a-03b7-4619-aa14-76667eb6461e · outbound
UQ: Assessing Language Models on Unsolved Questions MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cfd3e5b-372f-4122-b6cc-97d91e78eacc · outbound
UQ: Assessing Language Models on Unsolved Questions Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9effe6b5-3a3f-4cf8-b0bf-7c9cf841f69b · outbound
UQ: Assessing Language Models on Unsolved Questions GPT-4o System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d031044-a077-4dd3-af42-2e333a3d4d62 · outbound
UQ: Assessing Language Models on Unsolved Questions LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 971742a0-244e-4bb1-ac3d-08e9653d5389 · outbound
UQ: Assessing Language Models on Unsolved Questions When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa30218c-fcfd-4868-90e8-1ec7d03d2667 · outbound
UQ: Assessing Language Models on Unsolved Questions Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e32cb678-272d-4129-852e-ff4739ee86ff · outbound
UQ: Assessing Language Models on Unsolved Questions Verdict: A library for scaling judge-time compute
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 190333ed-45ec-463e-9577-a67c0a8fdb35 · outbound
UQ: Assessing Language Models on Unsolved Questions Llm-as-an-interviewer: Beyond static testing through dynamic llm evaluation, 2025
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 158af6cb-3842-49c6-a647-0dd64f7d1c0d · outbound
UQ: Assessing Language Models on Unsolved Questions Prometheus: Inducing fine-grained evaluation capability in language models, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 519f40fe-6840-4a3d-98a5-c3c18e81dd31 · outbound
UQ: Assessing Language Models on Unsolved Questions Prometheus 2: An open source language model specialized in evaluating other language models, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b93b5b2a-5278-4db2-b564-4f15cb79e6ef · outbound
UQ: Assessing Language Models on Unsolved Questions Scaling evaluation-time compute with reasoning models as process evaluators, 2025
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b77f1e55-ca6b-4450-ac93-948aef2b3bbc · outbound
UQ: Assessing Language Models on Unsolved Questions Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b75cb40a-367c-4ecf-9060-cdf0a7d783da · outbound
UQ: Assessing Language Models on Unsolved Questions The measurement of observer agreement for categorical data
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ddd119d8-450b-4eb6-b495-24bfb2c8b879 · outbound
UQ: Assessing Language Models on Unsolved Questions FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5edad0b-6467-463f-8ec4-68fefa0054bc · outbound
UQ: Assessing Language Models on Unsolved Questions Holistic Evaluation of Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 374897c1-dee4-4c24-be9c-51e809777187 · outbound
UQ: Assessing Language Models on Unsolved Questions Let's Verify Step by Step
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cfc8a7c-e9e9-4a2e-a6f8-b88c3c7ca7da · outbound
UQ: Assessing Language Models on Unsolved Questions WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5e8e7a-8fe5-4d40-98c1-2638cc7c8139 · outbound
UQ: Assessing Language Models on Unsolved Questions Evaluating Verifiability in Generative Search Engines
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ef7b271-c72c-4a61-8ea2-4d5fe30ccce6 · outbound
UQ: Assessing Language Models on Unsolved Questions The lean 4 theorem prover and programming language
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b9ea21ba-ee35-4da9-a8cb-eb220e3c4895 · outbound
UQ: Assessing Language Models on Unsolved Questions 2024 aime i
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c95d7ef9-422e-45a1-99ef-461c49e60bd0 · outbound
UQ: Assessing Language Models on Unsolved Questions List of open problems in sublinear algorithms
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4a628fec-1c3f-4286-8a62-7c19ff390b07 · outbound
UQ: Assessing Language Models on Unsolved Questions Introducing deep research
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6c25daa6-4998-412a-aae5-bf9e5fbdf080 · outbound
UQ: Assessing Language Models on Unsolved Questions Introducing OpenAI o1
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e3eceb3f-2fb3-406b-a739-7e4813bb1314 · outbound
UQ: Assessing Language Models on Unsolved Questions Introducing OpenAI o3 and o4‑mini
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation deceb3d9-7c72-4a20-9e7f-d229cbc500aa · outbound
UQ: Assessing Language Models on Unsolved Questions Llm evaluators recognize and favor their own generations
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 56d7c659-603b-4be9-bd3b-05300528dfcf · outbound
UQ: Assessing Language Models on Unsolved Questions For better or worse, benchmarks shape a field
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5c2371bb-a85c-4f62-af4c-210971e40705 · outbound
UQ: Assessing Language Models on Unsolved Questions KILT: a Benchmark for Knowledge Intensive Language Tasks
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2197aa18-381d-4f4b-8320-10369cff45bd · outbound
UQ: Assessing Language Models on Unsolved Questions Humanity's last exam, 2025
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd37a82f-a72d-47b4-bf2a-36bb7c1ad53f · outbound
UQ: Assessing Language Models on Unsolved Questions SQuAD: 100,000+ Questions for Machine Comprehension of Text
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6218fdfb-c2e4-45b0-9ea3-dc64640d69ae · outbound
UQ: Assessing Language Models on Unsolved Questions Gpqa: A graduate-level google-proof q&a benchmark
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 326f22e5-6d1b-4715-85f1-3cc44efa40da · outbound
UQ: Assessing Language Models on Unsolved Questions Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb828bae-2d44-495f-8143-88060708b8ea · outbound
UQ: Assessing Language Models on Unsolved Questions The Leaderboard Illusion
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1608eedb-7ee9-4894-8458-1124658ba64f · outbound
UQ: Assessing Language Models on Unsolved Questions Stack Exchange
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0410b2a1-1b7c-4dad-9a8f-d2539af0772e · outbound
UQ: Assessing Language Models on Unsolved Questions Terminal‑Bench : A benchmark for ai agents in terminal environments
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f163747-562a-4de4-97a2-59c4637bf9e5 · outbound
UQ: Assessing Language Models on Unsolved Questions Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61312f11-fa5f-47f3-b869-c30b2ab19004 · outbound
UQ: Assessing Language Models on Unsolved Questions Supergpqa: Scaling llm evaluation across 285 graduate disciplines, 2025
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5716fd16-daec-44b3-8206-cc4a9af77846 · outbound
UQ: Assessing Language Models on Unsolved Questions SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a02e5a-b954-4fb5-99a2-6e485959b2f8 · outbound
UQ: Assessing Language Models on Unsolved Questions GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af631803-152b-46b6-9303-cd34c843b03e · outbound
UQ: Assessing Language Models on Unsolved Questions Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f5b672-60a7-49ca-80f8-3ab96fef1669 · outbound
UQ: Assessing Language Models on Unsolved Questions PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f7bc9b-f90c-4732-aa86-290f9a53859c · outbound
UQ: Assessing Language Models on Unsolved Questions Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bbab3f8d-f82d-429a-85a0-c1756cd4f1b6 · outbound
UQ: Assessing Language Models on Unsolved Questions Self-Preference Bias in LLM-as-a-Judge
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9f21bba-2b91-4da7-a383-e50b111deb34 · outbound
UQ: Assessing Language Models on Unsolved Questions Measuring short-form factuality in large language models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a23df0-5f8f-4c0a-9cbf-4f3bcb23d3b6 · outbound
UQ: Assessing Language Models on Unsolved Questions BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7d7cb80-05cc-4274-92e7-bc59da4121d1 · outbound
UQ: Assessing Language Models on Unsolved Questions LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ecb45b0-6194-47f7-924f-ae0483dd783e · outbound
UQ: Assessing Language Models on Unsolved Questions Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922a6fd3-f114-4d7b-bcbc-f62e11b2bddc · outbound
UQ: Assessing Language Models on Unsolved Questions Jimenez, Alex L
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce5b0cba-cc9a-4e86-bbe2-5c0fc907b01e · outbound
UQ: Assessing Language Models on Unsolved Questions $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f237b55-1857-4f1e-adad-db3377df5413 · outbound
UQ: Assessing Language Models on Unsolved Questions Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66825604-8add-46ae-9958-6f17510d1834 · outbound
UQ: Assessing Language Models on Unsolved Questions HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd6385f5-84a6-49e2-88ce-eee3f128ed90 · outbound
UQ: Assessing Language Models on Unsolved Questions A careful examination of large language model performance on grade school arithmetic
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7adaf2fb-c55c-4003-be7f-b533fa00d918 · outbound
UQ: Assessing Language Models on Unsolved Questions Challenges in Trustworthy Human Evaluation of Chatbots
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe1d022-d3b2-45cf-8ff6-2d05aac7e3ba · outbound
UQ: Assessing Language Models on Unsolved Questions Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dde716a-4b75-4d50-b822-de3b1a0e9cca · outbound
UQ: Assessing Language Models on Unsolved Questions AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73bd8a55-7730-41ab-a739-ca47bfd026d9 · outbound
UQ: Assessing Language Models on Unsolved Questions LIMA: Less Is More for Alignment
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c801d62d-bb02-4a36-ab60-326a1d65e71a · outbound
UQ: Assessing Language Models on Unsolved Questions Reinforcing General Reasoning without Verifiers
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1adc20b5-9a0a-4593-926a-6c1846e92381 · outbound
UQ: Assessing Language Models on Unsolved Questions Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2025
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 51e8effd-fb12-4824-8bf6-19061be8cfac · inbound
MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval UQ: Assessing Language Models on Unsolved Questions
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.