Pith. sign in

Paper Citation Record · LEDGER

UQ: Assessing Language Models on Unsolved Questions

As of 18 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2508.17580.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17580 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:07:19.897165Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T04:09:46.616019Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:06:04.092759Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f01bbf6-b12f-46b1-b9ae-2e3b1a5845af · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

UQ: Assessing Language Models on Unsolved Questions Piqa: Reasoning about physical commonsense in natural language

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.844600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.612229Z digest=sha256:dc2fbc36681eac25b23de4ae30a801dea62d45fc33f02ab641eab5c5be30807a

Observation 5a54b432-f1ee-46b1-bbef-b1f27ff77a20 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

UQ: Assessing Language Models on Unsolved Questions Evaluating Large Language Models Trained on Code

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.617165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.617165Z digest=sha256:2f26fb718c1acd19c8349269e48778f20b8116dadd73c3c0d3b778d9e2030295

Observation 7c1e60d2-ce2f-4c88-b390-652795b68d92 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

UQ: Assessing Language Models on Unsolved Questions Chatbot arena: An open platform for evaluating llms by human preference

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.832642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.621341Z digest=sha256:c2d0753ad07fd44c6ee744693f2f392046d53785ea6fc5d7ef90b47394ff18b1

Observation e4e6170e-de5a-4d34-b437-6df8349d1a78 · outbound

This paper cites ARC Prize 2024: Technical Report.

UQ: Assessing Language Models on Unsolved Questions ARC Prize 2024: Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.625583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.625583Z digest=sha256:a0625f078f4aa36820d23e0e466f9dd6a82d539b4f623bedad96d1eb24facb4c

Observation 2176fa22-ebda-4452-93cd-285dcdc6d287 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

UQ: Assessing Language Models on Unsolved Questions Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.629700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.629700Z digest=sha256:7b94637b51e17dfc7fd73b2b1bd66124e9bbc02be38e1dc13769482713dc910e

Observation 5b463a90-bfec-45d1-b90e-d542c6e823f5 · outbound

This paper cites Chain-of-Verification Reduces Hallucination in Large Language Models.

UQ: Assessing Language Models on Unsolved Questions Chain-of-Verification Reduces Hallucination in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.634145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.634145Z digest=sha256:8ad5a2397b6c4d3e202d1f536e37ab7fe35ce700e19beb7179a772b2099d4aa7

Observation b0069902-9909-4c98-8a27-57bfafb52b46 · outbound

This paper cites AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback.

UQ: Assessing Language Models on Unsolved Questions AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.638212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.638212Z digest=sha256:43a7b7fc91d00dfb9d0ede7517b4fccc13ceb184947790390dcc97f10533341a

Observation 7f96f91e-c4d8-41ae-a264-112d60ea3305 · outbound

This paper cites Are We Done with MMLU?.

UQ: Assessing Language Models on Unsolved Questions Are We Done with MMLU?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.642257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.642257Z digest=sha256:878583315be9c242dd6d52bcb49f37ea9763c97d56f08d50b20577169ed93f67

Observation f1ab611a-d399-4435-a64c-f27784877ef4 · outbound

This paper cites Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024.

UQ: Assessing Language Models on Unsolved Questions Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.646336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.646336Z digest=sha256:481c29ce1d17ac73a72c36170b1d3ecc83388a7ad27785d66da9fb23e8edf604

Observation 593bde5f-a009-44e0-89ef-0bb417085f93 · outbound

This paper cites Great Models Think Alike and this Undermines AI Oversight.

UQ: Assessing Language Models on Unsolved Questions Great Models Think Alike and this Undermines AI Oversight

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.653111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.653111Z digest=sha256:44bf5fc92c8187edf0577e79d9601698228311e97e1c93c1dee51b9197aeeb94

Observation cd69e84c-1409-4aa7-8517-78fe5e4d119c · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

UQ: Assessing Language Models on Unsolved Questions Measuring Coding Challenge Competence With APPS

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.657157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.657157Z digest=sha256:d5f1a5f0326a17f6533a9ce14f176ee5d9fd5e3081aa64bbdf56f550731b975f

Observation a09cdff2-45c3-4aa9-bf4b-68c6d2d6426a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

UQ: Assessing Language Models on Unsolved Questions Measuring Massive Multitask Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.661079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.661079Z digest=sha256:b1130d2f86eb9d04f2e42bdc0ccef109b0d7b91c79bc44bfb652acdb51077164

Observation fcbcf013-b427-4ecb-95de-df825ee3cb95 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

UQ: Assessing Language Models on Unsolved Questions Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.664976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.664976Z digest=sha256:ad61fbaf89eba57525df3a4be3895e90cdf9fff817ebb11c3faed73dd603aa5d

Observation 79c6791a-03b7-4619-aa14-76667eb6461e · outbound

This paper cites MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations.

UQ: Assessing Language Models on Unsolved Questions MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.669042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.669042Z digest=sha256:09bfd48421af79454a64c3d6b38ddc7cd214556274b26c552b6a84f0a84d7f41

Observation 7cfd3e5b-372f-4122-b6cc-97d91e78eacc · outbound

This paper cites Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards.

UQ: Assessing Language Models on Unsolved Questions Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:07:20.365810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.673009Z digest=sha256:c0c962d4709ead3d50f316dc872ebd2b9c3f5ada07049c939537040549bd6575

Observation 9effe6b5-3a3f-4cf8-b0bf-7c9cf841f69b · outbound

This paper cites GPT-4o System Card.

UQ: Assessing Language Models on Unsolved Questions GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.677086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.677086Z digest=sha256:27606822305a369a774656aa714c2a08d5368dce148423219f41b192807bdd33

Observation 1d031044-a077-4dd3-af42-2e333a3d4d62 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

UQ: Assessing Language Models on Unsolved Questions LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.681246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.681246Z digest=sha256:0df1c245d603676f37c0778b429307fa635d05acdb7036b49c1fbbc9580f10d8

Observation 971742a0-244e-4bb1-ac3d-08e9653d5389 · outbound

This paper cites When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories.

UQ: Assessing Language Models on Unsolved Questions When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.685883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.685883Z digest=sha256:06d0c5f1f288ebb286f1a7d9577f6a3bef1ea7d1c94372b22465bea68fe44231

Observation aa30218c-fcfd-4868-90e8-1ec7d03d2667 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

UQ: Assessing Language Models on Unsolved Questions Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.690864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.690864Z digest=sha256:8b21722911e10f78c5d9833a94221671867da2e2b4fc520e43e3c2439ab140d2

Observation e32cb678-272d-4129-852e-ff4739ee86ff · outbound

This paper cites Verdict: A library for scaling judge-time compute.

UQ: Assessing Language Models on Unsolved Questions Verdict: A library for scaling judge-time compute

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.695712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.695712Z digest=sha256:46b8f8afad3a8e9ceef33933ae585ba67944222414cfeda22c5469c4bb793261

Observation 190333ed-45ec-463e-9577-a67c0a8fdb35 · outbound

This paper cites Llm-as-an-interviewer: Beyond static testing through dynamic llm evaluation, 2025.

UQ: Assessing Language Models on Unsolved Questions Llm-as-an-interviewer: Beyond static testing through dynamic llm evaluation, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.805990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.700099Z digest=sha256:b7c52845ba54ef83f6ce016f9e041d550c2e1a48a735875af066415f345f3604

Observation 158af6cb-3842-49c6-a647-0dd64f7d1c0d · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models, 2024.

UQ: Assessing Language Models on Unsolved Questions Prometheus: Inducing fine-grained evaluation capability in language models, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.703854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.703854Z digest=sha256:4dcd093bf5a48038306c336c2350d166bd1d4fbdd4c43c132018a522139d1b09

Observation 519f40fe-6840-4a3d-98a5-c3c18e81dd31 · outbound

This paper cites Prometheus 2: An open source language model specialized in evaluating other language models, 2024.

UQ: Assessing Language Models on Unsolved Questions Prometheus 2: An open source language model specialized in evaluating other language models, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.785977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.707492Z digest=sha256:9993787d4f804735a05b7955a16eef1ffd46f0873e7b101a5e1073e1fe9a8875

Observation b93b5b2a-5278-4db2-b564-4f15cb79e6ef · outbound

This paper cites Scaling evaluation-time compute with reasoning models as process evaluators, 2025.

UQ: Assessing Language Models on Unsolved Questions Scaling evaluation-time compute with reasoning models as process evaluators, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.773770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.711145Z digest=sha256:9d8201a915295d5d52409842b838919816f1157523077d8146c5b0a25d696f8c

Observation b77f1e55-ca6b-4450-ac93-948aef2b3bbc · outbound

This paper cites Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

UQ: Assessing Language Models on Unsolved Questions Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.761723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.714954Z digest=sha256:2ef8f0f8236a0f16e20c523a28c30454734faa3f6525960db8b095d28b6c3a10

Observation b75cb40a-367c-4ecf-9060-cdf0a7d783da · outbound

This paper cites The measurement of observer agreement for categorical data.

UQ: Assessing Language Models on Unsolved Questions The measurement of observer agreement for categorical data

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.749131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.718615Z digest=sha256:78c419695a9838031bc20846e69d68fa27b7c82b207a9aab76bfe1d7efd7b43d

Observation ddd119d8-450b-4eb6-b495-24bfb2c8b879 · outbound

This paper cites FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets.

UQ: Assessing Language Models on Unsolved Questions FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.722330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.722330Z digest=sha256:afeb1b4eb60600f56a49921ab730eff62e233e9dcb13eac2e6508cc9449d52aa

Observation c5edad0b-6467-463f-8ec4-68fefa0054bc · outbound

This paper cites Holistic Evaluation of Language Models.

UQ: Assessing Language Models on Unsolved Questions Holistic Evaluation of Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.726851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.726851Z digest=sha256:2cc27bac0d5def2b8451e1498f13c6c9bb8d385cc065501dd64c705ab78a6fa6

Observation 374897c1-dee4-4c24-be9c-51e809777187 · outbound

This paper cites Let's Verify Step by Step.

UQ: Assessing Language Models on Unsolved Questions Let's Verify Step by Step

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.730930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.730930Z digest=sha256:21674b9eca387e267267e02c506412a51ff3d1bb1829e4fccec15a7f5e03163c

Observation 8cfc8a7c-e9e9-4a2e-a6f8-b88c3c7ca7da · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

UQ: Assessing Language Models on Unsolved Questions WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.734688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.734688Z digest=sha256:6a83f0ba743d05e12139de2362feae188c1b6ebc381eb28c1d2cdc820ea8be3b

Observation 3f5e8e7a-8fe5-4d40-98c1-2638cc7c8139 · outbound

This paper cites Evaluating Verifiability in Generative Search Engines.

UQ: Assessing Language Models on Unsolved Questions Evaluating Verifiability in Generative Search Engines

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.738927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.738927Z digest=sha256:0c379279d2b47b97f03450f0bd26f9f05bc7bbd792b7d6a2465f079a3622c1eb

Observation 5ef7b271-c72c-4a61-8ea2-4d5fe30ccce6 · outbound

This paper cites The lean 4 theorem prover and programming language.

UQ: Assessing Language Models on Unsolved Questions The lean 4 theorem prover and programming language

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.736175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.742881Z digest=sha256:c82fe8d42e5ad237ff310f1ec39e4293c930b2cde442396f05d3fd874b7966c5

Observation b9ea21ba-ee35-4da9-a8cb-eb220e3c4895 · outbound

This paper cites 2024 aime i.

UQ: Assessing Language Models on Unsolved Questions 2024 aime i

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.723466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.747278Z digest=sha256:833b867d0fd9ae6c96f738f7676a1c79a8e90b06eb6764c6ef8db497b178e1b1

Observation c95d7ef9-422e-45a1-99ef-461c49e60bd0 · outbound

This paper cites List of open problems in sublinear algorithms.

UQ: Assessing Language Models on Unsolved Questions List of open problems in sublinear algorithms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.705112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.751098Z digest=sha256:073d5aaa8d7bcdd4f1ffd0685081ed1046befb676b569ae60dfb37d530662612

Observation 4a628fec-1c3f-4286-8a62-7c19ff390b07 · outbound

This paper cites Introducing deep research.

UQ: Assessing Language Models on Unsolved Questions Introducing deep research

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.687109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.754764Z digest=sha256:96f991a529acc890cdb039027f5a50e8170c6287f98fc3bf7768e51f355aebfe

Observation 6c25daa6-4998-412a-aae5-bf9e5fbdf080 · outbound

This paper cites Introducing OpenAI o1.

UQ: Assessing Language Models on Unsolved Questions Introducing OpenAI o1

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.671042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.758160Z digest=sha256:e5690c3f796586d57bed2f55dd9083cbf3ab5ebad3d24b4235b50444fda3d7af

Observation e3eceb3f-2fb3-406b-a739-7e4813bb1314 · outbound

This paper cites Introducing OpenAI o3 and o4‑mini.

UQ: Assessing Language Models on Unsolved Questions Introducing OpenAI o3 and o4‑mini

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.658233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.761596Z digest=sha256:f4426ac63d5b31635402e43793be559ea4c67f1e4a87677f50998074c43ff47d

Observation deceb3d9-7c72-4a20-9e7f-d229cbc500aa · outbound

This paper cites Llm evaluators recognize and favor their own generations.

UQ: Assessing Language Models on Unsolved Questions Llm evaluators recognize and favor their own generations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.646999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.765577Z digest=sha256:171268e85dc418304f82890fa137c580753a712556afefe84e7e3f4d1f685807

Observation 56d7c659-603b-4be9-bd3b-05300528dfcf · outbound

This paper cites For better or worse, benchmarks shape a field.

UQ: Assessing Language Models on Unsolved Questions For better or worse, benchmarks shape a field

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.635111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.769878Z digest=sha256:b78de2c0a4c187e8f4f9f347c9ed4b44fbd0c3eece8ae68804031241e229ee1c

Observation 5c2371bb-a85c-4f62-af4c-210971e40705 · outbound

This paper cites KILT: a Benchmark for Knowledge Intensive Language Tasks.

UQ: Assessing Language Models on Unsolved Questions KILT: a Benchmark for Knowledge Intensive Language Tasks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.774679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.774679Z digest=sha256:85dd8d92ea0c234cd67d2df3f429cbc9e1e15f6d460ef2053666504747573ef9

Observation 2197aa18-381d-4f4b-8320-10369cff45bd · outbound

This paper cites Humanity's last exam, 2025.

UQ: Assessing Language Models on Unsolved Questions Humanity's last exam, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.622441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.779121Z digest=sha256:45e6dad5b71c845fb1b9c6a04f5708e4236928b624954188e2518ef7908aad39

Observation fd37a82f-a72d-47b4-bf2a-36bb7c1ad53f · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

UQ: Assessing Language Models on Unsolved Questions SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.783306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.783306Z digest=sha256:06444f5cb2f3877b6b891174c0da60907dea040faf878904d967f21989b567fb

Observation 6218fdfb-c2e4-45b0-9ea3-dc64640d69ae · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

UQ: Assessing Language Models on Unsolved Questions Gpqa: A graduate-level google-proof q&a benchmark

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.609957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.787188Z digest=sha256:60006618b11951e8abf728773cfb79fd4324687dbf341a8fd82a27ec716bb4c2

Observation 326f22e5-6d1b-4715-85f1-3cc44efa40da · outbound

This paper cites Measurement to Meaning: A Validity-Centered Framework for AI Evaluation.

UQ: Assessing Language Models on Unsolved Questions Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.792055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.792055Z digest=sha256:c2d7f238d3d5e26ba7e0d4ee196164b4b979226be3be9d71e942a7bfd8ca19a0

Observation fb828bae-2d44-495f-8143-88060708b8ea · outbound

This paper cites The Leaderboard Illusion.

UQ: Assessing Language Models on Unsolved Questions The Leaderboard Illusion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.796329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.796329Z digest=sha256:4e9150760429b3abec39027fba4edda48d975866ba4fa270e91338a31652764c

Observation 1608eedb-7ee9-4894-8458-1124658ba64f · outbound

This paper cites Stack Exchange.

UQ: Assessing Language Models on Unsolved Questions Stack Exchange

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.598457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.800390Z digest=sha256:77f581b234b7a5dce0fd97721cc1ec631bfc82bf8cf2b8ae7adf392fa963c260

Observation 0410b2a1-1b7c-4dad-9a8f-d2539af0772e · outbound

This paper cites Terminal‑Bench : A benchmark for ai agents in terminal environments.

UQ: Assessing Language Models on Unsolved Questions Terminal‑Bench : A benchmark for ai agents in terminal environments

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.586555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.804671Z digest=sha256:800aa9ddd3dcc89560b7167fde9623fbfc67a9c5670807eaebe0d224cd72d790

Observation 0f163747-562a-4de4-97a2-59c4637bf9e5 · outbound

This paper cites Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead.

UQ: Assessing Language Models on Unsolved Questions Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.808716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.808716Z digest=sha256:45fde431d972c9680b286d60ba3f69236526681128404ac041a7fb517d3df3ce

Observation 61312f11-fa5f-47f3-b869-c30b2ab19004 · outbound

This paper cites Supergpqa: Scaling llm evaluation across 285 graduate disciplines, 2025.

UQ: Assessing Language Models on Unsolved Questions Supergpqa: Scaling llm evaluation across 285 graduate disciplines, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.572183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.812701Z digest=sha256:7491fc6c11f15937caf0a91d6c97876d51543f6710fd17f07f2ddbe76f698819

Observation 5716fd16-daec-44b3-8206-cc4a9af77846 · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.

UQ: Assessing Language Models on Unsolved Questions SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.816806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.816806Z digest=sha256:79814df4d7a4f9cb0ca6abd75d3648da4f735027d8825f0e90199bfaf02fc35c

Observation d5a02e5a-b954-4fb5-99a2-6e485959b2f8 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

UQ: Assessing Language Models on Unsolved Questions GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.820943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.820943Z digest=sha256:0aa6826e8c272c0d330a261ded74ee0844475a4859323694445910c4fdcd38f3

Observation af631803-152b-46b6-9303-cd34c843b03e · outbound

This paper cites an unresolved cited work.

UQ: Assessing Language Models on Unsolved Questions Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.824905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.824905Z digest=sha256:3019fbb2c1156c92e8e7f2fb06ff2a0e39983dc3ce8549dfe6e9bea2b89a617b

Observation 24f5b672-60a7-49ca-80f8-3ab96fef1669 · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

UQ: Assessing Language Models on Unsolved Questions PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.828751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.828751Z digest=sha256:579bab474381b4ed8ad92d5ad7c9909044ebc09c51641db3a6f0b5a942365399

Observation 77f7bc9b-f90c-4732-aa86-290f9a53859c · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

UQ: Assessing Language Models on Unsolved Questions Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.548388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.832809Z digest=sha256:d1ecd76f16453b85b2faf353f048bcd0c20f077c1d39f413ab7ec0b81fb8a522

Observation bbab3f8d-f82d-429a-85a0-c1756cd4f1b6 · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

UQ: Assessing Language Models on Unsolved Questions Self-Preference Bias in LLM-as-a-Judge

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.836673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.836673Z digest=sha256:acb74b48cfc6c95c3287eb5ee485e2d7bd7ccbe5a7c1d28b6e50776378250108

Observation c9f21bba-2b91-4da7-a383-e50b111deb34 · outbound

This paper cites Measuring short-form factuality in large language models.

UQ: Assessing Language Models on Unsolved Questions Measuring short-form factuality in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.840580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.840580Z digest=sha256:1d61af165b657af8d32c4462e371e90742f8623267f205bee2d58527c69514cc

Observation 18a23df0-5f8f-4c0a-9cbf-4f3bcb23d3b6 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

UQ: Assessing Language Models on Unsolved Questions BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.844366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.844366Z digest=sha256:156dad55193dd18f16e329a07206dfc53c5ccc60e65909d70b477c777763c548

Observation b7d7cb80-05cc-4274-92e7-bc59da4121d1 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

UQ: Assessing Language Models on Unsolved Questions LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.848296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.848296Z digest=sha256:80b65e3eefb9661759d807873c7b659d390e7e5ea5e114ff973e81829e7b1f42

Observation 7ecb45b0-6194-47f7-924f-ae0483dd783e · outbound

This paper cites Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement.

UQ: Assessing Language Models on Unsolved Questions Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.852288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.852288Z digest=sha256:268494c77fd1113a028a8a1a6cd263622c29e1f1b0720db5887c88b7765dda9b

Observation 922a6fd3-f114-4d7b-bcbc-f62e11b2bddc · outbound

This paper cites Jimenez, Alex L.

UQ: Assessing Language Models on Unsolved Questions Jimenez, Alex L

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.856397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.856397Z digest=sha256:7c9f8c2eab3d9785e06c25eb0fcbe7c6b38ed30d1378bbae5c182f5c3f6f1f4f

Observation ce5b0cba-cc9a-4e86-bbe2-5c0fc907b01e · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

UQ: Assessing Language Models on Unsolved Questions $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.859994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.859994Z digest=sha256:43d993b2fc854de776261ec152a94e1364236fa887ea80c6a5c5b0d6c55275a7

Observation 4f237b55-1857-4f1e-adad-db3377df5413 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

UQ: Assessing Language Models on Unsolved Questions Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.863897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.863897Z digest=sha256:b22edc7319ec6f16bf793bbfbaea96f8179d8aca52b2c3e507be4a2c864af100

Observation 66825604-8add-46ae-9958-6f17510d1834 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

UQ: Assessing Language Models on Unsolved Questions HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.868052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.868052Z digest=sha256:3c4a1ded4e092eef81beb43490356d8e0b1509635c1a831dd4e19fd98305b6cd

Observation fd6385f5-84a6-49e2-88ce-eee3f128ed90 · outbound

This paper cites A careful examination of large language model performance on grade school arithmetic.

UQ: Assessing Language Models on Unsolved Questions A careful examination of large language model performance on grade school arithmetic

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.529186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.872281Z digest=sha256:8865dad5bf0797e4ad13c259ffabfb9a2bbdb23cc4537a4ed14664b0dc155721

Observation 7adaf2fb-c55c-4003-be7f-b533fa00d918 · outbound

This paper cites Challenges in Trustworthy Human Evaluation of Chatbots.

UQ: Assessing Language Models on Unsolved Questions Challenges in Trustworthy Human Evaluation of Chatbots

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.876480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.876480Z digest=sha256:890cb7ca876df82091c969bc960b6074c27d0086a3a6ee9eb59bbbd539ccd6bf

Observation 2fe1d022-d3b2-45cf-8ff6-2d05aac7e3ba · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

UQ: Assessing Language Models on Unsolved Questions Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.880787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.880787Z digest=sha256:1a8296a2492a8b01980db017278f45b78a2878d265d9bce91ea0adbb7d62de0d

Observation 5dde716a-4b75-4d50-b822-de3b1a0e9cca · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

UQ: Assessing Language Models on Unsolved Questions AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.885288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.885288Z digest=sha256:f8d6e805f3c90966f1cf894c657358ee874523c0e6f3e8aea0dc99fadb010dcd

Observation 73bd8a55-7730-41ab-a739-ca47bfd026d9 · outbound

This paper cites LIMA: Less Is More for Alignment.

UQ: Assessing Language Models on Unsolved Questions LIMA: Less Is More for Alignment

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.889330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.889330Z digest=sha256:896139eacd5d8846b596f414170e52150fb1403be7db9b20c011da579c4483c5

Observation c801d62d-bb02-4a36-ab60-326a1d65e71a · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

UQ: Assessing Language Models on Unsolved Questions Reinforcing General Reasoning without Verifiers

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.893401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.893401Z digest=sha256:fb81abba2333d6d09d347bc6d77786b54f77461c96243d928969cf634233b805

Observation 1adc20b5-9a0a-4593-926a-6c1846e92381 · outbound

This paper cites Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2025.

UQ: Assessing Language Models on Unsolved Questions Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2025

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.516012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.897165Z digest=sha256:f0130c6b3f2f37de0baf769a7003d8f2dac0078b9749a4d3418a1f877f3c61f4

Pith citing papers

Observation 51e8effd-fb12-4824-8bf6-19061be8cfac · inbound

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval cites this paper.

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval UQ: Assessing Language Models on Unsolved Questions

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:06:04.097499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T04:09:46.616019Z digest=sha256:02ae2f945655042b00737af3e1eb40cef5612e52ab593320c1bdcfe4195d41dc