Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T22:31:50.438762Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 22 inbound Pith citation observations for arXiv:2602.16763.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T22:31:50.438762Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T00:44:43.810112Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
13 of 13 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation f3c806a3-43e1-43bd-abe2-0adaa5c2aa63 · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation Measuring Massive Multitask Language Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b58905e-cc33-4c8e-be5d-b61d9360f42b · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40a0510-7c23-4b12-afc3-ef30a00e6805 · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0417d2f9-5519-4c5f-b5bc-6099c4622281 · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b6ef3d5-a115-45dc-9148-6bee456752b1 · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation Instruction-Following Evaluation for Large Language Models
Reference 149
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06aebcd-19f3-4aa5-a0cf-5b7b1c528a6a · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation naacl-long.235/
Reference 235
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caeb7d52-bc16-4fac-a8f0-5664799f36bc · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 349
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5136f8e4-b6ad-4574-bb8c-b19f221d02e0 · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 391
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65a4a9a5-6d7b-4775-9165-ffad68427178 · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation acl-long.744/
Reference 744
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 525c5edd-54fd-4ea3-8071-1cf2c1945008 · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation Solving Quantitative Reasoning Problems with Language Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9185ab9c-7987-4f7e-bc58-a4d9312e495b · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2d6b555-5801-431f-b0a6-01762320e696 · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c9315ac-3235-43f6-9ad2-c51a545d6aa8 · outbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation ISBN 979-8-89176-189-6
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ff38621-70b3-4256-99ff-a37b2003e43a · inbound
CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dd1e38d5-1178-4239-81e7-289640d2962e · inbound
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 86092544-6e25-4ac5-b058-7c02897280f1 · inbound
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 763dee6d-0f07-42ea-8c09-275be3dac9dc · inbound
The Generalized Turing Test: A Foundation for Comparing Intelligence When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 71a81c9b-6d6b-4739-a23d-19f2c11535c1 · inbound
The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5c2ab9c-19c2-4394-b5b6-86a5fc0ecd4f · inbound
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge? When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 68d7a209-0531-400c-8e56-717a8098c3e6 · inbound
Next-Billion AI Index: The compass for AI utility and adoption in the global majority When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 11ff3c2c-543c-4597-97c8-eff206659c08 · inbound
Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b27cd14a-2eab-4087-a0e9-8ae7384aa83b · inbound
MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 768ab1bc-4dd9-4f56-856c-e3874910ec0a · inbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c69b331e-a486-4646-9b76-1644549805b2 · inbound
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f1fe2fe-1b85-4555-beb6-214553e819df · inbound
Life After Benchmark Saturation: A Case Study of CORE-Bench When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b1520a68-7236-423e-be5f-cdc32214bf66 · inbound
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a41c46dd-78f5-4178-a2c8-60645e838e5f · inbound
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0247edde-e7ae-4f33-85a2-4f787075ae99 · inbound
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8e411ce-4fa7-4a8a-99a5-e184655474aa · inbound
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 691aeb6c-63af-4bcb-95c6-a49321aeaea5 · inbound
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e9f13b0-9a8a-424b-9b86-33930f4e895e · inbound
Information-Theoretic Limits of Reliability and Scaling in Language Models When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74817a94-953e-4b35-8b49-f3ecd1261629 · inbound
Rethinking Transfer in Continual Learning: A Replay-Based Realisation When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 280fc8e5-0241-40c8-a4c8-56bde827c6d3 · inbound
Economic Evaluations of Language Models When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8cc822-fd72-44de-a898-d4e60d06251f · inbound
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b0b308-281c-4648-966f-60f6308eeddf · inbound
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.