Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T15:29:29.577032Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2607.18438.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T15:29:29.577032Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 199fe0a8-c4c4-4def-b0fc-d153adc8b48b · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains ARC Prize Leaderboard
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b12e11-c705-42d5-8d7e-4dd29783267c · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains τ 2-Bench Telecom Benchmark Leaderboard
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65cf5065-a338-470a-ab92-4fda93111e5c · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Humanity’s Last Exam Benchmark Leader- board
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5941784-cf22-451b-b93c-d91dc20f6a0d · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Introducing Claude Opus 4.7
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92a0fe3-69d0-4dbc-b7df-5392c0d50925 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Claude Mythos Preview System Card
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c37784-cc22-4005-b93c-d405ec4e3649 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25c5f1eb-6ec9-490e-8c4c-ecdc93a38316 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Introducing GPT-5.5
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e1502d-9276-42a8-bfd5-614bd3e637c9 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Grok 4.3 Beta
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb367314-54ce-45e8-a229-cb20f3c374f6 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Qwen3.7: The Agent Frontier
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2f134a-7b0c-48a9-8009-f1286765e82f · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d822aa-4ddb-4df3-bd1e-ec4c29589558 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf52d656-99a6-493f-ad55-56598de2dd88 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AgentFrontier: Expanding the Capabil- ity Frontier of LLM Agents with ZPD-Guided Data Synthesis
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be277fbd-77a6-48c9-8b0c-12d907e8234a · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Measuring Agents in Production
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34aec3da-e1c8-4667-92ca-5fb377f91ae9 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca1b3420-6791-41f2-bfa6-4db8fbad9a91 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 164f11a7-8477-4b31-afce-87a851f1deff · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1afa2c31-87b0-4c15-858f-f118faad26df · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6019c4f-e0c5-4d44-8dc9-dfb79436910d · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba79076e-70f6-42aa-b0c2-6211e74cc6e6 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains GPQA Diamond Benchmark Leaderboard
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a682bf96-755a-461d-9ae9-0677552ab193 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains MATH-500 Benchmark Leaderboard
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e94e7b-1cd9-4635-942e-8306e0a16a0c · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains https: //artificialanalysis.ai/evaluations/aime- 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61df446c-4935-4572-926a-f55994545daa · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0450d210-6410-4c0d-b99f-b0f43082cc93 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a02249a-32d0-481c-aaa8-e66eae75cdca · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e48715e-bf87-4070-8042-f448aa32867b · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75bdcedc-0a18-4811-958c-30d6e7c4cf80 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0e9443d-03ac-42f0-b36f-7d9d8a8b9786 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc74aea-868e-4afc-9ae7-3903b33d138c · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains GAIA: a benchmark for General AI Assistants
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6721ea98-8a24-4d57-944e-251f92ceba4c · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains APEX-Agents
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df722cae-32da-4264-9865-75bf17af0167 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Are Your LLMs Capable of Stable Reasoning?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e43b92a4-133a-4a15-8c65-8b8bb364b280 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Project Glasswing
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ca0af4-d1a1-45b7-a902-76095d03c743 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Claude Mythos Preview: Anthropic’s Frontier Model Explained
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ef6eb1-c340-4cc9-8b34-a0f704039f99 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Holistic Agent Leader- board: GAIA
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8223312-68e0-4dc1-b75a-1ce7a7689831 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd0ef2fa-538b-4ceb-ab22-42df4ed8a441 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832cdfff-61ee-400e-9c61-dcf0b491a644 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Investigating Data Contamination in Modern Benchmarks for Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2cbc7d4-c6f3-4dc7-9736-64b0026e5fb3 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Anthropic API documentation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb4d9cfe-0542-4d59-81fc-e38bc39c98dd · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Gemini API documentation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f09ac12-5d15-46c7-9323-d81a42746256 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains OpenAI API reference
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9015a179-9740-4a4e-9b9e-e292861079af · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Adaptive thinking
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff719fe2-494c-4888-88d9-75f23ea9945f · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Gemini thinking
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c86265-215c-42b4-98af-cade56917493 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Reasoning models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da1afee8-2254-41c8-a250-dc17ace86eb9 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AA-Omniscience: Evaluating Cross- Domain Knowledge Reliability in Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 686673b7-628e-455a-817a-91a3608e4013 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 924eb265-a4e8-4033-8bd0-665f4c8775e2 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AA-Omniscience: Knowledge and Halluci- nation Benchmark
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d31f90aa-ea5f-4814-b6eb-2506aa063eb5 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Why Language Models Hallucinate
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd7a8fc-2e9e-4786-b250-2b065eb18798 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Claude Fable 5 and Claude Mythos 5
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a38cb3-1ead-4f0f-9e68-19d31dca1107 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Artificial Analysis Intelligence Index
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7420a9-0ec0-4c21-92ab-0a5a112220d2 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 744271cb-af13-4ff9-bc2c-d1eb992b9f24 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Artificial Analysis Intelligence Index: Methodology
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b76a908b-d87a-4bce-a245-d025cda6e789 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AssetOpsBench: Stirrup Agent
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a169a86c-7b9a-4bc1-8c97-bfb852aaa9e5 · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d788fbd-d216-4eab-9658-8c918d5a007b · outbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Humanity's Last Exam
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.