Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:08:37.211593Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 20 inbound Pith citation observations for arXiv:2505.02018.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:08:37.211593Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:51:33.982308Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T09:37:00.794953Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 901ca0ab-2ceb-467d-b344-f9275ba63166 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ff9991-5074-4f4f-9977-3a4edb842d62 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046532ea-32f6-4276-881e-247740734ad6 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2119914a-948e-4405-8289-f2512103dfc1 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a32250-9098-4991-9df7-e3a447c5d0bd · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation A Survey on In-context Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16af64ba-4781-4d3f-b6ab-2f51c478a3df · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b045eb29-5e9b-494e-b4e9-0bf7a7b67910 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 896e554c-d5ea-4019-bd45-278dbc884eea · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd53163-92e2-4b33-b976-518b9fbc667c · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d74862-8704-4ecb-b646-251e24d0e306 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adbe2ce2-92a8-4fab-a393-dd4cbd880767 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Mistral 7B
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 890e79c4-5704-42ae-b873-ebda6e4631aa · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41852214-afa7-4f26-8aa4-cc9d7a1835d5 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation LLaVA-OneVision: Easy Visual Task Transfer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ba43c20-15dd-4169-adf9-9fd75bfeaa68 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Let's Verify Step by Step
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d59a197-b110-49bc-99c4-00097571d469 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5902fe21-f805-4922-9787-8c9d38a2ae0d · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23adfb42-5f99-4050-a933-43ce982bb974 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65169e20-b180-47a5-8c92-4855c2473aa7 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd68f00d-aeee-4e07-8292-06ac04f5ad16 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3688aef-2b7e-4890-847d-4a2b8e816409 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Proximal Policy Optimization Algorithms
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cfd9fce-d256-472f-833f-b20eb37d000f · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Gemini: A Family of Highly Capable Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 225db3c6-fa30-4ec1-ad14-cf98183f05d9 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ef78f6a-82d7-42e3-a7c8-b737338c977f · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation LLaMA: Open and Efficient Foundation Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e72334-2933-4ecd-a867-758b6510045c · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b90e2e02-5131-4f14-8a4f-884f98fc611e · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Qwen2.5 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa27026e-124e-4985-b15a-adda4d3d69e5 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c79360-668d-497f-a836-551a7e4c272d · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a27e704-5c6a-4d5a-98a7-d507c32dd65a · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 058a52be-0676-4ed5-9fd2-638bdbc36c49 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b58d27-633f-4c41-bc19-d45582ea3b95 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24dbde0d-2409-46a6-9159-ca2cb92aa486 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8806b74-2c98-4f0a-988b-7cefe9485945 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec6373ea-254b-449c-8dd0-4a08b6447722 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eede72e9-19c7-43c1-ab3e-ed9791cf2413 · outbound
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bb2e622-0bb7-408f-a4e1-f37537f1d806 · inbound
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3b2dc24-843c-46a1-9923-1f07e03446ec · inbound
RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4380e4f1-ea7e-40b0-b55e-c85b8b215dc0 · inbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bfe9a92-6741-43ec-9626-7cc730452232 · inbound
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2498c15-1ae6-4456-91ac-4b9828478eed · inbound
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0663a81-39ac-4ee8-b7e8-d3f351d37055 · inbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a84e9cb5-7ffb-4bf3-afd3-a6f9e5cf63c3 · inbound
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64549480-e240-4c10-8fa0-a9dfccbb5966 · inbound
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 797a0210-c121-42a9-957d-0453b00c769d · inbound
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7e63eb40-1a4e-4c5d-901b-a879d875aed6 · inbound
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bc41635-32bd-45f0-a7d2-1591273abba0 · inbound
IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58be337b-c331-4d8a-9e11-ca6c1b02246e · inbound
ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5463af0c-4957-439c-86c4-2a873252e8f6 · inbound
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b397e0b1-1b2e-4a2a-bbc7-90c5013a1e00 · inbound
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4f51371a-7990-4445-871a-7f81ef2b07f0 · inbound
SFBench: The SciFy Scientific Feasibility Benchmark R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1bac3f50-ff3e-4870-b37c-ed52de082641 · inbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0c812abd-0d67-434e-afb2-cdd9590282e0 · inbound
Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f756069-7f36-4809-bd2d-a468123d613e · inbound
Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a55390-e8ef-4e21-a9a3-17ef0e834002 · inbound
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e38f4f09-afe9-4b91-9cad-06f220e0180c · inbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.