Pith. sign in

Paper Citation Record · LEDGER

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

As of 21 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 20 inbound Pith citation observations for arXiv:2505.02018.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02018 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:08:37.211593Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:51:33.982308Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T09:37:00.794953Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 901ca0ab-2ceb-467d-b344-f9275ba63166 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.630151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.630151Z digest=sha256:bab68e17042079e68821870404c295b1c67c46da137b489a8d485a5e1d191a20

Observation 69ff9991-5074-4f4f-9977-3a4edb842d62 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.639033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.639033Z digest=sha256:489a8fdd39e587d821aad06b01dc8cf22283dc816f558fc2e8f1c00309629782

Observation 046532ea-32f6-4276-881e-247740734ad6 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.655360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.655360Z digest=sha256:633f63a5211c987d022ce59d857b04c60f2b159ea5358ab3f99f82aec49500cd

Observation 2119914a-948e-4405-8289-f2512103dfc1 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.731861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.731861Z digest=sha256:1af3705eacaff860cd25322be714fd99c50a6ec6d26712fdb532616fcc7968cf

Observation d1a32250-9098-4991-9df7-e3a447c5d0bd · outbound

This paper cites A Survey on In-context Learning.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation A Survey on In-context Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.822069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.822069Z digest=sha256:9ca6a1f9aa327c368129185f414ff86ab6f71fd860be558fb322e1907292b4af

Observation 16af64ba-4781-4d3f-b6ab-2f51c478a3df · outbound

This paper cites The Llama 3 Herd of Models.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.827017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.827017Z digest=sha256:2ed2f83af878f9ceb0a198a12dcc4540dcf86e5596ec5dd40fd0f04222671670

Observation b045eb29-5e9b-494e-b4e9-0bf7a7b67910 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.831096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.831096Z digest=sha256:2029354e827414b86dbffc212229a27dc07846bbafd2f8bbe7fdcca53b64ca43

Observation 896e554c-d5ea-4019-bd45-278dbc884eea · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.835479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.835479Z digest=sha256:b5e5ae315df8a3610bd074f01a33aa23b9d458afe42a95d03e069d782122495e

Observation 6bd53163-92e2-4b33-b976-518b9fbc667c · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.839708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.839708Z digest=sha256:6bc80d9e70026fbfba3c439a792ffc152dfbabddad3ef3d4a4e8214f94f2e4a3

Observation 87d74862-8704-4ecb-b646-251e24d0e306 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.844435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.844435Z digest=sha256:70930c777c9a3c6d0d1ef29a0e79a87c763cbb8b645d13359b092308212d45b8

Observation adbe2ce2-92a8-4fab-a393-dd4cbd880767 · outbound

This paper cites Mistral 7B.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.848362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.848362Z digest=sha256:cc2c28de32f9d06c8b43718fccbf820a6824b0213d75be1ec203b02344fc1310

Observation 890e79c4-5704-42ae-b873-ebda6e4631aa · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.854952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.854952Z digest=sha256:6df296843775e6994881f3051212744d9e3b1a5162d2782f66a9f8b05a7b245d

Observation 41852214-afa7-4f26-8aa4-cc9d7a1835d5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.907535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.907535Z digest=sha256:71f3f6d869e02d121a626d58fb5429aeac54c16c806c6977160e20762594ea7c

Observation 1ba43c20-15dd-4169-adf9-9fd75bfeaa68 · outbound

This paper cites Let's Verify Step by Step.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Let's Verify Step by Step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.929662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.929662Z digest=sha256:36a1630185b241ca34494493dea0a9fe34876b517998193c5b5937e26061f028

Observation 3d59a197-b110-49bc-99c4-00097571d469 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.974190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.974190Z digest=sha256:1d492877d2c726ff26a9f85e7d4e2e5c1dcdbe08da18c60c3d36a167931dc304

Observation 5902fe21-f805-4922-9787-8c9d38a2ae0d · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.978317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.978317Z digest=sha256:7affe0cfa2f3d4f7da678e8cb0cef406ba880b931a3e832096ee47d3efbddf60

Observation 23adfb42-5f99-4050-a933-43ce982bb974 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.982335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.982335Z digest=sha256:376c9ce2c870de47c758919b5081d9bcf0a353670a071260fd05fb5a813b56d3

Observation 65169e20-b180-47a5-8c92-4855c2473aa7 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.991973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.991973Z digest=sha256:99cb78c752852cf5398f84bb864ac7048164bb63cc3ac21912ec08b854bf61da

Observation dd68f00d-aeee-4e07-8292-06ac04f5ad16 · outbound

This paper cites Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.996085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.996085Z digest=sha256:4daf06667d5ce993c3bb0caa906463dd43d237dd11a2272490f681ae081fe0a1

Observation f3688aef-2b7e-4890-847d-4a2b8e816409 · outbound

This paper cites Proximal Policy Optimization Algorithms.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Proximal Policy Optimization Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:37.003737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:37.003737Z digest=sha256:817932c4b77dd0b103171e975c5a4a1b340a304613785b35bf25613d9ac7e1f8

Observation 4cfd9fce-d256-472f-833f-b20eb37d000f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Gemini: A Family of Highly Capable Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:37.007909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:37.007909Z digest=sha256:a6316892db5b5d34ae215cd0fd5b0cf945726de294f4a2748aef8ab5a871cf5b

Observation 225db3c6-fa30-4ec1-ad14-cf98183f05d9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:37.012124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:37.012124Z digest=sha256:f38882bd8c432f653206b23dfa100d03e4c361db94a0cf4522772cda3911235c

Observation 6ef78f6a-82d7-42e3-a7c8-b737338c977f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation LLaMA: Open and Efficient Foundation Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:37.031827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:37.031827Z digest=sha256:f053f9f778f6a3cf439062188b0a2cee7006fd16f2916f41ee164a74ab633025

Observation 68e72334-2933-4ecd-a867-758b6510045c · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:37.117285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:37.117285Z digest=sha256:ef0fdeb8ad9edee5a2717f8dd8c11aa55506f9df0c9bbfec6a3cf5d40a78346c

Observation b90e2e02-5131-4f14-8a4f-884f98fc611e · outbound

This paper cites Qwen2.5 Technical Report.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:37.173284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:37.173284Z digest=sha256:430861da3898469b06cb7bd8f376ca87fdaaf7ecf6573258d59507dc377a3688

Observation fa27026e-124e-4985-b15a-adda4d3d69e5 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:37.201782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:37.201782Z digest=sha256:2d8a70cc44cf03cd6ade46e1ba0badee2372247b350280a22a935217239ccf0d

Observation d4c79360-668d-497f-a836-551a7e4c272d · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:37.206420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:37.206420Z digest=sha256:b0c573468d301d1d9a55030dee14a6f6d9d1a962e5c9ae8061a209f8c0ed634d

Observation 6a27e704-5c6a-4d5a-98a7-d507c32dd65a · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:37.211593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:37.211593Z digest=sha256:ba2b31c8ec12b18a3f6cf80ebc1fd0e99091d2b5c553032a129aee87c58571c9

Observation 058a52be-0676-4ed5-9fd2-638bdbc36c49 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:37.054900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:37.054900Z digest=sha256:befbb379889ef02c710fe61a4b97de6f0e6a2099aa554b84e9d1d127a57f0b61

Observation 07b58d27-633f-4c41-bc19-d45582ea3b95 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.999701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.999701Z digest=sha256:9d2b4b7cd608dd0e6241dd065ad78babe80c1b01b5f8ea245ec0ba21ebacfff8

Observation 24dbde0d-2409-46a6-9159-ca2cb92aa486 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.679530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.679530Z digest=sha256:55aece73f0227387f47d1d43ef463483376deaed6c276bdae2fb65a434f0e24f

Observation f8806b74-2c98-4f0a-988b-7cefe9485945 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.987084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.987084Z digest=sha256:9e0a03df3c97fa30fc5889b0c9a87a45b47ab729705f30624380095f3cbe9405

Observation ec6373ea-254b-449c-8dd0-4a08b6447722 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.852041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.852041Z digest=sha256:193b2f2f2e2e72eba670298e35b165c436842fdec593f206d51329ba801bdfec

Observation eede72e9-19c7-43c1-ab3e-ed9791cf2413 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T04:08:36.634704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:08:36.634704Z digest=sha256:6e3db2a34a6dc3da3b63ce77c8e51bbe0201f064fbb9333bf3bc5c3ffb10d634

Pith citing papers

Observation 7bb2e622-0bb7-408f-a4e1-f37537f1d806 · inbound

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research cites this paper.

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:33.982308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:33.982308Z digest=sha256:c8aba2f01ba8c8395503880538737dae9efd773f6b825642b640f8e3aa74f62d

Observation a3b2dc24-843c-46a1-9923-1f07e03446ec · inbound

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs cites this paper.

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:42.512437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:42.512437Z digest=sha256:ffe35ff0ee4007d9a67b56fe2e15df12bb93858584b0252f4c89bb7b0febc16c

Observation 4380e4f1-ea7e-40b0-b55e-c85b8b215dc0 · inbound

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems cites this paper.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.587173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.587173Z digest=sha256:d962844c7b4976d8163ca0cf0ec58855978bc42e3fff4deb13a6542b9ff33b1c

Observation 8bfe9a92-6741-43ec-9626-7cc730452232 · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.710933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.710933Z digest=sha256:2bdc834802f0ee51ebf81bb3d37b997cd4f2affa68787f86e7b9665816c88f6c

Observation e2498c15-1ae6-4456-91ac-4b9828478eed · inbound

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models cites this paper.

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T18:53:05.031094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:53:05.031094Z digest=sha256:e7c92a1f3287cb6f891510f57f105c8de9bf12ddc63c9919e2a660cc0b0aaa2d

Observation b0663a81-39ac-4ee8-b7e8-d3f351d37055 · inbound

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents cites this paper.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.234154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.234154Z digest=sha256:2314be2bd195239e6145c30aac99ab98faed54d41eea957a4a4eb00e319b4feb

Observation a84e9cb5-7ffb-4bf3-afd3-a6f9e5cf63c3 · inbound

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents cites this paper.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.420473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.420473Z digest=sha256:b6e0195c5d0a73434102fa51ed740b7eee8f4660557d4f095a84b9f65b9c9903

Observation 64549480-e240-4c10-8fa0-a9dfccbb5966 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:16:34.617724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:b2ed76426ea4d0fcf93d4f86ff7b17bb33080c0d3c5ad36b63b3d66be93c648f

Observation 797a0210-c121-42a9-957d-0453b00c769d · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.093227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:f24dcb54be2d00a9c1c9d0461b9a045a2ac56ff66f0a98fe5084b0938ca7456c

Observation 7e63eb40-1a4e-4c5d-901b-a879d875aed6 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:05.988418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:05.988418Z digest=sha256:90b6ef1d1e799adffd8e58caa7ca84c72deb8168c49c11073716bbb811f3c9f3

Observation 5bc41635-32bd-45f0-a7d2-1591273abba0 · inbound

IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video cites this paper.

IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T23:47:18.378887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:47:18.378887Z digest=sha256:bd9e0024bed44aa6b4f55d25cce643bdb0305b3bc631d37612419be75aa8c38d

Observation 58be337b-c331-4d8a-9e11-ca6c1b02246e · inbound

ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control cites this paper.

ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:30.430849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T06:24:47.014162Z digest=sha256:e95727f3bf0f392c991a7fc4dac94a465006cdbb57e1b1d1fb920fa63f0cbee0

Observation 5463af0c-4957-439c-86c4-2a873252e8f6 · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.709984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:c4570ef01621e7e3964acb404a940e6a214357ec824a75dc11d0365e7c2ae71e

Observation b397e0b1-1b2e-4a2a-bbc7-90c5013a1e00 · inbound

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? cites this paper.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.590378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:f229fde4a499e57b823980c82dd38a45914b9fe599c12ee23966b0b597b1b5f9

Observation 4f51371a-7990-4445-871a-7f81ef2b07f0 · inbound

SFBench: The SciFy Scientific Feasibility Benchmark cites this paper.

SFBench: The SciFy Scientific Feasibility Benchmark R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:04:21.322289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T06:59:19.857066Z digest=sha256:5da0d521f1e26c7d0313b1b246a09a9bfe8638a9818685db74ef660e08dcd472

Observation 1bac3f50-ff3e-4870-b37c-ed52de082641 · inbound

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models cites this paper.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.796446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:5f2fa4f0b3e13e27b4251e0817f7da6f3dfcabbbedda68eabc9de90d835985ea

Observation 0c812abd-0d67-434e-afb2-cdd9590282e0 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-14T11:49:31.595873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:49:31.595873Z digest=sha256:d489ff2c833a253302f5807023a197a425e514159c6975b56ae2b62e74bb8e4a

Observation 9f756069-7f36-4809-bd2d-a468123d613e · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:38.710219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:38.710219Z digest=sha256:7386db98c55fe595c26c94eca8d26e06cd09fa6df1ff0f8e2ec5eaad09d5ec93

Observation 96a55390-e8ef-4e21-a9a3-17ef0e834002 · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T06:13:54.686646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:13:54.686646Z digest=sha256:6d3962bb39706456a31517c9e7459953d48227edbdd90a64e210014ea3bdcf14

Observation e38f4f09-afe9-4b91-9cad-06f220e0180c · inbound

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning cites this paper.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.422268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.422268Z digest=sha256:4001e7d2ff795a495a7c83704a3a6bdde4d0a77b44caadb8679bccc0da8ff970