Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:10.160986Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 30 inbound Pith citation observations for arXiv:2505.14652.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:10.160986Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:52.681204Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:58:32.537541Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8660dee9-4aa0-4f2b-8b3f-ccbb38e9523f · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f4d5688-991c-4ec3-b892-f229d8b214b6 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa17cf45-45ea-48df-8151-bfd00c7d1077 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f54a7f-a643-4f7e-8794-c1dc59c3c474 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36bcc3cf-19b4-4d15-9b86-948cdbcd9e05 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Deepcoder: A fully open-source 14b coder at o3-mini level, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7765d8d-b838-4ad8-aefc-8b5984f9f77a · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b41e973-c2c2-4820-94a8-6df9c61bc57d · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains s1: Simple test-time scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5396bf05-66ad-46fe-889f-81fc88969b09 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains MMLU-pro: A more robust and challenging multi-task language understanding benchmark
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f7ec4993-0070-4ba3-a03a-6d02924c8ddd · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains MAmmoTH2: Scaling instructions from the web
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 75a69ac9-c4f3-4bb9-90b5-f4bdcdbbc664 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a4f703-91d3-4723-845e-7053bd05a2bd · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 752dd351-798e-4dca-884f-00620bd47ab3 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains TheoremQA: A theorem-driven question answering dataset
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7abd6682-5d79-4dbe-8f1b-764533b579b5 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains BIG-Bench Extra Hard
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58e34ca-405b-4694-a0b3-f38922369352 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Measuring mathematical problem solving with the MATH dataset
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feedb8fe-64ce-4896-8fca-d57edbbb2586 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Training Verifiers to Solve Math Word Problems
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ea4c29-168b-4bbf-989f-6909cb26e6cf · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61b49554-a17f-4557-a281-7c0e1044a9b2 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f65fb8a5-15cb-4e2b-99fa-5e41bafe52cb · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20d67b80-af51-4d82-8d60-d8e885ce2f81 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Efficient Test-Time Scaling via Self-Calibration
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ae24153-daf7-4435-ad31-8f4ebbc6e3d8 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains OpenAI o1 System Card
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70039290-fb81-4070-9396-c52d1f418df3 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Qwen2.5 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62152091-6a38-487d-b1c9-8c1bf0e489ed · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains QwQ-32B: Embracing the power of reinforcement learning, March 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b7bfed-5928-4752-af39-907b057f3687 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 301f934e-4a3f-4208-96e2-f87d2b259427 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Training language models to follow instructions with human feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18bf4b9-14a5-4292-a002-0cab5289ebac · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7475d5f0-a5f1-4a9c-8f75-cff476ba9780 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Nemotron-CrossThink: Scaling self-learning beyond math reasoning, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e84209f-bbff-45ee-a5d2-77e813fae459 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec55bc7c-da4d-417c-8666-434d516fb66d · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Gemini: A Family of Highly Capable Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de6e2a22-695a-4287-8156-4b2377720a6e · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4488639-9137-4a1c-919d-2d22188fc6c2 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Solving quantitative reasoning problems with language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e7393aa3-7260-4634-8426-7e27861835c0 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains GPT-4o System Card
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cab01a3-205b-415b-86cd-173a2afd898f · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45cb9444-f20f-40f2-879a-2cba68c92911 · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0476a43f-538a-4ae5-9ce0-ab9aa34d982e · outbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Final Decision: Yes A.5 Detailed Hyper-Parameters We provide the detailed hyperparameters for training our General-Reasoner variants in Table 9
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e8604899-443b-4c3e-8189-b8cef03b172b · inbound
Reinforcing General Reasoning without Verifiers General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c458e86f-bf83-4911-8d45-be01a33a7d17 · inbound
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84d26e8-1849-4665-96d1-d46135f5a881 · inbound
Reinforcement Pre-Training General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78fb8e2-03c2-47d9-a4ab-11b796b2e74a · inbound
One Token to Fool LLM-as-a-Judge General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a004ef3-4a04-48a4-845c-295a6671fe8a · inbound
VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c1d8365-832a-456a-b54f-62345e323905 · inbound
MUR: Momentum Uncertainty guided Reasoning for Large Language Models General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f17ff87-35b6-4bee-b944-dab26b743bfa · inbound
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0242803c-7fac-4bad-bf40-2fb9cad42d41 · inbound
Hermes 4 Technical Report General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3cbd91c-f7c8-41f4-983f-73cbf847eeea · inbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5967dec-a513-4d0c-a844-d051323217f3 · inbound
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f02fdb-4ad0-4d2f-a77c-6f7a85099092 · inbound
MoCo: A One-Stop Shop for Model Collaboration Research General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 64427d66-ef48-4ef3-bb07-287120e20721 · inbound
Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 589
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed0e52c4-2ff1-4323-b54f-08c5820904c1 · inbound
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 797f6862-7acc-4dc6-bcd3-b4badfd13adf · inbound
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7794019f-94f7-48c4-803c-b6858482df27 · inbound
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f87208b-cc9f-464b-a250-c0f76a6b5909 · inbound
TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 96ad7fa4-1ced-4789-8816-d6e831d70999 · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8764e888-71d2-4741-8a44-3cee1f0a77c0 · inbound
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02260b02-af6e-4c34-b6e9-be7e7667b468 · inbound
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b561991d-0d04-4f2b-9389-a584b724ae35 · inbound
How Well Do LLMs Perform on the Simplest Long-Chain Reasoning Tasks: An Empirical Study on the Equivalence Class Problem General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c27bc82-1920-4710-8c03-1da9f3780b3d · inbound
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 153
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c98ad48-f85f-40c2-8df1-1f32771137ad · inbound
M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1da4d8c6-7b3a-4288-9ab6-9c323b7e5c41 · inbound
Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c8bd9bd-d072-436a-9cd3-e510aee9a90f · inbound
Harnessing LLM Agents with Skill Programs General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 817f0464-c865-4d97-bf12-1a6adf7a3269 · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d6bd743e-635f-48b5-a0e2-b434492a0924 · inbound
Trust Region On-Policy Distillation General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 180
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad228ab9-3f72-466f-adb9-f8cd1dd759d0 · inbound
ResMerge: Residual-based Spectral Merging of Large Language Models General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ea3f5d0-a361-44e1-bfef-4275aac5c3cf · inbound
Invariant Gradient Alignment for Robust Reasoning Distillation General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ab7a2902-ecef-4c65-ad83-8059c68c19e7 · inbound
Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e19bfea-0437-473c-8d2b-e9a4c1e01820 · inbound
ArchEval: Measuring AI Agents as Computer Architects General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.