Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:07:11.351063Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 3 inbound Pith citation observations for arXiv:2507.19766.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:07:11.351063Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:35:31.077918Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T08:46:26.428426Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5c735a32-6238-426a-975b-5fabe3121819 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2734aa74-a458-4f4d-93c9-0945d4426584 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities On-Policy RL with Optimal Reward Baseline
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a55ddd6-06f7-4b78-b3f1-ae50f8f80e8f · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Skywork Open Reasoner 1 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629a8ce9-f93a-499d-aa29-f067afaa48d1 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities OpenAI o1 System Card
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06fb2d2f-c07b-4e5a-ab10-7bd775961a99 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities YaRN: Efficient Context Window Extension of Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f1f2719-f220-49e0-9a92-fc578c72971d · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c388865f-7eac-4003-8b93-e5f5eeacb09b · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 659d4409-3b4c-430c-9528-b3798a32d09f · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abcb13e0-4ae3-4c20-9f7a-9f70e88cc368 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf226d8a-bc79-4441-8796-7c44343c35a5 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Qwen3 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df889235-2c8e-4aa3-b8cc-e72b7019f29d · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e386364b-93b7-4e54-9f92-cba37ce75c91 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities The surprising effectiveness of negative reinforcement in llm reasoning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cab312e2-9caa-465c-be13-b432e97a1a24 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Proximal Policy Optimization Algorithms
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 748b368d-2259-4f61-9894-70d0e2434516 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Unresolved cited work
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdc45cd7-dcb4-45d8-b03f-c4e476396495 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4dfdd6c-3709-41fc-b4ef-5d11d911b5c2 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28ec755-8b39-4d14-95f0-12869927dce3 · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities HybridFlow: A Flexible and Efficient RLHF Framework
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e135b1a-7ed4-4bf2-93b1-d2872429679d · outbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Reasoning with Exploration: An Entropy Perspective
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfd12a22-b3f9-4d70-9b11-dfb5301856d3 · inbound
Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa62aa5c-4434-43ff-90ea-d3e69ecc663c · inbound
StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c70c37d-ed93-455f-96a2-7c731d5cbd66 · inbound
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.