Pith. sign in

Paper Citation Record · LEDGER

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

As of 7 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 3 inbound Pith citation observations for arXiv:2507.19766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19766 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:07:11.351063Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:35:31.077918Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:46:26.428426Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c735a32-6238-426a-975b-5fabe3121819 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.543215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.543215Z digest=sha256:e771228a6f90fda35367819e5ed12ee51deef2957f3702bbe50480f605588845

Observation 2734aa74-a458-4f4d-93c9-0945d4426584 · outbound

This paper cites On-Policy RL with Optimal Reward Baseline.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities On-Policy RL with Optimal Reward Baseline

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.581791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.581791Z digest=sha256:885f21ac9e0b5021d428e77aebb855e7797fd81416cde2c5f6c1d9f8b5988179

Observation 0a55ddd6-06f7-4b78-b3f1-ae50f8f80e8f · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Skywork Open Reasoner 1 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.663365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.663365Z digest=sha256:fe292f8c6f44da2c250ddbbdf98d2916e519150ff75952cfbd6c039d9068ac58

Observation 629a8ce9-f93a-499d-aa29-f067afaa48d1 · outbound

This paper cites OpenAI o1 System Card.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities OpenAI o1 System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.716359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.716359Z digest=sha256:5629779fb60eaba114d4dc9f1af16e68e9b36fdf86bbc51c7446087a537fc957

Observation 06fb2d2f-c07b-4e5a-ab10-7bd775961a99 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities YaRN: Efficient Context Window Extension of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.758510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.758510Z digest=sha256:6df5ba576f7221f791c9ebc74e39a0b2d34cf60aca6f16373c387854c0dc7f51

Observation 8f1f2719-f220-49e0-9a92-fc578c72971d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.948538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.948538Z digest=sha256:e9cb30e79aeb8773b5638159b26a1b6806c03ffa1a4e9cd306d236602b8f6a40

Observation c388865f-7eac-4003-8b93-e5f5eeacb09b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.053622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.053622Z digest=sha256:941bc44c93bf6dabd8bdecad481b285bd362ccdfb8e0cd0130e4d5a38a0525e3

Observation 659d4409-3b4c-430c-9528-b3798a32d09f · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.097353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.097353Z digest=sha256:0eee47fdb3e9493013f513dccae2efc7af19d53b3acc414620542df2e6fe59a4

Observation abcb13e0-4ae3-4c20-9f7a-9f70e88cc368 · outbound

This paper cites Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.143364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.143364Z digest=sha256:aca72a389c091dea2171f781414eceeb5df331c6fb82f1cc78750384a36a06bf

Observation cf226d8a-bc79-4441-8796-7c44343c35a5 · outbound

This paper cites Qwen3 Technical Report.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Qwen3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.221814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.221814Z digest=sha256:4b751fd776cb38acac9f64a1f7da2aa8326df6b74e27e66927a34246d065df9c

Observation df889235-2c8e-4aa3-b8cc-e72b7019f29d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.266713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.266713Z digest=sha256:eb7b46bf695ac4a3924ce96beb54da273db99383adf63fa5ace69de83b52238f

Observation e386364b-93b7-4e54-9f92-cba37ce75c91 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities The surprising effectiveness of negative reinforcement in llm reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.351063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.351063Z digest=sha256:bd12c0f2f021eb2b23576169f1b0b1ce19fd9ecce288b0590892403273d849e8

Observation cab312e2-9caa-465c-be13-b432e97a1a24 · outbound

This paper cites Proximal Policy Optimization Algorithms.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.876183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.876183Z digest=sha256:779ba84e18f7e9b5e7ed52157a0e6e8281c2afeefa7ad5815be7488ed10680ff

Observation 748b368d-2259-4f61-9894-70d0e2434516 · outbound

This paper cites an unresolved cited work.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Unresolved cited work

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.904156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.904156Z digest=sha256:5bf8a9f96ab48f05cb6fdfee3dad4ebdaee3c199e09e9b2ea9c3834728ef0ea1

Observation cdc45cd7-dcb4-45d8-b03f-c4e476396495 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.324859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.324859Z digest=sha256:135d928a9afe9bd300e2090976ac9170e77d378608954560d335a575841d1c51

Observation b4dfdd6c-3709-41fc-b4ef-5d11d911b5c2 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.826129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.826129Z digest=sha256:039daeafe02fdc106554cc2d4d73cf6466a847a9c1e15fb7a18728e25447d540

Observation c28ec755-8b39-4d14-95f0-12869927dce3 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities HybridFlow: A Flexible and Efficient RLHF Framework

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.025563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.025563Z digest=sha256:2200e7af8a79cc34fde6b72afa609a2022e2069d07d092196ddb2d0190b0c431

Observation 6e135b1a-7ed4-4bf2-93b1-d2872429679d · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Reasoning with Exploration: An Entropy Perspective

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.485653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.485653Z digest=sha256:0f49f68c81928ad1f761fde52114606c688d94555f4eea0b8aa221a0a784cff1

Pith citing papers

Observation dfd12a22-b3f9-4d70-9b11-dfb5301856d3 · inbound

Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis cites this paper.

Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T04:35:31.077918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:35:31.077918Z digest=sha256:e8985929b1d2b82dc7da0ca29a4363ddede565fa773633bd3a7daa157d1cfd83

Observation aa62aa5c-4434-43ff-90ea-d3e69ecc663c · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:46.677712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:46.677712Z digest=sha256:e5f016c867dec0fa378de2f55acdec37fdfa7b473a59c492d00c27547c231d13

Observation 3c70c37d-ed93-455f-96a2-7c731d5cbd66 · inbound

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training cites this paper.

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.431711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T13:49:26.560459Z digest=sha256:28a28faed190fbefe61b5b74350009b88e8d400a56eb122021735b935fa1c21b