Pith. sign in

Paper Citation Record · LEDGER

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection

As of 18 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.12457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12457 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:37:48.272553Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:12:49.428954Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:16:19.329149Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d95cc04d-e0aa-41e9-955b-192db53b1496 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Training Verifiers to Solve Math Word Problems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.171875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.171875Z digest=sha256:d8869644a88105abe98e84233732448103d1ba37919d8d82ead30b43d7e70d9c

Observation c5a7ad9e-50e4-4ca1-9b14-dcd9186fb452 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.176875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.176875Z digest=sha256:39597e59dcd34ef19770c1a2e766e932b8e8eb6e4997e15418399af1e102212c

Observation 1b6d4fa0-e60f-4c01-84bb-bba903eea854 · outbound

This paper cites The Llama 3 Herd of Models.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.180658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.180658Z digest=sha256:e136edf9f45e0acd43d4e247e74c1fa47b574c23412fbcc6593b4dd133dbcf1d

Observation 02521e1c-2ff6-4669-8778-c8bc90c192cb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.184581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.184581Z digest=sha256:3257d8e3712c32de129d2ecfe1441f9694b0aa4314bf46842890a77e68eba972

Observation 92924e94-15bd-458b-815c-a9de3c6b790b · outbound

This paper cites Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.188646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.188646Z digest=sha256:10c46417a4caee7e287240eb2c16bbc0bd686607693278adce3ec95e54cabafb

Observation 92b57636-be46-4f6d-9d6b-f498b408a9e1 · outbound

This paper cites OpenAI o1 System Card.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection OpenAI o1 System Card

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.192475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.192475Z digest=sha256:94d620de9cc7a3e4a3526a1285074e52c5e55af48cb436c8861a550ca26bace8

Observation 6f7742b3-60cd-4201-9607-2a7a28ce9821 · outbound

This paper cites Mistral 7B.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Mistral 7B

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.196537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.196537Z digest=sha256:d93202493b8dc9fcfbb255d43c8d850f0e5a32e3355e1fbf7d4d0a6ace99cdfa

Observation 6a2f10d3-852e-4b8b-9e87-2f55cebfb27c · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Efficient memory management for large language model serving with pagedattention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.200716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.200716Z digest=sha256:0c1dfa71bf30751b3c5e4c4a54ea6f311422b40558f50accd5dc59698cde56a8

Observation e6b93777-e0c5-4553-b9bc-13ac382145f1 · outbound

This paper cites Reasoning Under 1 Billion: Memory-Augmented Reinforcement Learning for Large Language Models.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Reasoning Under 1 Billion: Memory-Augmented Reinforcement Learning for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.204008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.204008Z digest=sha256:2db722fcb619f3616295cdfbc24bd37ea911c2eb2460f19f5fc5b242f7418899

Observation 63cfa6f5-6848-4bea-8218-27a564635c80 · outbound

This paper cites LIMR: Less is More for RL Scaling.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection LIMR: Less is More for RL Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.207868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.207868Z digest=sha256:7cf9099ab06d827e504f3243bec2c8b6d5c9a41c873d0c136737d977e5662c18

Observation 927685ce-a55f-4d23-8b5a-65b4e8c71d1e · outbound

This paper cites Let's Verify Step by Step.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Let's Verify Step by Step

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.211604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.211604Z digest=sha256:de90845be867d8e211b0850794499a1bcf1bf14f96dc0851fff5d08acae0f7d3

Observation a04e5741-1d71-4a98-ad63-2ea995e919f7 · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.215470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.215470Z digest=sha256:adb0a357236d1be1c16086608c27d3f1d2bcb3ff5ca7d14e835cf5d4ff59e6a9

Observation f708a974-955d-4f41-acd2-61da25435205 · outbound

This paper cites Tinyzero.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Tinyzero

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:37:48.535566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:37:48.219278Z digest=sha256:20133c421c912de7bd3c71abbb67fb3c5cc3021d731f5f7ea504a72b16d31915

Observation ebf5aaa6-de94-492e-84a6-b8c3d422c460 · outbound

This paper cites Proximal Policy Optimization Algorithms.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Proximal Policy Optimization Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.222621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.222621Z digest=sha256:16d98ee1c2179a9b10cc04b16b571a416b67a329e2988ff7cfc0980f7ded6185

Observation da097bbe-05df-4389-8c2a-4ae76b6d7e58 · outbound

This paper cites Vygotsky’s zone of proximal development: Instructional implications and teachers’ professional development.English language teaching, 3(4):237–248, 2010.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Vygotsky’s zone of proximal development: Instructional implications and teachers’ professional development.English language teaching, 3(4):237–248, 2010

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:37:48.524889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:37:48.226773Z digest=sha256:5c92f98752d2ac54f29bdd4f2ba04670f7118a125ed39c20bb1fed0e47320ef8

Observation 592f6d01-00d6-4e05-804e-8b1b5fa7809e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.230112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.230112Z digest=sha256:322fcdaddf81ec04b70cf533a13e911cdc7298a95ba0321611ad95de40acbf43

Observation d12433b2-675b-4548-b109-4fd16cf88ee0 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.233491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.233491Z digest=sha256:3e3d3d8b918335fe18cbc04464ff822ba41e9c4932ec96a6dc05a3e52fdabcaf

Observation 90576817-ff7f-49ce-ad8d-3f8dd5d8a6dc · outbound

This paper cites A Survey on Post-training of Large Language Models.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection A Survey on Post-training of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.237252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.237252Z digest=sha256:62c1c7367eb6af1a378088e93ec61554b208b82c690a9d83383a4963fa0d3bca

Observation 987fb47c-4b71-49fe-a94a-b6cade2d47d0 · outbound

This paper cites LESS: Selecting Influential Data for Targeted Instruction Tuning.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection LESS: Selecting Influential Data for Targeted Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.240990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.240990Z digest=sha256:4c79d4724b394381d40a102bba11f29592fcc1bac9ffbd8b298e9230f7f1027b

Observation c3262b2d-cab7-45bc-97a6-a12482c87147 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.244456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.244456Z digest=sha256:8993c23ed53d78e7cf76b9cb5f79a006392c5a9f7af879a631a5bcaf9a81bc98

Observation 3ebd5619-df64-4559-90ac-b53a8910d502 · outbound

This paper cites Qwen2.5 Technical Report.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Qwen2.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.248342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.248342Z digest=sha256:16f6cb9e5b6e97eb444c59f1c0d9bbd38a66bab6a91126711459a3db40a890ae

Observation 37ab6da3-c9ad-4bfc-b41f-65b3ac8190eb · outbound

This paper cites LIMO: Less is More for Reasoning.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection LIMO: Less is More for Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.251846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.251846Z digest=sha256:a18d559b3c05b3ac4e524948408c0228143d84222364e25b22b15724f1e55084

Observation 8afb3c14-f9c5-48f4-ad36-7d9664a3c2a8 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.255513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.255513Z digest=sha256:64cc9356e20c50be3f303aefa90a3b9fea662c4cff79c5e73ea23d41d3ec2435

Observation 8ee0bd85-2670-4060-b2d1-5cf54eab7afe · outbound

This paper cites Supervised Fine-Tuning Achieve Rapid Task Adaption Via Alternating Attention Head Activation Patterns.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Supervised Fine-Tuning Achieve Rapid Task Adaption Via Alternating Attention Head Activation Patterns

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.259068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.259068Z digest=sha256:82c710c4aa9ebdd063c67539ca69e62909e08912850523b955368ff21fa2b2e8

Observation 8ff5a9fa-0fda-40b9-bc36-eb5cab39c37f · outbound

This paper cites Deci- phering the impact of pretraining data on large language models through machine unlearning.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Deci- phering the impact of pretraining data on large language models through machine unlearning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:37:48.513748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:37:48.262517Z digest=sha256:2ed5e0e5533449b24b9317348d13f869383e34c924eea2606349b8e527b81feb

Observation 9ee32d7c-cdd9-41a7-b324-4be260614525 · outbound

This paper cites Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:37:48.318647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:37:48.265885Z digest=sha256:579ac98720d2b2bc551e60702e1858c474df061317221d881992ca76c74b37ba

Observation 4f0bab86-5815-4175-9292-d5404ef1afc1 · outbound

This paper cites Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.269399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.269399Z digest=sha256:fb2c716e9d47903f6aa8c15847f592ee08286a6e989b903632a3104d003112b4

Observation 3d2e64d9-4610-4995-b082-3ee594b24872 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Fine-Tuning Language Models from Human Preferences

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.272553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.272553Z digest=sha256:efa1cb3a941c412368bb8294e39c2e6e155730a72023a8c24d32b8fc6c96b1b3

Pith citing papers

Observation 98825fa1-60da-4e9a-a865-5f7e05120346 · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.330644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:45b1796d9135b865a20624d0ff4bcb47f810c773695abced6dddf8dffadba3e2