Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

As of 7 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 36 inbound Pith citation observations for arXiv:2506.06122.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06122 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:33.099054Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:16:05.122488Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:45.796847Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab363d8d-7b2f-4975-9a8b-c0d38603374e · outbound

This paper cites https://openai.com/index/introducing-o3-and-o4-mini/, 2024.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library https://openai.com/index/introducing-o3-and-o4-mini/, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.446681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.804104Z digest=sha256:81b457c5df06f582e7d494c102bf06d9fd400e6eb7830edf638228e460047322

Observation 512580ab-0b0b-49a5-af3b-2db8b3a8261e · outbound

This paper cites URL https://huggingface.co/datasets/Maxwell-Jia/AIME_2024.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library URL https://huggingface.co/datasets/Maxwell-Jia/AIME_2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.429316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.811590Z digest=sha256:39afe843a9cb28184439d125545ebccfcc8edb4db5e068401d9cebffafa53646

Observation 1096a416-8100-4e99-963a-1e24ec5b7fe3 · outbound

This paper cites https://qwenlm.github.io/blog/qwq-32b/, 2025.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library https://qwenlm.github.io/blog/qwq-32b/, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.414382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.815959Z digest=sha256:4d1e3cb407c997fa363321e09db33e1f8806c25b6b94c974fdbd35be916679c0

Observation 63217e89-5d4d-4da0-80ae-13cef1e56627 · outbound

This paper cites LMRL gym: Benchmarks for multi-turn reinforcement learning with language models, 2025.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library LMRL gym: Benchmarks for multi-turn reinforcement learning with language models, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.400303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.820035Z digest=sha256:0011d46f77bd9a0cd22620f8f5452a2e1b7e50f54e5b20d31baf9f495d2aa95d

Observation eb9c277d-b119-4769-8a15-803007d4add3 · outbound

This paper cites Nemotron-crossthink: Scaling self-learning beyond math reasoning.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Nemotron-crossthink: Scaling self-learning beyond math reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.824360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.824360Z digest=sha256:7c97ba5350a59b8ead3dcab6d21756624ed3d30af57c649bc452ff8ca24cb284

Observation 3c4200e3-014b-46c1-bade-93478d5c1400 · outbound

This paper cites Deep Reinforcement Learning from Policy-Dependent Human Feedback.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Deep Reinforcement Learning from Policy-Dependent Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.828557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.828557Z digest=sha256:b86472c2c8b8c0864a557cbe02d6e6f3a91ef176a55537fcf8930584e233f11c

Observation 2b5dd761-4421-4663-b2f4-bef60f31666f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.834001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.834001Z digest=sha256:f6760b4f607f0950c7a6def6bfe9f6d2cad3d60bbe123f45076d4909701262ab

Observation cca0b717-263d-43b3-a1a4-5fc609de5d36 · outbound

This paper cites Gonzalez, and Ion Stoica.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Gonzalez, and Ion Stoica

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.838170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.838170Z digest=sha256:8ff8b38af4fb7bbdd72458ffced8b6ff2ad19b75b3028c45bfb54cd82e4691a2

Observation f0035524-e9bd-45e4-8d53-2f1331ef07ed · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Training Deep Nets with Sublinear Memory Cost

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.842355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.842355Z digest=sha256:9b9d3484ecaec0c356b33a12e3a1e8d7a7a81c4165febf9c5fa70f98a8d6274f

Observation 793eeeff-3b8a-4ae4-83cb-b2c1da129e97 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.846754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.846754Z digest=sha256:031d8abd2cb5d545c70cae00f3f1ed520279f02bf3d5fe58b58dc539703aa877

Observation 9ff9f6ff-9cde-4b24-8e38-d7d5895e15dd · outbound

This paper cites Multi-Programming Language Sandbox for LLMs.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Multi-Programming Language Sandbox for LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.853271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.853271Z digest=sha256:c8a1288c81469357ef7573876f1fd9a8a47cffa58261bf2d127f4987c9ed401a

Observation 96b3d1f0-3812-48c6-a12e-2b744886c69f · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Group-in-Group Policy Optimization for LLM Agent Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.857820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.857820Z digest=sha256:2218fbe04dd5aeb2eec115ba23f222fc4eb77a9f638175c9fc2ea580a7b818d6

Observation f2d3095f-0f99-4008-ad1c-ad9dbd9f0756 · outbound

This paper cites Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.862408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.862408Z digest=sha256:d53f7b3738e34642817562d0454678e73a44e324736133a4f56de81182690569

Observation 28b6c286-7e98-473b-8f10-a6e470496c92 · outbound

This paper cites NeMo: a toolkit for Conversational AI and Large Language Models , 2025.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library NeMo: a toolkit for Conversational AI and Large Language Models , 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.376089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.867118Z digest=sha256:d8bea7c9d6905888527574ad194fbb0603a22018e3a40aed3427b3dbf683addf

Observation d3e8f4f8-3665-4816-a5f4-b793aeb2c65c · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.870873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.870873Z digest=sha256:459d7e5c446a585051e6308a8233c82c60226faaeb81a747e53e1d064b676f0e

Observation 683195ac-343c-4526-8a60-3fb774b83e18 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.875169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.875169Z digest=sha256:b7cac937120b7bbc60e766c7974b83c3af9b651962a2c4ad0953647e28453334

Observation be8d13dc-eb44-4a95-90e6-5cf38a3a11eb · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.879473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.879473Z digest=sha256:91d61782721562638681fb1fc4c529e38ae20c50064f65f3d210dd9d88873df8

Observation 8651b302-d6ce-4d89-bb60-861847e44f6a · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.884352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.884352Z digest=sha256:4ea6c481eea09c04a63ed09f6ab9b4bcc90f23cf79d46f4f584c9b767631978f

Observation 056139c0-5239-4443-b06f-420c0720588a · outbound

This paper cites Reinforcement Learning Via Practice and Critique Advice.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Reinforcement Learning Via Practice and Critique Advice

Reference 19

Resolution
verified exact
doi, observed 2026-08-07T06:02:33.155236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.889336Z digest=sha256:a58cb305a8990b56f5e61bc2b4de440ba9a5993eec33a58b696521db101f1391

Observation b5f7a72e-29d9-4520-8d3e-4bb2cf6255b0 · outbound

This paper cites Bradley Knox.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Bradley Knox

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.351460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.893785Z digest=sha256:771f9ac827c778565571273d4e365900c1b59db08145e7926dcd7d405d035d91

Observation 90fb3b8c-b6d4-4915-930c-4261edaf2eb8 · outbound

This paper cites Bradley Knox and Peter Stone.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Bradley Knox and Peter Stone

Reference 21

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T06:02:33.696419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.898885Z digest=sha256:15638c517d8e12480ac7b2ad1b05565ae5af5b92a95dece6c761d4f51a593f5a

Observation ad75e3a8-6c0b-4028-a972-b3872e325934 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Gonzalez, Hao Zhang, and Ion Stoica

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.904314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.904314Z digest=sha256:c761e184f2b0363b56860ae4f335882a89be549432512685f114f70c76cd4a0f

Observation 615c330c-4bd4-4a25-b56c-5d755219493d · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.908648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.908648Z digest=sha256:a99343b3ef746daaa165e282e320f42bb12ea8c2543001bb6e8c330e04857096

Observation 8f00ffe9-e3b6-4ce6-a072-b46b04f10ab9 · outbound

This paper cites \ PUZZLE \ : Efficiently aligning large language models through \ Light-Weight \ context switch.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library \ PUZZLE \ : Efficiently aligning large language models through \ Light-Weight \ context switch

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.324242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.913814Z digest=sha256:2b0cf79389c300a3f8422e189e7dbf49c393963d79020f19e2527c38e2749ecc

Observation a51ce3ff-543c-4194-beae-29d86be9df15 · outbound

This paper cites DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.918029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.918029Z digest=sha256:c0461a29ac3ca356cac87796c98b7d41408b95a36aada4aa2338cf3fe8c14c70

Observation 4397b7e9-5a51-4462-97ae-4f0f935b4ab0 · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Sequence Parallelism: Long Sequence Training from System Perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.922834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.922834Z digest=sha256:f2b76edf28f019175d9c36408f58ac23d1898bc951cc34906100d7e7246d9230

Observation a44c0d5f-64d6-4a0a-9fc7-8da28e0b317c · outbound

This paper cites 2 D - DPO : Scaling direct preference optimization with 2-dimensional supervision.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library 2 D - DPO : Scaling direct preference optimization with 2-dimensional supervision

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.310187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.927878Z digest=sha256:a84ccf42e141a31aa4a7c1ffb76930bbfc15f2980fc1037aebbc9565d3debb90

Observation 406a2f89-b656-4de2-99cb-2610d173383f · outbound

This paper cites MMI ference: Accelerating pre-filling for long-context vlms via modality-aware permutation sparse attention.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library MMI ference: Accelerating pre-filling for long-context vlms via modality-aware permutation sparse attention

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.296688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.932348Z digest=sha256:cc296cf0fe213db5b9f49d3a8b19efb7bf5f41eab8618d654a816dbc24b55563

Observation e7b4d3f4-ebb4-4e22-9f10-77a51fd73506 · outbound

This paper cites Remax: A simple, effective, and efficient method for aligning large language models.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Remax: A simple, effective, and efficient method for aligning large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.282458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.936775Z digest=sha256:47fb54dd4935896e98829bd176f4109cdf4cecd33b5ede180541e265f64c595e

Observation a09bc472-20b0-4139-99ed-bb10412582e2 · outbound

This paper cites Agentbench: Evaluating LLM s as agents.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Agentbench: Evaluating LLM s as agents

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.266836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.940884Z digest=sha256:b24adaa1989f3163853733d931344a29d93416d814ac0a561b6b88ece90070e3

Observation 64a4cfb5-9d1c-418b-beef-f05dc3f30fd8 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.945286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.945286Z digest=sha256:163950ee1f8dbea45c403efd539c1c3eafcafd74c310ce71ed6b255404541853

Observation 5806a8e1-d95b-4688-b1ac-4adc6101a954 · outbound

This paper cites Giving advice about preferred actions to reinforcement learners via knowledge-based kernel regression.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Giving advice about preferred actions to reinforcement learners via knowledge-based kernel regression

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.253154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.949346Z digest=sha256:51be8e59813189fbbc5bd87012e0b1c95e7e159663aab6ff5b1cd6e83795b136

Observation fda7b7fa-5ba2-4757-93b0-eeb1cbb426da · outbound

This paper cites ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.953420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.953420Z digest=sha256:2bf7b909a93f795c50437ced06e245cb72d16656f8a446ea0662b7765d090479

Observation defc11c7-f3aa-457b-90db-c9975613fc73 · outbound

This paper cites Deepspeed autotuning.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Deepspeed autotuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.236231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.959359Z digest=sha256:066fc83466c95d466243ca868d3ecd167c6c78fa5edbf14bf061378d1d4db716

Observation 123395c2-3d21-432f-9930-228d98f2d3da · outbound

This paper cites Jordan, and Ion Stoica.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Jordan, and Ion Stoica

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.220809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.965937Z digest=sha256:448660f90e53d039b547f5645b2f77df25bcca26650d8e2816a421dba2361c3f

Observation 3a5065d3-99e7-4fd3-891e-2949c7603e78 · outbound

This paper cites Codeforces Dataset , 2025.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Codeforces Dataset , 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.206003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.970461Z digest=sha256:c78da765519c991766fd8f125a0ce156f151c3f326a26670dc7957626ca26270

Observation a1697028-1a66-4132-9bb1-b72fb554c0e7 · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Training language models to follow instructions with human feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.975022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.975022Z digest=sha256:1424f542065edbf2de969ac280d1912f757ffebc7e78f9b3b5e60b844b02e52c

Observation 27d00742-6fb2-48e7-9c2b-b7117f3999f9 · outbound

This paper cites Training Software Engineering Agents and Verifiers with SWE-Gym.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Training Software Engineering Agents and Verifiers with SWE-Gym

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.979143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.979143Z digest=sha256:6885c7d9ffab0828d7518a1db7730b7206c0f4ebec5b5f6103596d1a1133a2a1

Observation 633b5127-42d9-4779-a1e5-7cdf14c0ff53 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.983611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.983611Z digest=sha256:974ae10e4f97a83a87e84774168ca20ed782f6db21de4ce9470f2a22130269a4

Observation d18e43d9-aaca-4ef0-bb16-eeeeb14d299a · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Zero: Memory optimizations toward training trillion parameter models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.987821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.987821Z digest=sha256:58770237db100bae37d6ae3bf8213d952481db53437996313a4358c203c26a7f

Observation 46f32bd2-0b30-4ab9-a40c-980c200795e6 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.140846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.992239Z digest=sha256:5b28d7291230a8fb6d34efc7b1204a752f4194888cba03443620190342295979

Observation 927c602d-c083-4269-ab55-c88aa426d8c0 · outbound

This paper cites \ Zero-offload \ : Democratizing \ billion-scale \ model training.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library \ Zero-offload \ : Democratizing \ billion-scale \ model training

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.107266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:32.997027Z digest=sha256:b331051f206ff0e886be33cb5ccba882faaf620ba7d63b235e0dada8dc020783

Observation c66bf214-1d17-4cad-882e-9f3369c3361c · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.001110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.001110Z digest=sha256:b71416a7613cb8a7887ce24e657b23d798436e64fc4408a0874278c0167ae8b8

Observation da40d64a-8569-4411-9254-4304e9b63ba6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.005872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.005872Z digest=sha256:3a9f07d222849dd21b70d906092bff568d46a621c13f18900a37334846a2dc7d

Observation b291a80f-b68d-48ab-a4db-a12b72032a51 · outbound

This paper cites Seed-thinking-v1.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Seed-thinking-v1

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.010006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.010006Z digest=sha256:c2eafae467a8b25313aca3c3b434eacb09daa4978ba5ce7a71475798f5deb4e1

Observation 3e957aad-b38a-4f15-8153-d1dce9d3d95e · outbound

This paper cites Sglang: Fast serving framework for large language models.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Sglang: Fast serving framework for large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:34.081743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:33.014019Z digest=sha256:9ba8cb3957ae86f7a876feec4aee7ee16da64f65b68f8be1a973caaac90af356

Observation 370f16b0-649d-43df-b2dd-35ec2da6d49e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.017868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.017868Z digest=sha256:1d431b4993d97f0381a5a747cb82207881e8bc46a470957afb96b382bb7cb079

Observation b23fa2d8-71d9-4169-8408-6b79ee715869 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library HybridFlow: A Flexible and Efficient RLHF Framework

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.021832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.021832Z digest=sha256:ca972a711d6ae0da69374c8bd72effd65a7387c974035d3aae0bac4e13c3aaec

Observation a8bcd306-525c-4606-b0d6-0eb8a9030166 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.026092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.026092Z digest=sha256:e2d2bd9872704d66665f32ef9161bb99632e4448c60ae7181bdeb1369c56fb31

Observation 735ce3c0-d2bb-4b5f-be3c-410c65cf066e · outbound

This paper cites LLM-as-a-Judge & Reward Model: What They Can and Cannot Do.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.030635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.030635Z digest=sha256:9a081eb3425fe9e6a8aeddabf2c3a1f44c5724b1b8803f5b1d8a35808fa04ddf

Observation 4f20c883-72af-4be5-b670-ffd717f77947 · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.035346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.035346Z digest=sha256:1101a62958bc06283a59b4b7fa47f72daba2bf4c04f53f9d228f19b4604a322f

Observation ccacab87-c5a4-4bef-b8c4-7defe3dc7e0e · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.039639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.039639Z digest=sha256:cfa39866fdce4c694d7efa586cd0667e005bc96fc82928ab20fad38e5a995728

Observation de5044b5-9486-4bb7-84fc-3278acc94c7d · outbound

This paper cites Deep TAMER : Interactive Agent Shaping in High-Dimensional State Spaces.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Deep TAMER : Interactive Agent Shaping in High-Dimensional State Spaces

Reference 53

Resolution
verified exact
doi, observed 2026-08-07T06:02:33.140311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:33.043759Z digest=sha256:523f285253583bb279518917195cf169d0d33bf0fbf446d9eab5d0c044c99f24

Observation 93fe47af-8199-4c12-9e38-51a65f955481 · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.047913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.047913Z digest=sha256:1104e3db8962a798e3cfb45bbba7e56c3f10cd03dff261b57adedc568319c96b

Observation 2aa1046d-820a-47b7-93e7-aa2b8b71956e · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.052451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.052451Z digest=sha256:ac27d36d3d51f2d96b48ba950e7845148839522c478c94cb86b2a4ce48935b76

Observation 2efadbfd-edfb-4181-b9f7-132155bc0036 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Star: Bootstrapping reasoning with reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.056426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.056426Z digest=sha256:e9de822d1bd547c506637acc9857fc4e560b3f19c981db258214e0290a64abca

Observation 68bec4cf-22ee-441b-8a54-330e0d72bbc7 · outbound

This paper cites OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.060236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.060236Z digest=sha256:6fb851f753d5dec8da354862884e8e44e040d8b26aba07b16e700ef1d6056349

Observation 6dc93f18-6194-4cba-bd60-729ece3974ec · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.064390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.064390Z digest=sha256:54fd142e9b6e4b779b80f4d3a41ae30171acba90eadfd4514573df1997e3e72d

Observation cdaff5c4-cb20-4e3e-bf71-6da010ba40a2 · outbound

This paper cites Optimizing RLHF Training for Large Language Models with Stage Fusion.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Optimizing RLHF Training for Large Language Models with Stage Fusion

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.069126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.069126Z digest=sha256:1049fef4a1fe92021026cfadbca7f3ed64e31f4f1cd571d54b56089e62e3b0f1

Observation 4f11a428-3c5d-49a9-bcf4-d8d0d4410b22 · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.073265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.073265Z digest=sha256:19a021c631f83b7c6c98a43c3835c9a1cc227fbe9a3527f69ab3612688d8c9b3

Observation 9ac5e8ab-15e7-49d2-96fe-10d5b3a07f41 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.077517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.077517Z digest=sha256:146f8d3064399dde60af9493d53c2a90dc1553d73527b997b1297973c5af4fbd

Observation c54788d8-f16d-4631-8d1d-3d7393188dbe · outbound

This paper cites Ar CH er: Training language model agents via hierarchical multi-turn RL.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Ar CH er: Training language model agents via hierarchical multi-turn RL

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:33.985545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:02:33.081676Z digest=sha256:7894b813affcd45858d2188bd193eb0274ab073a59816ca6a561b144632967cf

Observation feaa7981-0591-492c-af26-a8a78b1dd2c2 · outbound

This paper cites write newline.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library write newline

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.085783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.085783Z digest=sha256:3a1dd17e6b5f36ecb40ed7f637f62a7ae9fe3e88cc337d4d0e8cf195d346acf2

Observation a5292bed-1d2d-4471-b832-a9d43a5904c0 · outbound

This paper cites @esa (Ref.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library @esa (Ref

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.090557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.090557Z digest=sha256:822ccbada619e749625ac8bfc7be96b4ce8a604ee37e9b83d14315637072d52f

Observation 78be3607-5a0c-4183-975a-66aa7cb91b35 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.094738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.094738Z digest=sha256:7932bf922eb40d29a2db78bec6566085e858bb57f2089ae26295d58b30f14065

Observation 0892558c-3855-4e65-96a6-48ba4c59a46c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.099054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.099054Z digest=sha256:55ed646f40ba7f2f5aa1eebea20c007b23ffaa7d43d600ea85592ebb6122eded

Pith citing papers

Observation c1576350-4c9a-4f54-bcae-7fcecb8cd1d3 · inbound

RecGPT Technical Report cites this paper.

RecGPT Technical Report Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:05.122488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:05.122488Z digest=sha256:19350237c7996f35da785a231f8f62ccf4836e5bf28d92941fa469d84802ef8b

Observation 897f391b-d021-49bd-bf3a-e1519725b756 · inbound

Agent Lightning: Train ANY AI Agents with Reinforcement Learning cites this paper.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:55.008086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:55.008086Z digest=sha256:3153566554aef3fed1900dd918dee1a32447d6efd8f0bd73b1d876e9f45b40f8

Observation 301f3916-28be-482d-97d3-d72a57f8f99b · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:42.401654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:42.401654Z digest=sha256:6c72e43b2c7d08b9ea16317316476382d177fe5f47a115010759dd929e4ebcfd

Observation ad7c4fc5-127c-4d41-af37-1e097e952e77 · inbound

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:36:11.236206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T09:33:57.037949Z digest=sha256:35ea726e3c9c5971471eba371640ad44a0faf4bc66df35e2fcb5fe30a0b0ed37

Observation ff131638-77fc-41eb-bd47-63dfda5b131e · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.781696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.781696Z digest=sha256:0d836a5d0efa8a5d41e6df7ecc25a147d9ef3afbc73494e7b9df5d5b253e4bf8

Observation 72950ae2-769d-4d8a-b64a-9ff5eaac8337 · inbound

Learning to Trust: Dynamic Utilization of Retrieval-Augmented Generation for E-commerce Search Relevance cites this paper.

Learning to Trust: Dynamic Utilization of Retrieval-Augmented Generation for E-commerce Search Relevance Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:11:07.151719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T08:06:23.479875Z digest=sha256:e67e67459e1a2c6f701d7b3580645a414baba2cea06e1376aaeae8e85214116e

Observation b00c5174-590a-4584-aaa1-0aed4cf90aa2 · inbound

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization cites this paper.

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T09:48:09.807998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:48:09.807998Z digest=sha256:b78021bcae603f720c7923c5c175d693bc5ffbd92b52858752812fe8f7ca6cd1

Observation 36e7d6b3-f6d2-4b37-bd48-56e3bb42f980 · inbound

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning cites this paper.

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:40:14.531558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:38:30.169363Z digest=sha256:287996a2702ee6fdbb32c09a2d817a80adb952ae5b9ae1665465eaf97a9f38b2

Observation e0190bf3-ac8c-4976-b820-cf6851901b3d · inbound

TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots cites this paper.

TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.134883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:29:39.744103Z digest=sha256:085a3147a9ea1fba42b9206720c8893372a09e5ab9700aa63781414b2413300e

Observation a677679c-d9c2-4e32-9af5-1c0ba159a767 · inbound

EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis cites this paper.

EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:08:04.507306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T16:05:32.355446Z digest=sha256:cdbdf7653ccb40ea9d4d9171b117763c4ecff50c259bd99e48be48ad930a6f72

Observation 1acb3228-e099-4eb7-9fee-cabf3f8bd5a7 · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:53.985066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:53.985066Z digest=sha256:a686f8af2bd1499a5911832a0bfe86459cca2a8a878a4d7d7fa30aabf985a18d

Observation a890e8d4-eb8d-409d-8cd3-39bcc6009c17 · inbound

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning cites this paper.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:07:47.084927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:3b3b389972a54cf5f35ce3ef8d454cf7027828cfa1f44fef9ded39b55517b1d5

Observation 54e0a443-0b89-41c7-81af-a6060d84df1e · inbound

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models cites this paper.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.042955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T14:09:26.842696Z digest=sha256:b44c221858e7c5b394b12caba88f7b64226d8205e13b3034ed1a4a3072d81d70

Observation 1d08131e-bc68-4cac-8b13-e6311204bd50 · inbound

Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks cites this paper.

Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:28:14.103827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:24:17.788697Z digest=sha256:1044108361e50f1cb9e23d0f6e97250272f98885020d571caa3717b2afa4dbe4

Observation 3a2c3f59-b823-4fc0-9825-6df5dccc7e83 · inbound

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale cites this paper.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:03.246392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:ed930bcec663b0a1c23ba095def598a4b03e93fcd3283696fb482f5c3be6bb6c

Observation a2a9e41c-0119-478c-b58c-56dc4825aef2 · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:42:06.689947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:6eb45ffd751fb7634fda1527544525226b862282f4d8e67a11cda1188d4117b3

Observation 99dfa058-3c03-46c6-8b34-8c311883effb · inbound

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning cites this paper.

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:52:14.069720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:27:15.270996Z digest=sha256:4b60ffeaa6c8c8aa3733a2235522c62c41c023b92f9e91cba9b923962419df73

Observation b1a486e4-e265-4b14-9808-e213a9aecbd7 · inbound

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training cites this paper.

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:11.139120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T06:36:33.914193Z digest=sha256:836e9a4ff755bfa8f521846f194d55c8134ad161bfa50623b81ad2c990c28d6f

Observation a8b294fb-cc58-4de5-9569-c9d875a851aa · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:16.305049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:ac6f572b780183395cc56d2e2a8cc9f1f8d65d27b2ff8aeb07617b233460c83e

Observation d2f7b234-e637-4a7a-8c4c-15610da0c84e · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.289906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:94a99a0b1d34678df8cb89d46360b99764049b71890149a353b66928c524043e

Observation 40ed7ce8-f643-4933-b370-36d34c3c9eff · inbound

UserGPT Technical Report cites this paper.

UserGPT Technical Report Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.683912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:10:51.555653Z digest=sha256:0ef8f4429856d83d1602596b009db36e62fa698943006dc2d9bf5520ff425e81

Observation 7a6902da-cc9d-44b4-9d2e-093d8d1a27e0 · inbound

Verifiable Process Rewards for Agentic Reasoning cites this paper.

Verifiable Process Rewards for Agentic Reasoning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:27.373550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:59:44.082011Z digest=sha256:d22de9c11916a7549a953fa478f68785fd9cb21e99e9ed8eb76e599ce55093a3

Observation b1f48d73-be66-4a87-b2f0-3d1e3bde7565 · inbound

Verifiable Process Rewards for Agentic Reasoning cites this paper.

Verifiable Process Rewards for Agentic Reasoning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:45:07.266701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:69eaa3743972e39660304dd33573791f89c18bca61c9609f31c0409bcd0b9225

Observation f3660f3f-d005-46a3-a023-10a31f34b7e2 · inbound

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction cites this paper.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.534275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:6672b0faaae25126ea16a7445512d84f6eec64e88832a2c8df1b8966b5b2e711

Observation 6518ee4a-f0b2-471c-af4f-1923bffdc160 · inbound

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding cites this paper.

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.446849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T06:58:07.996335Z digest=sha256:dbe32df63f9fafe01f45599776745a1a2a69333be087e8960013f88203243809

Observation 9d6eb88f-6a99-4b30-8264-ddb67750a865 · inbound

Libra: Efficient Resource Management for Agentic RL Post-Training cites this paper.

Libra: Efficient Resource Management for Agentic RL Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.171549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:02:00.385932Z digest=sha256:2602f4dab7f4f1f33a503ab1042a715b6159bb8c96d635cabc3f91f693bdbea6

Observation 3c641bf8-1063-4d73-a40c-366c69a3b6fc · inbound

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice cites this paper.

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:46:41.068185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T08:04:17.224987Z digest=sha256:a49caca4b5dbb08d07e3767eb69b38c2c57537ba99af9d357d7b85e8d586c7f1

Observation 9bcb4746-94ca-4dfc-8ebf-c5cdb6f82076 · inbound

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training cites this paper.

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.692270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:06:16.100762Z digest=sha256:448ec6c634e9c994a70ae313b23fa02436ff0ba8fddd45e91dc15c6a12c96e2b

Observation 262e3738-a5e4-427d-94f9-349e4919805e · inbound

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning cites this paper.

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.798796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:44:23.085858Z digest=sha256:6f3c0cf22ff5e03e1637f266387aa24f6f794c7df164c19e4b139c7327b60daf

Observation 40aadc8f-ea47-46b9-b722-25100d56df1f · inbound

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning cites this paper.

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:20.912894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:23:19.403334Z digest=sha256:84c5757835280fe207494b90ce2734b2dcaee37605a0322169ebe48988b6a847

Observation fe389f63-900c-4705-ab04-2215ee34ba30 · inbound

ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping cites this paper.

ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.883444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T03:56:46.899707Z digest=sha256:5b4007caf71d2de3a1270c3f669814d3e5349cc628e744ad73fe59ea85235362

Observation 8fde50f1-ed10-4e5f-97cc-876f04a51ffb · inbound

ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping cites this paper.

ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T09:19:06.709345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:19:06.709345Z digest=sha256:a87e2e760c2dec30dea074956e802b68426ed5c8590d77ec551ce7b1601c4059

Observation cc362501-73a8-4172-b17c-eddfad69ac7d · inbound

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis cites this paper.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:60cee1820d160ce0a7d08826ead3cbf19cd33b7d1c8c2c718880671ef49a611b

Observation 4f69aa17-7c57-4d73-a9b0-1eade8a3b828 · inbound

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training cites this paper.

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T04:42:35.589143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:42:35.589143Z digest=sha256:8ff73ac441bb82b9df443703f98d53edc23d15cf102b06629dde1ce9f27fa61d

Observation 8dcd30dd-5f01-4957-979c-a3a7d152910c · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.223313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.223313Z digest=sha256:1328432e204baa4cf4a48e546375977ed630177964e83795519b6c512cc65726

Observation a0af7cf3-3f4c-4b46-97c8-74e1d9486267 · inbound

Zing: Social Mind for LLMs cites this paper.

Zing: Social Mind for LLMs Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-30T14:24:27.787441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T14:24:27.787441Z digest=sha256:8c5e9ae40935f51c3ccb8a618a699fd124557ed92be708c871ed8230864773e5