Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:33.099054Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 36 inbound Pith citation observations for arXiv:2506.06122.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:33.099054Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:16:05.122488Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:29:45.796847Z
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ab363d8d-7b2f-4975-9a8b-c0d38603374e · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library https://openai.com/index/introducing-o3-and-o4-mini/, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 512580ab-0b0b-49a5-af3b-2db8b3a8261e · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library URL https://huggingface.co/datasets/Maxwell-Jia/AIME_2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1096a416-8100-4e99-963a-1e24ec5b7fe3 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library https://qwenlm.github.io/blog/qwq-32b/, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63217e89-5d4d-4da0-80ae-13cef1e56627 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library LMRL gym: Benchmarks for multi-turn reinforcement learning with language models, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb9c277d-b119-4769-8a15-803007d4add3 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Nemotron-crossthink: Scaling self-learning beyond math reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c4200e3-014b-46c1-bade-93478d5c1400 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Deep Reinforcement Learning from Policy-Dependent Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b5dd761-4421-4663-b2f4-bef60f31666f · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca0b717-263d-43b3-a1a4-5fc609de5d36 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Gonzalez, and Ion Stoica
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0035524-e9bd-45e4-8d53-2f1331ef07ed · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Training Deep Nets with Sublinear Memory Cost
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793eeeff-3b8a-4ae4-83cb-b2c1da129e97 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ff9f6ff-9cde-4b24-8e38-d7d5895e15dd · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Multi-Programming Language Sandbox for LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96b3d1f0-3812-48c6-a12e-2b744886c69f · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Group-in-Group Policy Optimization for LLM Agent Training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2d3095f-0f99-4008-ad1c-ad9dbd9f0756 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28b6c286-7e98-473b-8f10-a6e470496c92 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library NeMo: a toolkit for Conversational AI and Large Language Models , 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3e8f4f8-3665-4816-a5f4-b793aeb2c65c · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683195ac-343c-4526-8a60-3fb774b83e18 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8d13dc-eb44-4a95-90e6-5cf38a3a11eb · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8651b302-d6ce-4d89-bb60-861847e44f6a · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Gpipe: Efficient training of giant neural networks using pipeline parallelism
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 056139c0-5239-4443-b06f-420c0720588a · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Reinforcement Learning Via Practice and Critique Advice
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5f7a72e-29d9-4520-8d3e-4bb2cf6255b0 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Bradley Knox
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90fb3b8c-b6d4-4915-930c-4261edaf2eb8 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Bradley Knox and Peter Stone
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad75e3a8-6c0b-4028-a972-b3872e325934 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Gonzalez, Hao Zhang, and Ion Stoica
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615c330c-4bd4-4a25-b56c-5d755219493d · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f00ffe9-e3b6-4ce6-a072-b46b04f10ab9 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library \ PUZZLE \ : Efficiently aligning large language models through \ Light-Weight \ context switch
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a51ce3ff-543c-4194-beae-29d86be9df15 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4397b7e9-5a51-4462-97ae-4f0f935b4ab0 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Sequence Parallelism: Long Sequence Training from System Perspective
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a44c0d5f-64d6-4a0a-9fc7-8da28e0b317c · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library 2 D - DPO : Scaling direct preference optimization with 2-dimensional supervision
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 406a2f89-b656-4de2-99cb-2610d173383f · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library MMI ference: Accelerating pre-filling for long-context vlms via modality-aware permutation sparse attention
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7b4d3f4-ebb4-4e22-9f10-77a51fd73506 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Remax: A simple, effective, and efficient method for aligning large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a09bc472-20b0-4139-99ed-bb10412582e2 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Agentbench: Evaluating LLM s as agents
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64a4cfb5-9d1c-418b-beef-f05dc3f30fd8 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5806a8e1-d95b-4688-b1ac-4adc6101a954 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Giving advice about preferred actions to reinforcement learners via knowledge-based kernel regression
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fda7b7fa-5ba2-4757-93b0-eeb1cbb426da · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation defc11c7-f3aa-457b-90db-c9975613fc73 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Deepspeed autotuning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 123395c2-3d21-432f-9930-228d98f2d3da · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Jordan, and Ion Stoica
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a5065d3-99e7-4fd3-891e-2949c7603e78 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Codeforces Dataset , 2025
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1697028-1a66-4132-9bb1-b72fb554c0e7 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Training language models to follow instructions with human feedback
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d00742-6fb2-48e7-9c2b-b7117f3999f9 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Training Software Engineering Agents and Verifiers with SWE-Gym
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 633b5127-42d9-4779-a1e5-7cdf14c0ff53 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d18e43d9-aaca-4ef0-bb16-eeeeb14d299a · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Zero: Memory optimizations toward training trillion parameter models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f32bd2-0b30-4ab9-a40c-980c200795e6 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 927c602d-c083-4269-ab55-c88aa426d8c0 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library \ Zero-offload \ : Democratizing \ billion-scale \ model training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c66bf214-1d17-4cad-882e-9f3369c3361c · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da40d64a-8569-4411-9254-4304e9b63ba6 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Proximal Policy Optimization Algorithms
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b291a80f-b68d-48ab-a4db-a12b72032a51 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Seed-thinking-v1
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e957aad-b38a-4f15-8153-d1dce9d3d95e · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Sglang: Fast serving framework for large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 370f16b0-649d-43df-b2dd-35ec2da6d49e · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23fa2d8-71d9-4169-8408-6b79ee715869 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library HybridFlow: A Flexible and Efficient RLHF Framework
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8bcd306-525c-4606-b0d6-0eb8a9030166 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 735ce3c0-d2bb-4b5f-be3c-410c65cf066e · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f20c883-72af-4be5-b670-ffd717f77947 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccacab87-c5a4-4bef-b8c4-7defe3dc7e0e · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de5044b5-9486-4bb7-84fc-3278acc94c7d · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Deep TAMER : Interactive Agent Shaping in High-Dimensional State Spaces
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93fe47af-8199-4c12-9e38-51a65f955481 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa1046d-820a-47b7-93e7-aa2b8b71956e · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2efadbfd-edfb-4181-b9f7-132155bc0036 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Star: Bootstrapping reasoning with reasoning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68bec4cf-22ee-441b-8a54-330e0d72bbc7 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc93f18-6194-4cba-bd60-729ece3974ec · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdaff5c4-cb20-4e3e-bf71-6da010ba40a2 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Optimizing RLHF Training for Large Language Models with Stage Fusion
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f11a428-3c5d-49a9-bcf4-d8d0d4410b22 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac5e8ab-15e7-49d2-96fe-10d5b3a07f41 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c54788d8-f16d-4631-8d1d-3d7393188dbe · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Ar CH er: Training language model agents via hierarchical multi-turn RL
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation feaa7981-0591-492c-af26-a8a78b1dd2c2 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library write newline
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5292bed-1d2d-4471-b832-a9d43a5904c0 · outbound
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78be3607-5a0c-4183-975a-66aa7cb91b35 · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0892558c-3855-4e65-96a6-48ba4c59a46c · outbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Unresolved cited work
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1576350-4c9a-4f54-bcae-7fcecb8cd1d3 · inbound
RecGPT Technical Report Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 897f391b-d021-49bd-bf3a-e1519725b756 · inbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 301f3916-28be-482d-97d3-d72a57f8f99b · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 179
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad7c4fc5-127c-4d41-af37-1e097e952e77 · inbound
SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff131638-77fc-41eb-bd47-63dfda5b131e · inbound
TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72950ae2-769d-4d8a-b64a-9ff5eaac8337 · inbound
Learning to Trust: Dynamic Utilization of Retrieval-Augmented Generation for E-commerce Search Relevance Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b00c5174-590a-4584-aaa1-0aed4cf90aa2 · inbound
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36e7d6b3-f6d2-4b37-bd48-56e3bb42f980 · inbound
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0190bf3-ac8c-4976-b820-cf6851901b3d · inbound
TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a677679c-d9c2-4e32-9af5-1c0ba159a767 · inbound
EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1acb3228-e099-4eb7-9fee-cabf3f8bd5a7 · inbound
StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a890e8d4-eb8d-409d-8cd3-39bcc6009c17 · inbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54e0a443-0b89-41c7-81af-a6060d84df1e · inbound
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d08131e-bc68-4cac-8b13-e6311204bd50 · inbound
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a2c3f59-b823-4fc0-9825-6df5dccc7e83 · inbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2a9e41c-0119-478c-b58c-56dc4825aef2 · inbound
EasyVideoR1: Easier RL for Video Understanding Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99dfa058-3c03-46c6-8b34-8c311883effb · inbound
Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1a486e4-e265-4b14-9808-e213a9aecbd7 · inbound
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8b294fb-cc58-4de5-9569-c9d875a851aa · inbound
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2f7b234-e637-4a7a-8c4c-15610da0c84e · inbound
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40ed7ce8-f643-4933-b370-36d34c3c9eff · inbound
UserGPT Technical Report Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a6902da-cc9d-44b4-9d2e-093d8d1a27e0 · inbound
Verifiable Process Rewards for Agentic Reasoning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1f48d73-be66-4a87-b2f0-3d1e3bde7565 · inbound
Verifiable Process Rewards for Agentic Reasoning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3660f3f-d005-46a3-a023-10a31f34b7e2 · inbound
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6518ee4a-f0b2-471c-af4f-1923bffdc160 · inbound
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d6eb88f-6a99-4b30-8264-ddb67750a865 · inbound
Libra: Efficient Resource Management for Agentic RL Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c641bf8-1063-4d73-a40c-366c69a3b6fc · inbound
SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9bcb4746-94ca-4dfc-8ebf-c5cdb6f82076 · inbound
Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 262e3738-a5e4-427d-94f9-349e4919805e · inbound
Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40aadc8f-ea47-46b9-b722-25100d56df1f · inbound
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe389f63-900c-4705-ab04-2215ee34ba30 · inbound
ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fde50f1-ed10-4e5f-97cc-876f04a51ffb · inbound
ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc362501-73a8-4172-b17c-eddfad69ac7d · inbound
Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f69aa17-7c57-4d73-a9b0-1eade8a3b828 · inbound
Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dcd30dd-5f01-4957-979c-a3a7d152910c · inbound
QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0af7cf3-3f4c-4b46-97c8-74e1d9486267 · inbound
Zing: Social Mind for LLMs Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.