Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T05:31:55.864438Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 75 inbound Pith citation observations for arXiv:2601.05242.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T05:31:55.864438Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:41:36.440654Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
46 of 46 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 6db48e9a-15f4-4c36-8b2f-f582bbe38f7d · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 98645355-1915-40ff-b637-5116521dd9aa · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e9050e69-1f99-44df-aa7c-b1d83dad6a32 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Rule based rewards for language model safety.Advances in Neural Information Processing Systems, 37:108877–108901
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 06a78541-2ff2-4350-bb66-443d8c281e9d · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Grpo-care: Consistency- aware reinforcement learning for multimodal reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6a391917-9599-4436-b4e9-948ba9e4c19c · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ed488f19-4333-4581-beb7-800f1b83b99a · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Genderalign: An alignment dataset for mitigating gender bias in large language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 38b085b1-8e2d-4d99-81c2-5aa958f76739 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6f3fef15-f1c2-41a8-b637-901c74538654 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9b95cd5d-4473-4670-aa64-826cd34163cc · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Proximal Policy Optimization Algorithms
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation df9af16d-d487-48f9-928b-8c01186f69a1 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization ToolRL: Reward is All Tool Learning Needs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2409a2b5-9084-43ab-b4ce-af93af0edb0f · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Understanding R1-Zero-Like Training: A Critical Perspective
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e80133ac-8ee3-4a92-8171-df8c1e65c19a · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization ToolACE: Winning the Points of LLM Function Calling
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 242b7515-f9f1-49ed-94c7-38bfeb578881 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Hammer: Robust function-calling for on-device language models via function masking
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 04c557e9-7884-4102-bb50-8ebe3fb154e0 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization xlam: A family of large action models to empower ai agent systems
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cdd56d70-f0f5-4d42-9f29-4cf4218f59ef · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Qwen2.5 technical report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f57370dc-2a75-47a3-b204-03f02f21e24b · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization HybridFlow: A Flexible and Efficient RLHF Framework
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6a920fd8-6ceb-4f85-b353-2bf3c7d5806f · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2f7704fc-3149-410a-81a8-a48fdef83944 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Qwen3 Technical Report
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 033c03fb-2b0d-44ca-8f17-a6b1bf6c59f0 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9f051c1a-1581-4354-82fb-04c0e28b1265 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization American invitational mathematics examination - aime.In American Invitational Mathematics Examination - AIME 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 09c8627f-0971-4f0c-a6fe-bfe27007e94d · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization American invitational mathematics examination - amc.In American Invitational Mathematics Examination - AMC
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f407a6ab-91ef-4afc-89ac-98017634d462 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Measuring Mathematical Problem Solving With the MATH Dataset
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 03b2a86f-642e-497c-84f0-780900f95b91 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Solving quantitative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f8d4a327-e825-4350-965a-d96811a0614f · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a8543c89-6cde-49c1-98c4-33b94bd3428b · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Process Reinforcement through Implicit Rewards
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 51b9b410-57c9-4054-84c8-d173c5f2972f · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Measuring Coding Challenge Competence With APPS
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 42a6f5dd-d969-4846-8c27-afdf53712567 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Competition-level code generation with alphacode.Science, 378(6624):1092–1097
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 69f1b680-8e02-4369-9eeb-fa7e214222ba · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization TACO: Topics in Algorithmic COde generation dataset
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2a2a8774-33e8-46c5-bcb4-130998ac44ad · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5e188b9f-9df5-4567-8c83-33dd08db7df0 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Group Sequence Policy Optimization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9ece9e33-3cad-40c4-86b3-079318eeae4d · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 639e8c14-3980-4fb4-a234-08fec123b9f1 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d2b32c19-8c2a-4eef-8563-03b99e8c74bd · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization DLER: Doing length penalty right – incentivizing more intelligence per token via reinforcement learning.arXiv preprint arXiv:2510.15110
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3a6cb3e7-3953-4541-a511-8a7a39c5ae88 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1368f7b9-ca5c-4869-9a8a-610a1ad1915c · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f004d14b-2ee9-4297-aab5-2fd59480b0cd · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Alarm: Align language models via hierarchical rewards modeling
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e40f7de6-08a5-470b-ae19-82c9688f6ad7 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2fe4f056-f9d2-4b09-aae9-940072415879 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation df8e3d1b-6163-434b-bc98-6f5aefa008f6 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Training language models to reason efficiently
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 623a3f8b-50b5-4736-ba2f-d0a0329aa0fd · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ff693d59-b172-4e5b-a32a-fe37107681d4 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 56fffa5f-907e-49f8-a5be-9506a292801d · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4462bb69-b228-49e2-b8ed-655ed551cd64 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization name”: “Tool name
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation edbdaf2a-ca9b-4571-96ce-326e1a5fd59f · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Provide at least one of <tool_call> or <response>
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d4725689-072e-429a-96c7-dd007c65aacb · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization name” field and a “parameters
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f5d41f42-55fa-4189-8c08-eec6f960f937 · outbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 756b415a-fac6-475c-a928-5c0b93d09dbc · inbound
Real-Time Hardware-Free HIFU Interference Suppression via Teacher-Student Diffusion Framework GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ae3d4ec-5dae-4916-b904-92d217142dcb · inbound
Sharpness-Guided Group Relative Policy Optimization via Probability Shaping GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b765e8bd-e92c-4454-a086-c2147b294638 · inbound
Spatiotemporal Continual Learning for Mobile Edge UAV Networks: Mitigating Catastrophic Forgetting GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 54068fc5-9baa-4763-ae28-8dc667633fcd · inbound
HeaPA: Difficulty-Aware Heap Sampling and On-Policy Query Augmentation for LLM Reinforcement Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e380dba4-12c8-41f6-94d7-c2c62ee88766 · inbound
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2decc221-ac9f-49fc-b4dc-8aa7d8a64757 · inbound
SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ae12815-035b-41c9-8bc9-ecd39067c914 · inbound
SaFRO: Satisfaction-Aware Fusion via Dual-Relative Policy Optimization for Short-Video Search GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ab570eb-3ded-4986-8413-b22c95dddac7 · inbound
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad6fa33-eb62-42db-ae54-8afd890265f4 · inbound
HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31c6e195-efff-4d38-8380-3ebe6d7ef01d · inbound
Target Policy Optimization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 91013a7c-3150-4123-b73f-20a5cc6ad4ec · inbound
EasyVideoR1: Easier RL for Video Understanding GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 419c69a7-8e32-4422-a630-cdcfce22a5d6 · inbound
Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6d3afc50-cc49-44d5-8f5c-60079f14a76b · inbound
LASER: Learning Active Sensing for Continuum Field Reconstruction GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0e7e998a-d4b1-4d97-904a-c8028c083c3b · inbound
LASER: Learning Active Sensing for Continuum Field Reconstruction GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a73cf345-a0b9-4e67-b8c3-fe8e89420450 · inbound
AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1c290574-c2b5-4415-823e-fc33590f1616 · inbound
TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a0df11c5-c6e9-48c4-98c3-9f707f9547c9 · inbound
TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecf9a142-5195-43a7-a4b7-01ca095de9e6 · inbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9fd3a95b-8c5f-49cf-ada2-17da193727e2 · inbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a42b4f71-db54-4435-bbf6-d31c03f57ee6 · inbound
Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e828cad4-e01a-4dd7-871a-04f5f38d6f66 · inbound
RVPO: Risk-Sensitive Alignment via Variance Regularization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 314fdb32-8343-4a63-9f61-48325ea1d40e · inbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c3a0b8f1-76a3-4943-8dda-e27bc9f73d8d · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dec8714c-c459-440e-af63-b17846014260 · inbound
Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b86a72f0-c157-48e4-b015-4bf978b8cf12 · inbound
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 579ec141-6ada-49da-ac25-7be3f7e63d2f · inbound
LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation afd2f9c9-da37-4372-ab6b-b9f41bd8a7ce · inbound
MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a8795d1d-3643-44ca-b74a-161beaf87c63 · inbound
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2532e26e-3f01-4db6-909d-58dce24a829a · inbound
Scaling Retrieval-Augmented Reasoning with Parallel Search and Explicit Merging GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7649c6d4-b12b-4472-934c-3e8720022b87 · inbound
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1f25d252-ab46-4924-a118-18dae3533de2 · inbound
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 332a722f-8f81-487c-af95-163ff3f03bbd · inbound
Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8410e56d-cc67-4eb0-a5a7-d1f1bbebb74f · inbound
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fd05c426-13db-4602-8a4c-27135cbc2e9f · inbound
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eb2202f9-567a-4391-9d3d-71bbc11e88df · inbound
DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aa991f04-88ca-42f9-ad4f-2b0934d6c9e5 · inbound
Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9a425fd3-0c6f-422f-8a3e-f3c3e99c8355 · inbound
CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d750f5b7-4775-485f-8675-09b1042e36e7 · inbound
APE: Agentic Prompt Enhancer for Image Generation and Editing GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c430fc96-1c4a-42ce-980c-145aa79a3982 · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4c25ea9a-6e08-40ba-a01c-ad13555f3673 · inbound
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation acf14a0f-0a73-4baa-aec8-9768f1748718 · inbound
ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d2e61737-9652-43b1-a6b3-f07670e7865f · inbound
It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ee104ebc-bcec-4998-a31f-4cc87155e90e · inbound
It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776c23a1-4bd0-4ab8-8a5f-c5953a6d375d · inbound
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 43be9fb9-a16a-4356-afee-f3671210174d · inbound
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 063e3d13-47bb-489c-9bf6-cfbafbad0576 · inbound
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ad7b8ad7-8902-4870-9053-e2d5accc0660 · inbound
DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0aa6f82d-6ef2-4736-9625-bb22cdbf692e · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b381c90d-44ed-4f2d-874d-647402c74b37 · inbound
ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5e87097d-e921-4943-94b8-b5255ecbeeea · inbound
Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0bbf67cd-b306-4aec-8879-54cb24cae39b · inbound
Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 86d15dec-f82c-43a5-9b4d-b8afaf7f053c · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c6844963-29d2-4bb7-b6cb-c8d0b251f8e9 · inbound
Recommendation as Generation: Unifying Personalized Video Generation and Recommendation at Industrial Scale GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 50c3f134-735c-4433-a934-a630ec25c35a · inbound
Scaling Multi-Reference Image Generation with Dynamic Reward Optimization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b4df026d-1a6d-4169-b402-ea447b561c76 · inbound
Qwen-Image-2.0-RL Technical Report GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 22060474-f3aa-4c8a-af79-5bb78a40b770 · inbound
Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c8456204-8654-4122-b66c-c2548f31fcae · inbound
Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 03ca1633-f4ab-4994-81b7-0608bac1dbe4 · inbound
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c52bfc50-600d-4071-9ee1-3ce319ec6ee1 · inbound
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 067ffb3f-8381-4700-9906-be33b54cbd13 · inbound
TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9e5e5603-4bdb-422c-b1c4-72f21a964894 · inbound
Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 11334e03-e355-4ae6-924e-e86d06c2c4c8 · inbound
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b8049c65-21f1-4c5a-878c-252c59991c2f · inbound
Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e26078-28c1-476d-81af-d8a08b4a3e21 · inbound
The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cf67799-577a-4114-afb6-2495a684fd25 · inbound
Test-Time Scaling via Error Localization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b357547-316a-48f5-a73c-4ec5b56ddc10 · inbound
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b4066c-9a62-4e75-9310-a28ca5601676 · inbound
HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4465a515-2790-4087-a521-d354a93bb977 · inbound
A Formalism-Aware Reward Loop for Handwritten UML-to-PlantUML Generation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e8983d-011f-4b77-9d09-18d24677f944 · inbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bc3d340-d0ad-46f6-886e-05ad422f23b4 · inbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b293e87d-f5db-43d3-bbab-abb4f9fe0c65 · inbound
Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6883944a-b635-4ce0-836c-bbb1f219f393 · inbound
Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c6746d9-388b-4d77-992a-e3c437860635 · inbound
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16f86c57-e7c9-4d71-9c00-ed02c5ebba99 · inbound
Conversational Orchestration for Organic 6G GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bd78f97-a4eb-41da-9db7-a23278644693 · inbound
FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.