Pith. sign in

Paper Citation Record · LEDGER

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

As of 15 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 74 inbound Pith citation observations for arXiv:2601.05242.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.05242 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T05:31:55.864438Z

measured 120 of 120 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 74 of 74 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:53.701026Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact28
  • verified fuzzy14
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6db48e9a-15f4-4c36-8b2f-f582bbe38f7d · outbound

This paper cites Learn to Reason Efficiently with Adaptive Length-based Reward Shaping.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:55.906712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:7883648b591fa6b9f5f8746afc62a112359798f457152e8fc7d4c8dfcc52d375

Observation 98645355-1915-40ff-b637-5116521dd9aa · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:55.913738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:64eb20316314a15d70bbdba470044b38c931a95c4c0f212307570eeba04f23de

Observation e9050e69-1f99-44df-aa7c-b1d83dad6a32 · outbound

This paper cites Rule based rewards for language model safety.Advances in Neural Information Processing Systems, 37:108877–108901.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Rule based rewards for language model safety.Advances in Neural Information Processing Systems, 37:108877–108901

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.106902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:9a15bc11a361602ec5dd83970aa50dc56c167aadf6bee0c0e1708c663b9c87b0

Observation 06a78541-2ff2-4350-bb66-443d8c281e9d · outbound

This paper cites Grpo-care: Consistency- aware reinforcement learning for multimodal reasoning.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Grpo-care: Consistency- aware reinforcement learning for multimodal reasoning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.111856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:c5c8f67c2e72a4d72f0a4038a7f0d61ead155f28c97ba9669d1bf6967d1abb62

Observation 6a391917-9599-4436-b4e9-948ba9e4c19c · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:56.042488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:90e94345da99bd3929870595481d6101b6dac38af43a4944dc0a3f5dab7d8da6

Observation ed488f19-4333-4581-beb7-800f1b83b99a · outbound

This paper cites Genderalign: An alignment dataset for mitigating gender bias in large language models.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Genderalign: An alignment dataset for mitigating gender bias in large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.122088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:3f58165b11018af4f51aaf8fbae2cb4c919fa64a5d82219721b1baf03cd9d096

Observation 38b085b1-8e2d-4d99-81c2-5aa958f76739 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:56.049954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:4898928ed905826560be2f9aa217bba8c51ca63656e5a453736d321626c86e32

Observation 6f3fef15-f1c2-41a8-b637-901c74538654 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:56.056480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:8c9147bb4d5015ba982cdce24bcf36c275dd354d4138002086b37f4592a7ae73

Observation 9b95cd5d-4473-4670-aa64-826cd34163cc · outbound

This paper cites Proximal Policy Optimization Algorithms.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Proximal Policy Optimization Algorithms

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:56.062395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:fe59dfeb6105d8eba0ebde6cc535eef3023269e41625e1ef6c6eabe402e6d819

Observation df9af16d-d487-48f9-928b-8c01186f69a1 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization ToolRL: Reward is All Tool Learning Needs

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:56.068850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:7b9b42194dfca5f0a274de47c48cd20fe36031238ae61da00e955af70177db10

Observation 2409a2b5-9084-43ab-b4ce-af93af0edb0f · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:56.074559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:4f7493fc2f78f05b3e859dd251c34abbbbeff58e7eca930abc263781396ed90c

Observation e80133ac-8ee3-4a92-8171-df8c1e65c19a · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization ToolACE: Winning the Points of LLM Function Calling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.080671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:48071f347e62bab760abdbbb0eabd84e589818078b20b395e0e07c882c429387

Observation 242b7515-f9f1-49ed-94c7-38bfeb578881 · outbound

This paper cites Hammer: Robust function-calling for on-device language models via function masking.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Hammer: Robust function-calling for on-device language models via function masking

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.088678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:de03e041dd89120378d4f50ebb711847a11ca8af430173c4408de822613cbc2b

Observation 04c557e9-7884-4102-bb50-8ebe3fb154e0 · outbound

This paper cites xlam: A family of large action models to empower ai agent systems.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization xlam: A family of large action models to empower ai agent systems

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.165210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:6e4aa7b70e4ae6c4cd49a776c3ce8acc308d571a12fdd7ef1daa84607c43e976

Observation cdd56d70-f0f5-4d42-9f29-4cf4218f59ef · outbound

This paper cites Qwen2.5 technical report.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Qwen2.5 technical report

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.170294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:9625a1726e9f31e959a60729c2df857543a85be3cb04d6cac155358d2be6e9a5

Observation f57370dc-2a75-47a3-b204-03f02f21e24b · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization HybridFlow: A Flexible and Efficient RLHF Framework

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:56.094169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:6ccfd00dbba391fcdc26404127ed9d6723a001e5f45823366b0ca2061e2ac02b

Observation 6a920fd8-6ceb-4f85-b353-2bf3c7d5806f · outbound

This paper cites The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.179757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:31ccedde5845b9ec54432cca14af98715672abfb7a44c2332465109806bc4a65

Observation 2f7704fc-3149-410a-81a8-a48fdef83944 · outbound

This paper cites Qwen3 Technical Report.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Qwen3 Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:56.100665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:de8354c6025f085514c990b1bd3304e89105ba4e47874b58c9dff392d1bbc89a

Observation 033c03fb-2b0d-44ca-8f17-a6b1bf6c59f0 · outbound

This paper cites an unresolved cited work.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-15T05:31:56.189029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:4e4a57f481c2808c6228c5cf781ff6559c798eac98efecf6794c416edeec688b

Observation 9f051c1a-1581-4354-82fb-04c0e28b1265 · outbound

This paper cites American invitational mathematics examination - aime.In American Invitational Mathematics Examination - AIME 2024.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization American invitational mathematics examination - aime.In American Invitational Mathematics Examination - AIME 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.116947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:bdb5d207ec4146352b70dbb016d2d2f2787bada2d2b72bb15392980d4f3fa73e

Observation 09c8627f-0971-4f0c-a6fe-bfe27007e94d · outbound

This paper cites American invitational mathematics examination - amc.In American Invitational Mathematics Examination - AMC.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization American invitational mathematics examination - amc.In American Invitational Mathematics Examination - AMC

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.126711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:2433e417c26b4dd8b11a64d9db1a91b77360919ae252a1aa70308a79b27a7b61

Observation f407a6ab-91ef-4afc-89ac-98017634d462 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Measuring Mathematical Problem Solving With the MATH Dataset

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:55.920604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:0e59c7a1fe593e9fb6b41318bba49f114540a530dad9dc22d3bc6d194b440212

Observation 03b2a86f-642e-497c-84f0-780900f95b91 · outbound

This paper cites Solving quantitative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Solving quantitative reasoning problems with language models.Advances in neural information processing systems, 35:3843–3857

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.140030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:f2354e9f37f4d000307c1c0f957cc1db4379c5b7e4fc4f8621bc41db9981b2f5

Observation f8d4a327-e825-4350-965a-d96811a0614f · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:55.928839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:542b79256be181d6b60de3a825353c682d6c691b6ca219cf3bf7c80d1b1c3fa4

Observation a8543c89-6cde-49c1-98c4-33b94bd3428b · outbound

This paper cites Process Reinforcement through Implicit Rewards.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Process Reinforcement through Implicit Rewards

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:55.936117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:557d12094d9b8447334631fd28a2c91b3c42efb52e35c0b52cd848356eb6d0b1

Observation 51b9b410-57c9-4054-84c8-d173c5f2972f · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Measuring Coding Challenge Competence With APPS

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:55.942554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:cf083f28b9ec312dc6e18744a036abda74caa5941a5aef2a01108d72d466dd8c

Observation 42a6f5dd-d969-4846-8c27-afdf53712567 · outbound

This paper cites Competition-level code generation with alphacode.Science, 378(6624):1092–1097.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Competition-level code generation with alphacode.Science, 378(6624):1092–1097

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.160213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:c6090ebf049d51f189ea9c550db116400a6187bf4cb06b33179450c567d220c5

Observation 69f1b680-8e02-4369-9eeb-fa7e214222ba · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization TACO: Topics in Algorithmic COde generation dataset

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:55.949971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:7fab236c3443a9a8743274964af90101fda372a5a8744cfb51cf1aade85c5af2

Observation 2a2a8774-33e8-46c5-bcb4-130998ac44ad · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:55.957074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:a7099e961888544322853638dc7c47c499994657b4cc3f7b0b166af04e3efcfd

Observation 5e188b9f-9df5-4567-8c83-33dd08db7df0 · outbound

This paper cites Group Sequence Policy Optimization.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Group Sequence Policy Optimization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:55.963435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:4683eba0e9dd689532ad5c2a19427b06a042a6645574e62c81614b3ebd48d1db

Observation 9ece9e33-3cad-40c4-86b3-079318eeae4d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:55.970172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:dd9f2215697b69087ebc7ffb0507e58f4940ed954d89f1183574eaa1e7aba4c1

Observation 639e8c14-3980-4fb4-a234-08fec123b9f1 · outbound

This paper cites Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:55.977831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:74251396767bef603e8d3ccdd505f907161fa3f93c886a14bda73b322ecdf3c6

Observation d2b32c19-8c2a-4eef-8563-03b99e8c74bd · outbound

This paper cites DLER: Doing length penalty right – incentivizing more intelligence per token via reinforcement learning.arXiv preprint arXiv:2510.15110.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization DLER: Doing length penalty right – incentivizing more intelligence per token via reinforcement learning.arXiv preprint arXiv:2510.15110

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:55.984403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:92a7edfacd20ae41679672870f4e289d6d21df0927609f6126cff5669b1cd17b

Observation 3a6cb3e7-3953-4541-a511-8a7a39c5ae88 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:31:55.991381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:74d471f21276046df65cbc701dea17090c864c8631e71b2f36561c21d81c4838

Observation 1368f7b9-ca5c-4869-9a8a-610a1ad1915c · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:31:55.998279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:514f3c5b03f2adc5598ef7731357cbbd6d4acc760cd4016c9f37394b69806fcb

Observation f004d14b-2ee9-4297-aab5-2fd59480b0cd · outbound

This paper cites Alarm: Align language models via hierarchical rewards modeling.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Alarm: Align language models via hierarchical rewards modeling

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.134533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:faa431643855d29265a4a486d799fbb57ac35503cdacd4c3d30e58791a773546

Observation e40f7de6-08a5-470b-ae19-82c9688f6ad7 · outbound

This paper cites an unresolved cited work.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-15T05:31:56.144847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:4e0bcafd752e8b1b8a83213dd0d78536ccedd29b4edf590e8eeb7944f3235b5f

Observation 2fe4f056-f9d2-4b09-aae9-940072415879 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.006812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:7c86236bbac2f569ffd0c49b3aa4b85a16f81572413a621156b2797ea829eddd

Observation df8e3d1b-6163-434b-bc98-6f5aefa008f6 · outbound

This paper cites Training language models to reason efficiently.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Training language models to reason efficiently

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.014140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:a5dade0f7a965f564167ea175630ee909e138688b99d60ba7d2555119e701793

Observation 623a3f8b-50b5-4736-ba2f-d0a0329aa0fd · outbound

This paper cites Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.021874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:baadd1db5892b626819b9c53e35f8fd30d35d6f50d13cb3018b82e568481bb61

Observation ff693d59-b172-4e5b-a32a-fe37107681d4 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:0fe77aadc2d344925ad1bf18e1147f775c8476803d48d38773bc8bd52d264851

Observation 56fffa5f-907e-49f8-a5be-9506a292801d · outbound

This paper cites Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.036833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:cd72f18d5777b6c49f6f84849c67a8dcf0796bad6d922f491b5f78a5695557ff

Observation 4462bb69-b228-49e2-b8ed-655ed551cd64 · outbound

This paper cites name”: “Tool name.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization name”: “Tool name

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.155044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:8dd440cfba7fdff34fb3ef1b9e38a65b0b314cfdbea6d337d07aef3bdd3101a6

Observation edbdaf2a-ca9b-4571-96ce-326e1a5fd59f · outbound

This paper cites Provide at least one of <tool_call> or <response>.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Provide at least one of <tool_call> or <response>

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.174769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:636d498e6f0f25e87945bcf96cd43a4fdf1a0ed0fde7a6e9fca80fab7582d847

Observation d4725689-072e-429a-96c7-dd007c65aacb · outbound

This paper cites name” field and a “parameters.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization name” field and a “parameters

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T05:31:56.184654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:1066c3b3e3f2797ea976498688e4afac3f152877d1cd9d079eef3c443ed200f5

Observation f5d41f42-55fa-4189-8c08-eec6f960f937 · outbound

This paper cites an unresolved cited work.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Unresolved cited work

Reference 49

Resolution
malformed identifier
raw_fallback, observed 2026-05-15T05:31:56.150053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:1cdc775728ae96b4445e05cbc1a6b531d6b4e2382622e72213085ca8ad981c3d

Pith citing papers

Observation 756b415a-fac6-475c-a928-5c0b93d09dbc · inbound

Real-Time Hardware-Free HIFU Interference Suppression via Teacher-Student Diffusion Framework cites this paper.

Real-Time Hardware-Free HIFU Interference Suppression via Teacher-Student Diffusion Framework GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T12:30:58.223702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:30:58.223702Z digest=sha256:987009ca6fc729df968af2093bb7cbb77fd15dffba0fbe4567f42d50ad6a8aea

Observation 5ae3d4ec-5dae-4916-b904-92d217142dcb · inbound

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping cites this paper.

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T03:12:22.352863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T03:10:57.839146Z digest=sha256:5b640cbc7a4f4fc2c2ce739d5b264e9e5cde62ae9967c3c21724a35fcd453555

Observation b765e8bd-e92c-4454-a086-c2147b294638 · inbound

Spatiotemporal Continual Learning for Mobile Edge UAV Networks: Mitigating Catastrophic Forgetting cites this paper.

Spatiotemporal Continual Learning for Mobile Edge UAV Networks: Mitigating Catastrophic Forgetting GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:40:48.910497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T09:38:25.954499Z digest=sha256:802551247ade2211a4d65ea71a0d51b587f019d0d6fa684e6413ee05b47f8aae

Observation e380dba4-12c8-41f6-94d7-c2c62ee88766 · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.504798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.504798Z digest=sha256:c02ed05a4780afae1384f39da6a1d606f5b1dd52ad90187a10031f2490902430

Observation 2decc221-ac9f-49fc-b4dc-8aa7d8a64757 · inbound

SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems cites this paper.

SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T19:10:40.676299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:10:40.676299Z digest=sha256:54759d26dfad5984b0a77f12f7bee774e6970f610c1480cb0d708d7552796bd5

Observation 6ae12815-035b-41c9-8bc9-ecd39067c914 · inbound

SaFRO: Satisfaction-Aware Fusion via Dual-Relative Policy Optimization for Short-Video Search cites this paper.

SaFRO: Satisfaction-Aware Fusion via Dual-Relative Policy Optimization for Short-Video Search GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:59.197748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:59.197748Z digest=sha256:d91194eb02f2c0c8f941bd7777d590d6722b656af27d14a1501576b5e3745448

Observation 6ab570eb-3ded-4986-8413-b22c95dddac7 · inbound

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration cites this paper.

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T16:50:50.552104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:50:50.552104Z digest=sha256:5b0e0a67e978ca4751818d68c25cd1be94dfe39a52de7d10b898f24fa2a1087f

Observation cad6fa33-eb62-42db-ae54-8afd890265f4 · inbound

HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving cites this paper.

HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T12:58:58.057084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:58:58.057084Z digest=sha256:4d75464fd8b9fbd354488edba6de1fb86229b95db21d800cf04af9986d53e4e2

Observation 31c6e195-efff-4d38-8380-3ebe6d7ef01d · inbound

Target Policy Optimization cites this paper.

Target Policy Optimization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:21:37.744591Z digest=sha256:ead5d03e3d90b12b8d199e5c5fde3c36235e79217db7d385b032ef85883c825a

Observation 91013a7c-3150-4123-b73f-20a5cc6ad4ec · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:207432b922c8ffb41bb12249f21d1472099228f1103d27c9413ac9f12e930918

Observation 419c69a7-8e32-4422-a630-cdcfce22a5d6 · inbound

Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models cites this paper.

Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T03:41:27.425176Z digest=sha256:1b15c2593c360cc1c243d9ed7d20cc5eac428d3b84b1621694570d87bbbb35d4

Observation 6d3afc50-cc49-44d5-8f5c-60079f14a76b · inbound

LASER: Learning Active Sensing for Continuum Field Reconstruction cites this paper.

LASER: Learning Active Sensing for Continuum Field Reconstruction GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T03:35:23.240677Z digest=sha256:29bd1cb8a2f7f9e248c4bf0786fc7e676815754f1030903c11f1b7b82d53b328

Observation 0e7e998a-d4b1-4d97-904a-c8028c083c3b · inbound

LASER: Learning Active Sensing for Continuum Field Reconstruction cites this paper.

LASER: Learning Active Sensing for Continuum Field Reconstruction GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T08:30:53.002288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-05T08:29:56.797997Z digest=sha256:619a43718a02d221130b46ca436aa46a7c11a4ea86b35966a3691c5e6fe425bd

Observation a73cf345-a0b9-4e67-b8c3-fe8e89420450 · inbound

AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards cites this paper.

AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T02:18:05.770211Z digest=sha256:dfabdb59f3382cb10860b79ec6471ba1c0f39a58bdcbce7f20437b12957fcf24

Observation 1c290574-c2b5-4415-823e-fc33590f1616 · inbound

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning cites this paper.

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T07:15:28.648869Z digest=sha256:2856bae5c2288b821316f26cca33b984e2e210eaaf2192a1a0e6725ef832c1b9

Observation a0df11c5-c6e9-48c4-98c3-9f707f9547c9 · inbound

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning cites this paper.

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T02:26:14.136231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:26:14.136231Z digest=sha256:5585b0290b1be6bea918798644dbb674b7d92801c5da7fa061acd1a13ac21e99

Observation ecf9a142-5195-43a7-a4b7-01ca095de9e6 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T17:57:08.606559Z digest=sha256:d602c463635c37141599af453a75f4cacf0a232b616ae4cfe6dee637f92c7fd7

Observation 9fd3a95b-8c5f-49cf-ada2-17da193727e2 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.648398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:c02af9a8328a033ab411e5c1a2c9027f485edb6c2002680818506f0716840987

Observation a42b4f71-db54-4435-bbf6-d31c03f57ee6 · inbound

Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis cites this paper.

Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T17:48:09.078064Z digest=sha256:fe01807c772a1953c8c68b8863ba553a99dc6a8e757301e0bced948c3899e7f8

Observation e828cad4-e01a-4dd7-871a-04f5f38d6f66 · inbound

RVPO: Risk-Sensitive Alignment via Variance Regularization cites this paper.

RVPO: Risk-Sensitive Alignment via Variance Regularization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T15:00:11.237293Z digest=sha256:665c7e304af2d799a0490a0f445d89443262cdf743126b97136913f74a8c7f82

Observation 314fdb32-8343-4a63-9f61-48325ea1d40e · inbound

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients cites this paper.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:38b4a26c3c8887bb95b5587dc453aa28acbaafdbcecd9cb3bac6ec65ac9a4502

Observation c3a0b8f1-76a3-4943-8dda-e27bc9f73d8d · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:f3c8a265949ecade903da7106736aa90d95a5a3e8a8e8007e25911e0caa0c1c2

Observation dec8714c-c459-440e-af63-b17846014260 · inbound

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair cites this paper.

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T01:21:43.166304Z digest=sha256:b23d9dc99adcb8f82ff5176f3f6a82cc4fcc986f67e818b8fcf7781b8f68cbb4

Observation b86a72f0-c157-48e4-b015-4bf978b8cf12 · inbound

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning cites this paper.

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T01:52:27.764121Z digest=sha256:5f3a1ca7885fe9954ad20e60b1ac670782ae44d25faabb1ac84e363d96e00cbf

Observation 579ec141-6ada-49da-ac25-7be3f7e63d2f · inbound

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models cites this paper.

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T02:30:42.407934Z digest=sha256:2f57d0c8f51aede6f93bdc92a5223ea0e7173f20c98ea332288650db404be6a5

Observation afd2f9c9-da37-4372-ab6b-b9f41bd8a7ce · inbound

MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading cites this paper.

MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T05:22:43.330891Z digest=sha256:7c6ea840a9e405e38d45e852cf82a65e66f4b72198221790d21bbbace3722e35

Observation a8795d1d-3643-44ca-b74a-161beaf87c63 · inbound

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning cites this paper.

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T19:20:32.435135Z digest=sha256:640e94b9c90a99930d8bd4246c5cc9214eded898a8f476ae7a29b60059835474

Observation 2532e26e-3f01-4db6-909d-58dce24a829a · inbound

Scaling Retrieval-Augmented Reasoning with Parallel Search and Explicit Merging cites this paper.

Scaling Retrieval-Augmented Reasoning with Parallel Search and Explicit Merging GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T18:52:42.989888Z digest=sha256:d723d5ee0930cc73e991e990753c6dbf1292ad26c8308a1f22cdec882cd2168b

Observation 7649c6d4-b12b-4472-934c-3e8720022b87 · inbound

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization cites this paper.

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T19:24:05.375951Z digest=sha256:29417f3d934c4c42b4bb24a6c614f61baa3951e2e0b9482059dd6e0208ff5661

Observation 1f25d252-ab46-4924-a118-18dae3533de2 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:5582f39cbd22600298cb93943e7198c04e6f9e7a07fdeae6f80143e8ff1d68f2

Observation 332a722f-8f81-487c-af95-163ff3f03bbd · inbound

Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target cites this paper.

Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 18

Resolution
malformed identifier
local_arxiv, observed 2026-05-20T14:18:21.446022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T14:13:49.666793Z digest=sha256:ec0e246e65c8da280cf1ce3b9da51ac90cbd5d2f50f5bd803445318b1b00a5cf

Observation 8410e56d-cc67-4eb0-a5a7-d1f1bbebb74f · inbound

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR cites this paper.

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:03:03.550716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T05:02:35.271960Z digest=sha256:e5200797905c4663476e4ea39953f499d9ad1de6e23d8413e058e0fb34415e75

Observation fd05c426-13db-4602-8a4c-27135cbc2e9f · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T06:19:41.937876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:b288fc4d0baa39774f6ecd5794f0eac58f16a470c513ffcfa6f77b2708a42f3e

Observation eb2202f9-567a-4391-9d3d-71bbc11e88df · inbound

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection cites this paper.

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:30:19.865498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-25T04:29:16.600159Z digest=sha256:e7e221c1b2aa92c15899bd50e7bb6b9581aaf1b42a12c1c75cf535221a13bebc

Observation aa991f04-88ca-42f9-ad4f-2b0934d6c9e5 · inbound

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models cites this paper.

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:24:00.427611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T22:18:10.508034Z digest=sha256:bf650cc49b1e416607827917640f879bc59bd6b8c4abccbdfa000477c5881c06

Observation 9a425fd3-0c6f-422f-8a3e-f3c3e99c8355 · inbound

CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents cites this paper.

CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T21:53:59.318824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T21:50:04.277333Z digest=sha256:ba8af899b0c7d370cbcf101356aa264910d88cb5ac5d4619ceb0031515b65d97

Observation d750f5b7-4775-485f-8675-09b1042e36e7 · inbound

APE: Agentic Prompt Enhancer for Image Generation and Editing cites this paper.

APE: Agentic Prompt Enhancer for Image Generation and Editing GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:02:46.693518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T22:58:17.807988Z digest=sha256:86d5982e037adc0c6ce295d14c86e55bcb53c6575ab438089466ad98d28ca93a

Observation c430fc96-1c4a-42ce-980c-145aa79a3982 · inbound

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts cites this paper.

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:02:34.057899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T19:01:14.340754Z digest=sha256:91c483e6c3eb2eed7628b17f81a4ee595b2e0d57edb6fd2eff9ccb7b4ceac6bf

Observation 4c25ea9a-6e08-40ba-a01c-ad13555f3673 · inbound

Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills cites this paper.

Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:47:19.966898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T21:16:14.267471Z digest=sha256:d97b5394e794fcc06bcb699f1735cf51a9bc3f6c78f28927b0c0621462b1d817

Observation acf14a0f-0a73-4baa-aec8-9768f1748718 · inbound

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning cites this paper.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.875133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:c2316567987bc9644fc422b283fd7eea3ffb592f33bdefdf3b7f152734f593b5

Observation d2e61737-9652-43b1-a6b3-f07670e7865f · inbound

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO cites this paper.

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T05:27:39.778455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T13:18:38.450423Z digest=sha256:5b6ffdcc4843b203f565b1cc1b7d9fb7663e5b4499fecfe4c348a65be21455e2

Observation ee104ebc-bcec-4998-a31f-4cc87155e90e · inbound

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO cites this paper.

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T11:54:45.594653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:54:45.594653Z digest=sha256:93e7fcf688bddd3a0eeaec26a03ee618109a4d12a93ee8ebd2a4851b68949cea

Observation 776c23a1-4bd0-4ab8-8a5f-c5953a6d375d · inbound

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models cites this paper.

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:57:38.713059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T14:20:02.171646Z digest=sha256:c42d5ea53b5a8af15de4a7188a42f7baa11c542e84c9bb1249f205afd1da794b

Observation 43be9fb9-a16a-4356-afee-f3671210174d · inbound

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models cites this paper.

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:54:36.275829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T10:52:23.199230Z digest=sha256:be17af8cd3e29c49fac2fe816d916b7ea123532c46ade2c5286ef243b0a27cf0

Observation 063e3d13-47bb-489c-9bf6-cfbafbad0576 · inbound

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models cites this paper.

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-27T13:20:57.362956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T12:53:00.569655Z digest=sha256:91c540dcccc8902c7966367071315ef3548d3ea61faace6fb23ab83084d3ffdb

Observation ad7b8ad7-8902-4870-9053-e2d5accc0660 · inbound

DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models cites this paper.

DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:28:48.894727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T03:04:13.554098Z digest=sha256:d7408e215840802f5c9b292f4579de6d5816a8e1ca95a389ad1fa32983d9831b

Observation 0aa6f82d-6ef2-4736-9625-bb22cdbf692e · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:56.134917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:5e8f77994b9b8ee822a3702ee0fee68860633236641003afb6bd62d48aed9cbc

Observation b381c90d-44ed-4f2d-874d-647402c74b37 · inbound

ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL cites this paper.

ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:19:13.532151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T21:17:04.368521Z digest=sha256:063e4beb9deae3dc5e8a9ce334fa52485750a629fe8d6e5caa5f6b94efbacf60

Observation 5e87097d-e921-4943-94b8-b5255ecbeeea · inbound

Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning cites this paper.

Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:37.817747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T14:24:49.321248Z digest=sha256:30a00ec25f823a7c0c61cf4b6711e2230cd920234639ebaa854d5601cb8c6290

Observation 0bbf67cd-b306-4aec-8879-54cb24cae39b · inbound

Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning cites this paper.

Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:19:02.328520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-03T23:16:52.446951Z digest=sha256:fa0e22911a79ef43b2dc7808db746bcb87bb40b8c80d6997d10eb6f49ef74f53

Observation 86d15dec-f82c-43a5-9b4d-b8afaf7f053c · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 120

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:40.651987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:ec0c30a9718049e9daa35d7b110bf26f26c0bbff86eaff0bcfb7950bb3ed2c9f

Observation c6844963-29d2-4bb7-b6cb-c8d0b251f8e9 · inbound

Recommendation as Generation: Unifying Personalized Video Generation and Recommendation at Industrial Scale cites this paper.

Recommendation as Generation: Unifying Personalized Video Generation and Recommendation at Industrial Scale GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:20:07.231762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-25T20:21:09.611166Z digest=sha256:a2f9828673e61c58666b38b092bdf8b4d055b8500a00ba620610dcb92fbdd2d5

Observation 50c3f134-735c-4433-a934-a630ec25c35a · inbound

Scaling Multi-Reference Image Generation with Dynamic Reward Optimization cites this paper.

Scaling Multi-Reference Image Generation with Dynamic Reward Optimization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:39:50.262046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T05:06:12.721122Z digest=sha256:0827901fe3667eaaf61b082eb7dd21b5f2f5ac566aaeec6247dbe1479649ab36

Observation b4df026d-1a6d-4169-b402-ea447b561c76 · inbound

Qwen-Image-2.0-RL Technical Report cites this paper.

Qwen-Image-2.0-RL Technical Report GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T19:06:02.841539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T01:11:12.757656Z digest=sha256:546e0f9d722e2d41431ef5b64b6715812922a0bfccdb826856a19861d9595657

Observation 22060474-f3aa-4c8a-af79-5bb78a40b770 · inbound

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index cites this paper.

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:25:42.399885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T05:27:22.527220Z digest=sha256:34d41b882f7580c09a7b3c1294d4822912dccf253d3db0c31c6d0979dc2858dc

Observation c8456204-8654-4122-b66c-c2548f31fcae · inbound

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist cites this paper.

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:45:42.054069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T05:15:40.555929Z digest=sha256:0c2561c554be7d9dd45b634e0187eb6d01bfa8eda4ace4b03e4a368e8a14adcf

Observation 03ca1633-f4ab-4994-81b7-0608bac1dbe4 · inbound

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning cites this paper.

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:26:58.475026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-02T13:24:17.538850Z digest=sha256:cb88e6ec6f6f3979ad8c5639a2d04243bbc84739bab9f4bf2cce2c44a5bd1352

Observation c52bfc50-600d-4071-9ee1-3ce319ec6ee1 · inbound

World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments cites this paper.

World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:55.849272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-03T20:14:46.147637Z digest=sha256:c7dff048cac3a14ebd289626aabb8c9e98f4c13b1c39a744145982e25b89b316

Observation 067ffb3f-8381-4700-9906-be33b54cbd13 · inbound

TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B cites this paper.

TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.915329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-03T15:09:14.993998Z digest=sha256:6e9e36cde5c39b61c19e9f10c28234ee3943ac60f0d750930dc16183ddd00b9a

Observation 9e5e5603-4bdb-422c-b1c4-72f21a964894 · inbound

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment cites this paper.

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-08T03:14:31.674801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T03:13:07.963343Z digest=sha256:863948046ae3409a85df04e41a489792d35f47187521474866d4047176b441ac

Observation 11334e03-e355-4ae6-924e-e86d06c2c4c8 · inbound

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence cites this paper.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.271067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:db9d34c368b80b0b153d91b1f9955aba51b58bf6dce5c0c42af497cf28f23e62

Observation b8049c65-21f1-4c5a-878c-252c59991c2f · inbound

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering cites this paper.

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T14:37:32.540554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:37:32.540554Z digest=sha256:d56cf70c48d4922a39a9d3340254bad50b4b23eab80f2d5e0d925acfa75edb51

Observation f9e26078-28c1-476d-81af-d8a08b4a3e21 · inbound

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works cites this paper.

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T08:01:22.671613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:01:22.671613Z digest=sha256:081f56f5b6d273ceb8f479e4245c66b535af122ddc8b40cd0bf9ad03d6b27567

Observation 9cf67799-577a-4114-afb6-2495a684fd25 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:17.804777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:17.804777Z digest=sha256:467452b9c482a4d741305179e5a4c1d2da581870b47cc6f09feba4f4e20a80dd

Observation 4b357547-316a-48f5-a73c-4ec5b56ddc10 · inbound

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization cites this paper.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.477551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.477551Z digest=sha256:45ae37bb099b018e71cb56c6b2e2c90ec968b2bcfb964733a63e7f309d2205e2

Observation e8b4066c-9a62-4e75-9310-a28ca5601676 · inbound

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models cites this paper.

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T14:11:20.934737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:11:20.934737Z digest=sha256:b8a5c9cace1a233a411ffb034f06204809a27a3615a92dc9bde7083172be8904

Observation 4465a515-2790-4087-a521-d354a93bb977 · inbound

A Formalism-Aware Reward Loop for Handwritten UML-to-PlantUML Generation cites this paper.

A Formalism-Aware Reward Loop for Handwritten UML-to-PlantUML Generation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T15:58:06.400280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:58:06.400280Z digest=sha256:053da3ded48671fb480ebbe02f7cb7a55da415f2fb21bcae0d729fef3dcdc53c

Observation 91e8983d-011f-4b77-9d09-18d24677f944 · inbound

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL cites this paper.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.281195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.281195Z digest=sha256:4a0c953d862530bd7fca24886ea46922113fe02fdd9740c25e30c0f4107bfe7b

Observation 3bc3d340-d0ad-46f6-886e-05ad422f23b4 · inbound

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation cites this paper.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.701026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.701026Z digest=sha256:f07f9b42c8713e3262ff489ea3617aee0bba94cbd3048112ce1b9bd4637a662a

Observation b293e87d-f5db-43d3-bbab-abb4f9fe0c65 · inbound

Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence cites this paper.

Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-10T21:04:21.686902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:04:21.686902Z digest=sha256:398cc1f951f07856cdee3078f0d50d642c9308c67e59c96ae85651a309b55c7e

Observation 6883944a-b635-4ce0-836c-bbb1f219f393 · inbound

Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation cites this paper.

Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T00:38:21.946354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:38:21.946354Z digest=sha256:a141463386e4d68bd652b31e2971a79edce1858231748ed1161b915317c5c8b3

Observation 8c6746d9-388b-4d77-992a-e3c437860635 · inbound

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning cites this paper.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.633399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.633399Z digest=sha256:396bfa3d59a30506ddb3af47ccdb1ec36a359735b6b616870a475e03bef55513

Observation 16f86c57-e7c9-4d71-9c00-ed02c5ebba99 · inbound

Conversational Orchestration for Organic 6G cites this paper.

Conversational Orchestration for Organic 6G GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:39.915223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:39.915223Z digest=sha256:14bee1d1ac76345f9afc76c6f3b17ce2a2a89b32928e5ee2a9a7a707039f7daf

Observation 7bd78f97-a4eb-41da-9db7-a23278644693 · inbound

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding cites this paper.

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T17:56:48.463595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:56:48.463595Z digest=sha256:d5b1bd1c9a7a6fee3b029c6fb9f4408f5e039c5117f00fc2de102ebc27d347b6