Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:07.633786Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 4 inbound Pith citation observations for arXiv:2509.22047.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:07.633786Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T08:01:22.596796Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
35 of 35 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation a3726d08-fc94-415a-b2fb-927bf1841e39 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems URL: " 'urlintro :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66e50fcd-521d-457e-9ccd-3ec88077a324 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a190610-c58f-4f4b-88be-babca967dd9d · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48324b8d-8d99-46a1-ad87-f61a763f4819 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Concrete Problems in AI Safety
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc09a5fc-0f1d-4b49-a885-e643a163392e · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 760ba96c-ecea-41e6-a84b-43d6098f94d1 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4c476ffd-6f71-47da-9b68-d7ec0b870b1e · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Alegre, Ann Nowe, Ana Bazzan, El Ghazali Talbi, Gr\' e goire Danoy, and Bruno C
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c99a6e90-0a57-4def-a4a1-a03e316fb8e0 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation da0ff5f9-e48f-4e8f-8f72-50680748e675 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4129b698-f09a-4d28-8c40-669c6998dace · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a123be4e-6328-4b17-89c9-46ee84a7d3a8 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0cdecf5c-ab15-49e9-90d5-0e3b2c68a704 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7cb53439-b37c-46ac-a318-7ddb1aed75c9 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee9dc9f2-ed03-451f-b8e3-6b8b7c43b9b7 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e810e80a-1e2c-41ed-9040-29ecc7dfa711 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3cde26d9-8acb-466c-b53a-8b8d9b873093 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e021a919-49f1-4e17-a82c-b6afdea50566 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d8f6086f-edbb-444a-ad56-e6c81bfe27ab · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems DeepSeek-V3 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d896acf7-ec33-4129-aec3-2e9a403b4625 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Understanding R1-Zero-Like Training: A Critical Perspective
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c222c6-1e6b-4426-a944-ae9d3b266aa9 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e23f06e1-02aa-418a-9697-73c43e8af065 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Bradley Knox, Chelsea Finn, and Scott Niekum
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation eb683d59-e733-45c2-b54a-fd10b31088b4 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Magistral
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 514b1e38-43e0-4bd5-a606-2a4150296a47 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a6ce09f-1bb0-455f-a487-bf80d7821883 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6771ef64-3ba7-43f3-b4e3-647754b8e4b1 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 14c1781f-c34e-4d4d-ae53-b8619391bfc8 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9d1b11e7-5133-4d5a-b163-084082d8ac14 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6d7cbe3a-90be-465c-9e47-0fa424104b69 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 968d5be7-7c4c-4c75-bfca-aba6ac703d52 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2d50b4-5f92-4b7b-9444-e13a3ae5bbe5 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems BNPO: Beta Normalization Policy Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5654c1-11e8-495d-aa80-613796a2a4f0 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Qwen3 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c491fc-cc26-41cc-820f-ab52ce302f17 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Group Sequence Policy Optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e784c5-8604-4ff7-aa4a-d9fa2cb420a2 · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ec15d00e-3a95-4f00-8c22-c978c59cd61b · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc5925d-4d6f-4cfe-acad-35db3d87ebff · outbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Fine-Tuning Language Models from Human Preferences
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48dbd911-a145-4704-a5fb-8372531e2bb0 · inbound
Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7406bf45-c10e-4699-a55d-e05fd9c9b6e3 · inbound
LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6a1c1148-0f65-4735-b945-e1745b9a9113 · inbound
CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5e56dbc7-b428-44ad-acbf-d3500a46a56b · inbound
The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.