Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T07:28:09.142557Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 2 inbound Pith citation observations for arXiv:2606.04923.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T07:28:09.142557Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:38:21.327505Z
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c5d30149-2e59-4ef0-8a2e-e1e49bc48d1d · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Process Reinforcement through Implicit Rewards
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad52b475-2f43-40fa-8336-a6c0647e623e · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e4e8840-30f9-4b58-a5c2-a0bbfe3639b0 · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Reinforcement Learning with Rubric Anchors
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 59e0214f-51f5-4fa5-823b-493cf344dea2 · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 116c76d7-88de-4864-9b42-03a1213eeced · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Anas Mahmoud, MohammadHossein Rezaei, Zihao Wang, Anisha Gunjal, Bing Liu, and Yunzhong He
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13a3e2c0-5970-4d88-b2e2-3c943be9f22d · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Reward Hacking in Rubric-Based Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 21f02363-1a98-43df-8bac-c410e831d4b1 · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning LMSYS Arena technical blog and evaluation suite
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c048db2-ceb2-4b26-9b35-27f32a6b8f03 · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning LLM Evaluators Recognize and Favor Their Own Generations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 20efc244-71ef-4cca-9c38-d5b0e1d48a46 · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69298f13-a195-40c1-a79c-7c4ee426a8c6 · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Rl tango: Reinforcing generator and verifier together for language reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a74f70af-1f51-4db6-bf16-5eb7f8d15b90 · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Chaining the evidence: Robust reinforcement learning for deep search agents with citation-aware rubric rewards
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c43b644e-7461-4f5e-9e52-fc598aa0b036 · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e33b69e7-b985-4035-8a2b-6466063fd67c · outbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b18520-87ef-48b5-bbd2-b0891ff74f9c · inbound
Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71c07a45-7329-4c20-9d61-7033f69c75d2 · inbound
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.