Pith. sign in

Paper Citation Record · LEDGER

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 2 inbound Pith citation observations for arXiv:2606.04923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.04923 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:28:09.142557Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:38:21.327505Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c5d30149-2e59-4ef0-8a2e-e1e49bc48d1d · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:16:43.760659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:1cfbff98aeeae6de7def6677223030d65cbbf031d00e8af448ffd617bc1a341b

Observation ad52b475-2f43-40fa-8336-a6c0647e623e · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:16:43.763323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:59965fed22ee954168a354268c0bbdcd58d8acaaea8f7dc8cb3b10c252cff7ab

Observation 5e4e8840-30f9-4b58-a5c2-a0bbfe3639b0 · outbound

This paper cites Reinforcement Learning with Rubric Anchors.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Reinforcement Learning with Rubric Anchors

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:16:43.765955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:7e221bc2fc9a6f646efafeadcb59d7b654dac15fac12b4715959e16cfb5b217b

Observation 59e0214f-51f5-4fa5-823b-493cf344dea2 · outbound

This paper cites Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:16:43.768695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:1e9aa51a856c3991d548a3783dfbb0ff5120bd5828bb98b71ba2889ab54db8f3

Observation 116c76d7-88de-4864-9b42-03a1213eeced · outbound

This paper cites Anas Mahmoud, MohammadHossein Rezaei, Zihao Wang, Anisha Gunjal, Bing Liu, and Yunzhong He.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Anas Mahmoud, MohammadHossein Rezaei, Zihao Wang, Anisha Gunjal, Bing Liu, and Yunzhong He

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:16:43.752573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:4a537fbd779f40dc3c0d1bb9a213b2d818103ca3c1a88701cb360a932cf084cb

Observation 13a3e2c0-5970-4d88-b2e2-3c943be9f22d · outbound

This paper cites Reward Hacking in Rubric-Based Reinforcement Learning.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Reward Hacking in Rubric-Based Reinforcement Learning

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:16:43.741991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:2b0ab307e2eec5039efbae6e9c6ae80627977a167dc735ae0c9d62c38bbdb3b5

Observation 21f02363-1a98-43df-8bac-c410e831d4b1 · outbound

This paper cites LMSYS Arena technical blog and evaluation suite.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning LMSYS Arena technical blog and evaluation suite

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T07:28:09.142557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:73bbde3138cdb1f8efcaed6b7c147b83e2efda4c6919b1cd2585d9a92f2a3b2f

Observation 0c048db2-ceb2-4b26-9b35-27f32a6b8f03 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning LLM Evaluators Recognize and Favor Their Own Generations

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:16:43.747093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:8cca0db9a2e0c3dc90a37c4a2c0a5f3531fc4d2f7833d1221a81bcac33960b17

Observation 20efc244-71ef-4cca-9c38-d5b0e1d48a46 · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:16:43.744531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:32a75d8a03573f62a9e1e6572626e45b3818bfba06f9a7bfdbe6978565343ef1

Observation 69298f13-a195-40c1-a79c-7c4ee426a8c6 · outbound

This paper cites Rl tango: Reinforcing generator and verifier together for language reasoning.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Rl tango: Reinforcing generator and verifier together for language reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:16:43.755400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:be695e1812348f9f0c65ee028d71dce3b52824ae54c734d3d056c231ad2abaa7

Observation a74f70af-1f51-4db6-bf16-5eb7f8d15b90 · outbound

This paper cites Chaining the evidence: Robust reinforcement learning for deep search agents with citation-aware rubric rewards.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Chaining the evidence: Robust reinforcement learning for deep search agents with citation-aware rubric rewards

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:16:43.749812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:813167d7ab8d3402b77c84912b709b96e91c6a371dbd7f9e6d514883c73cc057

Observation c43b644e-7461-4f5e-9e52-fc598aa0b036 · outbound

This paper cites Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-08-04T02:24:30.324324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:8aa27d744f5cbf104693c07ccf58020bb6f649f8ebe9047ff17379b9464724c2

Observation e33b69e7-b985-4035-8a2b-6466063fd67c · outbound

This paper cites an unresolved cited work.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T07:28:09.142557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:771feef9a37ba7e86116fb15cfa148254d4c708f574239ef3e139629f913e744

Pith citing papers

Observation 39b18520-87ef-48b5-bbd2-b0891ff74f9c · inbound

Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG cites this paper.

Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T05:47:14.970921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:47:14.970921Z digest=sha256:be7e925ade865c577fd77e28bf00d6c5dc7b3a4c97935e3b45c714874a15f82c

Observation 71c07a45-7329-4c20-9d61-7033f69c75d2 · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:21.327505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:21.327505Z digest=sha256:c113ad101986fb9599a4dff089774840220e26f30135915849266c2dac581cd0