Pith. sign in

Paper Citation Record · LEDGER

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges

As of 19 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2607.05904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05904 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T21:20:07.569587Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:30:04.547511Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch10

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 506e4a48-45dd-4232-bdb2-632c0437beae · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Constitutional AI: Harmlessness from AI Feedback

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T21:25:38.646181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:59256c32db2143960293e322a9fc10c00ddb429fb9c60a767c19545e4dd8a4c2

Observation a2a38675-5ead-4f4a-89e7-1e1b34693671 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.654700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:be4292d7e46aca7d7e0599e0962795dbc73dcf67d15ab39e7a9f21e5687d3a08

Observation e426a20c-f573-4c4b-a2db-4c6fccabcff7 · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Reward Model Ensembles Help Mitigate Overoptimization

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.636176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:50bafb88eac1ff10f018fa73118ba19b7fc0bd7ab1d94508eb877ed1866618c4

Observation 6539054c-9dac-4621-b592-f3be852adeb6 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Scaling Laws for Reward Model Overoptimization

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.660138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:03be55719f8be7c2692e4dd6220412a3adffda303f15a6bf2d9251ff69989f85

Observation 5f5cc307-ba46-4aab-bcad-d2edb1acc691 · outbound

This paper cites On scalable oversight with weak LLMs judging strong LLMs.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges On scalable oversight with weak LLMs judging strong LLMs

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.643507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:9bb49334e6752c1deaa7cdb5abd3956177c8d94cfadae9825a8ef509b0b84e7f

Observation 3aa99a92-a834-4240-aaea-2d299ae90004 · outbound

This paper cites When Does Verification Pay Off? A Closer Look at LLMs as Solution Verifiers.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges When Does Verification Pay Off? A Closer Look at LLMs as Solution Verifiers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-08T21:25:38.633754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:28aa517afab57e345bad85c2238c47ef951f72697ea0cea8d106d185ec7ccf6a

Observation 69d997df-802e-42d1-aafe-6cd36dd11edf · outbound

This paper cites Spontaneous Reward Hacking in Iterative Self-Refinement.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Spontaneous Reward Hacking in Iterative Self-Refinement

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.621167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:e02c3e373c88369a7f789943494d5f6d2f819d1c7e3681330891518982d7787c

Observation c8f2e35b-8718-4d47-a563-965679ca8cff · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-08T21:25:38.651895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:84873513c5225b462e37907a75320252c078ddb13f8e8c2eaff1cbfa00f8293f

Observation bb7db7e6-05e1-4192-86a6-93676208df0a · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Towards Understanding Sycophancy in Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-08T21:25:38.623673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:e1c024a9310eca2e75b98a57d250718187d3476166764495108681d3b9a115e9

Observation 74981873-42cf-49ce-8db6-a0512b782060 · outbound

This paper cites RLSR: Reinforcement Learning from Self Reward.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges RLSR: Reinforcement Learning from Self Reward

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.648851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:74fe9469157e60a2b2a511b43beb8cb51d947a9d658b4f55e41a13472bf0c985

Observation fad3a143-a4bd-4dc4-9ba6-2eac61ce336e · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges A Long Way to Go: Investigating Length Correlations in RLHF

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T21:25:38.638530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:9f09e6433f9719f744d2bd89ecb212ff30d79cc8cdfc8d8aa272b2f54a7a5dd3

Observation 8ecccce0-e3b5-4121-8e76-605d931b0bb1 · outbound

This paper cites Language Models Learn to Mislead Humans via RLHF.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Language Models Learn to Mislead Humans via RLHF

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.626147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:f31fe7a62e649d6608a1ac494e139a6632c4473381d6e7d17de6d2e58bc3eaa7

Observation 3f3cf442-54df-48bf-9996-6e8ce059e79d · outbound

This paper cites Self-Rewarding Language Models.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Self-Rewarding Language Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.628686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:998aec2e14e6c1c9d9c6cbd713a7e17df36e91be6d7b2a32d82477385859d4ca

Observation 673ffde6-90b0-4138-9369-886aaa53056e · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.631126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:00e17219b81881fde0663b7ececb289a88758a83e32bf172598fb69b6447f370

Observation 23382bfc-ab7b-4f1b-95f5-22557924239b · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.657623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:97797aeea844298daa7f443439fc506dac7693abd16a5c2676036ec085379c86

Observation 9f584040-5632-4e7d-8caa-742fc0662be9 · outbound

This paper cites My answer:.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges My answer:

Reference 16

Resolution
malformed identifier
raw_fallback, observed 2026-07-08T21:25:38.776124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:68831ba9961a27914a9f3e0cfc0efaf124f3bd3c73c7e25fdc217512ce7e9c2a

Observation dd5d18f7-6d21-4e60-851b-7e93d26be757 · outbound

This paper cites an unresolved cited work.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Unresolved cited work

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-08T21:25:38.640934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:5b8b8e5c98abf0bcc2cce85105e53d1016130c148b2edacc4634ed1319049074

Pith citing papers

Observation 6ae011a9-146e-4fe1-beee-c62444a77711 · inbound

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR cites this paper.

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T07:36:47.336670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:36:47.336670Z digest=sha256:5c8f68c3dedcdcb622be3be276b0cd2a9a493c8763316b32cd3339fd797b64c7

Observation c00c37c5-f72f-4540-aa54-4c93eb9684bd · inbound

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition cites this paper.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:04.547511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:04.547511Z digest=sha256:fb3b426bfe3e97bd9e7e99b6e6f28dbadcc923feb5919d86dfecebfe63a356eb