Pith. sign in

Paper Citation Record · LEDGER

Value-Free Policy Optimization via Reward Partitioning

As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.13702.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13702 v4

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:49.368412Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:20.906356Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d1e438d-cf68-4314-9294-e3aba8e89a37 · outbound

This paper cites Rrhf: Rank responses to align language models with human feedback,.

Value-Free Policy Optimization via Reward Partitioning Rrhf: Rank responses to align language models with human feedback,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.821253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.229111Z digest=sha256:68b4a145099ae5a1375118402f2487dbfb04d84702d27ac444d46f2c4117c838

Observation 2be3843c-13a7-4a96-be30-6551e33c15e0 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Value-Free Policy Optimization via Reward Partitioning RLHF Workflow: From Reward Modeling to Online RLHF

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.233090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.233090Z digest=sha256:822d16a99971dcffedc9efe6f542307c5022b41b3a5ca5d7971da3331e125bb0

Observation c6740c56-52c9-4774-bd8a-e41a14d51045 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms,.

Value-Free Policy Optimization via Reward Partitioning Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.811145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.236673Z digest=sha256:1b3338e8fea2fe3a341eaa74e921aee318657b498b3fa59013749ba7399fae38

Observation 0364f190-1d28-4505-b39d-323425002d42 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Value-Free Policy Optimization via Reward Partitioning Direct preference optimization: Your language model is secretly a reward model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.241247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.241247Z digest=sha256:93fbfa196c2a9ed2970b1f0cbf0baceead60569b653064044d307a8a62d11cd0

Observation 170bbaa3-093d-4e6d-a9ea-1ab59660b784 · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment,.

Value-Free Policy Optimization via Reward Partitioning Generalized preference optimization: A unified approach to offline alignment,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.795028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.244822Z digest=sha256:405bcf2fe980d9a8d432c8b56334d68f3fa051d3b972cd146d6712e47a330c09

Observation 848849fa-a948-449d-976a-015b8f7fb801 · outbound

This paper cites Offline Regularised Reinforcement Learning for Large Language Models Alignment.

Value-Free Policy Optimization via Reward Partitioning Offline Regularised Reinforcement Learning for Large Language Models Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.248057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.248057Z digest=sha256:03efd572a3b54deca8f0600bdc2b3f533ac48e8bb81fe516d6d8be0b18ea14a3

Observation ffe4e4d3-d6bc-4aa1-8966-aed06214946b · outbound

This paper cites Model alignment as prospect theoretic optimization,.

Value-Free Policy Optimization via Reward Partitioning Model alignment as prospect theoretic optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.785448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.251658Z digest=sha256:96ae9cc1a2205b116618c24a7a2c5de712b052f887ca54671fff107334294f76

Observation 6699bbdc-bc91-4330-9be9-e3013b84bd57 · outbound

This paper cites Instruction tuning for large language models: A survey,.

Value-Free Policy Optimization via Reward Partitioning Instruction tuning for large language models: A survey,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.254776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.254776Z digest=sha256:fb8b2e204df681b297f7e39257a7299bcaeb03f42186c29146bfd7f4ce322540

Observation fb0b05c5-085b-4d2c-b061-a4849a9ca341 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Value-Free Policy Optimization via Reward Partitioning Proximal Policy Optimization Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.259019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.259019Z digest=sha256:36cbc1c11c6d0142626a0f6ded7bc8ff49f93071bcf3ff1c1acb397a9a906165

Observation ff63cb40-f6da-4a37-bc7b-9a7815d04dff · outbound

This paper cites Trust region policy optimization,.

Value-Free Policy Optimization via Reward Partitioning Trust region policy optimization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.775562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.263025Z digest=sha256:69b10eaeedfd7133fc858672985f3d5671b74e86c1de99f29fc9734c1b80039c

Observation 3b6d7598-5178-4239-ba4a-ae55ba16e2ab · outbound

This paper cites Training language models to follow instructions with human feedback,.

Value-Free Policy Optimization via Reward Partitioning Training language models to follow instructions with human feedback,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.266419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.266419Z digest=sha256:f93a19ef3f77212c48b6db7d842eda732a030ce8bc9b230182f05b2a85dc47d3

Observation bc503428-b3a4-4786-8dab-a4ce942ed59c · outbound

This paper cites GPT-4 Technical Report.

Value-Free Policy Optimization via Reward Partitioning GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.269741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.269741Z digest=sha256:2bbf1bcf6b5240164d33507e521ec1d6774ed2061c1b598b9fc34bf87c7da43f

Observation 41a6f198-8d6c-4b8d-bc1c-481155b2fd4c · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

Value-Free Policy Optimization via Reward Partitioning The claude 3 model family: Opus, sonnet, haiku,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.759152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.272989Z digest=sha256:b7f19e89ac76ddec918d5d17a49bafec6ae40547f777f561c3f096bf7ea36656

Observation 3056ed27-b9f2-4632-b3e7-ae35c551fc55 · outbound

This paper cites Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing,.

Value-Free Policy Optimization via Reward Partitioning Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.748634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.277532Z digest=sha256:7dd6f9ec8891ce15df5566874d5450b7b825bbfac6fd54bfd98fac591ff488af

Observation cff56b3c-ca0c-458e-92b1-3cd249e28eca · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models,.

Value-Free Policy Optimization via Reward Partitioning Guiding pretraining in reinforcement learning with large language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.738544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.280587Z digest=sha256:5aa2bf00a9e59783aab2db80efa8ed4ba5cd6295dfc03f6545306e608408c790

Observation 0d3ab8c7-82c6-472a-85b1-493e848b096b · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Value-Free Policy Optimization via Reward Partitioning RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.283692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.283692Z digest=sha256:f61ada2ee79a916f08948f3ef7ddda61a4f0993a36be71a62eeb71a26ef9efce

Observation 8e48b80a-f4a5-4b8c-a791-f04e80f8181f · outbound

This paper cites CREAM: consistency regularized self-rewarding language models,.

Value-Free Policy Optimization via Reward Partitioning CREAM: consistency regularized self-rewarding language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.727405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.287664Z digest=sha256:20a904f6c0111d767d08831c302dc7f268e2afa100830dfbe185ca0f37020752

Observation 0267a425-915e-45a6-a8be-60f29fa65428 · outbound

This paper cites Self-rewarding language models,.

Value-Free Policy Optimization via Reward Partitioning Self-rewarding language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.717989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.290911Z digest=sha256:892a83dc08ca696c9df7e5c73580cab626f2607c4e6aef79d581ae318c682ec3

Observation 586052fa-c7b1-42f5-b7d7-8e467d6e9ebc · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Value-Free Policy Optimization via Reward Partitioning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.293738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.293738Z digest=sha256:2d6870fde898679e1f0bd08be9fdc923c69f917b7e7ab2d51e785bb7f04b76b6

Observation 15bae0be-35c2-44a2-bd98-8b0a8a551dbd · outbound

This paper cites The Llama 3 Herd of Models.

Value-Free Policy Optimization via Reward Partitioning The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.297393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.297393Z digest=sha256:d1f2742bfe620d11dda8320d6b71f45d98019174ce4eee3faf0cc5e1bda18a1c

Observation 163ec618-6992-4fcf-ae7d-625c5dd4cf0a · outbound

This paper cites Qwen3 Technical Report.

Value-Free Policy Optimization via Reward Partitioning Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.300277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.300277Z digest=sha256:06a29f8ce4a5f49fcfe9067e30ddfb4d4f13db83b47a8234faa2280da04cc63b

Observation 2acf63e6-e1f2-4458-9c6f-427069ec6d62 · outbound

This paper cites Nemotron-4 340B Technical Report.

Value-Free Policy Optimization via Reward Partitioning Nemotron-4 340B Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.303322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.303322Z digest=sha256:623629a8cf5b43b0f4e98307166eb6728d03279d73ff15359ef90e3be7da2bcf

Observation 4d4fdb80-5a78-4520-b10c-85b41aeb9bdf · outbound

This paper cites Rank analysis of incomplete block designs: I. the method of paired comparisons,.

Value-Free Policy Optimization via Reward Partitioning Rank analysis of incomplete block designs: I. the method of paired comparisons,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.306523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.306523Z digest=sha256:58cd6dc05d9670dc7167e4edfa917befa35081ddf5abe15733361766286e7716

Observation b0633016-44d7-4956-a4d3-a56c310edf3e · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences,.

Value-Free Policy Optimization via Reward Partitioning A general theoretical paradigm to understand learning from human preferences,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.702083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.309895Z digest=sha256:9cc18668167ee0a670d771bf6d381de65a699db5f13ae8be77956c42e9a7f533

Observation f496c456-0247-4c12-b42d-338ec07bae57 · outbound

This paper cites Advances in prospect theory: Cumulative representation of uncertainty,.

Value-Free Policy Optimization via Reward Partitioning Advances in prospect theory: Cumulative representation of uncertainty,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.692625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.312840Z digest=sha256:9c4b584ff8d59bac092b21fd14da8efd62e382cda071d0b0fe18bb8c2639af95

Observation efb414d2-4669-4b7d-a5eb-e8b65446ba0b · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Value-Free Policy Optimization via Reward Partitioning UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.315854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.315854Z digest=sha256:dfc7345783cd51bfb27392cda71cc2b741fa1a1643616a1b6e33e344f7653d49

Observation 7262b461-5e87-474b-85dd-4960bc1d8d80 · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models,.

Value-Free Policy Optimization via Reward Partitioning Alpacaeval: An automatic evaluator of instruction-following models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.683149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.319398Z digest=sha256:2ab530dda7762869e7366a3ba76269d56ae0ccd2bae44ddb88fca3b4bdd0b57b

Observation 29968375-32d7-49df-bf73-acb155c2e158 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Value-Free Policy Optimization via Reward Partitioning Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.322892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.322892Z digest=sha256:29964f85a217b8225cbbbd3a77567917b23cd05c155a9224ceadee2aa3b81d23

Observation 3b78be03-5ee5-44bc-9e33-ffbaca2aee97 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Value-Free Policy Optimization via Reward Partitioning Instruction-Following Evaluation for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.326677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.326677Z digest=sha256:6a38618b5390cb40d21b085e0608c15d28bfa0c7a76817dac1a8bfd32c67c0ad

Observation 2b07c684-5b46-4e78-b412-79e35815b2c1 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Value-Free Policy Optimization via Reward Partitioning Training Verifiers to Solve Math Word Problems

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.330155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.330155Z digest=sha256:30cc3a0f697aedbddce8f81fc92cdd86de0f09f845dbaad4cad551ba6b5420f0

Observation 17fce266-5c00-4cb7-852f-8059491a6f1c · outbound

This paper cites Scaling up models and data with t5x and seqio,.

Value-Free Policy Optimization via Reward Partitioning Scaling up models and data with t5x and seqio,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.667368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.333178Z digest=sha256:0704e2684ebf666d67f4177ce2217336cb7c42e5a5fb03f63e04b6e453dc0027

Observation 7fb1bfa7-74ea-4b4e-93ff-1c7f6ca8f999 · outbound

This paper cites Mistral 7b,.

Value-Free Policy Optimization via Reward Partitioning Mistral 7b,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.336652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.336652Z digest=sha256:b7d2849e34924bdd3033a42c8ff0c47833bc9a46df51db229f6ced3c91aa2316

Observation ab1ceb4b-8ef4-47cf-9ba4-9e85d1442e9d · outbound

This paper cites Llama: Open and efficient foundation language models,.

Value-Free Policy Optimization via Reward Partitioning Llama: Open and efficient foundation language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.650316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.340728Z digest=sha256:ab99a96b953ab02b1198a6c76ea193838c0938884e9f494b7ac120e5936c6d0a

Observation edba6570-de0e-4a9b-9180-4ac294105a90 · outbound

This paper cites Qwen2 Technical Report.

Value-Free Policy Optimization via Reward Partitioning Qwen2 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.344252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.344252Z digest=sha256:9ce58fb380a37e96bb24c3c31dd3a7b55dde4c816d946e04e30c45acd23dbd44

Observation 526d418b-dd00-4c0e-8f9d-32f81c3ca79a · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Value-Free Policy Optimization via Reward Partitioning BERTScore: Evaluating Text Generation with BERT

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.347457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.347457Z digest=sha256:7f77805837471606d5a494027ec916dd868804aa345af0e30224d562e46f185b

Observation 762e49c1-11ac-4aac-bf6e-ea267e0d46be · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Value-Free Policy Optimization via Reward Partitioning Rouge: A package for automatic evaluation of summaries,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.350864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.350864Z digest=sha256:0aad0e5a887e9ded2c8954cdb6425b2b78482e606a52493b195ff45c69a6e293

Observation 1c2fe268-4e4a-4beb-85cc-8fae2cd7f522 · outbound

This paper cites A diversity-promoting objective function for neural conversation models,.

Value-Free Policy Optimization via Reward Partitioning A diversity-promoting objective function for neural conversation models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.634832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.354796Z digest=sha256:26670d7413b6d70d693773c52e5a75b794f9738fede2e7ebadabab9b31e5ed41

Observation c84dfb56-bb9d-4b13-8610-5d333745a163 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Value-Free Policy Optimization via Reward Partitioning A Survey on LLM-as-a-Judge

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.358212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.358212Z digest=sha256:aa503f2b7b5869eccb377be50b70edf18ebedc26b365b20222c12fb2d16143a7

Observation 0c8e74f6-81cf-4fd3-9c7a-df18274cf243 · outbound

This paper cites GPT-4o System Card.

Value-Free Policy Optimization via Reward Partitioning GPT-4o System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.361713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.361713Z digest=sha256:745b00efdc12f0380e45e32d001dcfc82265e606e6011ec5ae3c276cf7d3a173

Observation d8c59ea9-5c99-45a5-bbd7-2c6f3921eecc · outbound

This paper cites Claude 3.5 sonnet.

Value-Free Policy Optimization via Reward Partitioning Claude 3.5 sonnet

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.623894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:31:49.365419Z digest=sha256:74395d9e136d8b39373e83c0230b5b2977a788ec370614c817edf1f42e6d3703

Observation 9a806028-570f-490e-bdc4-e44b439ba7ac · outbound

This paper cites Decoupled weight decay regularization,.

Value-Free Policy Optimization via Reward Partitioning Decoupled weight decay regularization,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.368412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.368412Z digest=sha256:9a809565ede960fb1aac579802444ba9483c4a4d2d8b8ea026ff2bb5929b4ad4

Pith citing papers

Observation 58037a65-8d50-481a-8f54-57238aecf6dd · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Value-Free Policy Optimization via Reward Partitioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.906356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.906356Z digest=sha256:cfa50e684b550a00d7290103b4bf22fc50615eafe753277ac59cb36ac473ee84