Pith. sign in

Paper Citation Record · LEDGER

Value-Free Policy Optimization via Reward Partitioning

As of 22 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.13702.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13702 v4

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:49.368412Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:20.906356Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d1e438d-cf68-4314-9294-e3aba8e89a37 · outbound

This paper cites Rrhf: Rank responses to align language models with human feedback,.

Value-Free Policy Optimization via Reward Partitioning Rrhf: Rank responses to align language models with human feedback,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.821253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.229111Z digest=sha256:9cac47df384a689023c91de4d8a861b4883e36fbb3678162f756c8695d20d4c8

Observation 2be3843c-13a7-4a96-be30-6551e33c15e0 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Value-Free Policy Optimization via Reward Partitioning RLHF Workflow: From Reward Modeling to Online RLHF

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.233090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.233090Z digest=sha256:ce6c004a12b1c594fe35a2053b21a63fdf70e7636ecc9e9cf856659278e1405e

Observation c6740c56-52c9-4774-bd8a-e41a14d51045 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms,.

Value-Free Policy Optimization via Reward Partitioning Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.811145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.236673Z digest=sha256:c462c649652bfe894448356617fa39dc8e235200bc9ab1faf5dd12e9ff840d6b

Observation 0364f190-1d28-4505-b39d-323425002d42 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Value-Free Policy Optimization via Reward Partitioning Direct preference optimization: Your language model is secretly a reward model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.241247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.241247Z digest=sha256:179d9f247c9df02cb74267086da08faf106db8014180143bf823f79264c00e6b

Observation 170bbaa3-093d-4e6d-a9ea-1ab59660b784 · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment,.

Value-Free Policy Optimization via Reward Partitioning Generalized preference optimization: A unified approach to offline alignment,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.795028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.244822Z digest=sha256:4cbfef1d2c8be41b70e47e4b0e72350bfc4716be83f7fc9eb18b03f5cebb07eb

Observation 848849fa-a948-449d-976a-015b8f7fb801 · outbound

This paper cites Offline Regularised Reinforcement Learning for Large Language Models Alignment.

Value-Free Policy Optimization via Reward Partitioning Offline Regularised Reinforcement Learning for Large Language Models Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.248057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.248057Z digest=sha256:1c13a13ce56bfa815f7003fc843582e500dc93b8784536b77b38ec25af93c1bf

Observation ffe4e4d3-d6bc-4aa1-8966-aed06214946b · outbound

This paper cites Model alignment as prospect theoretic optimization,.

Value-Free Policy Optimization via Reward Partitioning Model alignment as prospect theoretic optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.785448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.251658Z digest=sha256:5c4b6aee4e7faf641a18a4797b5e01759dc123e762deb0cf5dbb3dceebc5363e

Observation 6699bbdc-bc91-4330-9be9-e3013b84bd57 · outbound

This paper cites Instruction tuning for large language models: A survey,.

Value-Free Policy Optimization via Reward Partitioning Instruction tuning for large language models: A survey,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.254776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.254776Z digest=sha256:56efae9abf838574f067bc56af429d7fa887dfa7904ddc33b5d7c90dcd2514a3

Observation fb0b05c5-085b-4d2c-b061-a4849a9ca341 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Value-Free Policy Optimization via Reward Partitioning Proximal Policy Optimization Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.259019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.259019Z digest=sha256:b4507e8ae94a64a0d5818c2af7cb0aebb0b25d8db8a30fa583bb82f38a8a6444

Observation ff63cb40-f6da-4a37-bc7b-9a7815d04dff · outbound

This paper cites Trust region policy optimization,.

Value-Free Policy Optimization via Reward Partitioning Trust region policy optimization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.775562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.263025Z digest=sha256:2ec8d111ae38bf8e5fed6c640248972c8d39a78635ca8b6b37258a2e0e560522

Observation 3b6d7598-5178-4239-ba4a-ae55ba16e2ab · outbound

This paper cites Training language models to follow instructions with human feedback,.

Value-Free Policy Optimization via Reward Partitioning Training language models to follow instructions with human feedback,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.266419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.266419Z digest=sha256:6fb2a7c2570d83738c5d7e2772ea44f5af2a03c20fa7b2c0fd4bd75a317b8647

Observation bc503428-b3a4-4786-8dab-a4ce942ed59c · outbound

This paper cites GPT-4 Technical Report.

Value-Free Policy Optimization via Reward Partitioning GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.269741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.269741Z digest=sha256:4446c91cb7ba3ddd6a35de685ad7719874f8e3c2a6f707e7a9f52d8c7ca56d01

Observation 41a6f198-8d6c-4b8d-bc1c-481155b2fd4c · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

Value-Free Policy Optimization via Reward Partitioning The claude 3 model family: Opus, sonnet, haiku,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.759152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.272989Z digest=sha256:6ff9f0643654abb980c20827c06907523990d30de8ed424d9041b2335a371e3b

Observation 3056ed27-b9f2-4632-b3e7-ae35c551fc55 · outbound

This paper cites Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing,.

Value-Free Policy Optimization via Reward Partitioning Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.748634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.277532Z digest=sha256:2b82e62efe024a20b7e7828d0dee6bb0ae3b2e1e5f6e9b73550b6eb71a0292af

Observation cff56b3c-ca0c-458e-92b1-3cd249e28eca · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models,.

Value-Free Policy Optimization via Reward Partitioning Guiding pretraining in reinforcement learning with large language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.738544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.280587Z digest=sha256:efb5a2fa716d2d86d02cd3610b6e6ce3f96650a48083bd84b4bdf603e963441d

Observation 0d3ab8c7-82c6-472a-85b1-493e848b096b · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Value-Free Policy Optimization via Reward Partitioning RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.283692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.283692Z digest=sha256:3225a082b3edc75583d306ca940e400fb7b8d2f77014e43e58b00af25cd31368

Observation 8e48b80a-f4a5-4b8c-a791-f04e80f8181f · outbound

This paper cites CREAM: consistency regularized self-rewarding language models,.

Value-Free Policy Optimization via Reward Partitioning CREAM: consistency regularized self-rewarding language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.727405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.287664Z digest=sha256:85274d244783b3fa2328844fba9e86a8b683e8b8ac49529f769f0a92b00effa2

Observation 0267a425-915e-45a6-a8be-60f29fa65428 · outbound

This paper cites Self-rewarding language models,.

Value-Free Policy Optimization via Reward Partitioning Self-rewarding language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.717989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.290911Z digest=sha256:d12e4457c6d387c0236230b19a4003b6e6b77e4df670e181f6c216124b9f15c0

Observation 586052fa-c7b1-42f5-b7d7-8e467d6e9ebc · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Value-Free Policy Optimization via Reward Partitioning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.293738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.293738Z digest=sha256:67356a66008acf627255e4d483de33bcdb28b7eafe9673689dff2f4d177a795f

Observation 15bae0be-35c2-44a2-bd98-8b0a8a551dbd · outbound

This paper cites The Llama 3 Herd of Models.

Value-Free Policy Optimization via Reward Partitioning The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.297393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.297393Z digest=sha256:d9785fa9ae5945325652ce153b8607e247cdf697d97ffa4748b8a3abbc31efe6

Observation 163ec618-6992-4fcf-ae7d-625c5dd4cf0a · outbound

This paper cites Qwen3 Technical Report.

Value-Free Policy Optimization via Reward Partitioning Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.300277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.300277Z digest=sha256:adebd4ffd22fcb1548a04c110868fb1f0638e08f0c377dd49edb3b87eb4b2fa5

Observation 2acf63e6-e1f2-4458-9c6f-427069ec6d62 · outbound

This paper cites Nemotron-4 340B Technical Report.

Value-Free Policy Optimization via Reward Partitioning Nemotron-4 340B Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.303322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.303322Z digest=sha256:2e46ed2988b3707f69b2c0a9c88c599be55df52380e317d6496246c95a394afd

Observation 4d4fdb80-5a78-4520-b10c-85b41aeb9bdf · outbound

This paper cites Rank analysis of incomplete block designs: I. the method of paired comparisons,.

Value-Free Policy Optimization via Reward Partitioning Rank analysis of incomplete block designs: I. the method of paired comparisons,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.306523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.306523Z digest=sha256:48b48aee38fc4cc4ae961d898ab50887bd6b1d189b4050267949c1efd40680ca

Observation b0633016-44d7-4956-a4d3-a56c310edf3e · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences,.

Value-Free Policy Optimization via Reward Partitioning A general theoretical paradigm to understand learning from human preferences,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.702083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.309895Z digest=sha256:a4cc33f02fa4f12af794b63fb2dd5736d3aeedea812a951d80bf51fd83a8b2ce

Observation f496c456-0247-4c12-b42d-338ec07bae57 · outbound

This paper cites Advances in prospect theory: Cumulative representation of uncertainty,.

Value-Free Policy Optimization via Reward Partitioning Advances in prospect theory: Cumulative representation of uncertainty,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.692625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.312840Z digest=sha256:45db5eaa6531abc751e974d8da3b781156e130031afee1eb99be559073a7ea03

Observation efb414d2-4669-4b7d-a5eb-e8b65446ba0b · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Value-Free Policy Optimization via Reward Partitioning UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.315854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.315854Z digest=sha256:b8aac8b04c8643ed1c800ed2f3edfdf0de46d7ff52099f3e70975512617b3965

Observation 7262b461-5e87-474b-85dd-4960bc1d8d80 · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models,.

Value-Free Policy Optimization via Reward Partitioning Alpacaeval: An automatic evaluator of instruction-following models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.683149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.319398Z digest=sha256:15b56791f06b05f67b7920f1951d7880b8ce9e4d4a8879eaf47b93baaf145c54

Observation 29968375-32d7-49df-bf73-acb155c2e158 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Value-Free Policy Optimization via Reward Partitioning Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.322892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.322892Z digest=sha256:81fc415b003cb7815a09fadf5d8d2a218c8da1915149d6f7b4da399cc9854518

Observation 3b78be03-5ee5-44bc-9e33-ffbaca2aee97 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Value-Free Policy Optimization via Reward Partitioning Instruction-Following Evaluation for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.326677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.326677Z digest=sha256:d2588c96b6666385ab993287687007be1602915164eb8207cec7445f2e6e973b

Observation 2b07c684-5b46-4e78-b412-79e35815b2c1 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Value-Free Policy Optimization via Reward Partitioning Training Verifiers to Solve Math Word Problems

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.330155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.330155Z digest=sha256:3379ac8c81b6dbeb3787a854716556d5f6bac0db71cf4e0804a0f4afc6226712

Observation 17fce266-5c00-4cb7-852f-8059491a6f1c · outbound

This paper cites Scaling up models and data with t5x and seqio,.

Value-Free Policy Optimization via Reward Partitioning Scaling up models and data with t5x and seqio,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.667368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.333178Z digest=sha256:341964a0ce9ddb7716f94c35714c7734a9c310c532e789a73d3d7dfcb7ad2420

Observation 7fb1bfa7-74ea-4b4e-93ff-1c7f6ca8f999 · outbound

This paper cites Mistral 7b,.

Value-Free Policy Optimization via Reward Partitioning Mistral 7b,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.336652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.336652Z digest=sha256:9081dbe987501b1232b25cbc02954f2f1728f69e8517916760b809958e963751

Observation ab1ceb4b-8ef4-47cf-9ba4-9e85d1442e9d · outbound

This paper cites Llama: Open and efficient foundation language models,.

Value-Free Policy Optimization via Reward Partitioning Llama: Open and efficient foundation language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.650316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.340728Z digest=sha256:ee4049145c367199a57dbd569887ec8080aa2f53130a9f7205ac4cb172956d19

Observation edba6570-de0e-4a9b-9180-4ac294105a90 · outbound

This paper cites Qwen2 Technical Report.

Value-Free Policy Optimization via Reward Partitioning Qwen2 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.344252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.344252Z digest=sha256:036265455971cb0a41e81e54525f54d891fbc235f4651ce6cb15e057d5fc968c

Observation 526d418b-dd00-4c0e-8f9d-32f81c3ca79a · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Value-Free Policy Optimization via Reward Partitioning BERTScore: Evaluating Text Generation with BERT

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.347457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.347457Z digest=sha256:93754f3f7931b830e4c4d8734bed591558cdf3ddcdbbe0fd9ddb0025064784a8

Observation 762e49c1-11ac-4aac-bf6e-ea267e0d46be · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Value-Free Policy Optimization via Reward Partitioning Rouge: A package for automatic evaluation of summaries,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.350864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.350864Z digest=sha256:ca1bc73a12554051d18225932e59cafbb93e36365d54bc4653af75d9ffdc7bed

Observation 1c2fe268-4e4a-4beb-85cc-8fae2cd7f522 · outbound

This paper cites A diversity-promoting objective function for neural conversation models,.

Value-Free Policy Optimization via Reward Partitioning A diversity-promoting objective function for neural conversation models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.634832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.354796Z digest=sha256:e185f8b1f9da08e4da04ac195f984aad8a77050d5de995fcb9155e34e6681e4e

Observation c84dfb56-bb9d-4b13-8610-5d333745a163 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Value-Free Policy Optimization via Reward Partitioning A Survey on LLM-as-a-Judge

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.358212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.358212Z digest=sha256:b81297f74a8486dbdb84939a9730772f4e4c492602d638a6724541d970a44817

Observation 0c8e74f6-81cf-4fd3-9c7a-df18274cf243 · outbound

This paper cites GPT-4o System Card.

Value-Free Policy Optimization via Reward Partitioning GPT-4o System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.361713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.361713Z digest=sha256:165cba0cd1c7c02633ed8db71cea7453cfde86c0a11ecdecc34004645b69e6fc

Observation d8c59ea9-5c99-45a5-bbd7-2c6f3921eecc · outbound

This paper cites Claude 3.5 sonnet.

Value-Free Policy Optimization via Reward Partitioning Claude 3.5 sonnet

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.623894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:31:49.365419Z digest=sha256:03f8af56fd86345600cb38515f10bc04f529e310d9890a1a181e0aba21763451

Observation 9a806028-570f-490e-bdc4-e44b439ba7ac · outbound

This paper cites Decoupled weight decay regularization,.

Value-Free Policy Optimization via Reward Partitioning Decoupled weight decay regularization,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.368412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.368412Z digest=sha256:3b8e2a3fe1e1c72942d920af40e2ab101e21af565cf6c465b6755deaeeaf21a9

Pith citing papers

Observation 58037a65-8d50-481a-8f54-57238aecf6dd · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Value-Free Policy Optimization via Reward Partitioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.906356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.906356Z digest=sha256:82f47fdf1ab83fb162801122f2140aaeea0c20964037bbc9021f96c63e53afb4