Pith. sign in

Paper Citation Record · LEDGER

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.29246.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29246 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:57:59.957126Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2252a65d-8797-4a2d-88ac-2bce9c40056d · outbound

This paper cites an unresolved cited work.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:55.107801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:55.107801Z digest=sha256:681b277b3f40287361d479f1fea0efebf20c382132855378ad6ab3d92fcf6f02

Observation c56373e8-8d14-4135-aa16-9e83e4560f72 · outbound

This paper cites GPT-4 Technical Report.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:55.270507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:55.270507Z digest=sha256:1b56051c66d9cf2b14bc7ee1d97edcea606ef31d693b7a9e203c07020a430a49

Observation 3c428343-9b71-46ba-97f9-9d108374649c · outbound

This paper cites Constrained policy optimization.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Constrained policy optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:55.474776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:55.474776Z digest=sha256:2765529775708afa8dedf6c8916ffc5c052d3cabd6693021710eafe943226043

Observation 710a68b9-ed44-4579-a283-d54ab2267edc · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL A General Language Assistant as a Laboratory for Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:55.639962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:55.639962Z digest=sha256:a023ee83139880941254dd475534c433777eff9eafada393be8c301d75e1cb5a

Observation c886ddbb-9dfd-40bf-878c-c32ef9e2f118 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:55.776508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:55.776508Z digest=sha256:ef80d4fa762186360beaecb728788e5556366f9d9fbfd54ee9106cd6e1e74fd8

Observation f536a981-51e5-491c-8815-21f3cb3dd3b8 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Safe rlhf: Safe reinforcement learning from human feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:55.906800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:55.906800Z digest=sha256:7da08c56d00d99476fe0a4c1237bbf9b3c3da0b8834f1005407535adcdb2d6fc

Observation 00b266f5-8f52-4085-974c-087954635c6e · outbound

This paper cites Dynamic multi-reward weighting for multi-style controllable generation.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dynamic multi-reward weighting for multi-style controllable generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:56.053973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:56.053973Z digest=sha256:2120cfc345b1765ad65a7e18f026ca3b370f751d52820263f959ad621c393cb7

Observation 36b1c5c6-eaf4-459d-b99d-87d16faeede5 · outbound

This paper cites Controlled Text Generation via Language Model Arithmetic.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Controlled Text Generation via Language Model Arithmetic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:56.255359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:56.255359Z digest=sha256:da58ec547e8038f809b343bb3773cdd3921ebb4124d57ce66b15cc7de9b8b14e

Observation 9b35f7ae-9383-43d2-871d-d0bd26217ed8 · outbound

This paper cites Sciknoweval: Evaluating multi-level scientific knowledge of large language models.arXiv preprint arXiv:2406.09098, 2024.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Sciknoweval: Evaluating multi-level scientific knowledge of large language models.arXiv preprint arXiv:2406.09098, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:56.421922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:56.421922Z digest=sha256:186d9dae1698666665556e5c3f1bc2b142d9689bc551661724d25e6ae47c3d22

Observation 76e35d0d-e00e-4726-a1b9-7acb6c8fad02 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:56.478012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:56.478012Z digest=sha256:8a0790c1d97c941e7a434d46990e537e1cbdbd12570a1e75762a3e8f346cb2a4

Observation 0aeeb0ac-8a3e-44c7-b16d-220ba9594999 · outbound

This paper cites Training products of experts by minimizing contrastive divergence.Neural computation, 14(8): 1771–1800, 2002.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Training products of experts by minimizing contrastive divergence.Neural computation, 14(8): 1771–1800, 2002

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:56.556638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:56.556638Z digest=sha256:de8a31d48a6cc0e740e3cb4107bfe76ffce229d3dc844b12b182ff6c3bc80350

Observation d09ef992-9976-4b7a-ad03-d54637359612 · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:56.614290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:56.614290Z digest=sha256:8f0fba619688d2d2b8c0db97e55c25bcfa4a12eaf2bb0ea5753c9bcca7794d6b

Observation 283c1b82-022e-4dba-9eee-eede598c0b17 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:56.722704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:56.722704Z digest=sha256:d9a33561da54fe8dc02a919e1c236776c06381d9f956f01742460e32209f61e5

Observation 2c0ac92d-9ac7-47ee-aece-ac1a95b36dc2 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL The power of scale for parameter-efficient prompt tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:56.805771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:56.805771Z digest=sha256:8e7f2495dccb8ca6f882d086f378723f7ad295a8761ee9f345c0f0efc6ca56a7

Observation c8f23978-ec55-455c-a541-bddc6d2be2bc · outbound

This paper cites Gradient-adaptive policy optimization: Towards multi-objective alignment of large language models.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Gradient-adaptive policy optimization: Towards multi-objective alignment of large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:56.890897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:56.890897Z digest=sha256:dfa575bd2183547e31d9838de45dcc8071660954a5307188159376ed5e123b58

Observation 2338c79f-231e-4e2f-87db-97c88a8d3dd1 · outbound

This paper cites Prefix-tuning: Optimizing continuous prompts for generation.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Prefix-tuning: Optimizing continuous prompts for generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:56.946072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:56.946072Z digest=sha256:fa87b2ff59645e36497acdd2df835571ee0e192fce4396fa08ca246c7c392a93

Observation c44bc475-1f86-4cf5-814a-33156b0f5d9a · outbound

This paper cites Dichotomous Diffusion Policy Optimization.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dichotomous Diffusion Policy Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.027963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.027963Z digest=sha256:c7329547b690b60e1b42516327f8eb7b56d53efb80053e93dff2edbd18f8e73e

Observation f4338bba-2e77-45b0-bc39-7ddafa50bf5d · outbound

This paper cites Mitigating the alignment tax of rlhf.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Mitigating the alignment tax of rlhf

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.111192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.111192Z digest=sha256:b947108ddc7a9c5fcb034563d348d4e5f1cf2d97bbe36b3dea7ffc5484733b89

Observation 43580a14-00a4-4046-a1fe-0e31a0e10efd · outbound

This paper cites DeepSeek-V3 Technical Report.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL DeepSeek-V3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.170681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.170681Z digest=sha256:01c6b1fdca7951b399156c47501c30d33e3f3bc7202e0cbdd39be92faa852424

Observation f00b3543-9e56-4633-99c3-50c5359b8567 · outbound

This paper cites Dexperts: Decoding-time controlled text generation with experts and anti-experts.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dexperts: Decoding-time controlled text generation with experts and anti-experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.209594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.209594Z digest=sha256:171839aa9d4cb89497f67273e784f11ae305ca89e4ec0b2f9281bf5302a83c14

Observation 91e8983d-011f-4b77-9d09-18d24677f944 · outbound

This paper cites GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.281195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.281195Z digest=sha256:4a0c953d862530bd7fca24886ea46922113fe02fdd9740c25e30c0f4107bfe7b

Observation d09a2950-006e-4929-a5c3-06a546bc80b5 · outbound

This paper cites Contrastive Decoding Improves Reasoning in Large Language Models.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Contrastive Decoding Improves Reasoning in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.414627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.414627Z digest=sha256:72ed4e547e46b0da991bbc94ddb9c38aa4e96f5387ae65ad3c6177ee6958bfa8

Observation 519e1b21-465a-48ff-8af6-285222c09a1f · outbound

This paper cites Training language models to follow instructions with human feedback.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.550904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.550904Z digest=sha256:3177039c695ae916818e314d34edbc0866e0a372953068fb45bd13baa7e3f314

Observation 79d03c10-79d7-4257-9e56-85f8bade94ba · outbound

This paper cites Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.658858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.658858Z digest=sha256:4802e4ff4bf071d33b5001d2f3512549312abe5ad19da46b15c18325d0c04f61

Observation 60fae710-b15f-47ad-aaab-3b0a667bc766 · outbound

This paper cites Efficiently Scaling Transformer Inference.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Efficiently Scaling Transformer Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.766979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.766979Z digest=sha256:4b67e729a60399550451049b87561143eb5a3813d178d3c9d1bab49a79f314f3

Observation f4dbf456-c303-4424-8266-63279c71fa1d · outbound

This paper cites Toolrl: Reward is all tool learning needs.Advances in Neural Information Processing Systems, 38:105523–105553, 2026.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Toolrl: Reward is all tool learning needs.Advances in Neural Information Processing Systems, 38:105523–105553, 2026

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.931652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.931652Z digest=sha256:322eea107e2c1a93f56a387766f9ca4c7eb7226aff22dff4ad1277fa0d893217

Observation 18a6c740-2c46-41a6-bc98-84658d90fbc1 · outbound

This paper cites Dmoerm: Recipes of mixture-of-experts for effective reward modeling.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dmoerm: Recipes of mixture-of-experts for effective reward modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.038209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.038209Z digest=sha256:6983a8e1d630ad0d11c844ea56024a2f082506693b8113e748443be066fec89a

Observation 933e819f-9aaa-4f02-9e7d-0e57b9f5a0f9 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.146712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.146712Z digest=sha256:b09a5ece720e859f3d9f2e897728f82739bdf31d3e871fd80bd9f08691a65be9

Observation d164ed71-f4cc-46b1-90dc-f45e60b3ff22 · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL WARM: On the Benefits of Weight Averaged Reward Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.293232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.293232Z digest=sha256:a890bcd5c4ccc7f01bbe4ba69866cd287b6a1115d1fc2e6ea546f62ae5550bf2

Observation ddb18dcd-5389-493e-9ace-2059d88a841b · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.404799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.404799Z digest=sha256:94a586f8d6ed42e5edd52ea31e45da512ae7fd424057662ade0dba951cd01964

Observation 8628402e-3a1b-475d-9b05-e8d74a491354 · outbound

This paper cites A survey of multi-objective sequential decision-making.Journal of Artificial Intelligence Research, 48:67–113, 2013.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL A survey of multi-objective sequential decision-making.Journal of Artificial Intelligence Research, 48:67–113, 2013

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.515088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.515088Z digest=sha256:d1249ada1d0350505779ee6829da1fd7a5f2e2dffb3a4802b5070de370524af5

Observation 357c0212-c5f2-4562-9d4a-08842e2ba2c8 · outbound

This paper cites Scienceqa: A novel resource for question answering on scholarly articles.International Journal on Digital Libraries, 23(3):289–301, 2022.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Scienceqa: A novel resource for question answering on scholarly articles.International Journal on Digital Libraries, 23(3):289–301, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.543147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.543147Z digest=sha256:49fe016ef747e9e11babe62820b297f9096d769372b3dce5c36dc0b9dd98c15f

Observation cf1a97c0-b3a4-4c86-a02c-b6ddc315f087 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Proximal Policy Optimization Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.640700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.640700Z digest=sha256:69a75aaed2b940f4d3d83fe9d2a536f39d74651202213f6772a176cfecfdb7fa

Observation 8885a5af-f4aa-4145-b41c-ea459b4ef695 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.769275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.769275Z digest=sha256:e8009927a24a3d65b8994b85fdec75db3f47ca7812f674d327a7302ac7b4307b

Observation 59c2228c-b2e3-4d21-969e-27d1be6177fe · outbound

This paper cites MIT press Cambridge, 1998.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL MIT press Cambridge, 1998

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.923934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.923934Z digest=sha256:8272723edbecf0353eed0cdfba495c38546e065c73f4f67bd53f33387c184de0

Observation d9549157-ac81-43e7-9ea4-cae6e2590bed · outbound

This paper cites Hashimoto.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Hashimoto

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.072272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.072272Z digest=sha256:719a4c332242de47eed62cbaf72d875bdaf0eecf2051bdd9e342566083e03677

Observation 869403b3-bd32-4e28-859a-bedd41bbcac0 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Qwen2.5: A party of foundation models, September 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.184106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.184106Z digest=sha256:a9514220a17126411bdf7941107804d4eef78fd8daacf8a70a65b390eadb87b0

Observation 15785910-7208-433a-9312-1489c8754843 · outbound

This paper cites Interpretable preferences via multi- objective reward modeling and mixture-of-experts.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Interpretable preferences via multi- objective reward modeling and mixture-of-experts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.278322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.278322Z digest=sha256:b8e0c9ed928864d44d4dedce9dc7da71e7cb1e86fa4dd399c2bab911c6f384c7

Observation 2059eb48-5b8f-409d-95e3-8eda809e0613 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Aligning Large Language Models with Human: A Survey

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.358866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.358866Z digest=sha256:7e0464f26d41ce777ce12585f857fe8b9f30b7f9f8cd148334de2db9edca002b

Observation 53e6144e-33e8-454b-b059-9793794acb39 · outbound

This paper cites Multi-objective Reinforcement learning from AI Feedback.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Multi-objective Reinforcement learning from AI Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.414952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.414952Z digest=sha256:79d21763ed75f619b9e114f5e084a68320e932ef8c7c8fe41658266c3d07a1b4

Observation b2f1d597-9f39-42f4-9f2b-210396983f6f · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.453124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.453124Z digest=sha256:fd9bf736c88a34a0e73f4661e96c6dddad3ab790ca62f1a626901bc0733c501f

Observation 688eef7b-a850-4b9d-9991-2ddcb722e34f · outbound

This paper cites Rewards-in-context: Multi-objective alignment of foundation models with dynamic preference adjustment.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Rewards-in-context: Multi-objective alignment of foundation models with dynamic preference adjustment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.464943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.464943Z digest=sha256:eb661720a6625fb1ee808a19b5cc32a48de08686eab622170d7fc3203c26a032

Observation c6f28c77-d3a1-45e7-a0ca-68f330b176c2 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale.Advances in Neural Information Processing Systems, 38:113222–113244, 2026.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dapo: An open-source llm reinforcement learning system at scale.Advances in Neural Information Processing Systems, 38:113222–113244, 2026

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.498440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.498440Z digest=sha256:95e42b17db1f14f2a7a39216ca4d3eb4f7081cc8399b835ad925744e8fb4d440

Observation 236c7969-d734-4214-84c0-fce3c1900fdd · outbound

This paper cites Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.532178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.532178Z digest=sha256:00b7bb781a170285700bebb3c02d0e477e3a6e40823ff94506f61af1ee918a8b

Observation c432441f-4f02-4915-8c54-66fd1168a07a · outbound

This paper cites Group Sequence Policy Optimization.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Group Sequence Policy Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.570265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.570265Z digest=sha256:cdcd4bec6fbcb1a73f15b0408075f405964f2c280f086831e24e7e69740131e9

Observation fd437aff-d4f8-4944-b63b-fbeec4b8b18c · outbound

This paper cites Beyond one-preference- fits-all alignment: Multi-objective direct preference optimization.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Beyond one-preference- fits-all alignment: Multi-objective direct preference optimization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.603375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.603375Z digest=sha256:13d3d77482efcb740a94033958a5124691749148cf31b401452c0deda7587b00

Observation 4a908d3b-3316-4cfe-92db-d48ff77ec12e · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Fine-Tuning Language Models from Human Preferences

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.636424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.636424Z digest=sha256:5d2cef4cde488832efa026de539acce0c2b6a986650ad261b5fb7c3a3071b00a

Observation c3f791c3-5068-46c1-b111-0ee663959998 · outbound

This paper cites an unresolved cited work.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.669547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.669547Z digest=sha256:57dd6c647f8306751fba2cacf4f83fed66a7f3a66cdf5b7a89b0cba45cf1da6a

Observation 2e665d08-12e3-4708-b27b-49ad6f654631 · outbound

This paper cites an unresolved cited work.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.703705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.703705Z digest=sha256:01f1d65e1ccf2ab7a2108a402723b74503ec6f16c606efccfc531d3ca37384bb

Observation 6edc1baf-0eab-471f-a361-e479bf5a0652 · outbound

This paper cites an unresolved cited work.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.735618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.735618Z digest=sha256:de936e49ca3a66e64cbc7095957d9b55ac5c4c13e96ab79f40d63d98e92c81df

Observation a0baae10-9d9a-48d0-a9b6-0f1a15126da5 · outbound

This paper cites an unresolved cited work.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.773858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.773858Z digest=sha256:e2172965044489c742b1448434043d70340f9ea46c03dd89a45058ba448a751d

Observation 40321411-77e6-48c6-ae9f-75596065d0c4 · outbound

This paper cites A", "B",.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL A", "B",

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.791574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.791574Z digest=sha256:595c0bfee42c75b16a5b166aab25cbe4c1b4249e1fd1d508684b63ddf87aa3e5

Observation ce02a277-1602-412e-abba-2d5046ed19cf · outbound

This paper cites Provide at least one of<tool_call> or <response>.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Provide at least one of<tool_call> or <response>

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.829818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.829818Z digest=sha256:350e224823fbe1d97cccbccaa8718873b811b58c06d948257485cd088ebaa2ea

Observation a755a497-cc37-4d2a-88fe-99d37282fd9d · outbound

This paper cites name" field and a.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL name" field and a

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.876040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.876040Z digest=sha256:04119fc82fa6e11a889d076ef1d37533593ee6691d120b64b72a9f54027f22e5

Observation fe4842b0-bf39-4b9c-b21d-b92771163819 · outbound

This paper cites an unresolved cited work.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.909007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.909007Z digest=sha256:5a157db185c60940cd782ccdfa38dd02d8d894c2e96bb5f9d3625976c11a4114

Observation 9a64dbc0-e2c5-4ab3-911e-4f6f06e01b83 · outbound

This paper cites an unresolved cited work.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.942109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.942109Z digest=sha256:bfa3578a5d1b6485c372d6816e6d4f0435fc34448748fb3b1a33060bdbe32431

Observation 1c8af99f-cb91-4ad2-8e21-6ff0c6274339 · outbound

This paper cites name ":.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL name ":

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:59.957126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:59.957126Z digest=sha256:d98c4ae0231bd6a959fc33ddd2c6b4552e83394df1e57ac68378327af049d4fe

Pith citing papers

No inbound Pith citation observations are available.